Practice — Multimodal LLMs and Vision (4 questions)
Intermediate
Open
Free
How Vision-Language Models Distinguish Between Multiple Objects Permalink →
A multimodal LLM like GPT-4V/GPT-4o is shown a photo containing several different objects (say, a cat, a dog, and a bowl on a table). When you ask "what color is the dog's collar," how does the model know which item in the image is the dog, as opposed to the cat or the bowl? Does it "detect" objects the way a model like YOLO does?
Share this question