Paths Subjects Questions Quizzes Pricing Search
Intermediate Open Pro

Why VLMs Struggle to Count Objects Accurately

You ask a vision-language model "how many people are in this photo" for an image with 12 people, several partially overlapping. The model confidently answers "8." A classical object detector run on the same image correctly outputs 12 bounding boxes. Explain why counting is a characteristic weak point for VLMs even though object-presence questions ("is there a person in this photo?") are usually reliable.

Share this question

← Back to Multimodal LLMs and Vision practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.