Paths Subjects Questions Quizzes Pricing Search
Advanced Open Pro

Set-of-Mark Prompting for Reliable Object Reference

You're building a product where a user photographs their desk and asks "what's the model number on the third device from the left?" Plain "ask the VLM the question" gives inconsistent answers because the model sometimes picks the wrong device. Describe a technique to make this reliable, and explain why it works better than just improving the prompt wording.

Share this question

← Back to Multimodal LLMs and Vision practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.