Advanced
Open
Pro
Why CLIP-I Alone Isn't Enough for Identity Evaluation
Part of the AI Engineer Interview path →
Part of the ML System Design Interview path →
Part of the Generative Vision & Image AI System Design path →
Your headshot product's evaluation dashboard reports CLIP-I scores (cosine similarity between the CLIP image embedding of a generated headshot and the user's reference photos) as its identity-fidelity metric. A product manager notices that several generated headshots score high CLIP-I but, when a human looks at them, clearly do not resemble the user closely enough to ship.
- Explain, mechanistically, why CLIP-I can score a wrong-looking result highly.
- Explain what DINO is trained to do differently, and why that makes it a stronger identity-fidelity proxy.
- Propose a revised evaluation approach using more than one metric, and explain what each one is actually catching.
Share this question