Advanced
Open
Pro
Handling Scarce Video-Text Data: Two Strategies, Not One
Part of the ML System Design Interview path →
Part of the Generative Vision & Image AI System Design path →
Your team has access to hundreds of millions of image-caption pairs (from the pretrained text-to-image system) but only a much smaller pool of high-quality video-caption pairs. A teammate proposes: "let's just train only on the video data we have — it's the only data that actually teaches the model what we need, and mixing in image data will confuse it about the time dimension."
- Explain why paired video-text data is scarcer than paired image-text data in the first place.
- Describe both mitigation strategies this subject covers, and evaluate the teammate's "video-only" proposal against each.
- Which strategy would you default to, and what's the argument for choosing it over the alternative?
Share this question