Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Handling Scarce Video-Text Data: Two Strategies, Not One

Your team has access to hundreds of millions of image-caption pairs (from the pretrained text-to-image system) but only a much smaller pool of high-quality video-caption pairs. A teammate proposes: "let's just train only on the video data we have — it's the only data that actually teaches the model what we need, and mixing in image data will confuse it about the time dimension."

  1. Explain why paired video-text data is scarcer than paired image-text data in the first place.
  2. Describe both mitigation strategies this subject covers, and evaluate the teammate's "video-only" proposal against each.
  3. Which strategy would you default to, and what's the argument for choosing it over the alternative?

Share this question

← Back to Text-to-Video Generation practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.