Text-to-Video Generation
text-to-image-diffusion-models built a system that can turn a prompt into a single, high-quality, prompt-aligned image. This subject is the interview question people mean when they say "design Sora" or "design Movie Gen": do the same thing, except the output is a short video — a sequence of frames that has to look right individually and move right together. Every piece of vocabulary from subject 2 is still load-bearing here: diffusion training (predict the noise, MSE loss, conditioning on a caption embedding and timestep), classifier-free guidance and DDIM-style sampling, CLIP and CLIPScore, U-Net and DiT, FID and Inception Score. None of that gets re-taught. What this subject adds is a single new axis — time — and the consequence of that axis is that the compute budget subject 2 treated as "expensive but manageable" becomes the central design constraint of everything that follows.
The right posture for this material, in an interview and in this subject, is "same as text-to-image, except…" — reach for that phrase constantly. A candidate who re-derives diffusion training or re-explains CFG from scratch here is wasting the clock on material the interviewer already knows they know (or should have learned in the text-to-image round); a candidate who says "same noise-prediction objective as image diffusion, except the model also has to reason across frames, and here's specifically how" is demonstrating the thing this whole track has been building toward: composable understanding, not memorized subject boundaries.
Clarifying Requirements
The same clarify-first discipline from every subject in this track, but the numbers this time are the whole ballgame:
- Duration, resolution, frame rate. A concrete anchor: 5-second clips, 720p (1280x720), 24 frames per second. Multiply those out and you get the number every downstream decision has to answer to:
5 seconds x 24 fps = 120 frames. The system isn't generating an image — it's generating 120 of them, coherently. - Latency. "A few minutes of processing time" is a reasonable starting budget for video, in sharp contrast to subject 2's "near real-time, single-digit-to-low-double-digit seconds." State this explicitly, because it's the number that will later justify why video generation can tolerate a heavier pipeline (visual decoding, spatial and temporal super-resolution) than subject 2's image pipeline needed.
- Domain and language. Broad domain (no narrow category restriction) and English-only to start — the same scoping moves subject 2 made, for the same reasons.
- Audio. Out of scope for a first version — silent video generation, with audio flagged as a plausible future enhancement rather than a day-one requirement. Naming this explicitly avoids a scope creep the interviewer didn't ask for.
- Data scale and compute budget. On the order of 100 million video-caption pairs, with the caveat that some captions are noisy or non-English — the data-prep section below exists because of that caveat. And a compute budget large enough to be worth stating out loud: thousands of high-end GPUs (H100-class) dedicated to training, because — as the next section makes concrete — video training is not a "slightly more expensive" version of image training.
- A pretrained text-to-image model as a starting point. A completely reasonable assumption to negotiate for in an interview, and the one this subject leans on throughout: rather than solving text-to-video from a blank slate, extend subject 2's system.
- Safety. Same first-class requirement as subject 2 — the system needs safeguards against generating offensive or harmful content, not a follow-up feature.