Text-to-Video Generation
Design a Sora/Movie Gen-style system by extending text-to-image diffusion across a time axis — latent diffusion's temporal compression, U-Net/DiT extended with temporal attention and 3D patches, and the Fréchet Video Distance metric that catches what frame-level FID can't
The Sora/Movie Gen interview, built as a delta from `text-to-image-diffusion-models` rather than a re-teach: why naive per-frame diffusion is computationally untenable for a 5-second 720p clip, latent diffusion as the headline fix (a compression network shrinking both the temporal and spatial dimensions, worked through to its ~512x-cheaper number), data prep at 100M-video scale with a concrete ~200TB latent-caching calculation, extending U-Net and DiT across the time axis (temporal attention, temporal/3D convolution, 3D patches, RoPE), why current frontier systems chose DiT over U-Net, the two mitigations for scarce video-text data, the cost levers that make training feasible (including a video-specific spatial-and-temporal super-resolution cascade), Fréchet Video Distance as the genuinely new evaluation metric this subject teaches, and the two video-specific stages — a visual decoder and temporal super-resolution — layered onto subject 2's inference chain.
Practice questions (6)
-
View →
The Compute Argument: Why LDM Isn't Optional for Video
Advanced · Free -
View →
What Temporal Layers Add to an Image-Only U-Net
Advanced · Free -
View →
Handling Scarce Video-Text Data: Two Strategies, Not One
Advanced -
View →
What FVD Catches That Averaged Frame-Level FID Cannot
Advanced -
View →
The Two Stages Subject 2's Inference Chain Doesn't Have
Advanced