Match a job Paths Subjects Questions Quizzes Pricing
Machine Learning Advanced Pro

Text-to-Video Generation

Design a Sora/Movie Gen-style system by extending text-to-image diffusion across a time axis — latent diffusion's temporal compression, U-Net/DiT extended with temporal attention and 3D patches, and the Fréchet Video Distance metric that catches what frame-level FID can't

30 min read 17 views

The Sora/Movie Gen interview, built as a delta from `text-to-image-diffusion-models` rather than a re-teach: why naive per-frame diffusion is computationally untenable for a 5-second 720p clip, latent diffusion as the headline fix (a compression network shrinking both the temporal and spatial dimensions, worked through to its ~512x-cheaper number), data prep at 100M-video scale with a concrete ~200TB latent-caching calculation, extending U-Net and DiT across the time axis (temporal attention, temporal/3D convolution, 3D patches, RoPE), why current frontier systems chose DiT over U-Net, the two mitigations for scarce video-text data, the cost levers that make training feasible (including a video-specific spatial-and-temporal super-resolution cascade), Fréchet Video Distance as the genuinely new evaluation metric this subject teaches, and the two video-specific stages — a visual decoder and temporal super-resolution — layered onto subject 2's inference chain.

Practice questions (6)

  • The Compute Argument: Why LDM Isn't Optional for Video

    Advanced · Free
    View →
  • What Temporal Layers Add to an Image-Only U-Net

    Advanced · Free
    View →
  • Handling Scarce Video-Text Data: Two Strategies, Not One

    Advanced
    View →
  • What FVD Catches That Averaged Frame-Level FID Cannot

    Advanced
    View →
  • The Two Stages Subject 2's Inference Chain Doesn't Have

    Advanced
    View →
See all 6 questions →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.