Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Text-to-Video Generation (6 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Advanced Open Free

The Compute Argument: Why LDM Isn't Optional for Video Permalink →

A teammate proposes shipping a text-to-video system that runs subject 2's diffusion pipeline directly in pixel space, one frame at a time, with a lightweight mechanism bolted on afterward to smooth transitions between frames: "we already have a working text-to-image diffusion model — let's just run it 120 times per video and blend the results."

  1. Using the numbers from this subject's requirements (5s, 24fps, 720p), work out why this proposal is untenable for a shippable product, with the arithmetic shown.
  2. Explain what latent diffusion (LDM) changes about where the diffusion process actually runs, and why that's the fix.
  3. Even setting the "120 independent generations" framing aside, name the specific quality problem this proposal would still have that LDM alone does not solve.

Share this question

Advanced Open Free

What Temporal Layers Add to an Image-Only U-Net Permalink →

Your team's video generator is a subject-2-style U-Net with cross- attention for text conditioning, applied frame-by-frame to a video latent. Generated clips have good per-frame quality but visibly flickering, inconsistent motion — an object's shape or color shifts noticeably between adjacent frames even when the prompt describes smooth, continuous motion.

  1. Explain, mechanistically, why a subject-2-style U-Net (2D convolutions + cross-attention) produces exactly this failure mode.
  2. Describe temporal attention and temporal (3D) convolution, and explain what each one specifically fixes.
  3. Where do these new layers get added relative to the U-Net's existing structure, and what stays unchanged?

Share this question

Advanced Open Pro

Handling Scarce Video-Text Data: Two Strategies, Not One

Unlock this question →
Advanced Open Pro

What FVD Catches That Averaged Frame-Level FID Cannot

Unlock this question →
Advanced Open Pro

The Two Stages Subject 2's Inference Chain Doesn't Have

Unlock this question →
Advanced Open Pro

Reproducing LDM's ~512x Compression Number From Scratch

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.