Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

The Compute Argument: Why LDM Isn't Optional for Video

A teammate proposes shipping a text-to-video system that runs subject 2's diffusion pipeline directly in pixel space, one frame at a time, with a lightweight mechanism bolted on afterward to smooth transitions between frames: "we already have a working text-to-image diffusion model — let's just run it 120 times per video and blend the results."

  1. Using the numbers from this subject's requirements (5s, 24fps, 720p), work out why this proposal is untenable for a shippable product, with the arithmetic shown.
  2. Explain what latent diffusion (LDM) changes about where the diffusion process actually runs, and why that's the fix.
  3. Even setting the "120 independent generations" framing aside, name the specific quality problem this proposal would still have that LDM alone does not solve.
Solution

1. Why the naive proposal is untenable

5 seconds at 24fps is 120 frames. Each frame at 720p (1280x720) has about 3.6x as many pixels as the 512x512 image subject 2's diffusion model was benchmarked against, and a single such image takes roughly a second to generate on a high-end GPU. Scaling per-frame cost up for resolution and then multiplying by 120 frames pushes total generation time for one clip into the range of several minutes — two to three orders of magnitude past subject 2's "near real-time" single-digit- seconds budget, and this is before accounting for the fact that running 120 independent image generations and blending afterward doesn't even attempt temporal coherence; it would need to run through the full diffusion process once per frame, in full pixel-space resolution, with no shared computation across frames at all.

2. What LDM changes

LDM adds a separately trained compression network (a VAE-style visual encoder/decoder) ahead of the diffusion process. The diffusion model never touches raw pixels — it adds noise to, and denoises, a much smaller latent representation that compresses both the temporal dimension (fewer frames) and the spatial dimensions (lower resolution) of the video. Only at the very end does a visual decoder map the denoised latent back to pixel space. Because diffusion's cost scales with how much data it processes at every one of its many denoising steps, working in a representation that's roughly 512x smaller (an 8x temporal reduction times an 8x-by-8x spatial reduction) makes the entire process roughly 512x cheaper — turning "several minutes, naively" back into something shippable.

3. The quality problem LDM alone doesn't solve

Even with LDM's compute savings, running 120 independent per-frame generations and blending them afterward does nothing to address temporal consistency — nothing in that pipeline lets information flow between frames during generation, so objects can flicker, change identity, or move incoherently between frames even if each individual frame is high quality and cheap to produce. That's precisely why this subject's architecture section adds temporal attention and temporal convolution (U-Net) or 3D patches (DiT) on top of LDM — LDM fixes the compute problem, but reasoning across frames is a separate, architectural fix that a blend-after-the-fact approach doesn't provide.

Share this question

← Back to Text-to-Video Generation practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.