Practice — Text-to-Video Generation (6 questions)
The Compute Argument: Why LDM Isn't Optional for Video Permalink →
A teammate proposes shipping a text-to-video system that runs subject 2's diffusion pipeline directly in pixel space, one frame at a time, with a lightweight mechanism bolted on afterward to smooth transitions between frames: "we already have a working text-to-image diffusion model — let's just run it 120 times per video and blend the results."
- Using the numbers from this subject's requirements (5s, 24fps, 720p), work out why this proposal is untenable for a shippable product, with the arithmetic shown.
- Explain what latent diffusion (LDM) changes about where the diffusion process actually runs, and why that's the fix.
- Even setting the "120 independent generations" framing aside, name the specific quality problem this proposal would still have that LDM alone does not solve.
Share this question
What Temporal Layers Add to an Image-Only U-Net Permalink →
Your team's video generator is a subject-2-style U-Net with cross- attention for text conditioning, applied frame-by-frame to a video latent. Generated clips have good per-frame quality but visibly flickering, inconsistent motion — an object's shape or color shifts noticeably between adjacent frames even when the prompt describes smooth, continuous motion.
- Explain, mechanistically, why a subject-2-style U-Net (2D convolutions + cross-attention) produces exactly this failure mode.
- Describe temporal attention and temporal (3D) convolution, and explain what each one specifically fixes.
- Where do these new layers get added relative to the U-Net's existing structure, and what stays unchanged?
Share this question