Advanced
Open
Pro
Reproducing LDM's ~512x Compression Number From Scratch
Part of the ML System Design Interview path →
Part of the Generative Vision & Image AI System Design path →
An interviewer asks you to justify, with arithmetic rather than a cited fact, why latent diffusion is described as making video generation roughly 512x cheaper for a 5-second, 24fps, 720p clip compressed by a factor of 8 in each of the temporal and spatial dimensions.
- Compute the total pixel-value count for the uncompressed clip and for its latent representation, showing your work.
- Derive the compression ratio and explain, in one sentence, why this translates into roughly proportional training/generation cost savings.
- A colleague asks: "if 8x-8x-8x gives ~512x, wouldn't a 16x compression factor per dimension give an even better ~4096x, so why not just compress more aggressively?" What's the tradeoff being ignored in that question?
Share this question