Advanced
Open
Pro
The Two Stages Subject 2's Inference Chain Doesn't Have
Part of the ML System Design Interview path →
Part of the Generative Vision & Image AI System Design path →
An engineer familiar with text-to-image-diffusion-models's inference
chain (prompt safety -> prompt enhancement -> generation -> harm
detection -> super-resolution cascade) proposes reusing that exact
five-stage chain unmodified for the text-to-video product, arguing "the
video model is just a diffusion model like the image one, so the same
pipeline should work."
- Identify the two stages this subject adds to that chain, and explain why each one is required rather than optional for a text-to-video system built on LDM.
- Where does each new stage sit in the ordered chain, and why does that placement matter?
- Is the engineer's "it's just a diffusion model" framing wrong, right, or partially right? Justify your answer precisely.
Share this question