Advanced
Open
Pro
Diagnosing a Bottleneck Across Generation, Decoding, and Super-Resolution
Part of the ML System Design Interview path →
Part of the Generative Vision & Image AI System Design path →
Your image generation product deploys the generator, the tokenizer's decoder, and a super-resolution model as one combined service that scales as a single unit. Under load, p95 latency spikes badly, and an engineer notices the decoding step (a single Transformer-free forward pass) is waiting behind a backlog of generation requests (256 sequential Transformer steps each) even though decoding itself is fast. Separately, someone on the team asks whether the tokenizer's encoder should be added to this service "in case we need it."
- Diagnose why bundling these into one scaled unit causes the observed latency spike, and propose the fix.
- Answer the encoder question directly, with justification.
- What would change about how you'd scale each of the three components if you split them?
Share this question