Advanced
Open
Pro
Fixing a Personal-Mode Latency SLA Breach
Part of the ML System Design Interview path →
Part of the Generative Vision & Image AI System Design path →
Your personal-mode try-on service currently routes every request through the diffusion cascade to maximize quality. Monitoring shows p50 latency of ~8 seconds, well past the 5-second interactive budget from Step 1, and session-abandonment rate on the try-on screen is rising.
- Diagnose the root design mistake, referencing Step 3's architecture decision.
- Propose a concrete fix using the cascade pattern from the scalability deep-dive, and describe exactly what happens on the request path for a typical user versus one who lingers.
- After the fix ships, what would you expect to happen to average garment fidelity for the typical fast-path user, and is that an acceptable trade? Justify your answer against Step 1's assumptions.
Share this question