Advanced
Open
Pro
Worked Example: Sampling a 2048x2048 Image
Part of the ML System Design Interview path →
Part of the Generative Vision & Image AI System Design path →
Your image generator was trained with a tokenizer where each visual token corresponds to a 64x64 pixel chunk. The product now needs 2048x2048 output, but you'd rather not retrain the generator at that resolution.
- If you generated natively at 2048x2048 with 64px tokens, how many sequential generation steps would that take, and why is that a latency concern worth flagging?
- Propose an alternative that avoids retraining the generator at the higher resolution, and name the general pattern it's an instance of.
- Walk through the full sampling process (seed to final pixels) for your proposed approach, including where top-p sampling fits in and what it's trading off.
Share this question