Autoregressive Image Generation
"Design a system that generates high-resolution images" sounds like it should reuse everything you know about GANs and VAEs, and the first thing a strong candidate does is explain why it can't. multimodal-llms-and-vision in this track covers image understanding — how a vision-language model reads patches of an existing image and reasons about them in text. This subject is the mirror image: image synthesis — how a model produces pixels that didn't exist before, one design decision at a time, in the same clarify → frame → prepare data → build → train → evaluate → design loop the rest of this track uses for any ML system.
The specific angle here is autoregressive image generation: treat an image as a sequence, the same way a language model treats a sentence, and generate it token by token. It's the same core idea GPT-style text models use, adapted to pixels — and it's the natural on-ramp to the harder, more common interview topic, diffusion, which gets its own subject next. Understanding autoregressive generation first buys you two things this track leans on repeatedly: the evaluation vocabulary (FID, Inception Score) and the super-resolution cascade pattern, both of which resurface in every later generative-image and generative-video subject.
Clarifying Requirements
As with any ML system design interview, resist the urge to jump to architecture. A candidate-style exchange for this prompt typically nails down:
- Scope: unconditional generation first (no text prompt) — "generate a plausible image," not "generate this specific image." Conditioning on text is a later extension, not the starting design.
- Domain: a bounded category (natural scenes, product photos) rather than "anything," at least to start — training data and evaluation both get harder as the domain widens.
- Resolution: a concrete target, e.g. 1024×1024 or 2048×2048, stated explicitly because it drives the tokenization design below.
- Latency: "a few seconds per image" is a very different budget from "real time," and it's the number that ultimately justifies picking autoregressive generation over vanilla diffusion in this design.
- Data: a large corpus of unlabeled images — millions, not thousands — since nothing here needs human labels; the images are the supervision.
State these out loud before drawing a single box. Everything that follows is a consequence of the numbers you pin down here.