Match a job Paths Subjects Questions Quizzes Pricing

Autoregressive Image Generation

Design an unconditional high-resolution image generator by treating pixels as a token sequence — VQ-VAE tokenization and decoder-only Transformer generation

Overview Read

Autoregressive Image Generation

"Design a system that generates high-resolution images" sounds like it should reuse everything you know about GANs and VAEs, and the first thing a strong candidate does is explain why it can't. multimodal-llms-and-vision in this track covers image understanding — how a vision-language model reads patches of an existing image and reasons about them in text. This subject is the mirror image: image synthesis — how a model produces pixels that didn't exist before, one design decision at a time, in the same clarify → frame → prepare data → build → train → evaluate → design loop the rest of this track uses for any ML system.

The specific angle here is autoregressive image generation: treat an image as a sequence, the same way a language model treats a sentence, and generate it token by token. It's the same core idea GPT-style text models use, adapted to pixels — and it's the natural on-ramp to the harder, more common interview topic, diffusion, which gets its own subject next. Understanding autoregressive generation first buys you two things this track leans on repeatedly: the evaluation vocabulary (FID, Inception Score) and the super-resolution cascade pattern, both of which resurface in every later generative-image and generative-video subject.


Clarifying Requirements

As with any ML system design interview, resist the urge to jump to architecture. A candidate-style exchange for this prompt typically nails down:

  • Scope: unconditional generation first (no text prompt) — "generate a plausible image," not "generate this specific image." Conditioning on text is a later extension, not the starting design.
  • Domain: a bounded category (natural scenes, product photos) rather than "anything," at least to start — training data and evaluation both get harder as the domain widens.
  • Resolution: a concrete target, e.g. 1024×1024 or 2048×2048, stated explicitly because it drives the tokenization design below.
  • Latency: "a few seconds per image" is a very different budget from "real time," and it's the number that ultimately justifies picking autoregressive generation over vanilla diffusion in this design.
  • Data: a large corpus of unlabeled images — millions, not thousands — since nothing here needs human labels; the images are the supervision.

State these out loud before drawing a single box. Everything that follows is a consequence of the numbers you pin down here.


Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.