Autoregressive Image Generation
Design an unconditional high-resolution image generator by treating pixels as a token sequence — VQ-VAE tokenization and decoder-only Transformer generation
Design a high-resolution image synthesis system the way ByteByteGo-style interviews frame it: why VAEs and GANs stall past a few hundred pixels, why autoregressive beats diffusion on raw sampling speed, how a VQ-VAE/VQGAN tokenizer turns an image into a sequence of discrete tokens, how a decoder-only Transformer generates that sequence, the two-stage training pipeline and its four tokenizer losses, top-p sampling with a worked 1024x1024 example, and the evaluation and service-separation decisions interviewers probe.
Practice questions (6)
-
View →
Why Quantize? Posterior Collapse and the VQ-VAE Codebook
Advanced · Free -
View →
Why a Decoder-Only Transformer for the Image Generator
Advanced -
View →
Diagnosing a Blurry VQGAN Reconstruction
Advanced -
View →
Worked Example: Sampling a 2048x2048 Image
Advanced -
View →
FID vs. Inception Score: What Would Move Each
Advanced