Practice — Autoregressive Image Generation (6 questions)
Advanced
Open
Free
Why Quantize? Posterior Collapse and the VQ-VAE Codebook Permalink →
A teammate proposes skipping the quantizer in a VQ-VAE-style image tokenizer: "just use the encoder's continuous output directly as the latent representation, like a standard VAE — it's simpler and avoids the awkward non-differentiable lookup."
- Explain, mechanistically, why this is likely to hurt image quality at high resolution specifically, not just "make training harder."
- Name the two distinct reasons VQ-VAE introduces a codebook, and explain how each addresses a different problem.
- The quantizer's nearest-codebook lookup has no defined gradient. Describe, at a concept level, how the model is still trained end-to-end despite this.
Share this question
Advanced
Open
Pro