Advanced
Open
Pro
Why a Decoder-Only Transformer for the Image Generator
Part of the ML System Design Interview path →
Part of the Generative Vision & Image AI System Design path →
A candidate proposes using a bidirectional (encoder-style) Transformer or a convolutional recurrent network for the image generator instead of a decoder-only Transformer, arguing "images aren't inherently left-to-right like text, so why force a sequential architecture on them?"
- Give the two reasons a decoder-only Transformer is the standard choice here, and address the candidate's objection directly.
- Describe, at the component level, what happens to a visual token as it moves from "sampled token ID" to "input the Transformer can attend over."
- This subject explicitly bridges to
transformers-for-ai-engineers. In your own words, what is actually shared between a text-generating and an image-generating decoder-only Transformer, and what changes?
Share this question