Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Why a Decoder-Only Transformer for the Image Generator

A candidate proposes using a bidirectional (encoder-style) Transformer or a convolutional recurrent network for the image generator instead of a decoder-only Transformer, arguing "images aren't inherently left-to-right like text, so why force a sequential architecture on them?"

  1. Give the two reasons a decoder-only Transformer is the standard choice here, and address the candidate's objection directly.
  2. Describe, at the component level, what happens to a visual token as it moves from "sampled token ID" to "input the Transformer can attend over."
  3. This subject explicitly bridges to transformers-for-ai-engineers. In your own words, what is actually shared between a text-generating and an image-generating decoder-only Transformer, and what changes?

Share this question

← Back to Autoregressive Image Generation practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.