Text-to-Image Generation with Diffusion Models
Design a DALL-E/Imagen/Stable Diffusion-style system — noise schedules, U-Net and DiT cross-attention, classifier-free guidance, DDIM, and the CLIP-based evaluation vocabulary the rest of the generative-image track leans on
The deepest subject in this track's generative-image sequence: why diffusion, not autoregressive generation, is the default choice for text-to-image quality and a tunable steps-vs-quality dial; caption engineering at 500M-pair scale (BLIP-style re-captioning, CLIP-score filtering); the U-Net (downsampling/upsampling blocks with cross-attention) and DiT (patchify -> Transformer -> unpatchify) architectures; the forward/backward diffusion process and the predict-the-noise training objective at interview altitude; classifier-free guidance and DDIM step reduction as the two sampling techniques worth having ready; CLIP and CLIPScore introduced properly, once, for this subject and its two upcoming siblings to reuse; and the full production system — data, training, optimization, and inference pipelines, including the prompt-safety-to-super-resolution inference chain.
Practice questions (6)
-
View →
Choosing Diffusion Over Autoregressive for a Consumer T2I Product
Advanced · Free -
View →
Debugging a U-Net That Ignores the Prompt
Advanced -
View →
CFG and DDIM: What Each One Actually Buys You
Advanced -
View →
CLIP and CLIPScore: One Model, Two Jobs in the Same Pipeline
Advanced -
View →
Why the Inference Chain Is Ordered the Way It Is
Advanced