Generative Adversarial Networks: From Minimax to StyleGAN
autoregressive-image-generation and text-to-image-diffusion-models both answer the same interview shape — "generate an image from noise or a prompt" — with two architectures that treat generation as, respectively, next-token prediction and iterative denoising. Neither is a GAN, and if that's the only image-generation vocabulary you bring to an interview, you'll lose the follow-up almost every senior interviewer eventually asks: "why not a GAN here?" This subject exists to close that gap, and it is deliberately positioned first in this track's business-case sequence, not last — every case study that follows it (virtual try-on, generative fill, product photography, super-resolution) makes an explicit GAN-vs-diffusion-vs-discriminative call in its own framing step, and all of them lean on the comparison table built here rather than re-deriving it.
The angle is ByteByteGo's "design a realistic face generator" prompt, but the goal isn't just to reproduce StyleGAN's architecture diagram. It's to leave you able to argue, with specifics, why a technology whose "GANs are dead, diffusion won" narrative is common in casual conversation is still what ships inside some of the most latency-sensitive and highest-volume image products running today. This is a concept subject, not a case study — it uses text-to-image-diffusion-models's structure (state a loss and what each term buys, no derivations) rather than the seven-step case-study arc the business-case subjects that follow it use.