Generative Vision & Image AI System Design
From how images are generated — GANs, autoregressive tokenizers, diffusion, personalization, video — to how image AI actually ships as a business: virtual try-on, generative fill, catalog photography, super-resolution, synthetic training data, factory defect detection and deepfake detection. Each case study ends with how the real product solved scale, what it costs, and what an interviewer will push on.
1 of 14 subjects free
Who it's for
Built for anyone who wants a structured, ordered path through Generative Vision & Image AI System Design — 14 subjects, free to start, at your own pace.
Sign up free to save your progress through this path.
What you'll learn
-
1
ML System Design Interview Framework
Learn how ML system design interviews are scored, a 7-step framework from requirements to monitoring, back-of-envelope estimation for QPS, embeddings and GPU cost, common problem framings, and the mistakes that sink strong candidates.
Free Start → -
2
Multimodal LLMs and Vision
Learn how multimodal LLMs actually process images: patch tokenization, vision-encoder pretraining, and how the model distinguishes between multiple objects in a scene through attention rather than bounding-box regression. Covers the practical gap between vision-language models and classical object detectors, common failure modes (counting, fine-grained discrimination, spatial relations), and prompting techniques (referring expressions, crops, set-of-mark) that make multi-object questions reliable in production.
Pro Start → -
3
Generative Adversarial Networks: From Minimax to StyleGAN
The foundation subject the rest of this track's business cases depend on: why a single forward pass still matters in 2026 (real-time serving, the adversarial loss hiding inside every VAE/VQGAN decoder, GAN-based distillation of diffusion models, still-SOTA perceptual super-resolution and fast inpainting) argued against diffusion and autoregressive generation on latency, quality, diversity, training stability and controllability; the generator-vs-discriminator minimax game and the non-saturating loss at interview altitude; the failure-mode ladder — mode collapse, vanishing gradients, high-resolution instability — and each fix (minibatch discrimination, WGAN/WGAN-GP, progressive growing); StyleGAN's mapping network, AdaIN/modulated convolutions, per-layer style injection, noise inputs, style mixing and the truncation trick, with StyleGAN2's fixes at one-paragraph altitude; controllability as a product feature (InterFaceGAN-style latent directions, GAN inversion, identity preservation); FFHQ-style data and the output-diversity requirement; FID/IS plus precision/recall for generative models and a bias audit; and a synchronous generation-service design with latent storage, moderation and the consent/deepfake question.
Pro Start → -
4
Autoregressive Image Generation
Design a high-resolution image synthesis system the way ByteByteGo-style interviews frame it: why VAEs and GANs stall past a few hundred pixels, why autoregressive beats diffusion on raw sampling speed, how a VQ-VAE/VQGAN tokenizer turns an image into a sequence of discrete tokens, how a decoder-only Transformer generates that sequence, the two-stage training pipeline and its four tokenizer losses, top-p sampling with a worked 1024x1024 example, and the evaluation and service-separation decisions interviewers probe.
Pro Start → -
5
Text-to-Image Generation with Diffusion Models
The deepest subject in this track's generative-image sequence: why diffusion, not autoregressive generation, is the default choice for text-to-image quality and a tunable steps-vs-quality dial; caption engineering at 500M-pair scale (BLIP-style re-captioning, CLIP-score filtering); the U-Net (downsampling/upsampling blocks with cross-attention) and DiT (patchify -> Transformer -> unpatchify) architectures; the forward/backward diffusion process and the predict-the-noise training objective at interview altitude; classifier-free guidance and DDIM step reduction as the two sampling techniques worth having ready; CLIP and CLIPScore introduced properly, once, for this subject and its two upcoming siblings to reuse; and the full production system — data, training, optimization, and inference pipelines, including the prompt-safety-to-super-resolution inference chain.
Pro Start → -
6
Personalizing Image Generation: DreamBooth, LoRA and Textual Inversion
The most product-shaped subject in this track's generative-image sequence: personalizing a pretrained diffusion model (`text-to-image-diffusion-models`) to one specific subject via three tuning methods on a single trade-off ladder — textual inversion (one new token embedding, cheap and weak), DreamBooth (full-model fine-tuning with a rare-token identifier and a class-specific prior preservation loss against overfitting and catastrophic forgetting, best fidelity), and LoRA (low-rank adapters, the same trick `fine-tuning-sft-lora-rlhf-dpo` teaches for LLMs, applied here to a diffusion U-Net) — argued and resolved for an AI-headshots product. Covers quality-gated data prep for a handful of user photos, the combined reconstruction-plus-prior-preservation training objective, hand-engineered sampling prompts, identity-fidelity evaluation (CLIP-I, DINO and dedicated face-recognition similarity, with the CLIP-vs-DINO reasoning spelled out), the asynchronous fine-tune-then-generate-then-verify production pipeline contrasted with subject 2's synchronous inference chain, and a deepfake/consent/PII safety sidebar.
Pro Start → -
7
Text-to-Video Generation
The Sora/Movie Gen interview, built as a delta from `text-to-image-diffusion-models` rather than a re-teach: why naive per-frame diffusion is computationally untenable for a 5-second 720p clip, latent diffusion as the headline fix (a compression network shrinking both the temporal and spatial dimensions, worked through to its ~512x-cheaper number), data prep at 100M-video scale with a concrete ~200TB latent-caching calculation, extending U-Net and DiT across the time axis (temporal attention, temporal/3D convolution, 3D patches, RoPE), why current frontier systems chose DiT over U-Net, the two mitigations for scarce video-text data, the cost levers that make training feasible (including a video-specific spatial-and-temporal super-resolution cascade), Fréchet Video Distance as the genuinely new evaluation metric this subject teaches, and the two video-specific stages — a visual decoder and temporal super-resolution — layered onto subject 2's inference chain.
Pro Start → -
8
Case Study: Virtual Try-On for Fashion E-commerce
Model interview answer for virtual try-on (Zalando/ASOS/Walmart-class fashion retailer): the catalog-mode/personal-mode split as the case's signature design fork; classical warping plus GAN (CP-VTON/VITON-HD lineage) versus Google's TryOnDiffusion cascaded parallel-UNet cross-attention approach versus a DreamBooth-style personalization shortcut that fails at 1M-SKU scale, argued against the GAN-vs-diffusion comparison table from this track's GAN foundation subject; paired on-model/flat-lay data and pseudo-pair construction; garment-agnostic person representation and the losses that protect print fidelity; a garment-fidelity evaluation suite (FID is not enough) including OCR-based print-text consistency and the selection-bias trap in try-on-usage-vs-return-rate analysis; precompute-the-catalog and cascade-preview-then-refine as the scaling patterns; and the cost-per-try-on-vs-cost-of-one-return ROI argument.
Pro Start → -
9
Case Study: Generative Fill — Inpainting, Outpainting and Object Removal
Model interview answer for Generative Fill (remove / extend / generate-with-prompt): why the right design is a cascade — a LaMa-style single-pass GAN for object removal, mask-conditioned latent diffusion for prompt-driven fill, outpainting as inpainting with the mask outside the frame — argued against the GAN-vs-diffusion comparison table from this track's GAN foundation subject; crop-around-the-mask as the scaling trick that makes 8K interactive editing possible; seam-blending and a masked-region-specific evaluation suite (FID is not enough); and a system design with content-credential signing as a first-class pipeline stage.
Pro Start → -
10
Case Study: AI Product Photography for a Million-SKU Catalog
Model interview answer for AI product photography (Amazon ad-image generator / Shopify Magic / Photoroom-class): why the right design is segment-first (a U²-Net/BiRefNet-class matting model, discriminative and cheap) then background generation conditioned on the cutout, never end-to-end regeneration of the product, argued from a fidelity constraint the GAN-vs-diffusion comparison table alone doesn't resolve; product-region invariance as the case's non-negotiable evaluation gate; mask-distribution and template-library vocabulary reused from the generative-fill case; and a genuinely batch-offline scalability story — throughput-optimized GPU batching, template reuse, spot instances and an automatic regenerate loop — that is deliberately unlike the interactive cases (virtual try-on, generative fill) already written in this track.
Pro Start → -
11
Case Study: Super-Resolution and Photo Restoration
Model interview answer for super-resolution and restoration (Pixel Super Res Zoom / Adobe Super Resolution / Real-ESRGAN-class systems): two product points designed side by side — on-device 2-4x gallery upscale under a phone NPU budget and cloud 8x+ face/scratch/colour restoration — with NVIDIA DLSS named and scoped out as a third, differently-constrained regime; the perception-distortion trade-off (PSNR-optimized CNNs vs. GAN-based perceptual SR vs. diffusion SR) argued against the GAN-vs-diffusion comparison table from this track's GAN foundation subject; degradation modelling as the data problem that actually decides whether a model works on real photos; RRDB/ESRGAN-style generators, the relativistic discriminator, and the GAN-loss weight as a tunable hallucination dial; a hallucination audit as this case's non-negotiable evaluation gate alongside PSNR/SSIM-vs-LPIPS/NIQE disagreement; and the edge-vs-cloud scalability story — tiling with overlap, INT8 quantization and distillation, a cheap-CNN-everywhere-plus-GAN/diffusion-on-request cascade, and the DLSS lesson that real-time speed is bought by constraints, not by a smaller version of the same model.
Pro Start → -
12
Case Study: Synthetic Training Data for Computer Vision
Model interview answer for synthetic training-data generation (autonomous driving, retail shelf detection, medical imaging-class systems): why this case is the track's internal flywheel pattern — generation as a data engine feeding a detector's own retraining loop, not a user-facing product, with recall lift per dollar of compute as the ROI story instead of engagement or conversion. The approach ladder from classical augmentation through simulation/rendering (Omniverse Replicator / Unity Perception-class domain randomization), CycleGAN-style unpaired sim-to-real translation, conditional diffusion augmentation, and fully layout-conditioned generation, argued against the GAN-vs-diffusion comparison table from this track's GAN foundation subject. The failure-slice mining loop, label preservation under editing, dedup/leakage checks against the test set, and the real-data-anchor filter that guards against model collapse from training on your own outputs. Train-on-synthetic-test-on-real as the only metric that matters, with FID demoted to a diagnostic; a telemetry-to-promotion pipeline as the system design; and a back-of-envelope ROI argument comparing recall lift per dollar of compute against the cost of collecting and labelling the same failure slice in the real world.
Pro Start → -
13
Case Study: Visual Defect Detection on a Production Line
Model interview answer for industrial visual inspection (electronics, automotive, pharma manufacturing lines, Landing-AI-class vendors): why this is the block's detection-not-generation and edge case, argued against every other case study in the series. The approach ladder from impossible supervised detection to one-class anomaly detection trained on normal parts only — reconstruction-based methods (AnoGAN, f-AnoGAN, grounded in this track's GAN vocabulary) versus pretrained-feature memory-bank methods (PaDiM, PatchCore) — and why memory-bank methods are the current pragmatic default on MVTec-AD-style benchmarks. Golden-sample capture, alignment and lighting normalization, synthetic defects for validation only (cross-referencing this track's synthetic-data-for-CV case), pixel-level anomaly maps pooled into an image-level score, and image-level AUROC plus pixel-level PRO/AUROC as the evaluation pair FID cannot replace. A decision-policy section mirroring the fraud-detection case's asymmetric-cost framing, but for escape cost versus scrap cost. A system design where inference happens on an edge box synchronised to a PLC reject signal, and a scalability deep-dive built entirely around this case's signature problem: one model (or memory bank) per SKU per line, a model-zoo registry, physical drift fixed by re-capturing golden samples, and onboarding a new SKU in hours with 50 images as the real scaling unit.
Pro Start → -
14
Case Study: Detecting Deepfakes and AI-Generated Media at Platform Scale
Model interview answer for platform-scale synthetic-media detection (the trust-and-safety systems behind Meta, YouTube, TikTok-class platforms, and the KYC-selfie-liveness vendors banks buy from): the block's adversarial case and the deliberate mirror of the GAN foundation subject — the very StyleGAN internals that subject teaches to generate are what a frequency-domain forensic detector here learns to spot. A defence-in-depth cascade argued as the thesis: C2PA content credentials and SynthID-class invisible watermarks prove an asset is known-real or known-generated cheaply and deterministically (provenance), which is a different problem from learned detectors proving an unlabelled asset is fake (detection), and a platform needs both. The DFDC's published generalization trap — leaderboard detectors that collapse on an unseen black-box generator — as the reason an in-house generator zoo and unseen-generator AUROC splits are the only honest way to build and grade a detector. Frequency-aware CNNs, face-crop pipelines, video frame-sampling budgets, multimodal lip-sync detectors, and calibration for a score rather than a verdict. A decision-policy section mirroring the fraud-detection case's review-capacity framing, and a scalability deep-dive built around the cascade-as-architecture pattern and continuous retraining as an arms race against new generator families.
Pro Start →