Match a job Paths Subjects Questions Quizzes Pricing
← All paths

ML System Design Interview

What an ML engineer must know to take a model to production — and to explain it in an interview. Starts with the interview framework and the generic system-design vocabulary, then covers data pipelines and feature stores, training and experimentation, serving and deployment, monitoring and drift, ranking and recommendation architecture, LLM applications, and multimodal vision-language models. Finishes with six end-to-end business cases modelled on real products: video recommendation, ad click prediction, payment fraud detection, search ranking, virtual try-on, and visual anomaly detection in manufacturing.

2 of 20 subjects free

Who it's for

Built for anyone who wants a structured, ordered path through ML System Design Interview — 20 subjects, free to start, at your own pace.

0 of 20 subjects complete 0%
Start: ML System Design Interview Framework →

Sign up free to save your progress through this path.

What you'll learn

  1. 1

    ML System Design Interview Framework

    Learn how ML system design interviews are scored, a 7-step framework from requirements to monitoring, back-of-envelope estimation for QPS, embeddings and GPU cost, common problem framings, and the mistakes that sink strong candidates.

    Free Start →
  2. 2

    System Design Basics

    Learn the fundamental principles behind building large-scale, reliable systems: scalability, availability, latency, and the key trade-offs that drive real-world architecture decisions.

    Free Start →
  3. 3

    ML Data Pipelines & Feature Stores

    Design the data side of an ML system: logging for train/serve parity, labelling strategies, point-in-time joins, negative sampling, batch vs streaming features, feature stores, training–serving skew, data validation, lineage and a worked notification-click pipeline.

    Pro Start →
  4. 4

    Model Training & Experimentation at Scale

    Learn how to answer the training half of an ML system design interview: baselines, model family choice, temporal offline evaluation, tuning budgets, experiment tracking, distributed training, GPU cost estimation, embedding tables, retraining cadence, and off-policy evaluation.

    Pro Start →
  5. 5

    Model Serving & Deployment

    Learn to design the serving side of an ML system: batch vs online vs streaming vs hybrid inference, latency budgets with numbers, model server patterns, compression, canary and shadow rollouts, feature-store consistency, autoscaling, and safe fallbacks.

    Pro Start →
  6. 6

    ML Monitoring, Drift & Retraining

    Learn why deployed models decay, how to monitor them in layers from system health to business KPIs, compute PSI and other drift statistics with worked numbers, handle delayed labels, and design retraining triggers with safe validation gates and rollback.

    Pro Start →
  7. 7

    Ranking & Recommendation System Architecture

    Learn the multi-stage recommendation architecture: candidate generation with two-tower models and ANN search, pointwise/pairwise/listwise rankers, multi-task value formulas, position debiasing, cold start, re-ranking policy, and how offline metrics relate to online A/B results.

    Pro Start →
  8. 8

    LLM Application System Design

    Learn to design LLM-powered systems for interviews and production: prompt vs RAG vs fine-tuning, end-to-end RAG architecture, token and cost budgeting, evaluation, guardrails, agents, caching and routing, with a worked support-assistant design.

    Pro Start →
  9. 9

    Multimodal LLMs and Vision

    Learn how multimodal LLMs actually process images: patch tokenization, vision-encoder pretraining, and how the model distinguishes between multiple objects in a scene through attention rather than bounding-box regression. Covers the practical gap between vision-language models and classical object detectors, common failure modes (counting, fine-grained discrimination, spatial relations), and prompting techniques (referring expressions, crops, set-of-mark) that make multi-object questions reliable in production.

    Pro Start →
  10. 10

    Generative Adversarial Networks: From Minimax to StyleGAN

    The foundation subject the rest of this track's business cases depend on: why a single forward pass still matters in 2026 (real-time serving, the adversarial loss hiding inside every VAE/VQGAN decoder, GAN-based distillation of diffusion models, still-SOTA perceptual super-resolution and fast inpainting) argued against diffusion and autoregressive generation on latency, quality, diversity, training stability and controllability; the generator-vs-discriminator minimax game and the non-saturating loss at interview altitude; the failure-mode ladder — mode collapse, vanishing gradients, high-resolution instability — and each fix (minibatch discrimination, WGAN/WGAN-GP, progressive growing); StyleGAN's mapping network, AdaIN/modulated convolutions, per-layer style injection, noise inputs, style mixing and the truncation trick, with StyleGAN2's fixes at one-paragraph altitude; controllability as a product feature (InterFaceGAN-style latent directions, GAN inversion, identity preservation); FFHQ-style data and the output-diversity requirement; FID/IS plus precision/recall for generative models and a bias audit; and a synchronous generation-service design with latent storage, moderation and the consent/deepfake question.

    Pro Start →
  11. 11

    Autoregressive Image Generation

    Design a high-resolution image synthesis system the way ByteByteGo-style interviews frame it: why VAEs and GANs stall past a few hundred pixels, why autoregressive beats diffusion on raw sampling speed, how a VQ-VAE/VQGAN tokenizer turns an image into a sequence of discrete tokens, how a decoder-only Transformer generates that sequence, the two-stage training pipeline and its four tokenizer losses, top-p sampling with a worked 1024x1024 example, and the evaluation and service-separation decisions interviewers probe.

    Pro Start →
  12. 12

    Text-to-Image Generation with Diffusion Models

    The deepest subject in this track's generative-image sequence: why diffusion, not autoregressive generation, is the default choice for text-to-image quality and a tunable steps-vs-quality dial; caption engineering at 500M-pair scale (BLIP-style re-captioning, CLIP-score filtering); the U-Net (downsampling/upsampling blocks with cross-attention) and DiT (patchify -> Transformer -> unpatchify) architectures; the forward/backward diffusion process and the predict-the-noise training objective at interview altitude; classifier-free guidance and DDIM step reduction as the two sampling techniques worth having ready; CLIP and CLIPScore introduced properly, once, for this subject and its two upcoming siblings to reuse; and the full production system — data, training, optimization, and inference pipelines, including the prompt-safety-to-super-resolution inference chain.

    Pro Start →
  13. 13

    Personalizing Image Generation: DreamBooth, LoRA and Textual Inversion

    The most product-shaped subject in this track's generative-image sequence: personalizing a pretrained diffusion model (`text-to-image-diffusion-models`) to one specific subject via three tuning methods on a single trade-off ladder — textual inversion (one new token embedding, cheap and weak), DreamBooth (full-model fine-tuning with a rare-token identifier and a class-specific prior preservation loss against overfitting and catastrophic forgetting, best fidelity), and LoRA (low-rank adapters, the same trick `fine-tuning-sft-lora-rlhf-dpo` teaches for LLMs, applied here to a diffusion U-Net) — argued and resolved for an AI-headshots product. Covers quality-gated data prep for a handful of user photos, the combined reconstruction-plus-prior-preservation training objective, hand-engineered sampling prompts, identity-fidelity evaluation (CLIP-I, DINO and dedicated face-recognition similarity, with the CLIP-vs-DINO reasoning spelled out), the asynchronous fine-tune-then-generate-then-verify production pipeline contrasted with subject 2's synchronous inference chain, and a deepfake/consent/PII safety sidebar.

    Pro Start →
  14. 14

    Text-to-Video Generation

    The Sora/Movie Gen interview, built as a delta from `text-to-image-diffusion-models` rather than a re-teach: why naive per-frame diffusion is computationally untenable for a 5-second 720p clip, latent diffusion as the headline fix (a compression network shrinking both the temporal and spatial dimensions, worked through to its ~512x-cheaper number), data prep at 100M-video scale with a concrete ~200TB latent-caching calculation, extending U-Net and DiT across the time axis (temporal attention, temporal/3D convolution, 3D patches, RoPE), why current frontier systems chose DiT over U-Net, the two mitigations for scarce video-text data, the cost levers that make training feasible (including a video-specific spatial-and-temporal super-resolution cascade), Fréchet Video Distance as the genuinely new evaluation metric this subject teaches, and the two video-specific stages — a visual decoder and temporal super-resolution — layered onto subject 2's inference chain.

    Pro Start →
  15. 15

    Case Study: Video Recommendation (YouTube / Netflix-style Homepage)

    Walk through a model ML system design interview answer for a video recommendation homepage: objectives, implicit-feedback labels, two-tower retrieval, multi-task ranking, position bias, cold start, evaluation, A/B testing, serving at scale, and monitoring.

    Pro Start →
  16. 16

    Case Study: Ad Click-Through Rate Prediction (Meta / Google Ads-style)

    Model interview answer for ad click-through rate prediction: why calibrated pCTR drives the auction, delayed labels and negative downsampling, sparse ID features, LR-to-DLRM model evolution, calibration monitoring, online learning and a low-latency serving architecture.

    Pro Start →
  17. 17

    Case Study: Real-Time Payment Fraud Detection (Stripe / PayPal-style)

    Model interview answer for real-time payment fraud detection: sub-100 ms decisions, 0.1 % positives with asymmetric costs, delayed chargeback labels, point-in-time features and velocity counters, GBDT plus graph signals, PR-AUC and dollar-weighted evaluation, cost-based decision policy and a streaming serving architecture.

    Pro Start →
  18. 18

    Case Study: Search Ranking (Airbnb / E-commerce Marketplace Search)

    Model interview answer for marketplace search ranking: two-sided objectives, retrieval vs learned ranking, position-bias-corrected labels, LambdaMART vs neural rankers, NDCG worked example, cold start, diversity re-ranking, serving architecture and monitoring.

    Pro Start →
  19. 19

    Case Study: Virtual Try-On for Fashion E-commerce

    Model interview answer for virtual try-on (Zalando/ASOS/Walmart-class fashion retailer): the catalog-mode/personal-mode split as the case's signature design fork; classical warping plus GAN (CP-VTON/VITON-HD lineage) versus Google's TryOnDiffusion cascaded parallel-UNet cross-attention approach versus a DreamBooth-style personalization shortcut that fails at 1M-SKU scale, argued against the GAN-vs-diffusion comparison table from this track's GAN foundation subject; paired on-model/flat-lay data and pseudo-pair construction; garment-agnostic person representation and the losses that protect print fidelity; a garment-fidelity evaluation suite (FID is not enough) including OCR-based print-text consistency and the selection-bias trap in try-on-usage-vs-return-rate analysis; precompute-the-catalog and cascade-preview-then-refine as the scaling patterns; and the cost-per-try-on-vs-cost-of-one-return ROI argument.

    Pro Start →
  20. 20

    Case Study: Visual Defect Detection on a Production Line

    Model interview answer for industrial visual inspection (electronics, automotive, pharma manufacturing lines, Landing-AI-class vendors): why this is the block's detection-not-generation and edge case, argued against every other case study in the series. The approach ladder from impossible supervised detection to one-class anomaly detection trained on normal parts only — reconstruction-based methods (AnoGAN, f-AnoGAN, grounded in this track's GAN vocabulary) versus pretrained-feature memory-bank methods (PaDiM, PatchCore) — and why memory-bank methods are the current pragmatic default on MVTec-AD-style benchmarks. Golden-sample capture, alignment and lighting normalization, synthetic defects for validation only (cross-referencing this track's synthetic-data-for-CV case), pixel-level anomaly maps pooled into an image-level score, and image-level AUROC plus pixel-level PRO/AUROC as the evaluation pair FID cannot replace. A decision-policy section mirroring the fraud-detection case's asymmetric-cost framing, but for escape cost versus scrap cost. A system design where inference happens on an edge box synchronised to a PLC reject signal, and a scalability deep-dive built entirely around this case's signature problem: one model (or memory bank) per SKU per line, a model-zoo registry, physical drift fixed by re-capturing golden samples, and onboarding a new SKU in hours with 50 images as the real scaling unit.

    Pro Start →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.