Match a job Paths Subjects Questions Quizzes Pricing
🤖

Machine Learning

ML algorithms, model training, evaluation, and production ML systems

52 subjects · browse with filters

Pro Machine Learning
Intermediate

Decision Trees

Learn how decision trees partition feature space using impurity measures, how recursive binary splitting works, which hyperparameters control overfitting, and how feature importance is calculated — the foundation for understanding gradient boosting models like LightGBM and XGBoost.

20 min 1 enrolled
Pro Machine Learning
Advanced

LightGBM

Understand how LightGBM builds sequential ensembles of trees, why leaf-wise growth outperforms level-wise, what every key hyperparameter controls, why a slow learning rate with more trees generalizes better, and how to diagnose and fix overfitting through regularization.

25 min 1 enrolled
Pro Machine Learning
Advanced

LLM Application System Design

Learn to design LLM-powered systems for interviews and production: prompt vs RAG vs fine-tuning, end-to-end RAG architecture, token and cost budgeting, evaluation, guardrails, agents, caching and routing, with a worked support-assistant design.

30 min 1 enrolled
Pro Machine Learning
Intermediate

Bias–Variance Trade-off & Cross-Validation

Master the bias-variance decomposition, learning and validation curves, k-fold, stratified, group and time-series cross-validation, nested CV for honest tuning, and the data-leakage traps that make validation scores lie.

28 min 1 enrolled
Pro Machine Learning
Intermediate

Multimodal LLMs and Vision

Learn how multimodal LLMs actually process images: patch tokenization, vision-encoder pretraining, and how the model distinguishes between multiple objects in a scene through attention rather than bounding-box regression. Covers the practical gap between vision-language models and classical object detectors, common failure modes (counting, fine-grained discrimination, spatial relations), and prompting techniques (referring expressions, crops, set-of-mark) that make multi-object questions reliable in production.

25 min 1 enrolled
Pro Machine Learning
Advanced

Case Study: Generative Fill — Inpainting, Outpainting and Object Removal

Model interview answer for Generative Fill (remove / extend / generate-with-prompt): why the right design is a cascade — a LaMa-style single-pass GAN for object removal, mask-conditioned latent diffusion for prompt-driven fill, outpainting as inpainting with the mask outside the frame — argued against the GAN-vs-diffusion comparison table from this track's GAN foundation subject; crop-around-the-mask as the scaling trick that makes 8K interactive editing possible; seam-blending and a masked-region-specific evaluation suite (FID is not enough); and a system design with content-credential signing as a first-class pipeline stage.

30 min 1 enrolled
Pro Machine Learning
Advanced

MLOps Model Lifecycle & Governance

Learn the complete model lifecycle as a single governed pipeline, why every stage transition needs an owner and a sign-off, how to version data/code/model/environment so any prediction is reproducible months later, and what a model card documents and why regulated teams cannot ship without one.

30 min 1 enrolled
Free Machine Learning
Beginner

ML System Design Interview Framework

Learn how ML system design interviews are scored, a 7-step framework from requirements to monitoring, back-of-envelope estimation for QPS, embeddings and GPU cost, common problem framings, and the mistakes that sink strong candidates.

30 min
Pro Machine Learning
Intermediate

ML Data Pipelines & Feature Stores

Design the data side of an ML system: logging for train/serve parity, labelling strategies, point-in-time joins, negative sampling, batch vs streaming features, feature stores, training–serving skew, data validation, lineage and a worked notification-click pipeline.

30 min
Pro Machine Learning
Advanced

Generative Adversarial Networks: From Minimax to StyleGAN

The foundation subject the rest of this track's business cases depend on: why a single forward pass still matters in 2026 (real-time serving, the adversarial loss hiding inside every VAE/VQGAN decoder, GAN-based distillation of diffusion models, still-SOTA perceptual super-resolution and fast inpainting) argued against diffusion and autoregressive generation on latency, quality, diversity, training stability and controllability; the generator-vs-discriminator minimax game and the non-saturating loss at interview altitude; the failure-mode ladder — mode collapse, vanishing gradients, high-resolution instability — and each fix (minibatch discrimination, WGAN/WGAN-GP, progressive growing); StyleGAN's mapping network, AdaIN/modulated convolutions, per-layer style injection, noise inputs, style mixing and the truncation trick, with StyleGAN2's fixes at one-paragraph altitude; controllability as a product feature (InterFaceGAN-style latent directions, GAN inversion, identity preservation); FFHQ-style data and the output-diversity requirement; FID/IS plus precision/recall for generative models and a bias audit; and a synchronous generation-service design with latent storage, moderation and the consent/deepfake question.

30 min 1 enrolled
Pro Machine Learning
Advanced

Case Study: AI Product Photography for a Million-SKU Catalog

Model interview answer for AI product photography (Amazon ad-image generator / Shopify Magic / Photoroom-class): why the right design is segment-first (a U²-Net/BiRefNet-class matting model, discriminative and cheap) then background generation conditioned on the cutout, never end-to-end regeneration of the product, argued from a fidelity constraint the GAN-vs-diffusion comparison table alone doesn't resolve; product-region invariance as the case's non-negotiable evaluation gate; mask-distribution and template-library vocabulary reused from the generative-fill case; and a genuinely batch-offline scalability story — throughput-optimized GPU batching, template reuse, spot instances and an automatic regenerate loop — that is deliberately unlike the interactive cases (virtual try-on, generative fill) already written in this track.

30 min 1 enrolled
Pro Machine Learning
Intermediate

Model Serving & Deployment

Learn to design the serving side of an ML system: batch vs online vs streaming vs hybrid inference, latency budgets with numbers, model server patterns, compression, canary and shadow rollouts, feature-store consistency, autoscaling, and safe fallbacks.

28 min
Pro Machine Learning
Intermediate

ML Monitoring, Drift & Retraining

Learn why deployed models decay, how to monitor them in layers from system health to business KPIs, compute PSI and other drift statistics with worked numbers, handle delayed labels, and design retraining triggers with safe validation gates and rollback.

28 min
Pro Machine Learning
Intermediate

Feature Engineering

Learn the feature engineering toolkit data scientists are tested on: scaling and transforms per model family, categorical encodings including leakage-safe target encoding, missing-value strategies, cyclical time features, point-in-time aggregations, feature selection, and sklearn pipelines that prevent training/serving skew.

30 min
Pro Machine Learning
Advanced

Case Study: Real-Time Payment Fraud Detection (Stripe / PayPal-style)

Model interview answer for real-time payment fraud detection: sub-100 ms decisions, 0.1 % positives with asymmetric costs, delayed chargeback labels, point-in-time features and velocity counters, GBDT plus graph signals, PR-AUC and dollar-weighted evaluation, cost-based decision policy and a streaming serving architecture.

30 min
Pro Machine Learning
Advanced

Case Study: Video Recommendation (YouTube / Netflix-style Homepage)

Walk through a model ML system design interview answer for a video recommendation homepage: objectives, implicit-feedback labels, two-tower retrieval, multi-task ranking, position bias, cold start, evaluation, A/B testing, serving at scale, and monitoring.

30 min
Pro Machine Learning
Intermediate

Clustering & Dimensionality Reduction

Learn k-means, hierarchical clustering, DBSCAN and Gaussian mixtures, how to choose k and evaluate clusters without labels, and how PCA, t-SNE and UMAP compress high-dimensional data — with a worked PCA example and interview traps.

28 min
Pro Machine Learning
Advanced

Case Study: Search Ranking (Airbnb / E-commerce Marketplace Search)

Model interview answer for marketplace search ranking: two-sided objectives, retrieval vs learned ranking, position-bias-corrected labels, LambdaMART vs neural rankers, NDCG worked example, cold start, diversity re-ranking, serving architecture and monitoring.

30 min
Pro Machine Learning
Advanced

Off-Policy Evaluation & Offline RL

Learn why you cannot safely A/B test an untested RL policy in production, importance sampling and per-decision importance sampling for off-policy evaluation, IPS/SNIPS variance reduction, the doubly-robust estimator, the distribution-shift problem in offline RL and how Conservative Q-Learning addresses it conceptually, and why off-policy evaluation is the mandatory gate before any production RL launch.

30 min
Pro Machine Learning
Advanced

Case Study: Virtual Try-On for Fashion E-commerce

Model interview answer for virtual try-on (Zalando/ASOS/Walmart-class fashion retailer): the catalog-mode/personal-mode split as the case's signature design fork; classical warping plus GAN (CP-VTON/VITON-HD lineage) versus Google's TryOnDiffusion cascaded parallel-UNet cross-attention approach versus a DreamBooth-style personalization shortcut that fails at 1M-SKU scale, argued against the GAN-vs-diffusion comparison table from this track's GAN foundation subject; paired on-model/flat-lay data and pseudo-pair construction; garment-agnostic person representation and the losses that protect print fidelity; a garment-fidelity evaluation suite (FID is not enough) including OCR-based print-text consistency and the selection-bias trap in try-on-usage-vs-return-rate analysis; precompute-the-catalog and cascade-preview-then-refine as the scaling patterns; and the cost-per-try-on-vs-cost-of-one-return ROI argument.

30 min 1 enrolled

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.