Match a job Paths Subjects Questions Quizzes Pricing

ML Data Pipelines & Feature Stores

Turn raw events into leak-free training sets and consistent serving features — the part of the ML interview that separates builders from talkers

Overview Read

ML Data Pipelines & Feature Stores

Ask any team what broke their last model in production and the answer is rarely "the architecture". It is a label that arrived late, a feature computed one way offline and another way online, a join that leaked tomorrow's data into today's row, or a schema change upstream nobody noticed. In the ML system design interview this maps directly onto Step 3 (data & labels) and Step 4 (features & pipeline) of the framework from the ML System Design Interview Framework subject — and interviewers weight those steps heavily because they are where real systems fail.

This subject teaches you to design that layer end to end: what to log, how to turn events into labels, how to build a training set that respects time, how to compute features in batch and streaming, what a feature store actually is, and how to keep training and serving consistent. It ends with a worked design — the training set for a "will the user click this notification" model — that you can reproduce on a whiteboard.

Boundaries: model training, hyperparameter search and experiment tracking are covered in the Model Training & Experimentation subject; detecting drift once the model is live is covered in ML Monitoring & Drift. Generic stream-processing and storage infrastructure (Kafka, object stores, partitioning) is assumed from the System Design Interview track.


The Data Flywheel

Production ML is a loop, not a pipeline. The model's outputs change what users see, which changes what gets logged, which changes the next training set.

graph LR
    USERS["users"] --> SERVING["serving<br/>(model vN)"]
    SERVING --> LOGS["event logs<br/>(impr, click,<br/>features)"]
    LOGS --> LABELS["labels +<br/>training set"]
    LABELS --> TRAIN["model vN+1<br/>(training)"]
    TRAIN --> SERVING

Two consequences you should say out loud in an interview:

  • Selection bias: you only observe outcomes for items the current model chose to show. Without some exploration traffic (a small random or epsilon-greedy slice), the training set never contains evidence that the model is wrong about items it never shows.
  • The pipeline is a product: every model improvement is capped by label quality, feature freshness and logging fidelity. Investing in the flywheel compounds; investing in the model alone plateaus.

Logging and Event Collection

The training set is built from what you logged, so logging is a design decision, not an afterthought.

What to log at serve time

Log Why
Request context (user id, session, device, timestamp, page, slot/position) Context features and position-bias correction
The exact feature values the model consumed Training/serving parity — you train on what the model saw, not a re-derived approximation
Model id/version and raw score Debugging, calibration checks, counterfactual analysis
Items shown (impressions), including those not clicked You need negatives; "clicked" alone is useless
Exploration flag / propensity For unbiased evaluation and off-policy learning
Outcome events (click, purchase, dismiss) with their own timestamps Labels

The single most valuable habit: log the feature vector at inference time ("feature logging" / "log-and-wait"). It sidesteps most point-in-time-join problems because the row already reflects the world as the model saw it. The cost is storage — a 200-feature row at ~1 KB, at 10⁹ impressions/day, is ~1 TB/day — so teams downsample negatives at log time and keep all positives.

Event schema basics

Give every event a stable event_id, an event_time (when it happened on the client) and an ingest_time (when the backend received it). Late-arriving events are normal (offline devices, retries); every downstream job must be defined in terms of event_time with an explicit lateness allowance, or it silently drops or double-counts.


Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.