ML Data Pipelines & Feature Stores
Ask any team what broke their last model in production and the answer is rarely "the architecture". It is a label that arrived late, a feature computed one way offline and another way online, a join that leaked tomorrow's data into today's row, or a schema change upstream nobody noticed. In the ML system design interview this maps directly onto Step 3 (data & labels) and Step 4 (features & pipeline) of the framework from the ML System Design Interview Framework subject — and interviewers weight those steps heavily because they are where real systems fail.
This subject teaches you to design that layer end to end: what to log, how to turn events into labels, how to build a training set that respects time, how to compute features in batch and streaming, what a feature store actually is, and how to keep training and serving consistent. It ends with a worked design — the training set for a "will the user click this notification" model — that you can reproduce on a whiteboard.
Boundaries: model training, hyperparameter search and experiment tracking are covered in the Model Training & Experimentation subject; detecting drift once the model is live is covered in ML Monitoring & Drift. Generic stream-processing and storage infrastructure (Kafka, object stores, partitioning) is assumed from the System Design Interview track.
The Data Flywheel
Production ML is a loop, not a pipeline. The model's outputs change what users see, which changes what gets logged, which changes the next training set.
graph LR
USERS["users"] --> SERVING["serving<br/>(model vN)"]
SERVING --> LOGS["event logs<br/>(impr, click,<br/>features)"]
LOGS --> LABELS["labels +<br/>training set"]
LABELS --> TRAIN["model vN+1<br/>(training)"]
TRAIN --> SERVING
Two consequences you should say out loud in an interview:
- Selection bias: you only observe outcomes for items the current model chose to show. Without some exploration traffic (a small random or epsilon-greedy slice), the training set never contains evidence that the model is wrong about items it never shows.
- The pipeline is a product: every model improvement is capped by label quality, feature freshness and logging fidelity. Investing in the flywheel compounds; investing in the model alone plateaus.
Logging and Event Collection
The training set is built from what you logged, so logging is a design decision, not an afterthought.
What to log at serve time
| Log | Why |
|---|---|
| Request context (user id, session, device, timestamp, page, slot/position) | Context features and position-bias correction |
| The exact feature values the model consumed | Training/serving parity — you train on what the model saw, not a re-derived approximation |
| Model id/version and raw score | Debugging, calibration checks, counterfactual analysis |
| Items shown (impressions), including those not clicked | You need negatives; "clicked" alone is useless |
| Exploration flag / propensity | For unbiased evaluation and off-policy learning |
| Outcome events (click, purchase, dismiss) with their own timestamps | Labels |
The single most valuable habit: log the feature vector at inference time ("feature logging" / "log-and-wait"). It sidesteps most point-in-time-join problems because the row already reflects the world as the model saw it. The cost is storage — a 200-feature row at ~1 KB, at 10⁹ impressions/day, is ~1 TB/day — so teams downsample negatives at log time and keep all positives.
Event schema basics
Give every event a stable event_id, an event_time (when it happened on the client) and an ingest_time (when the backend received it). Late-arriving events are normal (offline devices, retries); every downstream job must be defined in terms of event_time with an explicit lateness allowance, or it silently drops or double-counts.