Machine Learning
Intermediate
Pro
ML Data Pipelines & Feature Stores
Turn raw events into leak-free training sets and consistent serving features — the part of the ML interview that separates builders from talkers
30 min read
7 views
Design the data side of an ML system: logging for train/serve parity, labelling strategies, point-in-time joins, negative sampling, batch vs streaming features, feature stores, training–serving skew, data validation, lineage and a worked notification-click pipeline.
Practice questions (5)
-
View →
Diagnosing a Point-in-Time Leak
Intermediate · Free -
View →
Choosing an Attribution Window for Delayed Conversions
Intermediate -
View →
Negative Sampling and Probability Recalibration
Intermediate -
View →
Investigating Training–Serving Skew
Advanced -
View →
Data Validation Gates and a Feature Backfill
Intermediate