Practice — ML Data Pipelines & Feature Stores (5 questions)
Intermediate
Open
Free
Diagnosing a Point-in-Time Leak Permalink →
A team trains a churn model on rows of the form
(user_id, snapshot_date, label = churned within 30 days), joining a
user_features table that is fully recomputed every Friday night with
columns such as sessions_last_30d, support_tickets_last_30d and
days_since_last_login. Offline PR-AUC is 0.91; in production the
first month's PR-AUC is 0.52.
- Explain the most likely cause of the gap, with a concrete example of how a training row is contaminated.
- Rewrite the join so that it is point-in-time correct (SQL or pseudocode is fine).
- One of the three features has a second leakage problem unrelated to the join. Which one, and what would you do about it?
Share this question