Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Fix a Biased Churn Training Set Built from Censored Data

A data scientist builds a 30-day churn classifier by pulling "all users currently active or churned" from the warehouse each Monday, labeling each user 1 (churned) if they have cancelled as of today and 0 (retained) otherwise, using whatever tenure they currently have as a feature. The team notices the model's predicted churn rate is suspiciously low right after marketing pushes that spike new signups, and suspiciously higher during slow signup weeks.

  1. Explain precisely why this labeling scheme produces exactly that symptom.
  2. Redesign the label construction so that it does not have this bias, being specific about which rows are eligible to be used as training examples on a given Monday.
  3. Your proposed fix reduces the number of usable, labeled rows substantially compared to the original approach. Explain why this is an acceptable and necessary trade, and what you would do to partially compensate for the smaller effective training set.

Share this question

← Back to Long-Term Value & Delayed Reward Systems practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.