Paths Subjects Questions Quizzes Pricing Search
Advanced Open Pro

Offline Lift That Vanished Online

A new ranking model improved offline NDCG@10 by 4 % over the production model when evaluated on last month's logs. In a 5 % A/B test the online engagement metric was flat, within noise.

  1. Give three plausible reasons for the offline–online gap in this setting.
  2. Explain in words what inverse propensity scoring (IPS) would have estimated here, what must be true of the logging system for it to work, and why its estimate can be very noisy.
  3. What two changes would you make to the experimentation process so the next model iteration is evaluated more reliably before its A/B?

Share this question

← Back to Model Training & Experimentation at Scale practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.