Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Design a Doubly Robust OPE Pipeline for a Recommender Launch

A video recommender team wants to evaluate a new ranking policy offline before any online test, using six months of logged impressions (each with the recommended item, the logging policy's propensity for that item, and whether the user watched it). They ask you to design the OPE pipeline.

  1. Describe the two components you would build (and what each one is trained or computed from) to produce a doubly robust estimate, and explain in your own words why combining them is more robust than using either alone.
  2. The reward model component performs very well (low error) on popular items that were shown often, but has almost no training signal for items that were rarely surfaced by the logging policy — exactly the items the new policy tends to rank highly. What does this imply about how much you should trust the resulting DR estimate, and what would you do about it?
  3. Propose two concrete checks you would run before trusting the final DR point estimate enough to recommend a launch decision.

Share this question

← Back to Off-Policy Evaluation & Offline RL practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.