Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Evaluate a Proposed Offline Test of a New Bandit Policy

Your team has six months of logs from a deployed LinUCB send-time policy: for every impression, you have the context, the action taken, and the observed reward — but propensities (the probability the deployed policy assigned to the chosen action) were never logged. A data scientist proposes: "let's just replay this log against our new candidate policy — for every logged impression, if the new policy would have chosen the same action as the old one did, count the logged reward toward the new policy's estimated performance."

  1. Explain what this proposed method is doing, and what quantity it is trying to estimate.
  2. Identify the specific way the missing propensities make this approach worse than it needs to be, even setting aside the general bias concerns of off-policy evaluation.
  3. What is the minimum change to the logging pipeline that would most improve future ability to evaluate candidate policies offline, and why does it matter even though it doesn't fix the past six months of data?

Share this question

← Back to Contextual Bandits for Personalization practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.