Advanced
Open
Pro
An Off-Policy Evaluation Estimate With a Wide Confidence Interval
Before a canary launch, your team runs an off-policy evaluation of a new RL notification policy against three weeks of logs collected under the current bandit. The point estimate shows the new policy's expected shaped return is 6% better than the current production policy — but the 95% confidence interval is [-9%, +21%], driven by a small number of logged decisions with very low propensities getting very large importance weights.
- Explain, mechanically, why decisions with low logged propensities produce this kind of wide interval, and why you cannot just "trust the point estimate" here.
- Propose two different concrete changes — one to how data is collected, one to how the estimate is computed — that would tighten this interval, and explain the trade-off each one introduces.
- Suppose leadership is impatient and proposes skipping straight to a full-traffic launch "since the point estimate is positive and we're behind schedule." Give the strongest argument against that, specific to this system (not a generic "always test first" answer).
Share this question