Practice — Off-Policy Evaluation & Offline RL (5 questions)
Advanced
Open
Free
Compute IPS and SNIPS for a Candidate Ranking Policy Permalink →
A search-ranking team logged the following data under behavior policy \pi_b (a heuristic ranker) and wants to evaluate a candidate learned policy \pi_e before considering any live test:
| i | \pi_b(a_i\mid x_i) | \pi_e(a_i\mid x_i) | r_i (clicked = 1) |
|---|---|---|---|
| 1 | 0.40 | 0.30 | 1 |
| 2 | 0.10 | 0.05 | 0 |
| 3 | 0.60 | 0.65 | 1 |
| 4 | 0.02 | 0.40 | 1 |
| 5 | 0.25 | 0.20 | 0 |
- Compute the importance weight w_i for each row and the raw IPS estimate \hat V_{\text{IPS}}(\pi_e).
- Compute the SNIPS estimate \hat V_{\text{SNIPS}}(\pi_e) and the effective sample size n_{\text{eff}}. What do these two numbers together tell you about how much to trust this evaluation?
- A colleague argues "IPS is unbiased, so it's the more rigorous estimator — we should report \hat V_{\text{IPS}} to leadership, not SNIPS." Explain what is right and what is misleading about this argument in the context of a launch decision.
Share this question
Advanced
Open
Pro
Design a Doubly Robust OPE Pipeline for a Recommender Launch
Unlock this question →
Advanced
Open
Pro
Diagnose Extrapolation Error in an Offline-Trained Pricing Agent
Unlock this question →
Advanced
Open
Pro