Compute IPS and SNIPS for a Candidate Ranking Policy
A search-ranking team logged the following data under behavior policy \pi_b (a heuristic ranker) and wants to evaluate a candidate learned policy \pi_e before considering any live test:
| i | \pi_b(a_i\mid x_i) | \pi_e(a_i\mid x_i) | r_i (clicked = 1) |
|---|---|---|---|
| 1 | 0.40 | 0.30 | 1 |
| 2 | 0.10 | 0.05 | 0 |
| 3 | 0.60 | 0.65 | 1 |
| 4 | 0.02 | 0.40 | 1 |
| 5 | 0.25 | 0.20 | 0 |
- Compute the importance weight w_i for each row and the raw IPS estimate \hat V_{\text{IPS}}(\pi_e).
- Compute the SNIPS estimate \hat V_{\text{SNIPS}}(\pi_e) and the effective sample size n_{\text{eff}}. What do these two numbers together tell you about how much to trust this evaluation?
- A colleague argues "IPS is unbiased, so it's the more rigorous estimator — we should report \hat V_{\text{IPS}} to leadership, not SNIPS." Explain what is right and what is misleading about this argument in the context of a launch decision.
1. Weights and raw IPS
| i | w_i=\pi_e/\pi_b | r_i | w_i r_i |
|---|---|---|---|
| 1 | 0.75 | 1 | 0.75 |
| 2 | 0.50 | 0 | 0.00 |
| 3 | 1.083 | 1 | 1.083 |
| 4 | 20.00 | 1 | 20.00 |
| 5 | 0.80 | 0 | 0.00 |
\hat V_{\text{IPS}} = (0.75+0.00+1.083+20.00+0.00)/5 = 21.833/5 \approx 4.367 — an estimate above 1 for a reward bounded in [0,1], an immediate red flag that this raw estimate should not be trusted at face value.
2. SNIPS and effective sample size
\sum w_i = 0.75+0.50+1.083+20.00+0.80 = 23.133. \hat V_{\text{SNIPS}} = 21.833 / 23.133 \approx 0.944.
\sum w_i^2 = 0.5625+0.25+1.173+400+0.64 = 402.63. n_{\text{eff}} = (23.133)^2 / 402.63 = 535.14/402.63 \approx 1.33.
SNIPS gives a plausible, bounded estimate (≈0.944), but the effective sample size of about 1.33 out of 5 nominal rows says the estimate is driven almost entirely by row 4 alone (weight 20, by far the largest). Together, these numbers say: the point estimate itself is not unreasonable, but it should not be trusted as a stable, well-supported estimate — it would likely change substantially with a different sample of logged rows, since one row with an unusually low logging propensity (0.02) is doing almost all the work. This is not enough data to gate a launch decision on; more logged data (specifically, more coverage of the action row 4 represents) is needed before committing.
3. What's right and what's misleading about the colleague's argument
It is correct that IPS is unbiased in expectation while SNIPS is not (SNIPS is a ratio of random quantities and is only asymptotically unbiased/consistent) — that part of the claim is true. What is misleading is treating "unbiased" as equivalent to "more trustworthy for this decision": an estimator's bias is only one half of the relevant statistical picture, and here IPS's estimate (4.367) is obviously implausible on its face for a bounded reward, entirely because of its enormous variance driven by one extreme weight — an unbiased estimator with variance this large produces individual estimates that are frequently far from the truth, which is exactly the opposite of what a launch decision needs. In practice, for a finite, launch-gating sample, a slightly biased but low-variance, bounded estimate (SNIPS, or ideally a doubly robust estimate) that comes with a tight, defensible confidence interval is far more useful than a theoretically unbiased estimate with a confidence interval so wide it cannot distinguish "ship it" from "don't." The right answer to present to leadership is the SNIPS/DR estimate together with the effective sample size or a bootstrapped confidence interval — not raw IPS presented as if "unbiased" settled the question.
Share this question