Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Compute IPS and SNIPS for a Candidate Ranking Policy

A search-ranking team logged the following data under behavior policy \pi_b (a heuristic ranker) and wants to evaluate a candidate learned policy \pi_e before considering any live test:

i \pi_b(a_i\mid x_i) \pi_e(a_i\mid x_i) r_i (clicked = 1)
1 0.40 0.30 1
2 0.10 0.05 0
3 0.60 0.65 1
4 0.02 0.40 1
5 0.25 0.20 0
  1. Compute the importance weight w_i for each row and the raw IPS estimate \hat V_{\text{IPS}}(\pi_e).
  2. Compute the SNIPS estimate \hat V_{\text{SNIPS}}(\pi_e) and the effective sample size n_{\text{eff}}. What do these two numbers together tell you about how much to trust this evaluation?
  3. A colleague argues "IPS is unbiased, so it's the more rigorous estimator — we should report \hat V_{\text{IPS}} to leadership, not SNIPS." Explain what is right and what is misleading about this argument in the context of a launch decision.
Solution

1. Weights and raw IPS

i w_i=\pi_e/\pi_b r_i w_i r_i
1 0.75 1 0.75
2 0.50 0 0.00
3 1.083 1 1.083
4 20.00 1 20.00
5 0.80 0 0.00

\hat V_{\text{IPS}} = (0.75+0.00+1.083+20.00+0.00)/5 = 21.833/5 \approx 4.367 — an estimate above 1 for a reward bounded in [0,1], an immediate red flag that this raw estimate should not be trusted at face value.

2. SNIPS and effective sample size

\sum w_i = 0.75+0.50+1.083+20.00+0.80 = 23.133. \hat V_{\text{SNIPS}} = 21.833 / 23.133 \approx 0.944.

\sum w_i^2 = 0.5625+0.25+1.173+400+0.64 = 402.63. n_{\text{eff}} = (23.133)^2 / 402.63 = 535.14/402.63 \approx 1.33.

SNIPS gives a plausible, bounded estimate (≈0.944), but the effective sample size of about 1.33 out of 5 nominal rows says the estimate is driven almost entirely by row 4 alone (weight 20, by far the largest). Together, these numbers say: the point estimate itself is not unreasonable, but it should not be trusted as a stable, well-supported estimate — it would likely change substantially with a different sample of logged rows, since one row with an unusually low logging propensity (0.02) is doing almost all the work. This is not enough data to gate a launch decision on; more logged data (specifically, more coverage of the action row 4 represents) is needed before committing.

3. What's right and what's misleading about the colleague's argument

It is correct that IPS is unbiased in expectation while SNIPS is not (SNIPS is a ratio of random quantities and is only asymptotically unbiased/consistent) — that part of the claim is true. What is misleading is treating "unbiased" as equivalent to "more trustworthy for this decision": an estimator's bias is only one half of the relevant statistical picture, and here IPS's estimate (4.367) is obviously implausible on its face for a bounded reward, entirely because of its enormous variance driven by one extreme weight — an unbiased estimator with variance this large produces individual estimates that are frequently far from the truth, which is exactly the opposite of what a launch decision needs. In practice, for a finite, launch-gating sample, a slightly biased but low-variance, bounded estimate (SNIPS, or ideally a doubly robust estimate) that comes with a tight, defensible confidence interval is far more useful than a theoretically unbiased estimate with a confidence interval so wide it cannot distinguish "ship it" from "don't." The right answer to present to leadership is the SNIPS/DR estimate together with the effective sample size or a bootstrapped confidence interval — not raw IPS presented as if "unbiased" settled the question.

Share this question

← Back to Off-Policy Evaluation & Offline RL practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.