Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Off-Policy Evaluation & Offline RL (5 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Advanced Open Free

Compute IPS and SNIPS for a Candidate Ranking Policy Permalink →

A search-ranking team logged the following data under behavior policy \pi_b (a heuristic ranker) and wants to evaluate a candidate learned policy \pi_e before considering any live test:

i \pi_b(a_i\mid x_i) \pi_e(a_i\mid x_i) r_i (clicked = 1)
1 0.40 0.30 1
2 0.10 0.05 0
3 0.60 0.65 1
4 0.02 0.40 1
5 0.25 0.20 0
  1. Compute the importance weight w_i for each row and the raw IPS estimate \hat V_{\text{IPS}}(\pi_e).
  2. Compute the SNIPS estimate \hat V_{\text{SNIPS}}(\pi_e) and the effective sample size n_{\text{eff}}. What do these two numbers together tell you about how much to trust this evaluation?
  3. A colleague argues "IPS is unbiased, so it's the more rigorous estimator — we should report \hat V_{\text{IPS}} to leadership, not SNIPS." Explain what is right and what is misleading about this argument in the context of a launch decision.

Share this question

Advanced Open Pro

Design a Doubly Robust OPE Pipeline for a Recommender Launch

Unlock this question →
Advanced Open Pro

Diagnose Extrapolation Error in an Offline-Trained Pricing Agent

Unlock this question →
Advanced Open Pro

Apply Per-Decision Importance Sampling to a Multi-Step Session

Unlock this question →
Advanced Open Pro

Design the Full Pre-Launch Gate for a New Bidding Policy

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.