Off-Policy Evaluation & Offline RL
Estimate how a new policy would perform before it ever sees a real user, and know when offline data alone is enough to train one
Learn why you cannot safely A/B test an untested RL policy in production, importance sampling and per-decision importance sampling for off-policy evaluation, IPS/SNIPS variance reduction, the doubly-robust estimator, the distribution-shift problem in offline RL and how Conservative Q-Learning addresses it conceptually, and why off-policy evaluation is the mandatory gate before any production RL launch.
Practice questions (5)
-
View →
Compute IPS and SNIPS for a Candidate Ranking Policy
Advanced · Free -
View →
Design a Doubly Robust OPE Pipeline for a Recommender Launch
Advanced -
View →
Diagnose Extrapolation Error in an Offline-Trained Pricing Agent
Advanced -
View →
Apply Per-Decision Importance Sampling to a Multi-Step Session
Advanced -
View →
Design the Full Pre-Launch Gate for a New Bidding Policy
Advanced