Off-Policy Evaluation & Offline RL
A pricing team trains a reinforcement-learning agent to set discounts. In simulation it looks excellent — expected revenue up 12%. Someone proposes the obvious next step: run it as a live A/B test. Someone else points out the problem: a policy that has never touched real customers, evaluated only in a simulator that is necessarily an imperfect model of real demand, could just as easily set a discount that empties the margin on every order in a segment of the business for the two weeks before anyone notices and can call it off. Unlike a supervised classifier, whose worst-case failure is a bad prediction, a bad RL policy takes bad actions against real users, real inventory, and real money, and every one of those actions is live the instant the experiment starts. This is the problem this subject solves: how do you estimate what a new policy would do, using only data collected under an old policy, well enough to decide whether it is safe to test at all — before it ever gets a real A/B test.
This is the natural third subject in this track's arc. Reward Design & Delayed Credit Assignment covers how to shape and assign credit for a reward once you have one; Long-Term Value & Delayed Reward Systems covers what that reward should measure. This subject assumes you have a candidate policy — trained with the value-based, policy-gradient, or model-based methods from Value-Based Methods: Q-Learning to DQN, Policy Gradients & Actor-Critic, PPO & Modern Policy Optimization, or Model-Based RL & Planning — and answers the question that must be answered before it ever touches a real user: is it good, and how do you know, without the risk of finding out live? The forward link matters too: everything here consumes logged propensities — the probability the logging policy assigned to each action it actually took — and RL in Production: Safe Exploration & Serving is where that logging infrastructure is built and where the safe-rollout mechanics (canaries, guardrails, kill switches) that come after an OPE gate is passed are covered. This subject is the analytical gate; that one is the operational one.