Match a job Paths Subjects Questions Quizzes Pricing

RL in Production: Safe Exploration & Serving

Log propensities, cap the blast radius, and serve a learned policy without betting the business on it

Overview Read

RL in Production: Safe Exploration & Serving

Imagine a messaging product's notification system is upgraded from a rules engine ("send at 9am if the user hasn't opened the app in 2 days") to a learned policy that picks, for every user every day, whether to send a notification and which of several messages to send, trained to maximize 30-day retention. The offline numbers look great. It ships to 100% of traffic on a Friday. By Monday, the on-call engineer discovers the policy has learned that sending six notifications a day to a specific segment of anxious re-engagement-prone users modestly increases next-day opens, and it has been doing that to two million people for 60 hours. Nothing crashed. No error budget was burned. The policy did exactly what it was trained to do — it just did not know that "modest short-term lift, catastrophic unsubscribe rate" was a bad trade, because nobody told it that trade was off the table.

This is the central fact that distinguishes production reinforcement learning from serving an ordinary classifier, and it is worth stating before anything else: a supervised model that is wrong makes a bad prediction; an RL policy that is wrong takes a bad action. A prediction sits in a log until someone reads it. An action is executed in the world — a notification is sent, a price is set, a bid is placed, a video autoplays — and it cannot be un-executed. The other half of the same fact is that an action also changes what data you will have tomorrow: the policy chose what to try, so the policy's own history determines which parts of the world it ever gets evidence about. Get the serving discipline wrong and you do not just get a wrong number — you get a system that damages users at scale while quietly poisoning its own ability to learn from the mistake.

This subject is the systems companion to the rest of the reinforcement learning track: it assumes you already know what a policy, a reward and an MDP are (the RL Foundations subject), and it does not re-derive bandit or policy-gradient algorithms (the Multi-Armed Bandits, Contextual Bandits, Value-Based Methods and Policy Gradients subjects own that). What it teaches is what an interviewer means when they ask "okay, you have a trained policy — how do you actually run it in production?" The four pillars are propensity logging, guardrails, bounded exploration, and closing the loop between what you serve and what you retrain on. The Off-Policy Evaluation & Offline RL subject builds directly on the logging discipline taught here.


Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.