Advanced
Open
Pro
Design the Full Pre-Launch Gate for a New Bidding Policy
An ad-bidding RL agent has been retrained on 4 months of logged auction data. Leadership wants to know the process between "training finished" and "policy live on 100% of traffic." A naive proposal on the table is: "run OPE, and if the estimate looks better than the current policy, ship it to 100% of traffic next week."
- Explain concretely, with reference to this subject's tools, why "OPE looked good" is necessary but not sufficient to justify shipping straight to 100% of traffic.
- Design the actual gate sequence you would insist on, naming each stage and what decision criterion it uses.
- Suppose the OPE stage produces a doubly robust estimate showing the new policy is better by a wide margin, but the overlap check shows the new policy strongly favors a bidding range the logging policy almost never used. How does this change your confidence in the result, and what would you do differently before proceeding?
Share this question