Advanced
Open
Pro
Diagnose Extrapolation Error in an Offline-Trained Pricing Agent
A team trains a discount-pricing agent purely offline (no live environment interaction) on 18 months of logged pricing decisions and outcomes, using standard deep Q-learning with the logged data as a replay buffer. The historical policy almost always chose discounts between 5% and 20%; discounts above 30% appear in fewer than 0.1% of logged rows, mostly from a short-lived, poorly-targeted test months earlier. After training, the learned policy's Q-function ranks a 45% discount as the best action in a wide range of contexts, even though nothing resembling strong evidence for that in the data.
- Explain the specific mechanism by which vanilla Q-learning applied offline produced this outcome, referencing the role of the max operator and the lack of an environment to correct it.
- Explain, conceptually, how Conservative Q-Learning (CQL) would change the training objective to prevent this specific failure, and what its Q-function would look like differently for the 45% discount action compared to vanilla Q-learning.
- Even after switching to CQL, what would you still want to see before trusting this policy in production, and why does CQL alone not fully substitute for that?
Share this question