Practice — RL in Production: Safe Exploration & Serving (5 questions)
Advanced
Open
Free
Diagnose a Missing Propensity in Production Logs Permalink →
A ride-pricing team deployed a contextual bandit six months ago to choose a surge multiplier (1.0x, 1.2x, 1.5x, 2.0x) per city-zone every 5 minutes. The serving log contains: timestamp, zone features, the multiplier that was served, and the resulting ride-acceptance rate. It does not contain the probability the policy assigned to the served multiplier.
A new pricing research team now wants to answer, without running a new live experiment: "if we had used a slightly more conservative policy last quarter, how would acceptance rate have changed?"
- Explain precisely why this log cannot answer that question, even though it records the action taken and the outcome.
- The team proposes reconstructing propensities retroactively by re-running last quarter's model checkpoint against the logged features. What has to be true for that reconstruction to be valid, and name two concrete ways production systems violate it.
- Going forward, specify exactly what should be added to each log line (be specific — not just "log the propensity") so this situation does not recur, including how it should handle any post-hoc business overrides on the multiplier.
Share this question
Advanced
Open
Pro
Reward Penalty vs. Structural Guardrail for a Hard Notification Cap
Unlock this question →
Advanced
Open
Pro