Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Diagnose a Missing Propensity in Production Logs

A ride-pricing team deployed a contextual bandit six months ago to choose a surge multiplier (1.0x, 1.2x, 1.5x, 2.0x) per city-zone every 5 minutes. The serving log contains: timestamp, zone features, the multiplier that was served, and the resulting ride-acceptance rate. It does not contain the probability the policy assigned to the served multiplier.

A new pricing research team now wants to answer, without running a new live experiment: "if we had used a slightly more conservative policy last quarter, how would acceptance rate have changed?"

  1. Explain precisely why this log cannot answer that question, even though it records the action taken and the outcome.
  2. The team proposes reconstructing propensities retroactively by re-running last quarter's model checkpoint against the logged features. What has to be true for that reconstruction to be valid, and name two concrete ways production systems violate it.
  3. Going forward, specify exactly what should be added to each log line (be specific — not just "log the propensity") so this situation does not recur, including how it should handle any post-hoc business overrides on the multiplier.
Solution

1. Why the existing log cannot answer the question

Off-policy evaluation works by re-weighting each logged (state, action, reward) tuple by how much more or less likely the candidate policy would have been to take that same action compared to the policy that actually served it — an importance weight of the form π_candidate(a|s) / π_logging(a|s). The denominator, the logging policy's propensity, is not in the data. Knowing that multiplier 1.5x was served and that acceptance was, say, 62% tells you what happened under whatever the policy's degree of confidence was in 1.5x — but not how confident it was, so there is no way to say how much weight this row should carry when asking "what if a different policy had been more or less likely to pick 1.5x here." Without the propensity the log supports only on-policy questions (what did this policy do and what happened) not counterfactual ones.

2. Conditions for valid retroactive reconstruction, and how they break

Reconstruction is only valid if replaying the checkpoint against the logged features reproduces exactly the probability distribution the live system actually sampled from at serving time — i.e., the checkpoint, the feature values, and the sampling logic must all be byte-for-byte what was live. Two realistic ways this breaks: (a) features are recomputed from a warehouse table after the fact rather than replayed from what was logged at serving time, and the warehouse values differ slightly from what was live (a classic point-in-time correctness failure); (b) the live system applied a post-decision guardrail (e.g., a pricing cap during a public-safety event) that changed the effective probability of the served action, and the checkpoint alone — without knowledge of that override — reconstructs the raw model's propensity, not the propensity of the composite system that actually served the price.

3. What to log going forward

Each log line should contain: the full context/state snapshot (or a hash sufficient to reconstruct it exactly, logged at serving time, not recomputed later), the candidate action set available at that decision (the multiplier menu can vary by zone/regulation), the action served, the effective propensity of the action served — meaning the probability produced by the full serving pipeline including any post-decision override, not just the raw bandit's output — the policy version ID, and a flag plus rule ID for any guardrail that fired. If a pricing-cap override deterministically forces the multiplier to 1.2x regardless of what the bandit wanted, the effective propensity logged for that decision is 1.0 for 1.2x, not the bandit's raw score for 1.2x. This makes every future off-policy evaluation query answerable without a new live experiment, and correctly reflects the region where overrides bind most (public-safety events, regulatory caps) rather than silently mis-crediting that region to the learned policy.

Share this question

← Back to RL in Production: Safe Exploration & Serving practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.