Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Detect and Fix a Closing Feedback Loop

A customer-support routing policy learns, per incoming ticket, which of 4 support queues to route it to (general, billing, technical, retention-save). It is retrained weekly on the last 4 weeks of logged routing decisions and outcomes (resolution time, customer satisfaction score). Aggregate CSAT has been flat and acceptable for months. A new analyst, digging into the logs, finds that the "retention-save" queue has received under 0.5% of eligible tickets for the last 6 weeks, down from roughly 8% a year ago, even though the eligibility criteria for that queue haven't changed.

  1. Explain the mechanism by which weekly retraining on the policy's own logs produced this outcome, even though nobody changed the eligibility rule.
  2. Why is flat aggregate CSAT not reassuring here, and what monitoring should have caught this months earlier?
  3. Propose a concrete remediation, including how you'd safely re-introduce evidence about the retention-save queue without repeating the same collapse.

Share this question

← Back to RL in Production: Safe Exploration & Serving practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.