Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Interpret and Compare Regret Curves for Two Deployed Policies

Your company ran two live bandit deployments for the same subject-line-selection problem (5 arms) over 6 months and logged cumulative regret weekly. Policy 1 (fixed ε = 0.2 ε-greedy) shows regret climbing at a roughly constant rate every week for all 26 weeks. Policy 2 (Thompson sampling) shows regret climbing quickly in the first few weeks, then visibly flattening out by week 10 onward.

  1. Explain what each regret curve shape implies about each policy's long-run behavior, using the linear vs. sublinear regret distinction.
  2. A stakeholder looks at total cumulative regret at week 26 and notes Policy 2's is lower, and asks: "does that mean Thompson sampling is strictly better in every week?" Answer precisely.
  3. Estimate, qualitatively, what would happen to each policy's regret curve shape if the best arm's true mean started drifting slowly starting in week 15 (a novelty-decay effect on the previously-best subject line).

Share this question

← Back to Multi-Armed Bandits & Exploration practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.