Match a job Paths Subjects Questions Quizzes Pricing
Intermediate Open Pro

Compute LinUCB Scores by Hand for Two Arms

Two notification creatives, Urgent and Friendly, are scored by LinUCB with 2-dimensional context x = [\text{days\_since\_last\_open (normalized)}, 1] (bias term), \lambda = 1, \alpha = 2.0.

Current state: Urgent has \hat\theta_{\text{Urgent}} = [0.5, 0.1] and, for the query context x_t = [0.6, 1], an uncertainty term \sqrt{x_t^\top A_{\text{Urgent}}^{-1} x_t} = 0.12. Friendly has \hat\theta_{\text{Friendly}} = [0.2, 0.3] and, at the same context, uncertainty term 0.30.

  1. Compute each arm's predicted mean reward and UCB score at x_t = [0.6, 1].
  2. Which arm does LinUCB select, and is it the arm with the higher predicted mean? Explain the mechanism.
  3. Suppose Friendly is a brand-new creative added yesterday with only 8 impressions so far, while Urgent has 3,000. Explain qualitatively why their uncertainty terms differ the way they do, and what you'd expect to happen to Friendly's uncertainty term over the next few hundred impressions if its reward pattern is genuinely similar to Urgent's.

Share this question

← Back to Contextual Bandits for Personalization practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.