Intermediate
Open
Pro
Compute LinUCB Scores by Hand for Two Arms
Two notification creatives, Urgent and Friendly, are scored by
LinUCB with 2-dimensional context x = [\text{days\_since\_last\_open
(normalized)}, 1] (bias term), \lambda = 1, \alpha = 2.0.
Current state: Urgent has \hat\theta_{\text{Urgent}} = [0.5,
0.1] and, for the query context x_t = [0.6, 1], an uncertainty
term \sqrt{x_t^\top A_{\text{Urgent}}^{-1} x_t} = 0.12. Friendly
has \hat\theta_{\text{Friendly}} = [0.2, 0.3] and, at the same
context, uncertainty term 0.30.
- Compute each arm's predicted mean reward and UCB score at x_t = [0.6, 1].
- Which arm does LinUCB select, and is it the arm with the higher predicted mean? Explain the mechanism.
- Suppose
Friendlyis a brand-new creative added yesterday with only 8 impressions so far, whileUrgenthas 3,000. Explain qualitatively why their uncertainty terms differ the way they do, and what you'd expect to happen toFriendly's uncertainty term over the next few hundred impressions if its reward pattern is genuinely similar toUrgent's.
Share this question