Contextual Bandits for Personalization
LinUCB, neural bandits, and action features — the workhorse algorithm behind notification and content personalization
Learn how contextual bandits extend multi-armed bandits by conditioning the policy on features of the user and situation: the LinUCB algorithm and why a linear reward model gives closed-form confidence bounds, neural/deep contextual bandits for richer reward functions, action features vs. arm identity for scaling to large or changing catalogues, and position bias in bandit-served lists.
Practice questions (5)
-
View →
Justify a Contextual Bandit for Send-Time Personalization
Intermediate · Free -
View →
Compute LinUCB Scores by Hand for Two Arms
Intermediate -
View →
Design Cold Start for a Large, Turning-Over Creative Catalogue
Intermediate -
View →
Diagnose and Fix Position Bias in a Bandit-Ranked Carousel
Advanced -
View →
Evaluate a Proposed Offline Test of a New Bandit Policy
Advanced