Match a job Paths Subjects Questions Quizzes Pricing

Contextual Bandits for Personalization

LinUCB, neural bandits, and action features — the workhorse algorithm behind notification and content personalization

Overview Read

Contextual Bandits for Personalization

Multi-Armed Bandits & Exploration built the machinery to answer "which of these five subject lines is best?" — and it deliberately ignored one thing every real personalization system depends on: who is asking. A 22-year-old who opens the app at midnight and a 55-year-old who opens it at 7am are not the same decision problem, but a plain bandit treats them identically, learns one global compromise arm, and serves it to everyone. If the best send-time truly differs by user segment — and in almost every real product it does — a context-free bandit's single best-in-hindsight arm is mediocre for practically every individual, even though it is, by construction, the best single choice averaged across the whole population.

A contextual bandit fixes this by conditioning the policy on a feature vector — the context — describing the current user and situation, and learning a separate (but jointly parameterized) reward estimate per action as a function of that context. This is the single most direct upgrade from "the simplest RL you'll ship" to "the RL that's actually doing the personalization work" in a production system: contextual bandits are the algorithm quietly picking which push notification to send at which time, which article thumbnail to show, which onboarding flow variant fits a new user, and which piece of creative to serve to which shopper. If you remember one sentence from this subject for an interview, make it this: contextual bandits condition the policy on features of the user/situation; plain (context-free) bandits do not — everything else follows from that one distinction.


From Bandits to Contextual Bandits: The Formalism

Recall the plain bandit setup: K arms, each with a fixed unknown mean \mu_a, and the algorithm's job is to find and exploit the best one. The contextual bandit generalizes this by making the reward a function of a context x_t \in \mathbb{R}^d observed at the start of each round:

\text{At round } t: \text{ observe } x_t \;\rightarrow\; \text{choose arm } A_t \;\rightarrow\; \text{observe reward } r_t \sim \text{(distribution depending on } x_t, A_t)

For notification send-time selection, x_t might encode: hour of day, day of week, device type, days since last open, historical response-rate-by-hour profile, and current session context. The arms are the send-time buckets (e.g., "now," "in 2h," "in 6h," "tomorrow morning"). The reward model is no longer a single scalar \mu_a per arm — it's a function f_a(x), typically parameterized, that the algorithm has to learn jointly with deciding which arm to pull.

The regret definition generalizes accordingly: instead of comparing to the single best fixed arm, contextual regret compares to the best arm given the context on each round:

\text{Regret}(T) = \sum_{t=1}^{T} \Big( \max_a \mathbb{E}[r \mid x_t, a] - \mathbb{E}[r \mid x_t, A_t] \Big)

This is a strictly harder objective than plain-bandit regret — the "oracle" you're being compared against is now allowed to pick a different best arm every round depending on context, not the one best arm overall. A contextual bandit that ignored context and just played the best context-free arm could easily have terrible contextual regret even while looking fine by plain-bandit standards, which is precisely the personalization gap this subject exists to close.


Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.