Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Trace a Manual Q-Learning Update for a Cloud Autoscaler

A GPU autoscaling agent uses tabular Q-learning with three states — Low, Medium, High queue depth — and three actions — scale_up, hold, scale_down. The current Q-table is:

State scale_up hold scale_down
Low −2.0 −0.5 −1.0
Medium −3.0 −2.5 −4.0
High −7.0 −6.0 −9.0

The agent is in state Medium, takes action scale_up (chosen by ε-greedy exploration), observes reward r = -1.0, and transitions to state Low. Use \alpha = 0.2 and \gamma = 0.95.

  1. Compute the updated value of Q(\text{Medium}, \text{scale\_up}), showing the TD target and TD error explicitly.
  2. Explain, in terms of the update rule itself, why this is considered off-policy learning — specifically, what would have to change about the update for it to become on-policy (SARSA)?
  3. After this single update, would the agent's greedy policy in state Medium change? Would it change in state Low? Justify from the table.
Solution

1. The update

\max_{a'} Q(\text{Low}, a') = \max(-2.0, -0.5, -1.0) = -0.5
\text{TD target} = r + \gamma \max_{a'} Q(s', a') = -1.0 + 0.95 \times (-0.5) = -1.0 - 0.475 = -1.475
\delta = \text{target} - Q(\text{Medium}, \text{scale\_up}) = -1.475 - (-3.0) = 1.525
Q(\text{Medium}, \text{scale\_up}) \leftarrow -3.0 + 0.2 \times 1.525 = -3.0 + 0.305 = -2.695

The updated value is ≈ −2.695 — the action looks meaningfully better than the old estimate of −3.0, because the transition landed in a cheap state (Low) with a small penalty.

2. Why this is off-policy

The TD target uses \max_{a'} Q(\text{Low}, a') — the value of the best action available in Low (hold, at −0.5) — regardless of which action the agent actually takes next. The agent might, on the next step, explore into scale_down instead; the update does not care. This is what makes Q-learning off-policy: the update target assumes the greedy target policy from the next state onward, while the behavior policy generating the trajectory (ε-greedy here) can pick something else. To make it on-policy (SARSA), you would replace \max_{a'} Q(s', a') with Q(s', a'_{\text{actual}}), where a'_{\text{actual}} is whatever action the agent genuinely takes next according to its behavior policy — the update would then have to wait until that next action is chosen before it could be applied.

3. Effect on the greedy policy

In Medium, before the update the greedy action was hold (−2.5, the least negative of −3.0, −2.5, −4.0). After the update, scale_up is now −2.695, still worse than hold at −2.5 — so the greedy policy in Medium does not change after this single update, though the gap narrowed from 0.5 to about 0.195, and a couple more updates like this one could flip it. In Low, nothing in that row was touched by this update at all (only Q(Medium, scale_up) changed), so the greedy action there remains hold (−0.5) — unchanged, trivially, because no value in that row was modified.

Share this question

← Back to Value-Based Methods: Q-Learning to DQN practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.