Trace a Manual Q-Learning Update for a Cloud Autoscaler
A GPU autoscaling agent uses tabular Q-learning with three states —
Low, Medium, High queue depth — and three actions —
scale_up, hold, scale_down. The current Q-table is:
| State | scale_up | hold | scale_down |
|---|---|---|---|
| Low | −2.0 | −0.5 | −1.0 |
| Medium | −3.0 | −2.5 | −4.0 |
| High | −7.0 | −6.0 | −9.0 |
The agent is in state Medium, takes action scale_up (chosen by
ε-greedy exploration), observes reward r = -1.0, and transitions to
state Low. Use \alpha = 0.2 and \gamma = 0.95.
- Compute the updated value of Q(\text{Medium}, \text{scale\_up}), showing the TD target and TD error explicitly.
- Explain, in terms of the update rule itself, why this is considered off-policy learning — specifically, what would have to change about the update for it to become on-policy (SARSA)?
- After this single update, would the agent's greedy policy in state
Mediumchange? Would it change in stateLow? Justify from the table.
1. The update
The updated value is ≈ −2.695 — the action looks meaningfully
better than the old estimate of −3.0, because the transition landed
in a cheap state (Low) with a small penalty.
2. Why this is off-policy
The TD target uses \max_{a'} Q(\text{Low}, a') — the value of the
best action available in Low (hold, at −0.5) — regardless of
which action the agent actually takes next. The agent might, on the
next step, explore into scale_down instead; the update does not
care. This is what makes Q-learning off-policy: the update target
assumes the greedy target policy from the next state onward, while
the behavior policy generating the trajectory (ε-greedy here) can
pick something else. To make it on-policy (SARSA), you would replace
\max_{a'} Q(s', a') with Q(s', a'_{\text{actual}}), where
a'_{\text{actual}} is whatever action the agent genuinely takes
next according to its behavior policy — the update would then have
to wait until that next action is chosen before it could be applied.
3. Effect on the greedy policy
In Medium, before the update the greedy action was hold (−2.5,
the least negative of −3.0, −2.5, −4.0). After the update,
scale_up is now −2.695, still worse than hold at −2.5 — so the
greedy policy in Medium does not change after this single
update, though the gap narrowed from 0.5 to about 0.195, and a
couple more updates like this one could flip it. In Low, nothing
in that row was touched by this update at all (only
Q(Medium, scale_up) changed), so the greedy action there remains
hold (−0.5) — unchanged, trivially, because no value in that row
was modified.
Share this question