Practice — Value-Based Methods: Q-Learning to DQN (6 questions)
Trace a Manual Q-Learning Update for a Cloud Autoscaler Permalink →
A GPU autoscaling agent uses tabular Q-learning with three states —
Low, Medium, High queue depth — and three actions —
scale_up, hold, scale_down. The current Q-table is:
| State | scale_up | hold | scale_down |
|---|---|---|---|
| Low | −2.0 | −0.5 | −1.0 |
| Medium | −3.0 | −2.5 | −4.0 |
| High | −7.0 | −6.0 | −9.0 |
The agent is in state Medium, takes action scale_up (chosen by
ε-greedy exploration), observes reward r = -1.0, and transitions to
state Low. Use \alpha = 0.2 and \gamma = 0.95.
- Compute the updated value of Q(\text{Medium}, \text{scale\_up}), showing the TD target and TD error explicitly.
- Explain, in terms of the update rule itself, why this is considered off-policy learning — specifically, what would have to change about the update for it to become on-policy (SARSA)?
- After this single update, would the agent's greedy policy in state
Mediumchange? Would it change in stateLow? Justify from the table.
Share this question
Design and Debug a Replay Buffer for a Recommendation Agent
Unlock this question →Why Q-Learning Falls Off the Cliff More Than SARSA Permalink →
The classic "Cliff Walking" gridworld: an agent starts at one corner of a grid and must reach a goal in the opposite corner. A row of cells along the bottom edge is a cliff — stepping into it gives a reward of -100 and sends the agent back to start. Every other step costs -1. The shortest path runs directly along the cliff edge; a longer, safer path stays one row away from it.
You train two agents on this task with the same \varepsilon-greedy exploration schedule (\varepsilon = 0.1 throughout, never annealed): one with tabular Q-learning, one with tabular SARSA. Both eventually converge to stable value estimates.
During training itself — while \varepsilon-greedy exploration is still active and both agents are still occasionally taking random actions — which agent earns the higher average reward per episode?
Share this question