Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Value-Based Methods: Q-Learning to DQN (6 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Advanced Open Free

Trace a Manual Q-Learning Update for a Cloud Autoscaler Permalink →

A GPU autoscaling agent uses tabular Q-learning with three states — Low, Medium, High queue depth — and three actions — scale_up, hold, scale_down. The current Q-table is:

State scale_up hold scale_down
Low −2.0 −0.5 −1.0
Medium −3.0 −2.5 −4.0
High −7.0 −6.0 −9.0

The agent is in state Medium, takes action scale_up (chosen by ε-greedy exploration), observes reward r = -1.0, and transitions to state Low. Use \alpha = 0.2 and \gamma = 0.95.

  1. Compute the updated value of Q(\text{Medium}, \text{scale\_up}), showing the TD target and TD error explicitly.
  2. Explain, in terms of the update rule itself, why this is considered off-policy learning — specifically, what would have to change about the update for it to become on-policy (SARSA)?
  3. After this single update, would the agent's greedy policy in state Medium change? Would it change in state Low? Justify from the table.

Share this question

Advanced Open Pro

Migrate a Tabular Agent to Function Approximation

Unlock this question →
Advanced Open Pro

Diagnose a Diverging Training Run

Unlock this question →
Advanced Open Pro

Quantify and Fix Overestimation Bias in a Trading Agent

Unlock this question →
Advanced Open Pro

Design and Debug a Replay Buffer for a Recommendation Agent

Unlock this question →
Advanced Open Free

Why Q-Learning Falls Off the Cliff More Than SARSA Permalink →

The classic "Cliff Walking" gridworld: an agent starts at one corner of a grid and must reach a goal in the opposite corner. A row of cells along the bottom edge is a cliff — stepping into it gives a reward of -100 and sends the agent back to start. Every other step costs -1. The shortest path runs directly along the cliff edge; a longer, safer path stays one row away from it.

You train two agents on this task with the same \varepsilon-greedy exploration schedule (\varepsilon = 0.1 throughout, never annealed): one with tabular Q-learning, one with tabular SARSA. Both eventually converge to stable value estimates.

During training itself — while \varepsilon-greedy exploration is still active and both agents are still occasionally taking random actions — which agent earns the higher average reward per episode?

Share this question

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.