Advanced
Open
Pro
Quantify and Fix Overestimation Bias in a Trading Agent
A DQN-based trading agent chooses among four discrete actions at
each state: buy, sell, hold, hedge. At a particular next
state s', the true optimal value is the same for all four actions,
Q^*(s', a) = 10.0, but the trained network's noisy estimates are:
\hat{Q}(s', \cdot) = [10.4,\ 9.6,\ 11.1,\ 9.8]
- Compute the overestimation bias that vanilla DQN's target would introduce at this state, relative to the true value.
- A second, independently-noisy target network gives estimates \hat{Q}_{\theta^-}(s', \cdot) = [9.7,\ 10.3,\ 9.5,\ 10.6] for the same four actions. Walk through the Double DQN target computation and compare its bias to vanilla DQN's.
- Explain why this bias is not fixed by simply training the network longer or with more data, and why it specifically grows with the number of actions.
Share this question