Advanced
Open
Pro
Choose an n for a Delayed-Reward Trading Agent
An RL agent makes a sequence of trading decisions over a 100-step episode. Reward is zero at every step except the last, where it receives the realized profit or loss for the whole episode. You are deciding between n=1 (TD(0)), n=10, and n=100 (full Monte Carlo) n-step returns for the value-function update, with discount \gamma = 0.99.
- For each of the three choices, describe qualitatively what the bias and variance of the return estimate look like early in training, when the value function V is still a poor estimate.
- Given \gamma=0.99, compute \gamma^{10} and \gamma^{100} and use them to explain quantitatively how much each choice of n discounts the bootstrapped tail relative to the real observed rewards it includes.
- Recommend a choice (or a scheme) and justify it in terms of the credit-assignment problem specific to this "reward only at the very last step" structure.
Share this question