Advanced
Open
Pro
Design a Policy Head for a Continuous Pricing Action
A dynamic-pricing team wants to switch from a discrete price-tier DQN agent (12 discrete price points) to an A2C agent that outputs a continuous price multiplier in the range [0.5, 1.5] relative to a base price.
- Describe the policy head you would design (distribution family, outputs, and how you would enforce the [0.5, 1.5] bound), and contrast it with how the old DQN agent selected an action.
- Write the log-probability term \nabla_\theta \log \pi_\theta(a_t\mid s_t) you would need for the actor update, for your chosen distribution (a closed form for the Gaussian case is sufficient).
- A team member worries that "continuous actions mean we lose the DQN agent's Double-DQN-style protection against overestimation bias — doesn't that make the new agent worse?" Evaluate this concern.
Share this question