Advanced
Open
Pro
Debug a Collapsing Shared-Trunk Actor-Critic
A team implements A2C for a resource-allocation agent (continuous allocation fractions across services) using a shared trunk with two heads: a policy head (Gaussian mean and log-std) and a value head. Early in training, the critic's loss is very large (the value head was initialized far from the true returns), and shortly after, the team observes the policy's output standard deviation collapsing toward zero and the actor essentially always outputting the same allocation regardless of state.
- Explain the likely causal chain from the critic's large early loss to the policy's collapse, given the shared-trunk architecture.
- Propose two independent fixes, one architectural and one related to the loss function, and explain the mechanism of each.
- The team considers just lowering the overall learning rate as a third fix. Would this address the root cause? What would it cost?
Share this question