Advanced
Open
Pro
Migrate a Tabular Agent to Function Approximation
An ad-pricing team has a tabular Q-learning agent working well in a simulator with 200 discrete states (10 price buckets × 20 time-of-day/inventory buckets) and 5 discrete price actions. They now want to add three new continuous features — competitor price, remaining daily budget, and a rolling conversion-rate estimate — and expect the effective state space to explode.
- Quantitatively, why does naive discretization of the new features break the tabular approach? Give a rough number.
- Propose the function-approximation architecture you would use instead, and explain what changes about the update rule (write the semi-gradient update).
- A team member says "let's just keep the table but make it sparse (a hash map from visited (state, action) to value) — problem solved." What does this fix, and what does it not fix?
Share this question