Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Migrate a Tabular Agent to Function Approximation

An ad-pricing team has a tabular Q-learning agent working well in a simulator with 200 discrete states (10 price buckets × 20 time-of-day/inventory buckets) and 5 discrete price actions. They now want to add three new continuous features — competitor price, remaining daily budget, and a rolling conversion-rate estimate — and expect the effective state space to explode.

  1. Quantitatively, why does naive discretization of the new features break the tabular approach? Give a rough number.
  2. Propose the function-approximation architecture you would use instead, and explain what changes about the update rule (write the semi-gradient update).
  3. A team member says "let's just keep the table but make it sparse (a hash map from visited (state, action) to value) — problem solved." What does this fix, and what does it not fix?

Share this question

← Back to Value-Based Methods: Q-Learning to DQN practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.