Advanced
Open
Pro
Tree Search Over Actions When Some Actions Have Side Effects
A customer-remediation agent has read-only tools — get_order(id),
get_policy(topic), search_tickets(q) — and side-effecting tools
— issue_refund(order, amount), send_email(to, body),
create_ticket(...). The ReAct version has a known failure: it
commits at step 2 to a remediation type (refund / replacement /
escalate) and then gathers only evidence that fits the choice it
already made; 18% of cases end in the wrong remediation. The team
proposes a LATS-style search: branching factor 3, depth 5, every
candidate action scored by a self-evaluation call, expanding only
the best-scoring branch at each level.
- Using the lesson's cost accounting, estimate the LLM calls per case for the proposed search versus the ReAct version's ≈6, and say what would have to be true for that premium to be justified.
- Give a concrete trace in which branch A calls
issue_refund, branch B callsget_policy, and the evaluator prefers B. State what has happened in the world, and connect it to the lesson's caveat about forkable, comparable state. - Redesign the search so it is applied only where it earns its keep: which decision is search-shaped, which actions may appear as candidates inside a branch, how side-effecting actions are handled, and what measurement would justify — or kill — the whole idea given the 18% figure.
Share this question