Advanced
Open
Pro
Reflexion Retries or Parallel Voting: Spending Three Attempts on a Failing Test Suite
A coding agent (tools: read_file, run_tests, edit_file) fixes
failing CI tests; a single attempt succeeds 55% of the time, and
success is unambiguous because the suite either passes or not. The
budget is about three attempts' worth of compute per task. Two
proposals:
- V: run 3 independent attempts in parallel and accept any one that passes (voting, with the test suite as the scorer).
- R: run attempts sequentially with Reflexion — after each failure, generate a verbal reflection, store it as episodic memory, and include it in the next attempt; stop at the first pass or after 3 attempts.
A pilot measured: under V, P(at least one of 3 passes) = 0.70. Under R, a second attempt passes 60% of the time and a third 50% of the time, conditional on reaching it.
- Compute each proposal's success rate within budget, explain why V lands so far below the independence prediction, and explain — using what the lesson says voting and Reflexion each buy — why R can beat it here.
- Compare expected cost and latency for V and R, and name the conditions under which V would be the right choice after all.
- In some R runs the reflection blames the wrong cause ("the fixture is missing" when the edit had a syntax error) and the next attempt does worse than a fresh one would. Explain why the mechanism makes this possible, why the unambiguous pass/fail signal matters, and propose a mitigation plus a hybrid of V and R that hedges against it.
Share this question