Advanced
Open
Pro
Improving a Harness With Evidence Instead of Intuition
Your coding agent passes 68% of an internal 300-task benchmark. The team's instinct is to rewrite the system prompt — that is where the last three improvement attempts went, each producing a change within noise. You have full traces for every benchmark run: tool calls, results, permission decisions, token counts, and outcomes.
- Propose a systematic improvement loop that uses the traces, and say what makes each change falsifiable rather than a guess.
- Where would you look first, and what does the available evidence suggest about prompt rewrites versus other harness components?
- What would convince you an improvement is real rather than benchmark-specific overfitting?
Share this question