Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Improving a Harness With Evidence Instead of Intuition

Your coding agent passes 68% of an internal 300-task benchmark. The team's instinct is to rewrite the system prompt — that is where the last three improvement attempts went, each producing a change within noise. You have full traces for every benchmark run: tool calls, results, permission decisions, token counts, and outcomes.

  1. Propose a systematic improvement loop that uses the traces, and say what makes each change falsifiable rather than a guess.
  2. Where would you look first, and what does the available evidence suggest about prompt rewrites versus other harness components?
  3. What would convince you an improvement is real rather than benchmark-specific overfitting?

Share this question

← Back to Agent Harness Engineering practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.