Advanced
Open
Pro
Four Budgets, One Stuck Run: Diagnosing a Six-Hour Step
A harness implements the worked design's four budgets — 40 steps, a token cap, a 30-minute wall clock, a $2.00 cost cap — plus a 300-second per-command timeout. A run's trace ends like this:
step 22 (elapsed 24m, cost $0.90, tokens 38% of cap)
run_command("pytest tests/ -q") → auto-approved
[span opened 14:02:11 — never closed]
...
20:03 orchestrator job-level timeout (6h) kills the container
final status: unknown cost recorded: $0.90
branch: lost with the container
Investigation shows: one test opened a socket to an external host;
the sandbox's egress policy silently drops denied packets, so the
connect never failed; at 300 s the executor killed the pytest
parent process, but a worker child survived and kept the stdout
pipe open, and the executor's blocking read on that pipe never
returned.
- Diagnose which budgets were healthy, which control failed, and why. Explain why "three of four budgets were fine" is precisely the trap the subject warns about, and what the $0.90 figure hides.
- Redesign timeout and budget enforcement across the executor, the dispatcher, and the sandbox so that this run stops at minute 30 with its work preserved. State the exact status the run should return.
- The orchestrator's job timeout eventually ended the run. Argue why that is not a substitute for harness-level budgets, using the subject's layer test, and name the trace fields that would let you catch this class of failure in minutes rather than hours.
Share this question