Replacing Linters With a Reviewer Model: Costing a Sensor Stack
A coding-agent harness currently runs three sensors automatically
after every edit_file: ruff, the type checker, and the test
suite — roughly zero tokens and about 2 seconds combined per edit.
A proposal replaces all three with a single reviewer model step
that reads each diff and returns pass/fail with feedback, on the
grounds that "it catches everything the linters catch plus design
problems."
Measured facts: the reviewer costs about 6,000 input tokens (diff, surrounding file context, rubric) and 400 output tokens per edit at $3 per million input and $15 per million output, with ~8 s latency. A typical run makes 18 edits and operates under the worked design's $2.00 cost cap. In a pilot, the reviewer agreed with the linter on 91% of lint findings and returned a false "fail" on 7% of clean diffs.
- Classify each control involved as a guide or a sensor, and each sensor as computational or inferential. Compute the per-run cost and added latency of the proposal and compare it with the current stack.
- Explain why "catches everything the linters catch" is the wrong basis for the comparison, using the subject's design bias and the pilot numbers — including a failure the reviewer has that no linter can have.
- Propose the sensor stack you would actually ship, with the reviewer's placement and expected cost, and say how you would make the change falsifiable in the sense of the subject's harness-evolution evidence.
Share this question