Agent Harness Engineering
The runtime shell around the model: tool executors, sandboxes, permission tiers, context plumbing, and the observability that makes an agent auditable
A vendor-neutral treatment of the harness — everything in an agent that is not the model. Covers the Agent = Model + Harness decomposition and why the harness, not the loop, is the security and reliability boundary; the five components (executor, sandbox, permission and approval tiering, context/memory plumbing, observability); a traced walkthrough of one tool call through every layer; a worked coding-agent harness design with real permission tiers and budgets; the discipline boundary between harness engineering, prompt engineering, and orchestration; harness-specific failure modes (permission gaps, validation/execution mismatch, trust-boundary carryover, approval fatigue, unobserved cost blowup, sandbox escape) with mitigations and arithmetic; and how harnesses are audited and evolved from trajectory data.
Practice questions (13)
-
View →
Which Layer Owns It? Triaging a Post-Incident Action List
Advanced · Free -
View →
Designing Permission Tiers for a Database Migration Agent
Advanced · Free -
View →
Costing an Uncapped Tool Result
Advanced -
View →
Fixing an Approval Gate Nobody Reads
Advanced -
View →
A Gate That Checked the Wrong Thing
Advanced