Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Which Layer Owns It? Triaging a Post-Incident Action List

An internal agent that files and updates tickets caused an incident: it closed 340 tickets in one run, using a credential that also had delete permissions, and nobody noticed for two hours. The postmortem produced a mixed list of proposed fixes:

  1. "Give the agent a ticket-system credential that can comment and update but not delete."
  2. "Reword the system prompt to say the agent should never close more than a handful of tickets without asking."
  3. "Cap any tool result at 2,000 tokens."
  4. "Require a human confirmation before any bulk operation touching more than 10 records."
  5. "Route ticket triage through a fixed classification step first, and only send genuinely ambiguous tickets to the agent."
  6. "Emit a live per-run event stream so on-call sees an agent making hundreds of state changes while it is happening."

For each item: name which layer owns it (model/prompt, harness, or orchestration), and say whether it would actually have prevented or limited this incident. Then state which single fix you would ship first and why.

Solution

Layer assignment and effectiveness, item by item

  1. Scoped credential — harness (sandbox/credential axis). Would have limited the incident's worst case but not prevented it: the damage here was mass closing, which a comment-and- update credential still permits. Ship it anyway — it removes the far worse version of this incident (mass deletion) — but do not let it be counted as the fix for what actually happened. This is a good example of least privilege bounding blast radius without addressing the specific failure.
  2. Reword the system prompt — model/prompt layer. Would not reliably prevent it. A prompt instruction is a preference the model may or may not honor, and it is reachable by anything in the context — including ticket text the agent read. The whole point of the preference-vs-control distinction is that a rule you actually need to hold must live outside the model.
  3. Cap tool results — harness (executor result bounding). Correct hygiene, and it prevents a real and expensive failure mode, but it is unrelated to this incident: nothing here failed because a result was too large. Including it in the postmortem's top fixes would be cargo-culting a good practice into the wrong slot.
  4. Confirm before bulk operations >10 records — harness (permission/approval layer, gating on an effect rather than a tool name). This is the fix that directly prevents the incident: the agent would have stopped at ticket 11 and asked. Note it is expressed as an effect threshold, which means a new tool that also closes tickets inherits the gate instead of escaping it.
  5. Fixed classification step first — orchestration. It reduces how often the agent runs at all, which lowers exposure, but it does not bound what the agent does on the runs that still happen. Worth doing on cost and reliability grounds (this is the workflow-vs-agent argument), not as a safety control.
  6. Live per-run event stream — harness (observability). Would not have prevented the incident, but it is the reason "nobody noticed for two hours" was possible, and it is what converts a two-hour incident into a two-minute one next time. Detection is not prevention, but an undetectable failure class is strictly worse than a detectable one.

What to ship first

Item 4, the effect-based bulk-operation gate, because it is the only proposal that would have stopped this incident at the point of action, and it does so with a control outside the model's judgment. Item 6 (observability) is the immediate second, because it closes the detection gap and because without it you cannot verify that item 4 is actually firing. Items 1 and 3 are good harness hygiene to schedule; item 5 is a cost/reliability improvement; item 2 is acceptable as a belt-and-braces nudge but must never be presented as the control — an interviewer listening for the preference-vs-control distinction is specifically testing whether you rank it last.

Share this question

← Back to Agent Harness Engineering practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.