Advanced
Open
Pro
Auditing a Harness Whose Answers Are All Correct
Your team runs an internal agent over customer-support data. Its offline eval suite reports 94% task success, its outputs pass a quality rubric, and no user has complained. Security asks a question the eval cannot answer: "over the last 10,000 runs, did this agent ever read data the requesting user was not entitled to, or carry information from one customer's context into another's?"
- Explain why the existing evaluation setup structurally cannot answer this, and what property of harness failures makes it so.
- Design what you would instrument and check to be able to answer it — for future runs and, as far as possible, for the 10,000 that already happened.
- The agent is about to be extended into a multi-agent design (a retriever agent and a drafting agent). Say what that changes about the risk, and what you would require before shipping it.
Share this question