Advanced
Open
Pro
The Limits of a Single Classifier
Your team ships an injection classifier that scores 97% precision and 94% recall on your internal red-team test set, and leadership wants to announce "our assistant is protected against prompt injection."
- Explain, concretely, why this announcement overstates what a 94% recall classifier evaluated on a fixed test set can promise.
- Design what you would actually tell leadership to say, and what ongoing process (not a one-time number) would back it up.
- Given that no classifier eliminates the risk, what property of the system should leadership be asked to invest in instead, and why does it hold even when the classifier fails?
Share this question