Advanced
Open
Pro
Reading a First-Run Regression Report
You run all 40 golden sets, unchanged prompts, against the new model. Results: 19 pass with no meaningful change, 9 show format drift (correct content, wrong shape), 6 show tone/style drift, 4 show a genuine capability drop on a sub-task, and 2 show a refusal-behavior change.
- Explain why lumping all 21 regressions into one "fix the prompt" bucket is a mistake.
- For each of the four regression types, describe the fix you'd reach for and how confident you are it will fully close the gap.
- Which of the four types most needs a non-engineering sign-off before you touch it, and why?
Share this question