Intermediate
Open
Pro
A Metric Regressed but the Prompt Diff Is Clean — Now What?
Your team's summarization prompt has been stable at v22 for two
months, with no prompt changes deployed in that window. This week, the
weekly golden-set score dropped from a steady ~93% to 86%, and the
production guardrail dashboard shows average response length
increased by 30% starting three days ago. Git history confirms no
prompt file has changed in over two months. The model is called via a
provider API using a version alias (not a pinned model snapshot).
- What is the most likely explanation, and what evidence would confirm or rule it out?
- Walk through exactly what you'd check in your logging/tracing setup to investigate this, and explain why each piece of logged data is necessary.
- Once confirmed, what are your two realistic options, and what does each cost you?
Share this question