Paths Subjects Questions Quizzes Pricing Search
Intermediate Open Pro

A Metric Regressed but the Prompt Diff Is Clean — Now What?

Your team's summarization prompt has been stable at v22 for two months, with no prompt changes deployed in that window. This week, the weekly golden-set score dropped from a steady ~93% to 86%, and the production guardrail dashboard shows average response length increased by 30% starting three days ago. Git history confirms no prompt file has changed in over two months. The model is called via a provider API using a version alias (not a pinned model snapshot).

  1. What is the most likely explanation, and what evidence would confirm or rule it out?
  2. Walk through exactly what you'd check in your logging/tracing setup to investigate this, and explain why each piece of logged data is necessary.
  3. Once confirmed, what are your two realistic options, and what does each cost you?

Share this question

← Back to Prompt Evaluation and Versioning practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.