Advanced
Open
Pro
The Migration Looks Done. What Could Still Go Wrong?
All 40 prompts are at 100% traffic on the new model. Canary metrics were clean throughout the ramp, and the team wants to tear down the migration dashboards and call it closed.
- Describe the judge-drift failure mode: how can an LLM-as-judge evaluation pipeline produce a misleading signal that has nothing to do with the prompts actually being migrated?
- Describe a second failure mode that can appear weeks after a clean 100% rollout, and why offline evals and canary guardrails both structurally miss it.
- What would you tell the team about closing out the migration, concretely — what stays running and for how long?
Share this question