Advanced
Open
Pro
A/B Testing a Ranking Change
You've trained a new ranker that raises offline NDCG@30 by 4% versus the current production model. You want to A/B test it.
- Design the experiment: allocation, randomisation unit, and duration, with justification.
- List the primary metric and at least four guardrail metrics, and explain what each guardrail protects against.
- The two-week test shows the new ranker wins on watch time per DAU by 2% with no guardrail regressions. Explain why you might still not fully trust this result, and what you'd do next.
Share this question