Design the Release Plan for a High-Stakes Fraud Model
Your payments platform processes 50M transactions/day. The current fraud model declines 1.1% of transactions. A challenger trained on a new feature set shows a 6-point offline recall improvement. Fraud losses run into millions of dollars per percentage point of missed recall, and a false-positive decline directly angers a legitimate customer and can cause churn.
- Design the full release sequence from "we have a promising challenger" to "fully promoted," naming each stage, its duration, and what you are trying to learn or de-risk at each one.
- At the canary stage, decline rate on the challenger's traffic slice comes in at 1.35% against a pre-agreed guardrail band of [0.9%, 1.3%]. Is this an automatic-rollback condition or a human-gated one? Justify your answer and describe what the human (or automation) should check before deciding.
- Explain specifically why shadow mode alone would have been insufficient to validate this release, even if it showed excellent agreement with the champion.
1. Release sequence
- Shadow, several days. The challenger scores every live transaction with zero effect on decisions. Goal: catch infra bugs (timeouts, crashes under real load), gross miscalibration (score distribution wildly different from the champion's), and confirm latency is acceptable — all for free, before any user is exposed.
- Small random canary (1-2%), held long enough to span a full weekly cycle (fraud and legitimate-transaction patterns both vary by day of week — a canary that only sees weekdays is a biased sample). Goal: get a real, outcome-relevant read on decline rate and downstream signals with bounded exposure — at 50M transactions/day even 1-2% is 500K-1M real transactions/day, enough to accumulate statistical signal quickly despite fraud being a rare-event metric.
- Ramp: 5% → 15% → 40% → 100%, each held 2-3 days, with guardrail checks (decline rate band, manual-review queue capacity, latency) at every step, and a human checkpoint given the high-stakes / high-volume risk profile.
- Full ramp, old champion kept warm (blue/green-style) for an additional period as the instant-rollback target, then archived per the lifecycle's versioning/retention discipline.
2. The 1.35% breach
This is a human-gated decision, not automatic. It is a real breach of the pre-agreed band, but it is a borderline, ambiguous signal on a metric that can plausibly move in a good direction — a higher decline rate could mean the challenger is correctly catching more real fraud (exactly what the 6-point recall improvement predicted) or it could mean it is false-positiving on legitimate transactions. An automatic rollback treats every breach the same regardless of direction-of-cause, which is wrong here; the situation needs interpretation. Before deciding, the reviewer should check: is the increase concentrated in a specific segment the new features specifically targeted (consistent with intended improvement) or spread broadly (more consistent with a general miscalibration); has any downstream signal — early manual-review outcomes, customer complaint rate — started to differentiate whether the extra declines are correct; and is the manual-review queue capacity being exceeded by the extra volume (an operational guardrail independent of whether the decisions themselves are correct).
3. Why shadow mode alone is insufficient
Shadow mode's decisions never take effect, so it can only measure agreement between the challenger and champion, not which one is right when they disagree. If the challenger would have declined a transaction the champion approved, shadow mode logs that disagreement but the transaction is still approved either way — you never learn whether the challenger's decline would have been correct (caught real fraud) or a false positive (annoyed a legitimate customer), because that transaction's true fraud/not-fraud outcome is never tested against the challenger's actual decision. Only a canary, where the challenger's decisions on its traffic slice are real, can generate the outcome-based evidence (recall, false-positive rate on live traffic) that actually validates whether the 6-point offline improvement holds up in production.
Share this question