Practice — Data Quality, Testing & Observability (4 questions)
Diagnosing a Dashboard That Broke Days After the Actual Cause Permalink →
Your company's executive monthly_recurring_revenue dashboard, built
on top of a dbt model, is fed by a daily pipeline that pulls
subscription data from a third-party billing provider's API. On
Tuesday, the billing provider ships an unannounced API change: a
plan_price field that used to be an integer in cents (4900 for
$49.00) becomes a decimal in dollars (49.00). The field's name and
type are unchanged, so it still passes every existing schema check
and NOT NULL constraint. Your transformation logic, which divides
the raw value by 100 to convert cents to dollars, keeps running
without error — it just now produces numbers that are 100x too small.
Nobody notices until Friday, when a sales leader flags in Slack that
"the MRR chart looks broken."
- Explain precisely why none of your existing checks (schema
validation,
NOT NULL,unique) caught this, tying your answer to the specific category of schema change involved. - Design the monitoring that would have caught this within a day instead of three, and explain mechanistically why each check you propose would have fired.
- Walk through your incident response once the sales leader's message comes in: what do you do first, second, third, and why does the order matter?
- Propose a longer-term fix to the relationship with the billing provider itself, and explain why a purely technical fix (better monitoring) isn't sufficient on its own.
Share this question