Practice — Data Engineering in Production (5 questions)
Advanced
Open
Free
Diagnosing and Preventing a Runaway Warehouse Cost Spike Permalink →
Finance flags that your team's Snowflake bill for last month is 4x the
usual amount. You pull WAREHOUSE_METERING_HISTORY and find that a
single warehouse, ANALYTICS_XL, was active nearly continuously for
11 days straight, even overnight and on weekends, at an X-LARGE size.
Nobody remembers scheduling anything unusual. Digging further, you find
an engineer ran a one-off backfill on that warehouse 11 days ago,
manually resized it to X-LARGE "to make it finish faster," and the
warehouse's AUTO_SUSPEND was set to 3600 (one hour) months earlier
"because cold starts were annoying during a demo."
- Explain exactly how these two settings combined to produce an 11-day near-continuous billing window, not just "auto-suspend was too long."
- Propose the concrete configuration and process fixes — not just "be more careful" — that prevent this specific failure mode from recurring, for this warehouse and others like it.
- Design a detection mechanism that would have caught this within hours instead of a full billing cycle later, and explain what signal it watches for.
Share this question