Intermediate
Open
Pro
Where Does General Observability End and ML Monitoring Begin?
Your team just launched a fraud-scoring model behind a new serving endpoint. A teammate proposes the following monitoring plan: "Let's just watch p99 latency and error rate — that's the standard stuff — and separately have the data science team compute PSI on all 40 input features daily and email the team if anything looks drifted."
- What general-observability elements (from this subject) are missing from this plan, independent of anything ML-specific?
- Explain why "watch prediction-distribution health and feature freshness on the same dashboard, with the same alerting mechanics, as RED metrics" is a better integration than "PSI in a separate daily email," even though PSI itself is out of scope for this subject.
- Where exactly is the boundary — what would a good interviewer want to hear you say is covered here versus in ML Monitoring, Drift & Retraining, if asked to design monitoring for this endpoint?
Share this question