Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

The Feature That's 'Too Good' Breaks the Model

You're fitting a logistic regression to predict loan default from 3 features, including an internal risk_flag. In your training sample, every applicant with risk_flag = 1 defaulted, and every applicant with risk_flag = 0 did not — a perfect split. You fit with ordinary maximum likelihood, no regularization.

What happens to the coefficient on risk_flag?

Solution

It diverges toward infinity and never converges to a finite value.

This is complete separation: when some linear combination of the features (here, just risk_flag on its own) perfectly separates the two classes in the training data, the log-likelihood keeps increasing as that coefficient's magnitude grows, with no finite maximum. Pushing \beta \to \infty drives the predicted probability toward 1 for every default and toward 0 for every non-default, pushing log-loss arbitrarily close to its best possible value — but "arbitrarily close" is achieved only in the limit, never at a finite, well-defined stationary point. Standard MLE has no maximum to converge to.

In practice, the solver doesn't error — it just runs until an iteration limit or numerical overflow, spitting out some enormous, arbitrary coefficient with an equally enormous, meaningless standard error (any Wald test on it is nonsense). The genuinely counter- intuitive part: this model reports a fantastic fit on the training data it saw — perfect separation looks like perfect accuracy — while being fundamentally broken, not a good, well-calibrated classifier.

Two independent fixes address it, and it's worth knowing they attack different problems: (1) any regularization (L2, or exact remedies like Firth's penalized likelihood) adds a term that caps how far the coefficient can grow, restoring a finite, well-behaved solution; (2) separately, ask why one feature perfectly predicts the label — a rule-derived flag like risk_flag that already encodes the outcome it's meant to predict is a leakage/proxy problem, not a genuinely predictive signal, and regularizing it into a "reasonable" number doesn't fix that it shouldn't be a training feature at all.

Share this question

← Back to Logistic Regression & Linear Classifiers practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.