The Feature That's 'Too Good' Breaks the Model
You're fitting a logistic regression to predict loan default from 3
features, including an internal risk_flag. In your training sample,
every applicant with risk_flag = 1 defaulted, and every
applicant with risk_flag = 0 did not — a perfect split. You fit with
ordinary maximum likelihood, no regularization.
What happens to the coefficient on risk_flag?
It diverges toward infinity and never converges to a finite value.
This is complete separation: when some linear combination of the
features (here, just risk_flag on its own) perfectly separates the
two classes in the training data, the log-likelihood keeps increasing
as that coefficient's magnitude grows, with no finite maximum. Pushing
\beta \to \infty drives the predicted probability toward 1 for every
default and toward 0 for every non-default, pushing log-loss
arbitrarily close to its best possible value — but "arbitrarily
close" is achieved only in the limit, never at a finite, well-defined
stationary point. Standard MLE has no maximum to converge to.
In practice, the solver doesn't error — it just runs until an iteration limit or numerical overflow, spitting out some enormous, arbitrary coefficient with an equally enormous, meaningless standard error (any Wald test on it is nonsense). The genuinely counter- intuitive part: this model reports a fantastic fit on the training data it saw — perfect separation looks like perfect accuracy — while being fundamentally broken, not a good, well-calibrated classifier.
Two independent fixes address it, and it's worth knowing they attack
different problems: (1) any regularization (L2, or exact remedies
like Firth's penalized likelihood) adds a term that caps how far the
coefficient can grow, restoring a finite, well-behaved solution; (2)
separately, ask why one feature perfectly predicts the label — a
rule-derived flag like risk_flag that already encodes the outcome
it's meant to predict is a leakage/proxy problem, not a genuinely
predictive signal, and regularizing it into a "reasonable" number
doesn't fix that it shouldn't be a training feature at all.
Share this question