Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

One New Feature Took CV AUC From 0.81 to 0.97

You add merchant_id (50,000 unique values) to a LightGBM fraud model using mean target encoding: for each merchant, replace the ID with that merchant's historical fraud rate, computed once over the entire training set before doing anything else. You then evaluate with standard 5-fold cross-validation. CV AUC jumps from 0.81 to 0.97. In production, AUC drops back to roughly 0.82.

What actually happened?

Solution

Target leakage — each row's own label leaked into the encoded feature it's being scored on, and the CV split happened too late to catch it.

Mean target encoding computed over the full dataset bakes each row's own outcome into the very feature value used to predict that row: for a merchant with few transactions, a single fraudulent one can shift that merchant's whole encoded value toward "highly predictive" — partly of itself. Even for high-volume merchants, every row's contribution to the shared per-merchant average constitutes a small leak from label into feature, for that exact row.

The critical detail is when this happens relative to cross- validation: the encoding was computed on the full dataset before splitting into folds. So every fold's "held-out" rows already had their own labels mixed into the feature they're being evaluated on — 5-fold CV done afterward faithfully validates a model that has partially memorized labels via merchant_id, which is exactly why the CV number looks implausibly good and doesn't survive contact with genuinely new merchants and transactions in production. The leakage happened upstream of the CV split, so CV structurally could not catch it — it isn't a matter of using "more" cross-validation.

Fix: treat target encoding as any other label-dependent transform — fit it inside each CV fold, using only that fold's training rows, and apply it to the held-out rows without recomputing on them (or use a leakage-safe scheme like leave-one-out target encoding, or additive smoothing toward the global mean weighted by per-category count, computed strictly within each fold).

Share this question

← Back to LightGBM practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.