Practice — Feature Engineering (7 questions)
Spotting the Leak in a Target Encoding Pipeline Permalink →
A colleague encodes a merchant_id column (40,000 distinct values) for a
fraud model like this, run on the full training set before the
train/test split:
merchant_rate = df.groupby("merchant_id")["is_fraud"].transform("mean")
df["merchant_te"] = merchant_rate
Cross-validated AUC on this feature alone is 0.97. In production, AUC drops to 0.71.
- Explain exactly why this code leaks the label, using a merchant with 3 transactions as a concrete example.
- Rewrite the encoding so it is safe to use inside 5-fold cross-validation.
- Why does smoothing (shrinking toward the global rate) matter here even after you fix the leak?
Share this question
Missing Values: Indicator, Impute, or Let the Model Handle It
Unlock this question →Does Standardizing 'income' Change a Gradient-Boosted Tree's Predictions? Permalink →
A data scientist standardizes every numeric feature (z-score: subtract the mean, divide by the standard deviation) before fitting two models on the same data: a logistic regression and a gradient-boosted tree ensemble (LightGBM). Their reasoning: "scaling is always good practice, it can only help."
For the logistic regression, this is true — unscaled features with very different ranges distort the effective regularization strength per coefficient and can slow optimizer convergence. For the gradient-boosted tree model specifically, what effect does standardizing a single numeric feature (a strictly monotonic transform: same relative order, just a different scale) have on the trained model's actual predictions?
Share this question