Does Standardizing 'income' Change a Gradient-Boosted Tree's Predictions?
A data scientist standardizes every numeric feature (z-score: subtract the mean, divide by the standard deviation) before fitting two models on the same data: a logistic regression and a gradient-boosted tree ensemble (LightGBM). Their reasoning: "scaling is always good practice, it can only help."
For the logistic regression, this is true — unscaled features with very different ranges distort the effective regularization strength per coefficient and can slow optimizer convergence. For the gradient-boosted tree model specifically, what effect does standardizing a single numeric feature (a strictly monotonic transform: same relative order, just a different scale) have on the trained model's actual predictions?
No effect — trees are invariant to strictly monotonic transforms.
A decision tree split is a threshold test: "is income <= x?" Every
split the tree-growing algorithm considers is chosen by scanning the
sorted order of a feature's values and picking the split point that
best separates the target — it never uses the feature's absolute
scale, only which examples fall above versus below a candidate
threshold. Standardizing income (subtracting its mean, dividing by
its standard deviation) is a strictly monotonic transform: if
x_i < x_j before scaling, the same ordering holds after scaling.
Every split the tree would have chosen on the raw values has an
exact counterpart on the scaled values that produces the identical
partition of the training data — the threshold's number changes
(say, from "income ≤ 52,000" to "z ≤ 0.31"), but which rows land on
which side of every split, and therefore the entire tree structure
and every prediction, is unchanged.
This generalizes: any strictly monotonic transform of a single feature — scaling, mean-centering, \log, square root, or a rank transform — leaves a tree-based model's predictions bit-for-bit identical, because monotonic transforms preserve order and trees only ever depend on order. (It does not generalize to non-monotonic transforms — binning a continuous feature into a small number of buckets, for instance, throws away information a tree could otherwise split on, and can change predictions.)
The practical takeaway: scaling numeric features before a tree-based model is not wrong, but it is also not doing anything — any perceived improvement from scaling a tree model's inputs came from somewhere else (a code change, a different random seed, a genuinely non-monotonic transform smuggled in alongside the "scaling"). Skipping it for tree models is a free simplification, not a missed optimization.
Share this question