Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Why More Trees Don't Kill a Random Forest's Variance

A colleague argues: "our random forest is still overfitting a bit — let's just keep adding trees until the variance disappears." You have B bootstrap trees, each individually with prediction variance \sigma^2, and pairwise correlation \rho between any two trees' predictions (nonzero because every tree is grown from overlapping bootstrap samples of the same training set and can pick similar splits).

The variance of the forest's averaged prediction is:

\text{Var}(\bar{f}) = \rho\sigma^2 + \frac{(1-\rho)\sigma^2}{B}

As you take B \to \infty (grow arbitrarily many trees), what happens to the forest's variance?

Solution

It approaches a floor of \rho\sigma^2, set by inter-tree correlation — not zero.

As B \to \infty, the second term (1-\rho)\sigma^2/B \to 0, but the first term \rho\sigma^2 does not depend on B at all — it survives no matter how many trees you add. Averaging only cancels out the part of each tree's error that is independent across trees; the part that all trees share, because they were all grown from the same underlying data and tend to make similar mistakes on the same hard examples, cannot be averaged away by adding more copies of a correlated estimator.

Concretely: with \rho = 0.3 and \sigma^2 = 1, going from B=10 to B=1000 shrinks variance from about 0.3 + 0.07 = 0.37 to 0.3 + 0.0007 \approx 0.30 — almost all of the achievable reduction happens in the first few dozen trees, and the remaining 0.30 is permanent. Adding the 10,000th tree buys essentially nothing.

This is exactly why random forests bother with feature subsampling at each split (choosing the best split from only a random subset of features, not all of them) on top of bootstrap sampling of rows: it's specifically a mechanism to lower \rho, the thing more trees alone cannot fix. Bagging alone (bootstrap rows, full feature set at every split) tends to grow highly correlated trees, especially when one or two features are strongly dominant — every tree keeps splitting on the same feature first, making them similar to each other regardless of B. Random forest's feature-subsampling step decorrelates the trees, which lowers the \rho\sigma^2 floor itself — the actual lever for further variance reduction, once adding more trees has stopped helping.

Share this question

← Back to Bias–Variance Trade-off & Cross-Validation practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.