Practice — Bias–Variance Trade-off & Cross-Validation (6 questions)
Diagnosing a Model From Learning and Validation Curves Permalink →
You train a gradient boosting classifier on 50,000 rows. Training AUC
is 0.995, validation AUC is 0.81. You plot a learning curve (error vs
training-set size) and both the training and validation curves are
still converging toward each other as you add more rows, with a
visible gap remaining at the full dataset size. You also plot a
validation curve sweeping max_depth from 2 to 12, and validation AUC
peaks at max_depth=4 then declines.
- What do the two curves individually tell you, and what is the combined diagnosis?
- Rank the following interventions by expected usefulness here, and
justify the ranking: (a) collect more training rows, (b) reduce
max_depthto 4, (c) add L2 regularisation on leaf weights, (d) engineer better features. - After applying your top interventions, validation AUC rises to 0.87 but training AUC drops to 0.90. Is this an improvement? Explain.
Share this question
Why More Trees Don't Kill a Random Forest's Variance Permalink →
A colleague argues: "our random forest is still overfitting a bit — let's just keep adding trees until the variance disappears." You have B bootstrap trees, each individually with prediction variance \sigma^2, and pairwise correlation \rho between any two trees' predictions (nonzero because every tree is grown from overlapping bootstrap samples of the same training set and can pick similar splits).
The variance of the forest's averaged prediction is:
As you take B \to \infty (grow arbitrarily many trees), what happens to the forest's variance?
Share this question