Intermediate
Open
Pro
Diagnosing and Fixing Overfitting in a Decision Tree
You train a DecisionTreeClassifier on a dataset with 10,000 samples and
50 features. The results are:
- Train accuracy: 99.8%
- Validation accuracy: 74.1%
- What is happening and why?
- Which hyperparameters would you tune, in what direction, and why does each one help?
- Is there anything structurally wrong with using a single decision tree for this problem?
Share this question