Advanced
Open
Pro
Sizing a Model Under a Fixed Training-Compute Budget
Your team has a fixed pretraining compute budget of roughly
C ≈ 3 × 10^23 FLOPs and is deciding between two allocations, using
the approximation C ≈ 6 × N × D (N = parameters, D = training
tokens):
- Option A: N = 40B parameters
- Option B: N = 8B parameters
- Solve for the training-token count D each option would use to spend the full budget C.
- Without doing any further training, explain which option is more likely to be undertrained relative to its size, and why that's a meaningful distinction from "which option has more capacity."
- Name one reason a team might deliberately choose the smaller, more-trained-per-parameter option even if a slightly larger model trained on fewer tokens scored marginally higher on a training-time loss metric.
Share this question