Practice — Model Training & Experimentation at Scale (7 questions)
Intermediate
Open
Free
Baselines Before the Deep Model Permalink →
A team is building a "recommended for you" module for a mid-sized e-commerce site (2 M monthly users, 300 k products). Their first plan is to train a two-tower neural network with user- and item-history encoders and ship it if offline recall@50 exceeds 0.30.
- Which baselines should they build first, and what does each one tell them?
- The two-tower model reaches recall@50 = 0.31. Popularity ranking reaches 0.27. How would you use these numbers in a go/no-go discussion?
- Name one operational reason, unrelated to accuracy, why the baseline should still be maintained after the neural model ships.
Share this question
Intermediate
Open
Free
GPU Memory Needed to Train an LLM Permalink →
GPT-2 XL has 1.5 B parameters. Its weights in FP16 take ~3 GB on disk.
Roughly how much GPU memory is needed to train it (full fine-tuning, standard mixed-precision Adam) on a single GPU?
Show the per-parameter memory breakdown that gets you to your answer, and name two techniques that would bring the requirement down.
Share this question