Practice — How LLMs Are Built: Pretraining to Chatbot (5 questions)
Advanced
Open
Free
Why Dedup Instead of Just Training for Fewer Epochs? Permalink →
A teammate proposes skipping the deduplication step in a pretraining pipeline to save engineering time, arguing "we can just train for fewer epochs on the raw corpus instead — that reduces how many times the model sees any given document anyway."
- Explain why fewer epochs on an undeduplicated corpus is not an adequate substitute for deduplication. Name the two specific failure modes deduplication addresses that epoch count doesn't fix.
- Give one concrete scenario where skipping dedup would silently inflate a benchmark score without any real capability gain.
- What kind of technique does deduplication require at web scale, and why can't it just be exact string matching?
Share this question
Advanced
Open
Pro