Why Dedup Instead of Just Training for Fewer Epochs?
A teammate proposes skipping the deduplication step in a pretraining pipeline to save engineering time, arguing "we can just train for fewer epochs on the raw corpus instead — that reduces how many times the model sees any given document anyway."
- Explain why fewer epochs on an undeduplicated corpus is not an adequate substitute for deduplication. Name the two specific failure modes deduplication addresses that epoch count doesn't fix.
- Give one concrete scenario where skipping dedup would silently inflate a benchmark score without any real capability gain.
- What kind of technique does deduplication require at web scale, and why can't it just be exact string matching?
1. Why fewer epochs doesn't substitute for dedup
Reducing epoch count controls how many times the model sees the whole corpus, but it does nothing about the fact that near-duplicate content is overrepresented within a single pass through an undeduplicated corpus relative to how much genuinely distinct information it carries — a document mirrored across a dozen sites still appears a dozen times in one epoch, consuming gradient signal and limited model capacity disproportionately to its actual information content, even at one epoch. The two specific failure modes: (a) memorization risk — repeated exact or near-exact exposure to the same sequence, even within one epoch if it's duplicated many times across the corpus, measurably increases the chance the model memorizes and can regurgitate it verbatim; and (b) benchmark contamination — if a benchmark question or a close paraphrase leaked into the crawl, a single exposure is already enough to inflate the model's score on that benchmark, and epoch count is irrelevant to whether that one exposure happened at all.
2. A concrete contamination scenario
Say a popular blog post that discusses and answers a well-known reasoning benchmark's example questions (a common occurrence for widely-cited benchmarks) got crawled along with several mirrors and reposts of the same post. Without decontamination checking against known benchmark content, those benchmark answers effectively enter the training corpus, duplicated across the mirrors. The model's subsequent high score on that benchmark then reflects partial memorization of leaked answer text, not a genuine improvement in reasoning capability — and reducing epochs to one wouldn't have prevented the one exposure that mattered.
3. Why approximate techniques, not exact matching
At trillions of tokens across a web-scale crawl, near-duplicates vastly outnumber byte-for-byte exact duplicates — the same article with a different ad banner, a different site's boilerplate wrapper, or minor reformatting is a different string but semantically the same content, and exact string matching would miss all of it while also being computationally infeasible to run as an all-pairs comparison across a corpus this large. Dedup pipelines instead use approximate, distributed techniques — MinHash / locality-sensitive hashing to find documents that are highly similar (not necessarily identical) cheaply, without an all-pairs comparison — which is why deduplication at this scale is its own systems-engineering effort rather than a one-line script.
Share this question