Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — How LLMs Are Built: Pretraining to Chatbot (5 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Advanced Open Free

Why Dedup Instead of Just Training for Fewer Epochs? Permalink →

A teammate proposes skipping the deduplication step in a pretraining pipeline to save engineering time, arguing "we can just train for fewer epochs on the raw corpus instead — that reduces how many times the model sees any given document anyway."

  1. Explain why fewer epochs on an undeduplicated corpus is not an adequate substitute for deduplication. Name the two specific failure modes deduplication addresses that epoch count doesn't fix.
  2. Give one concrete scenario where skipping dedup would silently inflate a benchmark score without any real capability gain.
  3. What kind of technique does deduplication require at web scale, and why can't it just be exact string matching?

Share this question

Advanced Open Pro

Sizing a Model Under a Fixed Training-Compute Budget

Unlock this question →
Advanced Open Pro

Why Not Skip SFT and Go Straight From Base Model to DPO?

Unlock this question →
Advanced Open Pro

Does Mixture-of-Experts Actually Save You Memory?

Unlock this question →
Advanced Open Pro

A Data-Residency Constraint That Collapses the Decision Before Capability Even Matters

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.