Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Fine-Tuning: SFT, LoRA/QLoRA, RLHF and DPO (9 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Advanced Open Free

Picking the Right Technique for Three Different Requests Permalink →

Product brings you three separate requests in the same week:

  1. "Our JSON tool-call outputs are schema-valid only 85% of the time even with five few-shot examples in the prompt. We need 99%+."
  2. "We want the model to consistently prefer shorter, more direct answers over long hedging ones — we have pairs of past answers where our reviewers picked the better one, about 6,000 pairs."
  3. "We want to teach the model our company's refund policy so it stops asking customers to check the help centre."

For each, name the technique (SFT, LoRA/QLoRA, RLHF, DPO, or "none of these — use something else instead") you would recommend, and justify it in one or two sentences per case.

Share this question

Advanced Open Pro

Diagnosing a Fine-Tuned Model That Memorized the Wrong Things

Unlock this question →
Advanced Open Pro

Serving Per-Customer Fine-Tunes Without Per-Customer Deployments

Unlock this question →
Advanced Open Pro

Justifying RLHF vs. DPO to a Skeptical Staff Engineer

Unlock this question →
Advanced Open Pro

Designing the Evaluation Gate Before a Fine-Tune Ships

Unlock this question →
Advanced Open Free

What Raising LoRA Rank Actually Buys You Permalink →

A team is fine-tuning a LoRA adapter for a narrow classification-style task and finds validation loss has plateaued. Someone proposes raising the LoRA rank from 8 to 64 to fix it, reasoning "more rank means more capacity, so it should generalize better too." Holding every other hyperparameter fixed, what is the single most accurate description of what increasing LoRA rank primarily changes?

A. It increases the number of trainable parameters in the low-rank update matrices, raising the adapter's capacity to fit the training data — it does not, by itself, guarantee better generalization. B. It increases the effective rank of the frozen base-model weights, letting the base model itself adapt more to the new task. C. It reduces the KL divergence penalty against the reference model, allowing larger deviations from the base policy. D. It changes the number of transformer layers the adapter is applied to, spreading the update across more of the network.

Share this question

Advanced Open Pro

Why DPO Keeps a Frozen Reference Copy of the Policy

Unlock this question →
Advanced Open Pro

The Cheapest First Lever Against Catastrophic Forgetting

Unlock this question →
Advanced Open Pro

Naming the Failure Mode When Reward Climbs but Quality Drops

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.