Practice — Fine-Tuning: SFT, LoRA/QLoRA, RLHF and DPO (9 questions)
Picking the Right Technique for Three Different Requests Permalink →
Product brings you three separate requests in the same week:
- "Our JSON tool-call outputs are schema-valid only 85% of the time even with five few-shot examples in the prompt. We need 99%+."
- "We want the model to consistently prefer shorter, more direct answers over long hedging ones — we have pairs of past answers where our reviewers picked the better one, about 6,000 pairs."
- "We want to teach the model our company's refund policy so it stops asking customers to check the help centre."
For each, name the technique (SFT, LoRA/QLoRA, RLHF, DPO, or "none of these — use something else instead") you would recommend, and justify it in one or two sentences per case.
Share this question
Diagnosing a Fine-Tuned Model That Memorized the Wrong Things
Unlock this question →Serving Per-Customer Fine-Tunes Without Per-Customer Deployments
Unlock this question →What Raising LoRA Rank Actually Buys You Permalink →
A team is fine-tuning a LoRA adapter for a narrow classification-style task and finds validation loss has plateaued. Someone proposes raising the LoRA rank from 8 to 64 to fix it, reasoning "more rank means more capacity, so it should generalize better too." Holding every other hyperparameter fixed, what is the single most accurate description of what increasing LoRA rank primarily changes?
A. It increases the number of trainable parameters in the low-rank update matrices, raising the adapter's capacity to fit the training data — it does not, by itself, guarantee better generalization. B. It increases the effective rank of the frozen base-model weights, letting the base model itself adapt more to the new task. C. It reduces the KL divergence penalty against the reference model, allowing larger deviations from the base policy. D. It changes the number of transformer layers the adapter is applied to, spreading the update across more of the network.
Share this question