Advanced
Open
Pro
Designing the Evaluation Gate Before a Fine-Tune Ships
Your team just finished training a LoRA adapter (SFT) intended to improve tool-call argument formatting for an internal coding assistant. The target metric — schema-valid tool calls on a held-out set — went from 91% (prompted baseline) to 99.2% (fine-tuned). An engineering lead says "great, target metric passed, let's ship it today." Design the evaluation gate you would insist on before agreeing, and explain what could still be wrong despite the strong target-metric result.
Share this question