Advanced
Open
Pro
Distilling a Reasoning Teacher into a 3B On-Device Triage Model
You want a small (~3B parameter) model that runs on-device to reason through simple medical-intake triage questions ("based on these symptoms, what urgency category applies") using a fixed, checkable set of triage rules. Two options: (a) generate a large R1-style teacher model's reasoning traces on your triage question distribution and fine-tune the 3B model on them as ordinary SFT data (R1-Distill-style); (b) run the full STaR + RLVR pipeline directly on the 3B model from scratch, using the same checkable triage rules as the verifiable reward.
- Compare the two options on training cost, and name the mechanism that makes one structurally cheaper.
- Name the specific capability gap distillation leaves that running full RLVR on the 3B model directly would not have, and explain why that gap exists mechanistically rather than just asserting "distillation is worse."
- Design a concrete evaluation to detect whether the distilled model has actually fallen into that gap for this triage use case, before shipping it.
Share this question