Intermediate
Open
Pro
Pushing Back on 'Just Use the Bigger Model'
Your model is underperforming on a domain-specific extraction task (pulling structured fields out of contracts). Leadership's proposed fix: "swap the 8B model we're using for the 70B version of the same model family — bigger models are better, this should just work."
- Is parameter count alone a good predictor of whether this fixes the problem? Explain what it does and doesn't tell you.
- What two or three questions would you actually want answered before recommending for or against the swap?
- Give a concrete alternative explanation for the underperformance that a bigger model wouldn't fix, and how you'd distinguish it from "the model is just too small for this task."
Share this question