Paths Subjects Questions Quizzes Pricing Search
Advanced Open Pro

Serving Per-Customer Fine-Tunes Without Per-Customer Deployments

Your product lets enterprise customers customize the assistant's writing style to match their brand voice. Sales wants this available to any customer on the enterprise tier — today that's 40 customers, growing toward several hundred within a year. An engineer proposes full-parameter fine-tuning a separate model per customer.

  1. Explain what's wrong with that proposal as it scales, in terms of both training and serving cost.
  2. Propose the alternative architecture, and describe concretely how a single request gets served under it.
  3. What data would each customer's fine-tune need, and what's the evaluation gate before enabling a new customer's adapter in production?

Share this question

← Back to Fine-Tuning: SFT, LoRA/QLoRA, RLHF and DPO practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.