Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Does Mixture-of-Experts Actually Save You Memory?

A team is comparing two model options for a self-hosted deployment: a dense model with 30B total (and active) parameters, and an MoE model with 120B total parameters but only 15B active parameters per token (an 8-expert-total, 1-expert-active-per-token routing setup, roughly). A team member argues "the MoE model only activates 15B parameters per token, so it needs less than half the GPU memory of the dense 30B model."

  1. Is that memory claim correct? Explain what actually determines a model's memory footprint versus what determines its per-token compute cost.
  2. Given that, which of the two options is actually cheaper to host in terms of GPU memory, and which is cheaper in terms of per-token inference compute?
  3. Name one piece of engineering complexity the MoE option carries that the dense option doesn't, independent of the memory question.

Share this question

← Back to How LLMs Are Built: Pretraining to Chatbot practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.