Advanced
Open
Pro
Does Mixture-of-Experts Actually Save You Memory?
A team is comparing two model options for a self-hosted deployment: a dense model with 30B total (and active) parameters, and an MoE model with 120B total parameters but only 15B active parameters per token (an 8-expert-total, 1-expert-active-per-token routing setup, roughly). A team member argues "the MoE model only activates 15B parameters per token, so it needs less than half the GPU memory of the dense 30B model."
- Is that memory claim correct? Explain what actually determines a model's memory footprint versus what determines its per-token compute cost.
- Given that, which of the two options is actually cheaper to host in terms of GPU memory, and which is cheaper in terms of per-token inference compute?
- Name one piece of engineering complexity the MoE option carries that the dense option doesn't, independent of the memory question.
Share this question