← all articles

Mixture of experts and why parameter counts mislead

Every time a lab drops a new model, the headline number is total parameters. 47 billion. 236 billion. 671 billion. It reads like a horsepower spec, bigger is better, and for years that shorthand roughly held. Then mixture of experts architectures went mainstream, and the shorthand broke. A mixture of experts LLM can carry a parameter count that looks enormous on paper while running inference at a cost closer to a model a fraction of its size. If you’re picking models based on the number in the announcement post, you’re optimizing for the wrong variable.

What a mixture of experts model actually is

A standard dense transformer uses every parameter on every token. Feed it one word, and the full weight matrix in every feedforward layer gets multiplied through, whether that computation was useful for this particular token or not. That’s simple and it’s why dense models are easy to reason about, but it means cost scales directly with size. Double the parameters, roughly double the compute per token.

A mixture of experts model swaps the single feedforward block in each transformer layer for a bank of several smaller feedforward blocks, called experts, plus a small router network. For every token, the router looks at the token’s representation and picks a handful of experts to actually run, typically the top 2 out of 8, or top 8 out of a much larger pool in newer designs. The rest of the experts sit there unused for that token. Attention layers still run in full, but the feedforward layers, which hold most of the parameters, only fire a slice of themselves.

That’s the whole trick. You get a huge pool of parameters to draw from, so the model can specialize different experts for different kinds of input, but any single forward pass only touches a fraction of that pool.

Total parameters versus active parameters

This is the number that actually matters, and it’s usually buried a paragraph or two into the model card instead of in the headline.

Mistral’s Mixtral 8x7B has 8 experts per layer with top-2 routing. Its published architecture puts total parameters around 47 billion, but a given token only activates roughly 13 billion of those. DeepSeek-V3’s technical report describes a much larger split: 671 billion total parameters, but only about 37 billion active per token, using a fine-grained pool of routed experts plus a couple of always-on shared experts. In both cases the total figure is what gets quoted in coverage, and the active figure is what actually predicts how expensive and how fast the model is to run.

This is why a mixture of experts LLM with a total parameter count several times larger than a dense competitor can still come in cheaper per token to serve. Compute cost tracks active parameters far more closely than it tracks the total. The 671B number tells you how much specialized capacity the model has to draw on. The 37B number tells you what you’re actually paying compute for on any given request.

Why the compute picture isn’t the full picture

Here’s the part that doesn’t make it into the marketing copy: sparse activation saves you compute, not memory.

Every expert has to be loaded and ready, because the router’s choice changes token by token, sometimes token to token within the same sequence. You can’t predict in advance which experts a given request will need, so you can’t just page the unused ones out. That means VRAM or system RAM requirements track the total parameter count, not the active one. A 671B total parameter model needs the memory footprint of a 671B model, full stop, even though it’s only doing the arithmetic of a 37B model on each token. If you’re renting GPU capacity by the hour to self-host, the memory bill and the compute bill decouple from each other in a way dense models never made you think about.

This is also why MoE models are unusual to quantize and fine-tune well. Quantization error behaves differently across experts that see wildly different amounts of training signal, since some experts get routed to far more often than others. And fine-tuning risks starving experts that were rarely activated during pretraining, since gradient updates for a token only flow through the couple of experts the router picked for it. None of that shows up in a parameter count either.

The routing problem nobody puts on the landing page

The router is a small learned network, and it has a known failure mode: without correction, it converges toward using its favorite handful of experts and starves the rest. This is usually called expert collapse. Labs counter it with auxiliary load-balancing losses during training that penalize the router for uneven usage, and DeepSeek-V3’s report specifically calls out an auxiliary-loss-free balancing strategy as a design choice, which tells you this is still an active engineering problem, not a solved one. Get load balancing wrong and you end up with a model that has hundreds of billions of total parameters on paper, most of which are undertrained dead weight, while a small cluster of experts does almost all the actual work. In that failure mode, the total parameter count is even more misleading than usual, because the “extra capacity” the count implies never got meaningfully trained.

Routing also isn’t free at inference time. The router itself has to run on every token before the gating decision is made, and depending on serving infrastructure, tokens routed to different experts can end up needing different GPUs or nodes, which introduces communication overhead that a dense model simply doesn’t have. This is part of why MoE serving stacks are more complex than dense model serving stacks, and why not every provider that hosts a dense 70B model can serve a comparably capable MoE model with the same ease.

What to actually check instead of the headline number

When you’re picking between models for a coding assistant, a RAG pipeline, or anything else you’re paying API calls for, the total parameter count on its own tells you almost nothing about your bill or your latency. What actually predicts those things:

Active parameters per token, if the vendor publishes it. This is the number that correlates with compute cost and, generally, with inference latency.

Whether the provider bills you per token or per compute unit. Most API pricing is per input and output token regardless of architecture, which means a mixture of experts LLM’s efficiency advantage shows up as the vendor being able to offer a lower price for a given quality tier, not as a line item you see directly. If two models are priced identically per token but one is MoE with a much lower active parameter count, that’s a signal about margin and scalability on the vendor’s end, not necessarily about your bill today.

Context window and throughput behavior under load, which you have to actually test with your own traffic pattern, since routing overhead and expert-parallelism setups affect latency differently across providers and hardware.

None of this is a reason to distrust MoE as an architecture. It’s a genuinely useful way to get more specialized capacity without paying for it on every single token, and it’s why some of the strongest open-weight models released in the last two years use it. It’s a reason to stop treating the total parameter count as a proxy for cost or speed, because for MoE models it flatly isn’t one anymore. Read past the headline number to the active parameter count, and treat everything else, memory footprint, quantization behavior, routing stability, as a separate question you still have to ask.

If you want more breakdowns like this one on how the models you’re actually paying for work under the hood, AI Tool Gazette covers it without the vendor spin.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →