Mixture of experts: how MoE models work
A mixture of experts (MoE) model contains many parallel sub-networks called experts and uses a router to send each token through only a few of them.
What MoE is
In a standard dense transformer, every token passes through every weight in every layer. In a mixture of experts model, the feed-forward part of each layer is replaced by a set of experts, often dozens or hundreds of them, plus a small router. For each token at each layer, the router picks a handful of experts (commonly 2 to 8), and only those run. The attention layers are typically still shared by all tokens.
This lets a model have a very large total parameter count while doing the computation of a much smaller one. Many current open models, including large DeepSeek, Qwen, GLM and Mixtral releases, use this design.
Total vs active parameters
MoE models are described with two numbers. A model named like "30B-A3B" has 30 billion parameters in total but about 3 billion active per token. The two numbers matter for different things:
- Speed follows the active count. A 30B-A3B model generates roughly as fast as a dense 3B model, given enough memory bandwidth.
- Memory follows the total count. All experts must be loaded, since any token might need any expert.
- Quality sits in between: usually much better than a dense model the size of the active count, often a little below a dense model the size of the total.
How the routing works
The router is a small learned layer that scores every expert for the current token and keeps the top K. Training includes balancing tricks so tokens are spread across experts rather than all going to a favorite few. Experts do not cleanly specialize in human topics like "history" or "poetry"; their specializations are statistical and often about token patterns or syntax.
What it means for roleplay
For a hosted API user, MoE mostly shows up as lower cost and higher speed for a given level of quality. Large MoE models can also carry a lot of world knowledge, which helps with settings and genre detail. Some users find MoE models a little less consistent in style over long chats than dense models of similar quality, but that varies by model more than by architecture. For local users, MoE models are attractive because you can keep experts in system RAM and still get usable speed.
MoE on Wild West API
You do not configure experts through the API; you pick a model by ID and the routing happens inside it. Sampler settings work the same on MoE and dense models. See the models page for each model's context length and features.
FAQ
What does A3B mean in a model name?
It means about 3 billion active parameters per token. The other number, such as 30B, is the total parameter count.
Are MoE models better than dense models?
Not automatically. They are cheaper and faster per unit of quality, but a dense model with the same total size is often slightly stronger.