What is a mixture-of-experts AI model?
A mixture of experts (MoE) is an AI model design that splits part of the network into many sub-networks, called experts. For each piece of input, a router picks only a few of them, so only part of the model works at each step.
Also known as: MoE, mixture-of-experts, sparse mixture of experts
Researched and fact-checked by AI, with no human review. 7 sources listed below. How we verify
Last updated
How a mixture of experts works
Google's machine learning glossary describes a mixture of experts as a way to make a neural network more efficient. The network uses only a subset of its parameters, known as an expert, to process a given piece of input. A second part, the gating network, routes each input to the proper experts.
A 2017 paper by Noam Shazeer and six co-authors described a layer with up to thousands of expert sub-networks. A trainable gating network chose a small set of them for each example. The aim was a large rise in model capacity without a proportional rise in computation.
Mistral's Mixtral 8x7B, released in December 2023, applied the idea to a language model. Mistral says the model picks from 8 distinct groups of parameters. At every layer, a router chooses two of them for each token, a unit of text.
Why labs quote total and active parameters
Only some experts run at each step. So labs give two sizes for these models. Total parameters count the whole network. Active parameters count the part used for each token.
These are the figures each lab has published:
- Mixtral 8x7B: 46.7 billion total, with 12.9 billion used per token.
- Mistral Large 3: 675 billion total and 41 billion active, per Mistral's December 2025 announcement.
- DeepSeek-V4-Pro: 1.6 trillion total and 49 billion active, per DeepSeek's April 2026 release note.
- Mistral Large 4: 1.05 trillion total and 52 billion active, per Mistral's documentation on October 8, 2026.
What the design changes in practice
Mistral says the technique adds parameters to a model while keeping cost and latency under control. It says Mixtral handles input and output at the same speed and cost as a 12.9 billion-parameter model.
A small active count does not make a model small to host. The New Stack noted in October 2026 that serving the full Mistral Large 4 model will still require a substantial setup with several GPUs. GPUs are the chips used to run AI models.
Sources
- Machine Learning Glossary, Google for Developers
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer, arXiv (Shazeer et al.)
- Mixtral of experts, Mistral AI
- Introducing Mistral 3, Mistral AI
- DeepSeek V4 Preview Release, DeepSeek
- Mistral Large 4 - Mistral AI | Mistral Docs, Mistral AI
- Mistral's new AI tried to escape its test environment. In three weeks, anyone can download it, The New Stack