Models & Architectures
Active Parameters in plain English.
Also known as: activated parameters,total vs active parameters,sparse activation
The one-sentence version
In a mixture-of-experts model, the parameters actually used for each token — far fewer than the total, and what drives speed and cost.
Mixture-of-experts models are described with two numbers: total parameters and active parameters. Total is everything stored; active is how many are used for any given token, because a router sends each token to a few experts rather than all of them. Xiaomi's MiMo-V2.6-Pro, for example, has about 1.02 trillion total parameters but 42 billion active; DeepSeek V4.1 Flash is around 748 billion total with a 552 billion backbone. Active parameters determine how much compute each token costs, so they predict speed and price; total parameters determine how much knowledge and nuance the model can hold, and how much memory it takes to serve. This is why a model with a trillion total parameters can be cheaper per token than a 70-billion dense model, and why open-weight MoE giants are affordable to run as an API but impractical to self-host: you must hold every expert in memory even though you use a few at a time. When comparing models, read both numbers; "parameters" alone no longer tells you what you are paying for.