File 030 Technology 29 September 2026 2 min read
Declassified

The 20B model that fits where a 13B model will not

Total size and running size stopped being the same number, and that changes which model you can actually run.

The question“how do mixture of experts models work”

At a glance

GPT-OSS 20B
20B total, 3.6B active per token
GPT-OSS 120B
120B total, 5.1B active per token
What you pay for
memory buys knowledge; compute buys speed
What to check
both the total and the active parameter count

For years the arithmetic of local AI was simple: bigger model, more memory, slower. One number told you everything important.

That stopped being true, and the reason is the most useful thing to understand if you are choosing hardware.

Two numbers, not one

A mixture-of-experts model splits its network into many specialised sections, and for each token it uses only a few of them.

So there are now two sizes:

  • Total parameters — everything the model knows, which is what has to be stored
  • Active parameters — what actually runs for each token, which is what determines speed

GPT-OSS 20B has about 20 billion parameters in total and roughly 3.6 billion active. GPT-OSS 120B has around 120 billion, with about 5.1 billion active.

The 120B model has six times the knowledge of the 20B and runs only about 40% slower, because the amount of computation per token barely changed.

Why that is a bigger deal than it sounds

It breaks the trade-off people have been making since local models became possible.

Previously: to get more capability you needed more memory and you accepted that it would be slower. Both costs moved together.

Now the two have come apart. You pay memory for knowledge and compute for speed, and you can buy them separately.

Practical consequences:

  • A model that would not have fit now fits, and runs at a usable speed
  • Buying memory buys you knowledge rather than merely throughput
  • The best model for your machine is no longer simply the largest one that loads

What to check before you download

Two numbers, not one. Look for the total and the active parameter count on the model card.

A 30B mixture with 3B active will run like a small model and know considerably more than one. A 30B dense model will run like a 30B dense model, which on a laptop is slow enough to be annoying.

Most comparisons quote only the total, which is exactly the number that tells you least about the experience of using it.

The catch worth knowing

Mixture-of-experts trades memory efficiency for memory capacity. All those experts have to be resident even though most are idle at any moment.

So it is a wonderful technology if you have enough memory and not enough speed. If you are short on memory, it does not help you at all — it just makes the model bigger.

The trick is knowing which of those two you are.

Filed underailocalhardware

Every claim in this file is checked against primary sources — how we verify. Spotted an error? Tell us.

More Technology

All
TechnologyFILE 029

The most downloaded AI model nobody has heard of

It has 250 million downloads, it is 22 million parameters, and it will never write you a sentence. It may also be doing more real work than anything else on the hub.

Open the file2 min