File 031 Science 30 September 2026 2 min read
Declassified

A model with 82 million parameters does the voice

It is roughly one four-thousandth the size of a frontier language model, and it sounds better than systems a hundred times larger. Speech turned out to be an easier problem than anyone assumed.

The question“what is the best local text to speech model”

At a glance

Kokoro
82 million parameters
Downloads
over 11 million
Licence
Apache-2.0
Whisper Large v3 Turbo
about 1.6 GB at Q4, MIT licence

Kokoro is 82 million parameters.

For scale: the embedding model that powers most search is 22 million, and it only produces lists of numbers. Frontier language models are measured in hundreds of billions. Kokoro speaks, and it is good at it — better, by most accounts, than systems many times its size.

It has been downloaded over 11 million times and carries an Apache-2.0 licence.

Why speech is easier

Three reasons, and they explain a pattern that repeats across AI.

The output is small. A second of speech is a few thousand numbers. A paragraph of text is a few hundred tokens representing far more information. Producing audio is a narrower problem.

The rules are learnable. Pronunciation, rhythm and intonation follow patterns that are consistent and largely local. You do not need to reason about them; you need to have heard a great deal of speech.

Mistakes are forgiving. A slightly wrong word in a generated sentence can make the whole thing wrong. A slightly odd syllable in speech is barely noticeable, and the ear is remarkably tolerant.

The pattern this reveals

Across AI, capability does not scale evenly with size. What matters is how much reasoning a task requires.

TaskReasoning neededModel size that does well
Transcribing speechalmost none~800M
Speaking text aloudnone~80M
Embedding for searchnone~20M
Answering questionsa lot100B+

The tasks at the top of that table are the ones where small models win, and they are also the ones most people actually need.

What to do with this

If you have been assuming local AI means a worse experience, the speech models are where that assumption collapses first.

Whisper Large v3 Turbo transcribes audio to a standard that beats most paid services, with an MIT licence and about 1.6GB of memory.

Kokoro generates speech that is genuinely pleasant for a fraction of that.

Both run on hardware you already own, both work offline, and neither sends a recording of your voice anywhere.

The honest caveat

Small models are excellent at narrow, well-defined transformations and poor at judgement. Speech is on the right side of that line.

So the useful question is never how big is it? It is:

How much does this task actually require the model to think?

Most tasks require far less than the marketing implies, and that is where the small models quietly win.

Filed underaispeechmodels

Every claim in this file is checked against primary sources — how we verify. Spotted an error? Tell us.

More Science

All
ScienceFILE 011

The Noise That Turned Out to Be the Universe

Two physicists spent a year trying to remove a persistent hiss. They cleaned out pigeon droppings. It didn't help — because the noise was the Big Bang.

Open the file1 min
ScienceFILE 010

The Dinosaur Found With Its Last Meal Still Inside

A machine operator in Alberta hit something unusual. It turned out to be a 110-million-year-old armoured dinosaur — preserved in three dimensions, with its stomach contents intact.

Open the file1 min