A model with 82 million parameters does the voice
It is roughly one four-thousandth the size of a frontier language model, and it sounds better than systems a hundred times larger. Speech turned out to be an easier problem than anyone assumed.
At a glance
- Kokoro
- 82 million parameters
- Downloads
- over 11 million
- Licence
- Apache-2.0
- Whisper Large v3 Turbo
- about 1.6 GB at Q4, MIT licence
Kokoro is 82 million parameters.
For scale: the embedding model that powers most search is 22 million, and it only produces lists of numbers. Frontier language models are measured in hundreds of billions. Kokoro speaks, and it is good at it — better, by most accounts, than systems many times its size.
It has been downloaded over 11 million times and carries an Apache-2.0 licence.
Why speech is easier
Three reasons, and they explain a pattern that repeats across AI.
The output is small. A second of speech is a few thousand numbers. A paragraph of text is a few hundred tokens representing far more information. Producing audio is a narrower problem.
The rules are learnable. Pronunciation, rhythm and intonation follow patterns that are consistent and largely local. You do not need to reason about them; you need to have heard a great deal of speech.
Mistakes are forgiving. A slightly wrong word in a generated sentence can make the whole thing wrong. A slightly odd syllable in speech is barely noticeable, and the ear is remarkably tolerant.
The pattern this reveals
Across AI, capability does not scale evenly with size. What matters is how much reasoning a task requires.
| Task | Reasoning needed | Model size that does well |
|---|---|---|
| Transcribing speech | almost none | ~800M |
| Speaking text aloud | none | ~80M |
| Embedding for search | none | ~20M |
| Answering questions | a lot | 100B+ |
The tasks at the top of that table are the ones where small models win, and they are also the ones most people actually need.
What to do with this
If you have been assuming local AI means a worse experience, the speech models are where that assumption collapses first.
Whisper Large v3 Turbo transcribes audio to a standard that beats most paid services, with an MIT licence and about 1.6GB of memory.
Kokoro generates speech that is genuinely pleasant for a fraction of that.
Both run on hardware you already own, both work offline, and neither sends a recording of your voice anywhere.
The honest caveat
Small models are excellent at narrow, well-defined transformations and poor at judgement. Speech is on the right side of that line.
So the useful question is never how big is it? It is:
How much does this task actually require the model to think?
Most tasks require far less than the marketing implies, and that is where the small models quietly win.
Every claim in this file is checked against primary sources — how we verify. Spotted an error? Tell us.
Get the next file before it's public
One story a week, chosen from the archive. The kind you repeat at dinner and nobody believes.
One dispatch a week. Unsubscribe anytime.
More Science
AllThe Noise That Turned Out to Be the Universe
Two physicists spent a year trying to remove a persistent hiss. They cleaned out pigeon droppings. It didn't help — because the noise was the Big Bang.
The Dinosaur Found With Its Last Meal Still Inside
A machine operator in Alberta hit something unusual. It turned out to be a 110-million-year-old armoured dinosaur — preserved in three dimensions, with its stomach contents intact.