The same question never gets the same answer twice
Ask a model something, then ask again. The wording will differ, and occasionally the answer will too. That is not a bug being fixed.
At a glance
- Cause
- the model samples rather than looking up
- The setting
- temperature and top-p control how adventurous it is
- Consequence
- the same prompt can succeed once and fail the next time
- What helps
- ask twice, and save the answer you verified
Type the same question into the same model twice and you will get two different answers. Usually they agree. Sometimes they do not.
This surprises people, and it is worth understanding, because it explains a whole class of strange behaviour.
What is happening
When a model produces a response, it does not look up an answer. It works out, for each step, how likely each possible next piece of text is, then picks one.
If it always picked the single most likely option, every answer to the same question would be identical. So instead it samples — with a bias toward the likely options, but not a rule.
That is the source of the variation. It is a deliberate design choice, and it is the same mechanism that lets a model be asked for ten different headlines without producing the same one ten times.
Why it is not simply a flaw
A model that always produced the single most likely continuation would be repetitive and dull, and worse at anything creative. Ask for five names for a product and you would get one name five times.
The randomness is a feature with a cost. The cost is that answers are not reproducible.
What it means in practice
Accuracy varies run to run. A model that gets a question right nine times in ten will get it wrong on the tenth attempt. The failure is not a different model; it is the same model, sampled differently.
An answer is not evidence. Getting a good answer once does not mean you will get it again. If something must be right, ask more than once and compare.
Some tools reduce it. Most APIs let you lower the temperature — the setting that controls how adventurous the sampling is. At its lowest the model becomes nearly deterministic, and nearly repetitive.
It makes reproducibility hard. In any serious application this matters: the same input producing different output is a property most software does not have, and systems built on AI have to be designed around it.
What to do about it
Three habits that hold up:
1. Ask twice when it matters. Disagreement between two answers is a useful warning that the question is in uncertain territory. 2. Lower the temperature for factual work, if your tool exposes the setting. Determinism is worth more than variety when you want a correct answer. 3. Keep the version you checked. If you verified something, save the text. Re-asking does not reproduce it.
The thing to hold on to
A language model is not a database that returns a stored fact. It is a process that produces plausible text, and it runs slightly differently each time.
That is what makes it useful, and it is exactly why the output needs checking rather than trusting.
Every claim in this file is checked against primary sources — how we verify. Spotted an error? Tell us.
Get the next file before it's public
One story a week, chosen from the archive. The kind you repeat at dinner and nobody believes.
One dispatch a week. Unsubscribe anytime.
More Science
AllA model with 82 million parameters does the voice
It is roughly one four-thousandth the size of a frontier language model, and it sounds better than systems a hundred times larger. Speech turned out to be an easier problem than anyone assumed.
The Noise That Turned Out to Be the Universe
Two physicists spent a year trying to remove a persistent hiss. They cleaned out pigeon droppings. It didn't help — because the noise was the Big Bang.
The Dinosaur Found With Its Last Meal Still Inside
A machine operator in Alberta hit something unusual. It turned out to be a 110-million-year-old armoured dinosaur — preserved in three dimensions, with its stomach contents intact.