Why you pay half price to wait
Every major provider sells the same model twice, at two prices. The cheaper one is identical in every way except when you get the answer.
At a glance
- Standard rate
- $10 in / $50 out per million tokens
- Batch rate
- $5 in / $25 out
- The discount
- 50%, in exchange for not waiting on the answer
- Why it exists
- the provider fills idle capacity instead of buying more of it
On the price list, the two most expensive models in the world are listed twice:
| Input | Output | |
|---|---|---|
| Standard | $10.00 | $50.00 |
| Batch | $5.00 | $25.00 |
Same model. Same quality. Half the price. The only difference is that you do not get the answer immediately.
Why the discount exists
Because the provider's cost is not the same, and it is not about the model at all.
Running a model at scale is a scheduling problem. Graphics cards are expensive and sit idle between requests. If all your work arrives in unpredictable bursts, you have to buy enough capacity for the peaks and watch it sit unused the rest of the time.
A batch request lets the provider fill the gaps. Your work waits until there is spare capacity, and the provider sells something that would otherwise have been wasted. Both sides win, which is why the discount is so large.
It is the same logic as an off-peak train ticket. The seat costs the same to provide. Filling it at a lower price beats leaving it empty.
When the trade makes sense
Anything measured in thousands. Classifying a year of support tickets, extracting fields from an archive of invoices, summarising a document collection, translating a manual. These are jobs where you care about the result tomorrow, not in four seconds.
Any overnight work. If nobody is waiting, latency is worthless and the discount is free money.
When it does not
Anything interactive. A chat, an autocomplete, an agent working through a task step by step — all of these need answers now, because the next step depends on the last.
Anything small. Setting up a batch job for twelve requests takes longer than the saving is worth.
The number that decides it
Ask what the delay actually costs you.
If nothing downstream is waiting, the cost is zero and the discount is pure profit on the same work. If a person is watching, the cost is their attention, which is usually more than the money.
Waiting is the cheapest input you have. Most AI spending is on work that was never urgent in the first place, run at the urgent price because nobody changed a setting.
Where to find it
Most major providers offer a batch option, usually at 50% off, sometimes more. It is one parameter in the API call, and it is one of the largest single savings available to anyone spending real money on AI.
It is also the one almost nobody uses, because the default is the expensive path and changing it requires knowing that the choice exists.
Every claim in this file is checked against primary sources — how we verify. Spotted an error? Tell us.
Get the next file before it's public
One story a week, chosen from the archive. The kind you repeat at dinner and nobody believes.
One dispatch a week. Unsubscribe anytime.
More Money
AllTwenty-four AI models are free, and nobody is using them
There is no catch in the pricing column. There is a catch somewhere else, and it is the oldest one there is.
The cheapest AI is 8,824 times cheaper than the most expensive
Everyone calls it the same thing: using an AI model. The price list says these are not remotely the same product.
The two biggest AI models cost exactly the same
OpenAI's newest flagship and Anthropic's newest flagship are priced identically, to the cent. That is not a coincidence, and it tells you something about where the market has gone.