File 033 Money 2 October 2026 2 min read
Declassified

Why you pay half price to wait

Every major provider sells the same model twice, at two prices. The cheaper one is identical in every way except when you get the answer.

The question“what is batch pricing for AI models”

At a glance

Standard rate
$10 in / $50 out per million tokens
Batch rate
$5 in / $25 out
The discount
50%, in exchange for not waiting on the answer
Why it exists
the provider fills idle capacity instead of buying more of it

On the price list, the two most expensive models in the world are listed twice:

InputOutput
Standard$10.00$50.00
Batch$5.00$25.00

Same model. Same quality. Half the price. The only difference is that you do not get the answer immediately.

Why the discount exists

Because the provider's cost is not the same, and it is not about the model at all.

Running a model at scale is a scheduling problem. Graphics cards are expensive and sit idle between requests. If all your work arrives in unpredictable bursts, you have to buy enough capacity for the peaks and watch it sit unused the rest of the time.

A batch request lets the provider fill the gaps. Your work waits until there is spare capacity, and the provider sells something that would otherwise have been wasted. Both sides win, which is why the discount is so large.

It is the same logic as an off-peak train ticket. The seat costs the same to provide. Filling it at a lower price beats leaving it empty.

When the trade makes sense

Anything measured in thousands. Classifying a year of support tickets, extracting fields from an archive of invoices, summarising a document collection, translating a manual. These are jobs where you care about the result tomorrow, not in four seconds.

Any overnight work. If nobody is waiting, latency is worthless and the discount is free money.

When it does not

Anything interactive. A chat, an autocomplete, an agent working through a task step by step — all of these need answers now, because the next step depends on the last.

Anything small. Setting up a batch job for twelve requests takes longer than the saving is worth.

The number that decides it

Ask what the delay actually costs you.

If nothing downstream is waiting, the cost is zero and the discount is pure profit on the same work. If a person is watching, the cost is their attention, which is usually more than the money.

Waiting is the cheapest input you have. Most AI spending is on work that was never urgent in the first place, run at the urgent price because nobody changed a setting.

Where to find it

Most major providers offer a batch option, usually at 50% off, sometimes more. It is one parameter in the API call, and it is one of the largest single savings available to anyone spending real money on AI.

It is also the one almost nobody uses, because the default is the expensive path and changing it requires knowing that the choice exists.

Filed underaipricing

Every claim in this file is checked against primary sources — how we verify. Spotted an error? Tell us.

More Money

All
MoneyFILE 023

The two biggest AI models cost exactly the same

OpenAI's newest flagship and Anthropic's newest flagship are priced identically, to the cent. That is not a coincidence, and it tells you something about where the market has gone.

Open the file2 min