Every model page shows dollars per million tokens. Nobody sends a million tokens at a time. The number that matters is what your workload costs per day, and it is easy to get it wrong by three orders of magnitude if you read the rate card wrong.
A concrete workload
Take a realistic agent session: a coding or research agent doing small, frequent turns. A typical loop is:
- Input: 12,000 tokens (system prompt + tool definitions + conversation context)
- Output: 400 tokens (a tool call or a short reply)
- Cadence: 400 turns per active user per day. Agent loops are chatty; that is the point of them.
Priced against real numbers
Using live AtmoRouter rates for three strong mid-tier models — zai/glm-5.3-flash, mm/MiniMax-M2.7, and alicn/qwen3.8-omni-flash — here is one user-day:
| | glm-5.3-flash | MiniMax-M2.7 | qwen3.8-omni-flash | |---|---|---|---| | Input per turn | 12k tok | 12k tok | 12k tok | | Input price | $0.0004/M | $0.0008/M | $0.0004/M | | Output per turn | 400 tok | 400 tok | | Output price | $0.0013/M | $0.0030/M | $0.0012/M | | Per turn | ~$0.0048 | ~$0.0097 | ~$0.0048 | | Per user-day | ~$1.92 | ~$3.87 | ~$1.92 |
The precise math: at 12k input tokens, glm-5.3-flash costs 12,000 ÷ 1,000,000 × $0.0004 = $0.0048 per turn on input alone. Add 400 output tokens at $0.0013/M — $0.00000052, effectively nothing — and the session's cost is dominated entirely by input. In agent workloads, input is 95%+ of the bill.
Where the money actually goes
- 1.Context, not generation. The output above is cheap everywhere. What you pay for is re-sending the system prompt and conversation every single turn. This is why prompt caching matters more than any model choice:
glm-5.3-flashbills cached input at 10% of the input rate, so a cached agent loop costs roughly one-tenth of an uncached one. - 2.Retries, if your gateway bills them. At 400 turns/day, even a 5% retry rate is 20 extra turns. AtmoRouter never bills a request that errors; on gateways that do, retries are a silent 5% tax.
- 3.The model above your needs. The gap between a strong mid-tier model and a flagship is often 10-30× per token. For the inner loop of an agent — parsing tool calls, formatting, classification — the mid-tier models in the table above are indistinguishable in practice. Reserve flagships for the turns that need them.
The takeaway
Model pages advertise per-million prices because that is the unit of the upstream market. But the mental model that predicts your bill is: input tokens × turns. A model at $0.0004/M serving 12k-token turns costs about half a cent per turn — under two dollars a day for a heavy agent user — and the cheapest way to cut that in half is not a cheaper model, it is a cache hit.
Check any model's live rates on its AtmoRouter model page — input, output and cached input are printed per million tokens, next to the official list price.