How token costs are calculated
Each call combines uncached input, cached input, and billed output. Their rates can differ. The calculator adds these token charges, then multiplies by your planned call count.
+ cached input × cache-read rate
+ output × output rate) / 1,000,000
Use usage reported by your API for a representative call. Some APIs include cached tokens in total input, so subtract cached input first. Text estimates use a local heuristic: they do not run a model's official tokenizer or include every message-formatting token.
Comparison uses the same counts for every model. Different tokenizers, reasoning usage, and workflow behavior can change actual consumption. Monthly plans assume the same token mix for all calls. Cache-write surcharges, storage, tools, taxes, and other charges are excluded.
Make your budget reflect the workflow
Count every model call
One user task may involve planning, tool calls, and a final response. Use the average number of API calls per completed task.
Leave room for extra work
Add an allowance for retries and extra calls. A 10% allowance adds 10 calls per 100 planned calls; it is not a probability model for repeated failures.
Measure caching
Use observed cache-read counts where possible. The savings comparison treats cached input as uncached input at the same active tier.
Compare quality alongside price
The lowest token cost does not establish which model completes your task best. Test candidate models on your actual workflow.
Frequently Asked Questions (October 2026 Edition)
How does language affect token counts?
Token counts depend on the tokenizer, language, and text. CJK characters can split differently from English words. This calculator estimates text locally; use provider-reported counts for actual usage and compare real calls before making a budget.