✓ Copied to clipboard
{ } Token Budget
● Runs in your browserNo API key neededUSD pricing · October 2026

Know what your AI workflow costs.

Start with token usage, add your workflow volume, and compare the budget across models.

1 Tokens per call

Enter the counts for one representative API call. If input usage includes cached tokens, subtract those before filling in uncached input.

0 total input tokens · Cached input is billed at each model's cache-read rate.

2 Workflow volume

A 10% allowance adds 10 extra calls per 100 planned calls. All calls use the same token mix above.

Monthly API calls9,900

100 tasks × 3 calls × 30 days + 10% extra calls

Advanced pricing settings
DeepSeek pricing window

Cache-read savings are included. Cache-write surcharges, storage, tools, and other API charges are excluded.

Your estimate

Entered token counts

Monthly token cost

—

Enter tokens or try the example above.

Per API call

—

Includes input and output tokens

Cache-read savings

—

Compared with all input at the uncached rate

Token charges only. Rates depend on each model's input tier; provider-specific tokenization and extra API charges can change the final bill.

Compare model costs

Same workload, different rates. Click a model to use it in your budget above.

Rates: Oct 2026
Provider:
Sort by:
Model & Provider ↕
Context ↕
Input Cost ↕
Output Cost ↕
Monthly Cost ▲
In / Out ($/1M)
Pricing Data Sources & Audit Log (Verified October 2026)
Audit Anchor: October 9, 2026

Audited against official OpenAI API pricing announcements for GPT-6 Astra ($10.00/$50.00), GPT-6.1 Sol ($2.00/$10.00), and GPT-6 Luna ($0.10/$0.50). Reflects Standard rates and 90%–95% prompt caching discounts. Above 272k total input tokens (including cache hits), input/cache rates double and output rates rise by 50% for the entire request; the calculator applies this tier automatically.

Anthropic Claude Documentation docs.anthropic.com/models

Audited against Anthropic API pricing console for Claude 5.5 (Opus 5.5, Sonnet 5.5, Haiku 5.5). Includes cache-read discounts and context tiers. Cache-write surcharges are not included in the calculated totals.

Google Cloud · Global Standard Google Cloud pricing

Google Cloud global Standard PayGo rates for Gemini 3.8 Flash, Gemini 3.8 Flash Cyber, and Gemini 3.1 Pro Preview (including tiered pricing notes for prompts exceeding 200k tokens).

DeepSeek Open Platform platform.deepseek.com/pricing

Audited against DeepSeek API dynamic pricing schedules for DeepSeek-V4.1-Flash and DeepSeek-V4-Pro (including UTC 01:00-04:00 & 06:00-10:00 peak hours and 50% off-peak window).

Pricing note: This is a manually maintained snapshot. Check the linked provider pages before committing a budget. For high-volume enterprise SLA commitments or custom cloud partner discounts (AWS Bedrock, Azure OpenAI, Google Cloud Vertex), contact your vendor representative.

How token pricing works · Guides & FAQs +

How token costs are calculated

Each call combines uncached input, cached input, and billed output. Their rates can differ. The calculator adds these token charges, then multiplies by your planned call count.

Cost / call = (uncached input × input rate
+ cached input × cache-read rate
+ output × output rate) / 1,000,000

Use usage reported by your API for a representative call. Some APIs include cached tokens in total input, so subtract cached input first. Text estimates use a local heuristic: they do not run a model's official tokenizer or include every message-formatting token.

Comparison uses the same counts for every model. Different tokenizers, reasoning usage, and workflow behavior can change actual consumption. Monthly plans assume the same token mix for all calls. Cache-write surcharges, storage, tools, taxes, and other charges are excluded.

Make your budget reflect the workflow

Count every model call

One user task may involve planning, tool calls, and a final response. Use the average number of API calls per completed task.

Leave room for extra work

Add an allowance for retries and extra calls. A 10% allowance adds 10 calls per 100 planned calls; it is not a probability model for repeated failures.

Measure caching

Use observed cache-read counts where possible. The savings comparison treats cached input as uncached input at the same active tier.

Compare quality alongside price

The lowest token cost does not establish which model completes your task best. Test candidate models on your actual workflow.

Knowledge Base & FAQ

Frequently Asked Questions (October 2026 Edition)

How does language affect token counts?

Token counts depend on the tokenizer, language, and text. CJK characters can split differently from English words. This calculator estimates text locally; use provider-reported counts for actual usage and compare real calls before making a budget.

What are the official API pricing rates for Anthropic Claude 5.5 models?
As of October 2026: Claude Opus 5.5 is $4.00/1M input and $20.00/1M output ($0.20 cache read); Claude Sonnet 5.5 is $2.00/1M input and $10.00/1M output ($0.10 cache read); Claude Haiku 5.5 starts from $0.10/1M input and $0.50/1M output ($0.01 cache read). All three models support a 1M token context window.
What models make up the OpenAI GPT-6 family and what is their pricing?
The OpenAI GPT-6 series comprises: GPT-6 Astra ($10.00/1M input, $50.00/1M output, $1.00 cached) for frontier scientific research and high-stakes reasoning; GPT-6.1 Sol ($2.00/1M input, $10.00/1M output, $0.10 cached) as the high-efficiency agentic coding workhorse; and GPT-6 Luna ($0.10/1M input, $0.50/1M output, $0.01 cached) for ultra-fast, high-volume ingestion and classification. All three support 1.05M tokens context. These are Standard rates for prompts up to 272k input tokens. Above 272k total input tokens (including cache hits), the entire request uses 2x input/cache rates and 1.5x output rates.
What are the official API pricing rates for Google Gemini 3 models (Gemini 3.8 Flash & 3.1 Pro)?
As of October 2026, the Gemini 3 family offers: Gemini 3.8 Flash at $0.75/1M input ($0.075 cached) and $3.75/1M output (introductory pricing through Dec 31, 2026, including internal thinking process tokens, 1M context); Gemini 3.8 Flash Cyber at $1.50/1M input and $7.50/1M output ($0.15 cached, Google Cloud global Standard rates, 1M context); and Gemini 3.1 Pro Preview at $2.00/1M input ($0.20 cached) and $12.00/1M output with a 2M token context window.
How does DeepSeek-V4.1 Flash's Peak vs. Off-Peak pricing model work?
DeepSeek-V4.1-Flash charges Peak rates ($0.30/1M input, $1.20/1M output, $0.006 cache hit) during 01:00-04:00 and 06:00-10:00 UTC Monday-Friday. During Off-Peak hours (all other times and weekends), prices are discounted by 50% ($0.15/1M input, $0.60/1M output, $0.003 cache hit). You can toggle the Off-Peak switch in our simulator to see adjusted costs.
Is my prompt or confidential schema sent to any remote server?
No. The entire utility executes 100% locally within your browser. No prompts, code, API keys, or tool schemas are transmitted over the network or saved on remote servers. All custom model rates and editor preferences are stored strictly in your browser's localStorage. Shareable links encode configuration parameters safely in the client-side URL hash without server persistence.