Token Cost Explained: What You're Actually Paying For
A plain-English guide to AI tokens, how they're counted, and exactly how to calculate your real costs.
The article text carries the review date. The rate table below is rebuilt from the live catalogue on every deploy.
What Is a Token?
A token is about 0.75 words, or roughly 4 characters of English text, and you pay per million of them: from $0.05 per 1M input tokens on DeepSeek V4 Flash to $5 on GPT-5.5 and Claude Opus 5. A 1,000-word document is about 1,333 tokens, so sending it to GPT-5.4 ($2.50 per 1M input) costs about $0.003.
Tokenisation is how models process text: words are converted into numerical tokens before processing. Examples:
- "The quick brown fox" = about 4 tokens
- "Hello, world!" = 3 tokens
- A typical short email = 100–200 tokens
- A 1,000-word blog post = about 1,333 tokens
Token counts vary by model. GPT models use the tiktoken library; Claude, Gemini and DeepSeek use their own tokenisers. Non-English text, code and symbols usually tokenise less efficiently, so budget 20–30% more for those.
Input vs Output Tokens
Input tokens are what you send. Output tokens are what the model writes back. Output is almost always dearer, often by several times.
OpenAI GPT Latest example:
- Input: $2.00 per 1M tokens
- Output: $10.00 per 1M tokens
- A 1,000-token response costs $0.01 in output — a fraction of a cent
That asymmetry is why “be concise” is a cost control, not a style note. Capping output length is usually the fastest per-call saving available to you.
How to Calculate Monthly Cost
Work in millions of tokens per month, not per call. Per-call numbers are too small to reason about and hide the total.
Example calculation:
- You send 500 requests/day, each with 500 input tokens and 200 output tokens
- Monthly input: 500 × 500 × 30 = 7.5M tokens × $2.00/1M = $15.00
- Monthly output: 500 × 200 × 30 = 3M tokens × $10.00/1M = $30.00
- Total: $45.00/month on OpenAI GPT Latest
Same workload on OpenAI GPT Mini Latest:
- Monthly input: 7.5M × $0.75/1M = $5.63
- Monthly output: 3M × $4.50/1M = $13.50
- Total: $19.13/month — 57% cheaper
And on a high-throughput budget row such as DeepSeek V4 Flash ($0.04/$0.08 per 1M), the same month lands near $0.54. The model you route to matters far more than the prompt tricks you apply to it.
OpenAI API — pay-per-token, no surprises
Once you understand token pricing, per-1M rates make budgeting exact. GPT-5.4 nano at $0.20 input / $1.25 output per 1M is the cheapest OpenAI row, and GPT-5.4 mini at $0.75 / $4.50 is the workhorse for most production traffic.
Context Window and Why It Matters for Cost
Every model has a context window — the maximum number of tokens it can process at once. Most current flagships list 1M tokens: GPT-5.4, GPT-5.5, Claude Sonnet 5, Claude Opus 5 and Gemini 3.1 Pro. DeepSeek V4 Flash and Llama 4 Scout list 1.3M; Claude Haiku 4.5 lists 200K; GPT-5.4 mini and nano list 400K.
A big window is a cost risk, not a cost saving. Every token in the window is billed as input on every request, so an agent that keeps 300K tokens of files in context pays 300,000 × $2.50 / 1M = $0.75 per step on GPT-5.4 — $15 for a 20-step run before any output is counted.
Optimisation: send only the last N messages of history, summarise older context instead of including raw text, and let agents retrieve files on demand rather than preloading a whole repository.
Cached Tokens and Discounts
Anthropic, OpenAI and Google all offer prompt caching — re-using the processing of an identical prompt prefix. Cached input reads cost roughly 10% of the list input rate on Anthropic and OpenAI and roughly 25% on Gemini; Anthropic charges about 125% of list to write the cache the first time.
Worked example: a 5,000-token system prompt on Claude Sonnet 5 ($2 per 1M input) costs $0.01 per request uncached. Cached, each later read is roughly $0.001. At 10,000 requests a month that is about $10 instead of $100.
Two more discounts worth knowing:
- Batch API: Anthropic and Google list batch rows at 50% of the standard rate (Claude Sonnet 5 batch: $1 / $5 per 1M; Gemini 3.8 Flash batch: $0.375 / $1.875). Use it for anything that can wait hours — backfills, evaluations, nightly classification.
- Tool-call loops: an agent re-sends its context on every step, so caching the stable prefix (system prompt, tool definitions, loaded files) is the single largest saving for agent workloads.
Most production applications should enable caching by default and route non-interactive work to batch.
Key Takeaways
- →1 token ≈ 0.75 words ≈ 4 characters; a 1,000-word document is about 1,333 tokens
- →Output costs 3–6× input on current models: GPT-5.4 is $2.50 in / $15 out per 1M, Claude Sonnet 5 is $2 / $10
- →Monthly cost = (input tokens × input rate) + (output tokens × output rate) — work in millions per month
- →Cached input reads are roughly 10% of list on OpenAI and Anthropic; batch API rows are 50% of list
- →A 1M-token context window is billed as input on every request — keep agent context lean
Editorial context
Who is this for?
Developers, startups, and teams who want to reduce their AI API or subscription costs without sacrificing quality.
When NOT to use this
Users who need real-time data, image generation, or proprietary enterprise integrations may need more specialised tools.
Pricing insights
AI pricing varies widely — some models charge per token while others use flat subscriptions. Token-based APIs are usually cheaper for moderate usage, while subscriptions suit power users with high and consistent volume.
Alternatives to consider
Consider DeepSeek V4 Flash for cost-effective coding and writing, Gemini 3.8 Flash for fast tasks, or Claude Haiku 4.5 for lightweight structured work. Use the calculator to compare your specific usage.
Final verdict
The cheapest AI tool is the one that fits your exact workload. Use the cost calculator and decision engine on this site to find your optimal stack — most users can cut AI spend by 50% or more.
Frequently Asked Questions
How much does 1 million tokens cost?
It depends on the model and direction. Input: $0.05 on DeepSeek V4 Flash, $0.75 on GPT-5.4 mini, $2 on Claude Sonnet 5, $2.50 on GPT-5.4, $5 on GPT-5.5. Output is 3–6× higher: $0.16, $4.50, $10, $15 and $30 respectively.
How many tokens are in 1,000 words?
About 1,333 tokens for English prose (roughly 0.75 words per token). Code, non-English text and symbols use more tokens per word.
Are input and output tokens priced differently?
Yes. Output is always dearer. GPT-5.4 lists $2.50 per 1M input and $15 per 1M output; Claude Sonnet 5 lists $2 and $10. Capping response length is the fastest per-call saving.
What is a cached token?
A token in a prompt prefix the provider has already processed. Cached reads cost roughly 10% of the list input rate on OpenAI and Anthropic and roughly 25% on Gemini, which matters most for long system prompts and agent loops.
Related
Free courses · no sign-up
Still deciding? Learn the basics first, then come back to the prices.