Architecture cost review
LLM Batch API Pricing: Where 50% Off Is Real
Half price is real, but only for text-only requests to one model per batch, with no streaming and a 24-hour window. Here is what each model costs in batch, and what silently stays at full price.
Direct answer
Batch APIs typically halve per-token prices for work that can wait up to 24 hours. OpenRouter's Batch API, launched in September 2026, brings that discount to more than 70 models under one key, but xAI models get only 20%. Web search and prompt caching are not discounted, and batches are text-only. Move every delay-tolerant job to batch first; it is the largest single saving most teams have not taken.
Decision summary
| Decision area | What matters |
|---|---|
| Typical discount | 50% of per-token price on input and output (OpenRouter, OpenAI, Anthropic) |
| xAI models | 20% off in batch — e.g. Grok 4.3 $1.00 / $2.00 vs $1.25 / $2.50 |
| Window | 24 hours, the only accepted value on OpenRouter |
| Shape | One model and one endpoint per batch; inline requests array; chat, Responses, Messages or embeddings |
| Not discounted | Web search calls; prompt-caching rates vary by model |
| Rejected | Audio and video input, non-text output, streaming, :online variants |
What changed in September
OpenAI, Anthropic and Google have offered batch discounts for some time. Using them meant a separate batch integration for each provider. In September 2026 OpenRouter launched its own Batch API, covering more than 70 models at launch. You submit an inline array of requests to POST /api/v1/batches and collect results within 24 hours.
OpenRouter bills batch requests at typically 50% of the model's per-token price, mirroring OpenAI's and Anthropic's own batch discounts. Each batch runs on a single provider chosen at submission. By default that is the cheapest eligible :batch endpoint after your provider allowlist, data policy and BYOK settings.
Batch prices for current models
Every model with batch support lists a :batch variant at the batch rate. GPT-6.1 Sol falls from $2 and $10 per million tokens to $1 and $5. Claude Opus 5.5 falls from $4 and $20 to $2 and $10. GPT-6 Astra and Claude Fable 5.1 fall from $10 and $50 to $5 and $25. At the cheap end, GPT-6 Luna drops to $0.05 and $0.25, and Gemini 3.8 Flash to $0.375 and $1.875.
xAI models are the exception at 20% off. OpenRouter's example is Grok 4.3 at $1.00 and $2.00 in batch, against $1.25 and $2.50 interactive. For xAI-heavy workloads, a cheaper model run interactively can beat a discounted xAI batch.
- GPT-6.1 Sol: $2 / $10 → $1 / $5
- Claude Opus 5.5: $4 / $20 → $2 / $10
- GPT-6 Astra, Claude Fable 5.1: $10 / $50 → $5 / $25
- GPT-6 Luna: $0.10 / $0.50 → $0.05 / $0.25
- Gemini 3.8 Flash: $0.75 / $3.75 → $0.375 / $1.875
Worked example: a nightly document job
Take 100,000 documents a night at about 2,000 input and 300 output tokens each: 200 million input and 30 million output tokens. On GPT-6.1 Sol interactive that is $400 plus $300, or $700 a night. In batch it is $350. Over a 30-night month the batch saving is $10,500, far larger than OpenRouter's 5.5% credit fee on the batch spend.
The same job on GPT-6 Luna costs $35 interactive and $17.50 in batch. If Luna's quality is acceptable for the job, the cheaper model saves more than batching the stronger one. Test quality first, then apply batch to whichever model wins.
What does not get cheaper
Non-token charges are not uniformly discounted. On OpenRouter, web-search calls bill at standard rates, and OpenRouter-orchestrated search is not available in batch at all. Prompt-caching rates vary by model, so a cache-heavy workload may save less than half. The model page is the source of truth.
Batches are text-only. Audio and video inputs are rejected, as are requests for non-text output, streaming, :online model variants and some provider beta features. Some requests fail validation after the 202 response, moving the whole batch to failed. Validate a small batch before submitting a large one.
Operational rules that change the cost
Each batch uses one model and one endpoint shape. A request may omit model to inherit the batch model, but cannot name a different one. On Google models every request must share the same response_format, so mixed schemas need separate batches. Put endpoint, model and provider before requests in the JSON body, or the API returns a 400.
Results arrive asynchronously within 24 hours. Your pipeline needs polling, retries for failed items and a fallback for anything urgent. If you use BYOK, batches route through your provider key automatically, and OpenRouter charges only its BYOK fee.
How to decide what to batch
Sort your workloads by how long a result can wait. Anything that can wait until morning — evaluations, enrichment, backfills, nightly summaries, embedding refreshes — belongs in batch. Anything a user is waiting for does not.
Then compare two options for each batchable workload: the strong model in batch, and a cheaper model interactively. Choose on cost per accepted item. For many jobs the right answer is a cheap model in batch, which can cost under 5% of a premium model run interactively.
Batch versus caching versus a cheaper model
Batch is one of three ways to cut a large token bill, and they combine differently. Prompt caching cuts repeated input, such as a long system prompt or shared document. It works interactively, but its rates vary in batch. Batch cuts every token, but only for work that can wait. A cheaper model cuts every token at all times, if its quality is good enough.
Use them in that order of investigation. First ask whether a cheaper model passes your quality bar. Then ask whether the remaining work can wait 24 hours. Then ask whether a large shared prefix makes caching worthwhile. For a nightly job on a long shared policy document, the answer is often all three: a cheap model, in batch, with a cached prefix where the provider supports it.
Record the per-item cost of each combination on a sample of a few hundred items before switching the whole job. The cheapest combination on paper sometimes fails validation or quality checks, and a failed batch costs a day.
Key takeaways
- →Batch APIs typically halve per-token prices for work that can wait up to 24 hours. OpenRouter's Batch API, launched in September 2026, brings that discount to more than 70 models under one key, but xAI models get only 20%. Web search and prompt caching are not discounted, and batches are text-only. Move every delay-tolerant job to batch first; it is the largest single saving most teams have not taken.
- →Split workloads by latency need, send every text-only job that can wait to batch, and compare batch on a strong model against interactive on a cheaper one.
- →Batch discounts and supported models vary by model and provider; OpenRouter says the model page is the source of truth, and promotional rates can change.
How this page was prepared
This September 2026 cluster uses OpenRouter's live models API, OpenRouter and TypeSafe documentation, and the Laya model cards, all checked on 1 October 2026. Vendor benchmark claims are attributed, third-party benchmarks are labelled as such, and every cost example states its token assumptions. We did not run a private benchmark for these pages.
Frequently asked questions
How much does a batch API save?
Typically 50% of the per-token price on both input and output for OpenRouter, OpenAI and Anthropic batch requests. xAI models are 20% off in OpenRouter batch.
How long do batch requests take?
OpenRouter's Batch API uses a 24-hour completion window, which is the only accepted value.
Is everything discounted in batch?
No. Web search bills at standard rates, prompt-caching rates vary by model, and batches are text-only — audio, video and streaming requests are rejected.
Can one batch use several models?
No. Each OpenRouter batch uses one model and one endpoint shape. Submit separate batches for different models or response schemas.