AI Model Routing: How to Cut Costs Without Losing Quality
How to route AI requests to cheaper models intelligently — the technique used by companies spending millions on AI to stay profitable.
The article text carries the review date. The rate table below is rebuilt from the live catalogue on every deploy.
What Is Model Routing?
Model routing sends different AI requests to different models based on their complexity. Simple tasks go to cheap, fast models such as GPT-6 Luna ($0.10 / $0.50 per 1M), DeepSeek V4.1 Flash ($0.30 / $1.20 at peak, $0.15 / $0.60 off-peak per OpenRouter's launch note) or Mercury 2.5 ($0.04 / $0.15 while listed at 80% off); complex tasks go to GPT-6.1 Sol or Claude Sonnet 5 (both $2 / $10), and only the hardest to Claude Opus 5.5 ($4 / $20) or GPT-6 Astra ($10 / $50). Rates are OpenRouter's as of 1 October 2026.
Done well, routing cuts AI costs by 60–80% with no visible quality drop. The arithmetic: if 70% of 10M monthly input tokens move from GPT-6.1 Sol ($20 for all 10M) to GPT-6 Luna, the input bill becomes 7M × $0.10 + 3M × $2 = $0.70 + $6.00 = $6.70, before touching the output side.
Rule-Based Routing
Rule-based routing is the cheapest thing to build and covers most of the saving. Start here before anything clever.
Examples:
- Short Q&A under 100 tokens → OpenAI GPT Mini Latest ($0.75/1M input)
- Bulk classification → DeepSeek V4 Flash ($0.04/1M input), the cheapest capable tier
- Long document analysis → a large-context row; Anthropic Claude Sonnet Latest and Google Gemini Pro Latest both list million-token context windows
- Code generation → Anthropic Claude Haiku Latest for routine work, escalating to Anthropic Claude Sonnet Latest only when review rejects the cheap output
- Complex reasoning → OpenAI GPT Latest or Anthropic: Claude Opus Latest ($4.00/1M input on Claude Opus 5.5) only, and only on an explicit rule
Every rule needs a measured rejection rate. A rule that sends work to a cheap model which then fails review twice is more expensive than never having written it.
LLM-Based Routing
When rules run out, classify the request with a cheap model before you spend on an expensive one.
1. Send the request to OpenAI GPT Mini Latest with a classification prompt ($0.75/1M input) 2. If it returns 'complex': route to OpenAI GPT Latest at $2.00/1M input 3. If it returns 'simple': answer in the cheap model and skip the $1.25/1M premium entirely
The classifier call is not free, so it only pays when the share of simple requests is high. Measure the split before you build it.
DeepSeek V4 Flash — the cheap-tier router target
Most routing strategies send 60–80% of traffic to a cheap workhorse. DeepSeek V4 Flash lists $0.05 input / $0.16 output per 1M — 50× cheaper than GPT-5.4 on input — while handling classification, extraction and short answers well.
Fallback Routing
Start cheap, escalate on failure:
1. Send the request to the Tier 1 model (cheap) 2. Evaluate the response against your criteria 3. If the threshold is not met, retry with the Tier 2 model
This works well when success is evaluable: code that compiles, JSON that parses, a classification with a confidence score. The retry is not free — a Tier 1 failure plus a Tier 2 retry costs Tier 1 plus Tier 2 — so measure the rejection rate. At a 20% rejection rate on GPT-5.4 mini ($0.75 input) escalating to GPT-5.4 ($2.50), the blended input rate is 0.8 × $0.75 + 0.2 × ($0.75 + $2.50) = $1.25 per 1M, still half the flagship rate.
For agents, apply the same idea per step: run the planning step on the strong model and the repetitive tool-call steps on the cheap one, cache the shared prefix (cached input is roughly 10% of list on OpenAI and Anthropic), and give every run a hard step and token budget so a loop cannot escalate itself into a large bill.
Tools for Routing
Several frameworks handle routing out of the box:
- LiteLLM: unified API wrapper with fallbacks, load balancing and per-key budgets
- RouteLLM: open-source router from Berkeley, tuned for cost-quality trade-offs
- Martian: managed routing service that learns from your quality feedback
- OpenRouter: one key across 500+ catalogue rows, with provider fallback built in
For most applications, LiteLLM or a simple custom router is sufficient. Route batchable work (backfills, evaluations, nightly classification) to the batch API as well — Anthropic and Google list batch rows at 50% of standard rates, and OpenRouter's Batch API (live since September 2026) bills most models at about 50% with results within 24 hours.
Jev Router: Picks the Model and the Effort, and Keeps the Cache
OpenRouter added typesafe/jev-router in September 2026. Before each turn, TypeSafe's Jev decision model scores the prompt on difficulty and precision, then checks whether a bigger model or more reasoning effort would help, whether a cheaper model is enough, and whether the task has changed. It chooses both the model and the reasoning effort.
What matters for cost:
- Router fee: OpenRouter lists Jev Router itself at $0 prompt and $0 completion. You pay the rate of the model each turn is routed to, so the price is variable and depends on your mix.
- Cache-aware: once a model works, Jev Router keeps it for the rest of the session and raises or lowers effort on that model so the conversation stays cached. It switches only when the expected gain beats the cost, including the cache it would lose.
- Fails, does not fall back: if the Jev call times out or returns invalid output, the request fails rather than falling back to another router. Your client needs its own retry or default model.
- Audit trail: every response carries routing metadata with the reason, and OpenRouter Chat shows the chosen model and scores per turn.
- Privacy and region: Jev reads conversation text only, never attachments, under zero data retention, and
zdr: truerequests work. Router models are not served on the us. and eu. regional domains. - Context: 1M tokens.
OpenRouter's own claim is that on four agent benchmarks Jev Router solved 237 of 423 tasks against 130 for its Auto Router, 82% more. That is the vendor's benchmark, not an independent one: run your own tasks and compare cost per solved task, not tasks solved alone. Third-party write-ups describe light (GPT-6 Luna, DeepSeek V4.1 Flash), standard (GPT-6 Sol, Claude Sonnet 5) and heavy (Claude Opus 5.5, GPT-6 Astra) tiers behind it; treat that as indicative, because the routing metadata on each response is the only record of what you actually paid for.
The Cost a Router Can Add: Losing the Cache
Prompt caching only pays while a conversation stays on one model. A router that switches models mid-session makes the new model read the whole history again at its uncached input rate.
Worked example, a 200K-token conversation:
- Switching to Claude Opus 5.5 re-reads 0.2M tokens at $4 per 1M: $0.80 of input per switch.
- Switching to GPT-6.1 Sol re-reads it at $2 per 1M: $0.40 per switch.
- Staying on GPT-6.1 Sol with the history cached reads it at $0.10 per 1M: $0.02.
One switch to Sol costs 20 times what staying cached costs, and five switches to Opus 5.5 in one long session add $4.00 of input before any output. That is why a router that changes effort before it changes model matters, and why a naive per-turn router on long agent sessions can cost more than sending everything to one mid-tier model.
OpenRouter Auto Router Now Takes a cost_tier
openrouter/auto now accepts a cost_tier parameter with five bands, from low to max. It caps how expensive a model Auto Router may pick, which is the price control the original Auto Router lacked. Like Jev Router it is listed at a variable price (you pay the routed model), it has a 2M context window, and it is not available on the us. and eu. regional domains.
Which to try:
- Auto Router with a
cost_tierwhen you want a hard ceiling on the price band, or need the 2M context. - Jev Router when sessions are long and cached, because it is built to avoid switches that throw the cache away; handle its fail-not-fallback behaviour in your client.
- Your own rules when you must stay in-region, since neither router runs on the regional domains.
Our Jev Router vs OpenRouter Auto Router comparison sets the two side by side.
Jev as a Cheap Classifier or Verifier
The LLM-based routing step above pays a chat model to classify. TypeSafe's Jev (typesafe/jev-1.13, launched 15 September 2026) is a decision model built for that job and does not write text. You send state (a message, a document or a pending tool call) plus typed questions: noul for yes/no, choice to pick one option, score for a number in a range. You get typed answers your code can branch on, with per-option probabilities and a confidence value on choice and score.
Price on OpenRouter: $0.042 per 1M input tokens, output free, a 32,000-token window for state and questions together, and a median round trip of about 0.24 seconds. Any OpenRouter key works.
Worked comparison, 1M routing decisions a month at 300 input tokens each:
- GPT-6 Luna as the classifier with 20 output tokens per answer: 300M input × $0.10 per 1M = $30, plus 20M output × $0.50 per 1M = $10, total $40.
- Jev: 300M input × $0.042 per 1M = $12.60, output free.
At 500 tokens of state, 1M Jev decisions cost $21.00, about $22.16 after OpenRouter's 5.5% credit fee or $22.68 on the 8% Business plan.
Two of OpenRouter's Jev cookbooks fit routing directly:
- Gate: ask Jev whether a request needs the expensive model before calling it.
- Verified cascade: a cheap model drafts, Jev checks the draft against the context, and you escalate only when the check fails. That stands in for the validator in the fallback pattern above on tasks with no compiler or schema to check against.
Limits: Jev reads instructions literally, so keep arithmetic and date comparisons in your own code and send only the state the question needs. It returns no explanations; pair it with a chat model if you need prose.
Key Takeaways
- →Model routing sends each request to the cheapest model that passes your quality bar
- →Rule-based routing is the simplest starting point — no ML required — and covers most of the saving
- →Typical savings are 60–80%: moving 70% of 10M input tokens from GPT-6.1 Sol ($2 per 1M) to GPT-6 Luna ($0.10) cuts that line from $20 to $6.70
- →Start with two tiers: a cheap default and an expensive escalation, and measure the rejection rate
- →For agents, route per step, cache the shared prefix and cap every run's steps and tokens
- →Managed routers bill the routed model's rate: Jev Router lists a $0 router fee and is built to avoid switches that lose the cache, which costs $0.80 per 200K-token re-read on Claude Opus 5.5
- →Jev at $0.042 per 1M input with free output is a cheaper classifier than a chat model: $12.60 against $40 on GPT-6 Luna for 1M decisions at 300 tokens
Editorial context
Who is this for?
Developers, startups, and teams who want to reduce their AI API or subscription costs without sacrificing quality.
When NOT to use this
Users who need real-time data, image generation, or proprietary enterprise integrations may need more specialised tools.
Pricing insights
AI pricing varies widely — some models charge per token while others use flat subscriptions. Token-based APIs are usually cheaper for moderate usage, while subscriptions suit power users with high and consistent volume.
Alternatives to consider
Consider DeepSeek V4 Flash for cost-effective coding and writing, Gemini 3.8 Flash for fast tasks, or Claude Haiku 4.5 for lightweight structured work. Use the calculator to compare your specific usage.
Final verdict
The cheapest AI tool is the one that fits your exact workload. Use the cost calculator and decision engine on this site to find your optimal stack — most users can cut AI spend by 50% or more.
Frequently Asked Questions
What is AI model routing?
Sending each request to the cheapest model that can handle it — for example DeepSeek V4 Flash at $0.05 per 1M input for classification and Claude Sonnet 5 at $2 for complex synthesis — instead of sending everything to one flagship.
How much does model routing save?
Typically 60–80% of API spend. The saving comes from the price spread: DeepSeek V4 Flash is 50× cheaper than GPT-5.4 on input and GPT-5.4 mini is 70% cheaper, so every request that passes review on the cheap tier is money saved.
Does routing hurt quality?
Not if you measure it. Route only tasks with a checkable outcome (JSON that parses, tests that pass, a confidence score) and escalate on failure. A rule with a high rejection rate costs more than no rule at all.
Does Jev Router charge a routing fee?
OpenRouter lists Jev Router at $0 for prompt and completion; you pay the rate of whichever model it routes each turn to, plus OpenRouter's usual credit fee. Where it can save money is the cache: it avoids switching models mid-session, and re-reading a 200K-token conversation uncached costs $0.80 on Claude Opus 5.5 or $0.40 on GPT-6.1 Sol per switch.
What happens if Jev Router cannot make a routing decision?
The request fails. If the Jev call times out or returns invalid output, Jev Router does not fall back to another router, so your client should retry or send the request to a default model itself. Jev Router and OpenRouter's Auto Router are also unavailable on the us. and eu. regional domains.
Related
Free courses · no sign-up
Still deciding? Learn the basics first, then come back to the prices.