One of the clearest pieces of evidence in recent systems work is also the least flashy: according to a 2025 systems study, a tested deployment achieved up to 40% lower operating cost and 30% lower energy use, with less than 1.5% BLEU degradation and a 5 to 7% latency increase. Lower spend is available, but only if quality is measured against the work the business actually sends through the system, not against a benchmark that flatters the model.
The larger point is not that small models are better than large ones. It is that most organizations are paying large-model prices for a workload that is a mix of lookup, extraction, classification, summarization, and the occasional hard case. When RouteLLM results summarized in the available literature reported more than 85% lower cost on MT-Bench while retaining about 95% of GPT-4 quality, the savings fell to about 45% on MMLU and 35% on GSM8K, which is a warning as much as a promise. Routing works, but not the same way on every task.
This problem has lasted because the process is simpler when every request goes to one general model. The data needed for routing is usually missing, because organizations do not keep representative traces showing request type, difficulty, correctness, token use, and downstream business outcome. The incentives are also backward: token spend is visible, while rework, customer abandonment, and human review are spread across operations, support, and risk teams.
Most of the cost sits in the wrong place
The common pattern is a production system that treats every request as if it deserves the same answer path. That is wasteful for routine tasks and risky for difficult ones. According to NVIDIA research cited in the available results, a 7B model can be 10 to 30 times cheaper in latency, energy, and compute than a 70B to 175B model, but only if the 7B model is asked to do work it can actually finish.
This is why average benchmark quality is the wrong executive metric. The useful question is cost per successful outcome, with latency percentiles, escalation rate, and correction effort alongside it. A cheap answer that triggers a manual fix, a second call, or a compliance review is not cheap.
The same mistake shows up in architecture reviews. Teams optimize the language model and leave the rest untouched, even though poor retrieval can make a larger model seem necessary when the real issue is stale or missing context in the document store, ticketing system, or policy library. The model is often blamed for a systems problem.
Routing works only when the system can say no
The strongest production pattern is capability-based routing. A classifier, learned router, or rules engine predicts whether a small or specialized model is enough, and sends uncertain, sensitive, or high-impact requests to a larger model or a review queue. RouteLLM-style approaches reported large cost reductions while preserving most benchmark quality, but the point is not the specific benchmark. The point is that the router has to know when not to take the cheap path.
That means the routing logic needs more than one signal. Request type, language, input length, required tools, data sensitivity, user tier, and historical failure patterns all matter. So do hard rules for cases where the business cannot tolerate uncertainty, including policy conflicts, missing evidence, failed validation, and low confidence. If the system does not have a reliable escalation path, routing is just underprovisioning with a dashboard.
When routing is built properly, the faster decision is not only model selection. It is the decision about whether this request needs context from the records system, whether it can be answered from the knowledge base, whether it must be approved by a person, or whether it belongs in the exception queue. That is where the operating model changes: not fewer controls, but better placed controls. If you want a reference design for that layer, see Enterprise RAG Systems.
- 01Source systemsCollect the business and operational records that define the request and the answer.
- Application logs
- Document repositories
- Transactional systems
- 02Permissioned contextFilter and prepare authoritative evidence before any model sees it.
- Identity and access checks
- Metadata filters
- Retrieval index
- 03Routing decisionChoose the cheapest acceptable path for each request based on task, risk, and confidence.
- Rules engine
- Learned router
- Confidence thresholds
- 04Model executionUse a small model for routine work and escalate hard or sensitive cases to a larger model.
- Small language model
- Large language model
- Speculative decoding
- 05Validation and reviewCheck structure, policy, and evidence before output is accepted or sent to a person.
- Schema validator
- Human approval queue
- Retry and fallback logic
- 06Audit and learningRecord what happened and feed real outcomes back into evaluation and routing policy.
- Audit log
- Evaluation pipeline
- Cost and latency metrics
- Identity and access controls apply before retrieval and before generation.
- High-impact or low-confidence requests must escalate deterministically to a larger model or human review.
- Audit logs capture route, evidence, model, latency, token use, and outcome for every consequential response.
- Cost limits and evaluation thresholds prevent routing changes from increasing correction effort or risk.
| Capability | Azure | AWS | Google Cloud |
|---|---|---|---|
| Object and document storage | Equivalent managed service | Equivalent managed service | Equivalent managed service |
| Event streaming and workflow orchestration | Equivalent managed service | Equivalent managed service | Equivalent managed service |
| Vector retrieval and search | Equivalent managed service | Equivalent managed service | Equivalent managed service |
| Model hosting and inference | Equivalent managed service | Equivalent managed service | Equivalent managed service |
| Logging, monitoring, and audit | Equivalent managed service | Equivalent managed service | Equivalent managed service |
The parts that matter are data, thresholds, and fallback
A usable reference design starts upstream. Application logs, ticket histories, document repositories, transactional systems, and identity data feed an event stream and retrieval index. The router uses that context to decide whether a small model, a large model, or a human should handle the request. The model is only one component in a chain that also includes validation rules, audit logging, and cost limits.
That chain has to be evaluated against the organization’s own workload. Generic quality scores are not enough. The acceptance set should include factuality against authoritative records, schema validity, refusal behavior, tool-selection accuracy, and downstream task completion. If the request comes from a regulated workflow, the same response also needs traceability back to the source record, the route taken, and the evidence used.
According to the available research, quantization and serving optimizations can help, but they are not substitutes for routing discipline. The available results suggest that W8A8 or W4A16 can offer stronger accuracy-latency trade-offs on tested workloads, while very low bit widths can hurt reasoning quality. That means precision reduction belongs in the same governance conversation as model choice, because a cheaper model that fails on arithmetic, code, or long-context inputs simply shifts the bill elsewhere.
Where the approach pays off first
The best candidates are repetitive, bounded, and easy to verify. Intent classification, entity extraction, moderation, structured-field generation, document routing, and short support responses are good places to start because they have explicit outputs and measurable failure modes. Distillation can make those paths cheaper still, especially when the smaller model is trained to reproduce tool calls, schema validity, and refusal behavior from a stronger teacher model.
Retrieval matters just as much. A smaller model with authoritative context from policies, manuals, contracts, tickets, or records can outperform a larger model that is forced to guess from an incomplete prompt. But retrieval only helps when the source data is fresh, authorized, and relevant. Duplicated or stale chunks can increase hallucination and leakage rather than reduce them.
That is why the governance layer is not overhead. It is the condition that lets routing be used at scale without turning every exception into a crisis. Route changes, prompt changes, retrieval changes, and model changes should not be released together if the organization wants to know what broke. Version control, evaluation pipelines, access controls, and incident logging are part of the system, not optional paperwork. If that control stack is missing, AI Governance is not a separate program. It is the only way to make the routing program credible.
The tradeoff is real, and it is usually acceptable
The honest counterpoint is that routing adds machinery. It introduces another decision point, another failure mode, and another set of thresholds to tune. It also requires labeled traces, workload analysis, and a fallback path that can absorb the difficult cases without degrading service. For some organizations, that is enough to delay adoption, especially when the current large-model setup is good enough and the workload is not yet understood.
But the alternative is not simplicity. It is paying maximum-model costs for minimum-model work, while still accepting silent regressions because no one can tell whether a miss came from the model, the retrieval layer, the prompt, or the input data. The business pays either way. The difference is whether the bill shows up as infrastructure spend or as repeated manual correction.
That makes the executive decision straightforward. Do not ask whether small models can replace large ones. Ask which requests should never reach the large model, which requests must always escalate, and which ones can be resolved by a smaller model with evidence and guardrails. QueryNow builds that workflow in your environment in two weeks, and you pay $10,000 only after it meets the acceptance criteria you signed off on. Start at /build.
Ready to ship AI in your organization?
We build one workflow into a working tool in two weeks. You pay $10,000 only after every acceptance criterion you signed off on is met.
One workflow · Two-week build · $10,000, paid on delivery
QueryNow
QueryNow deploys production AI for enterprises on Azure, AWS, or Google Cloud. Founded in 2014, we help pharma, healthcare, manufacturing, and financial services organizations deploy governed AI systems. We build it, you pay when it works.
Learn more about us →


