How to Cut LLM API Costs: Caching, Model Routing and Batching
Short answer
To cut LLM API costs, first measure cost per task and find the biggest drivers. Then apply, roughly in order of impact: prompt caching for repeated context, routing simple requests to smaller models, batch APIs for non-urgent work, trimming prompts and retrieved context, limiting output length, reducing agent steps, and semantic caching for repeated questions. Check every change against an evaluation set so savings don't come at the cost of quality.
Key takeaways
- Measure cost per task and per feature before optimising anything.
- Prompt caching and model routing usually deliver the biggest savings.
- Batch APIs are much cheaper for work that doesn't need an instant answer.
- Always verify quality with evals after a cost change.
LLM costs rarely explode overnight. They creep. A system prompt grows by a few hundred tokens with each new rule. Retrieval sends ten chunks instead of five. The agent gains two more tools and starts making six calls per task instead of three. A feature launches and traffic doubles.
Then the monthly invoice arrives.
The good news: LLM costs are among the most controllable costs in software, because most of them come from choices you make in code. This guide covers the techniques we use, in the order we usually apply them.
Where do LLM API costs come from?
For most providers, you pay per token, with separate prices for input and output tokens (output is usually several times more expensive). So the cost of a task is:
(input tokens × input price) + (output tokens × output price), summed over every model call in the task.
That gives you four levers:
- Fewer calls per task
- Fewer input tokens per call
- Fewer output tokens per call
- Lower price per token — cheaper models, caching discounts, batch discounts
Step 1: Measure cost per task before optimising
You can't fix what you can't see. Before changing anything, instrument:
- Tokens and cost per model call, tagged with feature, user or account, and prompt version
- Calls per task (especially for agents)
- Cost per successful outcome — per resolved ticket, processed invoice, answered question
This usually reveals that a few features, prompts or customers drive most of the spend. Our guide to LLM observability covers how to set it up.
Step 2: Use prompt caching for repeated context
Many applications send the same large prefix on every request: a long system prompt, tool definitions, a policy document, product catalogue or few-shot examples. Prompt caching lets the provider reuse its processing of that prefix across requests.
At the time of writing:
- Anthropic bills cache reads at a fraction (around 10%) of the normal input price, with a small premium to write to the cache.
- OpenAI applies caching automatically for longer prompts with repeated prefixes and discounts cached input tokens.
- Google Gemini offers both implicit and explicit context caching.
Check each provider's current pricing pages, as details change.
How to get the most from caching:
- Put stable content first (system prompt, tools, reference documents) and variable content (user question, retrieved chunks) last
- Keep the stable prefix byte-for-byte identical — timestamps or user names in the system prompt break caching
- Monitor cache hit rate as a metric
For applications with long, repeated context, caching is often the single biggest saving, and it reduces latency too.
Step 3: Route requests to the right model
Using a frontier model for every request is like sending every package by overnight courier. Many tasks — classification, extraction, routing, short summaries, simple Q&A — work just as well on small, fast models that cost a fraction as much.
Model routing options:
- By task: each step in your pipeline uses a fixed model chosen in evaluation (e.g. small model for intent classification, frontier model for final reasoning)
- By difficulty: a classifier or rules decide whether a request is simple or complex
- Cascade: try the small model first; if validation fails or confidence is low, escalate to a larger model
Evaluate each routing decision against your test set (how to evaluate LLM applications). Our guide to choosing an LLM explains how to compare models on your own tasks.
Step 4: Use batch APIs for non-urgent work
Not everything needs an answer in two seconds. Evaluations, document backlogs, nightly enrichment, report generation and bulk classification can wait.
OpenAI's Batch API and Anthropic's Message Batches API both price batch requests at roughly 50% of the standard rate, with results returned asynchronously (typically within 24 hours). Google offers batch modes as well. Combined with caching, the savings stack.
Step 5: Send fewer input tokens
- Trim system prompts. Remove duplicated instructions, outdated rules and verbose examples. Test that quality holds.
- Retrieve less, retrieve better. A reranker lets you send 5 highly relevant chunks instead of 15 mediocre ones — cheaper and more accurate.
- Summarise long conversation history instead of resending every turn.
- Send only the tools an agent needs for the current step; tool definitions count as input tokens.
- Strip boilerplate from documents: headers, footers, signatures, HTML.
Step 6: Generate fewer output tokens
Output tokens are the most expensive per unit.
- Set sensible max token limits
- Ask for concise formats: structured JSON fields instead of prose when the output feeds code
- Avoid asking the model to repeat the input back
- For reasoning models, set reasoning effort or thinking budgets to match the task's difficulty
Step 7: Reduce agent steps
Agents can be the biggest cost multiplier, because each step is a model call carrying the full context.
- Deterministic steps in code: if the next step is always the same, don't ask the model to decide it
- Better tools: one tool that returns everything needed beats three calls that each return a fragment
- Step and budget limits per task (see AI agent guardrails)
- Parallel tool calls where supported, to cut both latency and repeated context
Step 8: Cache answers to repeated questions
For assistants that get the same questions repeatedly ("what's your refund policy?"), a semantic cache stores past answers and returns them when a new question is similar enough. It removes the model call entirely.
Use it carefully: answers must still respect permissions and be invalidated when source documents change.
Step 9: Control retries and failures
Retries can quietly double costs. Use exponential backoff, cap retries, and fix the root causes of validation failures rather than retrying until the model gets lucky. When one provider is rate-limiting, fail over to another instead of hammering it.
Should you self-host an open-source model?
Self-hosting open-weight models (Llama, Mistral, Qwen and others) can lower per-token costs at high, steady volume, and helps when data must stay in your environment. But include the full cost: GPUs (often idle at night), engineering, scaling, monitoring and upgrades. At low or spiky volume, hosted APIs usually win.
A middle path: use managed open-model hosting providers, or self-host only a small model for one high-volume task.
Which cost optimisations have the most impact?
| Technique | Typical impact | Effort | Quality risk |
|---|---|---|---|
| Prompt caching | High for long repeated context | Low | None |
| Model routing | High | Medium | Medium — evaluate |
| Batch APIs | ~50% on eligible work | Low | None |
| Trimming prompts & context | Medium | Low–medium | Low — evaluate |
| Output limits | Medium | Low | Low |
| Fewer agent steps | High for agents | Medium | Low–medium |
| Semantic caching | High for repetitive questions | Medium | Medium — invalidation |
| Self-hosting | Varies | High | Varies |
How we keep LLM costs predictable at Keyved
Cost is a design input on every project, not an afterthought. Before launch we project monthly cost at your expected volume, and our platform foundation tracks cost per request, per feature and per customer from day one, with budgets and alerts.
Most of the systems we build use caching-friendly prompt structures, multi-model routing and provider failover by default. If your AI bill is growing faster than your usage, talk to us — a cost review often pays for itself within the first month. You can also see what typical builds cost in our AI agent cost breakdown.
Frequently asked questions
Why are my LLM API costs so high?
Common causes are long system prompts and retrieved context sent on every request, using a frontier model for simple tasks, agents making many calls per task, verbose outputs, retries, and a small number of heavy users or features driving most of the spend.
What is prompt caching?
Prompt caching lets a provider reuse the processed form of a repeated prompt prefix, such as a long system prompt or document, across requests. Cached input tokens are billed at a substantial discount and processed faster. Anthropic, OpenAI and Google all offer forms of it.
What is model routing?
Model routing sends each request to the cheapest model that can handle it well. Simple tasks like classification or extraction go to small, fast models; complex reasoning goes to a frontier model. A rules-based or classifier-based router makes the choice.
How much can batch APIs save?
OpenAI's Batch API and Anthropic's Message Batches API both price batch requests at roughly half the standard rate, in exchange for results arriving asynchronously, typically within 24 hours. They suit evaluations, document backlogs, enrichment and reports.
Is it cheaper to self-host an open-source model?
Sometimes, at high and steady volume, or when data must stay on your infrastructure. At low or spiky volume, hosted APIs are usually cheaper once you include GPU costs, engineering time and operations.