AI cost optimization: get a grip on your LLM bill before it surprises you
What starts as a proof of concept costing a few tens of euros can, within a few months, grow into monthly bills of thousands of euros from OpenAI, Anthropic or Google. We help engineering and finance teams make those LLM costs transparent, reduce them technically and keep them predictable, without compromising the quality of your AI product.
Request a cost review View strategiesWhy your LLM bill grows faster than your product adoption
On the surface, large language model pricing looks straightforward: a rate per million input tokens and a rate per million output tokens. In practice, costs still spiral because tokens accumulate in places engineers don't watch closely: system prompts sent with every call, retrieval context that grows without limit, agentic loops that call the model repeatedly and chat histories that balloon to tens of thousands of tokens.
A chatbot handling ten thousand conversations a day, with an average context of five thousand tokens and a top-tier model such as GPT-4o or Claude Sonnet as the only option, quickly exceeds a thousand euros a day. An agentic system making ten to twenty model calls per user action, for planning, tool selection, observation and reflection, multiplies that bill by a further factor of ten. Anyone who never measures a baseline only discovers what happened at the end of the month.
AI cost optimisation is not a one-off technical fix but an ongoing FinOps process for large language models: measuring, attributing costs by feature and team, optimising where most of the spend occurs and continuing to monitor as your volume grows. The right tool or technique depends on your workload; a batch pipeline needs something different from a real-time copilot.
The seven levers that genuinely reduce your LLM costs
Not every saving tactic works for every workload. Below are the technical levers we focus on, ranked by typical impact where they suit your use case.
Prompt caching
Anthropic and OpenAI both support prompt caching: with Claude, a cache hit gives up to 90% off input tokens, while with OpenAI it is roughly 50%. It works best for long system prompts, static context and RAG fragments that stay the same across multiple requests. We structure your prompts so the cacheable part comes first and make sure the TTL matches your traffic pattern.
Model routing
Most requests in a production pipeline are simple: classification, extraction, summarisation. Route those to a cheaper model such as Claude Haiku, GPT-4o-mini or Gemini Flash, and send only the complex reasoning tasks to Sonnet, GPT-4o or Opus. An LLM-as-router pattern, or dedicated routers such as Martian, Portkey or Eden AI, decides per request which model is the best fit.
Batch API
For anything that doesn't need a real-time response, such as embedding pipelines, document classification, nightly enrichment and dataset cleaning, OpenAI's Batch API offers a 50% discount with a 24-hour SLA, and Anthropic offers a similar Message Batches API. We split your workload into real-time and asynchronous layers and move everything that can be batched, with cron scheduling and resume logic.
Semantic caching
In chatbots and search interfaces, many user questions overlap semantically, even when the exact wording differs. A vector cache (Redis with embeddings, GPTCache or Portkey's semantic cache) matches similar queries and serves the existing answer. In support bots we regularly save 30 to 60 per cent of model calls this way, provided the relevance threshold is set correctly.
Context pruning and structured output
Long chat histories and RAG results full of irrelevant material are a silent cost killer. We trim the context down to what the model actually needs through summary roll-ups, top-k re-ranking and token budgets per conversation. Using JSON schemas or Anthropic tool use, we enforce shorter, structured output, rather than free-form prose where a list will do.
RAG instead of long context
You can put a million tokens into the context window, but it is expensive and slow. For knowledge questions, a well-built RAG pipeline (chunking, hybrid search, re-ranking) is almost always cheaper and more accurate than sending the entire document along. We build retrieval layers with pgvector, Qdrant or Elastic and evaluate them with RAGAS-style frameworks.
Distillation and fine-tuning
For recurring tasks with enough examples, we train a smaller model that mimics the behaviour of the large model, known as knowledge distillation. A fine-tuned Llama 3.1 8B or Mistral Small can run specific domains at a fraction of the cost, whether you go through Together AI, Fireworks or self-hosting.
Quantisation for on-premises
If you host models yourself, for compliance reasons or at sufficient volume, quantisation (AWQ, GPTQ, GGUF, FP8) lets you run the same model on smaller GPUs. A 70B model that normally needs two A100s can, once quantised, fit on a single H100 or even on consumer GPUs. We calculate where the break-even point between API costs and self-hosting lies for your traffic.
Observability and alerting
Without measurement, there is no optimisation. We roll out Helicone, Langfuse, OpenLLMetry or Vellum as a gateway for your LLM calls. For each feature, user and model, we track costs, latency, cache ratio and errors. Budget alerts and anomaly detection prevent a loop in production from costing you a thousand euros within an hour.
Our approach: from raw invoice to predictable unit cost
We work in four phases, and after each one you receive a concrete picture of the savings and risks. No lengthy project without interim results: your bill starts falling during the scan itself.
Cost audit
We link your provider billing and analyse where tokens are actually consumed — per feature, per endpoint, per user. Often twenty per cent of calls cause eighty per cent of the costs. The audit report identifies the hot spots concretely.
Quick wins
Within two to three weeks, we implement the low-hanging fruit optimisations: switching on prompt caching, shortening system prompts, and moving work into batches. These usually deliver savings of twenty to forty percent without changing the product.
Architectural interventions
After that, we tackle model routing, semantic caching, RAG redesign and, where appropriate, fine-tuning. This is where the structural savings lie. We build, A/B test and measure quality before and after, so you can be confident your output stays at the same level.
FinOps loop
We install observability, dashboards and budget alerts, and hand over the operation to your team. Each month or quarter, we review together to incorporate new models, lower prices and growing traffic.
Tooling we work with
A mature AI cost stack consists of three layers: a gateway layer for caching, routing and observability; an evaluation layer to keep quality in view during optimisation; and an orchestration layer for batch jobs and agent loops. We combine best-of-breed open source with provider-native features where that is scalable and cheaper.
The choice depends on your stack. If you work on AWS Bedrock, we use cross-region inference profiles and provisioned throughput for fixed pricing. On Azure OpenAI, we use Provisioned Throughput Units (PTUs) where volume justifies it. With direct Anthropic or OpenAI keys, we rely more heavily on gateways such as Portkey or LiteLLM. On-premise or in your own VPC, we work with vLLM, TGI or Ollama behind an internal gateway.
Test your idea first: a working prototype in 1 day
With OneDayBuild, we turn your idea into something tangible in one day for €1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.
Explore OneDayBuild →Three typical workloads and what can be saved on them
The right optimisation depends on how your application behaves. Below are three scenarios we often encounter, with the tactics that deliver the most for each type.
Chatbot / customer support copilot
Many users, recurring intents, long system prompts containing brand guidelines. The biggest levers are prompt caching for the system portion, semantic caching on the most frequently asked questions, a cheap default model with escalation to a more expensive one for intent classification, and tight chat history pruning. Achievable savings are often thirty to sixty percent.
Document pipeline / batch extraction
Processing tens of thousands of invoices, contracts or emails per day. The Batch API delivers a fifty percent discount straight away; embedding models for pre-classification keep expensive LLM calls away from trivial documents; structured output via JSON schema prevents rework. For high-volume tasks, a fine-tuned smaller model is often the final step.
Agent / multi-step orchestrator
Ten to thirty model calls per user action for planning, tool selection and reflection. Model routing per step works very well here: a cheap model for planning and tool selection, a more expensive one for the final reasoning. Caching on tool descriptions and strict loop limits with budget alerts prevent runaway costs.
What an audit looks like
An example set-up for the first week. We map every feature to costs, routing and optimisation potential — concretely and with figures.
| Feature | Current model | Proposed optimisation | Expected impact |
|---|---|---|---|
| Support chatbot | Claude Sonnet on every call | Prompt caching + Haiku for classification | Input tokens 90% cheaper, 70% of routes to Haiku |
| Nightly document tagging | GPT-4o real-time | OpenAI Batch + GPT-4o-mini | 50% batch discount + 15x lower per-token price |
| RAG knowledge base | 200K context per query | Hybrid search + top-8 re-ranking | Context reduced from 200K to an average of 4K tokens |
| Agent planner loop | Sonnet for every step | Haiku for planning, Sonnet for synthesis | ~60% lower cost per agent trace |
| Embedding pipeline | text-embedding-3-large | 3-small + dimension reduction to 1024 | around 60% lower embedding costs |
The figures in the table are typical orders of magnitude based on the public pricing of providers (Anthropic, OpenAI). Your exact saving depends on traffic patterns, quality requirements and latency budgets, which we will determine together during the audit phase.
Why choose Appfront for AI cost optimisation
We are a product development team, not a pure consultancy. Our recommendations can always be delivered in code, not just in slides.
Frequently asked questions about AI cost optimisation
Want to get a grip on your LLM costs?
Send us your provider billing or a week's worth of logs. We will carry out an initial analysis and show you where the greatest savings lie, free of charge and without obligation.
Book a cost scan