AI cost optimization: get a grip on your LLM bill before it surprises you

What starts as a proof of concept costing a few tens of euros can, within a few months, grow into monthly bills of thousands of euros from OpenAI, Anthropic or Google. We help engineering and finance teams make those LLM costs transparent, reduce them technically and keep them predictable, without compromising the quality of your AI product.

Prompt caching Model routing Batch API Semantic caching Observability FinOps for AI
Request a cost review View strategies
LLM cost / day orange = baseline green = optimized Cache hit-rate 76% hits — input tokens 90% goedkoper Routed to Haiku/4o-mini Batch API queue 50% korting · 24h SLA

Why your LLM bill grows faster than your product adoption

On the surface, large language model pricing looks straightforward: a rate per million input tokens and a rate per million output tokens. In practice, costs still spiral because tokens accumulate in places engineers don't watch closely: system prompts sent with every call, retrieval context that grows without limit, agentic loops that call the model repeatedly and chat histories that balloon to tens of thousands of tokens.

A chatbot handling ten thousand conversations a day, with an average context of five thousand tokens and a top-tier model such as GPT-4o or Claude Sonnet as the only option, quickly exceeds a thousand euros a day. An agentic system making ten to twenty model calls per user action, for planning, tool selection, observation and reflection, multiplies that bill by a further factor of ten. Anyone who never measures a baseline only discovers what happened at the end of the month.

AI cost optimisation is not a one-off technical fix but an ongoing FinOps process for large language models: measuring, attributing costs by feature and team, optimising where most of the spend occurs and continuing to monitor as your volume grows. The right tool or technique depends on your workload; a batch pipeline needs something different from a real-time copilot.

The seven levers that genuinely reduce your LLM costs

Not every saving tactic works for every workload. Below are the technical levers we focus on, ranked by typical impact where they suit your use case.

💾

Prompt caching

Anthropic and OpenAI both support prompt caching: with Claude, a cache hit gives up to 90% off input tokens, while with OpenAI it is roughly 50%. It works best for long system prompts, static context and RAG fragments that stay the same across multiple requests. We structure your prompts so the cacheable part comes first and make sure the TTL matches your traffic pattern.

🔀

Model routing

Most requests in a production pipeline are simple: classification, extraction, summarisation. Route those to a cheaper model such as Claude Haiku, GPT-4o-mini or Gemini Flash, and send only the complex reasoning tasks to Sonnet, GPT-4o or Opus. An LLM-as-router pattern, or dedicated routers such as Martian, Portkey or Eden AI, decides per request which model is the best fit.

📦

Batch API

For anything that doesn't need a real-time response, such as embedding pipelines, document classification, nightly enrichment and dataset cleaning, OpenAI's Batch API offers a 50% discount with a 24-hour SLA, and Anthropic offers a similar Message Batches API. We split your workload into real-time and asynchronous layers and move everything that can be batched, with cron scheduling and resume logic.

🧠

Semantic caching

In chatbots and search interfaces, many user questions overlap semantically, even when the exact wording differs. A vector cache (Redis with embeddings, GPTCache or Portkey's semantic cache) matches similar queries and serves the existing answer. In support bots we regularly save 30 to 60 per cent of model calls this way, provided the relevance threshold is set correctly.

✂

Context pruning and structured output

Long chat histories and RAG results full of irrelevant material are a silent cost killer. We trim the context down to what the model actually needs through summary roll-ups, top-k re-ranking and token budgets per conversation. Using JSON schemas or Anthropic tool use, we enforce shorter, structured output, rather than free-form prose where a list will do.

📚

RAG instead of long context

You can put a million tokens into the context window, but it is expensive and slow. For knowledge questions, a well-built RAG pipeline (chunking, hybrid search, re-ranking) is almost always cheaper and more accurate than sending the entire document along. We build retrieval layers with pgvector, Qdrant or Elastic and evaluate them with RAGAS-style frameworks.

🎯

Distillation and fine-tuning

For recurring tasks with enough examples, we train a smaller model that mimics the behaviour of the large model, known as knowledge distillation. A fine-tuned Llama 3.1 8B or Mistral Small can run specific domains at a fraction of the cost, whether you go through Together AI, Fireworks or self-hosting.

📉

Quantisation for on-premises

If you host models yourself, for compliance reasons or at sufficient volume, quantisation (AWQ, GPTQ, GGUF, FP8) lets you run the same model on smaller GPUs. A 70B model that normally needs two A100s can, once quantised, fit on a single H100 or even on consumer GPUs. We calculate where the break-even point between API costs and self-hosting lies for your traffic.

📰

Observability and alerting

Without measurement, there is no optimisation. We roll out Helicone, Langfuse, OpenLLMetry or Vellum as a gateway for your LLM calls. For each feature, user and model, we track costs, latency, cache ratio and errors. Budget alerts and anomaly detection prevent a loop in production from costing you a thousand euros within an hour.

Our approach: from raw invoice to predictable unit cost

We work in four phases, and after each one you receive a concrete picture of the savings and risks. No lengthy project without interim results: your bill starts falling during the scan itself.

Cost audit

We link your provider billing and analyse where tokens are actually consumed — per feature, per endpoint, per user. Often twenty per cent of calls cause eighty per cent of the costs. The audit report identifies the hot spots concretely.

Quick wins

Within two to three weeks, we implement the low-hanging fruit optimisations: switching on prompt caching, shortening system prompts, and moving work into batches. These usually deliver savings of twenty to forty percent without changing the product.

Architectural interventions

After that, we tackle model routing, semantic caching, RAG redesign and, where appropriate, fine-tuning. This is where the structural savings lie. We build, A/B test and measure quality before and after, so you can be confident your output stays at the same level.

FinOps loop

We install observability, dashboards and budget alerts, and hand over the operation to your team. Each month or quarter, we review together to incorporate new models, lower prices and growing traffic.

Tooling we work with

A mature AI cost stack consists of three layers: a gateway layer for caching, routing and observability; an evaluation layer to keep quality in view during optimisation; and an orchestration layer for batch jobs and agent loops. We combine best-of-breed open source with provider-native features where that is scalable and cheaper.

The choice depends on your stack. If you work on AWS Bedrock, we use cross-region inference profiles and provisioned throughput for fixed pricing. On Azure OpenAI, we use Provisioned Throughput Units (PTUs) where volume justifies it. With direct Anthropic or OpenAI keys, we rely more heavily on gateways such as Portkey or LiteLLM. On-premise or in your own VPC, we work with vLLM, TGI or Ollama behind an internal gateway.

Helicone Langfuse OpenLLMetry Vellum Portkey LiteLLM Martian Eden AI GPTCache Redis Vector pgvector Qdrant vLLM Together AI Fireworks AWS Bedrock Azure OpenAI Anthropic Batch OpenAI Batch RAGAS
Not yet sure about a large project?

Test your idea first: a working prototype in 1 day

With OneDayBuild, we turn your idea into something tangible in one day for €1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.

Explore OneDayBuild →

Three typical workloads and what can be saved on them

The right optimisation depends on how your application behaves. Below are three scenarios we often encounter, with the tactics that deliver the most for each type.

💬

Chatbot / customer support copilot

Many users, recurring intents, long system prompts containing brand guidelines. The biggest levers are prompt caching for the system portion, semantic caching on the most frequently asked questions, a cheap default model with escalation to a more expensive one for intent classification, and tight chat history pruning. Achievable savings are often thirty to sixty percent.

📄

Document pipeline / batch extraction

Processing tens of thousands of invoices, contracts or emails per day. The Batch API delivers a fifty percent discount straight away; embedding models for pre-classification keep expensive LLM calls away from trivial documents; structured output via JSON schema prevents rework. For high-volume tasks, a fine-tuned smaller model is often the final step.

🤖

Agent / multi-step orchestrator

Ten to thirty model calls per user action for planning, tool selection and reflection. Model routing per step works very well here: a cheap model for planning and tool selection, a more expensive one for the final reasoning. Caching on tool descriptions and strict loop limits with budget alerts prevent runaway costs.

What an audit looks like

An example set-up for the first week. We map every feature to costs, routing and optimisation potential — concretely and with figures.

FeatureCurrent modelProposed optimisationExpected impact
Support chatbotClaude Sonnet on every callPrompt caching + Haiku for classificationInput tokens 90% cheaper, 70% of routes to Haiku
Nightly document taggingGPT-4o real-timeOpenAI Batch + GPT-4o-mini50% batch discount + 15x lower per-token price
RAG knowledge base200K context per queryHybrid search + top-8 re-rankingContext reduced from 200K to an average of 4K tokens
Agent planner loopSonnet for every stepHaiku for planning, Sonnet for synthesis~60% lower cost per agent trace
Embedding pipelinetext-embedding-3-large3-small + dimension reduction to 1024around 60% lower embedding costs

The figures in the table are typical orders of magnitude based on the public pricing of providers (Anthropic, OpenAI). Your exact saving depends on traffic patterns, quality requirements and latency budgets, which we will determine together during the audit phase.

Why choose Appfront for AI cost optimisation

We are a product development team, not a pure consultancy. Our recommendations can always be delivered in code, not just in slides.

Engineering firstIf you wish, we build the optimisations ourselves: gateways, caching layers, routers, observability. Rather than handing you a report and walking away.
Provider-agnosticWe work with Anthropic, OpenAI, Google, Mistral, Cohere, AWS Bedrock and Azure OpenAI. Lock-in is not the goal; the right mix is.
FinOps for AIWe bring the tagging, attribution and alerting discipline of classic cloud FinOps to your LLM stack, including showback per team or feature.
Quality gatesAn optimisation that lowers the quality of your output is not an optimisation. We measure with evals and regression suites before and after every change.
EU hosting where neededFor clients with GDPR requirements, we route through EU regions of Bedrock or Azure, or through self-hosted vLLM clusters. Compliance always takes precedence over any saving.
Concrete unit economicsWe don't report "lower costs" in the abstract. We report cost per conversation, per ticket or per processed document, so you can justify your product pricing.

Frequently asked questions about AI cost optimisation

How much can I realistically save on my LLM costs?
For most teams that have never optimised in a structured way, a saving of thirty to seventy per cent is within reach, depending on the workload. Quick wins such as prompt caching and model routing typically deliver twenty to forty per cent within two to four weeks. Architectural changes (RAG redesign, distillation, batch migration) then add a further substantial step. We determine the actual headroom during the audit.
Does prompt caching also work for my use case?
Prompt caching delivers the most value when you have long, reusable context, such as an extensive system prompt, fixed tool definitions or static RAG fragments, that stays the same across many requests. For chatbots, copilots and agent systems it is almost always worthwhile. For one-off calls with constantly new content, much less so. Anthropic offers up to ninety per cent discount on cache hits for input tokens, and OpenAI roughly fifty per cent. During the audit we measure how much cacheable content you have and which TTL strategy suits it.
What exactly is model routing, and why is it so effective?
Model routing means not every request goes to your most expensive model. A simple classification or intent detection task can be handled just as well by Claude Haiku, GPT-4o-mini or Gemini Flash, at a fraction of the price. Only complex reasoning or synthesis tasks are sent to Sonnet, GPT-4o or Opus. We build routing either as rule-based (by intent, length or user type) or through an LLM-as-router pattern, where a cheap model makes the decision itself. Tools such as Martian, Portkey and Eden AI offer this as a managed service.
When are OpenAI Batch or Anthropic Message Batches suitable?
For anything that does not need to happen within seconds. OpenAI Batch gives fifty per cent off with a twenty-four-hour SLA, and Anthropic offers something comparable. Typical use cases include nightly document tagging, embedding runs, dataset cleaning, email classification and bulk enrichment. We build a wrapper that splits jobs, handles retries and delivers results back into your data warehouse or database.
When does fine-tuning or running your own open-source model pay off?
Fine-tuning pays off once a specific task has high volume, enough training data is available and variation between requests is limited. For classification, extraction or style-consistent writing tasks, a fine-tuned Llama 3.1 8B or Mistral Small running on Together AI or Fireworks can often cost an order of magnitude less than a frontier API. Self-hosting with vLLM on your own GPUs becomes relevant at very high volumes or with strict compliance requirements, and that is where we calculate the break-even point.
What is semantic caching, and how does it differ from prompt caching?
Prompt caching works on an exact token-prefix match and sits with the provider. Semantic caching works on your side and matches on meaning: two differently worded but similar questions receive the same cached answer. We do this through a vector store with a sensitivity threshold (cosine similarity). It works particularly well for support bots and knowledge-base FAQs. We tune the threshold and build invalidation rules so that outdated answers don't linger.
How do you ensure quality does not decline after optimisation?
For every intervention, we set up evals: a fixed set of representative inputs with expected outcomes, plus an LLM-as-judge or human spot-check for subjective criteria. We run that suite before and after every change and reject the change if a key score drops. We use frameworks such as RAGAS, Promptfoo and Langfuse evals as standard. Quality always comes before cost: a cheap chatbot that hallucinates will cost you more in support tickets than you save.
Does this also work with GDPR-sensitive data and EU hosting?
Yes. We route through EU regions of AWS Bedrock, Azure OpenAI or Google Vertex, or we host ourselves with vLLM, TGI or Ollama in an EU cluster. For clients in finance, healthcare or government, we build the gateway layer inside your own VPC, with logs that never leave the European mainland. Optimisations such as caching, routing and distillation work at least as well in a self-hosted setup, often better, because you have more control over the stack.

Want to get a grip on your LLM costs?

Send us your provider billing or a week's worth of logs. We will carry out an initial analysis and show you where the greatest savings lie, free of charge and without obligation.

Book a cost scan

Edit content