Service · Software development

Custom LLM integrations in your own software.

A Large Language Model integrated directly into your application, workflow or customer portal. Classify, summarise, search or map free-text input to structured fields, without your data leaving to a training set.

Claude · GPT · GeminiPrivate LLMRAG & embeddingsEU data residency

An LLM is a building block, not a product.

A Large Language Model — Claude, GPT-4o, Gemini, Llama — is, in itself, simply an API call. The value only comes once such a model is genuinely part of your working process: it reads your documents, classifies your incoming emails, searches your knowledge base, or maps customer input to the right fields in your CRM. That is where we come in: the integration, the context, and the monitoring around it.

We build LLM integrations as part of a broader software project or as a standalone capability within an existing application. An AI agent is a specific form within that — a multi-step system that can call tools. An LLM integration is broader and often simpler: a single well-prompted call, an embedding search with RAG, or a fine-tuned classifier running in production.

Three flavours of LLM integration.

Depending on what the model needs to do, how sensitive the data is and how much custom work surrounds the model. We advise which fits during the first conversation, and always check first whether a lighter option will do.

Compact project · fixed sprint budget

Single-call LLM feature

A defined feature in your application where a single well-prompted LLM call does the work: classifying an incoming email, summarising a long document, automatically assigning a sentiment label to a review, or turning free-text customer input into structured fields in your CRM. Cloud LLM from Anthropic, OpenAI or Google. Includes monitoring of token usage and quality.

Prompt engineeringClassificationSummarisationForm mapping
Mid-sized project · fixed sprint budget

RAG and semantic search

The LLM receives context drawn from your own data. We build a retrieval pipeline over your documents, knowledge base or product catalogue: embeddings, vector store, semantic search, and a chat or Q&A layer that answers only from the sources it finds. Source citations as standard, hallucinations kept in check. Suited to internal knowledge bases, customer support and product search features.

EmbeddingsVector storeHybrid searchSource citations
Larger project · fixed sprint budget

Private LLM on your own infrastructure

An open-source model — Llama, Mistral, or a specialised variant — running on your own cloud or on-premises. Full data control, no calls to external APIs, no uncertainty about the training set. We deploy it via vLLM, Ollama or a managed inference layer, with the associated monitoring, scaling and authorisation. More upfront work, but for compliance-sensitive sectors often the only route.

Llama / MistralvLLM · OllamaGPU infrastructureOn-premises

What you get at the end.

A production-ready LLM integration, plus everything needed to manage and steer it yourself on cost, quality and privacy.

  • The LLM integration itselfIn production and staging, running in your own cloud (GCP, AWS, Azure) or with us. For private LLMs, including the inference stack.
  • Prompt library and evaluation setVersion-controlled prompts, an evaluation set with test examples, and a report showing how well the model performs your task.
  • Codebase and architecture documentationFull source code, build and deployment instructions, and an architecture overview explaining model choice, caching strategy and data flow.
  • Cost monitoring and token dashboardVisibility into tokens per call, cost per user or feature, and alerts when usage rises unexpectedly. Caching set up where it makes sense.
  • Compliance packageDPIA where relevant, GDPR mapping of data flows, contractual agreements with the model provider on data retention and training, and an overview of how your integration relates to the EU AI Act.
  • Maintenance contract (optional)Monitoring of model quality, prompt maintenance, ongoing development, and migration should you wish to move to a newer model. Fixed monthly fee.

When an LLM integration is the right choice.

Four patterns where we guide clients — if you recognise one, we'd be happy to talk further. Not every use case calls for an LLM; sometimes a classic rule engine or a specialised model is cheaper and more reliable.

Document flows

Incoming volume is too large to handle manually

Hundreds of invoices, contracts, emails or requests a day are currently read, sorted and forwarded by hand. An LLM can classify, summarise and extract the relevant fields — leaving your team time for the cases that genuinely need attention.

Unlocking knowledge

Your knowledge base can't be found

Employees or customers search through a growing pile of documents and fail to find what they need. Semantic search with RAG lifts search results from keyword level to meaning level, with source citations so you know where an answer comes from.

Processing free-text input

Customers type freely, your system expects fields

Lead forms, support requests, product specifications — users write in free text, while your CRM or ERP expects structured data. An LLM integration maps free text to the right fields, with validation and a human check where needed.

Compliance & privacy

Data may not go to a cloud API

Do you work with medical, legal or financial data where GDPR and sector regulations are strict? Then a private LLM on your own infrastructure is the right route — more expensive to set up, but without the privacy trade-offs of an external API.

Where we deploy LLMs in practice.

The applications we build most often. The common thread: a well-defined task, a measurable quality threshold and a logging layer that shows what the model did.

Documents

Classification and extraction

Automatically label incoming invoices, contracts, emails or requests and extract the relevant fields — amount, date, party, type, priority. For a high volume of incoming post, this demonstrably reduces manual work and delivers consistent data in your back office.

Text & language

Summaries and translations

Condense long documents into a few paragraphs, or produce domain-specific translations that understand legal or technical jargon. Far more reliable than generic translation APIs once you give the model context about your field.

Sentiment & review

Understanding reviews and support tickets

Automatically label customer reviews, NPS comments or support requests by sentiment, topic and urgency. Produces a dashboard that lets teams spot trends before complaints escalate.

Search

Semantic search with RAG

A search function that understands what a user means, not just which words appear. Combines keyword search with embeddings and returns answers with source citations. A building block for both internal knowledge bases and product search features.

Forms

Free-text input to structured data

A customer types what they are looking for in free text, and the LLM maps it to the right fields in your CRM, ERP or order system. Includes a human check for low confidence scores so errors don't reach production.

Product data

Generating product descriptions

For e-commerce catalogues with thousands of SKUs: consistent, SEO-friendly descriptions based on specifications. Tone of voice is safeguarded via a prompt library; people stay in the loop for sampling and final editing.

Code & SQL

Code and SQL generation within your own tooling

An LLM that generates queries or code snippets within your application, for example a reporting tool in which non-technical users can ask questions in plain language, which the model converts into safe SQL.

Privacy

Anonymisation and redaction

Stripping personal data from logs, support ticket archives or training data before they are used further. An LLM recognises patterns that regex misses: names, addresses and policy numbers in free text.

Models and infrastructure we work with.

No vendor lock-in to a single model provider. For each use case we choose what fits best in terms of quality, price, latency and compliance, and we make sure that switching to a newer model later does not mean a full rebuild.

Cloud LLMs

Anthropic, OpenAI, Google, Cohere

Anthropic Claude (Opus, Sonnet, Haiku) for nuanced language work and long context. OpenAI GPT-4o and o1 for broad use and complex reasoning. Google Gemini via Vertex AI for those already on Google Cloud. Azure OpenAI for those running a Microsoft stack. Cohere for specific embedding and classification work.

ClaudeGPT-4o · o1GeminiAzure OpenAICohere
Open source & private

Llama, Mistral, vLLM, Ollama

For scenarios where data must not leave the premises. We run Llama 3 (70B and smaller), Mistral variants and specialised fine-tunes via vLLM on GPU instances, or via Ollama for smaller models. This includes the operational layer: scaling, monitoring, authorisation and cost control.

Llama 3MistralvLLMOllamaGPU orchestration
Retrieval, evaluation & caching

Vector stores, evals, prompt caching

Postgres with pgvector, Qdrant or Pinecone as the vector store, depending on scale and your existing stack. Evaluation frameworks to detect regressions when models are updated. Anthropic's prompt caching and OpenAI's cached prompts are used where sensible; a caching strategy can significantly reduce token costs in production.

pgvectorQdrantPrompt cachingEval sets

How an LLM integration project runs.

1

Introduction and use-case scan

A conversation in which we establish which task the model needs to perform, what data is involved, how sensitive that data is and what quality is acceptable. This is also where the first answer to the cloud-versus-private question takes shape.

2

Prototyping and model selection

We test several models (Claude Sonnet, GPT-4o, Gemini or a Llama variant) on your own examples. At the end you receive a comparison on quality, latency and cost, plus advice on which model suits which flow best.

3

Building in sprints

A working build every two weeks. We build the integration into your stack, set up an evaluation suite, implement caching and cost monitoring, and configure logging so that we can prove which prompt version produced which answer. This includes any integrations with your existing systems.

4

Rollout, evaluation and maintenance

A phased rollout, initially to a small group of users so that unwanted behaviour is caught before it scales. Followed by ongoing maintenance: prompts evolve, models are replaced, and your eval set grows along with them.

Frequently asked questions.

What clients typically want to know before we start, including the most frequently asked question: what does a private LLM cost?

How does an LLM work, in broad terms?
A Large Language Model is trained on vast amounts of text and learns patterns within it. For each question, it predicts, word by word, the statistically most plausible next piece of text. It does not reason the way a human does; it recognises patterns. For your integration, this means that the clearer you define the task and the more specific the context you provide, the more reliable the result. A well-built LLM integration rests on three pillars: careful prompts, relevant context (often via RAG), and an evaluation set that measures whether the model performs your task well.
How much does a private LLM cost compared with a cloud LLM?
There are no fixed figures, but there is a clear logic. A cloud LLM from Anthropic, OpenAI or Google charges you per token; for a well-defined use case this often remains limited, especially with caching. A private LLM on your own GPU infrastructure has higher fixed costs (hardware or GPU instances, management, monitoring) but no longer any variable token costs. The break-even point lies where token volume becomes so high that an in-house inference stack is cheaper, or, more often, where compliance requires a private route regardless of price. We calculate both scenarios for you during the first sprint so that the choice rests on facts.
Which model do we choose: Claude, GPT, Gemini or Llama?
That depends on the task, the data residency requirements and the balance between price and performance. Anthropic Claude is strong in nuanced text and long context; OpenAI GPT-4o is broadly applicable and multimodal; Google Gemini is worth considering within the Workspace ecosystem; Mistral and Llama are open source and suitable for private deployment. We test several models on your own examples before choosing, because a choice based on gut feeling or brand name is rarely the right one.
Does our data stay within Europe?
Yes, provided we choose for that. Anthropic and OpenAI offer EU-region deployments via their Bedrock and Azure variants respectively, Google Gemini runs via Vertex AI in EU regions, and a private LLM on your own European cloud or on-premise keeps data within Europe in any case. We record the choice contractually and document the data flows for your GDPR records.
What about the GDPR and the EU AI Act?
Under the GDPR, personal data that you have processed by an LLM requires a data processing agreement with the model provider, a DPIA where risks are higher, and clear documentation of which data ends up where. The EU AI Act adds a classification: depending on the application, your integration falls under minimal, limited or high-risk requirements. At the outset we map your use case against both frameworks and build the integration so that you demonstrably comply.
Will our data be used to train the model?
Not under the business API tiers we use. Anthropic, OpenAI (via the API, not the free ChatGPT app) and Google Vertex AI provide contractual guarantees that inputs and outputs are not used for training. We record those clauses explicitly and, when in doubt, choose a private LLM so that your data never reaches the model provider at all.
How long does a project like this take?
For a well-defined single-call feature, we can be in production within a few sprints. A RAG project with larger document corpora often requires a few additional sprints for the data pipeline and evaluation. A full private LLM setup including infrastructure is a multi-sprint project. We always work iteratively: after the first sprint you have a prototype to steer on.
Do you work together with our internal IT or data team?
Almost always. We carry out knowledge transfer throughout the project, deliver a runbook for incidents and agree clear responsibilities. For organisations that want to approach LLM work more structurally, we are happy to combine this integration with a light AI strategy, so that the first integration does not stand alone but forms part of a larger plan.

Talk to us about your LLM integration.

A free, no-obligation half-hour introductory call. We listen to your use case, ask questions about data, compliance and quality, and give you direction you can act on, even if the eventual advice is not to deploy an LLM at all. For larger projects, we can combine this conversation with an exploration of enterprise AI implementation.

Edit content