Custom RAG implementation: retrieval-augmented generation for enterprise AI

Your LLM hallucinates, does not know your products and cannot cite a source. Retrieval-augmented generation solves that without fine-tuning: real-time grounding on your own knowledge sources (documentation, contracts, tickets, wikis, product data), with citation grounding for accountability and hybrid retrieval for high precision. We build production RAG that measures what it knows and knows what it does not know.

Hybrid retrieval (BM25 + dense) Reranking (Cohere, ColBERT) Citation grounding RAGAS evaluation Multimodal (PDF, table, image) Agentic RAG
Discuss your RAG use case View architecture
Query Embedding + BM25 hybrid retrieval Vector-DB Pinecone / Qdrant Weaviate / pgvector Reranker Cohere Rerank ColBERT v2 LLM-synthesis met citations [doc-12 ยง3.2] [wiki-handbook] [ticket-4421]

When RAG, when fine-tuning, when prompt engineering alone

The three methods for giving an LLM domain knowledge have different cost structures, latency profiles and accountability characteristics. The right choice depends on how often your data changes, the volume of ground truth and your requirements for source attribution. RAG is the right starting point in nine out of ten enterprise cases.

Prompt engineering works as long as your knowledge fits within a few thousand tokens and rarely changes. Once your knowledge source exceeds the context window or changes daily (product catalogue, ticket history, legal updates), prompt engineering breaks down. Fine-tuning is meant to teach behaviour or style, not to bake in facts. A fine-tuned model that produces an invented fact is just as hard to correct as an error in a weight matrix; you cannot selectively forget it.

RAG decouples knowledge from the model. The LLM remains a general reasoning engine, while the truth lives in an external document store with version history, audit trail and access control. Updating a document is an upsert to the vector DB, not a twelve-hour retraining run. Withdrawing a document (GDPR, confidentiality obligations, an expired contract) is a delete; with fine-tuning that is effectively impossible. For enterprise applications where accountability, retractability and source attribution matter, RAG is the structurally correct architecture.

RAGReal-time grounding on external knowledge. Citations to source. Updates without retraining. The default for enterprise Q&A, support bots, legal assistants and production documentation.
Fine-tuningStyle, tone, output format, task-specific behaviour (classification, extraction). Not facts. Optionally combine with RAG: fine-tune for format, retrieve for content.
Prompt engineeringSmall, static context (under 50KB). Quick prototypes, internal tools. Stops scaling once your source covers more than a few documents or changes regularly.

RAG architecture: ingestion, retrieval, synthesis, observability

A production RAG system consists of two pipelines running on different cadences: ingestion (batch or streaming) and retrieval (per query, sub-second). Layered over both is an observability layer that measures whether the system still does what it was built to do.

๐Ÿ“ฅ

Ingestion pipeline

Document loaders for PDF, DOCX, HTML, Confluence, SharePoint, Notion, Jira, Zendesk, S3 and databases. Layout-aware parsing (Unstructured.io, LlamaParse, Azure Document Intelligence) so that tables, headers and figure captions are not lost. This is followed by chunking (semantic or fixed-size with overlap), metadata enrichment (author, date, confidentiality class, source URI) and batch embedding via OpenAI text-embedding-3-large, voyage-3, Cohere embed-v3 or open-source BGE/e5. Indexing into the vector DB with a hybrid index (dense + BM25/SPLADE).

๐Ÿ”

Retrieval pipeline

Query understanding (rewriting, expansion, HyDE), parallel hybrid retrieval (top 50 dense + top 50 BM25), reciprocal rank fusion, reranking to the top 5 with Cohere Rerank or ColBERT v2, and context assembly with deduplication. The synthesis prompt explicitly asks for citations per claim. Optionally, an agentic loop with tool use for multi-hop questions where a single retrieval step is not enough.

๐Ÿ“Š

Observability & evaluation

Tracing via Langfuse, Helicone, LangSmith or OpenTelemetry: every retrieval, every prompt and every answer is logged. Offline evaluation on golden sets with RAGAS (faithfulness, answer relevancy, context precision, context recall) and LlamaIndex Eval. Online evaluation via thumbs-up/down, escalation rate and human audits. Drift detection: if context recall drops or the citation-grounding score falls, an alert is triggered.

We deliberately keep these three pipelines separate. Ingestion can fail without bringing retrieval down; retrieval can be slow without eval runs waiting on it. Each component can be scaled, debugged and replaced independently โ€” swapping Pinecone for Qdrant or pgvector is then a config change to a single component, not a rebuild of the system.

Chunking strategy, embeddings and vector database: three choices that determine everything

The three technical decisions that most directly affect retrieval quality are: how you split documents into pieces, which embedding model you choose, and which vector store you index in. There is no standard advice โ€” the right choice depends on document length, query style, language and latency budget.

Chunking. Fixed-size chunking (for example 512 tokens with a 64-token overlap) is robust and inexpensive, but it often breaks mid-argument. Semantic chunking uses a sentence encoder to find logical break points, so chunk boundaries fall where the subject actually changes. For long legal or technical documents, hierarchical chunking works well: small chunks for retrieval precision, linked to larger parent chunks that are passed as context during synthesis. The recall score on your own golden set is the only reliable way to determine which strategy wins.

Embeddings. OpenAI text-embedding-3-large is a strong default for English and reasonable Dutch. Voyage-3 and voyage-multilingual consistently perform well on multilingual benchmarks. Cohere embed-v3 supports input-type conditioning (search_document vs. search_query), which benefits asymmetric retrieval. For self-hosted scenarios, BGE-large, e5-mistral and gte-large work well; they can be fine-tuned on your own domain if out-of-the-box accuracy falls short. Important: switching embedding model means rebuilding your entire index, so benchmark beforehand on your own data.

Vector database. Pinecone is managed, scales easily and has strong hybrid search, but cannot be self-hosted. Weaviate offers rich filtering, a GraphQL API, and self-hosted or cloud deployment. Qdrant is high-performance, Rust-based, with strong filtering and a good self-hosting story. pgvector adds vector search to an existing PostgreSQL database, which is handy if your metadata is already relational and you want to manage a single database. Milvus is built for enterprise scale (billions of vectors) but carries more operational complexity. The choice depends on scale, hosting requirements (EU-only?), filter complexity, and whether you want standalone infrastructure or an integrated stack.

Pinecone Weaviate Qdrant pgvector Milvus OpenAI text-embedding-3-large voyage-3 Cohere embed-v3 BGE-large e5-mistral SPLADE BM25

Hybrid retrieval, reranking and advanced RAG patterns

Pure dense retrieval (embeddings only) predictably fails on exact terms, product codes, proper nouns and negations. Pure BM25 misses semantic relationships. Production RAG combines the two โ€” and uses a reranker to filter the top 50 candidates down to the top 5 that actually make it into the prompt.

Hybrid retrieval (BM25 + dense)

Two parallel queries: BM25 (lexical) for exact matches and dense (semantic) retrieval for conceptual matches. Results are merged using Reciprocal Rank Fusion or a weighted combination. On most enterprise corpora, this delivers a 10-25% improvement in recall compared with dense-only retrieval, especially for queries containing product codes, client names and specific terms that were absent from the embedding model's training data.

Reranking (Cohere Rerank, ColBERT v2)

A cross-encoder that calculates a relevance score for each candidate document against the query. It is considerably more expensive than bi-encoder retrieval, so it is only applied to a top-K shortlist. Cohere Rerank is an API call; ColBERT v2 can be self-hosted and quantised. Reranking often adds another 5-15% to precision and firmly filters out chunks that look semantically similar but do not actually answer your question.

HyDE and query expansion

Hypothetical Document Embeddings: the LLM first generates a hypothetical answer to the query, and you then search using the embedding of that answer. This works surprisingly well for questions that match poorly with document style, such as a short question searched against lengthy manuals. Query expansion (synonyms, related terms) raises recall but must be handled with care, as overly broad expansion undermines your precision.

Agentic RAG and GraphRAG

For multi-hop questions ("which contracts expire in Q3 and have a minimum-volume clause?"), a single retrieval step is not enough. Agentic RAG gives the model tools to search iteratively, filter results and pose sub-questions. Microsoft GraphRAG builds an entity graph from the corpus in advance, so queries about relationships between entities can be answered directly via the graph, which is faster and more precise than repeated embedding searches.

Multi-modal RAG

Tables in PDFs, diagrams, screenshots and scanned contracts are out of reach for pure text embeddings. We use layout-aware parsers (Unstructured.io, Azure Document Intelligence, LlamaParse) together with multi-modal embeddings (CLIP, voyage-multimodal, Cohere embed-multimodal) to index tables as structured data and figures as visual embeddings. The result: a question about a specific row in a table lands on the correct cell rather than on a surrounding paragraph.

Citation grounding and attribution

The model is explicitly instructed to link every claim to a document ID and section. In the output, a citation marker appears alongside each statement, and the UI renders it as a clickable source. This is not a cosmetic detail: it is what makes audit requirements, legal review and user trust possible. A RAG system without citations is not production-ready.

Evaluation: RAGAS, golden sets and what it really means for your RAG to work

A RAG system that works in a demo says nothing about whether it will work in production. You need quantitative evaluation; otherwise every change is a guess, and you have no way of knowing whether a new embedding version is an improvement or a regression.

Faithfulness

The extent to which the answer is genuinely grounded in the retrieved context. A hallucinated-fact rate is the most insidious regression here: the system sounds convincing but makes things up. RAGAS scores this by checking, claim by claim, whether it can be inferred from the context.

Answer relevancy

Does the answer actually match the question? A correct but irrelevant answer ("the documentation says X") still counts as a failure if the user asked something more specific. This is often measured with an LLM judge that scores question-answer pairs.

Context precision and recall

Precision: of the retrieved chunks, how many are genuinely relevant. Recall: of all relevant chunks in the corpus, how many were retrieved. These two metrics give the purest measure of retrieval quality, independent of LLM behaviour. NDCG and MRR are classic information retrieval metrics that refine this further.

Golden sets and regression tests

50 to 200 question-answer pairs validated by domain experts, each linked to the specific documents that contain the answer. Every deployment runs against this set. Making a regression visible is half the work; without a golden set you are working on impressions.

We deliver every RAG implementation with an eval harness that you can run yourself. A new LLM version, a different embedding model, an adjusted chunking strategy: you compare before and after, and you decide on figures rather than gut feeling. Frameworks we use as standard: RAGAS, LlamaIndex Eval, custom pytest harnesses for specific claim types, and Langfuse for production tracing.

Not yet sure about a large project?

Test your idea first: a working prototype in 1 day

With OneDayBuild, we turn your idea into something tangible in one day for โ‚ฌ1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.

Explore OneDayBuild โ†’

How we build RAG implementations

Iterative, with measurable deliverables per phase. No large upfront architecture that falls apart in month three; instead, a working system in week four that you can validate yourself.

Discovery and corpus analysis

Which documents, what volume, what format, which confidentiality class. We inventory use cases and define 30 to 50 sample questions with domain experts. We build a first golden set.

Baseline RAG in two to three weeks

A working prototype: ingestion pipeline, hybrid retrieval, a basic UI with citation grounding. A first measurement against the golden set. This is your starting point; all subsequent work improves from here.

Iterative improvement

Adding a reranker, tuning the chunking strategy, testing an embedding model swap, query rewriting, and prompt engineering on the synthesis step. We measure every iteration on faithfulness, recall and relevancy.

Production and observability

Authentication, RBAC, audit logging, tracing (Langfuse), drift monitoring, and alerts on quality regressions. Documentation and knowledge transfer so that your own team can run the eval harness.

The stack we use as standard

We remain model- and vendor-agnostic. The right choice depends on your hosting requirements (EU-only, on-premises, hybrid), scale, budget and latency needs. The list below is what we have used in production, not what we sell.

Python FastAPI LangChain LlamaIndex Haystack DSPy Pinecone Weaviate Qdrant pgvector Milvus OpenAI Anthropic Claude Azure OpenAI Mistral Cohere Rerank ColBERT v2 Unstructured.io LlamaParse RAGAS Langfuse Helicone Docker Kubernetes Terraform

For clients with strict data residency requirements (financial institutions, healthcare, government), we run the full stack on Azure West Europe or a Dutch cloud provider, with self-hosted Qdrant or pgvector and open-source LLMs via vLLM or Azure OpenAI Service in the EU region. No US data egress, and no vendor lock-in to a single SaaS platform.

Concrete RAG applications we build

RAG is not a single product; it is a pattern that fits almost any knowledge-intensive process. A few recurring applications where the return on investment is directly measurable.

Internal knowledge assistant

Confluence, SharePoint, Notion, Google Drive, Slack archives, Jira tickets and support history, combined into one chat interface where new staff and specialists find answers without ringing six colleagues. With RBAC so everyone sees only what they are allowed to see, and citation grounding so that statements link to the source page.

Customer support copilot

An interface for your support staff that, based on a customer question, retrieves relevant knowledge base articles, previous tickets and product documentation, and drafts a reply. The employee checks and sends it, bringing average handling time down and quality and consistency up. Optionally, a tier-1 bot speaks directly to the customer, with escalation logic.

Legal and contract Q&A

Make contracts, NDAs, terms of supply, case law and internal legal memos searchable with citation grounding. "Which contracts contain a minimum-volume clause and expire in Q3?" โ€” a question that takes days by hand, answered in seconds with direct references to article numbers and page numbers.

Production and operational documentation

Make work instructions, machine manuals, safety procedures and QA protocols available to operators on the shop floor via a voice or touch interface. Multi-modal RAG that retrieves diagrams and step-by-step photo sequences. Applicable in pharma, manufacturing, logistics and healthcare.

Data governance, RBAC and GDPR in RAG

Document-level authorisation

RBAC propagates down to the retrieval layer: a user only receives chunks from documents they are authorised to access. We link to user claims (Azure AD, Okta, custom IdP) and filter on metadata in the vector DB. No leaks through a chat interface.

GDPR compliance and data residency

Documents containing personal data are handled in line with the GDPR: purpose limitation, data minimisation, retention. Where EU data residency is required, the entire stack runs in an EU region. We provide a subprocessor overview and data processing agreements as standard.

Audit trail and retractability

Every answer is logged with the query, retrieved context, citations and model version. Withdrawing a document means deleting it from the vector DB plus a verification run. Fine-tuning cannot do this โ€” an important argument for RAG in regulated sectors.

Frequently asked questions about RAG implementation

When should I choose RAG and when fine-tuning?
RAG for facts, fine-tuning for style and task behaviour. If your knowledge changes, is sensitive to version history or requires source citations, RAG is the right choice. Fine-tuning is meant to teach a model classification or to produce a specific output structure consistently, not to bake in facts. In practice we often combine the two: a fine-tuned model for format and tone, with RAG for the substantive grounding.
Which vector database should we choose?
It depends on scale, hosting requirements and filter complexity. Pinecone if you want a managed service and no self-hosting. Qdrant or Weaviate if you want to self-host and need native hybrid retrieval. pgvector if your metadata already lives in PostgreSQL and you don't want to manage an additional database. Milvus if you are heading towards billions of vectors. We benchmark on your own workload during the discovery phase โ€” the choice of stack determines latency budget and operational costs for years.
How do you prevent hallucinations in production?
Three mechanisms working together. One: prompt engineering that explicitly instructs the model to "only make claims that appear in the context, otherwise say that you don't know". Two: RAGAS faithfulness scoring on every deployment, plus continuous tracing in production. Three: citation grounding that links every claim to a source; claims without a citation are flagged as a warning in the UI. Hallucination cannot be eliminated entirely, but it can be measured down to a reported rate that you accept.
What is the difference between reranking and hybrid retrieval?
Hybrid retrieval combines two ways of fetching the top-K candidates: BM25 (lexical) and dense embeddings (semantic). Reranking is a second pass over that top-K with a more expensive, more accurate model (cross-encoder) that scores each candidate's relevance to the query. They are complementary steps. Hybrid retrieval improves recall (you don't miss relevant chunks); reranking improves precision (only the truly relevant chunks make it into the prompt).
How do we measure whether the system works?
With a golden set of 50 to 200 question-and-answer pairs validated by your domain experts, we measure the RAGAS metrics: faithfulness, answer relevancy, context precision and context recall. Online, we additionally monitor escalation rate, user feedback (thumbs up/down) and query pattern drift. With every deployment we run regression tests; if a metric drops below a threshold, the deployment is blocked or an alert is triggered.
Does RAG also work on tables, PDFs with complex layouts and scanned documents?
Yes, provided you use multi-modal RAG. Layout-aware parsers (Unstructured.io, LlamaParse, Azure Document Intelligence) extract tables as structured data and preserve hierarchy. For images and diagrams we use multi-modal embeddings (CLIP, voyage-multimodal). Scanned documents first pass through OCR (Azure, Tesseract, Textract) before entering the pipeline. Table Q&A often requires a separate retrieval strategy, where the table is kept intact as a whole alongside a structured schema for filtering.
How do we keep documents in the vector database in sync with the source?
There are two patterns. Batch: a nightly or hourly job detects changes in the source (last-modified, change feed) and upserts them into the vector database. Streaming: event-driven via webhooks, where every change in Confluence, SharePoint or SQL immediately triggers an ingestion job. For delete propagation, we log document IDs with version and timestamp. Withdrawing a document means a delete plus a verification query, which matters for GDPR rights such as the right to erasure.
What determines the lead time of a RAG implementation?
A baseline RAG with hybrid retrieval, citation grounding and an evaluation harness typically takes two to four weeks. What determines the total lead time is corpus complexity (how many formats, how much parsing work), integration requirements (how many source systems to connect), governance (RBAC, audit, compliance review), and how much access you have to domain experts for golden-set validation. We work in iterative sprints, so you see results along the way and can steer, rather than waiting months.

Ready to have a RAG implementation built?

Discuss your use case with us. We will analyse your corpus, define an initial golden set and deliver a working baseline within a few weeks: measurable, with citations, and production-ready.

Schedule a technical session

Edit content