Custom RAG implementation: retrieval-augmented generation for enterprise AI
Your LLM hallucinates, does not know your products and cannot cite a source. Retrieval-augmented generation solves that without fine-tuning: real-time grounding on your own knowledge sources (documentation, contracts, tickets, wikis, product data), with citation grounding for accountability and hybrid retrieval for high precision. We build production RAG that measures what it knows and knows what it does not know.
Discuss your RAG use case View architectureWhen RAG, when fine-tuning, when prompt engineering alone
The three methods for giving an LLM domain knowledge have different cost structures, latency profiles and accountability characteristics. The right choice depends on how often your data changes, the volume of ground truth and your requirements for source attribution. RAG is the right starting point in nine out of ten enterprise cases.
Prompt engineering works as long as your knowledge fits within a few thousand tokens and rarely changes. Once your knowledge source exceeds the context window or changes daily (product catalogue, ticket history, legal updates), prompt engineering breaks down. Fine-tuning is meant to teach behaviour or style, not to bake in facts. A fine-tuned model that produces an invented fact is just as hard to correct as an error in a weight matrix; you cannot selectively forget it.
RAG decouples knowledge from the model. The LLM remains a general reasoning engine, while the truth lives in an external document store with version history, audit trail and access control. Updating a document is an upsert to the vector DB, not a twelve-hour retraining run. Withdrawing a document (GDPR, confidentiality obligations, an expired contract) is a delete; with fine-tuning that is effectively impossible. For enterprise applications where accountability, retractability and source attribution matter, RAG is the structurally correct architecture.
RAG architecture: ingestion, retrieval, synthesis, observability
A production RAG system consists of two pipelines running on different cadences: ingestion (batch or streaming) and retrieval (per query, sub-second). Layered over both is an observability layer that measures whether the system still does what it was built to do.
Ingestion pipeline
Document loaders for PDF, DOCX, HTML, Confluence, SharePoint, Notion, Jira, Zendesk, S3 and databases. Layout-aware parsing (Unstructured.io, LlamaParse, Azure Document Intelligence) so that tables, headers and figure captions are not lost. This is followed by chunking (semantic or fixed-size with overlap), metadata enrichment (author, date, confidentiality class, source URI) and batch embedding via OpenAI text-embedding-3-large, voyage-3, Cohere embed-v3 or open-source BGE/e5. Indexing into the vector DB with a hybrid index (dense + BM25/SPLADE).
Retrieval pipeline
Query understanding (rewriting, expansion, HyDE), parallel hybrid retrieval (top 50 dense + top 50 BM25), reciprocal rank fusion, reranking to the top 5 with Cohere Rerank or ColBERT v2, and context assembly with deduplication. The synthesis prompt explicitly asks for citations per claim. Optionally, an agentic loop with tool use for multi-hop questions where a single retrieval step is not enough.
Observability & evaluation
Tracing via Langfuse, Helicone, LangSmith or OpenTelemetry: every retrieval, every prompt and every answer is logged. Offline evaluation on golden sets with RAGAS (faithfulness, answer relevancy, context precision, context recall) and LlamaIndex Eval. Online evaluation via thumbs-up/down, escalation rate and human audits. Drift detection: if context recall drops or the citation-grounding score falls, an alert is triggered.
We deliberately keep these three pipelines separate. Ingestion can fail without bringing retrieval down; retrieval can be slow without eval runs waiting on it. Each component can be scaled, debugged and replaced independently โ swapping Pinecone for Qdrant or pgvector is then a config change to a single component, not a rebuild of the system.
Chunking strategy, embeddings and vector database: three choices that determine everything
The three technical decisions that most directly affect retrieval quality are: how you split documents into pieces, which embedding model you choose, and which vector store you index in. There is no standard advice โ the right choice depends on document length, query style, language and latency budget.
Chunking. Fixed-size chunking (for example 512 tokens with a 64-token overlap) is robust and inexpensive, but it often breaks mid-argument. Semantic chunking uses a sentence encoder to find logical break points, so chunk boundaries fall where the subject actually changes. For long legal or technical documents, hierarchical chunking works well: small chunks for retrieval precision, linked to larger parent chunks that are passed as context during synthesis. The recall score on your own golden set is the only reliable way to determine which strategy wins.
Embeddings. OpenAI text-embedding-3-large is a strong default for English and reasonable Dutch. Voyage-3 and voyage-multilingual consistently perform well on multilingual benchmarks. Cohere embed-v3 supports input-type conditioning (search_document vs. search_query), which benefits asymmetric retrieval. For self-hosted scenarios, BGE-large, e5-mistral and gte-large work well; they can be fine-tuned on your own domain if out-of-the-box accuracy falls short. Important: switching embedding model means rebuilding your entire index, so benchmark beforehand on your own data.
Vector database. Pinecone is managed, scales easily and has strong hybrid search, but cannot be self-hosted. Weaviate offers rich filtering, a GraphQL API, and self-hosted or cloud deployment. Qdrant is high-performance, Rust-based, with strong filtering and a good self-hosting story. pgvector adds vector search to an existing PostgreSQL database, which is handy if your metadata is already relational and you want to manage a single database. Milvus is built for enterprise scale (billions of vectors) but carries more operational complexity. The choice depends on scale, hosting requirements (EU-only?), filter complexity, and whether you want standalone infrastructure or an integrated stack.
Hybrid retrieval, reranking and advanced RAG patterns
Pure dense retrieval (embeddings only) predictably fails on exact terms, product codes, proper nouns and negations. Pure BM25 misses semantic relationships. Production RAG combines the two โ and uses a reranker to filter the top 50 candidates down to the top 5 that actually make it into the prompt.
Hybrid retrieval (BM25 + dense)
Two parallel queries: BM25 (lexical) for exact matches and dense (semantic) retrieval for conceptual matches. Results are merged using Reciprocal Rank Fusion or a weighted combination. On most enterprise corpora, this delivers a 10-25% improvement in recall compared with dense-only retrieval, especially for queries containing product codes, client names and specific terms that were absent from the embedding model's training data.
Reranking (Cohere Rerank, ColBERT v2)
A cross-encoder that calculates a relevance score for each candidate document against the query. It is considerably more expensive than bi-encoder retrieval, so it is only applied to a top-K shortlist. Cohere Rerank is an API call; ColBERT v2 can be self-hosted and quantised. Reranking often adds another 5-15% to precision and firmly filters out chunks that look semantically similar but do not actually answer your question.
HyDE and query expansion
Hypothetical Document Embeddings: the LLM first generates a hypothetical answer to the query, and you then search using the embedding of that answer. This works surprisingly well for questions that match poorly with document style, such as a short question searched against lengthy manuals. Query expansion (synonyms, related terms) raises recall but must be handled with care, as overly broad expansion undermines your precision.
Agentic RAG and GraphRAG
For multi-hop questions ("which contracts expire in Q3 and have a minimum-volume clause?"), a single retrieval step is not enough. Agentic RAG gives the model tools to search iteratively, filter results and pose sub-questions. Microsoft GraphRAG builds an entity graph from the corpus in advance, so queries about relationships between entities can be answered directly via the graph, which is faster and more precise than repeated embedding searches.
Multi-modal RAG
Tables in PDFs, diagrams, screenshots and scanned contracts are out of reach for pure text embeddings. We use layout-aware parsers (Unstructured.io, Azure Document Intelligence, LlamaParse) together with multi-modal embeddings (CLIP, voyage-multimodal, Cohere embed-multimodal) to index tables as structured data and figures as visual embeddings. The result: a question about a specific row in a table lands on the correct cell rather than on a surrounding paragraph.
Citation grounding and attribution
The model is explicitly instructed to link every claim to a document ID and section. In the output, a citation marker appears alongside each statement, and the UI renders it as a clickable source. This is not a cosmetic detail: it is what makes audit requirements, legal review and user trust possible. A RAG system without citations is not production-ready.
Evaluation: RAGAS, golden sets and what it really means for your RAG to work
A RAG system that works in a demo says nothing about whether it will work in production. You need quantitative evaluation; otherwise every change is a guess, and you have no way of knowing whether a new embedding version is an improvement or a regression.
Faithfulness
The extent to which the answer is genuinely grounded in the retrieved context. A hallucinated-fact rate is the most insidious regression here: the system sounds convincing but makes things up. RAGAS scores this by checking, claim by claim, whether it can be inferred from the context.
Answer relevancy
Does the answer actually match the question? A correct but irrelevant answer ("the documentation says X") still counts as a failure if the user asked something more specific. This is often measured with an LLM judge that scores question-answer pairs.
Context precision and recall
Precision: of the retrieved chunks, how many are genuinely relevant. Recall: of all relevant chunks in the corpus, how many were retrieved. These two metrics give the purest measure of retrieval quality, independent of LLM behaviour. NDCG and MRR are classic information retrieval metrics that refine this further.
Golden sets and regression tests
50 to 200 question-answer pairs validated by domain experts, each linked to the specific documents that contain the answer. Every deployment runs against this set. Making a regression visible is half the work; without a golden set you are working on impressions.
We deliver every RAG implementation with an eval harness that you can run yourself. A new LLM version, a different embedding model, an adjusted chunking strategy: you compare before and after, and you decide on figures rather than gut feeling. Frameworks we use as standard: RAGAS, LlamaIndex Eval, custom pytest harnesses for specific claim types, and Langfuse for production tracing.
Test your idea first: a working prototype in 1 day
With OneDayBuild, we turn your idea into something tangible in one day for โฌ1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.
Explore OneDayBuild โHow we build RAG implementations
Iterative, with measurable deliverables per phase. No large upfront architecture that falls apart in month three; instead, a working system in week four that you can validate yourself.
Discovery and corpus analysis
Which documents, what volume, what format, which confidentiality class. We inventory use cases and define 30 to 50 sample questions with domain experts. We build a first golden set.
Baseline RAG in two to three weeks
A working prototype: ingestion pipeline, hybrid retrieval, a basic UI with citation grounding. A first measurement against the golden set. This is your starting point; all subsequent work improves from here.
Iterative improvement
Adding a reranker, tuning the chunking strategy, testing an embedding model swap, query rewriting, and prompt engineering on the synthesis step. We measure every iteration on faithfulness, recall and relevancy.
Production and observability
Authentication, RBAC, audit logging, tracing (Langfuse), drift monitoring, and alerts on quality regressions. Documentation and knowledge transfer so that your own team can run the eval harness.
The stack we use as standard
We remain model- and vendor-agnostic. The right choice depends on your hosting requirements (EU-only, on-premises, hybrid), scale, budget and latency needs. The list below is what we have used in production, not what we sell.
For clients with strict data residency requirements (financial institutions, healthcare, government), we run the full stack on Azure West Europe or a Dutch cloud provider, with self-hosted Qdrant or pgvector and open-source LLMs via vLLM or Azure OpenAI Service in the EU region. No US data egress, and no vendor lock-in to a single SaaS platform.
Concrete RAG applications we build
RAG is not a single product; it is a pattern that fits almost any knowledge-intensive process. A few recurring applications where the return on investment is directly measurable.
Internal knowledge assistant
Confluence, SharePoint, Notion, Google Drive, Slack archives, Jira tickets and support history, combined into one chat interface where new staff and specialists find answers without ringing six colleagues. With RBAC so everyone sees only what they are allowed to see, and citation grounding so that statements link to the source page.
Customer support copilot
An interface for your support staff that, based on a customer question, retrieves relevant knowledge base articles, previous tickets and product documentation, and drafts a reply. The employee checks and sends it, bringing average handling time down and quality and consistency up. Optionally, a tier-1 bot speaks directly to the customer, with escalation logic.
Legal and contract Q&A
Make contracts, NDAs, terms of supply, case law and internal legal memos searchable with citation grounding. "Which contracts contain a minimum-volume clause and expire in Q3?" โ a question that takes days by hand, answered in seconds with direct references to article numbers and page numbers.
Production and operational documentation
Make work instructions, machine manuals, safety procedures and QA protocols available to operators on the shop floor via a voice or touch interface. Multi-modal RAG that retrieves diagrams and step-by-step photo sequences. Applicable in pharma, manufacturing, logistics and healthcare.
Data governance, RBAC and GDPR in RAG
Document-level authorisation
RBAC propagates down to the retrieval layer: a user only receives chunks from documents they are authorised to access. We link to user claims (Azure AD, Okta, custom IdP) and filter on metadata in the vector DB. No leaks through a chat interface.
GDPR compliance and data residency
Documents containing personal data are handled in line with the GDPR: purpose limitation, data minimisation, retention. Where EU data residency is required, the entire stack runs in an EU region. We provide a subprocessor overview and data processing agreements as standard.
Audit trail and retractability
Every answer is logged with the query, retrieved context, citations and model version. Withdrawing a document means deleting it from the vector DB plus a verification run. Fine-tuning cannot do this โ an important argument for RAG in regulated sectors.
Frequently asked questions about RAG implementation
Ready to have a RAG implementation built?
Discuss your use case with us. We will analyse your corpus, define an initial golden set and deliver a working baseline within a few weeks: measurable, with citations, and production-ready.
Schedule a technical session