Private LLM on-premise: data-sovereign AI within your own walls
For organisations that want to use generative AI without taking sensitive data outside their own infrastructure. We design, implement and manage on-premise large language models (Llama 3.3, Qwen, Mistral, DeepSeek and Gemma) on your own GPU cluster or in a Dutch private cloud zone. No US cloud, no training leaks, no unexpected model swaps: you keep full control over data, models and audit trails.
Discuss your on-prem case View the architectureWhy on-premise or private-cloud LLMs rather than the public API
For most organisations, the story starts with the same question: may patient records, mortgage applications, due diligence documents or operational military data be processed by a US cloud LLM at all? The answer is usually more nuanced than a plain yes or no, but in four areas the floodgates are shut.
The first area is data sovereignty. Since Schrems II, the legal basis for transferring data to the United States has been fragile: organisations must assess, for each processing flow, whether standard contractual clauses, a transfer impact assessment and additional technical measures are sufficient. For healthcare providers under NEN 7510, financial institutions supervised by DNB and the AFM, government bodies under the BIO and defence-related organisations, that assessment almost always comes out negative in practice. On-premise, or a Dutch private-cloud zone with no transatlantic processing, makes the discussion simple: data never leaves the building.
The second area is no-train guarantees. Public API providers claim that enterprise data is not used for model training, but the legal enforceability is limited, telemetry often flows outside the EU, and every model change or policy update alters what actually happens to your data. A private LLM that you run yourself simply can never end up in a training batch by accident. For IP-sensitive sectors such as pharma, high tech, law firms and M&A advisers, that is a hard requirement.
The third area is latency and availability. Inference inside your own data centre delivers sub-100 ms time-to-first-token, rather than 400-1,500 ms over a transatlantic connection. For real-time use cases such as interactive copilots, agent loops with multiple tool calls per second and voice AI, that is the difference between workable and frustrating. You also avoid going down when the upstream provider has an outage or suddenly tightens its rate limits.
The fourth area is auditability. GDPR Art. 32 requires you to be able to demonstrate which data was processed, by which model, with what result and by which user. ENISA guidelines for AI systems reinforce this. On a private LLM you can log every prompt, every response, every model version and every weighted output in full, without depending on what an external provider chooses to include in its audit log.
Which open-weight models work on-premise today
Over the past eighteen months, open-weight models have become competitive with the largest closed-source models on many benchmarks. For most enterprise use cases, including RAG, summarisation, classification, code assistance and structured extraction, on-prem is now a fully viable alternative. Here is a brief orientation of the model landscape.
Llama 3.1 / 3.3 (Meta)
The Llama family delivers solid general-purpose performance in English and Dutch at 8B, 70B and 405B parameters. Llama 3.3 70B Instruct is the workhorse size for company chatbots and RAG: it quantizes well to AWQ-INT4 on a single H100 or L40S, with strong instruction following and tool use. The licence permits commercial use within the well-known thresholds.
Qwen 2.5 / 3 (Alibaba)
Qwen 2.5-72B and the more recent Qwen 3 series score particularly well on multilingual benchmarks, code and mathematics. Qwen 2.5-Coder-32B is a popular choice for on-prem code assistants. The Apache 2.0 licence on many variants is easier for enterprise legal teams to accept than the Llama community licence.
Mistral & Mixtral
Mistral Small/Medium and the Mixtral 8x22B mixture-of-experts models are European open-weight candidates with explicitly enterprise-oriented licences. Mixtral combines high effective capacity with relatively few active parameters per token, which keeps throughput favourable on more expensive GPUs.
DeepSeek (V3 / R1)
DeepSeek-V3 and the reasoning variant R1 have open weights of considerable size (671B MoE with ~37B active). For reasoning-heavy tasks β legal, scientific, financial analysis β a quantised DeepSeek distill is an interesting on-prem option, provided you accept that the base models themselves were trained in China: the weights are local and send nothing upstream.
Gemma 2 / 3 (Google)
Gemma 2-27B and Gemma 3 are relatively small, sharply calibrated models. Gemma 3, with multimodal vision and long context, is interesting for document processing pipelines that combine OCR and text. A good candidate for edge deployments with limited VRAM.
EU-specific and domain models
EuroLLM, Salamandra and the Aleph Alpha models explicitly position themselves as European alternatives. We are also seeing growing use of domain fine-tunes: medical Llama variants (Meditron, OpenBioLLM), legal fine-tunes and financial LLMs. We assess which variant fits your data domain and regulatory context.
The inference stack: from model weights to production API
Downloading an open-weight model is the easy part. The real challenge lies in building an inference layer that balances throughput, latency, memory and reliability. The choice of inference engine often determines half of your GPU bill.
vLLM has become the de facto standard for high-throughput LLM serving. Continuous batching, paged attention and KV-cache optimisation mean that an H100 running vLLM can handle 5-10x more concurrent users than a naΓ―ve transformers implementation. vLLM natively supports Llama, Qwen, Mistral, Gemma, DeepSeek and the AWQ, GPTQ, GGUF and FP8 quantisation formats, plus prefix caching for RAG workloads where the same context is used repeatedly.
Hugging Face Text Generation Inference (TGI) is a strong alternative, with good integration into the wider HF ecosystem and a solid Triton Inference Server integration. For organisations already on the NVIDIA stack, NVIDIA NIM (NVIDIA Inference Microservices) is an attractive option: ready-made containerised inference microservices with TensorRT-LLM optimisations, FP8 paths on Hopper, and Helm charts for Kubernetes. NIM delivers out-of-the-box throughput that is hard to match with a hand-built vLLM deployment.
For lighter or edge deployments, Ollama and llama.cpp are excellent. llama.cpp runs GGUF-quantised models on CPU, on smaller GPUs or even Apple Silicon, and is a good choice for laptop copilots or demo environments. Ollama builds a user-friendly API layer on top, with automatic model management. Not every use case needs an 8x H100 cluster β sometimes a Mac Studio or a single L40S is enough.
Quantisation is the pivot between model size and GPU budget. AWQ (activation-aware weight quantisation) and GPTQ deliver INT4 weights with <1% accuracy loss on most benchmarks. FP8 on Hopper GPUs (H100/H200) combines high throughput with better quality than INT4. GGUF is the common quantisation format for llama.cpp and offers 2-bit to 8-bit variants. For RAG pipelines, we also optimise with FlashAttention-2 or -3 for long contexts, and TensorRT-LLM engines for critical low-latency paths.
GPU choices and capacity planning
Your hardware choice depends on model size, concurrent users, the tokens per second you need per user and context lengths. Here is a sketch of what we see working in production.
NVIDIA H100 / H200
The Hopper generation remains the premium choice for 70B+ models. With 80GB HBM3 (H100) or 141GB HBM3e (H200), a Llama 3.3 70B Instruct in FP8 fits on a single card, with plenty of room left for the KV cache and long contexts. For 405B models or MoE models such as DeepSeek-V3, you deploy 4-8 cards in tensor parallelism over NVLink/NVSwitch.
NVIDIA L40S
The Ada-generation L40S with 48GB GDDR6 is the price-performance favourite for 7B-to-30B models and INT4-quantised 70B. It has no NVLink, so tensor parallelism over PCIe is suboptimal, but it works excellently for single-card inference or pipeline parallelism. Often the right choice for mid-market on-prem deployments.
NVIDIA A100
The previous Ampere generation remains relevant: 40GB and 80GB variants, strong FP16/BF16 performance and wide availability on the second-hand market. There is no FP8, so quantisation ends up at INT4-AWQ or INT8. For many clients it is a good stopgap when a new H100 allocation has a lead time of months.
AMD MI300X
With 192GB HBM3 on a single card, the MI300X is the only option on which you can run a 405B model in BF16 on just a few cards. The ROCm toolchain with vLLM and SGLang is genuinely production-ready in 2025. For organisations that want to avoid NVIDIA lock-in or urgently need shorter delivery times, it is a serious candidate.
Capacity planning starts with four questions: which model, which quantisation, how many concurrent sessions and which context length. In vLLM benchmarks, a Llama 3.3 70B AWQ-INT4 on a single H100 typically serves 30-60 concurrent users with acceptable latency at contexts up to 8k. Doubling that requires a second card or more aggressive quantisation. We calculate this for your use case based on measured tokens per second, not vendor marketing.
Cost vs cloud LLM: for steady-state workloads above ~100M tokens per month, the TCO tips in favour of owning GPUs over public APIs. For highly spiky or small workloads, pay-per-token APIs remain economical, unless data sovereignty rules them out. We make that calculation explicit before any hardware is ordered.
Sectors where a private on-premise LLM is often the only option
Not every organisation needs an on-prem LLM. But for these sectors, laws, regulations or contractual obligations have almost always settled the question already.
Healthcare (NEN 7510 / NEN 7512 / NEN 7513)
Patient records, EHR data, imaging reports and triage conversations count as special category personal data. A private LLM for clinical copilots, discharge letter drafting or triage support processes this data within the hospital network, with audit trails that meet NEN 7513 logging requirements. Deployment almost always requires a DPIA and coordination with the Data Protection Officer.
Finance (DNB / AFM / DORA)
Banks, insurers and pension funds supervised by DNB must explicitly manage outsourcing risks. Since 2025, DORA has added strict requirements for third-party ICT providers. A private LLM for compliance monitoring, anti-money-laundering screening or fraud detection keeps processing within your own risk perimeter. It can be combined with our AI fraud detection solutions and with domain fine-tunes on internal policy documents.
Government (BIO / GDPR / Open Government Act)
National and municipal organisations work under the Baseline Information Security Government (BIO) and additional sector baselines. An on-prem or Dutch private cloud LLM meets the BIO requirements on data classification and jurisdiction. See also our pages on AI for municipalities and government for sector-specific implementations.
Defence and critical infrastructure
Defence-related suppliers and energy, water and telecoms companies under NIS2 face classification requirements that rule out public APIs. For them, air-gapped private LLMs are a necessity, not a luxury. We design deployments in which models, data and logging remain physically separated from the internet.
M&A and due diligence
Data rooms, letters of intent and draft purchase agreements are extremely sensitive: a leak can cost you a deal or lead to price-sensitive issues. A private LLM that summarises documents, flags red flags and answers questionnaires within an isolated project environment is a natural fit. Once the transaction is complete, the environment is simply shut down.
Pharma, high-tech and legal
R&D protocols, patent files, clinical data, client files: IP and confidentiality obligations that leave no room for public APIs. A private LLM fine-tuned on internal corpora delivers better results than a generic model and meets your confidentiality obligations.
Test your idea first: a working prototype in 1 day
With OneDayBuild, we make your idea tangible in a single day for β¬1,150, so you know whether further development is worth the investment. Decide to go ahead with the full build? Then we deduct the cost in full.
View OneDayBuild βArchitecture: air-gapped, hybrid or multi-tenant
Three architecture patterns cover most enterprise cases. Which one suits you depends on your risk profile, scaling needs and existing infrastructure.
Air-gapped on-premise
Models, vector stores and logging run on infrastructure with no outbound internet connection. Updates to model weights and software go through an airlock procedure: download on a bastion host, virus scan, hash verification, transfer to the closed zone. This is the standard for defence, critical infrastructure and the highest government classifications.
Hybrid: on-prem + cloud burst
Sensitive processing stays on-prem, while non-sensitive bulk work goes through an EU-hosted private endpoint or a moderated public API. A classification layer at the gateway routes each prompt to the right backend based on data labels. Suited to organisations that want peak scalability without giving up data sovereignty.
API gateway for multiple apps
A single on-prem LLM cluster serves multiple applications through a central gateway with OAuth/OIDC, per-application rate limits, prompt filtering, PII redaction and quota monitoring. It works like an internal OpenAI API: developers get a token, while IT security keeps central oversight.
Observability and audit
Full prompt/response logging in an append-only datastore, with PII pseudonymisation before storage where necessary. Model versions, system prompts and tooling configuration are logged too, for reproducibility. OpenTelemetry traces down to the inference engine make latency regressions traceable.
Model update pipeline
New model versions go through a fixed pipeline: download, hash verification, eval suite (Dutch benchmark, RAG set, refusal tests), canary rollout to 5% of traffic, then full promotion. Rollback takes minutes if quality metrics degrade.
Fallback strategy
In the event of GPU failure or cluster maintenance, the gateway routes to a secondary pool, a smaller reserve model or a gracefully degraded mode that accepts only critical requests. For healthcare and finance applications, we set explicit RTO/RPO targets in a service-level objective document.
From first workshop to managed private LLM environment
Our approach to on-prem LLM projects follows four phases. Each phase delivers a testable result: no abstract architecture diagrams without a working system underneath.
Discovery and risk analysis
We map out use cases, data classifications, the regulatory framework (NEN 7510, BIO, DNB circulars, DORA, NIS2) and existing infrastructure. The result is a ranked list of candidate applications with their feasibility and risk profile.
Proof of value on pilot hardware
On a test setup, often a single L40S or H100, we run the selected candidate models against your own data. This includes an eval suite with Dutch benchmarks, RAG relevance tests and refusal tests, giving you a concrete go/no-go decision based on quality.
Production deployment
Cluster design, GPU procurement, vLLM/NIM orchestration on Kubernetes, a gateway with OAuth and logging, and a model update pipeline. Integration with your identity provider, SIEM and existing monitoring stack. Test, acceptance and production environments.
Management and further development
SLA-backed management: 24/7 monitoring, model version rollouts, capacity monitoring, security patches and regular refusal evals. Periodic reviews of whether newer models add value, because the model market moves fast.
Why choose Appfront for your private LLM project
A sovereign stack as the starting point
We have been building software for sectors where US-cloud-by-default is not an option for some time. Our defaults are EU hosting, no-egress architectures and records of processing that meet GDPR Article 32. That is not an extra service; it is simply how we work.
Deep LLM engineering
Quantization choices, KV cache tuning, paged-attention configuration, prefix caching for RAG and TensorRT-LLM engine builds are not tick-box items for us but daily work. We know when FP8 does or does not cost you quality, and how to get 2x throughput without extra hardware.
One partner from workshop to operations
Discovery, architecture, implementation, integration with your applications and ongoing management under one roof. No handover between consultancy and build team, and no loss of context at each phase. You keep a single point of contact for the entire lifecycle.
Frequently asked questions about on-premise private LLMs
A private LLM on your own terms
Discuss with us whether an on-premise or private-cloud LLM is feasible for your organisation. We map out the risks, model choice and TCO in concrete terms β non-binding and with no obligation.
Plan a discovery session