Private LLM on-premise: data-sovereign AI within your own walls

For organisations that want to use generative AI without taking sensitive data outside their own infrastructure. We design, implement and manage on-premise large language models (Llama 3.3, Qwen, Mistral, DeepSeek and Gemma) on your own GPU cluster or in a Dutch private cloud zone. No US cloud, no training leaks, no unexpected model swaps: you keep full control over data, models and audit trails.

vLLM & TGI NVIDIA NIM Air-gapped deployments Quantization (AWQ/GPTQ/FP8) NEN 7510 / BIO / GDPR
Discuss your on-prem case View the architecture
GPU GPU vLLM FP8 AWQ Audit log no-egress

Why on-premise or private-cloud LLMs rather than the public API

For most organisations, the story starts with the same question: may patient records, mortgage applications, due diligence documents or operational military data be processed by a US cloud LLM at all? The answer is usually more nuanced than a plain yes or no, but in four areas the floodgates are shut.

The first area is data sovereignty. Since Schrems II, the legal basis for transferring data to the United States has been fragile: organisations must assess, for each processing flow, whether standard contractual clauses, a transfer impact assessment and additional technical measures are sufficient. For healthcare providers under NEN 7510, financial institutions supervised by DNB and the AFM, government bodies under the BIO and defence-related organisations, that assessment almost always comes out negative in practice. On-premise, or a Dutch private-cloud zone with no transatlantic processing, makes the discussion simple: data never leaves the building.

The second area is no-train guarantees. Public API providers claim that enterprise data is not used for model training, but the legal enforceability is limited, telemetry often flows outside the EU, and every model change or policy update alters what actually happens to your data. A private LLM that you run yourself simply can never end up in a training batch by accident. For IP-sensitive sectors such as pharma, high tech, law firms and M&A advisers, that is a hard requirement.

The third area is latency and availability. Inference inside your own data centre delivers sub-100 ms time-to-first-token, rather than 400-1,500 ms over a transatlantic connection. For real-time use cases such as interactive copilots, agent loops with multiple tool calls per second and voice AI, that is the difference between workable and frustrating. You also avoid going down when the upstream provider has an outage or suddenly tightens its rate limits.

The fourth area is auditability. GDPR Art. 32 requires you to be able to demonstrate which data was processed, by which model, with what result and by which user. ENISA guidelines for AI systems reinforce this. On a private LLM you can log every prompt, every response, every model version and every weighted output in full, without depending on what an external provider chooses to include in its audit log.

Which open-weight models work on-premise today

Over the past eighteen months, open-weight models have become competitive with the largest closed-source models on many benchmarks. For most enterprise use cases, including RAG, summarisation, classification, code assistance and structured extraction, on-prem is now a fully viable alternative. Here is a brief orientation of the model landscape.

πŸ¦™

Llama 3.1 / 3.3 (Meta)

The Llama family delivers solid general-purpose performance in English and Dutch at 8B, 70B and 405B parameters. Llama 3.3 70B Instruct is the workhorse size for company chatbots and RAG: it quantizes well to AWQ-INT4 on a single H100 or L40S, with strong instruction following and tool use. The licence permits commercial use within the well-known thresholds.

πŸ‰

Qwen 2.5 / 3 (Alibaba)

Qwen 2.5-72B and the more recent Qwen 3 series score particularly well on multilingual benchmarks, code and mathematics. Qwen 2.5-Coder-32B is a popular choice for on-prem code assistants. The Apache 2.0 licence on many variants is easier for enterprise legal teams to accept than the Llama community licence.

🌬️

Mistral & Mixtral

Mistral Small/Medium and the Mixtral 8x22B mixture-of-experts models are European open-weight candidates with explicitly enterprise-oriented licences. Mixtral combines high effective capacity with relatively few active parameters per token, which keeps throughput favourable on more expensive GPUs.

🐳

DeepSeek (V3 / R1)

DeepSeek-V3 and the reasoning variant R1 have open weights of considerable size (671B MoE with ~37B active). For reasoning-heavy tasks β€” legal, scientific, financial analysis β€” a quantised DeepSeek distill is an interesting on-prem option, provided you accept that the base models themselves were trained in China: the weights are local and send nothing upstream.

πŸ’Ž

Gemma 2 / 3 (Google)

Gemma 2-27B and Gemma 3 are relatively small, sharply calibrated models. Gemma 3, with multimodal vision and long context, is interesting for document processing pipelines that combine OCR and text. A good candidate for edge deployments with limited VRAM.

πŸ‡ͺπŸ‡Ί

EU-specific and domain models

EuroLLM, Salamandra and the Aleph Alpha models explicitly position themselves as European alternatives. We are also seeing growing use of domain fine-tunes: medical Llama variants (Meditron, OpenBioLLM), legal fine-tunes and financial LLMs. We assess which variant fits your data domain and regulatory context.

The inference stack: from model weights to production API

Downloading an open-weight model is the easy part. The real challenge lies in building an inference layer that balances throughput, latency, memory and reliability. The choice of inference engine often determines half of your GPU bill.

vLLM has become the de facto standard for high-throughput LLM serving. Continuous batching, paged attention and KV-cache optimisation mean that an H100 running vLLM can handle 5-10x more concurrent users than a naΓ―ve transformers implementation. vLLM natively supports Llama, Qwen, Mistral, Gemma, DeepSeek and the AWQ, GPTQ, GGUF and FP8 quantisation formats, plus prefix caching for RAG workloads where the same context is used repeatedly.

Hugging Face Text Generation Inference (TGI) is a strong alternative, with good integration into the wider HF ecosystem and a solid Triton Inference Server integration. For organisations already on the NVIDIA stack, NVIDIA NIM (NVIDIA Inference Microservices) is an attractive option: ready-made containerised inference microservices with TensorRT-LLM optimisations, FP8 paths on Hopper, and Helm charts for Kubernetes. NIM delivers out-of-the-box throughput that is hard to match with a hand-built vLLM deployment.

For lighter or edge deployments, Ollama and llama.cpp are excellent. llama.cpp runs GGUF-quantised models on CPU, on smaller GPUs or even Apple Silicon, and is a good choice for laptop copilots or demo environments. Ollama builds a user-friendly API layer on top, with automatic model management. Not every use case needs an 8x H100 cluster β€” sometimes a Mac Studio or a single L40S is enough.

Quantisation is the pivot between model size and GPU budget. AWQ (activation-aware weight quantisation) and GPTQ deliver INT4 weights with <1% accuracy loss on most benchmarks. FP8 on Hopper GPUs (H100/H200) combines high throughput with better quality than INT4. GGUF is the common quantisation format for llama.cpp and offers 2-bit to 8-bit variants. For RAG pipelines, we also optimise with FlashAttention-2 or -3 for long contexts, and TensorRT-LLM engines for critical low-latency paths.

vLLM TGI NVIDIA NIM Triton Inference Server TensorRT-LLM FlashAttention-3 paged attention KV cache AWQ GPTQ GGUF FP8 llama.cpp Ollama Kubernetes Helm

GPU choices and capacity planning

Your hardware choice depends on model size, concurrent users, the tokens per second you need per user and context lengths. Here is a sketch of what we see working in production.

NVIDIA H100 / H200

The Hopper generation remains the premium choice for 70B+ models. With 80GB HBM3 (H100) or 141GB HBM3e (H200), a Llama 3.3 70B Instruct in FP8 fits on a single card, with plenty of room left for the KV cache and long contexts. For 405B models or MoE models such as DeepSeek-V3, you deploy 4-8 cards in tensor parallelism over NVLink/NVSwitch.

NVIDIA L40S

The Ada-generation L40S with 48GB GDDR6 is the price-performance favourite for 7B-to-30B models and INT4-quantised 70B. It has no NVLink, so tensor parallelism over PCIe is suboptimal, but it works excellently for single-card inference or pipeline parallelism. Often the right choice for mid-market on-prem deployments.

NVIDIA A100

The previous Ampere generation remains relevant: 40GB and 80GB variants, strong FP16/BF16 performance and wide availability on the second-hand market. There is no FP8, so quantisation ends up at INT4-AWQ or INT8. For many clients it is a good stopgap when a new H100 allocation has a lead time of months.

AMD MI300X

With 192GB HBM3 on a single card, the MI300X is the only option on which you can run a 405B model in BF16 on just a few cards. The ROCm toolchain with vLLM and SGLang is genuinely production-ready in 2025. For organisations that want to avoid NVIDIA lock-in or urgently need shorter delivery times, it is a serious candidate.

Capacity planning starts with four questions: which model, which quantisation, how many concurrent sessions and which context length. In vLLM benchmarks, a Llama 3.3 70B AWQ-INT4 on a single H100 typically serves 30-60 concurrent users with acceptable latency at contexts up to 8k. Doubling that requires a second card or more aggressive quantisation. We calculate this for your use case based on measured tokens per second, not vendor marketing.

Cost vs cloud LLM: for steady-state workloads above ~100M tokens per month, the TCO tips in favour of owning GPUs over public APIs. For highly spiky or small workloads, pay-per-token APIs remain economical, unless data sovereignty rules them out. We make that calculation explicit before any hardware is ordered.

Sectors where a private on-premise LLM is often the only option

Not every organisation needs an on-prem LLM. But for these sectors, laws, regulations or contractual obligations have almost always settled the question already.

Healthcare (NEN 7510 / NEN 7512 / NEN 7513)

Patient records, EHR data, imaging reports and triage conversations count as special category personal data. A private LLM for clinical copilots, discharge letter drafting or triage support processes this data within the hospital network, with audit trails that meet NEN 7513 logging requirements. Deployment almost always requires a DPIA and coordination with the Data Protection Officer.

Finance (DNB / AFM / DORA)

Banks, insurers and pension funds supervised by DNB must explicitly manage outsourcing risks. Since 2025, DORA has added strict requirements for third-party ICT providers. A private LLM for compliance monitoring, anti-money-laundering screening or fraud detection keeps processing within your own risk perimeter. It can be combined with our AI fraud detection solutions and with domain fine-tunes on internal policy documents.

Government (BIO / GDPR / Open Government Act)

National and municipal organisations work under the Baseline Information Security Government (BIO) and additional sector baselines. An on-prem or Dutch private cloud LLM meets the BIO requirements on data classification and jurisdiction. See also our pages on AI for municipalities and government for sector-specific implementations.

Defence and critical infrastructure

Defence-related suppliers and energy, water and telecoms companies under NIS2 face classification requirements that rule out public APIs. For them, air-gapped private LLMs are a necessity, not a luxury. We design deployments in which models, data and logging remain physically separated from the internet.

M&A and due diligence

Data rooms, letters of intent and draft purchase agreements are extremely sensitive: a leak can cost you a deal or lead to price-sensitive issues. A private LLM that summarises documents, flags red flags and answers questionnaires within an isolated project environment is a natural fit. Once the transaction is complete, the environment is simply shut down.

Pharma, high-tech and legal

R&D protocols, patent files, clinical data, client files: IP and confidentiality obligations that leave no room for public APIs. A private LLM fine-tuned on internal corpora delivers better results than a generic model and meets your confidentiality obligations.

Not sure about a large project yet?

Test your idea first: a working prototype in 1 day

With OneDayBuild, we make your idea tangible in a single day for €1,150, so you know whether further development is worth the investment. Decide to go ahead with the full build? Then we deduct the cost in full.

View OneDayBuild β†’

Architecture: air-gapped, hybrid or multi-tenant

Three architecture patterns cover most enterprise cases. Which one suits you depends on your risk profile, scaling needs and existing infrastructure.

Air-gapped on-premise

Models, vector stores and logging run on infrastructure with no outbound internet connection. Updates to model weights and software go through an airlock procedure: download on a bastion host, virus scan, hash verification, transfer to the closed zone. This is the standard for defence, critical infrastructure and the highest government classifications.

Hybrid: on-prem + cloud burst

Sensitive processing stays on-prem, while non-sensitive bulk work goes through an EU-hosted private endpoint or a moderated public API. A classification layer at the gateway routes each prompt to the right backend based on data labels. Suited to organisations that want peak scalability without giving up data sovereignty.

API gateway for multiple apps

A single on-prem LLM cluster serves multiple applications through a central gateway with OAuth/OIDC, per-application rate limits, prompt filtering, PII redaction and quota monitoring. It works like an internal OpenAI API: developers get a token, while IT security keeps central oversight.

Observability and audit

Full prompt/response logging in an append-only datastore, with PII pseudonymisation before storage where necessary. Model versions, system prompts and tooling configuration are logged too, for reproducibility. OpenTelemetry traces down to the inference engine make latency regressions traceable.

Model update pipeline

New model versions go through a fixed pipeline: download, hash verification, eval suite (Dutch benchmark, RAG set, refusal tests), canary rollout to 5% of traffic, then full promotion. Rollback takes minutes if quality metrics degrade.

Fallback strategy

In the event of GPU failure or cluster maintenance, the gateway routes to a secondary pool, a smaller reserve model or a gracefully degraded mode that accepts only critical requests. For healthcare and finance applications, we set explicit RTO/RPO targets in a service-level objective document.

From first workshop to managed private LLM environment

Our approach to on-prem LLM projects follows four phases. Each phase delivers a testable result: no abstract architecture diagrams without a working system underneath.

Discovery and risk analysis

We map out use cases, data classifications, the regulatory framework (NEN 7510, BIO, DNB circulars, DORA, NIS2) and existing infrastructure. The result is a ranked list of candidate applications with their feasibility and risk profile.

Proof of value on pilot hardware

On a test setup, often a single L40S or H100, we run the selected candidate models against your own data. This includes an eval suite with Dutch benchmarks, RAG relevance tests and refusal tests, giving you a concrete go/no-go decision based on quality.

Production deployment

Cluster design, GPU procurement, vLLM/NIM orchestration on Kubernetes, a gateway with OAuth and logging, and a model update pipeline. Integration with your identity provider, SIEM and existing monitoring stack. Test, acceptance and production environments.

Management and further development

SLA-backed management: 24/7 monitoring, model version rollouts, capacity monitoring, security patches and regular refusal evals. Periodic reviews of whether newer models add value, because the model market moves fast.

Why choose Appfront for your private LLM project

A sovereign stack as the starting point

We have been building software for sectors where US-cloud-by-default is not an option for some time. Our defaults are EU hosting, no-egress architectures and records of processing that meet GDPR Article 32. That is not an extra service; it is simply how we work.

Deep LLM engineering

Quantization choices, KV cache tuning, paged-attention configuration, prefix caching for RAG and TensorRT-LLM engine builds are not tick-box items for us but daily work. We know when FP8 does or does not cost you quality, and how to get 2x throughput without extra hardware.

One partner from workshop to operations

Discovery, architecture, implementation, integration with your applications and ongoing management under one roof. No handover between consultancy and build team, and no loss of context at each phase. You keep a single point of contact for the entire lifecycle.

Frequently asked questions about on-premise private LLMs

What is the difference between on-premise, private cloud and a public LLM API?
On-premise means the GPU hardware and the model sit physically in your own data centre or colocation facility, under your own network and access control. Private cloud is a dedicated tenant with a (Dutch) cloud provider, logically isolated but physically shared within the same region. A public LLM API sends prompts to a shared multi-tenant service run by an external vendor. The right choice depends on data classification, your need for control and the IT staff you have available.
Which open-weight model is currently best for Dutch-language enterprise use?
For general Dutch-language business applications, Llama 3.3 70B Instruct generally performs well, with Qwen 2.5-72B and Mistral Medium as strong alternatives. The right choice depends on your use case: code assistance calls for different fine-tunes than clinical documentation or contract analysis. We run an eval suite on your own data so the choice is well founded.
How much GPU capacity do I need for 100 concurrent users?
As a rough rule of thumb, Llama 3.3 70B in AWQ-INT4 on a single H100 can handle 30-60 concurrent sessions with vLLM at acceptable latency on contexts of up to 8k tokens. For 100 users with longer contexts or higher tokens-per-second requirements, plan on 2-4 H100s, or an MoE model on 4-8 cards. Capacity planning depends heavily on prompt length, output length and the latency you need. We benchmark with representative traffic profiles.
Does an on-premise LLM comply with the GDPR and sector baselines such as NEN 7510 or the BIO?
On-premise is an important precondition, not the whole solution. GDPR compliance also requires a DPIA, data minimisation, purpose limitation, a retention policy, audit trails and appropriate technical and organisational measures under Article 32. NEN 7510 and the BIO add sector-specific requirements around logging, segmentation and authorisation. We deliver a setup that can be audited against these frameworks, including documentation and involvement of your DPO/CISO.
Can I also fine-tune my private LLM on internal data?
Yes. For most use cases, RAG (retrieval-augmented generation) is the first answer: you keep the base model unchanged and retrieve your data from a vector store. For specialised domains β€” legal jargon, medical terminology, proprietary taxonomies β€” we add LoRA or full fine-tuning on your own GPU infrastructure. Training data stays on your premises; nothing is sent to an external provider.
What happens if a better open-weight model is released tomorrow?
Our deployments are built modularly using vLLM, TGI or NIM. Adding a new model means downloading the weights, verifying the hash, running it through the eval suite and a canary rollout. No redesign of the stack. The model market moves fast; we have built that pace of change explicitly into the architecture.
How does the TCO compare to public API costs?
For steady-state workloads above a few tens of millions of tokens per month, the calculation usually tips in favour of owning your own GPUs, even after factoring in depreciation, power, cooling and management. For highly spiky or very small workloads, pay-per-token remains more attractive β€” unless the regulatory framework rules that option out. We make the business case explicit before any hardware is ordered.
Can a private LLM run fully air-gapped?
Yes. We have deployments where the inference environment has no outbound connection. Model updates, software patches and CVE fixes go through an airlock procedure: a bastion host downloads, scans and hashes; transfer to the closed zone takes place via controlled media or a one-way connection. Logging stays entirely within the zone. This is standard for defence and the highest-classification government applications.

A private LLM on your own terms

Discuss with us whether an on-premise or private-cloud LLM is feasible for your organisation. We map out the risks, model choice and TCO in concrete terms β€” non-binding and with no obligation.

Plan a discovery session

Edit content