AI fine-tuning service: domain models that truly understand your data

Generic LLMs are powerful, but they rarely deliver outputs that get the tone of voice, industry jargon or structured output right for your organisation. We fine-tune open-source models (Llama, Mistral, Qwen, Gemma) and hosted models on your own instruction data, so that a smaller, cheaper model in production outperforms an expensive general-purpose API. From dataset curation and supervised fine-tuning through to DPO, evaluation and deployment on vLLM or Inferentia.

Supervised fine-tuning LoRA / QLoRA DPO Continued pretraining Evaluation & guardrails vLLM deployment
Discuss your fine-tuning case Fine-tune or not?
Base model Llama-3.1-8B + instructie-data SFT / LoRA + preferenties DPO / RLHF Fine-tuned model eval pass@1 87% cost / 1k tok -72% vLLM TGI GGUF

When to fine-tune, when to use RAG, when to rely on prompting alone

The biggest mistake in AI projects in 2024 and 2025: fine-tuning because it is possible, when a better prompt or a retrieval-augmented setup would deliver more. Fine-tuning is expensive, requires curated data and has to be repeated with every base model update. That is why we begin every project with a sober decision tree.

Fine-tuning is the right choice when you want to transfer patterns that do not fit into a prompt: a writing style, a specific output structure, industry jargon, or a chain of reasoning the base model does not produce on its own. For example, a publisher who wants summaries to match the house style exactly, or a legal team that wants memos drafted in a fixed structure. Prompting becomes endlessly brittle in those cases, whereas a few thousand curated examples in an SFT set solve the problem permanently.

RAG (retrieval-augmented generation) is the right choice when you want to add knowledge that changes regularly: product catalogues, documentation, policy texts, client files. A fine-tuned model that has "learned" facts becomes outdated; a RAG pipeline that retrieves the same facts from your own knowledge base stays current. In practice we combine both: a model fine-tuned on your writing style and output format, fed by RAG for up-to-date content. You will see this pattern in custom LLM integrations and in AI document processing.

Access and site conditionsApproachWhy
Fixed output structure or JSON schemaSupervised fine-tuning (SFT)A few hundred golden samples deliver tighter structured output than any prompt.
Specialist jargon and domain reasoningContinued pretraining + SFTFirst train unsupervised on a domain corpus, then instruction-tune for task behaviour.
Tone of voice or house styleSFT with paired examplesStyle is a pattern, not facts; fine-tuning embeds it.
Preferred behaviour between two good answersDPO (Direct Preference Optimisation)Cheaper and more stable than classic RLHF, with no reward model needed.
Current facts or frequently changing contentRAG (no fine-tuning)Fine-tuned facts go out of date; retrieval stays live.
One-off task, fewer than 100 examplesFew-shot promptingFine-tuning only pays off from roughly 500 to 1,000 quality samples.

Methods we apply

Fine-tuning is not a single technique. Which approach suits you depends on your budget, dataset size, hosting requirements and the kind of behaviour you want to transfer.

SFT

Supervised fine-tuning

The workhorse technique. We train the model on paired input-output examples, your instruction data. Suitable for structured output, task execution, format compliance and style transfer. Works on almost any model scale from 1B to 70B parameters.

LoRA

LoRA and QLoRA

Parameter-efficient fine-tuning in which we train small adapter matrices instead of the whole model. QLoRA adds 4-bit quantisation, allowing us to tune 70B models on a single H100 or even a 24GB consumer GPU. Dramatically cheaper than full fine-tuning, and the adapters are small enough to store separately per client or task.

DPO

Direct Preference Optimisation

An alignment technique for situations where two correct answers exist and you prefer one over the other. We use DPO as a successor to classic RLHF: no separate reward model, no unstable PPO runs, just direct preference data. Suitable for tone, politeness, refusal behaviour and hallucination suppression.

CPT

Continued pretraining

For highly domain-specific language (medical literature, legal case law, technical standards), we continue training the model on an unlabelled corpus before instruction-tuning. This builds a base model that understands your domain's vocabulary and reasoning style. It requires more GPU time but gives a fundamentally more competent starting point.

PEFT

Adapters and PEFT variants

Alongside LoRA there are adapters, prefix-tuning, IA3 and prompt-tuning. We choose the PEFT method based on model architecture, the volume of training data and deployment requirements. For multi-tenant deployments, adapters are particularly powerful: one base model in memory, with a separate adapter loaded on the fly for each client.

RLHF

RLHF where needed

Reinforcement learning from human feedback remains relevant for projects where continuous feedback from end users must steer model behaviour. We apply it selectively, usually only in phase 3 of a project after SFT and DPO, because the costs and complexity are considerable. For most use cases, DPO is the better cost-benefit choice.

Dataset preparation: where 80% of the work lies

A fine-tuned model is only as good as its training data. In every project we spend most of our time compiling, cleaning and validating the instruction set. A few thousand carefully curated examples consistently outperform tens of thousands of raw records.

We work with the concept of "golden samples": a core collection of a few hundred examples that demonstrate exactly the desired behaviour. Around these we build synthetic extensions through paraphrasing, back-translation and model-assisted generation, with spot checks by domain experts. For instruction-tuning we use Alpaca-, ShareGPT- or OpenAssistant-style JSONL formats, depending on the target architecture.

Catastrophic forgetting is a real risk: a model that is tuned too narrowly loses its general capabilities. We mitigate this by including a fraction of general-purpose data during SFT (for example a 5-10% Tulu or FLAN mix), and by evaluating at intervals on a held-out general-purpose benchmark. For preference data (DPO), we use paired samples in which human annotators or a stronger reference model choose between two candidate answers.

For clients with sensitive data, we run the full dataset pipeline in a controlled environment: your cloud, our VPC, or on-premises. No data leaves the agreed perimeter, and no training runs take place on shared infrastructure. This is not an optional feature; it is a prerequisite for projects in healthcare, finance and government. Read more about AI for banking and finance and AI in healthcare.

Our fine-tuning workflow

From dataset assessment to production deployment in four phases. Each phase delivers a testable intermediate result, with no months-long black-box process.

Data assessment and goal setting

We take stock of your data, determine whether fine-tuning is the right tool, and agree with you on the evaluation criteria: pass@k, BLEU, ROUGE, task-specific metrics and hallucination rate.

Dataset curation

Building golden samples, choosing the instruction format, synthetic expansion via a stronger model, and quality control by domain experts. Output: a train/validation/test split that meets your target metrics.

Training and evaluation

Selecting the base model, tuning hyperparameters, running training jobs (LoRA/QLoRA/full fine-tuning), and tracking via MLflow or Weights & Biases. Iterative evaluation on the held-out set until the model reaches the target metrics.

Deployment and monitoring

Exporting the model to vLLM, TGI, llama.cpp GGUF or a managed endpoint (Inferentia, Replicate). Drift monitoring, latency tracking, A/B evaluation against the previous model, and a retraining cadence.

Use cases that consistently pay off

Not every AI problem should be fine-tuned, but where it fits, it delivers measurable benefits in cost, latency and quality. A number of recurring patterns follow.

Domain-specific generation (legal, medical, financial)

A law firm that wants to draft memos in a fixed structure and with the correct legal terminology, or a medical team that generates triage summaries that comply with a specific protocol. A fine-tuned 7B-13B model is usually more accurate and more predictable here than a generic API and, crucially, can run on-premises.

Tone of voice for customer service and publications

A publisher that generates copy in a specific house style, or a customer service team that wants to respond consistently according to their communication handbook. Style transfer is a pattern that fine-tuning solves permanently; prompts remain unstable as soon as the input changes. Often combined with AI customer service automation.

Structured output and function calling

For pipelines that must produce JSON, XML or a custom DSL, for example in document extraction, ticket routing or agent tooling, fine-tuning gives stronger guarantees than schema prompting. A model that has seen 100k examples of your schema no longer violates it.

Cost reduction: a smaller fine-tuned model instead of GPT-4

A common scenario: a team runs a high-volume task on GPT-4 or Claude and watches the API bill explode. Targeted SFT on a Llama-3.1-8B or Mistral-7B often achieves 90-95% of the quality at a fraction of the cost and latency, hosted on vLLM in your own VPC.

Not yet sure about a large project?

Test your idea first: a working prototype in 1 day

With OneDayBuild, we turn your idea into something tangible in one day for €1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.

Explore OneDayBuild →

Our training and deployment stack

We are deliberately framework-agnostic: the right tool depends on model architecture, dataset size, available GPUs and deployment requirements. What follows is what we actually use, not what is trending on X.

For training, we work depending on the situation with the Hugging Face Trainer and the TRL library for SFT and DPO, with Axolotl for configurable YAML-driven runs on larger models, and with Llama-Factory for broad multi-architecture support. For highly efficient tuning on consumer GPUs we use Unsloth, which includes a number of optimisations that make LoRA runs on 7B models up to 2x faster. Tracking and experiment management go through MLflow or Weights & Biases, depending on your existing MLOps stack.

For compute, where possible we work in your own cloud (AWS, Azure, GCP). If you would rather not build up GPU quota, we use Modal, Replicate or RunPod for occasional training runs. For production inference we run vLLM or Hugging Face Text Generation Inference (TGI) on your own infrastructure, or we convert to GGUF for llama.cpp when CPU or edge deployment is appropriate. AWS Inferentia and Trainium come into play at high volumes where the unit economics would otherwise not add up.

PyTorch Hugging Face Transformers TRL PEFT Axolotl Llama-Factory Unsloth DeepSpeed vLLM TGI llama.cpp / GGUF MLflow Weights & Biases Modal RunPod AWS Inferentia Llama 3.1 / Mistral / Qwen / Gemma

Evaluation, eval frameworks and guardrails

A fine-tuned model that scores well on your eval set is a product. A fine-tuned model without an eval set is a gamble. We deliberately invest in evaluation infrastructure before we start training seriously.

Task-specific metrics

Alongside generic perplexity, we use metrics that match your task: BLEU and ROUGE for summarisation, pass@k for code, exact match and F1 for extraction, and LLM-as-judge for open-ended generation. We build every eval suite to be reproducible, with the same test set, the same prompts and the same decoding parameters, so that improvement between runs can be demonstrated.

Hallucination and safety evaluation

For production models, we measure the hallucination rate using a curated set with known correct answers, and run a safety evaluation (jailbreak resistance, refusal behaviour, toxicity). Results are recorded per model version, so regressions are immediately visible.

A/B evaluation against a baseline

Every fine-tuned version is compared with a clear baseline: the base model with few-shot prompts, or the previous production model. Without that comparison you cannot tell whether the fine-tuning delivered a return or whether a better prompt would have achieved the same result.

Production monitoring and drift detection

After deployment we log input-output pairs (anonymised), monitor latency percentiles, and run periodic regression evals against our fixed test set. Models are retrained when the output distribution shifts or when a newer base model scores significantly better.

Why Appfront for fine-tuning

An honest decision tree before we start

We will say "no fine-tuning" when a better prompt or a RAG pipeline solves the problem. That saves you months of work and considerable compute costs. We do not want to sell you a project that was not needed.

Full pipeline responsibility

From dataset curation through training and evaluation to production deployment and ongoing maintenance. No handover to an MLOps team you still have to find; one team guides the model from concept to retirement.

Privacy and hosting within your perimeter

Training in your cloud or VPC, deployment on your infrastructure, and no data sent via US APIs unless you deliberately choose that. For healthcare, finance and government, this is not a luxury but a GDPR and sector requirement.

Frequently asked questions about AI fine-tuning

When is fine-tuning better than RAG?
When you want to transfer patterns — writing style, output structure, reasoning chains or domain jargon — that do not fit stably into a prompt. RAG is better when you want to add facts or frequently changing content. In practice we combine both: fine-tuning for the how, RAG for the what.
How much data do I need for a fine-tuning project?
For SFT on an open-source model we usually see useful results from 500–1,000 carefully curated instruction pairs. For LoRA on a specific style transfer, 200–500 is sometimes enough. For continued pretraining or large-scale domain adaptation, you are talking about millions of tokens of unlabelled corpus.
What is the difference between LoRA and QLoRA?
LoRA trains small adapter matrices instead of the full model, saving a great deal of memory and compute. QLoRA combines that with 4-bit quantisation of the base model, so you can even fine-tune a 70B-parameter model on a single H100 or a 24GB consumer GPU. QLoRA has become the standard starting point for most modern projects.
When do I choose DPO over RLHF?
In almost all cases. DPO is cheaper, more stable and does not require a separate reward model or PPO loop. RLHF remains relevant for projects where continuous feedback from end users needs to steer model behaviour, but for most business fine-tuning DPO offers a better cost-benefit ratio.
Which base model do you recommend?
That depends on licence requirements, language and deployment context. For English and Dutch, Llama 3.1 (8B/70B), Mistral, Qwen 2.5 and Gemma work well. For strictly commercial licences without restrictions, we look at the Apache 2.0 models within those families. For something extremely small and CPU-deployable, Phi-3 or Gemma 2 often make the shortlist.
Does a fine-tuned model lose general knowledge (catastrophic forgetting)?
It can, especially with one-sided datasets and learning rates that are too high. We mitigate this by mixing a fraction of generic instruction data into training, setting conservative learning rates and evaluating periodically on a general benchmark. For strictly domain-only models some loss is acceptable — for general-purpose assistants it is not.
Can we host the fine-tuned model on-premises?
Yes. We deploy on vLLM or TGI within your own Kubernetes cluster, on AWS Inferentia/Trainium for unit-cost optimisation, or we convert to GGUF for llama.cpp on CPU or edge hardware. No mandatory use of external APIs, and no data leaves your perimeter.
How do we measure whether the fine-tuning worked?
Before we start, we agree which metrics define success — pass@k, BLEU/ROUGE, task-specific F1, hallucination rate, latency, cost per 1,000 tokens. We measure every run reproducibly against the same held-out set and compare explicitly with the baseline (the base model with few-shot prompts or your existing production model).
What happens when a newer base model is released?
This is a recurring rhythm. We maintain your evaluation set and run a baseline on the new base model. If the gain is significant, we repeat the fine-tuning on the new foundation. Because dataset curation is the heaviest work and can be reused, a base-model update is usually a matter of weeks, not months.

A fine-tuned model for your domain?

Discuss your use case, dataset and target metrics with us. We will give you an honest answer on whether fine-tuning is the right tool — and if so, how we get to a production model.

Schedule a conversation

Edit content