AI fine-tuning service: domain models that truly understand your data
Generic LLMs are powerful, but they rarely deliver outputs that get the tone of voice, industry jargon or structured output right for your organisation. We fine-tune open-source models (Llama, Mistral, Qwen, Gemma) and hosted models on your own instruction data, so that a smaller, cheaper model in production outperforms an expensive general-purpose API. From dataset curation and supervised fine-tuning through to DPO, evaluation and deployment on vLLM or Inferentia.
Discuss your fine-tuning case Fine-tune or not?When to fine-tune, when to use RAG, when to rely on prompting alone
The biggest mistake in AI projects in 2024 and 2025: fine-tuning because it is possible, when a better prompt or a retrieval-augmented setup would deliver more. Fine-tuning is expensive, requires curated data and has to be repeated with every base model update. That is why we begin every project with a sober decision tree.
Fine-tuning is the right choice when you want to transfer patterns that do not fit into a prompt: a writing style, a specific output structure, industry jargon, or a chain of reasoning the base model does not produce on its own. For example, a publisher who wants summaries to match the house style exactly, or a legal team that wants memos drafted in a fixed structure. Prompting becomes endlessly brittle in those cases, whereas a few thousand curated examples in an SFT set solve the problem permanently.
RAG (retrieval-augmented generation) is the right choice when you want to add knowledge that changes regularly: product catalogues, documentation, policy texts, client files. A fine-tuned model that has "learned" facts becomes outdated; a RAG pipeline that retrieves the same facts from your own knowledge base stays current. In practice we combine both: a model fine-tuned on your writing style and output format, fed by RAG for up-to-date content. You will see this pattern in custom LLM integrations and in AI document processing.
| Access and site conditions | Approach | Why |
|---|---|---|
| Fixed output structure or JSON schema | Supervised fine-tuning (SFT) | A few hundred golden samples deliver tighter structured output than any prompt. |
| Specialist jargon and domain reasoning | Continued pretraining + SFT | First train unsupervised on a domain corpus, then instruction-tune for task behaviour. |
| Tone of voice or house style | SFT with paired examples | Style is a pattern, not facts; fine-tuning embeds it. |
| Preferred behaviour between two good answers | DPO (Direct Preference Optimisation) | Cheaper and more stable than classic RLHF, with no reward model needed. |
| Current facts or frequently changing content | RAG (no fine-tuning) | Fine-tuned facts go out of date; retrieval stays live. |
| One-off task, fewer than 100 examples | Few-shot prompting | Fine-tuning only pays off from roughly 500 to 1,000 quality samples. |
Methods we apply
Fine-tuning is not a single technique. Which approach suits you depends on your budget, dataset size, hosting requirements and the kind of behaviour you want to transfer.
Supervised fine-tuning
The workhorse technique. We train the model on paired input-output examples, your instruction data. Suitable for structured output, task execution, format compliance and style transfer. Works on almost any model scale from 1B to 70B parameters.
LoRA and QLoRA
Parameter-efficient fine-tuning in which we train small adapter matrices instead of the whole model. QLoRA adds 4-bit quantisation, allowing us to tune 70B models on a single H100 or even a 24GB consumer GPU. Dramatically cheaper than full fine-tuning, and the adapters are small enough to store separately per client or task.
Direct Preference Optimisation
An alignment technique for situations where two correct answers exist and you prefer one over the other. We use DPO as a successor to classic RLHF: no separate reward model, no unstable PPO runs, just direct preference data. Suitable for tone, politeness, refusal behaviour and hallucination suppression.
Continued pretraining
For highly domain-specific language (medical literature, legal case law, technical standards), we continue training the model on an unlabelled corpus before instruction-tuning. This builds a base model that understands your domain's vocabulary and reasoning style. It requires more GPU time but gives a fundamentally more competent starting point.
Adapters and PEFT variants
Alongside LoRA there are adapters, prefix-tuning, IA3 and prompt-tuning. We choose the PEFT method based on model architecture, the volume of training data and deployment requirements. For multi-tenant deployments, adapters are particularly powerful: one base model in memory, with a separate adapter loaded on the fly for each client.
RLHF where needed
Reinforcement learning from human feedback remains relevant for projects where continuous feedback from end users must steer model behaviour. We apply it selectively, usually only in phase 3 of a project after SFT and DPO, because the costs and complexity are considerable. For most use cases, DPO is the better cost-benefit choice.
Dataset preparation: where 80% of the work lies
A fine-tuned model is only as good as its training data. In every project we spend most of our time compiling, cleaning and validating the instruction set. A few thousand carefully curated examples consistently outperform tens of thousands of raw records.
We work with the concept of "golden samples": a core collection of a few hundred examples that demonstrate exactly the desired behaviour. Around these we build synthetic extensions through paraphrasing, back-translation and model-assisted generation, with spot checks by domain experts. For instruction-tuning we use Alpaca-, ShareGPT- or OpenAssistant-style JSONL formats, depending on the target architecture.
Catastrophic forgetting is a real risk: a model that is tuned too narrowly loses its general capabilities. We mitigate this by including a fraction of general-purpose data during SFT (for example a 5-10% Tulu or FLAN mix), and by evaluating at intervals on a held-out general-purpose benchmark. For preference data (DPO), we use paired samples in which human annotators or a stronger reference model choose between two candidate answers.
For clients with sensitive data, we run the full dataset pipeline in a controlled environment: your cloud, our VPC, or on-premises. No data leaves the agreed perimeter, and no training runs take place on shared infrastructure. This is not an optional feature; it is a prerequisite for projects in healthcare, finance and government. Read more about AI for banking and finance and AI in healthcare.
Our fine-tuning workflow
From dataset assessment to production deployment in four phases. Each phase delivers a testable intermediate result, with no months-long black-box process.
Data assessment and goal setting
We take stock of your data, determine whether fine-tuning is the right tool, and agree with you on the evaluation criteria: pass@k, BLEU, ROUGE, task-specific metrics and hallucination rate.
Dataset curation
Building golden samples, choosing the instruction format, synthetic expansion via a stronger model, and quality control by domain experts. Output: a train/validation/test split that meets your target metrics.
Training and evaluation
Selecting the base model, tuning hyperparameters, running training jobs (LoRA/QLoRA/full fine-tuning), and tracking via MLflow or Weights & Biases. Iterative evaluation on the held-out set until the model reaches the target metrics.
Deployment and monitoring
Exporting the model to vLLM, TGI, llama.cpp GGUF or a managed endpoint (Inferentia, Replicate). Drift monitoring, latency tracking, A/B evaluation against the previous model, and a retraining cadence.
Use cases that consistently pay off
Not every AI problem should be fine-tuned, but where it fits, it delivers measurable benefits in cost, latency and quality. A number of recurring patterns follow.
Domain-specific generation (legal, medical, financial)
A law firm that wants to draft memos in a fixed structure and with the correct legal terminology, or a medical team that generates triage summaries that comply with a specific protocol. A fine-tuned 7B-13B model is usually more accurate and more predictable here than a generic API and, crucially, can run on-premises.
Tone of voice for customer service and publications
A publisher that generates copy in a specific house style, or a customer service team that wants to respond consistently according to their communication handbook. Style transfer is a pattern that fine-tuning solves permanently; prompts remain unstable as soon as the input changes. Often combined with AI customer service automation.
Structured output and function calling
For pipelines that must produce JSON, XML or a custom DSL, for example in document extraction, ticket routing or agent tooling, fine-tuning gives stronger guarantees than schema prompting. A model that has seen 100k examples of your schema no longer violates it.
Cost reduction: a smaller fine-tuned model instead of GPT-4
A common scenario: a team runs a high-volume task on GPT-4 or Claude and watches the API bill explode. Targeted SFT on a Llama-3.1-8B or Mistral-7B often achieves 90-95% of the quality at a fraction of the cost and latency, hosted on vLLM in your own VPC.
Test your idea first: a working prototype in 1 day
With OneDayBuild, we turn your idea into something tangible in one day for €1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.
Explore OneDayBuild →Our training and deployment stack
We are deliberately framework-agnostic: the right tool depends on model architecture, dataset size, available GPUs and deployment requirements. What follows is what we actually use, not what is trending on X.
For training, we work depending on the situation with the Hugging Face Trainer and the TRL library for SFT and DPO, with Axolotl for configurable YAML-driven runs on larger models, and with Llama-Factory for broad multi-architecture support. For highly efficient tuning on consumer GPUs we use Unsloth, which includes a number of optimisations that make LoRA runs on 7B models up to 2x faster. Tracking and experiment management go through MLflow or Weights & Biases, depending on your existing MLOps stack.
For compute, where possible we work in your own cloud (AWS, Azure, GCP). If you would rather not build up GPU quota, we use Modal, Replicate or RunPod for occasional training runs. For production inference we run vLLM or Hugging Face Text Generation Inference (TGI) on your own infrastructure, or we convert to GGUF for llama.cpp when CPU or edge deployment is appropriate. AWS Inferentia and Trainium come into play at high volumes where the unit economics would otherwise not add up.
Evaluation, eval frameworks and guardrails
A fine-tuned model that scores well on your eval set is a product. A fine-tuned model without an eval set is a gamble. We deliberately invest in evaluation infrastructure before we start training seriously.
Task-specific metrics
Alongside generic perplexity, we use metrics that match your task: BLEU and ROUGE for summarisation, pass@k for code, exact match and F1 for extraction, and LLM-as-judge for open-ended generation. We build every eval suite to be reproducible, with the same test set, the same prompts and the same decoding parameters, so that improvement between runs can be demonstrated.
Hallucination and safety evaluation
For production models, we measure the hallucination rate using a curated set with known correct answers, and run a safety evaluation (jailbreak resistance, refusal behaviour, toxicity). Results are recorded per model version, so regressions are immediately visible.
A/B evaluation against a baseline
Every fine-tuned version is compared with a clear baseline: the base model with few-shot prompts, or the previous production model. Without that comparison you cannot tell whether the fine-tuning delivered a return or whether a better prompt would have achieved the same result.
Production monitoring and drift detection
After deployment we log input-output pairs (anonymised), monitor latency percentiles, and run periodic regression evals against our fixed test set. Models are retrained when the output distribution shifts or when a newer base model scores significantly better.
Why Appfront for fine-tuning
An honest decision tree before we start
We will say "no fine-tuning" when a better prompt or a RAG pipeline solves the problem. That saves you months of work and considerable compute costs. We do not want to sell you a project that was not needed.
Full pipeline responsibility
From dataset curation through training and evaluation to production deployment and ongoing maintenance. No handover to an MLOps team you still have to find; one team guides the model from concept to retirement.
Privacy and hosting within your perimeter
Training in your cloud or VPC, deployment on your infrastructure, and no data sent via US APIs unless you deliberately choose that. For healthcare, finance and government, this is not a luxury but a GDPR and sector requirement.
Frequently asked questions about AI fine-tuning
A fine-tuned model for your domain?
Discuss your use case, dataset and target metrics with us. We will give you an honest answer on whether fine-tuning is the right tool — and if so, how we get to a production model.
Schedule a conversation