Custom AI translation software development: from e-commerce localisation to legal and medical translation

Scaling to multiple EU markets quickly means translating thousands of product titles, legal clauses, package inserts or support conversations, at a quality that respects brand consistency, terminology and compliance. Generic translation buttons don't cut it. We build bespoke AI translation software that combines NMT engines, LLMs, terminology databases and human post-editing into a pipeline suited to your product catalogue, legal domain or medical records workflow.

E-commerce localisation Legal translation Medical translation Live-chat translation Document translation Meeting transcription
Discuss your translation case View use cases
EN: Pre-order now NL: Nu vooraf bestellen DE: Jetzt vorbestellen FR: Precommandez ES: Reserva ahora QE 92 TBX OK Glossary COMET QE

Why a bespoke AI translation pipeline delivers more than a button in your CMS

Generic translation plugins work for standalone sentences, but they fall short once terminology must stay consistent across thousands of SKUs, once a legal contract must not lose the distinction between "agreement" and "contract", or once package-insert text must match the registered product information word for word.

EU businesses exporting to six or more markets face a constant flow: new product pages on the webshop, contract amendments, policy documents, knowledge-base articles, support tickets, customer conversations, annual reports. Manual translation doesn't scale, and raw machine output without review damages the brand. The gain lies in a hybrid pipeline where neural machine translation (NMT) handles the first 80%, a terminology database ensures product names and legal terms don't slip, a quality-estimation model flags which segments need a human eye, and a translator with domain expertise reviews only those segments.

We build that pipeline around your existing systems: Shopify, Magento, Salesforce Commerce, Akeneo PIM, Contentful, an in-house DMS or a contract management platform. The translation engine isn't the product; the integration with your content flow is. We set up the architecture so Dutch content flows into EN, DE, FR, ES, IT and beyond without editors or lawyers having to copy and paste every release by hand.

Application areas where AI translation software makes a real difference

From product content to legal clauses and medical records, each domain has its own quality criteria, terminology and risk profile. We tailor the pipeline architecture accordingly.

🛒

E-commerce product content i18n

Product titles, descriptions, variants, attribute values, SEO meta tags and category texts in five to ten languages. Pipeline: PIM export to XLIFF, NMT translation with a brand glossary, automatic SEO keyword mapping per market and publication back to Akeneo, Magento or Shopify. Includes delta detection so only changed segments are retranslated.

⚖️

Legal translation with terminology control

Contracts, NDAs, general terms and conditions, privacy statements and compliance documents. A TBX terminology database records how legal concepts ("ontbinding", "verzuim", "overmacht", "boetebeding") are translated consistently across all languages. DNT tags protect party names, amounts and clause numbers from being altered by the model.

🏥

Medical translation and compliance

Package inserts, IFUs, clinical protocols, patient forms and medical records. MDR- and EMA-compliant terminology via MedDRA and a bespoke TBX. Mandatory review by a medically trained linguist per segment, with an audit trail. Back-translation as an additional QA step where regulations require it.

💬

Live-chat translation in customer support

A Dutch customer support agent who chats in real time with German, French or Spanish customers. Bidirectional translation with sub-second latency, glossary injection for product names and a correction loop whenever the QE score drops below a threshold. Integration with Zendesk, Intercom, Salesforce Service Cloud or your own helpdesk.

🎙️

Meeting transcription and translation

Whisper for speech-to-text, followed by DeepL or Claude to translate the transcript into the languages of your participants. Speaker labelling, automatic summaries per participant language and export to Notion, SharePoint or your own knowledge base. Useful for international teams, customer calls and compliance meetings.

📄

Document translation with layout preservation

PDF, Word and InDesign files that keep their formatting: columns, tables, footnotes, headers, footers and images stay exactly where they are. Pipeline: extraction to XLIFF with segment IDs, translation, reinjection via a document engine and an automated QA check for text overflow. No re-typesetting and no shifted layouts.

How Appfront builds an AI translation pipeline

Our starting question is never "which model should we use?" but "where does the content sit today, where does it need to go, and who reviews what?". The technical choices (NMT engine, LLM, post-editing tool) follow from the content flow and the quality requirements, not the other way round. A furniture brand's product catalogue can tolerate far more machine output than the leaflet for an infusion solution, so the pipeline choices differ fundamentally.

In practice, we first map your source and target systems, identify content types by quality class, and for each class choose the right mix of NMT (DeepL Pro, Google Translate Advanced, Microsoft Translator), LLM translation (OpenAI GPT-4o, Claude for stylistic translation), terminology enforcement (TBX import, Phrase TMS, MemoQ) and human post-editing (MTPE). We then build the orchestration: webhooks from your CMS or PIM, queueing, retry logic, audit logging and dashboards that let you track turnaround times, machine cost and post-editor effort.

The end result is a system that works invisibly: an editor publishes Dutch content, a few minutes later the other language versions are ready for review, a post-editor works through the flagged segments and the site or document goes live. No manual export-import cycles, no copying and pasting into a spreadsheet, and no more loose emails to a translation agency.

Architecture: NMT, LLMs and the orchestration in between

An AI translation pipeline is not a single model call. Quality comes from the combination of translation engine, terminology layer, quality estimation and review routing. These are the building blocks we use by default and tailor to your situation.

NMT vs LLM-based translation

Neural machine translation (DeepL, Google, Microsoft) is fast, inexpensive and strong at sentence level, which makes it ideal for product content and support cases. LLMs such as GPT-4o and Claude excel at stylistic translation, marketing tone, idiomatic expressions and domain context. We combine both: NMT as the baseline, and an LLM for segments where tone of voice or cultural adaptation matters most.

Glossary, TM and TBX pipelines

A Translation Memory (TMX) reuses previously approved translations; fuzzy matches save costs and guarantee consistency. A TBX terminology database enforces that "warranty" is always rendered as the agreed term and never drifts. We feed both into NMT engines via glossary APIs or into LLM prompts via dynamic retrieval, so agreed terminology does not slip away.

Quality estimation with COMET

COMET (and, in older setups, BLEU/chrF) calculates a quality score per segment without needing a reference translation. Segments above a threshold are published directly; those below are routed to a post-editor. The result: less review time, a predictable turnaround and measurable quality per content type and language.

DNT tags, back-translation and QA

Do-Not-Translate markers protect brand names, codes, legal article numbers and units of measurement. Back-translation (translating the target back into the source) is an additional QA step for medical and legal content, where divergent interpretations carry risk. Automated QA checks (length overflow, missing placeholders, number mismatches) run on every segment before publication.

From initial audit to production pipeline

Our approach to AI translation projects runs in four phases. Each phase is clearly defined, delivers something tangible and allows you to adjust course along the way.

Content audit and language strategy

We take stock of content types, languages, volumes, quality requirements per type and existing TM/TBX assets. We map where your content currently lives (CMS, PIM, DMS, helpdesk) and where it needs to go.

Engine selection and proof of concept

We benchmark two or three NMT and LLM combinations on your own content (anonymised), measure BLEU/COMET scores and post-editing effort, and select the stack that offers the best balance of cost, speed and quality.

Integration and MTPE workflow

Building API integrations with your CMS, PIM or helpdesk, setting up TM/TBX pipelines, integrating with CAT tools (Trados, MemoQ, Phrase), and configuring post-editing queues for your in-house or external linguists.

Monitoring and ongoing development

Dashboards for turnaround time, machine cost, post-edit distance and QE trends. Periodic retraining of glossaries, expansion into new languages and refinement of QA rules based on post-editor feedback.

Not yet sure about a large project?

Test your idea first: a working prototype in 1 day

With OneDayBuild, we turn your idea into something tangible in one day for €1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.

Explore OneDayBuild →

Technology we use for AI translation software

The stack depends on content type, languages and compliance requirements. For pure e-commerce localisation we often choose a DeepL Pro API baseline with a Phrase TMS layer on top. For legal and medical content we shift to a setup with stronger terminology control and a bigger role for LLMs in stylistic and context-aware translation. For live chat, sub-second latency is a hard requirement, so we choose different models than for batch document translation.

All components preferably run on EU infrastructure: DeepL is hosted in Germany, Microsoft Translator offers EU data residency, and we route LLM calls through Azure OpenAI Service in West Europe or an EU-based Claude deployment where available. For sectors with the strictest requirements (healthcare, finance, government) we build an on-premises or private cloud variant with self-hosted models (NLLB, Marian, Mistral fine-tunes).

DeepL Pro API Google Translate API Microsoft Translator API OpenAI GPT-4o Claude (Anthropic) Whisper COMET QE BLEU / chrF TBX TMX XLIFF 2.1 Trados MemoQ Phrase TMS NLLB Marian NMT Hugging Face LangChain Python / FastAPI PostgreSQL

Translation data and privacy: what you need to arrange

Translation content contains personal data and trade secrets and, in legal and medical contexts, even special category personal data. We build pipelines that comply with the GDPR, NIS2 and sector-specific frameworks.

GDPR-compliant processor chain

For each external translation engine (DeepL, OpenAI, Google) we arrange the appropriate data processing agreement and data residency settings. Personal data is pseudonymised before transmission wherever possible, and we switch off logging on the provider side by default where the service allows it.

EU hosting and data residency

DeepL is originally EU-hosted, Microsoft Translator offers EU-only routing, and Azure OpenAI Service runs in West Europe. For clients with strict requirements we build an on-premises or private cloud variant with self-hosted models, so that no segment ever reaches an external API.

Medical and legal compliance

For MDR package leaflets and registered product information, we follow EMA templates and QRD guidelines. For legal translations, we set up a revision workflow with sign-off capability and a full audit trail per segment. No anonymous machine output without a responsible human.

Audit trail and explainability

Every translation is traceable: which model produced which segment, which glossary entries were active, what QE score it received, and who carried out post-editing. For compliance questions or disputes, we can reproduce the complete chain of decisions for each segment.

Concrete scenarios we build

AI translation software only delivers value when it fits a workable flow. These are examples of pipelines we can deliver for different sector profiles.

Webshop with 8,000 SKUs, five languages

An EU retailer with a Magento or Shopify store wants product content in Dutch, English, German, French and Spanish without manual export cycles. Pipeline: a PIM webhook fires on every SKU change, NMT with a brand glossary, automatic SEO meta mapping, post-editing only for segments with a QE score below the threshold, and publication back to the store within the hour. Delta detection ensures only changed fields go through the engine again.

Legal contract management for a law firm

An international practice translates NDAs, MSAs and shareholder agreements between Dutch, English and German. Pipeline with TBX terminology locking for legal terms, DNT tags on party names and amounts, and mandatory review by a bilingual lawyer. Audit trail per clause for liability purposes, and integration with a document management system such as HighQ or iManage.

Medical device: leaflets in 24 EU languages

An MDR-regulated manufacturer needs instructions for use (IFUs) in all EU languages for every product release. Pipeline with EMA template validation, MedDRA terminology via TBX, back-translation per market, and human review by a medically trained linguist. Audit trail in line with notified body requirements, and automatic version comparison when text changes.

Live chat translation for SaaS support

A Dutch SaaS provider supports B2B customers in Germany, France and Spain. Implementation: a translation proxy for the helpdesk (Zendesk or Intercom) that translates incoming customer messages into Dutch and outgoing agent replies into the customer's language. A glossary with product and feature names, and a correction loop whenever the COMET score falls below the threshold.

Why Appfront for AI translation software

Pipeline thinking, not plugin thinking

We don't build a loose translate button; we build an end-to-end content flow from source CMS to target market. The translation engine is a component, not the product. That difference determines whether you still have to intervene manually a year from now.

Domain expertise by sector

E-commerce localisation requires different choices from an MDR leaflet or a law firm's contract. We make those choices explicit and align the architecture with your quality and compliance requirements, rather than applying a one-size-fits-all stack.

EU-first infrastructure

For clients in the EU, we arrange EU hosting, data residency, GDPR-compliant data processing agreements as standard and, where needed, a private cloud or on-premise variant. No US cloud unless you explicitly choose it.

Frequently asked questions about AI translation software

What is the difference between NMT and LLM translation?
Neural machine translation (DeepL, Google, Microsoft) is optimised for sentence-level translation, fast and cost-effective. LLMs such as GPT-4o and Claude excel at stylistic translation, idioms, marketing tone and context across multiple sentences. For product catalogues and support cases, NMT is often sufficient; for marketing copy, legal context or cultural adaptation, we deploy LLMs. In practice we combine both.
How is terminology kept consistent across thousands of products?
Via a centrally managed TBX (Term Base eXchange) terminology database. NMT engines such as DeepL Pro accept glossary uploads that enforce that certain source terms always receive the same target terms. In LLM pipelines, we dynamically inject relevant glossary entries into the prompt. A Translation Memory (TMX) also stores previously approved translations for reuse.
Can AI translation handle MDR patient information leaflets or medical records?
Yes, provided the right pipeline architecture is in place. For MDR- and EMA-regulated content, we use EMA templates, MedDRA terminology, mandatory review by a medically qualified linguist, back-translation as a QA step, and a full audit trail per segment. The machine does the heavy lifting; accountability and final sign-off rest with a human.
What is COMET and why do you use it?
COMET (Crosslingual Optimised Metric for Evaluation of Translation) is a neural quality estimation metric that predicts a quality score per segment without a reference translation. We use COMET to automatically determine which segments can be published directly and which need to go to a post-editor. We sometimes also use BLEU and chrF for batch evaluation against a reference.
Does this work with Trados, MemoQ or Phrase?
Yes. CAT tools such as Trados Studio, MemoQ and Phrase TMS have mature APIs and support XLIFF, TMX and TBX. We build integrations in which segments automatically pass through the NMT or LLM pipeline, appear back in the CAT tool for post-editing, and are then published to your CMS, PIM or DMS.
How is data privacy handled with external translation engines?
For every external engine (DeepL, OpenAI, Microsoft, Google), we arrange the appropriate data processing agreement, choose EU data residency where available, and disable provider-side logging where supported. For clients with strict requirements, we build an on-premise or private cloud variant using self-hosted models such as NLLB or Marian, so no segment leaves your own infrastructure.
Can AI translate live chat in real time?
Yes. With DeepL or Microsoft Translator, we achieve sub-second latency for sentence-level translation. Architecture: a translation proxy between the customer and the agent in the helpdesk, with glossary injection for product names, a correction loop for low QE scores, and logging in the ticket history. Works with Zendesk, Intercom, Salesforce Service Cloud and custom helpdesk implementations.
What does an AI translation pipeline cost?
The investment depends on the number of languages, content volume, integrations and compliance requirements. An e-commerce pipeline with DeepL and a Shopify integration is a different project from an MDR-compliant medical pipeline with self-hosted models and an audit trail. We typically start with a clearly defined proof of concept to validate cost and quality before you commit to a larger project.
How long does it take for the pipeline to become operational?
A first working pipeline for a defined content type is usually up and running within four to eight weeks. Expansion to more languages, content types or compliance requirements happens in iterations. We work in sprints so you see tangible results along the way and can adjust the scope.

Bespoke AI translation software for your EU markets?

Discuss your content flow with us. We will map where AI translation delivers the most value for your product catalogue, legal practice or medical pipeline, free of charge and without obligation.

Schedule a conversation

On applatenmaken.com, our platform on custom software development, you can find more detail about Building translation software.

Edit content