Custom AI translation software development: from e-commerce localisation to legal and medical translation
Scaling to multiple EU markets quickly means translating thousands of product titles, legal clauses, package inserts or support conversations, at a quality that respects brand consistency, terminology and compliance. Generic translation buttons don't cut it. We build bespoke AI translation software that combines NMT engines, LLMs, terminology databases and human post-editing into a pipeline suited to your product catalogue, legal domain or medical records workflow.
Discuss your translation case View use casesWhy a bespoke AI translation pipeline delivers more than a button in your CMS
Generic translation plugins work for standalone sentences, but they fall short once terminology must stay consistent across thousands of SKUs, once a legal contract must not lose the distinction between "agreement" and "contract", or once package-insert text must match the registered product information word for word.
EU businesses exporting to six or more markets face a constant flow: new product pages on the webshop, contract amendments, policy documents, knowledge-base articles, support tickets, customer conversations, annual reports. Manual translation doesn't scale, and raw machine output without review damages the brand. The gain lies in a hybrid pipeline where neural machine translation (NMT) handles the first 80%, a terminology database ensures product names and legal terms don't slip, a quality-estimation model flags which segments need a human eye, and a translator with domain expertise reviews only those segments.
We build that pipeline around your existing systems: Shopify, Magento, Salesforce Commerce, Akeneo PIM, Contentful, an in-house DMS or a contract management platform. The translation engine isn't the product; the integration with your content flow is. We set up the architecture so Dutch content flows into EN, DE, FR, ES, IT and beyond without editors or lawyers having to copy and paste every release by hand.
Application areas where AI translation software makes a real difference
From product content to legal clauses and medical records, each domain has its own quality criteria, terminology and risk profile. We tailor the pipeline architecture accordingly.
E-commerce product content i18n
Product titles, descriptions, variants, attribute values, SEO meta tags and category texts in five to ten languages. Pipeline: PIM export to XLIFF, NMT translation with a brand glossary, automatic SEO keyword mapping per market and publication back to Akeneo, Magento or Shopify. Includes delta detection so only changed segments are retranslated.
Legal translation with terminology control
Contracts, NDAs, general terms and conditions, privacy statements and compliance documents. A TBX terminology database records how legal concepts ("ontbinding", "verzuim", "overmacht", "boetebeding") are translated consistently across all languages. DNT tags protect party names, amounts and clause numbers from being altered by the model.
Medical translation and compliance
Package inserts, IFUs, clinical protocols, patient forms and medical records. MDR- and EMA-compliant terminology via MedDRA and a bespoke TBX. Mandatory review by a medically trained linguist per segment, with an audit trail. Back-translation as an additional QA step where regulations require it.
Live-chat translation in customer support
A Dutch customer support agent who chats in real time with German, French or Spanish customers. Bidirectional translation with sub-second latency, glossary injection for product names and a correction loop whenever the QE score drops below a threshold. Integration with Zendesk, Intercom, Salesforce Service Cloud or your own helpdesk.
Meeting transcription and translation
Whisper for speech-to-text, followed by DeepL or Claude to translate the transcript into the languages of your participants. Speaker labelling, automatic summaries per participant language and export to Notion, SharePoint or your own knowledge base. Useful for international teams, customer calls and compliance meetings.
Document translation with layout preservation
PDF, Word and InDesign files that keep their formatting: columns, tables, footnotes, headers, footers and images stay exactly where they are. Pipeline: extraction to XLIFF with segment IDs, translation, reinjection via a document engine and an automated QA check for text overflow. No re-typesetting and no shifted layouts.
How Appfront builds an AI translation pipeline
Our starting question is never "which model should we use?" but "where does the content sit today, where does it need to go, and who reviews what?". The technical choices (NMT engine, LLM, post-editing tool) follow from the content flow and the quality requirements, not the other way round. A furniture brand's product catalogue can tolerate far more machine output than the leaflet for an infusion solution, so the pipeline choices differ fundamentally.
In practice, we first map your source and target systems, identify content types by quality class, and for each class choose the right mix of NMT (DeepL Pro, Google Translate Advanced, Microsoft Translator), LLM translation (OpenAI GPT-4o, Claude for stylistic translation), terminology enforcement (TBX import, Phrase TMS, MemoQ) and human post-editing (MTPE). We then build the orchestration: webhooks from your CMS or PIM, queueing, retry logic, audit logging and dashboards that let you track turnaround times, machine cost and post-editor effort.
The end result is a system that works invisibly: an editor publishes Dutch content, a few minutes later the other language versions are ready for review, a post-editor works through the flagged segments and the site or document goes live. No manual export-import cycles, no copying and pasting into a spreadsheet, and no more loose emails to a translation agency.
Architecture: NMT, LLMs and the orchestration in between
An AI translation pipeline is not a single model call. Quality comes from the combination of translation engine, terminology layer, quality estimation and review routing. These are the building blocks we use by default and tailor to your situation.
NMT vs LLM-based translation
Neural machine translation (DeepL, Google, Microsoft) is fast, inexpensive and strong at sentence level, which makes it ideal for product content and support cases. LLMs such as GPT-4o and Claude excel at stylistic translation, marketing tone, idiomatic expressions and domain context. We combine both: NMT as the baseline, and an LLM for segments where tone of voice or cultural adaptation matters most.
Glossary, TM and TBX pipelines
A Translation Memory (TMX) reuses previously approved translations; fuzzy matches save costs and guarantee consistency. A TBX terminology database enforces that "warranty" is always rendered as the agreed term and never drifts. We feed both into NMT engines via glossary APIs or into LLM prompts via dynamic retrieval, so agreed terminology does not slip away.
Quality estimation with COMET
COMET (and, in older setups, BLEU/chrF) calculates a quality score per segment without needing a reference translation. Segments above a threshold are published directly; those below are routed to a post-editor. The result: less review time, a predictable turnaround and measurable quality per content type and language.
DNT tags, back-translation and QA
Do-Not-Translate markers protect brand names, codes, legal article numbers and units of measurement. Back-translation (translating the target back into the source) is an additional QA step for medical and legal content, where divergent interpretations carry risk. Automated QA checks (length overflow, missing placeholders, number mismatches) run on every segment before publication.
From initial audit to production pipeline
Our approach to AI translation projects runs in four phases. Each phase is clearly defined, delivers something tangible and allows you to adjust course along the way.
Content audit and language strategy
We take stock of content types, languages, volumes, quality requirements per type and existing TM/TBX assets. We map where your content currently lives (CMS, PIM, DMS, helpdesk) and where it needs to go.
Engine selection and proof of concept
We benchmark two or three NMT and LLM combinations on your own content (anonymised), measure BLEU/COMET scores and post-editing effort, and select the stack that offers the best balance of cost, speed and quality.
Integration and MTPE workflow
Building API integrations with your CMS, PIM or helpdesk, setting up TM/TBX pipelines, integrating with CAT tools (Trados, MemoQ, Phrase), and configuring post-editing queues for your in-house or external linguists.
Monitoring and ongoing development
Dashboards for turnaround time, machine cost, post-edit distance and QE trends. Periodic retraining of glossaries, expansion into new languages and refinement of QA rules based on post-editor feedback.
Test your idea first: a working prototype in 1 day
With OneDayBuild, we turn your idea into something tangible in one day for €1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.
Explore OneDayBuild →Technology we use for AI translation software
The stack depends on content type, languages and compliance requirements. For pure e-commerce localisation we often choose a DeepL Pro API baseline with a Phrase TMS layer on top. For legal and medical content we shift to a setup with stronger terminology control and a bigger role for LLMs in stylistic and context-aware translation. For live chat, sub-second latency is a hard requirement, so we choose different models than for batch document translation.
All components preferably run on EU infrastructure: DeepL is hosted in Germany, Microsoft Translator offers EU data residency, and we route LLM calls through Azure OpenAI Service in West Europe or an EU-based Claude deployment where available. For sectors with the strictest requirements (healthcare, finance, government) we build an on-premises or private cloud variant with self-hosted models (NLLB, Marian, Mistral fine-tunes).
Translation data and privacy: what you need to arrange
Translation content contains personal data and trade secrets and, in legal and medical contexts, even special category personal data. We build pipelines that comply with the GDPR, NIS2 and sector-specific frameworks.
GDPR-compliant processor chain
For each external translation engine (DeepL, OpenAI, Google) we arrange the appropriate data processing agreement and data residency settings. Personal data is pseudonymised before transmission wherever possible, and we switch off logging on the provider side by default where the service allows it.
EU hosting and data residency
DeepL is originally EU-hosted, Microsoft Translator offers EU-only routing, and Azure OpenAI Service runs in West Europe. For clients with strict requirements we build an on-premises or private cloud variant with self-hosted models, so that no segment ever reaches an external API.
Medical and legal compliance
For MDR package leaflets and registered product information, we follow EMA templates and QRD guidelines. For legal translations, we set up a revision workflow with sign-off capability and a full audit trail per segment. No anonymous machine output without a responsible human.
Audit trail and explainability
Every translation is traceable: which model produced which segment, which glossary entries were active, what QE score it received, and who carried out post-editing. For compliance questions or disputes, we can reproduce the complete chain of decisions for each segment.
Concrete scenarios we build
AI translation software only delivers value when it fits a workable flow. These are examples of pipelines we can deliver for different sector profiles.
Webshop with 8,000 SKUs, five languages
An EU retailer with a Magento or Shopify store wants product content in Dutch, English, German, French and Spanish without manual export cycles. Pipeline: a PIM webhook fires on every SKU change, NMT with a brand glossary, automatic SEO meta mapping, post-editing only for segments with a QE score below the threshold, and publication back to the store within the hour. Delta detection ensures only changed fields go through the engine again.
Legal contract management for a law firm
An international practice translates NDAs, MSAs and shareholder agreements between Dutch, English and German. Pipeline with TBX terminology locking for legal terms, DNT tags on party names and amounts, and mandatory review by a bilingual lawyer. Audit trail per clause for liability purposes, and integration with a document management system such as HighQ or iManage.
Medical device: leaflets in 24 EU languages
An MDR-regulated manufacturer needs instructions for use (IFUs) in all EU languages for every product release. Pipeline with EMA template validation, MedDRA terminology via TBX, back-translation per market, and human review by a medically trained linguist. Audit trail in line with notified body requirements, and automatic version comparison when text changes.
Live chat translation for SaaS support
A Dutch SaaS provider supports B2B customers in Germany, France and Spain. Implementation: a translation proxy for the helpdesk (Zendesk or Intercom) that translates incoming customer messages into Dutch and outgoing agent replies into the customer's language. A glossary with product and feature names, and a correction loop whenever the COMET score falls below the threshold.
Why Appfront for AI translation software
Pipeline thinking, not plugin thinking
We don't build a loose translate button; we build an end-to-end content flow from source CMS to target market. The translation engine is a component, not the product. That difference determines whether you still have to intervene manually a year from now.
Domain expertise by sector
E-commerce localisation requires different choices from an MDR leaflet or a law firm's contract. We make those choices explicit and align the architecture with your quality and compliance requirements, rather than applying a one-size-fits-all stack.
EU-first infrastructure
For clients in the EU, we arrange EU hosting, data residency, GDPR-compliant data processing agreements as standard and, where needed, a private cloud or on-premise variant. No US cloud unless you explicitly choose it.
Frequently asked questions about AI translation software
Bespoke AI translation software for your EU markets?
Discuss your content flow with us. We will map where AI translation delivers the most value for your product catalogue, legal practice or medical pipeline, free of charge and without obligation.
Schedule a conversationOn applatenmaken.com, our platform on custom software development, you can find more detail about Building translation software.