Service · AI development

Custom speech recognition app development.

A custom app that converts speech to text, for medical dictation, minutes, call centre transcriptions, legal files or accessibility. Together we choose the right engine (Whisper, AssemblyAI, Deepgram, Azure Speech, Google STT, or on-device Apple/Android), build a workflow around it that fits your domain, and handle the GDPR, NEN 7510 and AI Act layer from the start. Not a demo that fails as soon as a doctor starts dictating in specialist jargon, but an app that works in the practice where it is used.

Whisper / Deepgram / AssemblyAIOn-device or cloudDiarisationGDPR & NEN 7510EHR & CRM integration

Speech recognition is more than an API key.

If you want to build a speech recognition app today, you can call a Whisper endpoint or create a Deepgram account and retrieve your first transcription within an hour. That is the easy half. The difficult half begins as soon as the app has to run for real users in a real working environment: a surgeon dictating while wearing a face mask, a note-taker in a meeting room with four speakers talking over each other, a call centre agent in an open-plan office, a solicitor dictating a memorandum in a moving car. In those settings, generic speech recognition packages fall short straight away.

A good speech recognition app accounts for dialect, accent, specialist jargon, background noise, several speakers at once, poor microphones and users who do not always speak Dutch. It does so within the limits the law sets out, namely the GDPR and, since 2026, the AI Act as well, and without audio containing patient information or customer data quietly travelling through a US cloud. That is exactly what we build every day in our AI development work.

What sets us apart: we only choose the speech recognition engine once we know where it has to work. For a medical scribing app, Whisper via OpenAI often falls away on data residency grounds, and we opt for Azure Speech in West Europe with custom model fine-tuning on medical terminology. For a Dutch-language minutes app with diarisation, Deepgram Nova outperforms generic alternatives in our benchmarks. For an accessibility app without a permanent connection, we choose Apple Speech or Android SpeechRecognizer on-device, even if that costs a few percentage points of accuracy, because users gain nothing if the transcription only appears once they're back on Wi-Fi.

Three typical forms.

In our experience, speech recognition apps fall into three categories, each with a different focus on accuracy, speed and privacy. Which form suits you depends on the user and the domain. You often start with the first and grow into a combination.

Real-time dictation · live transcription

Real-time speech-to-text

The user speaks and sees the text appear immediately. This is typical for medical dictation, accessibility (live captions for deaf and hard-of-hearing people in a conversation), legal dictation during a case review, or voice commands in a field app. We work with streaming APIs from Deepgram, Azure or Google with sub-second latency, or with on-device models where privacy or the absence of coverage requires it. We include domain dictionaries so that terms such as "PCI-DSS", "epicarditis" or "partnership with fiscal unity" do not come back unreadable.

Streaming STTSub-second latencyDomain dictionaryOn-device option
Batch transcription · diarisation and summarisation

Asynchronous transcription and analysis

The user records a conversation, meeting, podcast or customer service call, and the app returns the processed text, with speakers separated, timecodes included, and optionally a summary produced by an LLM layered on top. Used for minutes apps, call centre quality monitoring, legal case documentation and podcast production. We often pair a speech recognition layer with an LLM layer that turns the transcript into a list of decisions, an action overview, a compliance report or a customer sentiment report.

DiarisationLLM summaryTimecodesMulti-language
Domain-specific · healthcare, legal, call centre

Domain-specific scribe app

A deeply integrated scribe app for a single domain: for example, a medical scribing app that listens in during a consultation, fills in the SOAP structure, and places the text directly into the EHR (HiX, Epic, ChipSoft, CGM). Or a legal dictation app that structures the conversation into a memo template and connects to a DMS (NetDocuments, iManage, Legalsense). Or a call centre app that transcribes every call, scores it against compliance criteria, and connects to your CRM. For healthcare contexts we work within the framework outlined on our healthcare software page and ensure NEN 7510 compliance from day one.

EHR integrationSOAP structureNEN 7510DMS integration

What you get at the end.

A production-ready speech recognition app for the domain that matters — for your clinicians, lawyers, agents or end users — plus everything needed to manage, audit and extend it. No black box: you own the code, models and data. We only handle operations if you want us to.

  • Mobile and/or web appNative iOS / Android or cross-platform (Flutter / React Native), with a web version where it fits — tested on the actual microphones and headsets used in your domain.
  • Speech recognition layer with domain vocabularyAn integrated STT stack (Whisper, AssemblyAI, Deepgram, Azure or Google) with custom vocabulary for your field — medical, legal, financial or technical.
  • Diarisation and multi-speaker supportAutomatic speaker identification for meetings, consultations or customer calls, with an optional manual correction interface for the cases the engine does not separate well on its own.
  • Offline / on-device modeApple Speech or Android SpeechRecognizer for situations without a connection, or a lighter on-device Whisper model for offline batch processing, with cloud fallback when a connection is available.
  • LLM layer for structuringOptional post-processing with an LLM (Azure OpenAI, Anthropic, Mistral, or open-source in your own environment) to turn the transcript into a SOEP note, minutes or a memo.
  • Integrations with your stackEHRs (HiX, Epic, ChipSoft, CGM, Cura), DMS systems (NetDocuments, iManage), CRMs (Salesforce, HubSpot, Dynamics) and data platforms. Set up properly once, it then runs within your workflow.
  • EU residency and privacy layerData processing in Western Europe, configurable retention, encryption at rest and in transit, and the option to delete audio immediately after transcription if your domain requires it.
  • DPIA, AI Act compliance and audit trailDocumentation and logging needed to demonstrably account for GDPR, NEN 7510 and the AI Act as a healthcare provider or law firm — including bias evaluation across accents and languages.
  • Management contract (optional)Monitoring of transcription quality per cohort, model updates, security patches and further development based on user feedback.

Who we build speech recognition apps for.

Eight patterns we see again and again in our AI projects. If you recognise your organisation or use case in one of them, we'd be glad to talk further, even if you're still unsure whether a custom app is the right answer or whether an existing off-the-shelf solution will do.

Healthcare

Medical scribing and EHR dictation

Doctors who want to dictate consultation room notes during or shortly after a consultation, with automatic SOAP structuring and direct entry into the EHR. We work with HiX, Epic, ChipSoft and CGM, with the NEN 7510 layer built in. Depending on the organisation's policy, audio is either stored encrypted after transcription or deleted straight away. For sensitive specialisms such as psychiatry, oncology and paediatrics, we discuss explicitly what may and may not be passed to an external STT provider.

Minutes

Meeting and interview transcription

An app for secretariats, journalism, research or HR that turns recorded meetings or interviews into structured text, with speakers separated and action items listed separately. For Dutch-language content, we critically assess which speech-to-text engine performs best on accents and background noise. The differences are currently larger than many English-language reviews would suggest.

Call centre

Call centre transcription and quality monitoring

A layer on top of your telephony or contact centre platform (Genesys, Amazon Connect, Five9, Talkdesk) that transcribes every customer call, scores it against compliance criteria and sentiment, and passes the results on to your CRM. Suited to insurers, banks, energy suppliers and B2C brands whose call volumes are currently only spot-checked by a quality team.

Accessibility

Live captioning and accessibility app

An app that provides live captions for deaf and hard-of-hearing people, in a conversation at the table, a meeting, a lecture or an event. Often with separate "speaker" and "listener" roles. We take a close look at the balance between cloud accuracy and latency, as captions that lag by several seconds are useless in a conversation.

Field service

Voice commands for engineers and technicians

Field apps in which engineers can dictate hands-free status updates, photo captions or customer notes, which is handy when you're hanging in a meter cupboard or standing on a roof wearing gloves. Often integrated with an existing field app (Salesforce Field Service, ServiceNow FSM) or with a custom solution of the kind our AI consulting clients regularly build in-house.

Legal

Dictation for lawyers and notaries

A dictation app for memos, court documents, deed fragments and client meeting notes. It handles legal vocabulary and abbreviations ("BW", "Rv", "WWFT") correctly, includes a button to mark confidentiality, and features a DMS integration. For cases classified as high-risk under the AI Act, we carry out the impact assessment upfront.

Podcast / media

Production pipeline for podcasts and interviews

An internal tool that converts raw audio into transcripts with separated speakers and timecodes, plus an initial summary for show titles and show notes. It connects to an editorial workflow in which editors cut and paste based on the text rather than the audio timeline, cutting production time by a factor of several.

Education / languages

Language teaching and pronunciation feedback

Language learning apps (Dutch lessons for non-native speakers, English for schoolchildren) that not only recognise what a learner says but also how well they pronounce it. For this we combine speech recognition with a pronunciation scoring component. We work similarly on tools for speech therapy and rehabilitation.

Which speech recognition engine fits.

We don't pick a favourite in advance. We only choose once we know the use case, the domain, the languages and the privacy requirements. A brief overview of each candidate, which we go through with you in an initial session.

Cloud · multi-language · open-source core

Whisper (OpenAI / Azure / self-hosted)

Broad language coverage, strong on general Dutch speech, and available as an open-source model for on-device or self-hosted deployment. Quick to get running via OpenAI, but for healthcare and legal work we more often use Azure OpenAI in Western Europe, or self-host when audio must never leave your environment. Adding diarisation requires extra setup.

Cloud · diarisation · strong Dutch

AssemblyAI

Good at diarisation, summaries, sentiment and speaker labels. Works well for minutes, podcast and call centre applications. Available in EU regions; for healthcare or legal contexts we discuss data residency and sub-processors explicitly.

Real-time · low latency · Nova model

Deepgram

In our benchmarks, this is often the winner for real-time Dutch-language speech, with sub-second latency and competitive pricing per hour of audio. The Nova model recognises dialects and accents surprisingly well. Suitable for live captioning, call centre monitoring and streaming dictation. An EU server location is configurable.

EU residency · enterprise · custom models

Azure Speech

For healthcare and government clients, this is often the practical choice: EU residency in West Europe, integration with Azure AD, custom speech models for medical or legal terminology, and it fits neatly into an existing Microsoft stack. It excels on the compliance requirements we encounter every day with our healthcare clients.

Multi-language · Google Cloud Speech-to-Text

Google STT

Strong in languages that few other providers support well, and well integrated into a Google Cloud stack. For clients with a data platform on BigQuery and audio buckets in Google Cloud Storage, Google STT often sits closest to the data, which saves on architecture and data transfer.

On-device · privacy-first · zero cloud transfer

Apple Speech & Android SpeechRecognizer

For situations where audio must not leave the device, such as a psychiatrist with a patient recording, a lawyer with a conversation under professional privilege, or a field app without connectivity, we opt for the native on-device speech frameworks from Apple and Android. They are well suited to dictation fragments and commands, with the decisive advantage that no audio travels over a network.

Not yet sure about a large project?

Test your idea first: a working prototype in 1 day

With OneDayBuild, we turn your idea into something tangible in one day for €1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.

Explore OneDayBuild →

How a speech recognition project runs.

1

Introduction & use case definition

A conversation in which we understand who will use the app, in which environment, and under which privacy and compliance constraints. Where possible, we literally listen in: a doctor dictating, a minute-taker recording a meeting, a call centre agent handling a call. Audio quality, dialect, specialist jargon, number of speakers and background noise together determine the profile the speech recognition engine will have to cope with. Where possible, we start here with a short advisory phase, as a full-blown custom app is sometimes not even needed and a targeted tool or integration is the right first step. See also our page on AI advice.

2

Engine selection & benchmark

We run two to four candidate engines side by side on real audio from your domain: no demo audio, no English test corpus, just your actual work. A doctor dictates for ten minutes, a minute-taker supplies two meetings, a call centre provides a typical day's batch. We measure Word Error Rate per cohort, latency, cost per hour of audio, and qualitative errors that a human hears but Word Error Rate does not capture. The choice of engine lands in a short decision matrix that we work through together with you.

3

DPIA, AI Act & data architecture

For healthcare, legal and HR contexts, we start the DPIA in parallel with the build, not as a separate step at the end. For applications that would count as high-risk under the AI Act (such as medical diagnostic support or legal decision support), we carry out the AI Act impact assessment in advance and settle the logging requirements from the outset. Audio flow, retention policy, encryption and sub-processor selection come together in a single architecture document that can also be discussed with your compliance officer or works council.

4

Building in sprints

A working build every two weeks on TestFlight, the Google Play internal track or a staging environment. We build incrementally: first the core transcription flow, then diarisation, then a domain dictionary, then LLM structuring, then EHR or DMS integration. Continuous evaluation runs against a growing benchmark set of audio. When the engine receives a release update or we adjust the vocabulary, we know straight away whether quality improves or worsens for each cohort. Pricing and logic steps are covered by automated tests, because a speech recognition app that quietly mixes up the SOAP fields is worse than one that is occasionally a little slower.

5

Rollout, training & ongoing development

Phased rollout: start with a pilot group, then expand. A short training session, a guide with dos and don'ts (including how the user's own speaking style affects transcription quality), and ongoing maintenance for model updates, OS updates and security patches. We monitor transcription quality per cohort and keep adjusting. A speech recognition app that works well on day one but isn't maintained afterwards degrades faster than most other software.

Compliance is not an afterthought.

Speech is sensitive data. For healthcare and legal applications, a well-built speech recognition app is by definition also a well-documented, GDPR-compliant and AI Act-compliant app. We build that layer in from the first sprint, not as a paperwork exercise at the end.

GDPR & DPIA

Audio recordings are special category personal data once they involve a doctor-patient or employee-HR conversation. We carry out a targeted DPIA, arrange data processing agreements with sub-processors, and build the retention and deletion policy into the app.

NEN 7510 for healthcare

For healthcare organisations we follow the NEN 7510 control set from the outset: access management, logging, encryption and supplier management. We align our development with the hospital's existing ISMS. See also our page on NEN 7510-compliant software.

AI Act & high-risk classification

The AI Act has been in force since 2026. A general meeting-notes app is low risk, but once transcription becomes input for medical, legal or HR decisions, the system may fall into the high-risk category. We carry out the classification up front and build the logging, transparency and human-oversight requirements into the design.

EU residency & sub-processors

By default we choose cloud regions in Western Europe (Azure West Europe, AWS Frankfurt, GCP europe-west4). For more demanding clients, such as academic hospitals, law firms and government bodies, we work with on-premise or dedicated-cloud setups in which no sub-processor outside the EU has access to the audio.

Bias evaluation for accent and language

Speech-to-text models perform worse for non-native speakers, regional dialects and children's voices. For accessibility, call centre and education apps, we carry out a bias evaluation on the target group in advance and adjust the engine choice or vocabulary fine-tuning accordingly. We record our findings openly.

Audit trail & human oversight

Every transcription has traceable provenance: who recorded it, which model processed it, who edited it, and when it reached the EHR or DMS. For high-risk applications, we build explicit "human-in-the-loop" steps in before any text enters the workflow.

Frequently asked questions.

What clients usually want to know before we start on their speech recognition app.

Which speech recognition engine is best?
There is no universal best option. The right choice depends on the use case, language, audio environment and privacy requirements. For real-time Dutch dictation, Deepgram often scores well in our benchmarks; for diarisation and summarisation, AssemblyAI works nicely; for healthcare and government clients with EU residency requirements, Azure Speech is frequently the winner; and for applications where audio must not leave the device, we opt for on-device Apple Speech or Android SpeechRecognizer, or a self-hosted Whisper model. Our first step is always a short benchmark using your own real audio, not a recommendation based on a vendor brochure.
May we use cloud speech recognition for patient data?
In principle, yes, provided you have a proper data processing agreement, audio stays within the EU, sub-processors align with your GDPR policy, and a DPIA has been carried out. Some healthcare organisations opt for on-premise or dedicated cloud for broader reasons, for example with particular patient groups (psychiatry, paediatrics, oncology). We advise per project on what fits and work with the institution's information security officer, linking to the existing NEN 7510 controls.
Can a speech recognition app work offline?
Yes. We build offline mode using Apple Speech (iOS), Android SpeechRecognizer (Android) or a lighter Whisper model on-device. Accuracy is slightly lower than cloud speech-to-text for long recordings, but in many domains it is the right choice: short dictation snippets, voice commands, accessibility captions, and legal consultations where the audio must not leave the device. We can also work in a hybrid way: on-device when there is no connection, cloud when there is.
How does the app handle multiple speakers?
Diarisation is a separate task layered on top of speech recognition. Some engines (AssemblyAI, Deepgram, Azure Speech) handle it reasonably well out of the box. For meetings, podcasts and consultations, we often add a correction interface where a user can fix a misattributed speaker with a single tap. We measure diarisation quality separately from Word Error Rate in our benchmarks.
Does speech recognition work with specialist jargon and abbreviations?
Not good enough on its own, which is one of the reasons a custom app often pays off compared with a generic off-the-shelf solution. We work with domain dictionaries (custom vocabularies) in which we include specialist terms, abbreviations, medication names or legal references in advance. For more demanding applications, we fine-tune a custom speech model in Azure Speech or train our own Whisper variant. Users lose trust in the app if the terms that matter in their profession are consistently rendered incorrectly.
How do you handle the EU AI Act for our speech recognition?
We carry out an AI Act classification up front: low risk (general transcription), limited transparency obligation, or high risk (speech recognition as input for medical diagnosis, legal decision support, or HR assessments). For high-risk applications, we take care of the mandatory risk management documentation, data quality requirements, transparency towards the end user, and the human oversight mechanism. We align with an existing ISMS rather than insisting on reorganising everything.
Can you integrate with our EHR or DMS?
Yes. We have experience with HiX, Epic, ChipSoft, CGM and Cura on the EHR side, and with NetDocuments, iManage, Legalsense and SharePoint on the DMS side. The integration typically runs through an HL7 FHIR interface for healthcare or a REST API for DMS. For CRM integrations, we work with Salesforce, HubSpot, Dynamics and Dutch alternatives. If your system is exotic or built in-house, we integrate against a custom API, which is something we're good at as an agency.
Do you work with Suki, DeepScribe or Abridge?
We regularly see these packages come up in scribe projects. They are strong products for the US market, but in the Dutch healthcare context they present two structural challenges: their Dutch-language accuracy falls short of their English-language counterparts, and data residency does not always sit where Dutch healthcare organisations need it. For clients where it fits, we integrate with them; where it does not, we build the scribe layer to measure with an EU-resident engine and a NEN 7510-compliant architecture.
Who owns the code, models and data?
You do. We hand over the full source code, custom-trained models, build pipelines and deployment scripts. Audio and transcripts remain in your cloud or with a hosting provider of your choice. If you later want to continue with another agency or take it in-house, that is possible, as there is no technical lock-in. We earn our keep through good work that lasts, not by keeping clients locked in.
What determines the cost and lead time of an app like this?
The biggest cost drivers are the complexity of the domain (general transcription versus a deeply integrated scribe app), the choice between cloud and on-device models, the number of speakers and languages, and the compliance layer the domain requires (a high-risk AI Act application needs considerably more documentation and logging than a low-risk minutes app). We work in sprints with fixed sprint budgets, so you can steer on scope sprint by sprint. An initial conversation quickly gives a good picture of the range and a realistic timeline.

Talk to us about your speech recognition app.

A no-obligation introductory call of half an hour. Tell us who will use the app, in which working situation and within which compliance context. We'll think along with you, give direction, and be honest about whether a custom app is the right answer. Sometimes the right first step is a short AI advisory phase; sometimes it really is a custom scribe or dictation app with the full build around it. You can also email us directly at fabian.vandijk@appfront.nl.

Edit content