Which speech recognition engine is best?
There is no universal best option. The right choice depends on the use case, language, audio environment and privacy requirements. For real-time Dutch dictation, Deepgram often scores well in our benchmarks; for diarisation and summarisation, AssemblyAI works nicely; for healthcare and government clients with EU residency requirements, Azure Speech is frequently the winner; and for applications where audio must not leave the device, we opt for on-device Apple Speech or Android SpeechRecognizer, or a self-hosted Whisper model. Our first step is always a short benchmark using your own real audio, not a recommendation based on a vendor brochure.
May we use cloud speech recognition for patient data?
In principle, yes, provided you have a proper data processing agreement, audio stays within the EU, sub-processors align with your GDPR policy, and a DPIA has been carried out. Some healthcare organisations opt for on-premise or dedicated cloud for broader reasons, for example with particular patient groups (psychiatry, paediatrics, oncology). We advise per project on what fits and work with the institution's information security officer, linking to the existing NEN 7510 controls.
Can a speech recognition app work offline?
Yes. We build offline mode using Apple Speech (iOS), Android SpeechRecognizer (Android) or a lighter Whisper model on-device. Accuracy is slightly lower than cloud speech-to-text for long recordings, but in many domains it is the right choice: short dictation snippets, voice commands, accessibility captions, and legal consultations where the audio must not leave the device. We can also work in a hybrid way: on-device when there is no connection, cloud when there is.
How does the app handle multiple speakers?
Diarisation is a separate task layered on top of speech recognition. Some engines (AssemblyAI, Deepgram, Azure Speech) handle it reasonably well out of the box. For meetings, podcasts and consultations, we often add a correction interface where a user can fix a misattributed speaker with a single tap. We measure diarisation quality separately from Word Error Rate in our benchmarks.
Does speech recognition work with specialist jargon and abbreviations?
Not good enough on its own, which is one of the reasons a custom app often pays off compared with a generic off-the-shelf solution. We work with domain dictionaries (custom vocabularies) in which we include specialist terms, abbreviations, medication names or legal references in advance. For more demanding applications, we fine-tune a custom speech model in Azure Speech or train our own Whisper variant. Users lose trust in the app if the terms that matter in their profession are consistently rendered incorrectly.
How do you handle the EU AI Act for our speech recognition?
We carry out an AI Act classification up front: low risk (general transcription), limited transparency obligation, or high risk (speech recognition as input for medical diagnosis, legal decision support, or HR assessments). For high-risk applications, we take care of the mandatory risk management documentation, data quality requirements, transparency towards the end user, and the human oversight mechanism. We align with an existing ISMS rather than insisting on reorganising everything.
Can you integrate with our EHR or DMS?
Yes. We have experience with HiX, Epic, ChipSoft, CGM and Cura on the EHR side, and with NetDocuments, iManage, Legalsense and SharePoint on the DMS side. The integration typically runs through an HL7 FHIR interface for healthcare or a REST API for DMS. For CRM integrations, we work with Salesforce, HubSpot, Dynamics and Dutch alternatives. If your system is exotic or built in-house, we integrate against a custom API, which is something we're good at as an agency.
Do you work with Suki, DeepScribe or Abridge?
We regularly see these packages come up in scribe projects. They are strong products for the US market, but in the Dutch healthcare context they present two structural challenges: their Dutch-language accuracy falls short of their English-language counterparts, and data residency does not always sit where Dutch healthcare organisations need it. For clients where it fits, we integrate with them; where it does not, we build the scribe layer to measure with an EU-resident engine and a NEN 7510-compliant architecture.
Who owns the code, models and data?
You do. We hand over the full source code, custom-trained models, build pipelines and deployment scripts. Audio and transcripts remain in your cloud or with a hosting provider of your choice. If you later want to continue with another agency or take it in-house, that is possible, as there is no technical lock-in. We earn our keep through good work that lasts, not by keeping clients locked in.
What determines the cost and lead time of an app like this?
The biggest cost drivers are the complexity of the domain (general transcription versus a deeply integrated scribe app), the choice between cloud and on-device models, the number of speakers and languages, and the compliance layer the domain requires (a high-risk AI Act application needs considerably more documentation and logging than a low-risk minutes app). We work in sprints with fixed sprint budgets, so you can steer on scope sprint by sprint. An initial conversation quickly gives a good picture of the range and a realistic timeline.