Which text-to-speech voice sounds best in Dutch?
At present, ElevenLabs scores highest in our blind listening panels for brand and audiobook work. For functional applications (magazine audio, accessibility layers), Azure Speech, Google Cloud TTS and Amazon Polly perform very well at noticeably lower cost. For maximum privacy we look at self-hosted open-source TTS (Coqui, Bark, F5-TTS). Our first step is always a short listening benchmark on your own texts, not a recommendation based on a vendor brochure.
Can a read-aloud app work offline?
Yes. We build offline mode with Apple AVSpeechSynthesizer (iOS), Android TextToSpeech (Android) or a lighter on-device TTS model. The voice is slightly less rich than cloud TTS, but for aeroplanes, trains passing through tunnels, accessibility apps without a permanent connection or children's books on holiday, it is the right choice. We often work hybrid: on-device for the basic work, premium cloud voice for highlight moments.
How does the app handle proper names and specialist jargon?
We build a pronunciation dictionary in which editors or administrators can tune proper names, place names, specialist jargon or medication names, using phonetic spelling or explicit IPA pronunciation. For more demanding applications, we fine-tune a custom neural voice in Azure Speech or a bespoke open-source model. Readers quickly lose trust in a read-aloud app if a recurring author's name is consistently mispronounced.
How do you handle the Copyright Act when reading articles or books aloud?
We check in advance which rights the publisher already holds, arrange additional agreements where necessary, and build into the architecture that an article can only be read aloud once its rights status is 'green'. For educational use, the educational exception sometimes works differently, and in that case we advise you on how to stay within it. For voice cloning we always require written consent and attestation from the original speaker.
How do you handle the AI Act for synthetic voice?
We begin with an AI Act classification. A general read-aloud app typically falls under the "limited transparency obligation": the user must know that the voice is synthetic. We implement that disclosure subtly in the app (an "AI voice" label next to the voice selection, a note during onboarding) without spoiling the reading experience. Voice cloning requires an additional layer of transparency and consent.
Do you work with Readspeaker or existing read-aloud providers?
We come across them regularly. Where it fits, we integrate with them rather than building our own read-aloud layer. But for publishers who see their voice identity as a distinguishing feature, or who want voice cloning or a bespoke brand voice, we build the TTS layer to measure with a suitable voice engine, with the copyright and AI Act layer properly arranged.
Can you integrate with our editorial CMS or LMS?
Yes. We have experience with WoodWing, Drupal, Sanity, Storyblok, Prismic and Contentful on the editorial side, and with Moodle, Brightspace, Magister and Itslearning on the LMS side. The integration typically runs via a REST or GraphQL API, with webhooks for "new article" events so the read-aloud audio is ready as soon as the article is published.
Who owns the code, voice settings and data?
You do. We deliver the full source code, the pronunciation dictionary, build pipelines and deployment scripts. Text, audio and listening behaviour remain in your cloud or with a hosting partner of your choice. No technical lock-in. For voice-cloned brand voices, the rights and the recording are contractually assigned to you. We earn our keep through good work that keeps you coming back.
What determines the cost and lead time of an app like this?
The main cost drivers are the complexity of the domain (a general read-aloud layer versus a deeply integrated accessibility app with an EAA project), the choice between cloud and on-device TTS, the number of voices and languages, and the compliance layer the domain requires. We work in sprints with fixed sprint budgets, so you can steer scope sprint by sprint.