Service · App development

Custom read-aloud app development.

A custom app that turns text into natural-sounding speech, for publishers who want their articles read aloud, for organisations supporting people with dyslexia or visual impairments, for children's content brands, audiobook and podcast start-ups, and for education providers offering course material as audio. Together we choose the right voice engine (ElevenLabs, Azure Speech, Google Cloud TTS, Amazon Polly or on-device alternatives such as Coqui, Bark and F5-TTS), build a reading experience around it, and handle the GDPR, EAA/WCAG, copyright and AI Act layer from the outset. No robotic voice that tires after three sentences, but a Dutch voice you can listen to for hours.

ElevenLabs / Azure / PollyOn-device or cloudHighlighting during playbackGDPR & EAA/WCAGOffline mode

Reading aloud is more than running text through a TTS pipeline.

Building a read-aloud app seems simple: send text to a text-to-speech API and get audio back. In an afternoon you have a prototype. But as soon as real users start using it, every hairline crack shows. A publisher with a thousand articles a week, a parent reading to a dyslexic child, a student listening to a textbook on their bike, a visually impaired reader with a newspaper: all of them expose the gaps. After a quarter of an hour the voice becomes tiring, proper names are mispronounced, quotations are read no differently from running text, the app stops when the screen turns off, and it does not work offline. That is before you even consider the rights to the text, child safety and the labelling requirement for synthetic voices under the AI Act.

A good read-aloud app takes these details into account. It chooses the right voice for the right content, highlights the sentence being read at that moment, remembers where you left off, switches smoothly between articles, and respects how a text is structured: headings, quotations, dialogue. It does this within the limits set by law: GDPR for reading behaviour, EAA and WCAG for accessibility, copyright law for the text and, since 2026, the AI Act for the voice. And without text containing personal data quietly travelling through a US cloud. That is what we build day in, day out in our app development work.

What sets us apart: we choose the voice only once we know what it has to do. For a dyslexia app for children, we often opt for a calm-sounding Dutch voice with adjustable speed; for a specialist publisher with opinion pieces, a more resonant ElevenLabs voice works well; for a daily newspaper with regional sections, we look for a voice that also pronounces Limburgish place names or Frisian proper names correctly; for an app for visually impaired users without a permanent connection, we choose an on-device speech engine: less rich, but always available. For the reverse direction, speech to text, see our page on building speech recognition apps.

Three typical forms.

In our experience, read-aloud apps fall into three categories, each with a different focus: voice quality, integration and audience. Which one suits you depends on who is listening and where the text comes from.

Editorial · publisher · newspaper audio

Read-aloud layer on existing content

The articles or magazines you already publish are read aloud in the app, with a natural-sounding Dutch voice, sentence-by-sentence highlighting in the text, a "pick up where you left off" button across devices, and an offline mode for on the go. Typical for specialist publishers, regional newspapers and quality magazines. We integrate with your editorial CMS (WoodWing, Drupal, Sanity, Storyblok, Prismic or an in-house editorial system). The link with our page on magazine apps is an obvious one: for publishers without their own app, read-aloud can quickly become the most distinctive feature.

CMS integrationAuto-publish to audioHighlight during read-aloudOffline reading mode
Accessibility · dyslexia · visually impaired

Accessibility read-aloud app

An app for people who cannot read or find reading difficult: pupils with dyslexia, adults with low literacy, people with a visual impairment, and people recovering from a stroke. We work closely with the target group, speech therapists, teachers and accessibility experts. It is about the details you won't find in marketing brochures: a reading view with a dyslexia-friendly typeface (Dyslexie or OpenDyslexic), adjustable speed and pauses, skip back by sentence, contrast modes, screen reader compatibility and word-by-word highlighting. For schools, schoolbook publishers and care institutions under the European Accessibility Act and WCAG 2.2.

Dyslexia typefaceEAA/WCAG 2.2Word-by-word highlightSpeed & pauses
Children's content · audiobooks · education

Read-aloud app for children, audiobooks or education

A dedicated read-aloud app for an audiobook publisher, a children's content brand, an educational library or a museum. Often with several voice options per story (father's voice, mother's voice, child's voice, dialect version), atmospheric sound effects for evocative scenes, and a parent section where reading time and progress are visible without any child data leaving the device. For education, it works well as an extension of our education apps: teaching material libraries that add read-aloud features to PDF or EPUB course materials.

Multi-voiceSound effectsChild-safe designPupil progress

What you get at the end.

A production-ready read-aloud app for your readers, pupils or customers, plus everything needed to manage, audit and extend it. No black box: you own the code, voice choices, texts and listening behaviour. We only take on management if you want us to.

  • Mobile app for iOS and AndroidNative (Swift / Kotlin) or cross-platform via Flutter or React Native. For accessibility apps, we more often recommend native, because of the deeper integration with VoiceOver and TalkBack.
  • Web listening mode (optional)A web component with the same voice choices, highlight mechanism and bookmark state that syncs with the mobile app.
  • TTS layer with voice choiceAn integrated TTS stack (ElevenLabs, Azure Speech, Google Cloud TTS, Amazon Polly or on-device) with multiple Dutch voices, optional dialect variants and a fallback chain.
  • Highlight while reading aloudSentence- or word-level highlighting during playback, with adjustable text size, contrast and typeface (including dyslexia-friendly alternatives). Compatible with VoiceOver and TalkBack.
  • Audio streaming with HLSFor longer content, a streaming pipeline with HLS segments, partial download for offline listening and bandwidth adaptation.
  • Offline / on-device modeApple AVSpeechSynthesizer, Android TextToSpeech or a lighter on-device TTS model for situations without a connection.
  • Integrations with your stackEditorial CMSs (WoodWing, Drupal, Sanity, Storyblok, Prismic, Contentful), audiobook formats (EPUB 3, PDF/UA), subscription systems (Recurly, Stripe Billing).
  • Editorial workflow for voice choicesA management environment in which a publisher or accessibility officer can fine-tune the voice, the pronunciation of proper nouns and the pauses for each article or chapter, without needing developers.
  • EU residency and privacy layerProcessing in Western Europe, configurable retention, encryption at rest and in transit, and, for children's content, strictly separated parent and child data.
  • DPIA, AI Act and EAA complianceDocumentation and logging to demonstrably account for GDPR, the AI Act and the European Accessibility Act, including an accessibility statement.
  • Management contract (optional)Monitoring of listening time, completion rate per article, voice quality per cohort, model updates and ongoing development. Fixed monthly fee, with several response-time levels.

Who we build read-aloud apps for.

Eight patterns we see again and again in our app projects. If you recognise your organisation in one of them, we would be happy to talk further.

Publishers

Magazine and newspaper with an audio layer

Publishers, magazines, regional newspapers and industry bodies who want to offer their articles as audio too, often because their readers prefer listening to reading during the commute or while exercising. We integrate with your editorial CMS, choose a Dutch voice that suits your brand, and build a reading experience with highlighted text, "pick up where you left off" across devices, and an offline mode for the train.

Dyslexia

Textbook publishers and dyslexia organisations

Textbooks that are also available as audio for pupils with dyslexia, with a dyslexia-friendly typeface, a calm voice that reads along sentence by sentence, and the option to have a word pronounced again. We work with teachers and speech therapists during scoping and test with the pupils themselves, not just with focus groups, to make sure the app works for the child who uses it.

Accessibility

Visually impaired and low literacy

Apps that read full newspapers or magazines aloud for people with a visual impairment, where screen readers handle the basic interaction but the reading experience is thin. We build a complement: a more natural voice, better navigation between articles, word-by-word highlighting for users with residual vision, and a button to listen to the summary. For people with low literacy, the same principle applies without a screen-reader layer.

Children's content

Read-aloud stories and children's media

Read-aloud apps for publishers of children's and young adult books and for brands with children's content. Often with several voices per story (father, mother, narrator), sound effects at the right moments, a bedtime mode and attention to screen time. Strictly child-safe: minimal data collection, no advertising aimed at children, parental controls, and an architecture that respects the European framework for child-oriented services.

Audiobook start-ups

Audiobook and podcast start-ups

For early-stage publishers that do not want to run full audio production: a read-aloud app with TTS as the foundation, and human narration only for the most valuable titles. We build the catalogue management, licence monitoring, the subscription system (Stripe Billing or Recurly) and the playback app, and make sure the voice choice can easily be upgraded later.

Accessibility officers

EAA trajectory with an accessibility statement

Organisations that fall under the European Accessibility Act and need to make their digital content accessible. We don't just design the read-aloud app; we also handle the accessibility statement, an EAA conformity assessment, a WCAG 2.2 test of the app, and a process that automatically checks new content for read-aloud suitability.

Education

Audio learning materials and LMS integration

Educational institutions and LMS vendors who want to offer course material as audio, for learners with a learning disability, for language teaching, for distance learners. We integrate with Moodle, Brightspace, Magister or Itslearning and make sure listening time counts towards learner analytics. See also educational app development.

Culture

Museum audio and city tours

A variant in which exhibition texts, city walks or historical sources are read aloud, with location-aware playback, a multilingual audience and a design that fits the brand identity. For cultural institutions that want to extend their opening hours through an audio experience at home.

Which text-to-speech engine suits you?

We don't pick a favourite in advance. We only choose once we know your audience, content, languages and privacy requirements. A brief overview of each option follows.

Premium · expressive · voice cloning

ElevenLabs

The standout for expressive Dutch voices. Suitable for brand productions where voice quality makes the difference, for audiobook experiments, and for brands that want their own voice identity through voice cloning. Bear in mind the AI Act requirements for synthetic voices. For voice cloning we always require a signed consent form and attestation from the speaker.

EU residency · enterprise · custom voices

Azure Speech (Text-to-Speech)

For publishers, government clients and healthcare organisations with EU data residency requirements, this is often the practical choice: a West European region, custom neural voices for brand voices, integration with Azure AD and alignment with a Microsoft stack that is often already in place. Strong Dutch-language voices and good SSML control over pauses, emphasis and pronunciation.

Multi-language · Google Cloud TTS

Google Cloud TTS

Broad language coverage and good neural voices, tightly integrated with a Google Cloud stack. For clients with BigQuery and assets in Google Cloud Storage, Google TTS sits close to the data, which saves on architecture and data transfer. The Dutch WaveNet voices are good enough for magazine and accessibility applications.

Workhorse · AWS stack

Amazon Polly

A solid workhorse TTS for clients already running on AWS: neural voices in Dutch, easy to connect to an S3 archive. Not the most expressive voice, but more than sufficient for functional read-aloud applications and easy to manage.

OpenAI · multiple voices · easy API

OpenAI TTS & Whisper stack

OpenAI offers TTS voices that work quickly through the same API as its well-known LLMs, which is handy for start-ups that want to launch fast. We often combine this with OpenAI Whisper on the listening side and an LLM layer for summaries. For clients with data residency requirements, we opt for Azure OpenAI in West Europe as an alternative.

On-device · privacy-first · zero cloud transfer

Apple AVSpeechSynthesizer & Android TextToSpeech

For situations where text must not leave the device, such as a dyslexia app with personal notes, a children's book on a plane, or a visual-impairment app without a permanent connection, we opt for the native on-device TTS frameworks. Less rich than ElevenLabs, but with no network transfer and no per-character costs. Often with a cloud fallback for premium voices.

Open source · self-hosted · fine-tuning in Dutch

Coqui TTS, Bark & F5-TTS

For clients who want maximum control, or who would rather not pay an external vendor per audio volume, we run open-source TTS models self-hosted: Coqui TTS, Bark, or the newer F5-TTS, which offers strong zero-shot voice cloning. Where needed, we fine-tune on Dutch speech. Infrastructure requirements are higher, but you keep full ownership and no text leaves your private network.

Web-based · Read Aloud

Web Speech API and built-in read-aloud features

For web products where a simple listening mode is enough, you can rely on the Web Speech API or the built-in read-aloud features of iOS, Android, Windows and macOS. These don't amount to a full read-aloud app, but for small-scale accessibility add-ons they are often the right starting point. We use them regularly in the prototype phase.

Not yet sure about a large project?

Test your idea first: a working prototype in 1 day

With OneDayBuild, we turn your idea into something tangible in one day for €1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.

Explore OneDayBuild →

How an engagement works.

1

Introduction & use case definition

A conversation to understand who will use the app, in what context, and with what requirements for accessibility, privacy and rights. Where possible, we listen in directly: a pupil with dyslexia shows how they currently read, an editorial team shows an article they want read aloud, a visually impaired reader demonstrates their screen reader workflow. Voice quality expectations, target audience, dialect and content structure together define the profile the TTS engine has to meet. Sometimes a targeted extension of your existing app is enough. See also AI development for the wider context.

2

Voice selection and listening benchmark

We place two to four candidate TTS voices side by side, using real content from your domain: not demo text, but your actual articles, chapters or course material. We have a listening panel from the target audience (teachers, visually impaired users, children, editorial staff) listen in, and measure intelligibility, listening fatigue after half an hour, proper nouns, dialect accuracy, and how the voice handles quotes and headings. The choice of engine is summarised in a short decision matrix.

3

Rights, DPIA and AI Act layer

For publishers and education, we start the rights layer in parallel with the build rather than as a separate step at the end. Which texts may be read aloud? How do we handle quoted excerpts? We take care of the DPIA for reading behaviour and parent-child links, and the AI Act classification for the TTS layer, including labelling of synthetic voices. The architecture document comes in a single piece that can also be discussed with your lawyer or compliance officer.

4

Building in sprints

A working build every two weeks on TestFlight, the Google Play internal track or a staging web version. We build incrementally: the basic read-aloud flow and voice choice first, then highlighting while reading and bookmark state, then offline mode, then CMS or LMS integration, then the pronunciation editing environment. Continuous evaluation against a growing benchmark set of text. For accessibility apps we run WCAG tests in parallel with the build, with real screen reader users in a final sprint.

5

Rollout, training & ongoing development

Phased rollout: first a pilot group, then a wider release. Short training for editorial or administrative staff, a guide on which voice works for which genre, and ongoing management for model updates, OS updates, security patches and voice updates. We monitor listening time, completion rate per article, and which words get corrected in the pronunciation editor, which feeds directly into the next pronunciation dictionary.

Accessibility and compliance from day one.

A read-aloud app touches four legal areas at once: GDPR for user data, the EAA and WCAG for accessibility, the Dutch Copyright Act for the texts being read aloud, and the AI Act for the voice that is produced. We address those layers from the first sprint, not as a closing paperwork exercise.

GDPR and reading behaviour

What someone reads or listens to is sensitive information, especially in dyslexia apps, accessibility apps and children's content. We carry out a targeted DPIA, arrange processing agreements with TTS providers, build a retention policy that keeps reading behaviour for as short a time as possible, and give users simple access and deletion rights.

EAA and WCAG 2.2

The European Accessibility Act sets requirements for many digital products, including those from publishers, e-commerce, banks and government. The read-aloud app itself must meet WCAG 2.2: screen reader compatibility, sufficient contrast, adjustable text and a logical focus order. We test with TalkBack, VoiceOver and NVDA, and we draft an accessibility statement that follows the Dutch government model.

Copyright Act & related rights

Text-to-speech conversion of an article or book touches two rights: the text right (under the Dutch Copyright Act, Auteurswet) and the right to audio performance (a related right). We make sure the publisher holds the rights to convert the text into audio, that quoted passages are handled correctly, and that written consent is always obtained for voice cloning. For educational use, the educational exception sometimes works differently.

AI Act & synthetic voice

The AI Act has been in force since 2026. A synthetic voice that sounds human is deepfake-like output and must be made recognisable, which is a transparency obligation towards the end user. We implement that labelling in the app and record in the documentation that the voice is a TTS voice. Voice cloning requires an additional layer of consent and attestation.

Child protection

Additional protection applies to content for children: minimal data collection, no targeted advertising, parental controls and age-appropriate design. We build child modes in which user data is not shared beyond what is strictly necessary, and maintain a processing inventory that complies with the European framework for services aimed at children.

EU residency & sub-processors

By default we choose data centres in Western Europe (Azure West Europe, AWS Frankfurt, GCP europe-west4). For stricter clients, such as academic publishers, government bodies and organisations for which data minimisation is a core requirement, we work with on-premises or dedicated cloud setups in which no sub-processor outside the EU sees the text or audio.

Frequently asked questions.

What clients usually want to know before we start.

Which text-to-speech voice sounds best in Dutch?
At present, ElevenLabs scores highest in our blind listening panels for brand and audiobook work. For functional applications (magazine audio, accessibility layers), Azure Speech, Google Cloud TTS and Amazon Polly perform very well at noticeably lower cost. For maximum privacy we look at self-hosted open-source TTS (Coqui, Bark, F5-TTS). Our first step is always a short listening benchmark on your own texts, not a recommendation based on a vendor brochure.
Can a read-aloud app work offline?
Yes. We build offline mode with Apple AVSpeechSynthesizer (iOS), Android TextToSpeech (Android) or a lighter on-device TTS model. The voice is slightly less rich than cloud TTS, but for aeroplanes, trains passing through tunnels, accessibility apps without a permanent connection or children's books on holiday, it is the right choice. We often work hybrid: on-device for the basic work, premium cloud voice for highlight moments.
How does the app handle proper names and specialist jargon?
We build a pronunciation dictionary in which editors or administrators can tune proper names, place names, specialist jargon or medication names, using phonetic spelling or explicit IPA pronunciation. For more demanding applications, we fine-tune a custom neural voice in Azure Speech or a bespoke open-source model. Readers quickly lose trust in a read-aloud app if a recurring author's name is consistently mispronounced.
How do you handle the Copyright Act when reading articles or books aloud?
We check in advance which rights the publisher already holds, arrange additional agreements where necessary, and build into the architecture that an article can only be read aloud once its rights status is 'green'. For educational use, the educational exception sometimes works differently, and in that case we advise you on how to stay within it. For voice cloning we always require written consent and attestation from the original speaker.
How do you handle the AI Act for synthetic voice?
We begin with an AI Act classification. A general read-aloud app typically falls under the "limited transparency obligation": the user must know that the voice is synthetic. We implement that disclosure subtly in the app (an "AI voice" label next to the voice selection, a note during onboarding) without spoiling the reading experience. Voice cloning requires an additional layer of transparency and consent.
Do you work with Readspeaker or existing read-aloud providers?
We come across them regularly. Where it fits, we integrate with them rather than building our own read-aloud layer. But for publishers who see their voice identity as a distinguishing feature, or who want voice cloning or a bespoke brand voice, we build the TTS layer to measure with a suitable voice engine, with the copyright and AI Act layer properly arranged.
Can you integrate with our editorial CMS or LMS?
Yes. We have experience with WoodWing, Drupal, Sanity, Storyblok, Prismic and Contentful on the editorial side, and with Moodle, Brightspace, Magister and Itslearning on the LMS side. The integration typically runs via a REST or GraphQL API, with webhooks for "new article" events so the read-aloud audio is ready as soon as the article is published.
Who owns the code, voice settings and data?
You do. We deliver the full source code, the pronunciation dictionary, build pipelines and deployment scripts. Text, audio and listening behaviour remain in your cloud or with a hosting partner of your choice. No technical lock-in. For voice-cloned brand voices, the rights and the recording are contractually assigned to you. We earn our keep through good work that keeps you coming back.
What determines the cost and lead time of an app like this?
The main cost drivers are the complexity of the domain (a general read-aloud layer versus a deeply integrated accessibility app with an EAA project), the choice between cloud and on-device TTS, the number of voices and languages, and the compliance layer the domain requires. We work in sprints with fixed sprint budgets, so you can steer scope sprint by sprint.
Fabian van Dijk Business Developer · fabian.vandijk@appfront.nl

Talk to us about your read-aloud app.

A thirty-minute introductory call, no obligation. Tell us who will be listening: a reader with dyslexia, a subscriber in the car, a pupil on the move, a child falling asleep, along with the content and compliance context. We will think it through with you, give direction, and be honest about whether a dedicated read-aloud app is the right answer. Sometimes the right first step is an extension through our broader app development service. You can also email us directly at fabian.vandijk@appfront.nl.

Edit content