Service · AI development

Custom AI app development with LLMs as the product core.

A mobile or web application in which AI is not a feature but forms the experience itself. Chat UIs, voice interaction, vision, RAG on your own data — built on Claude, GPT-4o, Gemini or on-device models.

LLM chatVoice & visionRAGOn-device AI

An AI app is not an ordinary app with a chatbot stuck on.

In an AI app, the LLM is the product. That changes almost every design decision: the UI must make streamed output feel natural, the architecture must keep token costs under control, and the system must cope with answers that are sometimes wrong. Hallucination mitigation, prompt engineering and evaluation are part of the foundations, not a last-sprint afterthought. At the same time, App Store reviewers expect you to explain how the AI works and which risks you address — a good-looking chat UI isn't enough to get through submission.

We build AI apps for startups with AI as their core proposition, for B2B tools that take a leap with an AI layer, and for consumer apps exploring new interaction models — think voice coaches, vision helpers, AI tutors, productivity assistants and symptom checkers. Target groups range from parents who want to help their child learn to lawyers reviewing contracts on a mobile phone, and from mental health coaches to people who want to use AI to cook, exercise or learn a language. Always custom, because no two AI apps have the same combination of model, data, latency requirements and compliance context.

We help you choose between Claude, GPT-4o, Gemini, Mistral or an open-source model, between cloud and on-device, and between a streaming chat UI and something that goes well beyond chat. For the broader range of AI work — agents, wide-scale integrations, generative content or strategic programmes — please visit our other AI pages; this page focuses specifically on the app as the end product your users actually get in their hands.

Our experience in mobile app development also means we take the practical side of AI in an app into account: token budgets per user, offline behaviour when there is no network, App Store classification, content moderation, age verification, and the UX of waiting for a model response without it feeling slow. These are exactly the details that make the difference between a demo and a product people use every day.

Three types of AI app we build.

The right choice depends on your target audience, data sensitivity and how central AI is to the experience. We use the first conversation to sharpen this together.

Consumer app · AI as the core experience

AI-first consumer app

The AI is the product. Think of a vision app that analyses photos for people with visual impairments, an AI coach for mental wellbeing that helps with stress and sleep, an AI tutor that helps pupils with maths or languages, an AI writing assistant for content creators, or a translation app that translates voice in real time on screen. Streaming chat UI, optional voice, multi-modal input, premium subscription with a token budget and App Store billing.

React Native or Swift/KotlinStreaming chatVoice AIIn-app purchases
B2B tool · AI augmentation within a professional field

Vertical AI app for professionals

An AI layer on a specific work process: a legal-tech app for contract review on the phone, an AI meeting summary for consultants straight after a conversation, AI document review for accountants during a client visit, a symptom checker for primary care, or a customer service app in which an AI chatbot forms the first line of support. RAG over your own knowledge base, strict authorisation, audit trail, sector-specific compliance, and clear boundaries on what the AI may and may not decide on its own. Output is always verifiable, with source references wherever possible.

RAG over domain dataSSO & role managementAudit logHuman-in-the-loop
AI assistant · sidekick in an existing app

AI sidekick in your current app

You already have an app or platform and want to add an AI layer: a conversational interface alongside the existing navigation, smart suggestions that understand context, in-context question answering over your own data, or agentic features that automate tasks. We build that layer on top of what is there, without throwing away your existing UX. Your users experience the AI feature as a natural extension rather than a foreign element bolted on. Often combined with a strong onboarding flow so users understand what the AI can and cannot do.

In-app assistantFunction callingOwn knowledge baseExisting backend

What you get at the end.

A working AI app in production, plus everything around it to run it yourself, keep developing it and keep costs under control.

For agentic AI with multi-step tool use, we often build a separate backend architecture. See our page on building AI agents for the variants where the AI carries out tasks autonomously.

  • The app itself, iOS and AndroidCross-platform via React Native or Flutter, or native where the use case requires it. Publication on the App Store and Play Store included.
  • AI backend with streamingToken-by-token output to the app, queueing for long prompts, fallback model when the primary provider is down.
  • RAG pipeline over your dataEmbeddings, vector database (Pinecone, Qdrant or pgvector), re-ranking and context injection. With an admin UI for adding documents.
  • Prompt library and evaluationVersion-controlled prompts plus an eval suite that detects regressions in model behaviour before we push to production.
  • Cost dashboardLive insight into token consumption per user, per feature and per model. With alerts when a limit is reached.
  • Codebase, documentation and runbookFull source code, architecture overview, deployment instructions and a runbook for incidents such as provider outages or hallucination reports.
  • Ongoing development contract (optional)Monitoring, model updates, prompt tuning and new features. A predictable monthly budget based on sprint capacity.

When a custom AI app is the right choice.

Four patterns in which we guide clients. If you recognise one of them, we are happy to talk through what the right answer is for your situation.

Product positioning

AI is central to your pitch

Your proposition is that AI solves something other apps cannot. In that case the app should feel like an AI product from day one, not a traditional app with a chat button tucked in a corner. Streaming UI, voice where it makes sense, and an interaction model that suits what the model does well.

Domain expertise

You have proprietary data and expertise

Generic LLM answers are not good enough for your audience. A RAG pipeline built on your own content, combined with domain-specific prompts and an evaluation suite that tests against your field, delivers the relevance and accuracy your users expect, without you having to train a model yourself.

Privacy

Sensitive data, sensitive sector

Medical, legal or HR data cannot simply be sent to a US-based LLM provider. EU data residency through Anthropic EU or OpenAI on Azure EU, on-device models for the most sensitive steps, or a private deployment of an open-source model are realistic options that we weigh together.

End-user product

You are taking an app to market

The App Store and Play Store have specific guidelines for AI apps: transparency about model behaviour, age ratings, content moderation and in-app billing for AI credits. These need to be built in from the design stage, not addressed at review submission, where they lead to rejection.

Tech stack and key decisions.

We are not dogmatic about any one framework or LLM provider. The right choice depends on your use case, user numbers, data sensitivity and team. Below are the options we most often arrive at, with the reasoning behind each.

For enterprise projects where AI needs to land across more than one app, we align these choices with the roadmap on our page about enterprise AI implementation.

  • Mobile: React Native, Flutter or nativeReact Native for rapid iteration with a single team, Flutter when pixel-perfect UI across both platforms is critical, and native Swift or Kotlin when performance or platform-specific AI (Apple Intelligence, Core ML, ML Kit) is decisive.
  • LLM providers: Anthropic, OpenAI, Google, MistralClaude for long context and nuance, GPT-4o for multimodal work and mature tool use, Gemini for long context and Google integrations, Mistral and open-source models for self-hosting or EU data residency.
  • LLM frameworks: LangChain, LlamaIndex, Pydantic AINot always necessary. Sometimes a direct API call with good prompt engineering beats adding a framework layer. We choose pragmatically, based on the complexity of the flows and tool use.
  • Vector database: Pinecone, Weaviate, Qdrant or pgvectorPinecone for a quick managed setup, Qdrant or Weaviate when you want to self-host, pgvector when you already use PostgreSQL and do not want to maintain a separate service.
  • On-device AI: Apple Intelligence, Gemini NanoFor privacy-sensitive features, a smaller model runs locally. We combine this with a cloud fallback for heavier tasks, so users will barely notice the difference apart from speed and battery use.
  • Voice AI: Whisper, ElevenLabs, DeepgramWhisper for robust speech-to-text, ElevenLabs for natural text-to-speech, Deepgram for low-latency streaming transcription. We mix these depending on language support and cost per minute.
  • Backend: Node.js, Python or GoPython where there is a lot of AI orchestration and data processing, Node.js for fast streaming endpoints and TypeScript sharing with the frontend, Go where throughput performance is decisive.
  • Key decisionsCloud or on-device, EU data residency yes or no, how you limit hallucinations, how token costs stay under control as you grow, and what the AI Act classification of your app looks like. We factor these in early in the process.
Not yet sure about a large project?

Test your idea first: a working prototype in 1 day

With OneDayBuild, we turn your idea into something tangible in one day for €1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.

Explore OneDayBuild →

Compliance for an AI app — what we take care of.

AI apps are subject to more legislation and guidelines than regular apps. We map out early in the process what applies, rather than layering it on at the end.

GDPR

Privacy by design

Data minimisation, data processing agreements with LLM providers, tightly configured log retention and a privacy statement that matches what actually happens. A DPIA for sensitive use cases.

AI Act

Classification and transparency

Determining whether your app falls under minimal, limited or high risk. For limited risk, minimal transparency about AI output; for high risk, more extensive documentation, evaluation and human oversight. Also AI literacy for your team (Art. 4).

App Store & Play Store

Platform guidelines

Apple Intelligence privacy guidelines, Apple's AI disclosure requirements and Google's Generative AI policies. Content moderation, age rating and clear explanation of data use in the App Store form.

WCAG / EAA

Accessibility

From June 2026, the European Accessibility Act applies to many digital products. AI chat UIs must work well with screen readers, voice input is an accessibility opportunity, and streaming output must not cause a focus trap.

How an AI app project works with us.

1

Introductory meeting and sharpening the use case

A conversation in which we understand what the app should do, who will use it, what data is available and which models are suitable. Often also an initial discussion about cloud versus on-device and the compliance context.

2

Prototype with real prompts

We build a tangible prototype flow with real LLM calls — not a mockup, but a working slice that shows how the AI responds to your data and your users. Based on this, we fine-tune scope and model choice.

3

Building in sprints

A working build every two weeks, tested on iOS and Android. In parallel we work on the prompt library, the RAG pipeline and the evaluation suite. You test along, as does a small group of beta users.

4

App Store submission and launch

App Store Connect and Play Console submission, with all the correct AI disclosures. Soft launch in one market to gather feedback for the wider rollout.

5

Continuous development phase

AI models improve rapidly, and so do prompts. We monitor costs, quality and usage, and run sprints to integrate new models, address hallucination patterns and extend features.

Frequently asked questions.

What clients usually want to know before we start.

What is the difference between an AI app and a regular app with AI features?
In an AI app, the LLM is the core experience, and design, architecture and pricing revolve around token usage, streaming, hallucinations and model evolution. A regular app with an AI feature treats AI as an addition — useful, but not decisive for the architecture. The difference lies in where the design effort goes, not in a particular feature checklist.
Which LLM should we choose — Claude, GPT-4o, Gemini or Mistral?
It depends on the use case. Claude is strong in long context, nuance and code; GPT-4o in multimodal and tool use; Gemini in long context and Google integrations; Mistral and open-source models when you want to self-host. We benchmark during the prototype phase on your own prompts and data — this prevents the choice from being based on generic benchmarks that say nothing about your situation. For deeper integration decisions, we also refer you to our page on custom LLM integrations.
Cloud AI or on-device — what is the difference?
Cloud AI (Claude, GPT-4o, Gemini) gives you access to the most powerful models, but sends data to an external provider and charges per token. On-device AI (Apple Intelligence on iOS 18+, Gemini Nano on Android 14+) runs locally, costs nothing per use and keeps data on the device, but the models are smaller and less capable. For sensitive data, on-device is attractive; for demanding reasoning tasks, you would typically choose the cloud. Many apps do both: simple tasks on-device, more complex ones via an EU cloud endpoint.
What does the AI Act mean for our app?
The AI Act distinguishes between prohibited, high-risk, limited-risk and minimal-risk systems. Most consumer AI apps fall under limited risk and must be transparent that output has been generated by AI. High risk applies, among other things, to use in recruitment, credit scoring, education assessment or medical contexts, where heavier documentation and evaluation requirements apply. We help you classify your app and build the right compliance layer in from the start, rather than bolting it on afterwards. For strategic projects, see our page on enterprise AI implementation.
How do you deal with hallucinations?
Hallucinations can't be eliminated entirely, but they can be greatly reduced. We combine RAG (so the model draws on your own sources), grounding instructions in the prompt, output validation via a second model or rules, and clear UI signals when the AI is uncertain. For critical decisions, we build in human-in-the-loop so that a user confirms before any action is taken.
What determines the cost of an AI app?
Mainly three things: development hours (depending on scope, platforms and the number of AI features), recurring model costs (tokens per user per month), and infrastructure costs for the vector database, hosting and monitoring. We work with sprint budgets for the build and set up a cost dashboard from day one, so you can track usage per feature, model and user. That lets you decide when a cheaper model or a caching layer is worth it.
What about privacy and sensitive data?
By default we work with EU data residency where possible (Anthropic EU endpoints, OpenAI EU deployment via Azure). For truly sensitive data we opt for an EU private deployment or an on-device model. We sign a data processing agreement with you and with the LLM provider, carry out a DPIA where necessary, and keep log retention tight. Your data is never used by us for model training.
Will my data be used to train the model?
With Anthropic, OpenAI Enterprise and Google Vertex AI, it is standard that your data is excluded from model training. On OpenAI's consumer APIs this is different, and we do not use them for production. We explicitly document which provider gives which guarantee and include this in your privacy statement, so your end users also know where they stand.
Can you also manage and further develop the app?
Yes, and for AI apps that is almost always sensible. New model versions appear every few months, prompt engineering remains a living craft, and costs can shift. An ongoing development contract with fixed sprint capacity ensures the app keeps pace with developments in the AI landscape. For a broader range of generative applications, see also our page on generative AI solutions.

Talk to us about your AI app.

A no-obligation introductory call of half an hour. We listen to your use case, think along on model choice and privacy, and give direction you can act on, even if the right next step is not to build straight away but first a prototype or a short validation sprint. Whether you are pre-product, want to enrich an existing mobile app, or are building on an AI strategy already in place elsewhere, we are happy to start the conversation and would rather get involved early than only step in once everything is settled.

Edit content