Service · Web development

Custom AI agent development.

An AI agent does more than answer questions. It carries out tasks on behalf of your staff: retrieving data, proposing decisions, calling tools, choosing next steps. We build domain-specific agents that work on your data, your systems and your rules, and are model-agnostic.

An AI agent is not a chatbot.

A chatbot answers a question and stops. An AI agent is given a goal, such as "research this lead", "classify this ticket" or "check this invoice against our purchasing terms", and works out for itself which steps are needed. It consults your documents, calls your APIs, compares options, validates intermediate results and returns a substantiated outcome, often with a trail of sources and decisions.

That difference makes building one harder. The agent needs to know which tools are available, when it may act, when it must escalate, and how to account for its steps. A chatbot talks; an agent does something that matters. That means strict requirements for reliability, observability and cost control. We build AI agents for businesses that do all of this in a production-ready environment, with monitoring per agent run, cost budgets per task and human-in-the-loop for actions with impact.

Under the bonnet, agents rely on concepts such as Anthropic's tool use, OpenAI Assistants, Microsoft Copilot Studio and open-source frameworks like LangChain and LlamaIndex. The best combination depends on your use case, data residency requirements and the models you have access to. We work model-agnostically: Anthropic Claude, OpenAI GPT, Azure OpenAI, Google Gemini, or open-source models such as Llama and Mistral within your own VPC. No vendor lock-in, no dogma about a single framework.

Not every task suits an agent. For tightly defined, deterministic processes, a standard integration or a rule-based workflow is often cheaper and more reliable. An agent adds value where the work requires judgement, where data must be brought together from several sources, or where the input is so irregular that a traditional if-then flow cannot keep up. In every initial conversation, we help you weigh this up honestly.

What makes an agent different.

Tool use
The agent calls your APIs and databases
RAG
Your documents as context, not model training
Memory
Long-term customer and task context
Guardrails
A rules engine checks every output

Three types of AI agent.

Depending on how much autonomy the agent is given and how deeply it integrates with your existing systems.

Flavour 01

Assistant agent (read-only)

Compact project

The agent reads along, searches, summarises and suggests. No write actions on your production systems. Ideal as a starting point: low risk, quick value, easy to monitor. Think of a sales research agent that combines LinkedIn, your CRM and public sources into a briefing for a meeting, or a research agent that gathers market signals weekly and summarises them for your management team.

RAG pipelineTool useStreaming UIAudit log
Flavour 02

Action agent (with human-in-the-loop)

Medium-sized project

The agent proposes actions and carries them out once a human approves. For example, a support agent that classifies a ticket, drafts a reply with source references and waits for one-click approval before it goes to the customer. Or a purchasing agent that drafts a request for quote and only sends it after review. The review step is built in deliberately: trust is something you build up gradually, not something you switch on.

Multi-step reasoningApproval flowCRM writeNotifications
Option 03

Autonomous agent (production-critical)

Larger project

The agent acts autonomously within tight boundaries. For example, an operations agent that classifies monitoring alerts, runs standard runbooks and only escalates when there is a deviation or uncertainty. This requires strong guardrails, observability for each step, a fallback path for when the agent is unsure, and clear limits on which tools the agent may call in which context. Only deploy it for tasks where the real risk of an incorrect action is low or easily reversed.

Tool orchestrationRules engineObservabilityCost controls

What you get at the end.

A production-ready agent, running in your environment or hosted by us, with everything around it to manage, adjust and extend it.

The agent itself

Production and staging, in your cloud (EU region) or with us. Your own prompts, your own tools.

Codebase and prompts

Full source code, a versioned prompt library, and the tool definitions the agent uses.

Observability dashboard

Insight into which prompts, which tools, which costs and which errors occur per run.

Evaluation suite

A test set to catch regressions when you change prompts, models or tools.

Managed service (optional)

Model updates, prompt tuning, cost monitoring, adding new tools.

What we build agents for.

Eight use cases we guide clients through. The common thread: repetitive work that requires judgement, on data that lives across multiple systems.

Sales

Lead research agent

Combines CRM, LinkedIn, public sources and your own sales notes into a meeting briefing. Can draft a first email and update CRM fields.

Customer support

Support agent

Classifies incoming tickets, finds similar cases and proposes a reply with sources cited. Escalates automatically when signs of urgency or complaint appear.

Operations

Operations agent

Detects anomalies in monitoring data, links them to runbooks and suggests a follow-up action. Generates weekly reports based on your own KPIs.

Compliance

Compliance agent

Reviews documents against internal rules or external standards. Performs sanctions list checks and due diligence steps, and flags deviations for human review.

HR

HR agent

Answers questions about employment terms from your own documents. Screens CVs against role criteria without using bias-sensitive fields.

Procurement

Procurement agent

Searches supplier catalogues, compares prices and prepares quote requests. Takes your procurement policy and approved suppliers into account.

Finance

Finance agent

Codes incoming invoices, checks expense claims against expense policy, and flags possible fraud patterns. Posts to your accounting system once approved.

IT

IT agent

Triage of monitoring alerts, automatic execution of standard runbooks, and escalation routing to the right team with context handed over.

How an engagement works.

01Use case scoping→ 02Evaluation and prototype→ 03Build→ 04Rollout and management
1 hour

Use case scoping

We map out the task: the goal, the data, the tools and the risk boundaries.

A few sprints

Evaluation and prototype

We build a working prototype on a single model, along with an evaluation set to measure quality objectively.

A few sprints

Production build

Guardrails, observability, cost controls, RAG pipeline, and integrations with your systems.

Ongoing

Rollout and maintenance

A pilot with a small group, followed by a wider rollout. We keep tuning the models and prompts.

Not yet sure about a large project?

Test your idea first: a working prototype in 1 day

With OneDayBuild, we turn your idea into something tangible in one day for €1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.

Explore OneDayBuild →

Frameworks we work within as standard.

An agent that carries out tasks on behalf of your organisation touches on privacy, security and compliance. We identify those frameworks at the outset and build them into the design, rather than tick them off afterwards.

GDPR

Data minimisation and DPIA

The agent only has access to the data needed for the task. We carry out a DPIA for every project that processes personal data, and we set retention limits in code, not just in policy.

EU AI Act

Risk classification upfront

At the start of the project, we determine which category your agent falls into (minimal, limited or high risk) and align documentation, transparency towards users and monitoring accordingly. For high-risk use cases (HR, government, critical infrastructure), we build in additional audit trails and human escalation points.

Data residency

EU-only where required

For regulated data, everything runs within EU regions: Azure OpenAI in West Europe or Sweden Central, Anthropic via AWS Bedrock EU, or a self-hosted Llama deployment on your own premises. No model data, logs or embeddings need to leave the EU for the US.

Guardrails

Multi-layered output validation

The agent cannot do everything. A rules engine checks output for forbidden content, tool usage and cost budget. High-impact actions require a second model validation or human approval. We log all failures in full, including the prompt, response and tool trace, for later review.

Models and stack we work with.

For each project, we choose what best suits quality, latency, cost and data residency. We are model-agnostic, with no allegiance to a single vendor or framework. For specific smart API integrations between the agent and your systems, we work to the same principles. Where an agent is too heavy for the task, we fall back to a simpler solution, and where your use case goes well beyond a single agent, we guide you through a broader enterprise AI implementation.

Models
Anthropic ClaudeOpenAI GPT-4o / o1Azure OpenAIGoogle GeminiLlama (on-premises)Mistral (on-premises)
Agent frameworks
LangChainLlamaIndexOpenAI AssistantsAnthropic SDKVercel AI SDKCustom orchestration
Infrastructure & observability
pgvectorWeaviateQdrantLangSmithLangfuseAzure / GCP / AWS

Frequently asked questions.

What is the difference between an AI agent and a chatbot?
A chatbot answers questions based on a model and, where relevant, your documents. An AI agent is given a goal and determines the steps itself to achieve it: calling tools, making decisions, choosing next steps and checking output. It carries out tasks, not just conversations. For a focused question-and-answer flow, an AI chatbot is often sufficient; for tasks spanning multiple steps and systems, you would choose an agent.
Which AI model do you choose for my agent?
That depends on your task. For complex reasoning and large contexts, we often choose Anthropic Claude Sonnet or Opus. For fast and inexpensive work, Haiku, GPT-4o-mini or an open-source model. For strict data residency, everything runs within Azure OpenAI or a self-hosted Llama deployment on your own premises. We benchmark during the evaluation phase on your own tasks before finalising the choice.
Is our data kept safe, and is it used for training?
With commercial APIs from Anthropic, OpenAI and Google, data is by default not used for training; we record this per project in the contracts and API settings. For regulated data, we deploy Azure OpenAI or an on-premises model so that the data never leaves your VPC. We establish retention limits, audit logging and data residency in the design.
Can the agent run in an EU region?
Yes. Anthropic Claude is available in EU regions via AWS Bedrock, OpenAI via Azure West Europe / Sweden Central, and Google Gemini via Vertex AI in the EU. For open-source models, we run Llama or Mistral within your own VPC or with an EU cloud provider. No data needs to leave the EU for the US.
What does the EU AI Act mean for our agent?
The EU AI Act classifies AI systems by risk. Many business agents (sales, support, operations) fall into the "limited risk" category, with transparency obligations. Agents that support HR decisions or run in government may be "high risk", with stricter documentation and monitoring requirements. We map the classification at the start of the project so the design is right from the outset. For AI in municipalities and government, we already have specific experience.
Can the agent run on-premises without external APIs?
Yes. We deploy Llama or Mistral as models, with vLLM or Ollama as the runtime, on your own GPUs or in a private cloud tenant. Quality is slightly lower than with top-tier commercial models, but for well-defined tasks it is often perfectly adequate. For more demanding agents, we sometimes combine a local model for routine work with a hosted API for difficult steps.
What determines the cost of an AI agent?
Three factors. One: the build scope. A read-only assistant is a project of a few sprints; an autonomous agent with deep integrations is a project of several sprints. Two: the runtime cost of the chosen model multiplied by the expected volume. Three: ongoing management after go-live. We make these three items transparent and, during the build, steer against a cost budget per agent run. For broader context on AI implementation, see enterprise AI implementation.
How quickly can we go live?
A focused assistant agent can be delivered in a few sprints. An action agent with human-in-the-loop oversight and integrations with your systems is a project of several sprints. Autonomous production agents with strict guardrails and compliance requirements need more time in the evaluation and hardening phase. We work in two-week sprints so you see a working version every two weeks.
What happens after the MVP?
The agent first goes to a small group of users. We measure quality (eval set and human review), latency and cost. Based on that, we tune prompts, tools and, where needed, the model. Then we widen the rollout. Models change quickly, so we schedule periodic evaluations to check whether a newer or cheaper model can replace the current setup without loss of quality.
How do you measure whether the agent works well?
We build an eval suite of representative tasks with expected outputs. On that set we run automated scores (correctness, completeness, source attribution) plus sampled human review. In addition, we log every production run with prompt, tool calls, model, cost and latency, so regressions from a prompt change or model update surface immediately. What isn't measured can't be reliably improved, and without evals an agent is always a gamble.

Talk to us about your AI agent.

A free, no-obligation half-hour introductory conversation. We listen to the task you want to automate, ask sharp questions about data, risk and volume, and give you an honest view on whether an agent is the right answer, perhaps a chatbot, or no AI at all.

Want to start a project?

We'd love to hear from you.

Long story, or would you otherwise have emailed it? Choose Detailed briefing: headings and bullet lists, images in the text and attached files.

Mies

Get in touch with Mies

Business Developer

Get in touch with Martijn

Founder of Appfront

Martijn

Your message has been sent

Thank you for your interest! We'll get back to you as soon as possible, usually within 1 working day.

Edit content