What exactly is an AI agent?
An AI agent is software that uses a large language model (LLM) to carry out tasks independently. Not one question and one answer like a chatbot, but end to end: the agent receives a goal, makes a plan, calls tools, assesses the intermediate results and adjusts until the task is complete.
The distinction matters more than it might seem. A chatbot answers questions. A Robotic Process Automation (RPA) bot follows predefined click paths. An AI agent can work out for itself, in an unexpected situation, what the logical next step is, and then carry it out. Anthropic calls this agentic workflows: systems in which the language model not only generates text but also decides, plans and acts within predefined boundaries.
An AI agent is an LLM with hands: it can use tools, make decisions and plan multiple steps. It is not a question-and-answer machine but a goal-driven executor operating within clear boundaries.
How do AI agents work under the hood?
The architecture is almost always the same, regardless of the model used (Claude, GPT, Gemini or an open-source variant). Four components do the work:
- The brain (LLM): Claude, GPT-4, Gemini or a locally run Llama model. This is where reasoning, language understanding and planning take place.
- Tools: functions the agent can call: running a database query, calling an API, searching a document, sending an email, creating a ticket. Tools are the agent's "hands".
- Memory: short-term (context of the current conversation) and long-term (past interactions, user preferences, knowledge base data via embeddings).
- Planning loop: the agent chooses a tool, runs it, reviews the result, and decides whether the task is complete or a next step is needed. This loop runs until the goal is reached or a stop condition is triggered.
We usually build agents on the Anthropic stack or with a framework such as LangGraph, using your own organisation's tools and data sources. If you prefer a custom solution to a drop-in one, we explain what that looks like at /diensten/web-ontwikkeling/ai-agent-laten-bouwen.
Four types of AI agents in a business context
Not every agent is equally ambitious. In practice we distinguish four variants, increasing in autonomy:
1. Single-task agent
One specific task, one clear output. For example, a CV-screening agent that compares an incoming CV against a job profile and returns a match score. Low complexity, quickly measurable impact.
2. Multi-step workflow agent
Several steps in sequence, often with conditional logic. A sales-research agent that looks up a new lead on LinkedIn and the web, summarises the findings, updates the CRM and delivers a first draft email. This is where the power of agentic workflows really shows.
3. Conversational agent
A customer or employee chats with the agent, but unlike a traditional chatbot, this agent can also carry out actions: look up an order, reschedule an appointment, resend an invoice. The line between "providing information" and "doing something" blurs. For customer service cases we more often build a dedicated AI chatbot for your business with action capability.
4. Background autonomous agent
Runs continuously in the background, without anyone asking for anything. For example, it monitors incoming emails, sales orders or system logs and intervenes on signals. Produces a daily or weekly report. Highest autonomy, and therefore the strictest governance.
Practical examples by discipline
A few use cases we have seen with clients recently, organised by functional area:
| Discipline | Example agent |
|---|---|
| Sales | Lead-research agent: gathers public information from LinkedIn and company websites, enriches the CRM record and writes a personalised introduction message. |
| Marketing | Content-optimisation agent: analyses A/B test results, compares them with previous campaigns and suggests follow-up variants. |
| Support | Ticket triage: classifies incoming tickets, links the right knowledge base articles and drafts a first reply for the support agent. |
| HR | CV screening: compares CVs with the profile, assesses candidate fit and then arranges interview scheduling via the calendar API. |
| Finance | Expense policy check: reads incoming expense claims, checks them against the policy and automatically codes invoices to the correct cost centre. |
| Operations | Anomaly detection: monitors production or logistics data, recognises deviations and routes alerts to the right person responsible. |
| Legal | Contract review: compares incoming contracts with your own templates, flags deviating clauses and provides an initial risk assessment. |
What stands out: in every example the agent runs on data and systems you already use (CRM, ATS, ticketing, ERP). The value does not come from the language model itself, but from the combination of the LLM, your business data and the right tools.
What an AI agent is NOT
An honest account requires managing expectations. Four things we explicitly do not promise:
- Not fully autonomous. An agent running entirely without guardrails is a risk. For anything with customer or financial impact, we build in a human-in-the-loop step.
- Not flawless. Hallucinations, meaning the invention of plausible but incorrect information, remain a real risk. Output validation and guardrails are essential.
- Not free. Every call to the language model consumes tokens. An agent that reasons continuously can run up considerable monthly costs if you don't manage token budgets.
- Not plug-and-play. An agent that delivers real value requires clean data, integrations with your existing stack, well-designed prompts and monitoring. It is a development project, not an off-the-shelf product.
Ask every vendor how they handle hallucinations, tool permissions and cost control. Anyone without a clear answer to these questions has not yet built an agent that is running in production.
Security and governance: the seven core areas
This part is perhaps the most important, and the least glamorous. An agent without governance is a liability risk. Seven points that must be addressed in every implementation:
- Data privacy: the GDPR and the EU AI Act are not optional. Which data goes to which model, where is that model hosted, and do you have a data processing agreement in place? For some use cases, a European-hosted or even on-premises solution is the only route.
- Tool permissions (least privilege): the agent only has access to what it needs for its task. No agent should be able to delete the entire database in production simply because that is technically possible.
- Human-in-the-loop for high-impact actions: for actions with financial, legal or customer impact, human approval is always required. The agent proposes, a human approves.
- Audit trail: every prompt, every tool call and every decision must be traceable. Without logs, troubleshooting is impossible and compliance is unthinkable.
- Cost monitoring: a token budget per agent, alerts when it is exceeded, and automatic rate limiting. A runaway loop can rack up considerable costs.
- Output validation (guardrails): rules that validate the agent's output before it reaches a customer or system. For example, a check that a generated email does not contain customer data from another account.
- Model selection: not every language model should be allowed to see your data. EU-hosted variants, on-premises models or zero-retention contracts are often required in regulated sectors. See also our article on enterprise AI implementation for the wider context.
This layer is not optional but an integral part of any serious project. A notable Search Console observation: the query "ai-agent software veilige bedrijfsimplementaties" ranks in the top three on Google, a sign that decision-makers are actively looking for this security aspect, and that many vendors are still glossing over it too easily.
When is an AI agent worthwhile, and when is it not?
The difference between a successful implementation and an expensive pilot usually lies not in the technology but in use-case selection. A few rules of thumb:
A good candidate
- Recurring tasks with a recognisable pattern (think classification, summarisation, first-line responses).
- High volumes where manual work no longer scales.
- Combining multi-source data (customer data + product information + stock + history).
- Tasks where 80% good enough delivers a large time saving, and a human handles the remaining 20%.
Not a good candidate
- Rare edge cases where contextual judgement and experience outweigh pattern recognition.
- Legally or medically binding decisions without a human final judgement.
- Sensitive customer interactions where a single miscommunication could damage the relationship.
- Processes whose data is so messy that even a human can't make sense of it. An agent does not improve poor input.
The question "is this a good agent case?" is in many projects more valuable than "which model should we choose?"
How to start: a five-step roadmap
We work in staged phases so you don't spend months building something that has no impact.
- 1. Use-case selection. A workshop with the teams who carry out the task today. Which steps are repetitive? Where is the most friction? What can be measured?
- 2. MVP agent. One clearly defined task, one clear goal, a few sprints to get something working. Not the full stack, but a genuine proof of value.
- 3. Measuring impact. Compare against the baseline: how much time, how many errors, how much turnaround time before and after. Only then make decisions about scaling.
- 4. Scaling. Only once the MVP has proven its value: additional use cases, better models, more integrations. Not the other way round.
- 5. Institutionalising governance. An internal policy for agents, a review process for new use cases, and clear ownership across IT and compliance. This is what most agencies skip, and what you will need most urgently a year from now.
Cost considerations
Without specific figures, as these depend heavily on model choice, volume and complexity, here are four cost categories to include in your business case:
- Tokens. The LLM charges per thousand input and output tokens. With agents, costs can add up because a single task may require several model calls.
- Development time. Building the tools, prompts, integrations and observability: a serious development project spanning several sprints.
- Monitoring and maintenance. An agent in production requires active monitoring. Models receive updates, APIs change, and prompts need to be refined.
- Retraining or fine-tuning. Not always necessary, but if you want to perform well in your own domain, fine-tuning or a custom embedding layer may come into play.
If you are weighing up whether a custom agent is smarter than an existing SaaS solution, our article on build versus buy covers the same trade-off from a broader software perspective.
Frequently Asked Questions
What is the difference between an AI agent, a chatbot and an RPA bot?
A chatbot answers questions based on a fixed script or a language model, but does not carry out tasks. An RPA bot follows a predefined sequence of clicks and cannot improvise when something deviates. An AI agent combines language understanding with tools and can independently decide the next step, even in situations that were not explicitly anticipated.
Which language model should we choose: Claude, GPT, Gemini or open source?
That depends on your requirements for price, quality, latency, data residency and compliance. Claude (Anthropic) and GPT (OpenAI) currently deliver the strongest agent performance. Gemini (Google) is competitive and sits close to the Google stack. Open-source models such as Llama or Mistral are attractive where data residency or cost are decisive. We always advise building in a model-agnostic way, so you can switch if the market shifts.
How do we handle GDPR and the EU AI Act?
Start with classification: what risk level does your use case fall under the EU AI Act? High-risk applications (recruitment, credit assessment, critical infrastructure) require stricter documentation and human oversight. For GDPR, choose a model with a data processing agreement, zero retention where possible and, preferably, European hosting. For clients in regulated sectors we work with EU data residency as standard.
Can an AI agent run entirely on-premises?
Yes, this is possible with open-source models such as Llama or Mistral, hosted on your own infrastructure or a European private cloud. On many tasks the quality is now close to commercial API models. The trade-off is that you manage the infrastructure, GPU costs and updates yourself. For strict compliance requirements, those costs often outweigh the risk of sending data to a US provider.
How do we measure the ROI of an AI agent?
Example metrics we include as standard: time saved per task, reduction in errors, shorter turnaround time, qualitative assessments from the people working with the agent, and cost per completed task (including token costs). A baseline measurement is crucial before the agent goes live; otherwise everything looks plausible in hindsight and nothing is proven.
What do we do when the agent makes mistakes?
Expect it to happen and build for it. Specifically: output validation guardrails, a human veto for high-impact actions, detailed logging so you can reproduce the error, and a feedback loop that turns errors into prompt or tool adjustments. An agent without an error procedure is not ready for production.
How do we scale after a successful MVP?
Once a first agent has proven its value, we proceed in this order: extend the same agent with more tools, roll out similar patterns to other departments, and only then tackle entirely new use cases. At the same time, governance is scaled up: an agent policy, a review board and clear ownership within IT. In our AI training for businesses, we describe how to bring your team along in this phase.
Which factors determine the cost of an AI agent project?
The complexity of the task, the number of integrations with existing systems, the volume of calls, the chosen model, and the governance requirements. A single-task agent on an existing API costs considerably less than an autonomous multi-step agent with an audit trail and human-in-the-loop. After an initial conversation, we always give an indication of the order of magnitude before we start anything.