AI Document AI Workflow automation

AI document processing: from incoming pile to structured data

Invoices, contracts, forms, emails, receipts, reports: every Dutch business processes hundreds to thousands of documents a day. Many of them still get keyed in manually into spreadsheets, ERP fields or mailboxes. We build AI pipelines that automatically read, classify, validate and pass documents on to your existing systems, with human review where it matters.

Not an out-of-the-box SaaS that your documents have to fit into, but a solution that fits your document flows, your exceptions and your integrations. Whether it concerns Peppol invoices, scanned PDFs, photos from a field-service app or email attachments.

What AI document processing involves, beyond OCR

OCR is the oldest part of the chain: converting pixels into characters. But modern document processing is far more than that. A serious pipeline performs at least five steps in sequence, and each of those steps has its own model choice, failure mode and corresponding guardrail. We set them out below, not to impress, but to make clear where your existing tooling is likely to get stuck and where custom development pays off.

1. Document classification: before a single field is extracted, the system needs to know what type of document is coming in. A sales invoice, an incoming consignment note and a returns form require entirely different fields and validations. We train lightweight classifiers (often sufficient: a fine-tuned DistilBERT or a prompt-based LLM router) so that the correct downstream pipeline is triggered.

2. Layout analysis and OCR: for scans and photos we use, depending on the requirements, Tesseract (open source, free, fine for clean scans), AWS Textract or Google Document AI (stronger on tables and handwritten text), or newer vision-language models such as GPT-4o and Claude 3.5 Sonnet that can produce structured JSON directly from an image. The choice depends on volume, languages, costs and compliance, not on what is trendy.

3. Field extraction and entity recognition — pulling the right fields out of raw text: IBAN, KvK number, date, VAT number, invoice lines with VAT rates, signature fields. We do this with a combination of structured prompting (an LLM with a JSON schema) and classic regex/NER where that is cheaper and more reliable. Not everything needs to be an LLM: a good regex for postcodes is faster, more predictable and free.

4. Validation and business rules — checking extracted data against the real world. Does the VAT total match the line items? Does the supplier exist in the creditors ledger? Is the date within a reasonable window? Does the PO number match an open purchase order? Here, custom work almost always beats a SaaS product: your business rules are not your neighbour's.

5. Human-in-the-loop for exceptions — when confidence scores are insufficient or validations fail, the document is sent to a review queue where a person makes corrections. Those corrections then feed back into the model. A good pipeline takes over between 70% and 95% of the work without human intervention. Not 100%, and that is not a failure but a realistic expectation.

Concrete document flows we automate

The flows below are the ones we encounter most often at Dutch mid-sized and large organisations. We build pipelines for them that connect directly to the ERP, CRM and DMS systems you already use.

Inbound

Supplier invoices

PDF, UBL/Peppol, scanned paper invoice or photo from a field worker. We extract the invoice header, lines, VAT rates and match them to PO numbers. Output goes directly to Exact, AFAS, Twinfield, Unit4, Navision or your own ERP via REST/SOAP.

Inbound

Incoming contracts

NDAs, supplier agreements, lease contracts. We identify clauses (term, notice period, indexation, liability), record the key details in a contract register and set up alerts for expiry dates. Integration with DocuSign, PandaDoc or a custom DMS.

Inbound

Application and intake forms

Customer onboarding, claim reports, grant applications, job applications. Whether it is a digital form, a scanned paper form or a photo taken in a mobile app, we extract the fields, validate them against business rules and route them to the right department.

Inbound

Email classification and routing

We classify generic inboxes (info@, customerservice@) by subject, urgency and target audience. Attachments are taken into account in the assessment. Routing to Zendesk, Freshdesk, TOPdesk or a custom ticketing system, with auto-replies for the simpler cases.

Compliance

ID verification and KYC

Passport, driving licence, identity card: extraction of the MRZ, checks on holographic security features and comparison with a selfie. Useful in fintech, mobility and the staffing sector, with a fall-back to specialist partners (Onfido, iDIN, Veriff) for the more demanding compliance cases.

Analysis

Report and spreadsheet analysis

Annual accounts, market reports, due diligence files, quarterly reports. We build retrieval pipelines that make documents searchable by semantic question, with proper source attribution so that an end user can click through to the page where the answer came from.

Our typical technical stack, and when we choose what

No blind loyalty to a single vendor. We choose per project based on document type, volume, languages, latency requirements, cost and data residency. Below are the combinations we use most often.

OCR and layout

  • Tesseract / OCRmyPDF: when the scans are clean and the volume is high. Open source, no API costs, runs on-premise when needed.
  • AWS Textract: when tables and forms are central. Strong on structured layouts and question-answer pairs.
  • Google Document AI: when there is a lot of handwritten text, or for specific pre-trained processors (invoice, passport, W-9).
  • Azure AI Document Intelligence: suited to clients who are fully embedded in the Microsoft 365 / Azure stack and want their data to stay within the EU.

LLMs for extraction and classification

  • GPT-4o and GPT-4o-mini: excellent at structured JSON output, quick, and comparatively inexpensive in the mini version. They can read images directly.
  • Claude 3.5 Sonnet: strong with long documents and reasoning. Often our choice for contract analysis.
  • Mistral models via private endpoint: for when data must not leave the organisation (defence, healthcare, financial services).
  • DistilBERT / spaCy NER: for classifiers and entity extraction where latency and cost really matter. A fine-tuned BERT model is often ten times cheaper and faster than an LLM call.

Workflow and orchestration

  • Temporal / AWS Step Functions: for long-running, retry-safe document pipelines with human approval steps.
  • n8n / Make: when the business wants to be able to adjust things itself and the logic isn't too complex.
  • FastAPI + Celery + Redis: our go-to for custom pipelines that we take fully under management.
  • Apache Kafka: for event-driven architectures where document events also need to trigger other systems.

Storage and retrieval

  • S3 / Azure Blob for the raw documents, with versioning and lifecycle policies.
  • PostgreSQL with pgvector when the document volume is manageable: fewer moving parts.
  • Pinecone / Weaviate / Qdrant when you are genuinely dealing with enterprise-scale volumes.
  • Elasticsearch / OpenSearch for classic full-text search alongside semantic retrieval.

How we approach a document processing project

We work in four phases. These may seem straightforward, but the content of each phase depends heavily on your document types and volumes. We don't bring a pre-packaged methodology, but we do bring a body of experience from previous implementations.

Phase 1

Document audit

Two weeks. We gather 50 to 200 representative examples per document type, measure current processing times, map out exceptions and determine which fields genuinely need to be extracted. Outcome: a scope document with use cases and success criteria.

Phase 2

Proof of concept

Three to four weeks. We build a working pipeline on your busiest document stream with a simple review interface. The goal: to demonstrate that we can exceed an agreed accuracy threshold on real data, not just on a demo set.

Phase 3

Production and integration

Six to twelve weeks depending on scope. Integrations with ERP/CRM/DMS, monitoring (Datadog, Grafana or within your own stack), audit logging, user and role management, and a review interface suited to your end users.

Phase 4

Ongoing development

Ongoing. Adding new document types, revising models, monitoring accuracy via dashboards, adding new fields and supporting new languages. We keep models up to date: that is an ongoing responsibility, not a one-off task.

Compliance, privacy and data residency

AI document processing almost always touches personal data, commercially sensitive information or compliance requirements. That is built into our considerations from day one, not added as an afterthought.

GDPR and purpose limitation

We build pipelines that capture only the data necessary for the processing purpose. Personal data is pseudonymised or redacted where this suits the use case. Retention periods can be configured per document type. Data processing agreements are arranged as standard.

EU-only data residency

If your documents must not leave Europe, we use models and infrastructure that stay within the EU: Azure OpenAI in West Europe, Mistral in your own private endpoints, or self-hosted models on Dutch cloud infrastructure. No transfer through American SaaS where that is unacceptable.

Sector-specific regulations

For healthcare (NEN 7510, ISO 27001), financial services (DORA, PCI-DSS) and government (BIO, NIS2), we have previous implementations behind us. Audit logging, role segregation, encryption at rest and in transit, and penetration testing are standard components.

EU AI Act classification

We'll work with you to classify which risk category your use case falls under the EU AI Act. For medium- and high-risk applications, we build transparency, monitoring and human oversight requirements into the design.

What it delivers, and what it doesn't

We can't make claims like "60% faster" or "40% cost saving" out of thin air. Those figures depend entirely on your document volume, current turnaround time and the complexity of exceptions. What we can offer is an honest outline of what we've seen in comparable projects:

  • For high-volume structured documents (such as incoming invoices in a known format), we often reach 85–95% straight-through processing without any human intervention.
  • For unstructured documents (handwritten forms, photographs, contracts), this typically falls between 50–75%. That's still a substantial saving on manual processing, but not the 95% that SaaS vendors claim in their marketing.
  • The biggest gains are rarely pure FTE savings. More often they come from fewer errors (VAT mistakes, duplicate payments), faster turnaround (paying creditors sooner improves supplier relationships) and better data for reporting.
  • Implementations almost always pay for themselves within 12–18 months at volumes of around 500 documents per week or more. Below that volume, a SaaS solution is sometimes cheaper, and we're upfront about that.

How we differ from out-of-the-box vendors

There are good SaaS solutions for document AI, such as Klippa, Rossum, Hypatos, Docparser and Mendable. For businesses with standard document flows and low customisation requirements, they work very well, and we often recommend them rather than a custom build. We come into the picture when:

  • Your documents don't fall into standard categories, such as industry-specific forms, architectural drawings, legal deeds or shipping documents.
  • Your business rules are complex and unique, such as a bespoke pricing model for supplier invoices or a specific approval routing.
  • Your data absolutely may not leave the EU or your own infrastructure.
  • You need to integrate with legacy systems that have no modern API connections.
  • You want to own the model and the pipeline rather than be locked in to a vendor.
  • Your volume is so high that per-document SaaS pricing erodes the business case.

Frequently Asked Questions

How accurate is AI document processing really?

For clean, structured documents such as digital invoices, modern pipelines often exceed 98% field accuracy. For handwritten forms or poor-quality scans, that drops towards 80–90%. We test on your own data during the proof-of-concept phase, not on marketing samples, so the figures we put into the business case hold up.

Does this work on scanned PDFs and photographs, not just digital documents?

Yes. We combine OCR (Tesseract, Textract, Document AI) with layout detection and LLM-based extraction. Photos taken with a mobile app are automatically straightened and optimised before they enter the pipeline. The difference in accuracy between a scanned paper and a digital PDF is usually smaller than people expect.

How do you handle documents in multiple languages?

Modern LLMs and OCR models natively support Dutch, English, German, French and Spanish. For languages with other scripts (Polish, Russian, Arabic, Chinese) we often add a separate language detection step beforehand. We never restrict a pipeline to a single language if your actual document flow is multilingual.

Can you integrate with our existing ERP / DMS?

In almost all cases, yes. We have previously built integrations with Exact Online, AFAS, Twinfield, Unit4, SAP, Navision/Business Central, Salesforce, HubSpot, TOPdesk, M-Files, SharePoint and on-premises SharePoint. For systems without a public API, we work with file watching on SFTP exchange or database replication.

What if the model makes a mistake, who is then liable?

A document AI pipeline edits proposals, not decisions. The business logic determines which fields go into the accounts without human checking and which need a review. We design this together with you so that accountability and risk fit your governance, not just what is technically possible.

Can you also run on-premise without the cloud?

Yes. For clients whose data may not leave the premises (defence, some healthcare organisations, classic financials), we run self-hosted models on your own GPU infrastructure or a Dutch private cloud. This is more expensive than a SaaS solution, so we weigh up the trade-offs together with you.

Ready to automate your document processes?

A first conversation is no-obligation. We review your document flows, give you a realistic estimate of where AI pays off and where it doesn't, and we are honest when a SaaS solution is a better fit than custom development.

Edit content