AI document extraction and OCR: from paper and PDF to structured data

Back offices, shared service centres, accountancy firms, insurers and legal teams process thousands of invoices, policies, contracts, Chamber of Commerce extracts, valuations and email attachments every day. Traditional OCR recognises characters but leaves the interpretation to people. With modern AI document extraction, including vision LLMs, layout-aware models and structured-output schemas, software reads documents in any format and delivers validated JSON straight to your accounting, policy or case management system.

Invoice extraction (PEPPOL UBL) Chamber of Commerce extracts Policy and claims processing Contract clause extraction Passport/ID and KYC CVs and emails
Discuss your document workflow View applications
AI { "invoice_no": "F-2419" "vendor": "Acme BV" "vat_id": "NL8001…" "date": "2026-04-30" "total": 1842.50 "currency": "EUR" "line_items": […] "confidence": 0.97 }

What AI document extraction and modern OCR do

Document processing has fundamentally changed in the past five years. Where Tesseract and ABBYY once read characters and you had to write regular expressions and templates yourself, vision LLMs and layout-aware models now extract meaning from a document, including tables, handwritten notes, stamps and scanned copies of varying quality.

An invoice may arrive as a structured PEPPOL UBL file, as a PDF generated by an ERP, as a scanned paper invoice or as a photo taken on a phone. For each format, your system must correctly identify the supplier, the VAT number, the invoice lines, the final totals and the payment reference, and ideally also detect whether the posting lines make sense. Classic OCR produces text, not meaning. AI document extraction delivers a validated JSON object straight away that matches the schema your accounting package or case management system expects.

The same applies to Chamber of Commerce extracts, cadastral documents and deeds, policy schedules, claim files, purchase agreements, mortgage offers, passports and ID documents, CVs, receipts and even the content of incoming emails. Each document type has its own schema; Appfront designs that schema together with you, trains or fine-tunes the model, and builds the pipeline that receives, classifies, extracts, validates and forwards documents to your existing systems, with human-in-the-loop review for cases where the confidence score falls below your threshold.

The difference from traditional OCR lies in three things: layout awareness (models understand columns, tables and visual grouping), contextual understanding (the word "total" is weighted differently when it sits bottom right than in the header), and direct structure (output is a typed JSON schema, not a block of flat text you still have to parse yourself). That difference determines whether a document is processed in 30 seconds and then flows through automatically in 90% of cases, or whether every invoice has to be handled again by human hands.

Document types we process

We build extraction pipelines for almost any document that appears in a Dutch back office. Below are the six most common categories. For each one, we have a reference schema that we refine together with you.

📄

Invoices and receipts (AP automation)

PEPPOL UBL via Mijn Overheid or supplier portals, free-form PDF invoices, scanned paper invoices, till receipts and restaurant bills. We extract supplier, KvK and VAT numbers, IBAN, invoice number, invoice date, due date, line items with VAT rate, and final totals. Output aligns with Exact Online, AFAS, Twinfield, Yuki, Visma.net or a custom ERP. See also our invoice recognition page.

🏛️

KvK extracts and land registry documents

Extracts from the Trade Register, cadastral records, ownership information, mortgage deeds and transfer deeds. Alongside direct extraction via the KvK API and the Kadaster integration, we also process scanned historical files where no API source is available.

📋

Policies and claims files

Policy schedules, policy terms, claim forms (loss reports, expert reports), medical reports and correspondence from insurance intermediaries. The model recognises coverages, excess, premiums, policy numbers and claim items. Suitable for health insurers, property and casualty insurers, life insurers, reinsurers and managing general agents.

⚖️

Contracts and legal files

Purchase agreements, employment contracts, NDAs, SLAs, lease agreements and general terms and conditions. Clause extraction identifies notice periods, penalty clauses, governing law clauses, price indexation and termination conditions. Output goes to your contract management system or CLM tooling.

🆔

Passports, ID documents and KYC documents

Dutch and foreign passports, identity cards, driving licences, MRZ zone extraction, NFC reading where available, plus supplementary KYC evidence such as utility bills and BRP extracts. Part of broader KYC-AML compliance software for banks, insurers and notaries.

📧

CVs, emails and attachments

CVs in any format (PDF, Word, scanned), incoming emails with multiple attachments, transport documents, CMR forms and tender specification PDFs for construction and infrastructure projects. The pipeline first classifies the document type, selects the appropriate schema and then extracts. For the construction sector, see also our CMR/transport document page.

OCR, ICR, IDP and vision LLMs: what works for what

The field is full of overlapping jargon. A brief orientation helps determine which type of solution suits your document flow, and which combination you typically need in production.

Traditional OCR (Tesseract, ABBYY FineReader)

Classic optical character recognition reads pixels and returns text. It works well on clean scans of typed documents. Drawbacks: no understanding of layout, no interpretation of tables, and sensitivity to skewed scans and handwriting. Useful as a pre-processor for downstream NLP, but structure requires regex engineering or templates.

ICR (Intelligent Character Recognition)

A variant of OCR that also attempts to recognise handwritten characters. Older ICR engines are limited; modern handwriting recognition relies on transformer models. Relevant for claim forms, medical notes and historical archives.

IDP (Intelligent Document Processing)

Umbrella term for end-to-end platforms: Hyperscience, Rossum, AWS Textract, Azure Document Intelligence, Google Document AI. They combine OCR, layout analysis, classification and extraction. Well suited if you want a quick out-of-the-box route and process standard document types; less suited to unique schemas or strict EU data residency requirements.

Vision LLMs and layout-aware models

Donut, LayoutLMv3, Pix2Struct and current vision models from Anthropic Claude and OpenAI GPT-4o read documents holistically, taking image, text and layout into account at the same time. They deliver structured JSON output directly via structured output schemas. Advantages: no regex, no template engineering, rapid iteration. Points to note: cost per page and data residency require careful architectural choices.

In practice we often combine several layers: a lightweight OCR or layout analyser (for example via Azure Document Intelligence Read or an open-source stack) as pre-processing, a classifier that determines the document type, and then a specialised extraction model — sometimes a fine-tuned LayoutLMv3 for high-volume standard documents, sometimes a Vision LLM for the long tail and exceptions. In our experience, this combination strikes the best balance between cost, latency and accuracy.

The right approach for you depends on volume, data sensitivity, the desired straight-through processing rate and your existing infrastructure. In an intake conversation we map these variables and make a proposal with a proof of concept — see also our page on the AI discovery workshop.

Applications by sector

Document extraction has its own character in each sector. Below are four sectors where we regularly work; the patterns are widely applicable.

Accountancy and bookkeeping

For accountancy firms and finance departments we automate accounts payable: incoming invoices are read, matched to purchase orders, posted to the correct general ledger account (based on supplier and previous postings) and forwarded to the approver. Can be combined with our AI for accountancy page and accountants portal.

Insurance and insurance brokers

Policy intake and claims handling are document-heavy. The model reads policy documents from delegated underwriters, classifies claims based on loss reports and routes them to the right handler. See also AI for insurance and insurance app development.

Legal and notarial services

Legal teams use clause extraction for due diligence (data rooms with hundreds of contracts), contract review and risk flagging. For notary offices we extract deeds, land registry documents and Chamber of Commerce extracts. Output flows into your DMS or CLM.

Shared service centres and back office

For large organisations with a central back office (energy, telecoms, housing associations, healthcare providers), we process email post, customer letters and forms. The pipeline first classifies (quote, complaint, contract change), then extracts and routes each item to the right department. Often linked to workflow engines through our document workflow automation.

How we deliver a document extraction pipeline

From initial analysis to production in four phases. We work iteratively and validate each component with real documents from your own archive.

Document audit

You provide a representative set of several hundred documents. We analyse formats, quality and layout variation, and determine which document types represent the most volume. Result: a priority list and an initial schema proposal.

Schema and golden dataset

For each document type we define a JSON schema (required fields, types, validations). You label or validate a golden dataset of 100-500 documents. This becomes our ground truth for both training and testing.

Model selection and pipeline

We choose between a vision LLM, a fine-tuned LayoutLMv3 or Donut, or a hybrid approach. Around that, we build the pipeline: intake (email, API, upload), preprocessing, classification, extraction, schema validation and human-in-the-loop for low-confidence cases.

Production and monitoring

Deployment in your cloud or on-premises, integration with downstream systems, monitoring of accuracy and throughput, and a feedback loop in which corrected documents continuously improve the model. Includes alerting when accuracy falls below a threshold.

Not yet sure about a large project?

Test your idea first: a working prototype in 1 day

With OneDayBuild, we turn your idea into something tangible in one day for €1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.

Explore OneDayBuild →

Technology and tooling

We work vendor-agnostically and choose based on your requirements around data residency, cost and accuracy. Below is the stack we often use — your situation may call for a different combination.

Tesseract ABBYY FineReader AWS Textract Azure Document Intelligence Google Document AI Hyperscience Rossum LayoutLMv3 Donut Pix2Struct Claude Vision GPT-4o Hugging Face Transformers PyTorch LangChain LlamaIndex Pydantic FastAPI Docker PostgreSQL PEPPOL UBL NEN 2082 (e-archive)

Alongside model engineering, we pay close attention to structured output schemas (Pydantic, JSON Schema, Anthropic tool use or OpenAI function calling), to retry and validation logic, and to archival compliance under NEN 2082 for organisations that must retain documents in line with e-archiving legislation. eIDAS considerations (signature validation on contracts, integrity of deeds) also receive attention where relevant.

If you use Hyarchis for documents in the mortgage process, see our page on a Hyarchis integration.

If you would rather integrate an existing document recognition platform than build your own, see our page on a Klippa integration.

Architecture, accuracy and human-in-the-loop

A document extraction system that is 95% accurate will get the remaining 5% wrong. The question is how you spot, validate and feed back those errors into the model. Our architecture is designed for exactly that.

Confidence scores and thresholds

We deliver a confidence score for each field. Fields below your threshold (for example 0.90 for IBAN, 0.98 for total amount) go to a review queue. Fields above the threshold flow straight through. You decide the balance between straight-through processing and quality assurance.

Human-in-the-loop interface

A lightweight review interface shows the document alongside the extracted fields and their detected positions. Staff correct in seconds rather than minutes. Every correction is logged as an improvement signal for model retraining or prompt tuning.

Schema validation and business rules

Alongside confidence, we validate the output: does the sum of line items match the total, is the VAT rate consistent, does the IBAN check digit pass, does the KvK number exist? Business rules catch plausible but incorrect extractions that the model alone cannot see.

Continuous evaluation

A dashboard measures field-level accuracy, turnaround time, escalations and cost per document — over time and per document type. An alert follows any regression. We use golden datasets for regression testing with every model update.

For sectors with strict data residency requirements (healthcare, government, financial institutions), we build the extraction pipeline on your own infrastructure or a European cloud — no documents leave your controlled environment. For less sensitive flows, a vision LLM as a service within Europe can be used, provided there is a signed DPA and zero-retention is configured. We discuss which route suits you during the architecture phase.

Why Appfront for document extraction

Understanding document flows

We have built pipelines for accountancy firms, insurers, legal teams, housing associations and logistics service providers. That domain knowledge translates into pragmatic schemas and realistic expectations around accuracy.

No vendor lock-in

We choose models and cloud providers on a case-by-case basis. We won't force a Hyperscience or Rossum licence if an open-source LayoutLMv3 running on your own infrastructure is a better fit. Your schema and your data stay yours.

From POC to production

Many document AI projects stall between prototype and actual integration with ERP, DMS or policy administration systems. We support the entire journey: extraction, integration, monitoring and ongoing maintenance.

Frequently asked questions about AI document extraction

What is the difference between traditional OCR and AI document extraction?
Traditional OCR (Tesseract, ABBYY) converts pixels into text without understanding layout or meaning. AI document extraction combines layout analysis, classification and interpretation, and directly delivers a structured JSON object that matches your target schema. In practice, modern pipelines often combine both: OCR as pre-processing, and a vision LLM or layout-aware model for the extraction itself.
What level of accuracy is achievable?
That depends on the document type, the quality of the scans and the complexity of the schema. For standard invoices we typically achieve high accuracy on key fields such as supplier, total amount and date. For complex contracts or poorly scanned historical records, accuracy is lower. We measure and report field-level accuracy on a golden dataset; that is the honest benchmark.
Can this run on-premises for data residency reasons?
Yes. For sectors where data must not leave the premises (healthcare, government, some financial institutions), we build the pipeline on your own infrastructure or a European cloud. Open-source models such as LayoutLMv3, Donut or a local Llama model can run fully self-hosted.
How do you handle PEPPOL UBL and other structured formats?
PEPPOL UBL and e-invoice XML are already structured, so no extraction is needed, only a mapping to your internal schema. Our pipeline first classifies the format. For structured input it skips the extraction stage and validates straight away. For free-form PDFs or scans, the document goes through the extraction models. The result is the same schema, regardless of the input format.
Does this also work for handwritten damage claim forms?
Handwritten text is harder than typed text, but modern vision LLMs and specialised handwriting recognition models achieve reasonable results on legible forms. For critically important fields (for example the claim amount), we almost always recommend human-in-the-loop review.
How does this connect with our accounting or policy software?
We build integrations via REST APIs, webhooks or existing integration platforms. For accounting packages such as Exact Online, AFAS, Twinfield, Yuki and Visma.net we have standard mappers; for policy systems we work with Sequel, IDIT and in-house policy administrations. The extraction delivers the schema, the integration carries it through.
Does this comply with GDPR and NEN 2082?
GDPR: documents may contain personal data, so we set up the pipeline following privacy by design, with a DPA, zero retention at external model providers and clear retention periods. NEN 2082 (e-archiving): we implement metadata, integrity safeguards and durable storage where you are subject to archiving obligations. The legal context determines the exact implementation.
How long does an implementation typically take?
A proof of concept covering a single document type is usually ready within a few weeks. A production pipeline with multiple document types, integrations and monitoring takes longer, depending heavily on the complexity of your systems landscape. We work in iterative sprints with interim deliveries, so you can see early on whether the approach works.

Want to automate your document flow?

Discuss your document flow with Appfront. Together we'll analyse which document types take up the most time and where AI extraction will deliver the greatest gains, with no obligation whatsoever.

Schedule a conversation

Edit content