AI document extraction and OCR: from paper and PDF to structured data
Back offices, shared service centres, accountancy firms, insurers and legal teams process thousands of invoices, policies, contracts, Chamber of Commerce extracts, valuations and email attachments every day. Traditional OCR recognises characters but leaves the interpretation to people. With modern AI document extraction, including vision LLMs, layout-aware models and structured-output schemas, software reads documents in any format and delivers validated JSON straight to your accounting, policy or case management system.
Discuss your document workflow View applicationsWhat AI document extraction and modern OCR do
Document processing has fundamentally changed in the past five years. Where Tesseract and ABBYY once read characters and you had to write regular expressions and templates yourself, vision LLMs and layout-aware models now extract meaning from a document, including tables, handwritten notes, stamps and scanned copies of varying quality.
An invoice may arrive as a structured PEPPOL UBL file, as a PDF generated by an ERP, as a scanned paper invoice or as a photo taken on a phone. For each format, your system must correctly identify the supplier, the VAT number, the invoice lines, the final totals and the payment reference, and ideally also detect whether the posting lines make sense. Classic OCR produces text, not meaning. AI document extraction delivers a validated JSON object straight away that matches the schema your accounting package or case management system expects.
The same applies to Chamber of Commerce extracts, cadastral documents and deeds, policy schedules, claim files, purchase agreements, mortgage offers, passports and ID documents, CVs, receipts and even the content of incoming emails. Each document type has its own schema; Appfront designs that schema together with you, trains or fine-tunes the model, and builds the pipeline that receives, classifies, extracts, validates and forwards documents to your existing systems, with human-in-the-loop review for cases where the confidence score falls below your threshold.
The difference from traditional OCR lies in three things: layout awareness (models understand columns, tables and visual grouping), contextual understanding (the word "total" is weighted differently when it sits bottom right than in the header), and direct structure (output is a typed JSON schema, not a block of flat text you still have to parse yourself). That difference determines whether a document is processed in 30 seconds and then flows through automatically in 90% of cases, or whether every invoice has to be handled again by human hands.
Document types we process
We build extraction pipelines for almost any document that appears in a Dutch back office. Below are the six most common categories. For each one, we have a reference schema that we refine together with you.
Invoices and receipts (AP automation)
PEPPOL UBL via Mijn Overheid or supplier portals, free-form PDF invoices, scanned paper invoices, till receipts and restaurant bills. We extract supplier, KvK and VAT numbers, IBAN, invoice number, invoice date, due date, line items with VAT rate, and final totals. Output aligns with Exact Online, AFAS, Twinfield, Yuki, Visma.net or a custom ERP. See also our invoice recognition page.
KvK extracts and land registry documents
Extracts from the Trade Register, cadastral records, ownership information, mortgage deeds and transfer deeds. Alongside direct extraction via the KvK API and the Kadaster integration, we also process scanned historical files where no API source is available.
Policies and claims files
Policy schedules, policy terms, claim forms (loss reports, expert reports), medical reports and correspondence from insurance intermediaries. The model recognises coverages, excess, premiums, policy numbers and claim items. Suitable for health insurers, property and casualty insurers, life insurers, reinsurers and managing general agents.
Contracts and legal files
Purchase agreements, employment contracts, NDAs, SLAs, lease agreements and general terms and conditions. Clause extraction identifies notice periods, penalty clauses, governing law clauses, price indexation and termination conditions. Output goes to your contract management system or CLM tooling.
Passports, ID documents and KYC documents
Dutch and foreign passports, identity cards, driving licences, MRZ zone extraction, NFC reading where available, plus supplementary KYC evidence such as utility bills and BRP extracts. Part of broader KYC-AML compliance software for banks, insurers and notaries.
CVs, emails and attachments
CVs in any format (PDF, Word, scanned), incoming emails with multiple attachments, transport documents, CMR forms and tender specification PDFs for construction and infrastructure projects. The pipeline first classifies the document type, selects the appropriate schema and then extracts. For the construction sector, see also our CMR/transport document page.
OCR, ICR, IDP and vision LLMs: what works for what
The field is full of overlapping jargon. A brief orientation helps determine which type of solution suits your document flow, and which combination you typically need in production.
Traditional OCR (Tesseract, ABBYY FineReader)
Classic optical character recognition reads pixels and returns text. It works well on clean scans of typed documents. Drawbacks: no understanding of layout, no interpretation of tables, and sensitivity to skewed scans and handwriting. Useful as a pre-processor for downstream NLP, but structure requires regex engineering or templates.
ICR (Intelligent Character Recognition)
A variant of OCR that also attempts to recognise handwritten characters. Older ICR engines are limited; modern handwriting recognition relies on transformer models. Relevant for claim forms, medical notes and historical archives.
IDP (Intelligent Document Processing)
Umbrella term for end-to-end platforms: Hyperscience, Rossum, AWS Textract, Azure Document Intelligence, Google Document AI. They combine OCR, layout analysis, classification and extraction. Well suited if you want a quick out-of-the-box route and process standard document types; less suited to unique schemas or strict EU data residency requirements.
Vision LLMs and layout-aware models
Donut, LayoutLMv3, Pix2Struct and current vision models from Anthropic Claude and OpenAI GPT-4o read documents holistically, taking image, text and layout into account at the same time. They deliver structured JSON output directly via structured output schemas. Advantages: no regex, no template engineering, rapid iteration. Points to note: cost per page and data residency require careful architectural choices.
In practice we often combine several layers: a lightweight OCR or layout analyser (for example via Azure Document Intelligence Read or an open-source stack) as pre-processing, a classifier that determines the document type, and then a specialised extraction model — sometimes a fine-tuned LayoutLMv3 for high-volume standard documents, sometimes a Vision LLM for the long tail and exceptions. In our experience, this combination strikes the best balance between cost, latency and accuracy.
The right approach for you depends on volume, data sensitivity, the desired straight-through processing rate and your existing infrastructure. In an intake conversation we map these variables and make a proposal with a proof of concept — see also our page on the AI discovery workshop.
Applications by sector
Document extraction has its own character in each sector. Below are four sectors where we regularly work; the patterns are widely applicable.
Accountancy and bookkeeping
For accountancy firms and finance departments we automate accounts payable: incoming invoices are read, matched to purchase orders, posted to the correct general ledger account (based on supplier and previous postings) and forwarded to the approver. Can be combined with our AI for accountancy page and accountants portal.
Insurance and insurance brokers
Policy intake and claims handling are document-heavy. The model reads policy documents from delegated underwriters, classifies claims based on loss reports and routes them to the right handler. See also AI for insurance and insurance app development.
Legal and notarial services
Legal teams use clause extraction for due diligence (data rooms with hundreds of contracts), contract review and risk flagging. For notary offices we extract deeds, land registry documents and Chamber of Commerce extracts. Output flows into your DMS or CLM.
Shared service centres and back office
For large organisations with a central back office (energy, telecoms, housing associations, healthcare providers), we process email post, customer letters and forms. The pipeline first classifies (quote, complaint, contract change), then extracts and routes each item to the right department. Often linked to workflow engines through our document workflow automation.
How we deliver a document extraction pipeline
From initial analysis to production in four phases. We work iteratively and validate each component with real documents from your own archive.
Document audit
You provide a representative set of several hundred documents. We analyse formats, quality and layout variation, and determine which document types represent the most volume. Result: a priority list and an initial schema proposal.
Schema and golden dataset
For each document type we define a JSON schema (required fields, types, validations). You label or validate a golden dataset of 100-500 documents. This becomes our ground truth for both training and testing.
Model selection and pipeline
We choose between a vision LLM, a fine-tuned LayoutLMv3 or Donut, or a hybrid approach. Around that, we build the pipeline: intake (email, API, upload), preprocessing, classification, extraction, schema validation and human-in-the-loop for low-confidence cases.
Production and monitoring
Deployment in your cloud or on-premises, integration with downstream systems, monitoring of accuracy and throughput, and a feedback loop in which corrected documents continuously improve the model. Includes alerting when accuracy falls below a threshold.
Test your idea first: a working prototype in 1 day
With OneDayBuild, we turn your idea into something tangible in one day for €1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.
Explore OneDayBuild →Technology and tooling
We work vendor-agnostically and choose based on your requirements around data residency, cost and accuracy. Below is the stack we often use — your situation may call for a different combination.
Alongside model engineering, we pay close attention to structured output schemas (Pydantic, JSON Schema, Anthropic tool use or OpenAI function calling), to retry and validation logic, and to archival compliance under NEN 2082 for organisations that must retain documents in line with e-archiving legislation. eIDAS considerations (signature validation on contracts, integrity of deeds) also receive attention where relevant.
If you use Hyarchis for documents in the mortgage process, see our page on a Hyarchis integration.
If you would rather integrate an existing document recognition platform than build your own, see our page on a Klippa integration.
Architecture, accuracy and human-in-the-loop
A document extraction system that is 95% accurate will get the remaining 5% wrong. The question is how you spot, validate and feed back those errors into the model. Our architecture is designed for exactly that.
Confidence scores and thresholds
We deliver a confidence score for each field. Fields below your threshold (for example 0.90 for IBAN, 0.98 for total amount) go to a review queue. Fields above the threshold flow straight through. You decide the balance between straight-through processing and quality assurance.
Human-in-the-loop interface
A lightweight review interface shows the document alongside the extracted fields and their detected positions. Staff correct in seconds rather than minutes. Every correction is logged as an improvement signal for model retraining or prompt tuning.
Schema validation and business rules
Alongside confidence, we validate the output: does the sum of line items match the total, is the VAT rate consistent, does the IBAN check digit pass, does the KvK number exist? Business rules catch plausible but incorrect extractions that the model alone cannot see.
Continuous evaluation
A dashboard measures field-level accuracy, turnaround time, escalations and cost per document — over time and per document type. An alert follows any regression. We use golden datasets for regression testing with every model update.
For sectors with strict data residency requirements (healthcare, government, financial institutions), we build the extraction pipeline on your own infrastructure or a European cloud — no documents leave your controlled environment. For less sensitive flows, a vision LLM as a service within Europe can be used, provided there is a signed DPA and zero-retention is configured. We discuss which route suits you during the architecture phase.
Why Appfront for document extraction
Understanding document flows
We have built pipelines for accountancy firms, insurers, legal teams, housing associations and logistics service providers. That domain knowledge translates into pragmatic schemas and realistic expectations around accuracy.
No vendor lock-in
We choose models and cloud providers on a case-by-case basis. We won't force a Hyperscience or Rossum licence if an open-source LayoutLMv3 running on your own infrastructure is a better fit. Your schema and your data stay yours.
From POC to production
Many document AI projects stall between prototype and actual integration with ERP, DMS or policy administration systems. We support the entire journey: extraction, integration, monitoring and ongoing maintenance.
Frequently asked questions about AI document extraction
Want to automate your document flow?
Discuss your document flow with Appfront. Together we'll analyse which document types take up the most time and where AI extraction will deliver the greatest gains, with no obligation whatsoever.
Schedule a conversation