Intelligent document processing, built to be checked.
Extraction accuracy is the part everyone asks about and the part that matters least on its own. This guide walks through the full pipeline - classify, extract, validate, review, record - and the design decisions that determine whether the output can be trusted.
Written for finance, admin, compliance, and operations teams handling high volumes of invoices, certificates, contracts, or licences.
By Kanban Studios engineering · Last reviewed
The pipeline, end to end
Intelligent document processing - IDP - is often described as OCR with a model attached. That framing causes bad systems. A document pipeline that holds up in production has five distinct stages, and the interesting engineering is in the last three.
Ingest accepts the file from wherever it arrives and normalises it: a phone photo of a certificate, a scanned PDF, a native PDF with a real text layer, an email attachment. Classification determines what the document is, because everything downstream depends on knowing whether this is an invoice, a trade licence, or a delivery note. Extraction pulls the fields that matter for that document type. Validation checks the extracted values against rules and against other systems. Review puts anything uncertain in front of a person. Only then does the result get written anywhere that matters.
Ingest and normalise
De-skew, de-noise, handle multi-page and multi-document files, detect whether a real text layer exists or OCR is required. Poor input handling shows up later as unexplainable extraction failures.
Classify
Route by document type before extracting. A mis-classified document extracted against the wrong schema produces confidently wrong fields, which is worse than no output.
Extract
Per-type field schemas with a confidence score attached to every field, not just to the document. Confidence at document level hides exactly the fields you needed to check.
Validate
Format checks, arithmetic checks, cross-field consistency, and cross-system reconciliation against the record the document is supposed to match.
Review and record
A queue that shows only what needs a decision, with the source document beside the extracted values, and an audit entry for every accepted or corrected field.
Validation is what makes extraction useful
A field extracted with high confidence can still be wrong, and a field extracted with low confidence can still be right. Confidence tells you how sure the model is, not whether the value is correct. Validation is the layer that converts one into the other, and it is where most of the durable value in an IDP system sits.
Format validation checks the shape: a tax registration number has a known length, a date is a date, an amount parses. Arithmetic validation checks internal consistency: line items sum to the subtotal, tax is the expected proportion, the total matches. Cross-system validation is the strongest of the three: does this invoice reference a purchase order that exists, for a supplier on file, within the amount that was approved?
A document that passes all three can often proceed unattended. A document that fails any of them should go to a human with the specific failure surfaced - not simply flagged as needing attention.
Model confidence is an input to validation, never a substitute for it.
What UAE paperwork changes about the design
Several characteristics of documents circulating in the UAE genuinely change the engineering, as opposed to being marketing colour.
Bilingual documents are common: Arabic and English on the same page, sometimes in the same table, with right-to-left and left-to-right text interleaved. This affects OCR configuration, reading order, and how bounding boxes are handled - and it needs testing on real bilingual samples rather than assumed to work.
Expiry is a first-class concept. Trade licences, establishment cards, visas, insurance certificates, and equipment certifications all carry validity windows, and the operationally valuable output is often not the extraction at all but the alert thirty days before a document lapses. That means storing expiry as structured data and building a scheduler over it, not just filing the PDF.
Record retention matters. VAT and corporate tax regimes impose record-keeping obligations, so the archive, its retention period, and its retrievability are part of the system design rather than an afterthought. Confirm the specific obligations that apply to your entity with your own tax or legal advisers - they vary by activity and by free-zone status, and this guide is not a substitute for that advice.
Data residency and access are worth deciding explicitly: where documents are stored, where inference runs, who can retrieve an original, and how long anything is retained. These are architecture decisions with contractual consequences, best settled before a build rather than during one.
Designing the review queue
The review surface determines whether the system is adopted. It is not an admin table - it is the main product for the people who use it every day, and it should be designed with that seriousness.
The reviewer needs the source document and the extracted values side by side, with each extracted field linked to the region of the page it came from so a value can be verified with a glance rather than a search. Fields that failed validation should be first and should state why. Correcting a field should take one action, and the correction should be stored as training signal and as an audit entry.
Queue design matters as much as screen design. Items should be prioritised by consequence and deadline rather than arrival order, assignment should prevent two people reviewing the same item, and the queue should be visibly bounded so a reviewer can see the end of the work.
Field-level provenance
Every extracted value points back to where on the page it came from. Verification becomes a glance instead of a hunt.
Failures first
Sort by what failed validation and why, not by upload time. Reviewer attention is the scarce resource in the system.
One-action correction
Fix the value, accept, move on. Every extra click multiplies across thousands of documents and is where adoption is lost.
Corrections as data
Each correction is both an audit record and a measurement of where the extraction layer is actually weak.
Rolling it out without betting the process on it
The safe rollout is shadow mode. The pipeline runs against real documents in parallel with the existing manual process, and nobody acts on its output. You compare the two, measure disagreement per field and per document type, and find out where the system is genuinely weak on your documents rather than on a vendor's samples.
From there, thresholds are set per field from observed behaviour - typically conservative at first, so more goes to review than strictly needs to. As the correction rate on a given field settles, that field's threshold can be relaxed. Fields with legal or financial consequence often stay gated permanently, and that is a legitimate end state rather than a failure to fully automate.
Retrieval over the archive - answering questions across processed documents rather than one at a time - is worth adding only once extraction and validation are stable. Retrieval over an unreliable archive produces answers that are fluent, sourced, and wrong.
Questions we get asked.
How accurate is document extraction?
Accuracy depends entirely on document type, scan quality, layout consistency, and language mix, so any single figure quoted without those conditions is not meaningful. The useful approach is to measure it on your own documents in shadow mode before relying on it, per field rather than per document, and to design the validation and review layers on the assumption that some fields will be wrong.
Can it handle Arabic and English in the same document?
Bilingual documents are common in the UAE and need to be handled explicitly rather than assumed - mixed reading order, right-to-left text, and bilingual table headers all affect OCR configuration and extraction. It should be tested on real bilingual samples from your own document set as part of feasibility, before any commitment to build.
What happens to documents the system cannot read?
They route to the review queue with the specific failure surfaced - unreadable scan, unknown document type, failed validation rule. Nothing should silently proceed on a partial or low-confidence extraction, and nothing should be discarded; an unprocessable document is a work item, not an error to swallow.
Do documents have to leave our systems for this to work?
That is an architecture decision to make deliberately. Where documents are stored, where inference runs, who can retrieve an original, and how long anything is retained can all be constrained to meet your requirements - but the constraints need to be stated up front, because they materially affect the design and the choice of components.
The work behind this guide.
Services this covers
- AI Document IntelligenceReading, checking, and filing documents by hand is slow and error-prone. We build AI-assisted document systems that classify, read, validate, and extract from your documents - and route them into your workflow, with answers that stay traceable to their source.How we build it
- AI Workflow AutomationWe design AI-assisted workflows that remove the manual, repetitive steps slowing your team down - with a human in the loop wherever judgement matters. Built around how your business already works, not a template.How we build it
- Custom SoftwareWhen off-the-shelf software almost fits but never quite does, we build the system that does. Web apps, internal tools, portals, and APIs on dependable, maintainable foundations - software you own and can grow into.How we build it
Systems where this was built
- DocuMindA compliance document workspace that reads certificates, contracts and invoices - confidence-scored field extraction, expiring-document alerts, and a needs-a-human review queue before anything is approved.Read the case study
- MizanAn AI case officer for housing-loan arrears rescheduling built for a UAE ministry - deterministic policy enforcement, document intelligence, explainable recommendations, and human escalation for exceptional cases.Read the case study
- DeedFlowTransaction orchestration for fractional and tokenized property deals: KYC/AML checks, document verification, hard settlement gates, e-signature, and a full audit trail behind every stage.Read the case study