Document Intelligence

Paperwork turned into structured, checkable records.

Documents arrive in whatever shape the sender chose. The job is to turn that into a fixed structure your systems and your finance team can both trust.

Team member checking extracted field values against a stack of printed operational documents
The pipeline

Eight stages, one of which is a person.

Extraction alone is not the deliverable. Validation, the confidence threshold and the review queue are what make the output safe to post into a finance or operations system.

Corrections made by reviewers are captured, so the same failure pattern can be addressed during optimisation rather than repeated indefinitely.

  1. 01
    Ingest

    Documents arrive from email, upload, a shared folder or a scanning workflow.

  2. 02
    Classify

    The document type is identified so the correct field set and rules apply.

  3. 03
    Extract

    Agreed fields are read into a fixed structure, with the source location retained.

  4. 04
    Validate

    Format, arithmetic and cross-reference checks against your own records.

  5. 05
    Score

    Each field carries a confidence indication; the threshold is set by you.

  6. 06
    Route

    Clean items continue; low-confidence or incomplete items go to the review queue.

  7. 07
    Review

    A person confirms or corrects, and the correction is captured for tuning.

  8. 08
    Post

    The approved record is written into the system of record with an audit entry.

Operational documents, invoices and forms laid out and organised on a desk for structured review
Document types

Structured output only makes sense once the fields are agreed.

We define the field set per document type with the team that uses the data — not from a generic template.

Typical document types and fields

Indicative examples. Your actual field set is confirmed during scoping and written into the scope of work.

Supplier invoices

Fields
Supplier, invoice number, date, line values, VAT, total, PO reference

Purchase orders

Fields
PO number, requester, cost centre, items, approved value

Application forms

Fields
Applicant details, requested service, declared information, attachments

Delivery documentation

Fields
Reference, consignee, quantities, dates, exceptions noted

Contracts and agreements

Fields
Parties, term dates, renewal notice, values, key clauses flagged

Identity and licence documents

Fields
Document type, holder name, number, expiry — handled per your policy

Operational correspondence

Fields
Sender, subject, request type, referenced records, required action

Internal forms

Fields
Submitter, department, request category, approval path

Two colleagues reviewing a queue of flagged documents on screen at a shared desk
Confidence thresholds

The threshold is a commercial decision, and it is yours.

A stricter threshold sends more items to review and catches more errors. A looser one moves faster and accepts more risk. There is no universally correct setting — it depends on the value and consequence of the document.

We set it with you, monitor how much reaches the queue in the first cycles, and adjust deliberately rather than quietly.

We do not publish accuracy percentages. Extraction quality depends on your document quality, and any figure quoted before seeing your paperwork would be marketing, not measurement.

Colleagues reviewing a printed data-handling policy and marking approvals with a pen
Handling sensitive documents

What is stored, and what is deliberately not.

Identity documents, personal data and commercially sensitive contracts are handled under an agreed policy: retention period, access list, storage location and deletion route are all stated before processing starts.

Where a document class is too sensitive to process, we exclude it from the pipeline entirely rather than applying a weaker control to it.

Data and security approach →

Start with one document type.

Working on your real paperwork gives an honest read on quality, exception volume and reviewer effort.