AI document processing that ends with the data in your systems.
Our AI document processing services take the PDFs, scans and email attachments your team keys in by hand, pull out the fields you need, check them against your own records, and post them to the system they belong in. A person reviews anything the model is unsure of, and every value can be traced back to the page it came from.

On this page
Documents we automate
- Supplier invoices, receipts and credit notes, posted to QuickBooks, Xero, NetSuite or your ERP.
- Purchase orders, delivery notes and bills of lading, matched against what was ordered and received.
- Insurance claims, explanations of benefits and remittance advice.
- Patient intake forms and referral letters, filed to the right record.
- Bank and card statements, for reconciliation.
- Contracts and leases, where the job is to find specific clauses, dates and amounts.
- Applications and onboarding packs that arrive as a bundle of different documents.
What the pipeline does
Reading the text is the easy part. The work that saves your team time happens around it:
- Intake. Documents arrive from a mailbox, a scanner folder, a portal upload or an API, and each one gets an ID and a copy stored where you control it.
- Classification. The system works out what each document is, and splits a 40-page bundle into its parts.
- Extraction. It pulls out the fields you need, with a confidence score for each one and the location on the page it came from.
- Validation. The values are checked against your data: the supplier exists, the PO number is real, the lines add up to the total, the invoice isn't a duplicate.
- Review. Anything below the confidence threshold, or anything that fails a check, goes to a person with the page and the extracted value side by side.
- Posting. Clean records go into your system through its API, once, with a link back to the source document.
Most of the accuracy comes from step 4. A model that reads "1,280.00" correctly still can't tell you it belongs to a supplier you stopped using last year, and your own records can.
Cloud extraction service or a language model
There are two kinds of engine, and most of our builds use both. Microsoft and Amazon publish their per-page prices:
| Service | What it does | Published price |
|---|---|---|
| Azure Document Intelligence, Read | Text from any page | $1.50 per 1,000 pages, less at high volume |
| Azure Document Intelligence, prebuilt models (invoice, receipt, layout and others) | Standard fields from common document types | $10 per 1,000 pages |
| Azure Document Intelligence, custom extraction | Fields from your own document types, trained on your samples | $30 per 1,000 pages |
| Amazon Textract, text detection | Text from any page | $1.50 per 1,000 pages |
| Amazon Textract, expense analysis | Invoice and receipt fields | $10 per 1,000 pages |
| Amazon Textract, forms | Key-value pairs from forms | $50 per 1,000 pages |
Azure prices are Microsoft's East US list prices and Textract prices are for US East (N. Virginia) and the first million pages a month, both checked on 25 September 2026.
A language model such as Claude or Gemini can read the same documents and handle ones that no prebuilt model covers: a handwritten note on a referral, an invoice with its totals in a table footnote, a contract where the answer depends on reading two clauses together. It's priced by tokens. Anthropic says each PDF page typically uses 1,500 to 3,000 text tokens, and the page is also sent as an image, which adds more. At Claude Haiku 4.5's $1 per million input tokens, the text alone comes to $1.50 to $3 per 1,000 pages, before image and output tokens.
A typical split: the cloud service handles the high-volume standard documents cheaply, and the language model handles the exceptions and the documents that need judgment. Our guide to what AI agent development costs compares model prices in more detail.
How we measure accuracy
Vendors quote accuracy measured on their own test sets, which says little about your documents, so we measure it on yours before you commit to a build:
- We take a sample of your real documents, including the messy ones, and record the correct value for every field by hand.
- We run the pipeline on that sample and report accuracy field by field. A supplier name at 99% and a line-item total at 90% need different handling.
- We set the review threshold from those results, and report the share of documents that go straight through with no human touch.
Those numbers become part of the fixed scope, and they are measured again after launch on live documents.
Where your documents go
Invoices carry bank details and patient forms carry health information, so where each page travels is decided at the design stage. The extraction can run in your own Azure or AWS account, and for the most sensitive documents it can run on your own servers with an open-weight model, as described on our private AI page. For health information, every service that handles the documents needs to be covered by a Business Associate Agreement, and we can sign one.
We are certified to ISO/IEC 27001:2022. The certificate and its scope are published.
When you don't need a custom build
If your team processes a few dozen invoices a month, the bill and receipt capture built into most accounting packages will probably do the job, and it costs less than any project. A custom pipeline pays off when the volume is in the hundreds or thousands of pages a month, when the documents don't fit a prebuilt model, or when the data has to land in more than one system. The audit tells you which side of that line you are on.
For the rest of the back office, see AI automation for small businesses and our custom API integrations.
What it costs
A pipeline for one document type going into one system is a small project. Several document types, validation against more than one system of record and a review queue your team works from every day make a larger one. We fix the price after the audit, once your document types, volumes and target systems are written down. For market figures, see what custom software development costs.
The Automation Audit
One workflow you name, mapped end to end, with the arithmetic done before anyone writes code.
- Scope
- One workflow you choose, traced end to end, including the steps nobody documented.
- Duration
- Two weeks, fixed.
- Fee
- Fixed, and quoted in full before we start. No hourly drift.
- You get
- A written map: what can be automated, what it would save in hours, what building it would cost, and what we would leave alone.
- You keep it
- The map is yours whether or not you hire us to build anything.
And if the audit shows the automation will not pay for itself inside twelve months, we will tell you, and we will not quote the build.
Questions we get asked about document processing
How accurate will it be on our documents?
We don't know until we test it, and neither does anyone else. That is why the audit measures accuracy on a sample of your own documents before we quote the build.
Can it read handwriting and poor scans?
Often, yes. Printed text on a clean scan is the easy case, and handwriting and faxed pages are harder. The field-by-field test shows which fields can go straight through and which should always be checked by a person.
What happens when the model gets a field wrong?
The validation step catches most errors before they reach your system, because the value has to agree with your own records. What gets through is caught by the daily reconciliation, and corrections made in the review queue are used to tune the thresholds.
Can you work with documents in languages other than English?
Yes. Both the cloud services and the language models read most major languages. We include those documents in the accuracy test so you see the results for each language.


