Skip to content
Applied AI

AI document processing that ends with the data in your systems.

Our AI document processing services take the PDFs, scans and email attachments your team keys in by hand, pull out the fields you need, check them against your own records, and post them to the system they belong in. A person reviews anything the model is unsure of, and every value can be traced back to the page it came from.

A steel sorting rack with blank sheets filed in its pigeonholes and one blue folder.
On this page

Documents we automate

  • Supplier invoices, receipts and credit notes, posted to QuickBooks, Xero, NetSuite or your ERP.
  • Purchase orders, delivery notes and bills of lading, matched against what was ordered and received.
  • Insurance claims, explanations of benefits and remittance advice.
  • Patient intake forms and referral letters, filed to the right record.
  • Bank and card statements, for reconciliation.
  • Contracts and leases, where the job is to find specific clauses, dates and amounts.
  • Applications and onboarding packs that arrive as a bundle of different documents.

What the pipeline does

Reading the text is the easy part. The work that saves your team time happens around it:

  1. Intake. Documents arrive from a mailbox, a scanner folder, a portal upload or an API, and each one gets an ID and a copy stored where you control it.
  2. Classification. The system works out what each document is, and splits a 40-page bundle into its parts.
  3. Extraction. It pulls out the fields you need, with a confidence score for each one and the location on the page it came from.
  4. Validation. The values are checked against your data: the supplier exists, the PO number is real, the lines add up to the total, the invoice isn't a duplicate.
  5. Review. Anything below the confidence threshold, or anything that fails a check, goes to a person with the page and the extracted value side by side.
  6. Posting. Clean records go into your system through its API, once, with a link back to the source document.

Most of the accuracy comes from step 4. A model that reads "1,280.00" correctly still can't tell you it belongs to a supplier you stopped using last year, and your own records can.

Cloud extraction service or a language model

There are two kinds of engine, and most of our builds use both. Microsoft and Amazon publish their per-page prices:

ServiceWhat it doesPublished price
Azure Document Intelligence, ReadText from any page$1.50 per 1,000 pages, less at high volume
Azure Document Intelligence, prebuilt models (invoice, receipt, layout and others)Standard fields from common document types$10 per 1,000 pages
Azure Document Intelligence, custom extractionFields from your own document types, trained on your samples$30 per 1,000 pages
Amazon Textract, text detectionText from any page$1.50 per 1,000 pages
Amazon Textract, expense analysisInvoice and receipt fields$10 per 1,000 pages
Amazon Textract, formsKey-value pairs from forms$50 per 1,000 pages

Azure prices are Microsoft's East US list prices and Textract prices are for US East (N. Virginia) and the first million pages a month, both checked on 25 September 2026.

A language model such as Claude or Gemini can read the same documents and handle ones that no prebuilt model covers: a handwritten note on a referral, an invoice with its totals in a table footnote, a contract where the answer depends on reading two clauses together. It's priced by tokens. Anthropic says each PDF page typically uses 1,500 to 3,000 text tokens, and the page is also sent as an image, which adds more. At Claude Haiku 4.5's $1 per million input tokens, the text alone comes to $1.50 to $3 per 1,000 pages, before image and output tokens.

A typical split: the cloud service handles the high-volume standard documents cheaply, and the language model handles the exceptions and the documents that need judgment. Our guide to what AI agent development costs compares model prices in more detail.

How we measure accuracy

Vendors quote accuracy measured on their own test sets, which says little about your documents, so we measure it on yours before you commit to a build:

  • We take a sample of your real documents, including the messy ones, and record the correct value for every field by hand.
  • We run the pipeline on that sample and report accuracy field by field. A supplier name at 99% and a line-item total at 90% need different handling.
  • We set the review threshold from those results, and report the share of documents that go straight through with no human touch.

Those numbers become part of the fixed scope, and they are measured again after launch on live documents.

Where your documents go

Invoices carry bank details and patient forms carry health information, so where each page travels is decided at the design stage. The extraction can run in your own Azure or AWS account, and for the most sensitive documents it can run on your own servers with an open-weight model, as described on our private AI page. For health information, every service that handles the documents needs to be covered by a Business Associate Agreement, and we can sign one.

We are certified to ISO/IEC 27001:2022. The certificate and its scope are published.

When you don't need a custom build

If your team processes a few dozen invoices a month, the bill and receipt capture built into most accounting packages will probably do the job, and it costs less than any project. A custom pipeline pays off when the volume is in the hundreds or thousands of pages a month, when the documents don't fit a prebuilt model, or when the data has to land in more than one system. The audit tells you which side of that line you are on.

For the rest of the back office, see AI automation for small businesses and our custom API integrations.

What it costs

A pipeline for one document type going into one system is a small project. Several document types, validation against more than one system of record and a review queue your team works from every day make a larger one. We fix the price after the audit, once your document types, volumes and target systems are written down. For market figures, see what custom software development costs.

// FIXED-FEE EVALUATION

The Automation Audit

One workflow you name, mapped end to end, with the arithmetic done before anyone writes code.

Scope
One workflow you choose, traced end to end, including the steps nobody documented.
Duration
Two weeks, fixed.
Fee
Fixed, and quoted in full before we start. No hourly drift.
You get
A written map: what can be automated, what it would save in hours, what building it would cost, and what we would leave alone.
You keep it
The map is yours whether or not you hire us to build anything.
The TechLand Commitment

And if the audit shows the automation will not pay for itself inside twelve months, we will tell you, and we will not quote the build.

Start with the audit

Questions we get asked about document processing

How accurate will it be on our documents?

We don't know until we test it, and neither does anyone else. That is why the audit measures accuracy on a sample of your own documents before we quote the build.

Can it read handwriting and poor scans?

Often, yes. Printed text on a clean scan is the easy case, and handwriting and faxed pages are harder. The field-by-field test shows which fields can go straight through and which should always be checked by a person.

What happens when the model gets a field wrong?

The validation step catches most errors before they reach your system, because the value has to agree with your own records. What gets through is caught by the daily reconciliation, and corrections made in the review queue are used to tune the thresholds.

Can you work with documents in languages other than English?

Yes. Both the cloud services and the language models read most major languages. We include those documents in the accuracy test so you see the results for each language.

Sources