Skip to content

Accuracy benchmark · v1.0 · measured 3 October 2026

How accurately IPP reads invoices, measured

We ran 181 synthetic test documents through live production: clean PDFs, phone photos, scans, handwritten invoices, credit notes, several currencies and VAT layouts, and deliberately hard cases. 96.1% came back with every key header field right, and 98.7% of all key header fields were right.

Every document IPP returned at high confidence was right on every scored field. The 8 documents with something wrong all came back at medium or low confidence, and 4 of them were also flagged for review.

Measured on 3 October 2026 · 181 synthetic test documents, live production

Headline results, measured on 3 October 2026, 181 synthetic test documents, live production
MeasureResultPrevious run2 October 2026
Documents with every key header field right174 of 18196.1%92.8%
Key header fields right1,555 of 1,576 fields98.7%97.8%
Line items (mean F1, 1 = every line right)174 of 179 documents with every line right0.9830.977
Invoices whose printed figures do not add up, flagged5 of 5100%60%
Consistent documents wrongly flagged2 of 1721.2%4.1%
Estimated AI cost per documentAzure and OpenAI at list price$0.045—
Extraction time per document, median95th percentile 19.2 s; see the method for queue time11.2 s—

Key header fields: vendor, invoice number, invoice date, due date (where printed), currency, subtotal, tax, total and document type. IPP’s own benchmark on invented documents; your documents will differ. See the limits and the method.

What’s in the test set

181 documents built to be hard to read

Every document is synthetic. Every vendor, customer, invoice number and amount is invented; only the town names are real places. No customer document was read, copied or imitated to make them.

Each document is drawn from a specification that is also its answer key, so the right answer is known to the cent. The set leans hard on purpose: 71 of the 181 documents are graded hard. Photos and scans take each kind of damage at three strengths on the same page, so the results show how accuracy falls as the damage grows.

Clean digital PDFs

31 docs

Headers on the left, right, centred or in a banner; gridded and borderless tables; two columns; landscape. US, UK, EU, South African and long date formats, comma and dot decimals.

Currencies

19 docs

USD, ZAR, NAD, EUR, GBP and ZMW, printed as symbols or codes, including a Namibian invoice that prints a bare "$". Dual-currency invoices with a separate "amount payable" box, and lines priced in several currencies.

VAT layouts

14 docs

Prices with and without VAT, a VAT column per line, a "Sub Total" printed after VAT, zero-rated exports, percentage lines and withholding deducted below the total.

Credit notes

10 docs

"Credit Note", "Nota de Crédito", "R -" amounts, brackets, a "CR" suffix, and a credit note printed without any minus signs.

Multi-page invoices

10 docs

Two to eight pages, lines continuing across pages, carried and brought-forward rows, totals only on the last page.

Freight and forwarding

16 docs

Carrier invoices with fuel surcharges and accessorials, clearing and forwarding invoices mixing VAT-free disbursements with VAT-able fees, and transport invoices printing trip, load and truck references.

Scans

15 docs

Grayscale scans, noise, fax-quality black and white, punch holes, and stamps printed over the text.

Phone photos

45 docs

Perspective, rotation, upside-down pages, motion and focus blur, low light, shadow, heavy JPEG compression, low resolution, crumples and folds, coffee stains, pen marks, cropped edges, a photo of a screen, and a combined phone capture.

Handwritten invoices

8 docs

Pre-printed cash-book and delivery-note invoices filled in by hand, at three legibility grades.

Hard cases

13 docs

Several invoices in one PDF, supplier statements, an order number printed beside the invoice number, no invoice number at all, and invoices whose printed figures do not add up, which must be flagged rather than "fixed".

Six of the documents, and what IPP returned

Synthetic UK invoice photographed with severe motion blur; the text is smeared sideways

Severe motion blur

Misread, and flagged at medium confidence: the misread figures did not add up.

photo-motion-blur-severe

Synthetic Namibian delivery note and invoice with the customer, lines and totals filled in by hand

Poor handwriting

Fully correct, at high confidence.

handwritten-invoice-book-poor-6

Synthetic clearing and forwarding invoice with lines in US dollars and an amount payable box in South African rand

Two currencies

Fully correct: the payable currency (ZAR) and its total.

currency-dual-usd-zar-1

Synthetic Namibian invoice in a typewriter font whose total includes VAT, with the VAT shown underneath

VAT included in the total

Fully correct: the net subtotal worked out, not the printed total.

tax-inclusive-total-na

Synthetic UK invoice photographed after being crumpled and folded into quarters

Crumpled and folded

Fully correct.

photo-crumple-severe

Synthetic South African invoice whose printed total is larger than the subtotal plus VAT

Figures that do not add up

Read as printed and flagged: the total is not subtotal plus VAT.

hard-total-inflated

Top half of each page shown. Synthetic documents; any resemblance to a real business is coincidental.

Test it yourself, on IPP or anything else

17 documents from the set, with the right answer for each and the scoring rules, in one zip (495 KB). Run them through any invoice reader and compare.

Download the samples

Results

Where IPP reads well, and where it does not

“Fully correct” means every key header field the document is scored on was right. One wrong digit in a total makes the whole document wrong.

By category

Fully correct documents by category, 3 October 2026
CategoryFully correctShare
Clean digital PDFs31 / 31
100%
Currencies19 / 19
100%
VAT layouts14 / 14
100%
Credit notes10 / 10
100%
Multi-page invoices10 / 10
100%
Freight and forwarding16 / 16
100%
Scans15 / 15
100%
Phone photos42 / 45
93.3%
Handwritten invoices7 / 8
87.5%
Hard cases10 / 13
76.9%

By difficulty

Easy62 / 62
100%
Medium47 / 48
97.9%
Hard65 / 71
91.5%

Photos and scans, by damage

The same 20 pages at three strengths of blur, glare, folds, stains, cropping and the rest.

Mild20 / 20
100%
Moderate20 / 20
100%
Severe17 / 20
85%

Field by field

Key header fields correct, field by field, 3 October 2026
FieldRightShare
Vendor name181 / 181100%
Invoice number175 / 17997.8%
Invoice date175 / 17997.8%
Due date (where printed)138 / 14098.6%
Currency178 / 17999.4%
Subtotal (net)176 / 17998.3%
Tax175 / 17997.8%
Total176 / 17998.3%
Document type181 / 181100%

Benchmark history

The same 181 documents went through production on 2 October 2026 and again on 3 October 2026. Between the two runs, changes to how IPP reads VAT-inclusive layouts, credit notes, supplier statements and files holding several invoices were deployed. Both runs are scored by the same rules.

The same documents in two production runs
Measure2 October 20263 October 2026
Fully correct documents168 / 181174 / 181
Key header fields right97.8%98.7%
Line items, mean F10.9770.983
Invoices that do not add up, flagged3 / 55 / 5
Files with several invoices, flagged0 / 22 / 2
Consistent documents wrongly flagged7 / 1722 / 172
VAT layouts, fully correct11 / 1414 / 14
Credit notes, fully correct9 / 1010 / 10
Multi-page invoices, fully correct9 / 1010 / 10
Phone photos, fully correct41 / 4542 / 45

Categories not listed had the same result in both runs.

When IPP isn’t sure

A wrong number should never look certain

Reading most documents right matters less than knowing which ones to check. In this run the documents IPP got wrong were the ones it said to look at.

Confidence on every document

Each document comes back High, Medium or Low. All 104 documents returned at High were right on every scored field. All 8 documents with a wrong field came back at Medium (6) or Low (2). Medium is a request to look, not a verdict: 69 of the 77 documents at Medium or Low had nothing wrong.

Figures that do not add up are flagged

IPP checks subtotal, tax and total against each other and flags a mismatch with the difference, rather than “fixing” it. All 5 invoices printed with figures that do not add up were flagged. The 2 consistent documents that were flagged were severe-blur photos IPP misread: the flag caught the misreading.

Statements are not bills

A supplier statement uploaded as an invoice is recognised and held, so it cannot be approved or pushed to accounting as a bill. Both statements in the set (2 of 2) came back as “not an invoice”, at low confidence, flagged.

Several invoices in one file

A PDF holding several invoices is recognised and flagged, asking for the invoices as separate files. IPP does not split the file. Both such files in the set (2 of 2) came back flagged, at low confidence.

Of the 7 documents that should be flagged, 7 were. Nothing is pushed to QuickBooks, Xero or Zoho Books without an approval, and auto-approve stays off until a workspace switches it on.

Every miss in this run

The 7 documents that were not fully correct, and the one fully correct document with a line-item miss.

Documents with a wrong field, 3 October 2026
DocumentWhat went wrongConfidenceFlagged
photo-motion-blur-severePhone photosSevere motion blur. The invoice number, dates, currency, amounts and lines were misread; the misread figures did not add up, so it was flagged.mediumYes
photo-defocus-blur-severePhone photosSevere focus blur. The invoice number, dates, amounts and lines were misread; flagged because the misread figures did not add up.mediumYes
photo-crop-edges-severePhone photosSeverely cropped edges. The invoice number was read wrong; every other field was right.mediumNo
hard-multiple-invoices-2Hard casesTwo invoices in one PDF, scored against the first. The tax and lines were read wrong; the file was flagged.lowYes
hard-multiple-invoices-3Hard casesThree invoices in one PDF, scored against the first. The date, amounts and lines were read wrong; the file was flagged.lowYes
hard-missing-invoice-number-1Hard casesNo invoice number printed, only an order number. IPP returned the order number as the invoice number instead of leaving it empty.mediumNo
handwritten-invoice-book-neat-2Handwritten invoicesHandwritten delivery-note invoice. The invoice date was read as a different date; every other field was right.mediumNo
photo-motion-blur-moderatePhone photosModerate motion blur. Every header field was right; one of four lines was misread.mediumNo

Honest limits

What these numbers do not tell you

  • It is IPP’s own benchmark. We built the documents, ran them and scored them. Nobody independent has audited it. The sample documents are here so you can check the scoring yourself.
  • The documents are synthetic. They were generated to imitate common layouts and damage, not collected from real suppliers.
  • Real documents vary. Your suppliers’ layouts, languages, scanners and phones are not ours. The set covers the categories above and no others. Results on your documents will differ, in either direction: try it on them.
  • It is one run. IPP reads documents with Azure Document Intelligence and GPT-4o. On the hardest documents, such as severe blur and poor handwriting, those models can return a different reading on a different run, so the same document can be right one day and wrong the next.
  • The weakest areas are named above. Severely blurred or cropped photos, PDFs holding several invoices (flagged, not split), an invoice with no invoice number printed, and handwritten dates.
  • Cost and time are estimates from one run. Cost is computed at the providers’ list prices, not taken from an invoice. Time varies with load.

Method

How the documents were made and scored

Making the documents

Each document is generated from a specification that holds both what is drawn and what a correct reading returns. Names, numbers and amounts come from invented word lists. Handwriting uses a handwriting font with random per-letter jitter at three legibility grades. Photos and scans apply one kind of damage at three strengths to the same clean page.

What ran

The live production service on 3 October 2026, unchanged: 173 documents through IPP’s public API and 8 handwritten invoices uploaded through the app, the way a customer sends them. The documents went through the same reading, checks and confidence grading as a customer’s.

Scoring, field by field

  • Amounts (subtotal, tax, total): exact to the cent. The subtotal is the net amount before VAT, on VAT-inclusive layouts too. Credit notes compare the amounts without their sign.
  • Dates: the exact calendar date. The due date is scored only where one is printed.
  • Invoice number: exact, ignoring case, spaces, “#” and a leading label. Where none is printed, right only when left empty.
  • Vendor: the same name after dropping case, punctuation and words like Ltd, Pty or LLC, or a near match.
  • Currency: the ISO code; on a dual-currency invoice, the currency of the amount payable.
  • Document type: invoice, credit note, or “not an invoice” for a supplier statement.
  • Line items: each expected line matched by amount (within a cent) and description. F1 combines lines missed and lines invented; 1 means every line right and nothing extra.
  • Flags: an invoice whose printed figures do not add up, or a file holding several invoices, should be flagged; any other document should not. Several invoices in one file are scored against the first.

Time and cost

Extraction time is the time production reports for reading each document: median 11.2 s, 95th percentile 19.2 s. Upload to result took longer in this run (median 57.8 s) because the benchmark sent its 173 API documents in one batch and checked for results at intervals; that measures the benchmark’s own queue, not what a single upload waits. Cost is estimated: Azure from each document’s page count and GPT-4o from production’s own token count, both at list price; $8.08 for the 181 documents.

Run notes

  • One document, credit-r-minus-1, was re-run on its own after a worker restart left its first job unfinished.
  • The test set (v1) holds 191 documents. 10 of them test a document type that is in beta and are left out of every figure on this page, in both runs.
  • The previous run’s own report counted the two multi-invoice files among the documents that should not be flagged. The history above scores both runs by the current rules, so its figures can differ slightly from that report.

Versions

Test set
v1, seed 42, generator 1, digest 9d53a1f10774e1e4
Run
prod-combined-postfix-20261003 (2026-10-03)
Previous run
prod-combined-v1-s42 (2026-10-02)
Page
v1.0, published 2026-10-03

The only benchmark that matters is your documents.

Start a 14-day trial with 50 invoices and no credit card, or try one document without signing up. If you want to talk about your documents first, write to us.