the assay report

Accuracy and cost, measured

98.7%of every human-checkable field was read correctly — measured over 1,361 documents, computed live from the register, never asserted. 987‰ fine, in assay-office marks.
98.7%
invoice fields read correctly

the headline corpus — synthetic invoices, measured live below

97.2% · 82.8%
receipt amounts · receipt lines

real photographed receipts — amounts and menu lines graded separately

$0.000022
of electricity per document

measured GPU board energy while it read

the headline corpus — measured live

Invoices — the live headline

Synthetic invoices across 50 layouts, reconciled field by field against ground truth. Computed live from the register on every page load.

documents processed
1,361

distinct documents with a completed reading

fields auto-accepted
61.2%

fields the gate cleared ÷ all extracted fields

documents a person saw
93.0%

documents with ≥1 held field ÷ all processed

mean cost / document
$0.000022

power-metered GPU energy × $0.11/kWh

mean reading time
8.7s

mean over all recorded readings

register events
22,841

append-only, hash-chained, verifiable

accuracy for all 8 fields

accuracy per field — matched ground-truth verdicts ÷ all measurable verdicts, weakest first, published.

invoice number
93.1% · n=1196
buyer
99.1% · n=845
total
99.6% · n=1089
subtotal
99.7% · n=923
date
99.7% · n=1333
tax
99.8% · n=519
supplier
99.9% · n=928
currency
100.0% · n=1251
The invoice number is the weakest field — on some templates the reader keeps the printed label glued to the number. Real behaviour, recorded, and the target of the calibration stage.

Does the grade predict correctness?

Accuracy split by routing — if grading works, fields that passed automatically must be right more often than the held ones. They are.

passed automatically (hallmarked)
99.9% · n=5552
held for a person
96.2% · n=2532

every other measurement — its own corpus, never blended

The scale run — 1,001 documents

One recorded batch, 2026-07-05: 1,001 documents on both cards in 76 minutes, zero failures; batch accuracy 98.77% (5,873/5,946). Full method: docs/METRICS.md in the repository.

RTX 3080 Ti
0documents / hour
8.7s per document · 0.2763 Wh · $0.0000304 energy each
RTX 4070
0documents / hour
7.8s per document · 0.1268 Wh · $0.0000139 energy each

Receipts — real photos, a second measured type

Real photographed receipts from CORD-v2 (Indonesian, CC-BY-4.0) — this makes “multilingual, real-photo” a measured claim, not a promise. The printed amounts and the menu lines are graded separately: reading a photographed menu line by line is honestly harder than the totals, so the two numbers are kept apart — never averaged into one, and never blended into the invoice headline above.

receipts read
100

distinct receipts with a completed reading

amount accuracy
97.2%

matched verdicts ÷ measurable, receipt amount fields (header) only

line-item accuracy
82.8%

matched verdicts ÷ measurable, receipt line cells — the real-photo line number

per-field accuracy + the review load

accuracy per amount field — matched verdicts ÷ measurable, header amounts only.

total
94.7% · n=94
tax
94.7% · n=38
subtotal
98.5% · n=66
cash paid
98.5% · n=66
change due
100.0% · n=63

and per line column — every menu line pulled from the receipt (description, quantity, unit price), each cell graded on its own, aligned to the printed menu order.

description
78.8% · n=255
unit price
83.1% · n=254
quantity
86.8% · n=228
fields auto-accepted
46.4%

fields the gate cleared ÷ all receipt fields

receipts a person saw
96.0%

receipts with ≥1 held field ÷ all receipts — the hard menu cells are held, fail-closed

mean reading time
13.8s

mean over all receipt readings

Line items — read row by row

Beyond the header totals, the reader pulls every line of the table — description, quantity, unit price, line total — and grades each cell on its own. Measured on a synthetic invoice set with line tables (its own slice, never blended into the invoice headline). Flawless computer-drawn renders — a mechanism proof end to end, not a real-photo number: the honest real number is the receipt lines above.

documents
40

synthetic invoices with a line table, read end to end

field accuracy
100.0%

matched ground-truth verdicts ÷ all measurable, this slice (header + line cells)

fields auto-accepted
90.0%

fields the gate cleared ÷ all fields on this slice

documents a person saw
95.0%

documents with ≥1 held field ÷ all — the sum-mismatch traps hold

accuracy per line column

accuracy per line column — every line’s cells aggregated by column, weakest first.

description
100.0% · n=234
quantity
100.0% · n=234
unit price
100.0% · n=234
line total
100.0% · n=234

Sorting the pile — by type

Before it reads a document, the machine can decide what kind it is. Measured on a labelled mixed pile of 28 synthetic documents: every type was sorted correctly, and — the point — nothing that isn’t a built type was let through. A type it doesn’t handle yet, or can’t recognise, is held for a person — never guessed, fail closed.

100.0%sorted to the right type (28/28) — 12 passed on to be read, 16 held for a person. Measured 2026-07-18.
every class, and where it routed
invoicen=6read today
receiptn=6read today
purchase ordern=6held — target not built yet — Wave 2
delivery noten=6held — target not built yet — Wave 2
unknownn=4held — not recognised — a person types the type

Three-way matching — invoice against order and delivery

Beyond reading one document, the machine links three — the invoice, the purchase order, and the delivery note — and checks every line: billed no more than delivered, delivered no more than ordered, price no higher than agreed. A mismatch is held for a person with a plain reason. Measured on 12 fabricated sets (12 invoices): every verdict correct.

5/12matched clean; 7 held with a per-line reason, 0 wrong. A mechanism proof on clean renders (real-photo line accuracy is the receipts corpus above). Measured 2026-07-18.
every set, and how it was judged
#0matchedorder, delivery and bill all agree
#1matchedorder, delivery and bill all agree
#2matchedorder, delivery and bill all agree
#3matchedorder, delivery and bill all agree
#4matchedorder, delivery and bill all agree
#5heldbilled for more than was delivered; billed for more than was ordered
#6heldbilled for more than was delivered; billed for more than was ordered
#7heldunit price above the agreed order price
#8helda billed line matches nothing ordered or delivered
#9heldno matching purchase order on file
#10heldno purchase-order number on the invoice
#11heldno delivery note to check against

Accuracy over runs

Held at scale across all 50 layouts, 20 of them never seen before the run.

60-document proving run · 2026-07-0398.3%
1,001-document scale run · 2026-07-0598.8%