the assay report
Accuracy and cost, measured
the headline corpus — synthetic invoices, measured live below
real photographed receipts — amounts and menu lines graded separately
measured GPU board energy while it read
the headline corpus — measured live
Invoices — the live headline
Synthetic invoices across 50 layouts, reconciled field by field against ground truth. Computed live from the register on every page load.
- documents processed
- 1,361
- fields auto-accepted
- 61.2%
- documents a person saw
- 93.0%
- mean cost / document
- $0.000022
- mean reading time
- 8.7s
- register events
- 22,841
distinct documents with a completed reading
fields the gate cleared ÷ all extracted fields
documents with ≥1 held field ÷ all processed
power-metered GPU energy × $0.11/kWh
mean over all recorded readings
append-only, hash-chained, verifiable
accuracy for all 8 fields ↓close the detail ↑
accuracy per field — matched ground-truth verdicts ÷ all measurable verdicts, weakest first, published.
Does the grade predict correctness?
Accuracy split by routing — if grading works, fields that passed automatically must be right more often than the held ones. They are.
every other measurement — its own corpus, never blended
The scale run — 1,001 documents
One recorded batch, 2026-07-05: 1,001 documents on both cards in 76 minutes, zero failures; batch accuracy 98.77% (5,873/5,946). Full method: docs/METRICS.md in the repository.
Receipts — real photos, a second measured type
Real photographed receipts from CORD-v2 (Indonesian, CC-BY-4.0) — this makes “multilingual, real-photo” a measured claim, not a promise. The printed amounts and the menu lines are graded separately: reading a photographed menu line by line is honestly harder than the totals, so the two numbers are kept apart — never averaged into one, and never blended into the invoice headline above.
- receipts read
- 100
- amount accuracy
- 97.2%
- line-item accuracy
- 82.8%
distinct receipts with a completed reading
matched verdicts ÷ measurable, receipt amount fields (header) only
matched verdicts ÷ measurable, receipt line cells — the real-photo line number
per-field accuracy + the review load ↓close the detail ↑
accuracy per amount field — matched verdicts ÷ measurable, header amounts only.
and per line column — every menu line pulled from the receipt (description, quantity, unit price), each cell graded on its own, aligned to the printed menu order.
- fields auto-accepted
- 46.4%
- receipts a person saw
- 96.0%
- mean reading time
- 13.8s
fields the gate cleared ÷ all receipt fields
receipts with ≥1 held field ÷ all receipts — the hard menu cells are held, fail-closed
mean over all receipt readings
Line items — read row by row
Beyond the header totals, the reader pulls every line of the table — description, quantity, unit price, line total — and grades each cell on its own. Measured on a synthetic invoice set with line tables (its own slice, never blended into the invoice headline). Flawless computer-drawn renders — a mechanism proof end to end, not a real-photo number: the honest real number is the receipt lines above.
- documents
- 40
- field accuracy
- 100.0%
- fields auto-accepted
- 90.0%
- documents a person saw
- 95.0%
synthetic invoices with a line table, read end to end
matched ground-truth verdicts ÷ all measurable, this slice (header + line cells)
fields the gate cleared ÷ all fields on this slice
documents with ≥1 held field ÷ all — the sum-mismatch traps hold
accuracy per line column ↓close the detail ↑
accuracy per line column — every line’s cells aggregated by column, weakest first.
Sorting the pile — by type
Before it reads a document, the machine can decide what kind it is. Measured on a labelled mixed pile of 28 synthetic documents: every type was sorted correctly, and — the point — nothing that isn’t a built type was let through. A type it doesn’t handle yet, or can’t recognise, is held for a person — never guessed, fail closed.
every class, and where it routed ↓close the detail ↑
Three-way matching — invoice against order and delivery
Beyond reading one document, the machine links three — the invoice, the purchase order, and the delivery note — and checks every line: billed no more than delivered, delivered no more than ordered, price no higher than agreed. A mismatch is held for a person with a plain reason. Measured on 12 fabricated sets (12 invoices): every verdict correct.
every set, and how it was judged ↓close the detail ↑
Accuracy over runs
Held at scale across all 50 layouts, 20 of them never seen before the run.