Accuracy

Proven on bad scans, not clean PDFs

We took one real invoice and made it bad. Blur. Skew. Faded toner. Low-res. Thirteen ways. Then we ran each one through the real pipeline. All thirteen came back with every field exact.

13/13

degraded scan variants recovered every ground-truth field exact. Blur, skew to 5°, grain, faded toner, uneven light, 72dpi.

0

wrong numbers ever produced. When readability drops, the extractor returns nothing rather than guessing, and the validator routes it to review.

Clean scan, 300dpi14/14
Low resolution, 150dpi14/14
Low resolution, 100dpi14/14
Low resolution, 72dpi14/14
Skewed 1.5°14/14
Skewed 3°14/14
Skewed 5°14/14
Gaussian noise (σ28)14/14
Salt-and-pepper noise14/14
Faded toner (0.42×)14/14
Uneven lighting14/14
Heavy blur14/14
Worst combo (100dpi + skew + noise + fade)14/14

Fields exact out of 14 ground-truth fields, per variant.

Accuracy across scan conditions
Scan conditionFields exactVerdict
Clean scan, 300dpi14/14AUTO-EXPORT
Low resolution, 150dpi14/14AUTO-EXPORT
Low resolution, 100dpi14/14AUTO-EXPORT
Low resolution, 72dpi14/14AUTO-EXPORT
Skewed 1.5° / 3° / 5°14/14 ×3AUTO-EXPORT
Gaussian noise (σ28)14/14AUTO-EXPORT
Salt-and-pepper noise14/14AUTO-EXPORT
Faded toner (0.42×)14/14AUTO-EXPORT
Uneven lighting14/14AUTO-EXPORT
Heavy blur14/14AUTO-EXPORT
Worst combo (100dpi + skew + noise + fade)14/14AUTO-EXPORT

13 of 13 variants scored 14/14 against ground truth, all AUTO-EXPORTED. Sweep re-run 2026-09-01 on the production model (gemini-3.5-flash-lite). When readability does drop, the extractor returns nothing rather than guessing, and the validator routes line-incomplete invoices to review. Raw data: docs/research/ACCURACY.md.

How the sweep was run

No curated demo set. The fixture invoice was degraded deliberately, then scored field by field against ground truth.

Degrade on purpose

A script renders the fixture invoice at varying quality: resolution from 300dpi down to 72dpi, skew up to 5 degrees, Gaussian and salt-and-pepper noise, contrast faded to 0.42×, uneven lighting, heavy blur, and a worst-case combo of all of them.

Run the live pipeline

Every variant goes through the real path, not a mock: render, vision-model read, defensive normalization into the canonical invoice contract, then the full validator set.

Score field by field

Fourteen ground-truth fields per invoice, from seller GSTIN to grand total. Omission mode is the design contract: when readability drops, the extractor returns nothing rather than guessing, and incomplete invoices are routed to review instead of exported.

Watch the clock

Median latency is about 2.3 seconds a page end-to-end on the production model (gemini-3.5-flash-lite), and about 3.3 seconds across the degrade sweep. Processing is asynchronous, so volume runs as a background job rather than a stopwatch.

See what the validators do when a field is incomplete →