Proven on bad scans, not clean PDFs
We took one real invoice and made it bad. Blur. Skew. Faded toner. Low-res. Thirteen ways. Then we ran each one through the real pipeline. All thirteen came back with every field exact.
degraded scan variants recovered every ground-truth field exact. Blur, skew to 5°, grain, faded toner, uneven light, 72dpi.
wrong numbers ever produced. When readability drops, the extractor returns nothing rather than guessing, and the validator routes it to review.
Fields exact out of 14 ground-truth fields, per variant.
| Scan condition | Fields exact | Verdict |
|---|---|---|
| Clean scan, 300dpi | 14/14 | AUTO-EXPORT |
| Low resolution, 150dpi | 14/14 | AUTO-EXPORT |
| Low resolution, 100dpi | 14/14 | AUTO-EXPORT |
| Low resolution, 72dpi | 14/14 | AUTO-EXPORT |
| Skewed 1.5° / 3° / 5° | 14/14 ×3 | AUTO-EXPORT |
| Gaussian noise (σ28) | 14/14 | AUTO-EXPORT |
| Salt-and-pepper noise | 14/14 | AUTO-EXPORT |
| Faded toner (0.42×) | 14/14 | AUTO-EXPORT |
| Uneven lighting | 14/14 | AUTO-EXPORT |
| Heavy blur | 14/14 | AUTO-EXPORT |
| Worst combo (100dpi + skew + noise + fade) | 14/14 | AUTO-EXPORT |
13 of 13 variants scored 14/14 against ground truth, all AUTO-EXPORTED. Sweep re-run 2026-09-01 on the production model (gemini-3.5-flash-lite). When readability does drop, the extractor returns nothing rather than guessing, and the validator routes line-incomplete invoices to review. Raw data: docs/research/ACCURACY.md.
How the sweep was run
No curated demo set. The fixture invoice was degraded deliberately, then scored field by field against ground truth.
Degrade on purpose
A script renders the fixture invoice at varying quality: resolution from 300dpi down to 72dpi, skew up to 5 degrees, Gaussian and salt-and-pepper noise, contrast faded to 0.42×, uneven lighting, heavy blur, and a worst-case combo of all of them.
Run the live pipeline
Every variant goes through the real path, not a mock: render, vision-model read, defensive normalization into the canonical invoice contract, then the full validator set.
Score field by field
Fourteen ground-truth fields per invoice, from seller GSTIN to grand total. Omission mode is the design contract: when readability drops, the extractor returns nothing rather than guessing, and incomplete invoices are routed to review instead of exported.
Watch the clock
Median latency is about 2.3 seconds a page end-to-end on the production model (gemini-3.5-flash-lite), and about 3.3 seconds across the degrade sweep. Processing is asynchronous, so volume runs as a background job rather than a stopwatch.