The demo is seductive. Python script. Tesseract or a VLM. Clean PDF. GSTIN extracted. Total extracted. Looks like a product. The demo takes a weekend. The product takes six months. Here is what you would actually build.
OCR engine: Tesseract, RapidOCR, or a VLM. Each has strengths and weaknesses on Indian photocopies. Tesseract reads clean type well and fails on 150 dpi phone photos. VLMs read bad paper better but hallucinate more. You need to pick one, tune it, and live with its failure modes.
GSTIN extraction: regex for the 15-character format. Checksum validation: 7 lines of Python. See GSTIN checksum. This part is open. You can copy it. But you still need to run it on every extracted GSTIN, and you still need to decide what happens when it fails.
HSN extraction: regex for digit lengths. Directory membership: a bundled list. See HSN validation. The list is MIT-licensed. You can copy it. But the list lags. You still need to decide shadow vs live.
Arithmetic closure: lines sum to taxable, taxable times rate equals tax, taxes plus taxable equals total. Plus/minus 1 rupee tolerance. See totals. This is the hard part. Not because the math is hard. Because round-off rows, mixed rates, and Indian grouping make it hard in practice.
OCR + regex
Clean PDF. GSTIN extracted. Total extracted. Looks like a product.
OCR + gate + review + Tally mapping
Bad paper, failed rows, ledger names, batch logic, hosting, monitoring.
The parts nobody demos
Review UI: when the gate fails, a person needs to see the crop and the reason. A table of naked strings is not a review UI. Keyboard-first, 1366x768, evidence beside the field. See NEEDS_REVIEW.
Tally mapping: 200 suppliers, 200 ledger strings. Mapping built once. Reused every batch. See vendor mapping. This is the work nobody demos because it is not sexy. It is the work that makes the product stick.
Hosting: India-hosted, single tenant. See hosting. Data never trains shared models. That is a privacy commitment, not a feature list.
Monitoring: what broke, when, why. Pipeline alerts. Cost tracking. Batch reports. The infrastructure that keeps the product running when nobody is watching.
What you can verify
The validators are auditable. GSTIN checksum, HSN format, arithmetic, regime. Run them yourself on the validation page. But validators without extraction are a library, not a product. Extraction is the hard part. That stays ours.
When to build
You have a platform team. You have many doc types beyond GST invoices. You need custom workflows. You have the budget for six months of engineering. Build.
You have one doc type: GST photocopies into Tally. You have no platform team. You need it working in weeks. Buy the gate. The price is ₹1.40 a page. The build price is six months of engineering time plus hosting plus monitoring plus the review UI you will never demo but always need.
The honest math
40,000 pages a month at ₹1.40 is ₹56,000 a month. That is ₹6,72,000 a year. A senior ML engineer in India costs more than that. Two engineers cost more. Hosting costs more. The build is more expensive unless you have scale beyond 40,000 pages or doc types beyond GST invoices.
I will not invent a salary for a US ML team. Indian ML engineers cost what they cost in your city. Plug your numbers. The gate may still be cheaper. Or it may not. The answer depends on your scale.
- EntryLedgervalidation page
Run the checks on a real fixture. Validators auditable, extraction ours.
- EntryLedger/pricing
₹1.40/page. Published. Compare to your build cost.
Is EntryLedger open source?+
No. The validators are auditable on the validation page, but the code stays ours. The product is the gate, not the OCR.
Should I build?+
If you have a platform team and many doc types, maybe. If GST photocopies into Tally, the gate is cheaper.
What would I build?+
OCR, GSTIN regex, checksum, HSN, arithmetic, review UI, Tally mapping, hosting, monitoring. Six months minimum.