Your supplier sends the same invoice twice. Once as a photocopy, once as a phone photo. The first one has GSTIN 27AAPFU1234A1ZF. The second has 27AAPFU1234A1ZE. One digit is different because the second scan is blurrier. Both have the same invoice number. Both get posted. You pay twice.
This is not a hypothetical. It happens in every AP department that processes scanned invoices. The duplicate is not always identical. Sometimes the GSTIN is slightly different. Sometimes the total is rounded differently. Sometimes the date is typed as 15/01 instead of 01/15. The duplicate is there, but it does not look like a duplicate.
Three kinds of duplicates
| Type | What it looks like | How to catch it |
|---|---|---|
| Exact duplicate | Same invoice number, same GSTIN, same amount | Simple uniqueness check |
| Near-duplicate | Same invoice number, different GSTIN (typo in second scan) | Invoice number match + manual review |
| Amount split | Same supplier, same date, amounts that sum to the original | Supplier + date + amount pattern |
The exact duplicate is easy. A simple uniqueness check catches it: same invoice number, same GSTIN, same amount. The system flags it and a person confirms.
The near-duplicate is harder. The invoice number is the same, but the GSTIN is different. This happens when the same invoice is scanned twice, and the second scan is blurrier. The OCR reads a different digit in the GSTIN. The result is two rows with the same invoice number but different GSTINs. A simple uniqueness check on invoice number alone would catch this, but it would also flag legitimate cases where two different invoices have the same number (which happens with some suppliers who use sequential numbering that resets annually).
The amount split is the rarest. The supplier sends two invoices that together cover the original amount. Maybe the original was for ₹100,000 and the supplier sent two invoices for ₹60,000 and ₹40,000. Or maybe the supplier corrected a mistake and sent a credit note. These are not true duplicates, but they need to be investigated.
How EntryLedger catches them
EntryLedger uses conditional uniqueness. The database constraint is: the same invoice number plus the same seller GSTIN is a duplicate. But if the GSTIN is empty (unidentified invoice), multiple rows can coexist. This handles both the exact duplicate and the near-duplicate where the GSTIN was mistyped in one scan.
The conditional part is important. Not every invoice has a clear GSTIN. Some scans are too blurry to read the GSTIN. Some invoices from small suppliers do not carry a GSTIN at all. If the uniqueness check required a GSTIN, these invoices would be blocked even when they are legitimate. By making the constraint conditional (GSTIN must match only if both are non-empty), EntryLedger handles the real-world messiness of scanned invoices.
When a duplicate is found, the invoice goes to review. The scan is there. The reason is there. A person confirms whether it is a true duplicate or a different invoice with the same number. The confirmation takes seconds. The duplicate never reaches Tally.
Why this matters
Duplicate invoices cost money. If you pay the same invoice twice, the cost is the full invoice amount. If you post the same invoice twice to Tally, your books are wrong. The reconciliation shows a discrepancy. Someone spends time tracing it. The time costs money.
The cost is not just financial. Duplicate entries distort your AP aging. They inflate your expenses. They make your financial statements inaccurate. At audit time, duplicate entries raise questions. The questions take time to answer. The answers cost money.
Catching duplicates at entry time is cheap. Catching them at audit time is expensive. The difference is a few seconds of validation versus hours of investigation.
How EntryLedger catches them
EntryLedger uses conditional uniqueness. The database constraint is: the same invoice number plus the same seller GSTIN is a duplicate. But if the GSTIN is empty (unidentified invoice), multiple rows can coexist. This handles both the exact duplicate and the near-duplicate where the GSTIN was mistyped in one scan.
The conditional part is important. Not every invoice has a clear GSTIN. Some scans are too blurry to read the GSTIN. Some invoices from small suppliers do not carry a GSTIN at all. If the uniqueness check required a GSTIN, these invoices would be blocked even when they are legitimate. By making the constraint conditional (GSTIN must match only if both are non-empty), EntryLedger handles the real-world messiness of scanned invoices.
When a duplicate is found, the invoice goes to review. The scan is there. The reason is there. A person confirms whether it is a true duplicate or a different invoice with the same number. The confirmation takes seconds. The duplicate never reaches Tally.
What this looks like in practice
An invoice comes in as a photocopy. The invoice number is extracted. The GSTIN is extracted. EntryLedger checks: is there already an invoice with this number and this GSTIN in the database? If yes, the invoice goes to review with a duplicate flag. The person reviewing sees the original invoice and the new one side by side. They confirm it is a duplicate and discard it. Or they confirm it is a different invoice and approve it.
The process takes seconds. The duplicate is caught at entry time, before it reaches Tally. The books stay clean. The reconciliation stays clean. The audit stays clean.
The three kinds of duplicates
The exact duplicate is the easiest to catch. Same invoice number, same GSTIN, same amount. The conditional uniqueness check catches it immediately.
The near-duplicate is harder. Same invoice number, different GSTIN. This happens when the same invoice is scanned twice and the second scan is blurrier. The OCR reads a different digit in the GSTIN. The result is two rows with the same invoice number but different GSTINs. The conditional uniqueness check allows this (because the GSTINs are different), but the invoice number match flags it for review.
The amount split is the rarest. The supplier sends two invoices that together cover the original amount. These are not true duplicates, but they need to be investigated. The system does not catch these automatically. A person needs to review the pattern and decide.
What Tally will and will not save you from
Tally can reject a duplicate voucher number in some setups. It will not help if the second scan got a different invoice number because OCR read O as 0. It will not help if the clerk “fixed” a failed import by changing the number. Detection has to happen on the scan identity (number + seller GSTIN) before export, not after a human has already improvised.
Empty GSTIN is a special case. Unidentified invoices must be allowed to sit in the queue without crashing uniqueness. If you make GSTIN mandatory in the database, you either invent GSTINs (criminal) or you cannot ingest the scan. Conditional uniqueness is the boring engineering that matches paper. Block when both keys are present and equal. Allow when a key is still missing. Review the missing ones.
Payment runs are where duplicates become cash. AP that posts twice and pays twice is a recovery project. AP that flags the second scan in review is a click. Build the click. Do not wait for the bank statement.
If your volume is small and one person remembers every invoice, you may not need software for this. If two clerks share a pile and email-in is on, you do. Memory does not scale to 40,000 pages. Constraints do.
- EntryLedgerentryledger.shop
Duplicate handling: same invoice number and seller GSTIN, when both are present. Empty GSTIN rows are not forced unique.
How does duplicate detection work?+
Conditional uniqueness: same invoice number + seller GSTIN = duplicate. Empty GSTINs coexist.
What happens when a duplicate is found?+
The invoice goes to review with a duplicate flag. A person confirms whether it is true or a different invoice.