docs: as-built reference for the statement OCR capture
Gives receipt capture the same treatment policy OCR just got: a doc that records what is in the code, separate from the spec that records what was designed. RECEIPT_CAPTURE_SPEC.md §2 had accumulated three BUILT notes totalling ~120 lines of findings, which is the right place for the evidence but the wrong place to look up how the matcher picks a column. docs/STATEMENT_OCR.md covers the pipeline, the OCR seam and its text-layer-first rule, all eight parsers and the ordering constraints between them, the matcher's two governing rules and the scopedRefField table, confirm-through-BillingService, the learning write-back, and the API surface. Weight goes to the things that are load-bearing and invisible from the code shape: brand detection must run to completion before layout because Tijuana bills predial and zona federal off the same treasury header; scopedRefField is exported because three call sites must agree or a reference gets learned into a column nothing searches; FEDERAL_ZONE's accountNumber holds a peso amount, so it fails the null-guards as well as the lookup; a misread `$` is the dangerous failure, not a missing one. Also records that CFE/CESPT/Telnor have no unit suite — they predate the gas/predial extension and were only verified end to end. Cross-linked from the spec, POLICY_OCR.md, PLAN.md, README and RESUME.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
+4
-3
@@ -233,9 +233,10 @@ signal at all (which must yield no provider rather than a bad guess).
|
||||
|
||||
## Related
|
||||
|
||||
- [`RECEIPT_CAPTURE_SPEC.md`](RECEIPT_CAPTURE_SPEC.md) §2 — the utility
|
||||
statement pipeline this was lifted from, and the origin of three rules the
|
||||
policy parser applies: detect the provider by brand before layout, only
|
||||
- [`STATEMENT_OCR.md`](STATEMENT_OCR.md) — the utility statement pipeline this
|
||||
was lifted from, as built. [`RECEIPT_CAPTURE_SPEC.md`](RECEIPT_CAPTURE_SPEC.md)
|
||||
§2 is its design and the measured evidence behind it. Between them they are
|
||||
the origin of three rules the policy parser applies: detect the provider by brand before layout, only
|
||||
ever apply the Tesseract digit-confusion map (`O→0`, `S→5`, `B→8`, …) to
|
||||
fields known to be digits, and parse amounts by separator *position* rather
|
||||
than assuming `,` is thousands.
|
||||
|
||||
@@ -131,7 +131,12 @@ single-movement form.
|
||||
|
||||
## 2. PDF / OCR auto-capture
|
||||
|
||||
> **BUILT — 2026-08-01.** Implemented and verified end to end against real
|
||||
> **BUILT — 2026-08-01.** As-built reference:
|
||||
> [`STATEMENT_OCR.md`](STATEMENT_OCR.md) — what the shipped feature does, its
|
||||
> parsers, matcher rules and API surface. This section stays the *design* and
|
||||
> the evidence behind it; go there for what is in the code today.
|
||||
>
|
||||
> Implemented and verified end to end against real
|
||||
> scanned statements. `apps/api/src/statements/` holds the module: a swappable
|
||||
> `OcrProvider` seam with a self-hosted Tesseract implementation, per-provider
|
||||
> parsers for CFE / CESPT / Telnor / gas / predial, a scoped matcher, and a
|
||||
|
||||
@@ -0,0 +1,315 @@
|
||||
# Utility Statement OCR Capture (receipt capture)
|
||||
|
||||
Reads the stack of scanned utility bills the office pays every month, proposes
|
||||
the customer and the amount for each page, and posts the confirmed pages to the
|
||||
ledger as one batch against one check. Built 2026-08-01, live under `/recibos`.
|
||||
|
||||
This is the **as-built** record. The design and the reasoning behind it are
|
||||
[`RECEIPT_CAPTURE_SPEC.md`](RECEIPT_CAPTURE_SPEC.md) §2, which also carries the
|
||||
measured results and the four spec corrections the real scans forced. Read that
|
||||
for *why*; read this for *what is there*.
|
||||
|
||||
## The job it replaces
|
||||
|
||||
Staff receive 300+ pages per month per service provider — CFE electricity,
|
||||
CESPT water, Telnor phone, gas, municipal predial, federal zone — and key each
|
||||
one into the ledger by hand, as a charge against the customer whose property
|
||||
the bill belongs to. The whole stack is paid with one office check, so the
|
||||
capture is naturally a batch.
|
||||
|
||||
Auto-capture is a **mode of** the existing capture screen, not a separate
|
||||
feature: it is the same daily job with a scanner instead of a keyboard, and
|
||||
both modes post through the same ledger path.
|
||||
|
||||
## What ships
|
||||
|
||||
| Piece | Path |
|
||||
|---|---|
|
||||
| API module | `apps/api/src/statements/` (service, controller, DTOs, matcher, parsers) |
|
||||
| OCR seam | `apps/api/src/statements/ocr/` (interface + Tesseract), bound in `apps/api/src/ocr/ocr.module.ts` |
|
||||
| Tables | `statement_batches`, `statement_documents` (`20260731235721_statement_ocr_intake`) |
|
||||
| Web | `components/Captura.tsx` (tab shell), `StatementIntake.tsx` (upload + batch list), `/recibos/:id` (review queue) |
|
||||
| Abilities | `statement:ingest`, `statement:review` — both **STAFF** |
|
||||
|
||||
STAFF is deliberate: the review step is what makes machine capture safe at that
|
||||
tier, since nothing reaches the ledger unconfirmed.
|
||||
|
||||
## The screen
|
||||
|
||||
`Captura` is one screen with two ways in:
|
||||
|
||||
- `/estado-cuenta/lote` → **manual** tab (`ManualCheckCapture`, key each
|
||||
receipt against one check by hand)
|
||||
- `/recibos` → **automática (OCR)** tab (`StatementIntake`, upload the scans)
|
||||
- `/recibos/:id` → the batch review queue, page image beside the extracted
|
||||
fields
|
||||
|
||||
Both end in the same place — charges on customers' ledgers posted against one
|
||||
check — so they are modes of one screen. Staff pick by what is on the desk that
|
||||
morning. Either URL renders the same component, so old bookmarks land on the
|
||||
right tab.
|
||||
|
||||
## Pipeline
|
||||
|
||||
```
|
||||
upload PDFs (+ service kind) → store source → render pages → text layer? → parse → match → review → confirm → ledger
|
||||
```
|
||||
|
||||
**A batch is one service kind.** The uploader labels it (ELECTRIC, WATER,
|
||||
TELEPHONE, GAS, PROPERTY_TAX, FEDERAL_ZONE) and that label is enforced: if the
|
||||
parser reads a page as a different provider, the page is rejected as
|
||||
mis-sorted rather than matched. Posting a phone bill as a water charge is the
|
||||
failure being prevented.
|
||||
|
||||
**Processing is not awaited.** 300 pages of OCR is minutes of CPU, far past any
|
||||
HTTP timeout, so `POST /statements/batches` returns the batch id immediately and
|
||||
the client polls. That is also what lets the review queue show partial progress.
|
||||
|
||||
**One page = one document.** Statements arrive **bundled, one customer per
|
||||
page** — Telnor's own `Pág 3 de 6` is its internal pagination, not the office's
|
||||
scan — so every rendered page becomes its own `StatementDocument` and the
|
||||
parser runs per page. (The policy OCR feature inverts this; see "Sibling
|
||||
feature" below.)
|
||||
|
||||
**One unreadable page must not abandon the other 299.** A page that throws
|
||||
becomes a single `OCR_FAILED` row and the loop continues.
|
||||
|
||||
Both the source PDFs (`statement/{batchId}/source-N.pdf`) and every rendered
|
||||
page image (`page-N.png`) are stored. The source is the artifact the office
|
||||
received and the only way to re-run a corrected parser over the original; the
|
||||
page image is what the reviewer looks at, because "what the parser read" is
|
||||
only checkable against a picture of the paper.
|
||||
|
||||
## The OCR seam
|
||||
|
||||
`OcrProvider` (`ocr/ocr.provider.ts`) is the swap point. Four methods:
|
||||
`available()`, `renderPages()`, `recognize()`, `textPages()`.
|
||||
|
||||
Everything above it works in terms of page text and word boxes, so the engine
|
||||
is replaceable without touching the parsers, the matcher or the schema. The
|
||||
shipped implementation is **self-hosted Tesseract**, and that choice is
|
||||
evidence-based rather than assumed — see the spec's measured results. A managed
|
||||
API (Textract, Document Intelligence, Document AI) fits behind the same
|
||||
interface with no schema change; at 300+ pages/month/company it would carry
|
||||
real recurring cost for accuracy that is not the bottleneck.
|
||||
|
||||
The binding lives in `apps/api/src/ocr/ocr.module.ts`, extracted out of
|
||||
`StatementsModule` when [`POLICY_OCR.md`](POLICY_OCR.md) needed the same seam.
|
||||
`StatementsModule` imports it and binds nothing itself, so the engine decision
|
||||
is one line in one file for both features.
|
||||
|
||||
### Text layer first, OCR as the fallback
|
||||
|
||||
**Not every statement is a scan.** The gas company sends born-digital CFDI
|
||||
invoices whose text layer is already exact and already positioned.
|
||||
`textPages()` reads it (`pdftotext -bbox-layout`, same poppler package as
|
||||
`pdftoppm`) and OCR runs only where there is none.
|
||||
|
||||
Rasterising a born-digital page and re-recognising it can only lose
|
||||
information — one sample turned `MEDIDOR: VM01014426` into
|
||||
`ar (LTR): 014420` — while costing about a minute of CPU for the privilege.
|
||||
Positions come back in the same pixel space `recognize()` uses, so the parsers'
|
||||
geometric helpers work unchanged on either source. When the text layer is used
|
||||
the document's notes say so verbatim: *"texto leído del PDF original, sin
|
||||
OCR"*.
|
||||
|
||||
### Word boxes, not just text
|
||||
|
||||
`OcrPage` carries `words[]` with pixel boxes because several real layouts are
|
||||
**tables**: the CESPT "RECIBO" prints `No. DE CUENTA` as a column header with
|
||||
the value in the row beneath it, which line-oriented text cannot associate.
|
||||
Parsers fall back to geometry for exactly those fields.
|
||||
|
||||
## Parsers
|
||||
|
||||
Eight providers, dispatched by a `BRAND` table checked before a `LAYOUT` table:
|
||||
|
||||
| Provider | Service kind |
|
||||
|---|---|
|
||||
| `CFE` | ELECTRIC |
|
||||
| `CESPT` | WATER |
|
||||
| `TELNOR` | TELEPHONE |
|
||||
| `GAS TIJUANA` | GAS |
|
||||
| `PREDIAL TIJUANA` / `PREDIAL ROSARITO` / `PREDIAL ENSENADA` | PROPERTY_TAX |
|
||||
| `ZONA FEDERAL TIJUANA` | FEDERAL_ZONE |
|
||||
|
||||
Three predial parsers rather than one because Tijuana, Rosarito and Ensenada
|
||||
issue three completely different documents — same tax, nothing else in common.
|
||||
|
||||
Rules that are load-bearing and easy to break:
|
||||
|
||||
- **Brand before layout, and never interleaved.** Scanned logos read badly (a
|
||||
CESPT header came back as `E BAJA ES PAGO / EALIFORNIA`), which is why the
|
||||
layout fallback exists — but *every* brand rule runs first, because a Telnor
|
||||
page contains words a CFE structural rule would otherwise claim.
|
||||
- **Tijuana bills predial and zona federal from the same treasury.** Same
|
||||
header, same address, same `ATB-541201` RFC, so every predial discriminator
|
||||
matches a zona federal page too. The words only that layout prints are
|
||||
`Marítimo Terrestre`, so its rule is asked ahead of all three predial ones.
|
||||
**Order matters here in a way that is invisible from the code shape.**
|
||||
- **Parse amounts by separator position.** A real Telnor bill OCR'd as
|
||||
`$ 649,00`; stripping commas as thousands separators makes that $64,900.
|
||||
- **A misread `$` is the dangerous failure, not a missing one.** An Ensenada
|
||||
receipt for `$2,203.00` OCR'd as `82,203.00` — the sign read as an 8, which
|
||||
would post a charge 37× too large and look entirely ordinary in the ledger.
|
||||
Every predial amount therefore requires a literal `$`; a page that cannot
|
||||
produce one reports no amount and goes to review.
|
||||
- **Digit-confusion repair only on fields known to be digits** (`O→0`, `S→5`,
|
||||
`B→8`, …), never on free text.
|
||||
- **The clave catastral is not `[A-Z]{2}[0-9]{6}`.** Position three is a letter
|
||||
in fifteen of the 932 stored claves (`MMB01041`, `CGH52121`). Digitising the
|
||||
whole tail maps that `B` to an `8` and yields a key matching no property.
|
||||
- **Barcodes beat printed labels.** Where a provider prints a payment barcode
|
||||
it is preferred and the two are cross-checked; disagreement sets
|
||||
`crossChecked: false` and forces review, because which of the two was
|
||||
misread is a judgement call.
|
||||
|
||||
## Matching
|
||||
|
||||
`StatementMatcherService`. Two rules govern everything:
|
||||
|
||||
**Match on one scoped field, never fuzzily across all identifiers.** Each
|
||||
service kind has exactly one column its statements print, and only that column
|
||||
is consulted. A blanket search over accountNumber/meterNumber/route would let a
|
||||
water account number collide with an unrelated phone number, and the mis-post
|
||||
would look perfectly ordinary in the ledger.
|
||||
|
||||
**Never match on the customer name.** A CESPT receipt for account `5365218`
|
||||
prints `ARNAIZ ROSAS ELSA AURORA`; the office's book, corroborated by the
|
||||
clave, has `CATT, RANDY`. The name on a utility bill is the registrant, not the
|
||||
current owner. Names are shown to the reviewer and are never an input.
|
||||
|
||||
### `scopedRefField` — which column each kind actually prints
|
||||
|
||||
| Kind | Column | Why |
|
||||
|---|---|---|
|
||||
| ELECTRIC, WATER, TELEPHONE, CABLE | `accountNumber` | the legacy column holds the printed number |
|
||||
| GAS | `meterNumber` | the number lived in free-text notes; `accountNumber` was never populated |
|
||||
| PROPERTY_TAX | `meterNumber` | `accountNumber` holds `DATMEX.predial`, an office file number that is neither unique nor printed anywhere |
|
||||
| FEDERAL_ZONE | `meterNumber` | `accountNumber` holds `DATMEX.zfed`, which is a **peso amount**, not a reference |
|
||||
|
||||
`scopedRefField` is exported because three places must agree on the answer: the
|
||||
lookup, the blank-service fill on review, and the write-back on confirm. When
|
||||
they disagree a reference gets learned into a column nothing searches, and the
|
||||
same page returns to the review queue every month forever.
|
||||
|
||||
The `FEDERAL_ZONE` case is the sharpest instance of a trap this codebase hits
|
||||
repeatedly (see also `policies.total`): a legacy column whose *name* promises
|
||||
an identifier and whose *contents* are something else. Three of its 77 values
|
||||
carry cents and one is negative. Worse than never matching — because every row
|
||||
already has a value, the `[field]: null` guards on learning and on the
|
||||
blank-service fill would never fire either.
|
||||
|
||||
### The clave catastral is a rescue on some layouts and the primary key on others
|
||||
|
||||
CESPT bills print the clave as well as an account number, so it rescues a page
|
||||
whose account number did not OCR — which happened on real samples. There it
|
||||
stays a hint.
|
||||
|
||||
On Rosarito and Ensenada predial the receipt prints **nothing else**, so a
|
||||
unique clave hit is a real match and auto-matches. Tijuana predial prints no
|
||||
clave at all; its only identifier is an 8-digit municipal account carried in a
|
||||
32-digit payment barcode (`account(8) + DDMMYY + amount(9) + folio(9)`) that
|
||||
the legacy database never held, so those pages start cold and are taught by the
|
||||
first confirm.
|
||||
|
||||
Multiple hits are always surfaced, never auto-picked — duplicate account
|
||||
numbers do occur in the legacy data, and the office's own `DUPLICADOS` report
|
||||
existed for a reason.
|
||||
|
||||
## Confirm: what gets written
|
||||
|
||||
`confirmBatch` posts through **`BillingService.createBatch`** — the same method
|
||||
the manual Editor screen uses — rather than writing `Transaction` rows
|
||||
directly, so OCR-sourced and hand-keyed receipts share one write path, one
|
||||
validation path and one audit trail.
|
||||
|
||||
- `source: "OCR"` and a per-line `captureRef` of the document id feed the
|
||||
duplicate-post guard, so a batch confirmed twice cannot double-charge.
|
||||
- `items[i]` is positionally parallel to `lines[i]` (a documented seam
|
||||
guarantee), so the created rows zip straight back onto the documents that
|
||||
produced them via `postedTransactionId`.
|
||||
- **The sign is applied here.** Charges are negative in this ledger; the parser
|
||||
reads the printed positive figure, and `-Math.abs()` is applied at the single
|
||||
point where a statement becomes a ledger row.
|
||||
- A missing amount blocks the confirm with the offending page numbers, rather
|
||||
than silently posting zero.
|
||||
|
||||
### Learning: the cold start is a one-time cost
|
||||
|
||||
After posting, `learnAccountRefs` writes each confirmed reference back onto the
|
||||
`PropertyService` that matched — **only where the field was null**. Never
|
||||
overwrites a number already on file, which would let one misread page rewrite
|
||||
good reference data.
|
||||
|
||||
This is what turns gas (whose numbers the migration never populated) and
|
||||
Tijuana predial (whose municipal account the legacy database never held) from a
|
||||
permanent review queue into a one-time cost: next month's statement for the
|
||||
same account matches on its own.
|
||||
|
||||
### Discarding
|
||||
|
||||
Refused once any page is `POSTED` — those pages already wrote ledger rows
|
||||
against a check, and a "discarded" label on the batch would leave the charges
|
||||
unexplained. Reject the remaining pages individually instead.
|
||||
|
||||
## API surface
|
||||
|
||||
| Method | Route | Ability |
|
||||
|---|---|---|
|
||||
| `GET` | `/statements/status` (is OCR + storage available) | authenticated |
|
||||
| `GET` | `/statements/batches`, `/batches/:id`, `/batches/:id/documents` | authenticated |
|
||||
| `GET` | `/statements/documents/:id/page` (streams the page image) | authenticated |
|
||||
| `POST` | `/statements/batches` (upload + service kind) | `statement:ingest` |
|
||||
| `PATCH` | `/statements/documents/:id` (correct a field or the match) | `statement:review` |
|
||||
| `POST` | `/statements/documents/:id/reject` | `statement:review` |
|
||||
| `POST` | `/statements/batches/:id/discard` | `statement:review` |
|
||||
| `POST` | `/statements/batches/:id/confirm` | `statement:review` |
|
||||
|
||||
## Requirements
|
||||
|
||||
Object storage (`S3_ENDPOINT` + credentials) for the scans, and `tesseract-ocr`
|
||||
/ `tesseract-ocr-data-spa` / `poppler-utils` in the API image. Both are checked
|
||||
at upload rather than at the first write — a missing dependency should be a 400
|
||||
on the request, not a `FAILED` batch minutes later. `GET /statements/status`
|
||||
reports both and the upload card hides itself unless both hold.
|
||||
|
||||
## Tests
|
||||
|
||||
- `parsers/statement-parser.spec.ts` — `detectProvider`,
|
||||
`normalizeCadastralKey`, and one suite per newer parser
|
||||
(`parsePredialTijuana`, `parsePredialRosarito`, `parsePredialEnsenada`,
|
||||
`parseGas`, `parseZonaFederal`). Every fixture is a **verbatim OCR excerpt
|
||||
from a real receipt**, including the ones that bit: the `82,203.00` Ensenada
|
||||
dollar sign, the three-letter clave, the CESPT logo garbage.
|
||||
- `ocr/tesseract.provider.spec.ts` — `parseBboxLayout`, the
|
||||
`pdftotext -bbox-layout` reader that produces the text-layer `OcrPage`
|
||||
(positions included, which is what lets the geometric helpers work on
|
||||
born-digital input).
|
||||
|
||||
Note the CFE / CESPT / Telnor parsers themselves have **no unit suite** — they
|
||||
predate the gas/predial extension and were verified against the 46-page corpus
|
||||
end to end rather than in isolation. Worth closing if they are touched.
|
||||
|
||||
## Not built
|
||||
|
||||
- **Handwritten folder numbers.** Staff pencil a customer number on each bill
|
||||
(`9`, `405`); Tesseract read `405` as `205`. Handwriting is a review hint at
|
||||
best and is deliberately not an input to matching.
|
||||
- **Re-running a corrected parser over a stored batch.** The source PDFs are
|
||||
kept precisely so this is possible, but nothing exposes it yet.
|
||||
- **Providers beyond the eight above.** Adding one is a `BRAND` entry, an
|
||||
optional `LAYOUT` entry, and a parser function.
|
||||
|
||||
## Sibling feature
|
||||
|
||||
[`POLICY_OCR.md`](POLICY_OCR.md) — the same pipeline reading carrier policy
|
||||
PDFs into `Policy` rows, built out of this one. It reuses the seam, the
|
||||
provider-detection ordering, the digit-confusion map and the
|
||||
amount-by-separator rule.
|
||||
|
||||
**One assumption does not carry over.** Here a page *is* a document, because
|
||||
statements arrive one customer per page. A policy PDF is one document across
|
||||
several pages, so that feature concatenates the pages and parses once per file.
|
||||
If you are porting a change between the two, that is the difference to check
|
||||
first.
|
||||
Reference in New Issue
Block a user