docs: as-built reference for the statement OCR capture

Gives receipt capture the same treatment policy OCR just got: a doc that
records what is in the code, separate from the spec that records what was
designed. RECEIPT_CAPTURE_SPEC.md §2 had accumulated three BUILT notes
totalling ~120 lines of findings, which is the right place for the
evidence but the wrong place to look up how the matcher picks a column.

docs/STATEMENT_OCR.md covers the pipeline, the OCR seam and its
text-layer-first rule, all eight parsers and the ordering constraints
between them, the matcher's two governing rules and the scopedRefField
table, confirm-through-BillingService, the learning write-back, and the
API surface.

Weight goes to the things that are load-bearing and invisible from the
code shape: brand detection must run to completion before layout because
Tijuana bills predial and zona federal off the same treasury header;
scopedRefField is exported because three call sites must agree or a
reference gets learned into a column nothing searches; FEDERAL_ZONE's
accountNumber holds a peso amount, so it fails the null-guards as well
as the lookup; a misread `$` is the dangerous failure, not a missing one.

Also records that CFE/CESPT/Telnor have no unit suite — they predate the
gas/predial extension and were only verified end to end.

Cross-linked from the spec, POLICY_OCR.md, PLAN.md, README and RESUME.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-02 13:02:14 -07:00
co-authored by Claude Opus 5
parent 872a661051
commit ec139737be
6 changed files with 334 additions and 6 deletions
+3 -1
View File
@@ -134,7 +134,7 @@ Given the amount of near-duplicate/overlapping data across snapshot tables (mult
- **Receipt capture module — DONE** (2026-07-27). The legacy "Editor" replacement, built on the single-movement capture from step 6. Wires up the previously-unused `Transaction.outstanding` (NOPAGO): capture flag on `POST /billing`, `?outstanding=` list filter, `POST /billing/:id/resolve-outstanding` (gated `ledger:create`, not `ledger:void` — resolving *completes* a capture), and exclusion from every balance aggregate exactly as the legacy `SALDOS ULTIMO 0`'s `HAVING NOPAGO = 0` did. Adds `POST /billing/batch` (one `$transaction`, check-level fields shared, per-line customer/amount) and `GET /billing/by-check`, plus the `cheque-count` report replacing `REPORTE CHEQUE COUNT` / `REPORTE POR CHEQUE` / `EDITA CHEQUE ALF|COUNT|NUM` — print/PDF/CSV/XLSX come free from the existing `/reportes/:slug` machinery. Web: `/estado-cuenta/lote` (the actual "Editor" screen, with live reconciliation against the physical check amount), plus an "Estado de pago" filter, a "sin fondos" row tag and a Resolver dialog on `/estado-cuenta`. No new abilities. Verified end-to-end against dev, API + browser.
**Two pre-existing bugs found and fixed while building it:** (a) `statement()` filtered `legacySourceTable: { notIn: [...] }`, which compiles to SQL `NOT IN` — and `NULL NOT IN (…)` is NULL, so **every app-captured movement was invisible on the customer statement** (438 rows in the movement browser vs 392 on the statement) while still appearing everywhere else. This would have made the whole receipt-capture feature look broken to staff. Now NULL-safe. (b) The balances *count* query omitted the void filter its own page query applied, so the row count disagreed with the rows.
**OCR seam:** `BillingService.createBatch(dto, opts)` is the single multi-row write path and carries three contract guarantees for the step-11 OCR module to post through — `items[i]` maps to `lines[i]` (so `StatementDocument.postedTransactionId` can be zipped back on), `opts.refs[i]` stamps `captureRef` with a duplicate-post guard that a *voided* row deliberately does not block, and `opts.source` is service-level only so an HTTP client cannot label hand-keyed rows as machine-captured. Backed by a new `TransactionCaptureSource` enum (MANUAL/BATCH/OCR) + `captureRef`, both nullable so the 40,136 migrated rows stay NULL rather than being mislabelled.
- **PDF/OCR auto-capture — DONE** (2026-08-01). The ingest→split→OCR→match→review pipeline for the 300+/month/service-provider statements staff key in by hand, built in `apps/api/src/statements/` and posting through §1.2's `createBatch` seam with `source: "OCR"` and a per-document `captureRef`. Web: `/recibos` + `/recibos/:id`. Abilities `statement:ingest`/`statement:review` (STAFF — the review step is what makes machine capture safe at that tier). OCR is self-hosted **Tesseract** behind a swappable `OcrProvider` interface; `tesseract-ocr`, `tesseract-ocr-data-spa` and `poppler-utils` were added to the API image.
- **PDF/OCR auto-capture — DONE** (2026-08-01). As-built write-up in [`docs/STATEMENT_OCR.md`](docs/STATEMENT_OCR.md); the design and the measured evidence stay in the spec's §2. The ingest→split→OCR→match→review pipeline for the 300+/month/service-provider statements staff key in by hand, built in `apps/api/src/statements/` and posting through §1.2's `createBatch` seam with `source: "OCR"` and a per-document `captureRef`. Web: `/recibos` + `/recibos/:id`. Abilities `statement:ingest`/`statement:review` (STAFF — the review step is what makes machine capture safe at that tier). OCR is self-hosted **Tesseract** behind a swappable `OcrProvider` interface; `tesseract-ocr`, `tesseract-ocr-data-spa` and `poppler-utils` were added to the API image.
**Every decision was driven by 10 real scans (46 pages).** Shipped-parser results on them: provider 46/46, account ref 43/46, amount 42/46, due date 44/46 — and against the dev database **39/46 (85%) exact auto-match, 40/46 (87%) identified**, the rest genuine review cases. The scans are pure images (no text layer), so OCR is mandatory, and they arrive **bundled one customer per page**.
**The three gaps are closed, and two of them were mis-stated in the spec.** (a) `TELEPHONE` now exists and is backfilled from `Property.phone1` only — coverage is 534/18/1 across phone1/2/3, so phone is one billed line per property, not three. (b) **Clave catastral ≠ predial**: `DATMEX.clave` (934 rows, `KA903009`) is what CESPT and predial bills actually print, while `predial` — what `PROPERTY_TAX.accountNumber` holds — has only 663 distinct values across 1135 rows and appears on no statement; the clave now lives on `Property.cadastralKey` as the matcher's secondary key and predial is left untouched. (c) Gas was **not** a dead end: 160 of the 334 `DATMEX.gas` values are real account numbers (the rest are `ESTACIONARIO`/`CILINDRO` descriptors), all recovered into `GAS.meterNumber`.
**Matching is scoped per service kind and never reads the customer name** — a CESPT receipt prints `ARNAIZ ROSAS ELSA AURORA` for an account this office holds under `CATT, RANDY`, because the name on a utility bill is the registrant, not the current owner. Normalisation is per provider: CFE strips leading zeros off `NO. DE SERVICIO`, Telnor strips the 664 LADA down to the stored local 7 digits. Where a provider prints a payment barcode it is preferred over the printed label (one CFE label OCR'd a digit too many while its barcode was correct) and the two are cross-checked, with disagreement forcing review. Confirming a document whose service had no reference writes it back, so gas and any other cold start is a one-time cost.
@@ -168,6 +168,8 @@ Repo scaffolded at `jorgecuadros-platform/`: npm workspaces, NestJS API with a r
**Step 11 is now three-quarters built.** Receipt capture, the multi-bank chequera and PDF/OCR auto-capture are all done and verified; only customer-number recycling remains unbuilt. `docs/RECEIPT_CAPTURE_SPEC.md` carries a BUILT note per section recording what shipped and, for §2, the four things real scanned statements proved the spec had wrong or unknown.
Each of the two OCR intakes now has an as-built doc separate from its spec — `docs/STATEMENT_OCR.md` and `docs/POLICY_OCR.md`. The specs record what was designed and why; those record what is in the code. They share one `OcrProvider` seam (`apps/api/src/ocr/`), so the Tesseract-vs-managed-API decision is one line for both.
**It also produced a feature nobody planned.** The statement OCR pipeline generalised: the same render→OCR→parse→match→review shape reads **carrier policy PDFs** into `Policy` rows, which is `docs/POLICY_OCR.md` (built 2026-08-01, GMX only so far). It belongs to step 12's subject matter but to step 11's lineage, and it is in no spec — worth knowing before reading `INSURANCE_FEATURES_SPEC.md`, which does not mention it. It also partly overlaps what §4's carrier API was wanted for, and unlike that section it is not blocked on a phone call.
**Step 12 is one-quarter built.** `docs/INSURANCE_FEATURES_SPEC.md` covers the insurance half of the same meeting (renewal emails, liquidación batch, certificate + portal delivery, carrier APIs) — see Build sequencing step 12 above. Verified the same way, plus a live query of the dev DB for the counts it quotes (email coverage, pending liquidación, installment fill rates) and of the staged Parquet for the legacy settlement-slot usage. **§1 renewal emails is done** (2026-08-01/02) and carries a BUILT note recording the three places the build diverged from the spec; §2 liquidación is still the smallest remaining piece, since the per-policy fields are already wired end to end.
+1 -1
View File
@@ -50,7 +50,7 @@ renewal avisos; `/renovaciones` is an alias onto its Pólizas tab), `/reportes`,
`/catalogos`, `/operaciones` (DB ingest/backup, ADMIN), `/usuarios`, `/login`.
Two OCR intakes share one `OcrProvider` seam (`src/ocr/`, Tesseract today):
utility statements → ledger rows ([`docs/RECEIPT_CAPTURE_SPEC.md`](docs/RECEIPT_CAPTURE_SPEC.md) §2)
utility statements → ledger rows ([`docs/STATEMENT_OCR.md`](docs/STATEMENT_OCR.md))
and carrier policy PDFs → `Policy` rows ([`docs/POLICY_OCR.md`](docs/POLICY_OCR.md)).
Both need `tesseract-ocr`, `tesseract-ocr-data-spa`, `poppler-utils` and object
storage; each reports its own availability and disables only itself if either
+5
View File
@@ -447,6 +447,11 @@ for what's actually next.
## Statement OCR intake (`/recibos`) — DONE 2026-08-01
> As-built reference: **`docs/STATEMENT_OCR.md`** (written 2026-08-02) — the
> parsers, the matcher's scoped-field rules, confirm/learning semantics and the
> API surface. `docs/RECEIPT_CAPTURE_SPEC.md` §2 stays the design and the
> measured evidence. This section is the session record of building it.
Plan step 11 §2 (`docs/RECEIPT_CAPTURE_SPEC.md` §2). The last big utilities
feature: staff scan the month's utility bills and the machine proposes customer
+ amount per page, instead of keying 300+ statements per company by hand. Built
+4 -3
View File
@@ -233,9 +233,10 @@ signal at all (which must yield no provider rather than a bad guess).
## Related
- [`RECEIPT_CAPTURE_SPEC.md`](RECEIPT_CAPTURE_SPEC.md) §2 — the utility
statement pipeline this was lifted from, and the origin of three rules the
policy parser applies: detect the provider by brand before layout, only
- [`STATEMENT_OCR.md`](STATEMENT_OCR.md) — the utility statement pipeline this
was lifted from, as built. [`RECEIPT_CAPTURE_SPEC.md`](RECEIPT_CAPTURE_SPEC.md)
§2 is its design and the measured evidence behind it. Between them they are
the origin of three rules the policy parser applies: detect the provider by brand before layout, only
ever apply the Tesseract digit-confusion map (`O→0`, `S→5`, `B→8`, …) to
fields known to be digits, and parse amounts by separator *position* rather
than assuming `,` is thousands.
+6 -1
View File
@@ -131,7 +131,12 @@ single-movement form.
## 2. PDF / OCR auto-capture
> **BUILT — 2026-08-01.** Implemented and verified end to end against real
> **BUILT — 2026-08-01.** As-built reference:
> [`STATEMENT_OCR.md`](STATEMENT_OCR.md) — what the shipped feature does, its
> parsers, matcher rules and API surface. This section stays the *design* and
> the evidence behind it; go there for what is in the code today.
>
> Implemented and verified end to end against real
> scanned statements. `apps/api/src/statements/` holds the module: a swappable
> `OcrProvider` seam with a self-hosted Tesseract implementation, per-provider
> parsers for CFE / CESPT / Telnor / gas / predial, a scoped matcher, and a
+315
View File
@@ -0,0 +1,315 @@
# Utility Statement OCR Capture (receipt capture)
Reads the stack of scanned utility bills the office pays every month, proposes
the customer and the amount for each page, and posts the confirmed pages to the
ledger as one batch against one check. Built 2026-08-01, live under `/recibos`.
This is the **as-built** record. The design and the reasoning behind it are
[`RECEIPT_CAPTURE_SPEC.md`](RECEIPT_CAPTURE_SPEC.md) §2, which also carries the
measured results and the four spec corrections the real scans forced. Read that
for *why*; read this for *what is there*.
## The job it replaces
Staff receive 300+ pages per month per service provider — CFE electricity,
CESPT water, Telnor phone, gas, municipal predial, federal zone — and key each
one into the ledger by hand, as a charge against the customer whose property
the bill belongs to. The whole stack is paid with one office check, so the
capture is naturally a batch.
Auto-capture is a **mode of** the existing capture screen, not a separate
feature: it is the same daily job with a scanner instead of a keyboard, and
both modes post through the same ledger path.
## What ships
| Piece | Path |
|---|---|
| API module | `apps/api/src/statements/` (service, controller, DTOs, matcher, parsers) |
| OCR seam | `apps/api/src/statements/ocr/` (interface + Tesseract), bound in `apps/api/src/ocr/ocr.module.ts` |
| Tables | `statement_batches`, `statement_documents` (`20260731235721_statement_ocr_intake`) |
| Web | `components/Captura.tsx` (tab shell), `StatementIntake.tsx` (upload + batch list), `/recibos/:id` (review queue) |
| Abilities | `statement:ingest`, `statement:review` — both **STAFF** |
STAFF is deliberate: the review step is what makes machine capture safe at that
tier, since nothing reaches the ledger unconfirmed.
## The screen
`Captura` is one screen with two ways in:
- `/estado-cuenta/lote`**manual** tab (`ManualCheckCapture`, key each
receipt against one check by hand)
- `/recibos`**automática (OCR)** tab (`StatementIntake`, upload the scans)
- `/recibos/:id` → the batch review queue, page image beside the extracted
fields
Both end in the same place — charges on customers' ledgers posted against one
check — so they are modes of one screen. Staff pick by what is on the desk that
morning. Either URL renders the same component, so old bookmarks land on the
right tab.
## Pipeline
```
upload PDFs (+ service kind) → store source → render pages → text layer? → parse → match → review → confirm → ledger
```
**A batch is one service kind.** The uploader labels it (ELECTRIC, WATER,
TELEPHONE, GAS, PROPERTY_TAX, FEDERAL_ZONE) and that label is enforced: if the
parser reads a page as a different provider, the page is rejected as
mis-sorted rather than matched. Posting a phone bill as a water charge is the
failure being prevented.
**Processing is not awaited.** 300 pages of OCR is minutes of CPU, far past any
HTTP timeout, so `POST /statements/batches` returns the batch id immediately and
the client polls. That is also what lets the review queue show partial progress.
**One page = one document.** Statements arrive **bundled, one customer per
page** — Telnor's own `Pág 3 de 6` is its internal pagination, not the office's
scan — so every rendered page becomes its own `StatementDocument` and the
parser runs per page. (The policy OCR feature inverts this; see "Sibling
feature" below.)
**One unreadable page must not abandon the other 299.** A page that throws
becomes a single `OCR_FAILED` row and the loop continues.
Both the source PDFs (`statement/{batchId}/source-N.pdf`) and every rendered
page image (`page-N.png`) are stored. The source is the artifact the office
received and the only way to re-run a corrected parser over the original; the
page image is what the reviewer looks at, because "what the parser read" is
only checkable against a picture of the paper.
## The OCR seam
`OcrProvider` (`ocr/ocr.provider.ts`) is the swap point. Four methods:
`available()`, `renderPages()`, `recognize()`, `textPages()`.
Everything above it works in terms of page text and word boxes, so the engine
is replaceable without touching the parsers, the matcher or the schema. The
shipped implementation is **self-hosted Tesseract**, and that choice is
evidence-based rather than assumed — see the spec's measured results. A managed
API (Textract, Document Intelligence, Document AI) fits behind the same
interface with no schema change; at 300+ pages/month/company it would carry
real recurring cost for accuracy that is not the bottleneck.
The binding lives in `apps/api/src/ocr/ocr.module.ts`, extracted out of
`StatementsModule` when [`POLICY_OCR.md`](POLICY_OCR.md) needed the same seam.
`StatementsModule` imports it and binds nothing itself, so the engine decision
is one line in one file for both features.
### Text layer first, OCR as the fallback
**Not every statement is a scan.** The gas company sends born-digital CFDI
invoices whose text layer is already exact and already positioned.
`textPages()` reads it (`pdftotext -bbox-layout`, same poppler package as
`pdftoppm`) and OCR runs only where there is none.
Rasterising a born-digital page and re-recognising it can only lose
information — one sample turned `MEDIDOR: VM01014426` into
`ar (LTR): 014420` — while costing about a minute of CPU for the privilege.
Positions come back in the same pixel space `recognize()` uses, so the parsers'
geometric helpers work unchanged on either source. When the text layer is used
the document's notes say so verbatim: *"texto leído del PDF original, sin
OCR"*.
### Word boxes, not just text
`OcrPage` carries `words[]` with pixel boxes because several real layouts are
**tables**: the CESPT "RECIBO" prints `No. DE CUENTA` as a column header with
the value in the row beneath it, which line-oriented text cannot associate.
Parsers fall back to geometry for exactly those fields.
## Parsers
Eight providers, dispatched by a `BRAND` table checked before a `LAYOUT` table:
| Provider | Service kind |
|---|---|
| `CFE` | ELECTRIC |
| `CESPT` | WATER |
| `TELNOR` | TELEPHONE |
| `GAS TIJUANA` | GAS |
| `PREDIAL TIJUANA` / `PREDIAL ROSARITO` / `PREDIAL ENSENADA` | PROPERTY_TAX |
| `ZONA FEDERAL TIJUANA` | FEDERAL_ZONE |
Three predial parsers rather than one because Tijuana, Rosarito and Ensenada
issue three completely different documents — same tax, nothing else in common.
Rules that are load-bearing and easy to break:
- **Brand before layout, and never interleaved.** Scanned logos read badly (a
CESPT header came back as `E BAJA ES PAGO / EALIFORNIA`), which is why the
layout fallback exists — but *every* brand rule runs first, because a Telnor
page contains words a CFE structural rule would otherwise claim.
- **Tijuana bills predial and zona federal from the same treasury.** Same
header, same address, same `ATB-541201` RFC, so every predial discriminator
matches a zona federal page too. The words only that layout prints are
`Marítimo Terrestre`, so its rule is asked ahead of all three predial ones.
**Order matters here in a way that is invisible from the code shape.**
- **Parse amounts by separator position.** A real Telnor bill OCR'd as
`$ 649,00`; stripping commas as thousands separators makes that $64,900.
- **A misread `$` is the dangerous failure, not a missing one.** An Ensenada
receipt for `$2,203.00` OCR'd as `82,203.00` — the sign read as an 8, which
would post a charge 37× too large and look entirely ordinary in the ledger.
Every predial amount therefore requires a literal `$`; a page that cannot
produce one reports no amount and goes to review.
- **Digit-confusion repair only on fields known to be digits** (`O→0`, `S→5`,
`B→8`, …), never on free text.
- **The clave catastral is not `[A-Z]{2}[0-9]{6}`.** Position three is a letter
in fifteen of the 932 stored claves (`MMB01041`, `CGH52121`). Digitising the
whole tail maps that `B` to an `8` and yields a key matching no property.
- **Barcodes beat printed labels.** Where a provider prints a payment barcode
it is preferred and the two are cross-checked; disagreement sets
`crossChecked: false` and forces review, because which of the two was
misread is a judgement call.
## Matching
`StatementMatcherService`. Two rules govern everything:
**Match on one scoped field, never fuzzily across all identifiers.** Each
service kind has exactly one column its statements print, and only that column
is consulted. A blanket search over accountNumber/meterNumber/route would let a
water account number collide with an unrelated phone number, and the mis-post
would look perfectly ordinary in the ledger.
**Never match on the customer name.** A CESPT receipt for account `5365218`
prints `ARNAIZ ROSAS ELSA AURORA`; the office's book, corroborated by the
clave, has `CATT, RANDY`. The name on a utility bill is the registrant, not the
current owner. Names are shown to the reviewer and are never an input.
### `scopedRefField` — which column each kind actually prints
| Kind | Column | Why |
|---|---|---|
| ELECTRIC, WATER, TELEPHONE, CABLE | `accountNumber` | the legacy column holds the printed number |
| GAS | `meterNumber` | the number lived in free-text notes; `accountNumber` was never populated |
| PROPERTY_TAX | `meterNumber` | `accountNumber` holds `DATMEX.predial`, an office file number that is neither unique nor printed anywhere |
| FEDERAL_ZONE | `meterNumber` | `accountNumber` holds `DATMEX.zfed`, which is a **peso amount**, not a reference |
`scopedRefField` is exported because three places must agree on the answer: the
lookup, the blank-service fill on review, and the write-back on confirm. When
they disagree a reference gets learned into a column nothing searches, and the
same page returns to the review queue every month forever.
The `FEDERAL_ZONE` case is the sharpest instance of a trap this codebase hits
repeatedly (see also `policies.total`): a legacy column whose *name* promises
an identifier and whose *contents* are something else. Three of its 77 values
carry cents and one is negative. Worse than never matching — because every row
already has a value, the `[field]: null` guards on learning and on the
blank-service fill would never fire either.
### The clave catastral is a rescue on some layouts and the primary key on others
CESPT bills print the clave as well as an account number, so it rescues a page
whose account number did not OCR — which happened on real samples. There it
stays a hint.
On Rosarito and Ensenada predial the receipt prints **nothing else**, so a
unique clave hit is a real match and auto-matches. Tijuana predial prints no
clave at all; its only identifier is an 8-digit municipal account carried in a
32-digit payment barcode (`account(8) + DDMMYY + amount(9) + folio(9)`) that
the legacy database never held, so those pages start cold and are taught by the
first confirm.
Multiple hits are always surfaced, never auto-picked — duplicate account
numbers do occur in the legacy data, and the office's own `DUPLICADOS` report
existed for a reason.
## Confirm: what gets written
`confirmBatch` posts through **`BillingService.createBatch`** — the same method
the manual Editor screen uses — rather than writing `Transaction` rows
directly, so OCR-sourced and hand-keyed receipts share one write path, one
validation path and one audit trail.
- `source: "OCR"` and a per-line `captureRef` of the document id feed the
duplicate-post guard, so a batch confirmed twice cannot double-charge.
- `items[i]` is positionally parallel to `lines[i]` (a documented seam
guarantee), so the created rows zip straight back onto the documents that
produced them via `postedTransactionId`.
- **The sign is applied here.** Charges are negative in this ledger; the parser
reads the printed positive figure, and `-Math.abs()` is applied at the single
point where a statement becomes a ledger row.
- A missing amount blocks the confirm with the offending page numbers, rather
than silently posting zero.
### Learning: the cold start is a one-time cost
After posting, `learnAccountRefs` writes each confirmed reference back onto the
`PropertyService` that matched — **only where the field was null**. Never
overwrites a number already on file, which would let one misread page rewrite
good reference data.
This is what turns gas (whose numbers the migration never populated) and
Tijuana predial (whose municipal account the legacy database never held) from a
permanent review queue into a one-time cost: next month's statement for the
same account matches on its own.
### Discarding
Refused once any page is `POSTED` — those pages already wrote ledger rows
against a check, and a "discarded" label on the batch would leave the charges
unexplained. Reject the remaining pages individually instead.
## API surface
| Method | Route | Ability |
|---|---|---|
| `GET` | `/statements/status` (is OCR + storage available) | authenticated |
| `GET` | `/statements/batches`, `/batches/:id`, `/batches/:id/documents` | authenticated |
| `GET` | `/statements/documents/:id/page` (streams the page image) | authenticated |
| `POST` | `/statements/batches` (upload + service kind) | `statement:ingest` |
| `PATCH` | `/statements/documents/:id` (correct a field or the match) | `statement:review` |
| `POST` | `/statements/documents/:id/reject` | `statement:review` |
| `POST` | `/statements/batches/:id/discard` | `statement:review` |
| `POST` | `/statements/batches/:id/confirm` | `statement:review` |
## Requirements
Object storage (`S3_ENDPOINT` + credentials) for the scans, and `tesseract-ocr`
/ `tesseract-ocr-data-spa` / `poppler-utils` in the API image. Both are checked
at upload rather than at the first write — a missing dependency should be a 400
on the request, not a `FAILED` batch minutes later. `GET /statements/status`
reports both and the upload card hides itself unless both hold.
## Tests
- `parsers/statement-parser.spec.ts``detectProvider`,
`normalizeCadastralKey`, and one suite per newer parser
(`parsePredialTijuana`, `parsePredialRosarito`, `parsePredialEnsenada`,
`parseGas`, `parseZonaFederal`). Every fixture is a **verbatim OCR excerpt
from a real receipt**, including the ones that bit: the `82,203.00` Ensenada
dollar sign, the three-letter clave, the CESPT logo garbage.
- `ocr/tesseract.provider.spec.ts``parseBboxLayout`, the
`pdftotext -bbox-layout` reader that produces the text-layer `OcrPage`
(positions included, which is what lets the geometric helpers work on
born-digital input).
Note the CFE / CESPT / Telnor parsers themselves have **no unit suite** — they
predate the gas/predial extension and were verified against the 46-page corpus
end to end rather than in isolation. Worth closing if they are touched.
## Not built
- **Handwritten folder numbers.** Staff pencil a customer number on each bill
(`9`, `405`); Tesseract read `405` as `205`. Handwriting is a review hint at
best and is deliberately not an input to matching.
- **Re-running a corrected parser over a stored batch.** The source PDFs are
kept precisely so this is possible, but nothing exposes it yet.
- **Providers beyond the eight above.** Adding one is a `BRAND` entry, an
optional `LAYOUT` entry, and a parser function.
## Sibling feature
[`POLICY_OCR.md`](POLICY_OCR.md) — the same pipeline reading carrier policy
PDFs into `Policy` rows, built out of this one. It reuses the seam, the
provider-detection ordering, the digit-confusion map and the
amount-by-separator rule.
**One assumption does not carry over.** Here a page *is* a document, because
statements arrive one customer per page. A policy PDF is one document across
several pages, so that feature concatenates the pages and parses once per file.
If you are porting a change between the two, that is the difference to check
first.