Commit Graph
4 Commits
Author SHA1 Message Date
rmancinas 3125b52057 feat(ocr): discard abandoned capture batches
A bad scan, the wrong PDFs or a duplicate upload used to leave a batch
sitting in READY_FOR_REVIEW forever, because the only exits were confirm
(posts to the books) or rejecting every page one at a time. Add a
DISCARDED terminal status to both OCR domains and a single endpoint per
domain that rejects every page still pending in one shot.

Discarding is refused once anything has landed: statements once a page is
POSTED, policies once a page is APPLIED. Those batches did real work and
have to be settled page by page.

- POST /statements/batches/:id/discard
- POST /policy-ocr/batches/:id/discard
- shared DiscardBatchCard on both review screens, gated the same way
2026-08-02 02:00:02 -07:00
rmancinasandClaude Opus 5 d6501f1d74 feat(recibos): OCR capture for gas butano and municipal predial
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m50s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m8s
Adds four parsers to the statement intake — GAS TIJUANA plus one per
municipality, because Tijuana, Rosarito and Ensenada issue three
completely different predial documents — and a text-layer fast path for
the born-digital invoices the gas company sends.

Measured against a new corpus of 14 documents / 29 pages: provider read
on 29/29, amount on 26/29, and 21/29 auto-matched against the dev
database (22/29 identified). The eight review cases are all legitimate.

Five things the corpus forced:

- Not every statement is a scan. The gas invoices are born-digital CFDIs
  whose text layer is exact; rasterising them only loses information (one
  sample turned `MEDIDOR: VM01014426` into `ar (LTR): 014420`). The new
  `OcrProvider.textPages` reads the embedded layer via `pdftotext
  -bbox-layout` — same poppler package as `pdftoppm`, so no new
  dependency — and OCR stays the fallback for real scans. Poppler's own
  `<line>` grouping follows text flow rather than the page, so words are
  regrouped by vertical position; without that, a two-column header
  leaves every label separated from the value printed beside it.

- The clave catastral is not two letters and six digits. Position three
  is a letter in 15 of the 932 stored claves, and digitising the whole
  tail mapped a real `MMB01041` to a nonexistent `MM801041`.

- Tijuana predial prints no clave at all. Its only identifier is an
  8-digit municipal account carried in a 32-digit payment barcode, which
  the legacy database never held, so it goes in `meterNumber` alongside
  gas — `accountNumber` holds `DATMEX.predial`, which is not a
  per-property key and must not be overwritten. Those pages start cold
  and are taught by the first confirm.

- On Rosarito and Ensenada the clave is the primary key, not a fallback:
  those receipts print nothing else, so a unique hit auto-matches. On a
  utility bill that merely happens to print one it stays a review hint.

- A misread `$` is the dangerous failure. An Ensenada receipt for
  $2,203.00 OCR'd as `82,203.00`, which would post a charge 37x too large
  and look ordinary in the ledger. Predial amounts now require a literal
  `$` and a page that cannot produce one goes to review.

The scoped match field is now one exported function rather than three
copies of `kind === "GAS" ? ... : ...`, since the lookup, the
blank-service fill and the confirm write-back have to agree or a
reference gets learned into a column nothing searches.

First tests in this package: 23 specs over the parsers and the text-layer
reader, every fixture a verbatim OCR excerpt from a real receipt. Adds
the jest config they need and a build tsconfig so they stay out of dist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 12:52:20 -07:00
rmancinasandClaude Opus 5 b59abda895 feat(captura): fold recibo OCR into Captura as an automatic mode
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m46s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m17s
Scanning a stack of bills and keying them in are the same daily job, ending
in the same ledger path, so OCR intake becomes a mode of the capture screen
instead of a second menu entry:

- components/Captura.tsx holds the mode switch; the manual check form moves
  verbatim to components/ManualCheckCapture.tsx and the OCR intake to
  components/StatementIntake.tsx.
- /estado-cuenta/lote opens on manual, /recibos on automatic — both render
  Captura, so batch-review links and old bookmarks still land right.
- Nav drops "Recibos (OCR)"; "Captura" covers both, with a NavLink.aliases
  field so /recibos still highlights it.

Also fixes the "El almacenamiento de documentos no está configurado" failure
staff hit on upload. Uploading with no object storage configured used to
succeed, then die on the first put minutes later, leaving a FAILED batch
whose only explanation was that string. createBatch now refuses up front,
GET /statements/status reports storageAvailable alongside ocrAvailable, and
the intake tab explains the situation instead of offering an upload that
cannot work. S3_* documented in .env.example (deploy stacks already set it).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 01:04:09 -07:00
rmancinasandClaude Opus 5 4d5008b545 feat(statements): OCR intake for scanned utility bills
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m41s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m18s
Staff key 300+ utility statements per company per month by hand. This adds
the ingest -> split -> OCR -> match -> review pipeline that proposes customer
and amount per page instead (RECEIPT_CAPTURE_SPEC §2), posting through the
existing BillingService.createBatch seam with source=OCR and a per-document
captureRef so machine and hand capture share one write path and audit trail.

Everything was designed against 10 real scanned statements (46 pages of CFE,
CESPT and Telnor bills) rather than from the sample-free spec. The scans have
no text layer at all — they are camera images — so OCR is mandatory, and they
arrive bundled one customer per page. Measured on those pages the parser
identifies the provider 46/46 and reads an account reference 43/46; against
the dev database that is 39/46 (85%) exact auto-match, 40/46 identified, with
the rest genuine review cases. That closes the OCR-provider question in favour
of self-hosted Tesseract: it clears the bar for a queue where a human confirms
every row, and OcrProvider keeps a managed API a one-line swap.

The samples corrected three things the spec had wrong or unknown:

- Clave catastral is NOT predial. DATMEX.clave (934 rows) is what CESPT and
  predial bills print; DATMEX.predial, which PROPERTY_TAX.accountNumber holds,
  has 663 distinct values across 1135 rows and appears on no statement. The
  clave now lives on Property.cadastralKey as the matcher's secondary key;
  predial is left untouched. This had been blocking predial matching.
- Gas was recoverable: 160 of 334 DATMEX.gas values are real account numbers
  (the rest are ESTACIONARIO/CILINDRO descriptors), now in GAS.meterNumber.
- Phone is one billed line per property (534/18/1 across phone1/2/3), so the
  new TELEPHONE ServiceKind backfills from phone1 only, not three rows.

Matching is scoped to one column per service kind and never reads the customer
name — a CESPT receipt prints ARNAIZ ROSAS ELSA AURORA for an account this
office holds under CATT, RANDY, because the printed name is the registrant,
not the current owner. Where a provider prints a payment barcode it beats the
printed label (one CFE label OCR'd a digit too many while its barcode was
correct) and the two cross-check, with disagreement forcing review.

Confirming a document whose service had no reference writes it back, so gas
and any other cold start is a one-time cost rather than a permanent queue.

Verified end to end against the live dev API and MinIO: real scans uploaded
over HTTP, matched, confirmed against a check, and the resulting rows checked
in MySQL (negative amounts, captureSource=OCR, concept derived from the batch
kind, captureRef linking back to each page). Re-confirming a posted batch is
refused. Test data was removed afterwards.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 00:42:35 -07:00