feat(recibos): OCR capture for gas butano and municipal predial
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m50s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m8s

Adds four parsers to the statement intake — GAS TIJUANA plus one per
municipality, because Tijuana, Rosarito and Ensenada issue three
completely different predial documents — and a text-layer fast path for
the born-digital invoices the gas company sends.

Measured against a new corpus of 14 documents / 29 pages: provider read
on 29/29, amount on 26/29, and 21/29 auto-matched against the dev
database (22/29 identified). The eight review cases are all legitimate.

Five things the corpus forced:

- Not every statement is a scan. The gas invoices are born-digital CFDIs
  whose text layer is exact; rasterising them only loses information (one
  sample turned `MEDIDOR: VM01014426` into `ar (LTR): 014420`). The new
  `OcrProvider.textPages` reads the embedded layer via `pdftotext
  -bbox-layout` — same poppler package as `pdftoppm`, so no new
  dependency — and OCR stays the fallback for real scans. Poppler's own
  `<line>` grouping follows text flow rather than the page, so words are
  regrouped by vertical position; without that, a two-column header
  leaves every label separated from the value printed beside it.

- The clave catastral is not two letters and six digits. Position three
  is a letter in 15 of the 932 stored claves, and digitising the whole
  tail mapped a real `MMB01041` to a nonexistent `MM801041`.

- Tijuana predial prints no clave at all. Its only identifier is an
  8-digit municipal account carried in a 32-digit payment barcode, which
  the legacy database never held, so it goes in `meterNumber` alongside
  gas — `accountNumber` holds `DATMEX.predial`, which is not a
  per-property key and must not be overwritten. Those pages start cold
  and are taught by the first confirm.

- On Rosarito and Ensenada the clave is the primary key, not a fallback:
  those receipts print nothing else, so a unique hit auto-matches. On a
  utility bill that merely happens to print one it stays a review hint.

- A misread `$` is the dangerous failure. An Ensenada receipt for
  $2,203.00 OCR'd as `82,203.00`, which would post a charge 37x too large
  and look ordinary in the ledger. Predial amounts now require a literal
  `$` and a page that cannot produce one goes to review.

The scoped match field is now one exported function rather than three
copies of `kind === "GAS" ? ... : ...`, since the lookup, the
blank-service fill and the confirm write-back have to agree or a
reference gets learned into a column nothing searches.

First tests in this package: 23 specs over the parsers and the text-layer
reader, every fixture a verbatim OCR excerpt from a real receipt. Adds
the jest config they need and a build tsconfig so they stay out of dist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-01 12:52:20 -07:00
co-authored by Claude Opus 5
parent 216309190c
commit d6501f1d74
12 changed files with 940 additions and 58 deletions
+24 -6
View File
@@ -6,8 +6,10 @@
* The shipped implementation is self-hosted Tesseract (see tesseract.provider).
* That choice is evidence-based rather than assumed: run against 46 pages of
* real scanned CFE, CESPT and Telnor statements, it identified the provider on
* 46/46 and extracted a usable account reference on 43/46, which is well past
* the bar for a queue whose whole point is that a human confirms every row. A
* 46/46 and extracted a usable account reference on 43/46, and on a later
* corpus of 19 scanned municipal predial receipts it read the provider on
* 19/19 and an identifier on 18/19 — well past the bar for a queue whose whole
* point is that a human confirms every row. A
* managed document-extraction API (Textract, Document Intelligence, Document
* AI) fits behind this same interface if per-page accuracy ever proves
* insufficient, with no schema change — but at 300+ pages/month/company it
@@ -31,10 +33,10 @@ export interface OcrPage {
/** Full page text, reading order, newline-separated. */
text: string;
/**
* Word boxes. Needed because two of the three real layouts are *tables* —
* the CESPT "RECIBO" prints `No. DE CUENTA` as a column header with the
* value in the row beneath it, which line-oriented text cannot associate.
* Parsers fall back to geometry for exactly those fields.
* Word boxes. Needed because several of the real layouts are *tables* — the
* CESPT "RECIBO" prints `No. DE CUENTA` as a column header with the value in
* the row beneath it, which line-oriented text cannot associate. Parsers fall
* back to geometry for exactly those fields.
*/
words: OcrWord[];
/** Mean word confidence across the page, 0..1. */
@@ -48,6 +50,22 @@ export interface OcrProvider {
renderPages(pdf: Buffer): Promise<Buffer[]>;
/** OCR a single rendered page image. */
recognize(pageImage: Buffer): Promise<OcrPage>;
/**
* Read a PDF's own text layer, one entry per page, `null` where the page has
* none worth using.
*
* Not every statement is a scan. The gas company e-mails born-digital CFDI
* invoices whose text is already exact and already positioned — running those
* through a rasteriser and a character recogniser can only lose information
* (one sample turned `MEDIDOR: VM01014426` into `ar (LTR): 014420`) while
* costing about a minute of CPU per page for the privilege. Where the layer
* exists it is strictly better input for the same parsers, so it is tried
* first and OCR remains the fallback for genuine scans.
*
* Positions are reported in the same pixel space `recognize` uses, so the
* geometric helpers in the parsers work unchanged on either source.
*/
textPages(pdf: Buffer): Promise<(OcrPage | null)[]>;
}
export const OCR_PROVIDER = Symbol("OCR_PROVIDER");
@@ -0,0 +1,67 @@
import { parseBboxLayout } from "./tesseract.provider";
/**
* Shaped like real `pdftotext -bbox-layout` output: the gas invoice lays its
* header out as two columns of independent text flows, so poppler puts a label
* and the value printed beside it in *different* `<line>` elements. Trusting
* that grouping is what left `PERIODO FACTURADO` with no value next to it and
* every period field empty on a batch whose text was perfectly readable.
*/
function word(x: number, y: number, text: string): string {
return `<word xMin="${x}" yMin="${y}" xMax="${x + 20}" yMax="${y + 8}">${text}</word>`;
}
function doc(...lines: string[]): string {
return `<doc><page width="612" height="792">${lines
.map((l) => `<flow><block><line>${l}</line></block></flow>`)
.join("")}</page></doc>`;
}
/** Enough words on the page to clear the "is this a real text layer" floor. */
function padding(): string {
return Array.from({ length: 50 }, (_, i) => word(10, 400 + i * 10, `w${i}`)).join("");
}
describe("parseBboxLayout", () => {
it("rejoins a label with the value printed beside it in another flow", () => {
const [page] = parseBboxLayout(
doc(
word(20, 100, "PERIODO") + word(45, 100, "FACTURADO:"),
word(300, 100.4, "20260630-20260630"),
padding(),
),
1,
);
expect(page).not.toBeNull();
expect(page!.text).toContain("PERIODO FACTURADO: 20260630-20260630");
});
it("keeps genuinely separate lines apart", () => {
const [page] = parseBboxLayout(
doc(word(20, 100, "Cuenta:") + word(80, 100, "0900003463"), word(20, 130, "Nombre:"), padding()),
1,
);
expect(page!.text.split("\n")).toContain("Cuenta: 0900003463");
expect(page!.text.split("\n")).toContain("Nombre:");
});
it("scales point coordinates into the render's pixel space", () => {
// Word boxes have to land in the same coordinate space tesseract reports,
// or the geometric helpers the parsers share silently stop finding values.
const [page] = parseBboxLayout(doc(word(72, 144, "X") + padding()), 300 / 72);
const x = page!.words.find((w) => w.text === "X")!;
expect(x.left).toBeCloseTo(300);
expect(x.top).toBeCloseTo(600);
});
it("reports no text layer for a scan carrying a few stray glyphs", () => {
expect(parseBboxLayout(doc(word(10, 10, "3") + word(40, 10, "of") + word(60, 10, "5")), 1)).toEqual([
null,
]);
});
it("decodes the entities poppler escapes", () => {
const [page] = parseBboxLayout(doc(word(10, 10, "A&amp;B") + padding()), 1);
expect(page!.text).toContain("A&B");
});
});
@@ -105,6 +105,37 @@ export class TesseractOcrProvider implements OcrProvider {
});
}
/**
* `pdftotext -bbox-layout` — the same poppler package `pdftoppm` comes from,
* so this costs no extra dependency in the runtime image.
*
* A page is only accepted when it carries a real text layer. Scanned PDFs
* frequently contain a handful of stray glyphs (a scanner watermark, a page
* number stamped by the MFP), and treating those as the page's text would
* hand every parser an almost-empty string and silently take OCR out of the
* loop — so a floor of MIN_TEXT_WORDS words has to be present before the
* layer is believed.
*/
async textPages(pdf: Buffer): Promise<(OcrPage | null)[]> {
await this.require();
return this.scratch(async (dir) => {
const src = join(dir, "in.pdf");
await writeFile(src, pdf);
const out = join(dir, "out.html");
try {
await run("pdftotext", ["-bbox-layout", src, out]);
} catch (err) {
this.logger.warn(
`pdftotext failed; falling back to OCR for this file: ${(err as Error).message}`,
);
return [];
}
// Points to pixels at the render DPI, so word boxes from either source
// land in one coordinate space and `valueUnder`'s thresholds hold.
return parseBboxLayout(await readFile(out, "utf8"), this.dpi / 72);
});
}
async recognize(pageImage: Buffer): Promise<OcrPage> {
await this.require();
return this.scratch(async (dir) => {
@@ -136,6 +167,130 @@ export class TesseractOcrProvider implements OcrProvider {
}
}
/**
* Below this many words a "text layer" is scanner debris, not a document.
* The real born-digital samples carry 400+ words a page; the scanned ones
* carry none at all, so the exact threshold is not delicate.
*/
const MIN_TEXT_WORDS = 40;
const ENTITIES: Record<string, string> = {
amp: "&",
lt: "<",
gt: ">",
quot: '"',
apos: "'",
};
function decodeEntities(s: string): string {
return s.replace(/&(#x?[0-9a-fA-F]+|[a-z]+);/g, (whole, body: string) => {
if (body[0] === "#") {
const code =
body[1] === "x" || body[1] === "X"
? parseInt(body.slice(2), 16)
: parseInt(body.slice(1), 10);
return Number.isFinite(code) ? String.fromCodePoint(code) : whole;
}
return ENTITIES[body] ?? whole;
});
}
/**
* Turn `pdftotext -bbox-layout`'s XHTML into one OcrPage per PDF page.
*
* Parsed with regexes rather than an XML library on purpose: the output is
* machine-generated by poppler with a fixed element shape (`page` > `flow` >
* `block` > `line` > `word`), and the alternative is a parser dependency in
* the API for one file format read in one place. Only `page` and `word` are
* consulted — see below for why poppler's own `line` grouping is discarded.
*
* `confidence` is 1 for every word: these are the document's own characters,
* not a recognition guess.
*/
export function parseBboxLayout(xhtml: string, scale: number): (OcrPage | null)[] {
const pages: (OcrPage | null)[] = [];
for (const pageMatch of xhtml.matchAll(/<page\b[^>]*>([\s\S]*?)<\/page>/g)) {
const words: OcrWord[] = [];
for (const w of pageMatch[1].matchAll(
/<word\s+xMin="([\d.eE+-]+)"\s+yMin="([\d.eE+-]+)"\s+xMax="([\d.eE+-]+)"\s+yMax="([\d.eE+-]+)"\s*>([\s\S]*?)<\/word>/g,
)) {
const text = decodeEntities(w[5]).trim();
if (!text) continue;
const left = Number(w[1]) * scale;
const top = Number(w[2]) * scale;
words.push({
text,
left,
top,
width: Number(w[3]) * scale - left,
height: Number(w[4]) * scale - top,
confidence: 1,
});
}
pages.push(
words.length >= MIN_TEXT_WORDS
? { text: toVisualRows(words), words, confidence: 1 }
: null,
);
}
return pages;
}
/**
* Reassemble words into the rows a reader sees, left to right.
*
* Poppler's own `<line>` grouping cannot be used for this. It groups by text
* flow, and these invoices lay their fields out as two columns of independent
* flows — so `PERIODO FACTURADO:` and the `20260630-20260630` printed beside
* it end up in different `<line>` elements, and every label-then-value pattern
* in the parsers misses a value that is plainly there on the page. Regrouping
* by vertical position restores the adjacency, and matches what tesseract
* hands back for the scanned version of the same layout.
*
* Rows are cut when a word's vertical centre leaves the band established by
* the row's first word, which tolerates the sub-pixel baseline differences
* between fonts on one line without merging two genuinely separate lines.
*/
function toVisualRows(words: OcrWord[]): string {
const centre = (w: OcrWord) => w.top + w.height / 2;
const sorted = [...words].sort((a, b) => centre(a) - centre(b) || a.left - b.left);
const rows: OcrWord[][] = [];
let current: OcrWord[] = [];
let band = 0;
for (const w of sorted) {
if (!current.length) {
current = [w];
band = centre(w);
continue;
}
// Half the word's own height: tall headings and body text both sit within
// their own line's band, and neither reaches into the next one.
if (Math.abs(centre(w) - band) <= Math.max(w.height, current[0].height) / 2) {
current.push(w);
} else {
rows.push(current);
current = [w];
band = centre(w);
}
}
if (current.length) rows.push(current);
return rows
.map((r) =>
[...r]
.sort((a, b) => a.left - b.left)
.map((w) => w.text)
.join(" "),
)
.join("\n");
}
/**
* Turn tesseract's TSV into words plus reassembled text.
*