feat(recibos): OCR capture for gas butano and municipal predial
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m50s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m8s

Adds four parsers to the statement intake — GAS TIJUANA plus one per
municipality, because Tijuana, Rosarito and Ensenada issue three
completely different predial documents — and a text-layer fast path for
the born-digital invoices the gas company sends.

Measured against a new corpus of 14 documents / 29 pages: provider read
on 29/29, amount on 26/29, and 21/29 auto-matched against the dev
database (22/29 identified). The eight review cases are all legitimate.

Five things the corpus forced:

- Not every statement is a scan. The gas invoices are born-digital CFDIs
  whose text layer is exact; rasterising them only loses information (one
  sample turned `MEDIDOR: VM01014426` into `ar (LTR): 014420`). The new
  `OcrProvider.textPages` reads the embedded layer via `pdftotext
  -bbox-layout` — same poppler package as `pdftoppm`, so no new
  dependency — and OCR stays the fallback for real scans. Poppler's own
  `<line>` grouping follows text flow rather than the page, so words are
  regrouped by vertical position; without that, a two-column header
  leaves every label separated from the value printed beside it.

- The clave catastral is not two letters and six digits. Position three
  is a letter in 15 of the 932 stored claves, and digitising the whole
  tail mapped a real `MMB01041` to a nonexistent `MM801041`.

- Tijuana predial prints no clave at all. Its only identifier is an
  8-digit municipal account carried in a 32-digit payment barcode, which
  the legacy database never held, so it goes in `meterNumber` alongside
  gas — `accountNumber` holds `DATMEX.predial`, which is not a
  per-property key and must not be overwritten. Those pages start cold
  and are taught by the first confirm.

- On Rosarito and Ensenada the clave is the primary key, not a fallback:
  those receipts print nothing else, so a unique hit auto-matches. On a
  utility bill that merely happens to print one it stays a review hint.

- A misread `$` is the dangerous failure. An Ensenada receipt for
  $2,203.00 OCR'd as `82,203.00`, which would post a charge 37x too large
  and look ordinary in the ledger. Predial amounts now require a literal
  `$` and a page that cannot produce one goes to review.

The scoped match field is now one exported function rather than three
copies of `kind === "GAS" ? ... : ...`, since the lookup, the
blank-service fill and the confirm write-back have to agree or a
reference gets learned into a column nothing searches.

First tests in this package: 23 specs over the parsers and the text-layer
reader, every fixture a verbatim OCR excerpt from a real receipt. Adds
the jest config they need and a build tsconfig so they stay out of dist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-01 12:52:20 -07:00
co-authored by Claude Opus 5
parent 216309190c
commit d6501f1d74
12 changed files with 940 additions and 58 deletions
@@ -105,6 +105,37 @@ export class TesseractOcrProvider implements OcrProvider {
});
}
/**
* `pdftotext -bbox-layout` — the same poppler package `pdftoppm` comes from,
* so this costs no extra dependency in the runtime image.
*
* A page is only accepted when it carries a real text layer. Scanned PDFs
* frequently contain a handful of stray glyphs (a scanner watermark, a page
* number stamped by the MFP), and treating those as the page's text would
* hand every parser an almost-empty string and silently take OCR out of the
* loop — so a floor of MIN_TEXT_WORDS words has to be present before the
* layer is believed.
*/
async textPages(pdf: Buffer): Promise<(OcrPage | null)[]> {
await this.require();
return this.scratch(async (dir) => {
const src = join(dir, "in.pdf");
await writeFile(src, pdf);
const out = join(dir, "out.html");
try {
await run("pdftotext", ["-bbox-layout", src, out]);
} catch (err) {
this.logger.warn(
`pdftotext failed; falling back to OCR for this file: ${(err as Error).message}`,
);
return [];
}
// Points to pixels at the render DPI, so word boxes from either source
// land in one coordinate space and `valueUnder`'s thresholds hold.
return parseBboxLayout(await readFile(out, "utf8"), this.dpi / 72);
});
}
async recognize(pageImage: Buffer): Promise<OcrPage> {
await this.require();
return this.scratch(async (dir) => {
@@ -136,6 +167,130 @@ export class TesseractOcrProvider implements OcrProvider {
}
}
/**
* Below this many words a "text layer" is scanner debris, not a document.
* The real born-digital samples carry 400+ words a page; the scanned ones
* carry none at all, so the exact threshold is not delicate.
*/
const MIN_TEXT_WORDS = 40;
const ENTITIES: Record<string, string> = {
amp: "&",
lt: "<",
gt: ">",
quot: '"',
apos: "'",
};
function decodeEntities(s: string): string {
return s.replace(/&(#x?[0-9a-fA-F]+|[a-z]+);/g, (whole, body: string) => {
if (body[0] === "#") {
const code =
body[1] === "x" || body[1] === "X"
? parseInt(body.slice(2), 16)
: parseInt(body.slice(1), 10);
return Number.isFinite(code) ? String.fromCodePoint(code) : whole;
}
return ENTITIES[body] ?? whole;
});
}
/**
* Turn `pdftotext -bbox-layout`'s XHTML into one OcrPage per PDF page.
*
* Parsed with regexes rather than an XML library on purpose: the output is
* machine-generated by poppler with a fixed element shape (`page` > `flow` >
* `block` > `line` > `word`), and the alternative is a parser dependency in
* the API for one file format read in one place. Only `page` and `word` are
* consulted — see below for why poppler's own `line` grouping is discarded.
*
* `confidence` is 1 for every word: these are the document's own characters,
* not a recognition guess.
*/
export function parseBboxLayout(xhtml: string, scale: number): (OcrPage | null)[] {
const pages: (OcrPage | null)[] = [];
for (const pageMatch of xhtml.matchAll(/<page\b[^>]*>([\s\S]*?)<\/page>/g)) {
const words: OcrWord[] = [];
for (const w of pageMatch[1].matchAll(
/<word\s+xMin="([\d.eE+-]+)"\s+yMin="([\d.eE+-]+)"\s+xMax="([\d.eE+-]+)"\s+yMax="([\d.eE+-]+)"\s*>([\s\S]*?)<\/word>/g,
)) {
const text = decodeEntities(w[5]).trim();
if (!text) continue;
const left = Number(w[1]) * scale;
const top = Number(w[2]) * scale;
words.push({
text,
left,
top,
width: Number(w[3]) * scale - left,
height: Number(w[4]) * scale - top,
confidence: 1,
});
}
pages.push(
words.length >= MIN_TEXT_WORDS
? { text: toVisualRows(words), words, confidence: 1 }
: null,
);
}
return pages;
}
/**
* Reassemble words into the rows a reader sees, left to right.
*
* Poppler's own `<line>` grouping cannot be used for this. It groups by text
* flow, and these invoices lay their fields out as two columns of independent
* flows — so `PERIODO FACTURADO:` and the `20260630-20260630` printed beside
* it end up in different `<line>` elements, and every label-then-value pattern
* in the parsers misses a value that is plainly there on the page. Regrouping
* by vertical position restores the adjacency, and matches what tesseract
* hands back for the scanned version of the same layout.
*
* Rows are cut when a word's vertical centre leaves the band established by
* the row's first word, which tolerates the sub-pixel baseline differences
* between fonts on one line without merging two genuinely separate lines.
*/
function toVisualRows(words: OcrWord[]): string {
const centre = (w: OcrWord) => w.top + w.height / 2;
const sorted = [...words].sort((a, b) => centre(a) - centre(b) || a.left - b.left);
const rows: OcrWord[][] = [];
let current: OcrWord[] = [];
let band = 0;
for (const w of sorted) {
if (!current.length) {
current = [w];
band = centre(w);
continue;
}
// Half the word's own height: tall headings and body text both sit within
// their own line's band, and neither reaches into the next one.
if (Math.abs(centre(w) - band) <= Math.max(w.height, current[0].height) / 2) {
current.push(w);
} else {
rows.push(current);
current = [w];
band = centre(w);
}
}
if (current.length) rows.push(current);
return rows
.map((r) =>
[...r]
.sort((a, b) => a.left - b.left)
.map((w) => w.text)
.join(" "),
)
.join("\n");
}
/**
* Turn tesseract's TSV into words plus reassembled text.
*