Compare commits

..
107 Commits
Author SHA1 Message Date
gitea-actions 8c144fe8c4 chore(release): v1.0.22
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m6s
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m0s
Deploy on tag / Deploy to galactus (push) Successful in 9s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.22.
2026-08-15 20:06:28 +00:00
rmancinasandClaude Opus 5 4f2f064955 fix(web): OCR batch review returns to capture, not the policy list
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m1s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m3s
"Volver a pólizas" on /polizas/captura/[id] dropped the reviewer at
/polizas, so getting back to the batch queue meant navigating in again.
Points at /polizas/captura and reads "Volver a captura", matching the
statement review screen, which returns to /recibos.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 13:04:14 -07:00
rmancinasandClaude Opus 5 2f99bd5f98 fix(web): define the layout utilities the screens were already using
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m5s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m50s
The policy OCR review header read
"Para revisarPágina 1700489616· PAMELA DENISE WAGONERLICENCIASANA".

`.row`, `.stack`, `.tag`, `.page-sub` and `.state-warn` are used across the
app but no rule ever defined them. Without `display: flex` the `gap` those
call sites pass does nothing, and JSX drops the newline between sibling
elements, so the header's spans concatenated. `.tag` rendered as prose rather
than as a chip for the same reason.

Defined against the existing design tokens: `.tag` takes `.badge`'s shape,
`.state-warn` is the gold sibling of `.state-error`, and `p.page-sub` keeps
the block margin while the inline form drops it so a flex row still centres.

Also dropped the hand-rolled "· " separator and margin from the OCR header —
the flex gap does that now, and the literal dot was left floating in it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 12:58:25 -07:00
gitea-actions 458b2b272d chore(release): v1.0.21
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m13s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m22s
Deploy on tag / Deploy to galactus (push) Successful in 8s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.21.
2026-08-15 19:41:02 +00:00
rmancinasandClaude Opus 5 81938877ed feat(policy-ocr): suggest the customer from the printed insured name
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m57s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m5s
The office books customers surname-first ("WAGONER, PAMELA") and carriers
print them given-name-first ("PAMELA DENISE WAGONER"), so the review screen
made staff retype a name the parser had already read. Comparing normalized
token sets makes the two orderings the same thing.

Only on the zero-hit path, where the policy number found nothing and a human
has to pick a customer anyway. The suggestions are written to a new
`customerSuggestions` column rather than `matchCandidates`, which the review
screen reads as policy-number hits, and they never set `matchedCustomerId` or
`confident` — matching on `Policy.policyNumber` is unchanged.

Two tiers, drawn where the real book has cliffs: EXACT (identical token sets)
and PARTIAL (containment, >=2 shared tokens, surname present). Of 1536
customers, 1487 have a distinct token set, so EXACT cross-person collisions
are ~0; loosen to surname + first given name and 131 (8.5%) collide, and 185
surnames are shared by 524 customers, which is why one token is never enough
and the surname must be printed explicitly. Replaying every book row as a
carrier would print it: 97.9% top-ranked correct, 1.2% a different row, all
but two of those the same human on a duplicate or variant row.

Normalization folds accents (OCR's MUNOZ reaches the book's MUÑOZ), drops
initials, Spanish particles, JR/S.A. DE C.V., and any token with a digit —
ANA prints the phone hard against the name as `Ph.3102001538`. Names over 8
tokens or 80 characters are refused outright, because GMX's especificación
has no field labels and the parser has handed its whole first page over as
`insuredName`.

Not used for utility statements: there the registrant genuinely is not the
customer, so the same trick would be wrong rather than noisy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 12:37:23 -07:00
gitea-actions d854dff091 chore(release): v1.0.20
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m2s
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m3s
Deploy on tag / Deploy to galactus (push) Successful in 24s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.20.
2026-08-15 18:59:38 +00:00
rmancinasandClaude Opus 5 19864f16f2 fix(ocr): keep the printed layout when rebuilding text from word boxes
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m1s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m24s
The parsers are written against `pdftotext -layout`, and every policy-ocr
fixture is a verbatim excerpt of it. The runtime does not use it: it reads
`-bbox-layout` and rebuilds the page from word boxes, and that rebuild
collapsed all white space — no blank lines between blocks, one space
between columns. White space is the only thing marking a cell boundary on
these borderless forms, so the fixtures could not see any of it.

What it cost, on the GMX PVL especificación and the ANA driver's policy:

- `espectBlock` walks a wrapped cell until a blank line. With no blank
  line it ran to the end of the page, so the insured's name came back as
  the entire first page of the specification.
- `INSURED\s{2,}` and its siblings matched nothing, and the phone that
  shares the name cell rode along with it ("PAMELA DENISE WAGONER
  Ph.3102001538"), which matches no customer.
- `parseAnaDriverCoverages` splits SUM INSURED from PREMIUM by the
  header's own column offsets. Without offsets, every premium was filed
  as a sum insured.

So `toVisualRows` now emits a blank line where the reader sees one (a
vertical gap over 1.6 line heights — the two populations measure 0.3-1.1
and 2.1+, so the threshold sits in empty space) and pads each word to its
own column, using one space wherever words merely follow each other so
rounding drift cannot sprinkle false cell boundaries through prose.

Two independent guards, so neither failure can come back silently: the
ANA phone splits on a single space, and the especificación's cell walk is
capped at the one wrap the longest cell on that document actually uses.

Verified against the real PDFs: the especificación reads "EMMER .
KATHLEEN" with all 18 coverages named (they were "(sin nombre)"), and the
ten born-digital gas invoices parse byte-identically to before. The
scanned statements are untouched — they come through tesseract, not this
path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 11:56:23 -07:00
gitea-actions ca6432efc8 chore(release): v1.0.19
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m38s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m16s
Deploy on tag / Deploy to galactus (push) Successful in 37s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.19.
2026-08-15 08:33:01 +00:00
rmancinasandClaude Opus 5 022d1935ad feat(policy-ocr): set policyTypeId and insuranceProviderId on confirm
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m0s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m14s
The BACKLOG claimed this was blocked on incomplete `policy_types` rows.
Querying the dev database says otherwise: AUTO (1316 policies) and LICENCIAS
(306) are both live and healthy, so ANA's two faces were never blocked at
all. Three separate things had been conflated.

What the parser now emits is a NAME, not an id -- it is a pure function over
text and must not reach for the database:

  ANA AUTOMOBILE        -> AUTO
  ANA DRIVER'S POLICY   -> LICENCIAS
  GMX (both documents)  -> MULT

`resolveLookups()` turns that into a foreign key at confirm, and does the
same for the carrier off the parser's provider code. It resolves, never
creates: a missing `policy_types` row means a human deleted it, and silently
recreating it would undo that with no record. An explicit `policyTypeId` /
`insuranceProviderId` on the confirm payload always wins.

GMX is MULT rather than INCENDIO because the caratula's own header reads
"Multiple Policy / Home" and the especificación is "PVL Hogar" -- one product,
two artifacts. MULT is the live row carrying 769 of them; INCENDIO is fire-only
and no policy in the book has ever used it.

The parser's provider code is not the carrier's row name, so PROVIDER_ROW_NAME
maps ANA onto "ANA SEGUROS", which is where the office's 738 ANA policies
already are.

--- the actual defect underneath -----------------------------------------

`policies.policyTypeId`, `policies.insuranceProviderId` and
`claims.adjusterId` are all ON DELETE SET NULL, and the lookups screen deleted
unconditionally. So deleting a lookup row returned 200 and silently blanked
the field on every row referencing it -- no error, nothing in the UI. That is
how M_EMPR disappeared and left 5 policies with no ramo, found months later
only by querying.

All three deletes now refuse while the row is in use, naming it and the count
("El tipo de póliza «M_EMPR» está en uso por 5 póliza(s)"). The schema-level
`onDelete: Restrict` the spec once recommended is deliberately not used: a raw
FK error is not something the operator can act on.

`20260815160000_policy_type_repair` cleans up what already happened:

  - restores M_EMPR and re-points its 5 policies, scoped to
    `policyTypeId IS NULL AND legacySourceTable = 'm_empr'` so it can never
    claim a policy blanked for some other reason
  - merges the duplicate "ANA" carrier (1 policy) into "ANA SEGUROS" (738).
    OCR is about to start assigning the carrier automatically and two rows
    would keep splitting the book. Written as joins, not subqueries, so both
    statements are no-ops when either row is absent -- a subquery form would
    resolve to NULL and blank the carrier off every ANA policy.
  - does NOT restore INCENDIO. It is the other row the migration would have
    produced, but the legacy INCENDIO table has 1 row that never loaded, so
    the type has zero policies and restoring it would only put a dead option
    in the type picker.

Verified by running the repair against the real broken dev data inside a
transaction and rolling back: 5 orphans -> 0, ANA/ANA SEGUROS -> one row with
739, and a second run in the same transaction changes nothing. The DDL half
matches `prisma migrate diff` exactly.

186 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 01:28:26 -07:00
rmancinasandClaude Opus 5 5a277f4885 feat(deploy): apply migrations at api container start
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m3s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m32s
`prisma migrate deploy` ran in one place only: a workflow step on the Gitea
runner, which has to reach the target host's MySQL on 3306 directly. Two
paths went around it:

  - `skip_migrate=true`, the documented answer for when the runner cannot
    reach 3306, left the schema a release behind with nothing to catch it.
    The mismatch surfaced later as a column-not-found at runtime rather than
    as a failed deploy.
  - A container brought back by `restart: unless-stopped` after a host
    reboot, or a stack re-applied by hand in Portainer, never runs the
    workflow at all.

docker/api-entrypoint.sh becomes the api image's ENTRYPOINT: migrate, then
exec node. If the migration fails the container exits non-zero and the API
never listens — serving against a schema that does not match the code is
worse than being down, because the failures are partial and silent (a write
to a missing column breaks one feature while the rest looks healthy).

This does not replace the workflow step and is not a substitute for it. That
step still runs FIRST, while the old code is serving, which is the order
expand/contract migrations are designed around. `migrate deploy` is
idempotent, so on the normal path the container's run is a no-op query.

Behaviour:

  RUN_MIGRATIONS=false     skip and start anyway; plumbed through both app
                           stack files, for a schema moved by hand
  DATABASE_URL unset       refuse to start, and say why
  P1001 (unreachable)      retry, default 20 x 3s -- a cold db container, and
                           galactus's MagicDNS lookup right after a reboot
  anything else            exit at once; retrying a broken migration only
                           delays the same error. P3005 prints the
                           `migrate resolve --applied 0000_init` hint the
                           workflow step already printed.

Only P1001 retries, so a genuinely broken migration is not buried under a
minute of noise.

Both stacks are replicas: 1 and must stay so for an unrelated reason (the
servicios email sweep has no DB lock). The old comment claiming migrations
must not run per-container because "N replicas would race" is dropped: they
would not corrupt anything, since Prisma takes a database advisory lock and
the losers find nothing pending -- they would only each pay the wait.

The prisma CLI is already in the runtime layer (the image copies
/repo/node_modules wholesale), but which of the two plausible .bin paths
carries it is an implementation detail of pnpm's hoisted linker, so the
entrypoint accepts either and the Dockerfile asserts one exists at BUILD
time. A missing CLI breaks the image build, not a production boot.

Verified by running the entrypoint against stubbed prisma binaries: clean
run, P3005, P1001-to-exhaustion, P1001-then-recovery, RUN_MIGRATIONS=false,
missing DATABASE_URL, missing CLI.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 01:17:01 -07:00
rmancinasandClaude Opus 5 d645ba51d3 feat(policy-ocr): read A.N.A. Seguros' two policy faces
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m39s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m21s
A.N.A. is the Rosarito office's tourist auto book and the second carrier
the policy OCR pipeline reads. It ships two unrelated faces, and the split
is different from GMX's: GMX ships two documents about one policy, A.N.A.
ships two products.

  AUTOMOBILE (SPECIAL POLICY FOR TOURISTS)  insures a car; vehicle table,
                                            9 numbered sections, one
                                            LIMIT OF LIABILITY column
  DRIVER'S POLICY (the office: "licencia")  insures up to 5 named drivers;
                                            no vehicle at all, 6 unnumbered
                                            sections in a different order,
                                            SUM INSURED + PREMIUM columns

The four automobile products the office sells (amplia / responsabilidad
civil, annual / by-the-day) are the same layout with different numbers, so
they get one parser rather than four.

These are born-digital portal PDFs, so pdftotext -layout returns exact
columns and the driver's-policy parser uses that: its two value columns
print the same shape (100,000.00 usd. / 18.70 usd.) with no per-row label,
so horizontal position is the only thing separating them. The split comes
from the header's own offsets, not a constant, because they shift between
products; when it can't be read every amount is reported as a sum insured
and the reviewer is told, rather than half the premiums being filed as
coverage limits.

Three things the layout will punish a naive read for:

- Each PDF prints its face two or three times (ORIGINAL, AGENT COPY, then
  a receipt and three travel cards) and the pipeline concatenates every
  page before parsing. The coverage walk is bounded to the first copy and
  the driver list to the first POLICY HOLDER block. Unbounded, the licencia
  returns the same person three times, which reads as a three-driver policy
  rather than as a bug.
- The money row is read positionally off its header. An unused DISCOUNT
  prints as a bare "-", so "find the six amounts" shifts every value one
  column left on a discounted policy.
- Two five-digit numbers sit in the header band and only one is the agent
  clave; the agent's street address is "BENITO JUAREZ 25 No.50 INT 38",
  three lines above the No. cell holding the policy number.

Sections 6-8 print a PREMIUM where the others print a limit, so
ParsedCoverage gains an optional `premium` (GMX never fills it) and the
review table a column: $40 is what legal aid cost, not a $40 liability
limit. Exclusions follow the GMX rule and go in the risk label with a null
amount -- which matters more here, since a responsabilidad-civil policy
prints 0.00 for material damage and the two are identical on the page.

Also in this change:

- coveragePeriodDays is parsed and written. A.N.A. sells 3- and 4-day
  policies; Policy.coveragePeriodDays defaults to 365, so a weekend policy
  left at the default sits in the renewals window a year out. Derived from
  the dates, cross-checked against the printed DAYS cell, disagreement
  noted not resolved.
- Vehicles and named drivers are parsed, shown read-only in review, and
  written as Vehicle / InsuredDriver rows on confirm, skipping any already
  on the policy (VIN then plate; licence then name). The case that forces
  the skip is confirming a renewal onto an existing policy. Nothing is ever
  updated or deleted -- a changed plate lands as a second row for a human.
- Batch.provider is set from what the parsers actually claimed instead of
  being hardcoded "GMX", so a mixed upload is labelled as mixed and the
  header can never contradict its own documents. PolicyDocument.documentType
  follows the same rule (was hardcoded GMX_POLICY).
- matchNote becomes TEXT. It was VARCHAR(191) and the note trail was sliced
  to 190 chars, which cut the tail notes -- the "could not read X" ones.
- The policy detail page renders an array coveragesJson as a table. Both
  shapes have always been possible there, but the object renderer was the
  only one, so an OCR-confirmed policy showed a row per array index labelled
  "0", "1", "2" with [object Object] as the value. ANA makes that routine.

GMX is untouched behaviourally; its two parsers now spread a shared empty
base instead of listing every null field. 29 new parser cases against
verbatim pdftotext output of three real ANA PDFs, 53 in the suite.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 01:09:16 -07:00
rmancinasandClaude Opus 5 14c4d44acb docs(policy-ocr): vigencia/agente/prima are keyed in by hand on the PVL layout
Build and Push Images / Build jorgecuadros-web (push) Successful in 3m4s
Build and Push Images / Build jorgecuadros-api (push) Successful in 3m14s
Confirmed with Luz, who handles GMX policies at the office: the three
fields the especificación does not carry are entered manually. The review
screen already supports it — all three are editable and `postPremium`
enables off the typed premium, so no code change was needed.

The parser's note said "esos datos están en la carátula de la póliza",
which now sends the reviewer looking for the wrong document. It says
"captúrelos a mano" instead, and names the consequence of leaving the
vigencia blank: `Policy.policyTo` is nullable and the renewals window
filters `policyTo: { gte, lte }`, so a policy confirmed without one never
matches and never gets a renewal notice — silently, permanently, with
nothing downstream erroring.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 00:20:09 -07:00
rmancinasandClaude Opus 5 45be0ad77d feat(policy-ocr): read GMX's Spanish PVL especificación layout
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m16s
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m16s
GMX ships two unrelated documents for the same policy and the office
downloads both from the same portal. The parser only knew the English
caratula, so a `…-CondicionesParticulares.pdf` parsed to an almost
entirely empty row — including the policy number, which the matcher needs.

`parseGmx` becomes a dispatcher over `parseGmxCaratula` (unchanged
behaviour) and the new `parseGmxEspecificacion`. Both still report
`provider: "GMX"`: the matcher keys on the policy number alone and must
not care which artifact was uploaded.

The especificación has no tables. Coverages are found by anchoring on
`Límite … Responsabilidad:` and walking backwards for the heading, where a
heading is a short line *preceded by a blank line* — length alone cannot
tell one from the wrapped tail of the paragraph above it, and without that
condition coverages get named after the last word of the preceding prose.

Also fixed, both pre-existing:

- The policy number's group widths are not the same across the two
  families (`007-037-…-0000-02` vs `07-037-…-00000-01`). The pinned-width
  regex is replaced by a shape, so both read.
- The caratula's ZIP fallback pushed a note saying it had read the ZIP
  from the address, then never assigned it.

Verified against the full ten-page real document: all 17 coverages,
amounts, deductibles and the excluded earthquake section match what is
printed. 24 parser tests (was 8), four of them regressions for ways this
layout can silently attach the *wrong* value rather than none.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 00:10:26 -07:00
rmancinas 7be897ef2b chore(release): v1.0.18
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m59s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m35s
Deploy on tag / Deploy to galactus (push) Successful in 1m11s
2026-08-11 13:17:45 -07:00
rmancinasandClaude Opus 5 cf40cd22ef fix(deploy): stop pinning API_ORIGIN, and probe one origin not the list
Two leftovers from making the browser derive the API origin. The stack env still
injected API_ORIGIN from a repo secret, which pinned the origin again on every
deploy and would have re-broken an https front door with mixed active content.
Drop it from both env_data blocks; the secret stays, now purely as the URL the
verify step probes.

That verify step was also about to break on its own: WEB_ORIGIN is a
comma-separated CORS list now, and `curl "$WEB_ORIGIN/version"` on a list
retries thirty times and fails a deploy whose app is perfectly healthy. Probe
the first entry, so keep the runner-reachable origin first in the secret.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 13:14:04 -07:00
rmancinas 3b02c6944f chore(release): v1.0.17
Build and Push Images / Build jorgecuadros-api (push) Successful in 3m16s
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m31s
Deploy on tag / Deploy to galactus (push) Successful in 5m33s
2026-08-11 12:56:52 -07:00
rmancinasandClaude Opus 5 683fd37b08 docs(deploy): stop documenting API_ORIGIN as required
The swarm stack still hard-failed on an unset API_ORIGIN, and both the env
template and the README told the reader to pin it — the exact habit the derived
origin was meant to end. Make it an optional override everywhere, and say that
WEB_ORIGIN is now a list.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 12:56:50 -07:00
rmancinasandClaude Opus 5 14c6183aa2 feat(deploy): derive the API origin from the page, not from a pinned env var
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m55s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m35s
The browser hard-required API_ORIGIN, so every move of the server — tailnet
today, the 192.168.1.0 office LAN later, a temporary demo domain in between —
meant editing the deploy env and redeploying. Worse, an http:// API origin on a
page served over TLS is blocked outright as mixed active content, which is what
broke the demo on https://jorgecuadros.freakma.com.

The browser now derives the origin from window.location the way a PHP app
would: same host on port 3001 over plain HTTP, or the same-origin /api path
under https (the reverse proxy strips the prefix). API_ORIGIN survives as an
optional override for a deployment that genuinely splits the two hosts, and SSR
still reads process.env because a derived origin is browser-only.

WEB_ORIGIN becomes a comma-separated list to match: one deployment is now
reached under several origins, and a credentialed fetch from an unlisted one
gets no CORS headers and fails.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 12:55:32 -07:00
gitea-actions 5352d49ecf chore(release): v1.0.16
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m57s
Build and Push Images / Build jorgecuadros-api (push) Successful in 4m4s
Deploy on tag / Deploy to galactus (push) Successful in 1m57s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.16.
2026-08-07 05:06:42 +00:00
rmancinasandClaude Opus 5 2169ffa78d feat(ops): verify the replica against the master, not just its own status
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m51s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m9s
Every field the replication card showed was self-reported by the replica, and
the two most reassuring ones lie in the same failure. Seconds_Behind_Source
reads 0 when the I/O thread is disconnected — with no incoming event there is
nothing to measure staleness against — and Replica_IO_Running only says the
network thread is alive, not that it is receiving.

Two checks that ask the master instead:

- GTID drift, folded into the polled status. GTID_SUBTRACT(master, replica)
  counts transactions the master executed that the replica has not, so a silent
  disconnect shows up as a number that climbs instead of a lag that stays 0.
  It also isolates transactions carried under the replica's OWN server UUID —
  writes that exist nowhere on the master. There are currently 518 of them,
  residue of the seed dump load; inert while log_replica_updates is off, and a
  real divergence the day anyone promotes that box.

- A full row-by-row comparison behind a button, over the eight tables
  my.jorgecuadros.com reads. GTIDs prove the replica applied everything the
  master sent; they say nothing about rows changed here by another route, which
  is the one failure the rest of the card cannot see.

The comparison hashes CONVERT(col USING binary), not CAST(col AS CHAR). CAST
transcodes into the connection character set, and the two servers do not agree
on it: the client inside the master's container negotiates latin1, the replica's
utf8mb4. Every accented character in a Mexican name, street or note then hashes
differently and the tool reports a permanent mismatch on exactly the tables that
hold free text. Caught by building it and running it — customers.name gave
3344437324815 against 3339150372121 under CAST, and 3339150372121 on both under
CONVERT. All eight tables now match byte for byte.

Verify is POST and audited despite reading nothing: it full-scans both servers,
so a prefetch or a refresh must not be able to start one.

Tests cover the GTID interval arithmetic, which is inclusive at both ends and
easy to get wrong by one in the direction that hides a gap.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 21:38:22 -07:00
rmancinasandClaude Opus 5 17d83291c3 feat(migration): refuse a full re-import that would delete native rows
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m51s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m42s
A full run_all.py pass truncates and rebuilds every table it owns from the
Access extract. That was harmless while the platform was a read-only mirror --
every row came from the extract, so wiping and rebuilding lost nothing. It
stopped being harmless once the platform started minting rows Access has never
heard of: allocated portal NUMids, customers created in the staff UI,
OCR-captured policies, app-booked ledger rows, uploaded documents.

REIMPORT is a button in /operaciones, so that was one click away.

native_guard.py counts what only exists here and exits 3; run_all.py runs it
before the first truncate and stops. Detecting an allocated NUMid needs the
staged Parquet -- the customer holds an ordinary-looking (utilities, DATGRAL,
'1172') ref, so "customer has no refs" cannot see it and only comparing against
the extract can. Missing staging is therefore treated as blocking rather than
as "nothing to protect".

The guard does not teach full mode to preserve anything: --sync already upserts
legacy rows against the existing refs and leaves the rest alone, and rebuilding
that inside full mode would re-implement it. --force-full (checkbox in the
REIMPORT confirm, recorded in the audit log) deletes them deliberately.

Verified against dev: clean before, exit 3 listing utilities/1172 with a
synthetic ref present, clean again after removing it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 21:01:01 -07:00
rmancinasandClaude Opus 5 6a97242fc3 feat(customers): allocate portal NUMids, with an audit for reusable ones
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m59s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m13s
Customers created in the staff UI had no NUMid and so could not log in to
my.jorgecuadros.com at all: the id is a CustomerLegacyRef row, not a column,
and create() deliberately writes none.

Allocation is a staff action (POST /customers/:id/portal-access, MANAGER)
rather than part of create, because insurance is expected to move to the
platform before utilities and an insurance-only customer has no reason to
spend a utilities id.

The audit that decides which ids are reusable took three passes. "Owns no
rows" matches nobody -- migration gave all 1,171 NUMids a property and a
transaction. "No transaction in N years" also matches nobody -- every customer
carries a synthetic Jan-1 opening-balance row, so everyone looks active this
year. Subtracting that row is what makes dormancy measurable, and it leaves 4
never-used ids and 10 dormant ones on dev. Two further traps are encoded in the
queries: insurance/DATGRAL is a separate id space that reuses the sourceTable
name and runs past 4,000, and ACCOUNT CANCELED is a transaction line type, not
an account state -- all 8 customers carrying it have current-year activity.

Recycling ships switched off (numid.recycleEmpty, default false). Every
reusable id still exists in Access DATGRAL, and a --sync run reassigns refs
with ON DUPLICATE KEY UPDATE customerId, so an id recycled before the utilities
cutover is silently handed back to its Access owner.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 20:04:13 -07:00
gitea-actions 7981c715ce chore(release): v1.0.15
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m49s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m6s
Deploy on tag / Deploy to galactus (push) Successful in 23s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.15.
2026-08-05 07:37:01 +00:00
rmancinasandClaude Opus 5 d173c9e9a0 fix(billing): stop double-counting history a BALANCE FORWARD already carries
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m52s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m4s
BALANCE FORWARD rows are not movements. Access materialized one per
customer per year, dated Jan 1, holding the closing balance of everything
before it — that is what let the portal keep each year in its own table and
still show a correct running balance from one year's rows.

The platform imported those rows AND the real pre-cutover history they
summarize, and every balance aggregate summed the lot. NUMid 501 read
-10,469.29 on the receivables worklist against -14,065.29 on the customer's
own statement and on the legacy portal; the gap was two cash receipts from
2009 and 2012 that the 2026 opening balance had already absorbed.

The scale settles what it is: summed the old way the whole book came to
+20,605,447.86 MXN — the office owing its customers 20.6 million pesos.
Floored, it is -56,855.90. A receivables ledger cannot be 20M in credit.

Adds BALANCE_FLOOR_JOIN + NOT_SUPERSEDED and applies them to balances()
(page and count queries, which must agree), to stats()'s per-currency and
per-domain figures, and to the owing/in-credit split. The four stats()
aggregates moved from Prisma groupBy to raw SQL because groupBy cannot
express a per-customer floor.

statement() takes the same floor as a scalar, which is also what stops
FEE ANUAL and fee15 leaking in. Those are not in
STATEMENT_EXCLUDED_SOURCE_TABLES — that list reproduces legacy's
DATOS2-only `datosfreak` — and they were putting 2,092 pre-cutover fee rows
across 1,062 customers into the statement, skewing it by -5,129,764 against
the number those customers have been quoted for years. Dating rather than
source is the right test: a fee row *after* the opening balance is a real
charge and still counts.

movements() is deliberately left alone. It is a browser over captured rows
— "how much water did we capture in April" — and staff need the historical
rows visible, so it keeps totalling everything, the same asymmetry
NOT_OUTSTANDING already has.

stats() now separates the two questions it was mixing: movements,
ledgerCustomers, crossLineCustomers and the date range stay unfloored
inventory; everything under byCurrency/byDomain is a balance and is floored.

BillingService had no tests. Adds 13 covering the floor's failure modes —
it fails silently, so MIN-vs-MAX, `>` vs `>=`, the NULL branch for customers
with no opening balance, and the join/predicate alias pairing are each
pinned, plus the 501 arithmetic as a regression.

Verified through the real service against the live ledger: balances() and
statement() both return -14,065.29 for 501, matching the portal.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 00:34:53 -07:00
rmancinasandClaude Opus 5 e9a5ee9e90 fix(migration): carry NOPAGO into transactions.outstanding
datosfreak's NOPAGO is the legacy "still owed" flag, and the website reads
it directly — account.statement.php splits the statement on NOPAGO = 0 vs
NOPAGO = 1 and renders the latter as "Outstanding Bills Requiring
Attention". transform_transactions.py hardcoded 0, so all 40,421 rows came
across settled and that section renders empty for anyone served off the
platform. Not a missing column: a missing section, with no error.

Only the three DATOS2-shaped tables carry the flag (76 rows set in datos2,
0 in FEE ANUAL and fee15); the EFECTIVO/FM3 cash streams have no such
column and keep the 0 default. Sync mode gets outstanding=VALUES(...) too,
so an additive sync corrects rows already loaded rather than leaving them
settled forever.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 23:53:45 -07:00
gitea-actions 458e67340c chore(release): v1.0.14
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m46s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m45s
Deploy on tag / Deploy to galactus (push) Successful in 1m5s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.14.
2026-08-05 05:33:49 +00:00
rmancinasandClaude Opus 5 b12382b436 ci: move the galactus deploy chain to a tag-only workflow
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m39s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m17s
The deploy was a job inside build.yml gated by
`if: startsWith(github.ref, 'refs/tags/v')`. Gitea draws every job into the
run graph before it evaluates that `if`, so an ordinary push to master showed
a pending "Deploy to galactus" — indistinguishable from prod being about to be
redeployed off an unreleased commit, and the only safe reaction is to cancel
the run, which takes the images down with it.

The gate itself was never wrong (no deploy-galactus run has ever been created
from a branch ref), but a guarantee you cannot see is not much of a guarantee.
`on: push: tags: ["v*"]` in a workflow of its own makes it structural: the
deploy cannot appear on a master build because the workflow does not exist
there.

It replaces `needs: build` by polling the Actions API for the build.yml run at
this tag and requiring it green, so both images are still known to be in the
registry before anything is pulled. AUTO_DEPLOY_GALACTUS still cuts the chain.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 22:28:02 -07:00
gitea-actions 2620559975 chore(release): v1.0.13
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m1s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m15s
Build and Push Images / Deploy to galactus (push) Successful in 8s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.13.
2026-08-05 05:23:57 +00:00
rmancinasandClaude Opus 5 ed19f51a52 fix(ops): re-stage before an additive sync
Build and Push Images / Deploy to galactus (push) Canceled after 0s
Build and Push Images / Build jorgecuadros-web (push) Canceled after 1m26s
Build and Push Images / Build jorgecuadros-api (push) Canceled after 1m27s
SYNC ran `run_all.py --sync` without `--stage`, so it depended on staged
Parquet under migration/output. That directory is part of the image, not a
volume, so any redeploy wiped it and the job died on the first transform:

    FileNotFoundError: '/repo/migration/output/stg_utilities/datgral.parquet'

Re-staging is also what makes the job's own label true — without it a sync
would replay whatever upload staged last, not the files currently in the
ingest folder.

Staging now counts as a numbered step when it runs, so the Operaciones
progress bar moves during the slowest phase instead of sitting empty.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 22:20:30 -07:00
rmancinasandClaude Opus 5 e85db73dbc ci: deploy to galactus automatically when a tag build goes green
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m55s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m23s
Build and Push Images / Deploy to galactus (push) Skipped
Cutting a release then had one manual step left: watch build.yml and
dispatch "Deploy to galactus" by hand with the version. Chain it.

build.yml gains a `deploy` job, `needs: build` and gated on
refs/tags/v*, that dispatches deploy-galactus.yml against the tag with
tag=<version> scope=app bootstrap=false skip_migrate=false. `needs`
waits for both matrix legs, so api and web are both in the registry
before prod pulls either — deploy-galactus.yml only pulls, and a
half-pushed pair leaves prod running one new image and one old one.

A dispatch rather than a `workflow_run:` trigger (which Gitea has
supported since 1.24) because deploy-galactus.yml reads
github.event.inputs.* in ten places; under workflow_run all of them are
empty strings, so the deploy would run with no tag. The dispatch keeps
that workflow's contract intact and keeps it hand-runnable, which is
how rollbacks work.

The dispatch is confirmed the same way release.yml confirms the build
started: snapshot the existing deploy-galactus run ids first, then
require a new one to appear. An accepted dispatch that creates no run
is the failure mode that cost v1.0.3 its images, and a plain "is there
a deploy run" check would be satisfied by the previous release.

Kill switch: repo variable AUTO_DEPLOY_GALACTUS=false prints the manual
command instead of deploying. Needs the existing RELEASE_TOKEN secret;
preflight fails loudly and names the manual command if it is unset.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 21:48:23 -07:00
gitea-actions d38bbc52ec chore(release): v1.0.12
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m30s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m48s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.12.
2026-08-05 04:40:07 +00:00
rmancinasandClaude Opus 5 fe761e119e feat(ops): show relay apply progress on the replication card
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m0s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m15s
Seconds_Behind_Source cannot answer "is it moving?". While the SQL thread
works through one large transaction the lag counter holds still — often at
0 — even though the replica is not caught up. The relay backlog does move,
and it comes out of the SHOW REPLICA STATUS the panel already runs, so this
costs no extra query and no connection to the source.

Adds applyProgress(), which reads Source_Log_File / Read_Source_Log_Pos vs
Relay_Source_Log_File / Exec_Source_Log_Pos and reports the fetched-but-not-
applied byte delta plus a percentage. Both positions are source binlog
coordinates, so they are only comparable while the two threads are on the
same file; across files the delta is meaningless (positions restart at ~4 in
each new file) and is reported as null rather than as a huge negative number.

The percentage deliberately stops at 99.99 while any backlog remains —
binlog positions are large enough that a real backlog of a few KB rounds to
100% and would render a lagging replica as caught up.

Not folded into `healthy`: a non-zero backlog is the normal state of a
working replica between fetch and apply, so alarming on it would cry wolf.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 21:35:39 -07:00
gitea-actions 4a929f7e7c chore(release): v1.0.11
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m1s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m44s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.11.
2026-08-04 00:55:49 +00:00
rmancinasandClaude Opus 5 66d0d071b0 feat(ops): show step progress for reimport and sync jobs
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m42s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m18s
A REIMPORT takes ~110 seconds and, until now, showed only a scrolling log —
there was no way to tell "halfway" from "wedged", which mattered the day one
actually did wedge.

run_all.py emits "[paso i/N] name" before each step and the API derives
progress from the job log. Emitting the marker from the Python rather than
having the UI count STEPS itself means the step count is stated in exactly
one place; adding a step cannot desync the display. Progress is derived, not
stored, for the same reason: the log is already the record of what happened,
and a separate counter could contradict it, which is precisely the confusion
a progress display exists to remove.

While RUNNING, step i is IN PROGRESS rather than finished, so only i-1 count
as done. Counting i would show 100% while the final step was still working —
and the final step (blob_extract) is the slowest, so the bar would sit at
"100%" for the longest stretch of the job.

BACKUP and RESTORE are a single mysqldump with no steps and deliberately
render no bar; a fabricated percentage would be worse than none. The safety
backup that precedes a REIMPORT is likewise named explicitly instead of
showing 0%, which reads as stuck.

Pinned by job-progress.spec.ts, including the literal line run_all.py emits,
so a change to the Python format fails a test rather than silently blanking
the panel.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 17:53:10 -07:00
gitea-actions f269dc8bfa chore(release): v1.0.10
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m28s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m55s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.10.
2026-08-04 00:43:02 +00:00
rmancinasandClaude Opus 5 eef9a5f4c8 fix(ops): stop the replica field parser reading the next line
Build and Push Images / Build jorgecuadros-api (push) Canceled after 51s
Build and Push Images / Build jorgecuadros-web (push) Canceled after 50s
The panel reported "Error SQL: Replicate_Ignore_Server_Ids:" against a
replica that was healthy — both threads running, zero lag.

`\s` matches newlines in JavaScript, so `^\s*NAME:\s*(.*)$` let the `\s*`
after the colon walk past an EMPTY field's line break and capture the
following line. Last_SQL_Error is blank on a healthy replica and
Replicate_Ignore_Server_Ids happens to be printed immediately after it, so
the blank error field returned the next field's name as its value. Every
empty field was affected; the visible damage was that a healthy replica
rendered as broken, which is the worst direction for a health panel to fail.

Fixed with `[^\S\n]` — horizontal whitespace only — on both sides of the
field name.

Extracted as replicaField() and pinned by replication.spec.ts against the
verbatim output of the live replica, keeping the empty Last_SQL_Error
adjacent to Replicate_Ignore_Server_Ids because that exact adjacency is what
broke. Also covers the literal "NULL" lag surviving as a distinct value from
empty, and a field name that is a suffix of another (Last_Error vs
Last_SQL_Error) not matching the wrong line.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 17:40:49 -07:00
gitea-actions 7f1bfe906e chore(release): v1.0.9
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m4s
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m13s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.9.
2026-08-04 00:31:58 +00:00
rmancinasandClaude Opus 5 7797c45e9f fix(ops): fail orphaned RUNNING jobs at startup
Build and Push Images / Build jorgecuadros-api (push) Canceled after 1m21s
Build and Push Images / Build jorgecuadros-web (push) Canceled after 1m21s
Ops jobs run as a child of the API process, so no job can outlive it. When a
deploy landed 110 seconds into a REIMPORT, the child died and nothing was
left to finalize the row — it stayed RUNNING forever. Because startJob()
refuses to start while any RUNNING row exists, that one interrupted job
wedged the panel permanently with no way out from the UI; recovering it took
a manual UPDATE against the production database.

A fresh boot is proof that nothing survived, so this is unconditional rather
than filtered on age: "started recently" does not imply "still alive" here.

Rows are updated one at a time rather than with updateMany so the reason can
be APPENDED to the log. A job whose log simply stops mid-step with no
explanation is what made the first occurrence hard to diagnose.

Failure to reconcile is logged and swallowed: a wedged panel is bad, an API
that will not boot is worse.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 17:22:15 -07:00
rmancinasandClaude Opus 5 dac1f1982f feat(ops): show read-replica health on the Operaciones screen
my.jorgecuadros.com serves customer balances from the Oracle VPS replica.
A replica whose SQL thread has stopped does not error — it keeps answering,
with data frozen at the moment it stopped — so nothing on the customer site
looks wrong and the only signal is a customer complaining about a stale
balance. This puts the failure somewhere a human sees it.

Deliberately does not trust the two fields an operator reaches for first.
Replica_IO_Running reports Yes while the SQL thread is stopped, because the
network thread keeps downloading binlog it will never apply; verified by
stopping SQL_THREAD and watching IO stay Yes. Seconds_Behind_Source reads
NULL whenever EITHER thread is down, so the card renders "sin dato" rather
than "0 s" — showing zero there would report an outage as perfect health.
The problem string is resolved most-specific-first for the same reason.

Shells out to the mysql client because the API has no MySQL driver and the
image already ships one. --ssl is required (the replica sets
require_secure_transport); --ssl-verify-server-cert=0 is deliberate and is
NOT the trade-off the website makes: this hop never leaves Tailscale and the
replica's firewall admits only this host, so WireGuard authenticates the
peer, whereas the DreamHost leg crosses the public internet and pins the CA.

The account behind it holds REPLICATION CLIENT and nothing else — it cannot
read a single row. REPLICA_DB_* unset is a supported state and renders "no
configurada", which is correct in dev and before cutover.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 17:22:15 -07:00
rmancinasandClaude Opus 5 1d689d8f46 fix(migration): label EFECTIVO cash rows as CASH DEPOSIT
The EFECTIVO ledgers have no type column — in Access the transaction type
is implied by which table a row lives in — so unlike DATOS2 there was no
string to map and typeId came out NULL on all 13,496 rows.

That is not just a blank label. handleGetAccountDetails in
my.jorgecuadros.com identifies payments by matching TYPEOFTRX against
('PAYMENT THANK YOU', 'PAYPAL', 'CASH DEPOSIT', 'CHECK DEPOSIT') to reset
the running balance in mode=current. An unlabelled payment is not
recognised, so the balance silently diverges from legacy — 285 rows across
129 customers in the current year alone.

"CASH DEPOSIT" is measured, not chosen: matching the unlabelled rows to the
live site on (NUMid, date, amount) resolves unanimously to that label —
66/66 in the current-year `datosfreak` and 100/100 in the prior-year `2025`
table, the only two periods the site allowlists.

The FM3 fee streams (EFECTIVO FM3 627, CHEQUE FM3 157) have the same
missing-type problem and are deliberately left NULL: every row predates both
exposed periods, so nothing can be matched against a legacy label and none
can reach a customer. Guessing "CHECK DEPOSIT" there would feed the
payment-detection list on no evidence.

type_id_for(None) returns None, so call sites without a label are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 17:10:12 -07:00
rmancinasandClaude Opus 5 7bec2a13d8 feat(deploy): add a replication health check for the read replica
my.jorgecuadros.com reads customer data from the Oracle VPS replica, and a
replica that has silently stopped applying serves stale balances rather
than erroring — so "is it replicating" needed an answer that is not a
human squinting at SHOW REPLICA STATUS.

Runs entirely against the replica over ssh, so it needs no credentials for
the galactus master, and exits non-zero on failure so it can be driven from
cron or a monitor.

It deliberately does not trust the two fields an operator reaches for first.
Replica_IO_Running reports Yes while the SQL thread is stopped, because the
network thread is still downloading binlog it will never apply — verified by
stopping SQL_THREAD and watching IO stay Yes. Seconds_Behind_Source reads 0
both when there is nothing to apply and when nothing is connected. The
trustworthy signal is GTID_SUBTRACT(Retrieved, Executed): binlog fetched but
not applied.

NULL lag means either thread is down, so it is reported as "not applying"
rather than blamed on a specific thread — the thread fields above already
say which, and guessing there produced a wrong diagnosis.

Uses sed rather than `head -n1`; on this machine `head` resolves to LWP's
HTTP head(1), which mangles the pipeline instead of failing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 17:07:33 -07:00
gitea-actions 127eaa9689 chore(release): v1.0.8
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m42s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m14s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.8.
2026-08-03 20:47:55 +00:00
rmancinasandClaude Opus 5 7226772c22 fix(migration): recover transaction type labels and minimum balance
Two fields the customer-facing site reads were being dropped on the way in
from Access.

transform_transactions.py mapped DATOS2's type string through the Access
`TYPE OF TRX` table and stored NULL on a miss. That table is a stale
pick-list rather than a constraint — staff free-text straight into DATOS2 —
so 78 distinct values covering 3,939 rows never resolved, including
BALANCE FORWARD (1,188) and ANNUAL FEE (1,116). Nothing else on
`transactions` carries the type text, so those rows lost their label
outright and rendered blank. Now mints a type_transactions row from the
literal string when the lookup lacks it; nameEs stays NULL since only the
lookup has translations.

transform_customers.py never carried DATGRAL.TIPO, leaving
customers.minimumBalance empty on every row despite the column existing.
TIPO is the minimum-balance threshold (100/200/300/500; 1,017 of 1,172
customers carry one), not an account type as the name suggests — the
customer app shows it as `minBalance`. Added to the insert list and to the
ON DUPLICATE KEY UPDATE clause, without which --sync would silently skip
it on existing rows.

Both land on the next `run_all.py --sync` reload.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 12:30:03 -07:00
gitea-actions 2fa12890f5 chore(release): v1.0.7
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m8s
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m53s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.7.
2026-08-02 20:29:10 +00:00
rmancinasandClaude Opus 5 e77e5546d8 docs: SES secrets created, ship blocker cleared
Five documents asserted the SES_* secrets were unset in Gitea. They now
exist, so all five are corrected rather than leaving the claim to rot in
whichever one a reader opens first.

Replaces the blocker with the two things creating the secrets does NOT
establish, since both fail in ways that look identical to a missing
config: SES_FROM must be a verified identity in SES_REGION, and the
account must be out of the SES sandbox — in sandbox SES only delivers to
verified recipients, so a sweep across 815 policyholders would fail
almost every send while the configuration reads as correct.

Recommends running the first sweep with debug on, which diverts every
recipient and, on the pólizas side, leaves the avisos pending so a failed
test consumes nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 13:25:44 -07:00
rmancinasandClaude Opus 5 3e12597204 docs: add BACKLOG.md, one list of everything outstanding
Open work was spread across six documents: PLAN's per-step status,
RESUME §6, two specs' collected open questions, and the "Not built"
sections of the two OCR docs. Nothing tracked the two live data defects
except a paragraph inside INSURANCE_FEATURES_SPEC, and nothing at all
recorded that master is 14 commits and 5 migrations past the last tag.

Compiled by reading those six, then checking each claim against the code
and the dev database rather than trusting the prose — which is how the
dead-table finding surfaced and how both insurance defects were confirmed
still open.

Leads with the ship blocker: SES_* is unset in Gitea while the pólizas
sweep defaults to enabled at 06:00, so deploying current master gives a
nightly sweep that fails every run. Set the secrets or disable the
schedule before cutting v1.0.7.

Findings not previously written down anywhere:

- policy_types still holds only AUTO/LICENCIAS/MULT and 5 policies still
  have a NULL policyTypeId; policyTypeId is still `String?` with Prisma's
  default SetNull, so the spec's recommended Restrict was never applied.
- EmailTemplate / EmailCampaign / EmailLog have zero references in
  apps/api/src or apps/web/src. Scaffolded for step 10's "email
  campaigns"; notificaciones shipped against email_notification_log
  instead. Either wire them or drop them.
- Customer.customerNumber does not exist, so recycling is not merely
  unbuilt but unstarted at the schema level.

Linked from PLAN.md and README so it is findable from either entry point.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 13:14:25 -07:00
rmancinasandClaude Opus 5 ec139737be docs: as-built reference for the statement OCR capture
Gives receipt capture the same treatment policy OCR just got: a doc that
records what is in the code, separate from the spec that records what was
designed. RECEIPT_CAPTURE_SPEC.md §2 had accumulated three BUILT notes
totalling ~120 lines of findings, which is the right place for the
evidence but the wrong place to look up how the matcher picks a column.

docs/STATEMENT_OCR.md covers the pipeline, the OCR seam and its
text-layer-first rule, all eight parsers and the ordering constraints
between them, the matcher's two governing rules and the scopedRefField
table, confirm-through-BillingService, the learning write-back, and the
API surface.

Weight goes to the things that are load-bearing and invisible from the
code shape: brand detection must run to completion before layout because
Tijuana bills predial and zona federal off the same treasury header;
scopedRefField is exported because three call sites must agree or a
reference gets learned into a column nothing searches; FEDERAL_ZONE's
accountNumber holds a peso amount, so it fails the null-guards as well
as the lookup; a misread `$` is the dangerous failure, not a missing one.

Also records that CFE/CESPT/Telnor have no unit suite — they predate the
gas/predial extension and were only verified end to end.

Cross-linked from the spec, POLICY_OCR.md, PLAN.md, README and RESUME.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 13:02:14 -07:00
rmancinasandClaude Opus 5 872a661051 docs: document policy OCR capture, the feature no spec proposed
Policy OCR shipped 2026-08-01 (5e9cb12) and was documented nowhere. It is
not in INSURANCE_FEATURES_SPEC.md because it did not come from that
meeting — it came out of building the utility statement OCR pipeline in
RECEIPT_CAPTURE_SPEC.md §2 and noticing the same shape fits carrier
policy PDFs. A reader had no way to find that lineage.

New docs/POLICY_OCR.md covers it end to end, with weight on the three
things that are not obvious from the statement side:

- **One PDF = one policy.** Statements arrive bundled one customer per
  page, so there a page is a document. A GMX certificate is one policy
  across two pages, so the pages are concatenated and the parser runs
  once per file — which is why `pageNumber` is a file ordinal and
  `storageKey` is the source PDF, not a page image.
- **The GMX certificate carries no premium at all** — it lives on a
  separate recibo PDF. Hence the null-preserving confirm and the
  double-gated ledger write.
- **OcrModule was extracted out of StatementsModule to make this
  possible**, and that was blocking rather than cosmetic.

Cross-referenced from RECEIPT_CAPTURE_SPEC.md §2 (where it came from),
INSURANCE_FEATURES_SPEC.md (which never proposed it, and whose §4 carrier
API it partly overlaps), PLAN.md step 11, README and RESUME.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 12:56:25 -07:00
rmancinasandClaude Opus 5 6331481f82 docs: record notificaciones as built, flags global, schedules editable
The docs still described the state before the last five commits: the
insurance spec called for a `@Cron` literal and a manual mark-as-sent
mutation, PLAN.md had step 12 as "NOT STARTED", and README's module and
route lists predated seven modules.

- MASS_EMAIL_NOTIFICATIONS.md: new "Send flags", "API surface" and
  "Scheduled runs" sections; "Cron (future)" removed — it exists. The
  flags table says which flags apply where, and why a debug renewal send
  must skip both the RenewalNotice row and `lastSuccessfulAt`.
- INSURANCE_FEATURES_SPEC.md: §1 BUILT note listing the three places the
  build diverged from the spec; §1.1 and §1.4 marked superseded in place
  rather than deleted, so the reasoning stays readable.
- PLAN.md: step 12 renewal emails DONE with the divergences; status
  paragraph rewritten.
- README.md: current module/route lists, plus a "Scheduled jobs" section —
  a reader cloning this repo had no way to know the API sends mail on a
  timer.
- DEPLOY_AND_MIGRATIONS.md: the cadence lives in app_settings and survives
  an image rollback, and the servicios sweep has no multi-replica lock.
- RESUME.md: session record for the whole notificaciones arc.
- RENEWAL_NOTICES.md: pointer that this is the legacy record, not what
  shipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 12:43:19 -07:00
rmancinasandClaude Opus 5 89611da202 feat(notificaciones): global send flags + editable schedules
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m47s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m3s
The "Flags del envío" panel lived inside the Servicios tab and only
governed the four bulk jobs. The pólizas half had no debug at all, so
there was no way to test a renewal notice without mailing a real
customer. The panel now lives in the /notificaciones shell above the
tabs and both halves read it.

`debug` on the renewal path diverts to the same override inbox as the
servicios jobs and deliberately does NOT write the `RenewalNotice` row
or advance the sweep's `lastSuccessfulAt` — the customer was not
notified, so nothing may gate the letter they are still owed.
`ignoreDayRestriction` and `useEmailLimit` stay estado-de-cuenta-only
and are labelled as such.

Both automatic sweeps are now operator-editable. The renewal cadence
was a `@Cron("0 6 * * *")` literal and servicios had no automatic run
at all; both now resolve through `NotificationScheduleService`, which
stores the cadence in `app_settings` and reinstalls the cron job on
save — no redeploy, no restart. Defaults preserve current behaviour:
pólizas 06:00 daily, servicios off. A scheduled run never inherits the
UI flags; it always sends for real.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 12:23:19 -07:00
rmancinasandClaude Opus 5 a491ef3eed feat(notificaciones): edit summary recipients in the UI
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m32s
Build and Push Images / Build jorgecuadros-api (push) Successful in 3m28s
NOTIFICATION_ADMIN_EMAILS made "add Beto to the summaries" a redeploy —
the wrong unit of work for a list that changes when office staff change.

Adds `app_settings`, a key/value table for the configuration staff must
be able to change without a deploy, and `SettingsService`, which resolves
every key db -> env -> default and reports which of the three a value
came from. That ladder is what makes the move safe: a deployment behaves
exactly as before until somebody saves in the UI, and the screen can say
"this is still coming from the deployment" rather than implying somebody
chose it.

- new ability `setting:manage` (ADMIN) — deliberately above
  `notification:send`, since redirecting the audit summaries is how
  someone would quietly stop them being read
- GET/PUT /notifications/settings/admin-emails; read is open to any
  logged-in user so the UI can display the list, write is gated
- resolved per job, not cached at boot, or we would reintroduce exactly
  the restart-to-apply behaviour being removed
- a saved empty list means "nobody" and does NOT fall through to the env,
  or clearing the field would keep mailing the people just removed

Credentials stay in env — see the model doc for where the line is drawn.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 11:58:42 -07:00
rmancinasandClaude Opus 5 f4b92fa7a5 fix(deploy): pass SES config through to the app stack
The stack env is assembled from Gitea repo secrets by the deploy
workflows' `env_data` block — there is no .env file on the host for the
app stack. SES was in neither, so `MailService` came up unconfigured on
every deployment and, with NODE_ENV=production killing the stdout dev
fallback, every notification and renewal aviso failed.

Wire SES_REGION / SES_FROM / SES_FROM_NAME / SES_ACCESS_KEY /
SES_SECRET_KEY / SES_CONFIGURATION_SET / NOTIFICATION_ADMIN_EMAILS
through both galactus and cubex. No `_GALACTUS` suffix: one SES identity
serves every deployment.

Kept out of the required-secrets preflight — mail is not needed to boot,
and failing a deploy over it would be wrong. Preflight warns instead,
since the failure is otherwise invisible until someone clicks "Ejecutar".

Also corrects the comments added in the previous commit, which claimed
these belonged in a host env file rather than in CI secrets.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 10:59:45 -07:00
rmancinasandClaude Opus 5 33833c3af9 feat(notificaciones): one send log across servicios and pólizas
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m30s
Build and Push Images / Build jorgecuadros-api (push) Failing after 3h13m42s
Renewal avisos left behind only a `RenewalNotice` row, whose sole job is
gating: a row with `sentAt` drops the policy off the pending list. It
cannot represent a failed send or a customer with no address, so the
Pólizas tab had no "Registro de envíos" to show and a sent notice simply
vanished from the list.

Renewals now write `email_notification_log` — the same table the four
bulk jobs write — as `RENEWAL_NOTICE` / `POLICIES`, with rows for
failures and no-email skips too. `RenewalNotice` keeps its gating role
unchanged; the two are complementary, not redundant.

- extend `EmailNotificationType` (+RENEWAL_NOTICE) and
  `EmailNotificationServicio` (+POLICIES); `level` now carries the aviso
  generation on renewal rows, so every reader must branch on the type
  first (`notificationLevelLabel()` is the one place that lives)
- backfill emailed notices (`channel = 'EMAIL'`) into the log; MAIL-channel
  rows are legacy printed letters and are deliberately left out
- extract `NotificationLogService`/`NotificationLogModule` as the single
  writer, so a feature that sends mail records it without pulling the
  bulk-job pipelines into its module
- `GET /notifications/log` and `/stats` take a comma-separated `servicio`
  list; each tab reads its own slice. This also fixes the "Omitidos"
  view, which mapped to no filter at all and showed every row
- share one `NotificationLogPanel` between both tabs
- pass SES_* / NOTIFICATION_ADMIN_EMAILS through the galactus compose,
  which was missing them entirely — mail is runtime config, not a CI
  secret, and the prod image sets NODE_ENV=production so a blank config
  fails loudly instead of falling back to stdout

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 03:01:03 -07:00
rmancinas c0cc0d2ac2 feat(renovaciones): send renewal notices from the list, drop manual marking
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m45s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m24s
The Pólizas tab now sends. Each pending row gets an "Enviar aviso" button
backed by POST /renewals/send, which renders, mails and records the notice
through the same path the daily sweep uses — so a hand-sent letter is
marked exactly like a swept one and drops off the pending list.

Sending is now the only way a notice gets marked as sent. Remove the
manual "Marcar impreso" / "Marcar EMAIL" buttons and the endpoint behind
them (POST /policies/:id/renewal-notices, PoliciesService.markRenewalNotice,
MarkRenewalNoticeDto): they wrote a sentAt with no mail behind it, which
let the list claim a customer was notified when nothing was sent.

sendOne refuses a generation that already has a sentAt (409) so a double
click cannot mail the customer twice, and 400s when the customer has no
email on file. Sweep and single send share the new deliver() helper.
2026-08-02 02:40:38 -07:00
rmancinas 53a5fe8076 feat(notificaciones): ejecutar todos for servicios jobs
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m53s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m13s
Add POST /notifications/run-all: runs the four notification jobs
(outstanding, payment confirmation, account status, trust confirmation)
sequentially with one shared set of flags from "Flags del envío".

Sequential rather than parallel — the jobs share the SES transport and
account status can self-throttle via useEmailLimit. A job that throws is
captured and the sweep continues, so one bad query cannot swallow the
other three envíos; the aggregate response carries per-job results plus
summed sent/skipped/failed and an errors count.

Audited as a single notification.run-all.run entry so one staff click is
one audit row. UI adds the button to the flags card, with a confirm when
debug is off, and a per-job summary in "Última respuesta".
2026-08-02 02:32:13 -07:00
rmancinasandClaude Opus 5 0332292ae9 fix(notificaciones): merge renewals into one screen, fix MailModule DI
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m42s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m22s
MailModule's provider used a `useFactory` with no `inject`, so the factory
received `undefined` and `new MailService(config)` threw on `config.get`,
taking the whole API down at boot. The module also wasn't actually
`@Global()` even though both NotificationsModule and RenewalsModule inject
MailService without importing it — that would have failed next. Replaced the
factory with a plain provider (ConfigModule is already `isGlobal`) and marked
the module global.

On the web side, mass email and renewal notices were two menu entries doing
the same job — telling a customer something by email. They are now two tabs
of `/notificaciones` (Servicios and Pólizas), following the Captura pattern:
`/renovaciones` still resolves, opening the same screen on its Pólizas tab so
existing bookmarks keep working.

The notifications page was also the last screen written in raw inline styles,
with blue buttons and filter pills that appear nowhere else in the app. It now
uses the shared design system: btn-primary/btn-outline, the seg segmented
control, card, tx-table, pager, and the servicios/fideicomiso badges.

Two supporting fixes found on the way: NOTIFICATION_STATUS_COLORS hardcoded
hex instead of the theme's positive/negative/muted vars, and `.small` was
referenced in 19 places across the app but never defined in globals.css.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 02:21:51 -07:00
rmancinas ec0e9c2a5d Merge branch 'massive-email-notification' into master
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m50s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m7s
# Conflicts:
#	.env.example
#	apps/api/src/app.module.ts
2026-08-02 02:05:36 -07:00
rmancinas a52e59cbc5 feat(notificaciones): mass email notifications over SES
Replaces the four legacy PHP scripts under email.notifications/send*.php
with a single NestJS module. Four jobs (outstanding payments, payment
confirmations, account-status alerts with day-of-week gates, trust
payment confirmations) share one MailService modelled on StorageService:
env-driven SES client, null fallback in dev with console logging, refuses
to send in production when unconfigured.

Schema adds email_notification_log (every attempt, sent/failed/skipped)
and account_status_history (one row per threshold hit, Job 3). Enums
encode the legacy wire shape so external log scrapers keep parsing
notificationType keys verbatim.

Web adds /notificaciones with four trigger cards, a flags panel, and a
paginated log browser. New notification:send ability gates all four
endpoints at MANAGER, matching the renewal:send trust tier.
2026-08-02 02:04:14 -07:00
rmancinas 87d8743251 feat(renovaciones): renewal notification emails over SES
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m48s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m4s
INSURANCE_FEATURES_SPEC §1. The office printed and mailed renewal letters
from the legacy CONTROL <ramo> RENEW[2/3] paper log; 91% of policyholders
have an email on file, so send the notice instead and keep the paper log
as the fallback.

A daily cron (06:00 America/Tijuana) sweeps three generations off
policyTo — 30 and 15 days before expiry, 7 days after — sends each
through SES, and upserts RenewalNotice by [policyId, generation] so a
policy is never notified twice for the same milestone. RenewalNotice now
records providerMessageId, so a later bounce or complaint webhook can be
traced back to the row that sent it.

- customers.emailOptOut excludes a customer from every sweep; editable
  from the customer form
- scheduled_job_states holds the sweep's lock and last successful run;
  the window is widened to cover days the job did not run, so a weekend
  outage does not silently drop a generation
- SES unconfigured is not an error outside production — messages are
  logged and skipped, so dev and CI never send
- /renovaciones (renewal:send, MANAGER+) lists what is pending per
  generation, runs the sweep by hand, and marks a notice sent by mail
  for the customers with no email
- POST /policies/:id/renewal-notices records that manual mark
- the aviso-renovacion report and the emails now share one projection
  (reports/renewal-letter.ts) instead of two copies of the mapping
2026-08-02 02:00:02 -07:00
rmancinas 3125b52057 feat(ocr): discard abandoned capture batches
A bad scan, the wrong PDFs or a duplicate upload used to leave a batch
sitting in READY_FOR_REVIEW forever, because the only exits were confirm
(posts to the books) or rejecting every page one at a time. Add a
DISCARDED terminal status to both OCR domains and a single endpoint per
domain that rejects every page still pending in one shot.

Discarding is refused once anything has landed: statements once a page is
POSTED, policies once a page is APPLIED. Those batches did real work and
have to be settled page by page.

- POST /statements/batches/:id/discard
- POST /policy-ocr/batches/:id/discard
- shared DiscardBatchCard on both review screens, gated the same way
2026-08-02 02:00:02 -07:00
gitea-actions 905fa31e47 chore(release): v1.0.6
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m37s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m11s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.6.
2026-08-02 02:06:12 +00:00
rmancinasandClaude Opus 5 5e9cb12fba feat(polizas): OCR capture for insurance policy PDFs
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m43s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m0s
Mirrors the utility statement intake on the insurance side: a policy_ocr
batch/document pair of tables, a GMX parser, a matcher keyed on
Policy.policyNumber, and a "Captura" screen under /polizas that proposes
policy -> customer for staff to confirm.

Lifts the OCR seam out of StatementsModule into its own OcrModule so
PolicyOcrModule can inject OCR_PROVIDER without taking on the rest of
the statement pipeline; StatementsModule now imports it and binds
nothing itself.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:07:29 -07:00
rmancinasandClaude Opus 5 5bce0e4c94 feat(recibos): OCR capture for zona federal (ZOFEMAT Tijuana)
Adds the ZONA FEDERAL TIJUANA parser to the statement intake, measured
against 8 pages of real "Zona Federal Marítimo Terrestre" receipts — the
federal maritime-zone occupancy fee the municipality bills on beachfront
lots. Provider read on 8/8, amount on 8/8 (each verified against the
paper), concession clave on 6/8, period on 8/8, deadline on 2/8.

Four things the corpus forced:

- Tijuana bills predial and zona federal from the same treasury: same
  header, same Paseo del Centenario address, same ATB-541201 RFC. Every
  predial discriminator matches a zona federal page too, so whichever
  rule is asked first wins it. The only words exclusive to this layout
  are "Marítimo Terrestre", so its brand rule is asked ahead of all
  three predial ones — and its structural rule, anchored on the stub's
  "Derechos de ocupación", ahead of theirs.

- FEDERAL_ZONE.accountNumber is an amount, not a reference. It holds
  DATMEX.zfed, whose 77 values include 246.06, 2369.09, 22653.94 and a
  negative -1679, while the concession claves these receipts are keyed
  by appear nowhere in the database. Matching on that column could never
  hit — and because every row already has a value, the `[field]: null`
  guards on learnAccountRefs and on the review blank-service fill would
  never fire either, so every page would return to the queue every
  bimester forever. The clave moves to meterNumber, joining gas and
  Tijuana predial, and the first confirm teaches the match.

- The payable figure is not the printed subtotal. The municipality
  rounds to whole pesos and prints the difference on its own "Ajuste Ley
  Hacienda Mpal" line (-$0.05 against a 591.05 subtotal, $0.21 against
  2,872.79). The "Total a pagar" box carrying the rounded figure sits on
  a grey fill and OCR'd on 1 of 8 pages; the SubTotal row read on 8 of
  8. So the amount is the rounded subtotal, cross-checked against the
  printed box wherever it survives — where it did, it agreed.

- The clave is 2 digits, a letter and 3 digits (12-T -012), not the
  cadastral shape, and the letter is kept as printed: toDigits maps D to
  0, which turns a real 14-D -014 into 140014. It is printed twice,
  which rescued a page whose heading was struck through by the office's
  own highlighter — the failure mode behind both missing claves.

Deriving the deadline from the bimester is deliberately not attempted:
it is the 17th of the month after the bimester closes on a current bill,
but four of these eight are late (a $1,000 Multa) and print a
recalculated date, so a derived date would be wrong on exactly the pages
a human most wants to see.

Re-ran the earlier corpora (25 pages: predial Tijuana/Rosarito/Ensenada,
CFE, CESPT, Telnor) through detection to confirm the new rules steal
nothing — all 25 still read as their original provider, including the
five Tijuana predial pages that share the RFC.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:07:21 -07:00
rmancinasandClaude Opus 5 d6501f1d74 feat(recibos): OCR capture for gas butano and municipal predial
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m50s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m8s
Adds four parsers to the statement intake — GAS TIJUANA plus one per
municipality, because Tijuana, Rosarito and Ensenada issue three
completely different predial documents — and a text-layer fast path for
the born-digital invoices the gas company sends.

Measured against a new corpus of 14 documents / 29 pages: provider read
on 29/29, amount on 26/29, and 21/29 auto-matched against the dev
database (22/29 identified). The eight review cases are all legitimate.

Five things the corpus forced:

- Not every statement is a scan. The gas invoices are born-digital CFDIs
  whose text layer is exact; rasterising them only loses information (one
  sample turned `MEDIDOR: VM01014426` into `ar (LTR): 014420`). The new
  `OcrProvider.textPages` reads the embedded layer via `pdftotext
  -bbox-layout` — same poppler package as `pdftoppm`, so no new
  dependency — and OCR stays the fallback for real scans. Poppler's own
  `<line>` grouping follows text flow rather than the page, so words are
  regrouped by vertical position; without that, a two-column header
  leaves every label separated from the value printed beside it.

- The clave catastral is not two letters and six digits. Position three
  is a letter in 15 of the 932 stored claves, and digitising the whole
  tail mapped a real `MMB01041` to a nonexistent `MM801041`.

- Tijuana predial prints no clave at all. Its only identifier is an
  8-digit municipal account carried in a 32-digit payment barcode, which
  the legacy database never held, so it goes in `meterNumber` alongside
  gas — `accountNumber` holds `DATMEX.predial`, which is not a
  per-property key and must not be overwritten. Those pages start cold
  and are taught by the first confirm.

- On Rosarito and Ensenada the clave is the primary key, not a fallback:
  those receipts print nothing else, so a unique hit auto-matches. On a
  utility bill that merely happens to print one it stays a review hint.

- A misread `$` is the dangerous failure. An Ensenada receipt for
  $2,203.00 OCR'd as `82,203.00`, which would post a charge 37x too large
  and look ordinary in the ledger. Predial amounts now require a literal
  `$` and a page that cannot produce one goes to review.

The scoped match field is now one exported function rather than three
copies of `kind === "GAS" ? ... : ...`, since the lookup, the
blank-service fill and the confirm write-back have to agree or a
reference gets learned into a column nothing searches.

First tests in this package: 23 specs over the parsers and the text-layer
reader, every fixture a verbatim OCR excerpt from a real receipt. Adds
the jest config they need and a build tsconfig so they stay out of dist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 12:52:20 -07:00
rmancinas 216309190c feat(recibos): live OCR progress bar on review page
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m51s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m35s
Backend already returns per-status counts via byStatus; render a real
progress bar (X% / N de M / en cola) while PENDING_OCR pages remain,
using the existing progress-track CSS. Falls back to indeterminate
when no docs have been reported yet.
2026-08-01 02:43:19 -07:00
rmancinasandClaude Opus 5 e589bda28b ci(build): skip the redundant master build when a release is cut
Build and Push Images / Build jorgecuadros-api (push) Canceled after 1m10s
Build and Push Images / Build jorgecuadros-web (push) Canceled after 1m8s
release.yml pushes the release commit and its tag in a single `git push`,
so Gitea created two build.yml runs for the same commit. Only the tag run
matters: it emits the X.Y.Z and X.Y image tags, and since it is the same
commit it publishes `latest` and `sha-<short>` as well. The master run was
pure duplicate work that had to be waited out or cancelled by hand.

Guard the build job with an `if` that skips a branch push whose head commit
message starts with `chore(release):`. Ordinary pushes to master are
unaffected, and tag pushes and manual dispatches always build.

The skipped master run keeps the release commit's sha, which would have let
release.yml's "Verify build.yml started" check go green on it alone even if
the tag run were never created — the exact failure that check exists to
catch. It now also requires the run's ref to be the tag, falling back to the
sha match only when the API reports no ref.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 02:28:09 -07:00
gitea-actions 98f7aa8a2d chore(release): v1.0.5
Build and Push Images / Build jorgecuadros-api (push) Canceled after 0s
Build and Push Images / Build jorgecuadros-web (push) Canceled after 0s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.5.
2026-08-01 09:22:56 +00:00
rmancinasandClaude Opus 5 898cf48c80 fix(migration): re-import died on the last step because blob_extract required deploy/.env.prod
Every transform resolves its target through dbenv.database_url(), which lets a
DATABASE_URL in the process environment win — that is how the API container
drives a re-import against its own database with no deploy/ directory present.
blob_extract.py was the one step that bypassed it and called load_env()
directly for the MinIO credentials, so the "Operaciones" re-import loaded all
the data and then exited 1 on:

  missing /repo/deploy/.env.prod — deploy the 'prod' DB stack and write its
  .env first

Give the S3 settings the same resolution as the DB URL: load_env() now returns
{} for an absent file, and setting()/require() layer the process environment on
top of it. blob_extract reads S3_ENDPOINT / S3_BUCKET and accepts either
S3_ACCESS_KEY/S3_SECRET_KEY or MINIO_ROOT_USER/MINIO_ROOT_PASSWORD, matching
the fallback order in storage.service.ts and the vars the api service already
sets in deploy/galactus/jorgecuadros-app.compose.yml. A genuinely missing
setting still fails fast, now naming the variable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 02:21:38 -07:00
gitea-actions 70fe425043 chore(release): v1.0.4
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m2s
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m55s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.4.
2026-08-01 09:09:15 +00:00
rmancinasandClaude Opus 5 567b033c46 fix(docker): re-import failed because the Access CLI tools were never installed
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m37s
Build and Push Images / Build jorgecuadros-api (push) Successful in 3m22s
The API image installed Alpine's `mdbtools` package, which ships only the
shared library. The command-line tools that migration/extract.py actually
shells out to -- `mdb-tables` and `mdb-export` -- are in the separate
`mdbtools-utils` subpackage, so the build succeeded and the re-import in the
"Operaciones" admin panel failed at run time with:

    RuntimeError: mdbtools not found on PATH (need mdb-tables and mdb-export)

Install `mdbtools-utils` instead; it pulls the library in as a dependency.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 02:07:22 -07:00
rmancinasandClaude Opus 5 1934470d53 ci(release): dispatch the fallback build with a fully qualified ref
Gitea's workflow dispatch API 404s on a bare `v1.0.3` and accepts only
`refs/tags/v1.0.3`, so the fallback added in fdbe9fd would have failed
the release instead of rescuing it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 02:00:23 -07:00
rmancinasandClaude Opus 5 fdbe9fdb88 ci(release): fail the release when the build never starts
Gitea creates workflow runs from the post-receive hook. When that hook
errors the refs still land, git prints `remote: error: Internal Server
Error` and exits 0 — a post-receive failure does not fail a push. v1.0.3
was cut exactly that way: tag pushed, no build run created, no images
published, and the release step green. It surfaced two steps later as a
404 when the deploy tried to pull 1.0.3.

Capture the push output and warn on `remote: error`, then verify a
build.yml run actually exists for the new commit, dispatching it against
the tag if not. Fail the release if that does not take either, so a
release that publishes nothing is red instead of green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 01:58:10 -07:00
gitea-actions e082113640 chore(release): v1.0.3
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.3.
2026-08-01 08:45:42 +00:00
rmancinasandClaude Opus 5 860d483bad fix(ops): backup failed on the MariaDB client shipped in the API image
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m7s
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m7s
Every backup on galactus died with:

  mysqldump: unknown variable 'set-gtid-purged=OFF'
  respaldo incompleto eliminado

Alpine's mysql-client is MariaDB's, so `mysqldump` inside the API
container is a shim over `mariadb-dump`, which has no --set-gtid-purged.
That took out BACKUP and, because they take a safety dump first, SYNC
and REIMPORT too.

Probe `mysqldump --help` and pass the flag only when it is advertised,
calling `mariadb-dump` directly otherwise — MariaDB writes no GTID state
unless asked with --gtid, so there is nothing to suppress. Testing
whether mariadb-dump merely exists would be wrong: on a host carrying
both clients it would shadow a perfectly good MySQL mysqldump.

The probe uses a command substitution rather than `--help | grep -q`
because PIPEFAIL is in effect for these commands and grep closing the
pipe early would report a supported flag as unsupported.

pre-migrate-backup.mjs is unaffected — it dumps from a real mysql:8.4
image, not from the API container.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 01:43:25 -07:00
rmancinasandClaude Opus 5 783ec83464 feat(ops): show upload percent, speed and ETA for ingest files
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m46s
Build and Push Images / Build jorgecuadros-api (push) Successful in 3m3s
The ingest upload used fetch(), which cannot report request-body
progress, so the only feedback was a static "Cargando…" label — no way
to tell a stalled 2 GB upload from a working one.

Switch uploadFile() to XMLHttpRequest and expose an optional onProgress
callback reporting loaded/total bytes, a smoothed transfer rate and a
remaining-time estimate. The Operaciones ingest table renders a progress
bar row under the file being uploaded. Once the bytes are all sent the
server still has to write the file, so that tail reads "Procesando en el
servidor…" rather than parking at 100%.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 01:38:22 -07:00
gitea-actions a9b4aab7ec chore(release): v1.0.2
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m4s
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m12s
Cut by rmancinas via the "Cut release" workflow. Pushing the tag triggers build.yml; deploy separately with tag=1.0.2.
2026-08-01 08:23:07 +00:00
rmancinasandClaude Opus 5 a8afd87c3f ci: add a "Cut release" dispatch workflow
Stamps every package.json, commits chore(release): vX.Y.Z, tags and pushes
both refs in one dispatch — patch/minor/major, or an explicit number. Cutting
a release from a laptop is how a manifest bump gets forgotten or a tag lands
on an unpushed commit; the only input here is the number.

Guards: refuses a version that already exists as a tag (releases are
immutable), a no-op bump, a leading `v`, and a malformed number. Checkout is
full-depth because the duplicate-tag check is meaningless against a shallow
clone.

Pushes with a RELEASE_TOKEN PAT rather than the built-in Actions token —
whether a push made with that token re-triggers build.yml depends on the Gitea
version, and a release that quietly publishes no images is worse than one that
fails outright.

Builds and deploys stay separate: the tag push triggers build.yml, and
deploying remains a deliberate dispatch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 01:21:13 -07:00
rmancinasandClaude Opus 5 b59abda895 feat(captura): fold recibo OCR into Captura as an automatic mode
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m46s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m17s
Scanning a stack of bills and keying them in are the same daily job, ending
in the same ledger path, so OCR intake becomes a mode of the capture screen
instead of a second menu entry:

- components/Captura.tsx holds the mode switch; the manual check form moves
  verbatim to components/ManualCheckCapture.tsx and the OCR intake to
  components/StatementIntake.tsx.
- /estado-cuenta/lote opens on manual, /recibos on automatic — both render
  Captura, so batch-review links and old bookmarks still land right.
- Nav drops "Recibos (OCR)"; "Captura" covers both, with a NavLink.aliases
  field so /recibos still highlights it.

Also fixes the "El almacenamiento de documentos no está configurado" failure
staff hit on upload. Uploading with no object storage configured used to
succeed, then die on the first put minutes later, leaving a FAILED batch
whose only explanation was that string. createBatch now refuses up front,
GET /statements/status reports storageAvailable alongside ocrAvailable, and
the intake tab explains the situation instead of offering an upload that
cannot work. S3_* documented in .env.example (deploy stacks already set it).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 01:04:09 -07:00
rmancinasandClaude Opus 5 4d5008b545 feat(statements): OCR intake for scanned utility bills
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m41s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m18s
Staff key 300+ utility statements per company per month by hand. This adds
the ingest -> split -> OCR -> match -> review pipeline that proposes customer
and amount per page instead (RECEIPT_CAPTURE_SPEC §2), posting through the
existing BillingService.createBatch seam with source=OCR and a per-document
captureRef so machine and hand capture share one write path and audit trail.

Everything was designed against 10 real scanned statements (46 pages of CFE,
CESPT and Telnor bills) rather than from the sample-free spec. The scans have
no text layer at all — they are camera images — so OCR is mandatory, and they
arrive bundled one customer per page. Measured on those pages the parser
identifies the provider 46/46 and reads an account reference 43/46; against
the dev database that is 39/46 (85%) exact auto-match, 40/46 identified, with
the rest genuine review cases. That closes the OCR-provider question in favour
of self-hosted Tesseract: it clears the bar for a queue where a human confirms
every row, and OcrProvider keeps a managed API a one-line swap.

The samples corrected three things the spec had wrong or unknown:

- Clave catastral is NOT predial. DATMEX.clave (934 rows) is what CESPT and
  predial bills print; DATMEX.predial, which PROPERTY_TAX.accountNumber holds,
  has 663 distinct values across 1135 rows and appears on no statement. The
  clave now lives on Property.cadastralKey as the matcher's secondary key;
  predial is left untouched. This had been blocking predial matching.
- Gas was recoverable: 160 of 334 DATMEX.gas values are real account numbers
  (the rest are ESTACIONARIO/CILINDRO descriptors), now in GAS.meterNumber.
- Phone is one billed line per property (534/18/1 across phone1/2/3), so the
  new TELEPHONE ServiceKind backfills from phone1 only, not three rows.

Matching is scoped to one column per service kind and never reads the customer
name — a CESPT receipt prints ARNAIZ ROSAS ELSA AURORA for an account this
office holds under CATT, RANDY, because the printed name is the registrant,
not the current owner. Where a provider prints a payment barcode it beats the
printed label (one CFE label OCR'd a digit too many while its barcode was
correct) and the two cross-check, with disagreement forcing review.

Confirming a document whose service had no reference writes it back, so gas
and any other cold start is a one-time cost rather than a permanent queue.

Verified end to end against the live dev API and MinIO: real scans uploaded
over HTTP, matched, confirmed against a check, and the resulting rows checked
in MySQL (negative amounts, captureSource=OCR, concept derived from the batch
kind, captureRef linking back to each page). Re-confirming a posted batch is
refused. Test data was removed afterwards.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 00:42:35 -07:00
rmancinasandClaude Opus 5 121952fdc1 chore(release): v1.0.1
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m6s
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m22s
v1.0.0's images were built from db2bd54, which predates the full-hash
footer. Deploying 1.0.0 would ship the abbreviated footer, so the version
that actually goes to galactus is 1.0.1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 16:38:06 -07:00
rmancinasandClaude Opus 5 15f533b984 fix(web): show the full commit hash in the build footer
The footer abbreviated to 7 characters, so the line read
"v master · 19f0319". That line exists to be pasted into `git show` or
compared against a registry tag, and an abbreviation makes both a manual
step — while the full 40-char value was already baked into the image
(build.yml passes `github.sha` whole, and /version returns it untouched).

`shortSha` had no other caller, so it goes with it.

The span gets `overflow-wrap: anywhere` and `min-width: 0`: hex offers no
break opportunity, and the api/web mismatch branch renders two of these
hashes side by side, which would otherwise push a phone into horizontal
scroll. Measured at a simulated 360px with both hashes present — the span
wraps, and documentElement.scrollWidth stays equal to clientWidth.

Verified in the browser against the dev database: footer renders
"v1.0.0 · db2bd54c0ffee1234567890abcdef0123456789a", hash length 40.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 16:37:48 -07:00
rmancinasandClaude Opus 5 db2bd545a1 chore(release): v1.0.0
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m42s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m28s
Every manifest still read 0.1.0 while the deployed images were addressed by
the moving tag `latest`. That combination is what hid the stale-image bug:
a checkout could not be placed against a running container, and `latest`
silently kept serving two-commit-old web code through a green deploy.

Tagging v1.0.0 makes docker/metadata-action publish immutable `1.0.0` and
`1.0` image tags, so deploys can name a version instead of a moving target.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 16:21:35 -07:00
rmancinasandClaude Opus 5 30dfc7dc3e fix(ops): run backups as an admin login, and stop recording failed dumps as good
The Operaciones panel (backup, restore, sync, re-import) shelled out to
mysqldump as the application user, parsed straight out of DATABASE_URL.
`--single-transaction` issues FLUSH TABLES, which needs the global RELOAD
privilege, and the app user is granted only ALL ON jorgecuadros.* plus
USAGE ON *.*. BACKUP failed outright; SYNC and REIMPORT failed with it,
since both take a safety backup first.

An admin credential is now supplied out of band via OPS_DB_ADMIN_USER /
OPS_DB_ADMIN_PASSWORD, mirroring what deploy/scripts/pre-migrate-backup.mjs
already does, rather than permanently elevating the user the API serves
requests as. Host, port and database still come from DATABASE_URL, so the
override can only change who logs in, never which server. Unset, it falls
back to the DATABASE_URL credentials and warns — local development is
unaffected.

Two defects in the dumps themselves, both shared with the deploy backup
before it was rewritten:

- No --set-gtid-purged=OFF. The production server is the replication source
  with GTID on, so every dump embedded SET @@GLOBAL.GTID_PURGED and was
  unrestorable onto the server it came from — the one thing the restore
  screen is for.

- The pipeline's exit status was gzip's, and gzip succeeded. A mysqldump
  that died on its first statement left a small, perfectly valid archive
  that the job recorded as SUCCESS and the restore screen listed as an
  ordinary restore point. Dumps now run under `set -o pipefail`, assert a
  CREATE TABLE count, and delete their own output on failure. Verified with
  a stubbed mysqldump: a failing dump exits 1, surfaces the real error,
  removes the partial file, and — critically — stops SYNC/REIMPORT before
  the ETL touches anything.

Restores gained pipefail too: a corrupt archive made gunzip fail while
mysql, fed a truncated stream, could still exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 16:21:00 -07:00
rmancinasandClaude Opus 5 d5ebb86cae fix(deploy): dump from a dedicated container as root, and prove the dump is real
The pre-migrate backup ran INSIDE the API container, which made it depend on
that image's toolchain — and deadlocked: the running image shipped a MySQL
client that could not authenticate, so the backup failed, which blocked the
very deploy that would have replaced the broken image. A backup must not depend
on the thing being deployed.

The dump now runs in a throwaway container built from mysql:8.4 with the API's
backup volume mounted. The volume name is discovered from the API container's
mounts, so the file still lands where the Operaciones restore screen looks. As
a container rather than an exec, its logs can simply be read — no more failures
reported as a bare exit code. The image is pulled if the host lacks it, since a
scope:app deploy never touches the db stack.

Three further defects found while verifying, none of which would have surfaced
without dumping against the real database:

- The dump now runs as root. mysqldump --single-transaction issues FLUSH
  TABLES, needing the global RELOAD privilege; the MySQL image grants the
  application user only ALL ON `<db>`.*, and --skip-lock-tables does not avoid
  it. Elevating the app's own runtime user would have been the worse trade.

- --set-gtid-purged=OFF. galactus is the replication SOURCE with GTID on, so a
  default dump embeds SET @@GLOBAL.GTID_PURGED and is unrestorable onto the
  server it came from. Verified: 0 GTID_PURGED lines in the output.

- Verification was too weak to be worth having. `test -s` plus `gzip -t` passes
  on a 372-byte gzip containing no tables, which is exactly what a dump that
  died on its first statement produces. It now asserts a CREATE TABLE count and
  logs it. A failed attempt also deletes its own output, so a truncated file
  never appears in the restore list.

Verified against live prod, both paths: success writes a 31-table dump the API
container can see; a wrong password fails with mysqldump's own error quoted and
leaves the volume empty.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 16:03:40 -07:00
rmancinasandClaude Opus 5 19f03198d6 fix(docker): install the MySQL 8.4 auth plugin; report why a dump fails
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m56s
Build and Push Images / Build jorgecuadros-api (push) Successful in 3m3s
The pre-migrate backup failed with "mysqldump exited 2" and nothing else.
Reproduced on the host with stderr captured:

  ERROR 1045: Plugin caching_sha2_password could not be loaded:
    /usr/lib/mariadb/plugin/caching_sha2_password.so: No such file or directory

Alpine's `mysql-client` is MariaDB's client and ships an EMPTY plugin
directory, so it cannot perform caching_sha2_password — MySQL 8.4's default and
effectively only auth method. `mariadb-connector-c` provides the plugin.

This was never about the deploy backup alone. Every mysqldump/mysql call from
the API container was broken, which means the whole Operaciones panel — backup,
restore, sync, re-import — could not work in a container. It went unnoticed
because that feature had only ever been run with the API on a developer
machine, where the Oracle client is installed. Verified after the fix: dump
exits 0, gzip valid, 31 CREATE TABLEs.

Also fixed, both found while chasing the above:

- The backup script reported an exit code and nothing else, because a detached
  exec captures no output — which is precisely why this needed a manual
  reproduction. mysqldump's stderr is now redirected to a file and read back
  through a short attached exec on failure, so the deploy log states the cause.
  Verified against live prod: the log now carries the 1045 line itself.

- Listing ONLY 100.100.100.100 as the containers' resolver costs them public
  DNS, since MagicDNS does not forward upstream unless the tailnet defines
  global nameservers. Nothing at runtime needed it, but `apk` inside the
  container stopped resolving, and anything outbound would have too. A public
  fallback resolver is now listed after MagicDNS.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 15:52:03 -07:00
rmancinasandClaude Opus 5 b2cdcbe2cd fix(api): session cookie never issued over HTTP; ship the seed script
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m41s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m5s
Prod came up with nobody able to log in, in two separate ways.

1. No sign-in account exists. `prisma migrate deploy` creates tables, never
   rows, and nothing in the deploy path seeds one — deliberately, since making
   an administrator should not be a side effect of shipping code. But
   apps/api/scripts was not in the runtime image either, so the only way to
   create the first account was to run the script from a developer machine
   against a production DATABASE_URL. Ship scripts/ in the image so it can be
   run on the host with docker exec. Still never run automatically.

2. Login could not establish a session at all. cookie.secure followed NODE_ENV,
   the image sets NODE_ENV=production, and the app is served over plain HTTP —
   express-session then silently emits NO Set-Cookie header. POST /auth/login
   still answered 200 with the full user object, no session was created, every
   later request 403'd, and the UI would have looped back to /login. It reads
   as an auth bug and is really a transport mismatch.

   The flag is now driven by SESSION_COOKIE_SECURE, still defaulting to
   NODE_ENV. An EMPTY value counts as unset rather than false, because compose
   turns an absent `${SESSION_COOKIE_SECURE:-}` into the empty string and the
   naive check would have quietly dropped Secure on any deployment that merely
   passed the variable through.

   galactus sets it to "false". That is acceptable ONLY because the host is
   reachable exclusively over Tailscale, so WireGuard already encrypts the
   wire. It must go back to "true" when the app is served over TLS or exposed
   off-tailnet; behind a TLS-terminating proxy, set trust proxy instead.

Verified against live prod: seeded an admin, POST /auth/login returns 200 with
full ADMIN abilities, a wrong password is rejected with 401, and no Set-Cookie
was present before this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 15:41:55 -07:00
rmancinasandClaude Opus 5 7e3b530174 fix(deploy): pull images explicitly, and detect api/web drift by commit
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m46s
Build and Push Images / Build jorgecuadros-api (push) Successful in 1m58s
The first successful galactus deploy came up all-green while the web tier was
running a build from two commits earlier. The registry held web:latest from
3ff56e6; the host still had a web:latest cached from 4ee7ec7; the deploy
reported success and served the old one. The API was only current because it
had been pulled by hand during earlier debugging.

Two independent failures, both fixed here.

1. Images are not pulled. The deploy action's `pull: true` does not reliably
   refresh an already-cached moving tag on a standalone endpoint. Added a
   Pull images step (deploy/scripts/pull-images.mjs) that pulls each image
   through Portainer's Docker API with registry credentials and fails the
   deploy if a pull fails — note the endpoint answers 200 even when the pull
   errored, so the stream body has to be inspected, not just the status.

2. The drift check could not see it. Both the verify step and the web footer
   compared APP_VERSION, but on a branch build BOTH tiers report "master", so
   equality proved nothing. They now compare gitSha, which is the only field
   that differs between two builds of the same branch. api and web come from
   one matrix run, so a difference can only mean an image was not replaced.

   This needed a /version on the web tier too — previously its build identity
   was only readable by scraping window.__APP_BUILD__ out of the HTML.

pull-images.mjs builds the X-Registry-Auth header as URL-safe base64 WITH
padding: Node's "base64url" omits the padding and Portainer's Go decoder
rejects it with "Illegal base64 data at input byte N".

Verified against galactus: pulls both images, and exits non-zero on a
nonexistent tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 14:57:44 -07:00
rmancinasandClaude Opus 5 1cba9bfc32 fix(galactus): give containers Tailscale's resolver so MagicDNS names resolve
With the image fixed, the API got as far as connecting and then died with
Prisma P1001 "can't reach database server". The cause is DNS, not routing.

galactus runs systemd-resolved, whose 127.0.0.53 stub is unreachable from
inside a container, so Docker falls back to the upstream resolver in
/run/systemd/resolve/resolv.conf — the LAN router, which knows nothing about
the tailnet. Verified from a probe container on galactus: resolving
galactus.tail01aa2.ts.net fails outright, while `nc 100.103.77.46 3306` is
OPEN. Only the lookup was broken.

Pin the api and web services to Tailscale's own resolver (100.100.100.100,
the same anycast address on every tailnet) with this tailnet's search suffix.
Both are overridable via TAILSCALE_DNS / TAILNET_SUFFIX. db and minio need
nothing — they make no outbound calls.

Verified end to end: the published image, unmodified, with only these DNS
settings, boots on galactus against the real database and serves
  /health   {"status":"ok"}
  /version  {"service":"api","version":"master","gitSha":"3ff56e6b..."}

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 14:47:19 -07:00
rmancinasandClaude Opus 5 3ff56e6b72 fix(docker): API image could never boot — missing workspace link and Prisma engine
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m44s
Build and Push Images / Build jorgecuadros-api (push) Successful in 1m57s
Two independent defects in docker/api.Dockerfile, both found by booting the
published image on galactus rather than by reading it. Neither had ever been
observed because no deploy had previously got far enough to start the API.

1. "Cannot find module '@jorgecuadros/database'".
   node-linker=hoisted flattens EXTERNAL dependencies into /repo/node_modules,
   but the workspace dependency stays linked per-package at
   apps/api/node_modules/@jorgecuadros/database -> ../../../../packages/database.
   The runtime stage copied only /repo/node_modules, so the link was dropped.
   Copy the @jorgecuadros scope dir as well — not the whole directory, whose
   only other contents are devDependencies.

2. "Prisma Client could not locate the Query Engine for runtime
   linux-musl-openssl-3.0.x ... generated for linux-musl".
   Prisma picks its engine by sniffing the build environment. The build stage
   had no openssl so it generated for plain "linux-musl", while the runtime
   stage demanded the openssl-3.0.x variant and refused to start. Fixed at both
   ends: binaryTargets now names the musl target explicitly in schema.prisma,
   so the shipped engine no longer depends on what happens to be installed at
   build time, and openssl is installed in the deps stage (generate) and the
   runtime stage (Prisma needs it regardless).

Verified by running the published image on galactus with each fix patched in
by hand, against the real database, until it got past both failures.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 14:39:25 -07:00
rmancinasandClaude Opus 5 27f04f1073 fix(deploy): preflight missing secrets instead of failing opaquely
The first deploy attempt (run 705) died on "Input required and not supplied:
token", which names the action's input rather than the secret that was unset —
the repo had only REGISTRY_USERNAME and REGISTRY_PASSWORD, so every deploy
secret was missing on both workflows. That is also why the endpoint_id /
pull_image input-name bug had gone unnoticed: neither workflow had ever got
far enough to use them.

Both workflows now check their required secrets up front and fail listing the
ones that are empty. The scope=full-only secrets are only required when the
dispatch is actually scope=full.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 14:24:43 -07:00
rmancinasandClaude Opus 5 4ee7ec71f0 feat(deploy): prisma migration history, /version, galactus standalone deploy
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m49s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m2s
Closes the gap between "what tag did I deploy" and "what is actually running",
and gives the schema a history that can be reasoned about across releases.

Migrations
- Baseline the existing schema as 0000_init (migrate diff --from-empty). The
  schema had only ever been applied with `prisma db push`, so no history
  existed and schema state was disconnected from app version. Existing
  databases must be baselined once with `migrate resolve --applied 0000_init`;
  the workflows print this remedy on P3005.
- Run `prisma migrate deploy` as a deploy STEP, not the container CMD — as a
  CMD, N replicas would race each other applying the same migration.

Version reporting
- GET /version on the API reports the APP_VERSION / GIT_SHA / BUILD_DATE that
  build.yml already baked into both images but nothing ever read.
- The web footer shows the web build and flags an api/web mismatch. The two
  cannot drift at build time (one matrix run) but can at deploy time.
- Both deploy workflows now fail if the running API does not report the tag
  that was dispatched — a stack naming a tag is not proof of what is running.
- scripts/set-version.mjs stamps every package.json, which had all sat at
  0.1.0 while real releases shipped as v1.x.

Pre-migrate backup
- deploy/scripts/pre-migrate-backup.mjs dumps the database from INSIDE the
  still-running old API container over Portainer's Docker API, so the file
  lands in the volume the Operaciones restore screen reads. A dump taken on
  the CI runner would be unreachable by the only restore path we have.
  Verifies the artefact with `gzip -t` before letting the migration proceed.

galactus
- deploy/galactus/*.compose.yml: standalone-Docker ports of the Swarm stacks.
  Plain compose silently ignores `deploy:`, so restart_policy becomes
  `restart: unless-stopped` — without it nothing returns after a host reboot.
- .gitea/workflows/deploy-galactus.yml drives endpoint 3 with its own secrets.

Fixes
- deploy.yml passed `endpoint_id` and `pull_image` to
  cssnr/portainer-stack-deploy-action, which has no such inputs (they are
  `endpoint` and `pull`). The endpoint was silently never set.

docs/DEPLOY_AND_MIGRATIONS.md documents expand/contract as the rule for schema
changes: Prisma has no down-migrations, so a code rollback never rolls the
schema back, and restoring the replication master from a dump diverges every
replica.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 11:41:12 -07:00
rmancinasandClaude Opus 5 9ba5d2d09a feat(bank): multi-bank chequera — required bankAccountId, per-account scoping
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m43s
Build and Push Images / Build jorgecuadros-api (push) Successful in 1m59s
The office keeps more than one operating account (Utilities banks in MXN,
Seguros in USD), but bank_transactions was a single implicit MXN register by
design. Adds Bank/BankAccount and makes every read and write in the module
scoped to exactly one account.

Schema:
- Bank / BankAccount. Currency is fixed per account and BankTransaction has
  no currency column of its own — a movement inherits its account's, the way
  a real bank account doesn't mix currencies.
- BankTransaction.bankAccountId, required. A movement with no known account
  isn't reconcilable against a statement.
- @@index([bankAccountId, transactionDate]): every read now filters by
  account and orders/groups by date.

Migration:
- backfill_bank_accounts.py seeds Scotiabank + "Utilities — Scotiabank (MXN)"
  and backfills all 22,669 existing rows onto it, then promotes the column to
  NOT NULL and attaches the FK. Standalone because prisma db push cannot add
  a required column to a populated table. Idempotent; re-running once a second
  account exists does not re-point rows.
- run_all.py runs it (both modes) before transform_bank.py, which now resolves
  the account by label and fails fast if it is missing.

API:
- ?bankAccountId= required on list/stats/facets/summary — not optional with an
  "all accounts" default, since summing an MXN and a USD register repeats the
  currency-collapsing mistake the billing module exists to prevent. Missing is
  400, unknown is 404.
- facets() had no account clause at all and summary() has two raw-SQL rollups;
  all three are now parameterised. Scoping only one of summary's queries would
  leave the year list and its drill-down describing different books.
- New bank/accounts + bank/banks sub-resource under a MANAGER
  bank:manage-accounts ability. currency is absent from the update DTO: booked
  movements are denominated in it, so editing would re-denominate history.
  Capture into a closed account is rejected.

Web:
- /banco gains an account picker (remembered per browser) and reads every
  figure in the selected account's currency; the "single currency (MXN)"
  doc-comment and the hardcoded MXN formatting are gone.
- New /banco/cuentas for banks and accounts. Accounts are closed, never
  deleted — the FK is required, so deleting one would destroy its register.
- /inicio's chequera card names the account it is reading instead of implying
  a single register.

Verified against dev + browser: a second USD account showed full read/write
isolation from the MXN register, whose totals were unchanged (22,669
movements, net 1,014,266.97).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 23:54:16 -07:00
rmancinasandClaude Opus 5 c100dfa224 feat(web,api): scale spacing with text size, persist preference per account
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m13s
Build and Push Images / Build jorgecuadros-api (push) Successful in 3m16s
Two follow-ups to the text-size control.

Spacing now scales with the text. All padding, margin, gap and min-height
declarations in globals.css move from px to rem (263 declarations, converted
mechanically), so --ui-scale drives the whole layout rather than just the
glyphs. Deliberately left in px: border widths, which must stay hairlines;
box-shadow offsets; border-radius, which reads as bloated when scaled on large
cards; --shell-max, a container cap that must not outgrow the viewport; and
media-query breakpoints, which are conditions rather than declarations. With
spacing following along, the presets gain a 1.5 "Máximo" step and MAX_UI_SCALE
rises from 1.4.

The preference now lives on the account instead of only in one browser.
User.uiScale (Float, default 1) is added to the schema and to the safe select,
so it rides along on /auth/login and /auth/me. PATCH /auth/preferences writes
it, guarded by AuthenticatedGuard only — every role including VIEWER may set
their own, and the target is always the session's user id, never a body
parameter, so this cannot be used to touch another account. The global
ValidationPipe's whitelist rejects any extra field, so role cannot ride in
alongside uiScale.

localStorage stays, demoted to a pre-paint cache for the layout.tsx script;
AppShell reconciles it against the account once /auth/me answers, with the
account winning. FontScaleControl becomes a controlled component since the
same value is now edited from the appbar and the drawer.

Verified against the dev API: PATCH persists and is reflected by a subsequent
/auth/me, out-of-range values are rejected 400, and an extra "role" field in
the body is rejected 400.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 22:30:06 -07:00
rmancinasandClaude Opus 5 0bf97e6d2c feat(web): group top nav, add mobile drawer and app-wide text size control
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m37s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m18s
The appbar had grown to 11 flat links with no responsive behaviour, and
overflowed below ~1100px.

Nav is now 7 top-level entries: Inicio, Clientes, Pólizas, Propiedades,
Reportes stay one click away, while the movement screens (Captura, Estado de
cuenta, Chequera) and the admin screens (Catálogos, Usuarios, Operaciones)
collapse into "Cobranza" and "Admin" dropdowns. Groups are ability-filtered
and disappear entirely when the user can see none of their items, so VIEWER
never renders an empty Admin menu. activeHref now scans the flattened link
list, and a group trigger highlights while one of its children is current.

Below 980px the nav collapses to a burger drawer that lists every group
expanded, closing on navigation and on Escape.

Text size is user-adjustable app-wide. Every font-size in globals.css is
converted from px to rem (mechanically, 133 declarations) and the root size
becomes calc(100% * var(--ui-scale)), so one variable on <html> rescales the
whole UI. The preference persists in localStorage and is applied by a
pre-hydration script in layout.tsx to avoid a flash at the default size; the
Aa control lives in the appbar and, as a segmented row, in the drawer.
Spacing stays in px by design, which is why 1.3 is the largest preset.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 22:16:30 -07:00
rmancinasandClaude Opus 5 7df928c3ab feat(billing): receipt capture — outstanding workflow, batch by check, reconciliation
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m1s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m3s
Implements docs/RECEIPT_CAPTURE_SPEC.md §1, the legacy "Editor"
replacement, on top of the single-movement capture from plan step 6.
No new abilities: batching and resolving are both capturing.

- outstanding (legacy NOPAGO): capture flag, ?outstanding= filter, and
  POST /billing/:id/resolve-outstanding (gated ledger:create, not
  ledger:void — resolving completes a capture rather than reversing
  one). Outstanding rows are excluded from every balance aggregate,
  matching the legacy SALDOS ULTIMO 0 query's HAVING NOPAGO = 0, but
  still count in the movement browser's filtered totals.
- POST /billing/batch: many customers' receipts against one check, in
  one $transaction. Deliberately not a persisted batch entity —
  checkNumber is already a column and grouping by it answers every
  legacy by-check query.
- GET /billing/by-check + a cheque-count report, replacing REPORTE
  CHEQUE COUNT / REPORTE POR CHEQUE / EDITA CHEQUE ALF|COUNT|NUM. Print,
  PDF, CSV and XLSX come free from the existing /reportes/:slug machinery.
- Web: /estado-cuenta/lote (the Editor screen, with live reconciliation
  against the physical check amount), an "Estado de pago" filter, a
  "sin fondos" row tag and a Resolver dialog, plus a top-level "Captura"
  nav entry.

Integration seam for the OCR auto-capture module (spec §2), which is
required to post through createBatch rather than writing Transaction
rows itself: items[i] maps to lines[i] so postedTransactionId can be
zipped back on; opts.refs[i] stamps captureRef with a duplicate-post
guard that a voided row deliberately does not block; opts.source is
service-level only, so an HTTP client cannot label hand-keyed rows as
machine-captured. captureSource/captureRef are nullable so the 40,136
migrated rows stay NULL rather than being mislabelled.

Fixes two pre-existing bugs found while building this:

- statement() filtered legacySourceTable with `notIn`, which compiles to
  SQL NOT IN — and `NULL NOT IN (...)` is NULL, so every app-captured
  movement was invisible on the customer statement (438 rows in the
  movement browser vs 392 on the statement) while showing everywhere
  else. This would have made the whole capture feature look broken.
- The balances count query omitted the void filter its own page query
  applied, so the total disagreed with the rows.

Nav highlighting now resolves by longest match; the previous
first-startsWith logic lit up both the parent and any nested entry.

Verified end-to-end against the dev DB, API and browser; all test rows
removed afterwards. Also corrects RESUME.md, which documented the dev
ports as :3001/:3000 — they are :4501/:4500, from the env files.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 21:54:41 -07:00
rmancinasandClaude Opus 5 26a4faa33e docs(insurance): spec renewal emails, liquidación batch, certificate, carrier APIs
Companion to docs/RECEIPT_CAPTURE_SPEC.md — the insurance half of the
2026-07-25/26 meeting. Documentation only; no application code.

Verified against the code and a live query of the dev DB rather than
designed from the meeting notes alone, which changed several conclusions:

- Renewal emails and the liquidación batch are much smaller than they
  look. RenewalNotice + its @@unique([policyId, generation]) idempotency
  key and the aviso-renovacion letter body already exist; the per-policy
  liquidation fields are wired end to end. What's missing is a scheduler,
  a mail client, and the batch layer.
- Carrier research: ANA and GMX are one company (Grupo Valore). ANA
  exposes a live SOAP service with a published operation list; GMX
  publishes no machine interface at all. Every ANA operation serves
  new-business quoting/issuance, not "list my book" — so the direction
  question decides whether the feature is buildable.
- UTILSEG is unusable for Utilities↔Seguros reconciliation and the spec
  closes that long-standing open question: DATGRAL.[NUM UTIL] is
  authoritative (name match 298/563 vs 58/1024), and where the two
  sources overlap they contradict on 170 of 218 shared ids.

Also records two live defects found while verifying: policy_types is
missing its INCENDIO and M_EMPR rows (the FK is ON DELETE SET NULL, so 5
m_empr policies silently lost their ramo), and the legacy settlement
slots don't match the target model (MULT/INCENDIO carry two, M EMPR
carries four, Policy collapses to one).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 21:54:23 -07:00
rmancinasandClaude Sonnet 5 159dcc4963 docs(plan): add step 11 for receipt-capture + net-new ops features
Points PLAN.md at docs/RECEIPT_CAPTURE_SPEC.md and surfaces its open
design questions (OCR provider, Seguros bank details, clave catastral
vs. predial, recycling triggers) separately from the existing ops-only
open items list.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 23:28:13 -07:00
rmancinasandClaude Sonnet 5 9b9ee201c9 docs: add receipt capture, OCR, multi-bank & customer-recycling spec
Forward implementation spec covering the legacy "Editor" receipt-capture
workflow plus three net-new requests from the 2026-07-25/26 meeting with
Jorge: PDF/OCR auto-capture, multi-bank chequera support, and
customer-number recycling. Matching logic and data-model gaps for each
were verified against the actual migration scripts and API code, not
just the schema comments.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 23:26:23 -07:00
rmancinasandClaude Opus 5 6dbd4a319b fix(reports): render renewal-letter premium lines as booleans
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m44s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m1s
`r.netPremium` / `r.total` are strings, so an empty-string value leaked
""` into the JSX instead of rendering nothing. Wrap in Boolean() so the
guard is a real conditional.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 12:07:32 -07:00
rmancinasandClaude Sonnet 5 f7ae0d5342 feat(reports): parameterized renewal-notice report + legacy report reference
Build and Push Images / Build jorgecuadros-web (push) Failing after 59s
Build and Push Images / Build jorgecuadros-api (push) Successful in 1m59s
Replaces ~40 legacy Access renewal-notice report clones (one per carrier
per coverage tier, e.g. AMPL/RC/LIC RENEW X MES/VENCE ATLAS 13/2013) with
one parameterized aviso-renovacion report driven by real Policy/Vehicle/
coveragesJson data instead of hand-typed label text per clone.

- schema.prisma: add RenewalNotice, replacing the legacy CONTROL <ramo>
  RENEW[2/3] X MES paper log of which notice generation was sent
- reports: new "letter" ReportFormat + aviso-renovacion registry entry +
  LetterLayout renderer in ReportRunner.tsx
- docs/RENEWAL_NOTICES.md + migration/legacy_report_defs/: extracted (via
  Application.SaveAsText, since the VBA project wouldn't load) and
  documented the legacy report/query chain this replaces

Coveragesjson key names and a mark-as-sent mutation are still unverified/
unbuilt — see caveats in docs/RENEWAL_NOTICES.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-25 23:05:21 -07:00
rmancinasandClaude Opus 4.8 1b79b43a54 fix(migration): make Phase B additive sync actually work + verify end-to-end
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m37s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m9s
The --sync path had never been run and was broken in several ways. Fixed and
verified against the dev DB (two consecutive syncs, both exit 0, 32/32
assertions: stable PKs, manual-row preservation, changed-row updates,
legacy-delete, no child duplication, zero FK orphans; idempotent).

- policies/properties: reuse each legacy row's existing id (by provenance)
  BEFORE building child rows, so children no longer point at a discarded fresh
  uuid; rebuild legacy-owned children via scoped delete + reinsert.
- customers: replace zip(customers, refs) (mispaired almost every row) with a
  ref-grouped id remap; names now restore and no spurious customers appear.
- drop the invalid Vehicle @@unique(legacySourceTable, legacyId) — one legacy
  policy row carries up to 3 vehicles sharing a legacyId; handle via delete+reinsert.
- upsert lookup tables (policy_types, insurance_providers, type_transactions,
  adjusters) by natural name and remap child FKs instead of inserting fresh
  uuids that nothing points at.
- transactions: drop updatedAt=NOW() (no such column); guard report formatting
  on NULL legacySourceTable (manual rows). Same report guard in bank.
- add manual-safe prune (prune_empty_customers.py --sync, in SYNC_STEPS): prune
  only legacy-owned empties, never manually-added customers.

web: customer-detail mini tx list now strikes voided rows with an "(anulado)"
tag (was the last void-UI rendering gap; /estado-cuenta already handled it).

docs: RESUME.md updated — Phase B sync marked verified end-to-end, void-UI
browser pass recorded.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-24 13:37:22 -07:00
rmancinas 8802f08d4f feat(reports): reports module + /inicio + edo-cuenta-datos prefill
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m35s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m11s
- New reports backend (registry, service, controller, outputs, types)
  with catalog endpoint + slug/CSV/XLSX/PDF/print outputs.
- /reportes catalog + /reportes/[slug] runner; ReportRunner + ContextReports
  components wire pre-filtered links from domain pages.
- Fix: /reportes/[slug] now reads searchParams and forwards initialParams to
  ReportRunner so /reportes/edo-cuenta-datos?customerId=... auto-runs
  instead of dropping the id and forcing a manual customer search.
- /inicio landing page; root + login redirect to /inicio.
- Company header env vars + logo asset for PDF/print rendering.
- exceljs + pdfkit deps.
2026-07-23 23:20:41 -07:00
rmancinas 921a47cbaa feat(web): use company logo PNG in brand mark
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m2s
Build and Push Images / Build jorgecuadros-web (push) Successful in 3m7s
Replace the JC text mark in AppShell and login with the real
company_logo.png. Restyle .brand-mark to host the image (white
rounded bg, object-fit contain). Appbar 38px, login panel 46px.
2026-07-23 22:17:17 -07:00
rmancinas 27b3bd9efc feat: expand admin and data sync workflows
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m40s
Build and Push Images / Build jorgecuadros-api (push) Successful in 3m13s
2026-07-23 22:00:08 -07:00
rmancinasandClaude Opus 4.8 70911e7e62 feat(deploy): app stack + manual Portainer deploy workflow
Build and Push Images / Build jorgecuadros-web (push) Successful in 2m2s
Build and Push Images / Build jorgecuadros-api (push) Successful in 3m36s
Add the missing api/web deployment path on top of the existing image build CI.

- deploy/jorgecuadros-app.stack.yml: PROD app stack (api + web) pulling the
  git.mancinas.io registry images. Does not ship mysql/minio (separate stacks);
  API reaches them via DATABASE_URL / S3_ENDPOINT. API pinned to the
  jorgecuadros_db node for stable ingest/backup volumes; web is stateless.
- deploy/jorgecuadros-app.env.example: documented stack env template.
- .gitea/workflows/deploy.yml: manual (workflow_dispatch) deploy to Portainer
  via cssnr/portainer-stack-deploy-action. Inputs: image tag + scope
  (app = web+api, full = db+minio+app, applied db->minio->app).

Make the web API origin runtime-configurable instead of build-baked: the root
layout injects window.__API_ORIGIN__ from the API_ORIGIN env (force-dynamic) and
lib/api.ts resolves it at runtime, so one built image serves any deployment.

Also: dev.sh to run both dev servers (frees stale ports first) and move local
dev to ports web 4500 / api 4501.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-23 20:07:48 -07:00
rmancinasandClaude Opus 4.8 afe2411c86 feat(storage): wire MinIO/S3 document upload & download into the API + web
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m37s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m8s
The schema has carried `storageKey` pointers and the migration has written
blobs to MinIO since day one, but the API had no S3 client — documents could
only be deleted, never uploaded or retrieved. This adds the missing wiring.

API
- StorageModule/StorageService (@aws-sdk/client-s3, path-style for MinIO):
  put/getStream/delete, best-effort bucket ensure on boot, gracefully disabled
  when S3 env is absent (ServiceUnavailable on use).
- Reads S3_ENDPOINT/S3_BUCKET + S3_ACCESS_KEY/S3_SECRET_KEY, falling back to
  MINIO_ROOT_USER/MINIO_ROOT_PASSWORD so one credential set drives both the
  migration and the API.
- Property service documents: POST :id/documents (multipart), GET
  :id/documents/:childId/download (streamed), delete now also drops the blob.
- Policy documents: same upload/download/delete (previously had none).
- Keys stay under the service/<id>/… and policy/<id>/… prefixes the migration
  established.

Web
- api.ts: shared uploadFile() helper (uploadIngest refactored onto it),
  upload/download/remove helpers for property & policy documents.
- Servicios, polizas, clientes detail pages: real Descargar links and an
  upload control (gated by policy:update / property:update) replacing the
  "storage pending" notes.

Infra
- docker-compose: minio service (9000/9001, healthcheck, named volume) + S3
  env wired into the api service.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-23 19:19:57 -07:00
rmancinas 45afb824ef Merge feat/crud-rbac: CRUD/RBAC + ops panel + Docker CI/versioning
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m41s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m8s
2026-07-23 19:02:07 -07:00
214 changed files with 47036 additions and 934 deletions
+47
View File
@@ -3,3 +3,50 @@ DATABASE_URL=mysql://jorgecuadros:jorgecuadros@localhost:3306/jorgecuadros
SESSION_SECRET=change-me-to-a-random-string SESSION_SECRET=change-me-to-a-random-string
WEB_ORIGIN=http://localhost:3000 WEB_ORIGIN=http://localhost:3000
NEXT_PUBLIC_API_ORIGIN=http://localhost:3001 NEXT_PUBLIC_API_ORIGIN=http://localhost:3001
# Object storage (MinIO / S3) for document blobs and scanned receipt pages.
# Without S3_ENDPOINT + credentials the API still boots, but every document
# upload/download and the whole recibo OCR intake are disabled. Credentials fall
# back to MINIO_ROOT_USER / MINIO_ROOT_PASSWORD when the S3_* pair is unset.
S3_ENDPOINT=http://localhost:9000
S3_BUCKET=jorgecuadros-documents
S3_ACCESS_KEY=
S3_SECRET_KEY=
# Login the "Operaciones" screen runs mysqldump/mysql as. Optional locally: when
# unset it falls back to the DATABASE_URL credentials, which a dev MySQL usually
# grants enough for. Required in any deployment, where the application user has
# only ALL ON jorgecuadros.* and mysqldump --single-transaction needs the global
# RELOAD privilege. Host/port/database always come from DATABASE_URL.
OPS_DB_ADMIN_USER=
OPS_DB_ADMIN_PASSWORD=
# Company info — printed in the header of every report (PDF + browser
# print). Leave blank to use the placeholders. COMPANY_LOGO_PATH is
# optional; when unset the API falls back to apps/api/assets/company_logo.png.
COMPANY_NAME=Jorge Cuadros & Asociados
COMPANY_ADDRESS_LINE1=
COMPANY_ADDRESS_LINE2=
COMPANY_CITY_STATE=
COMPANY_PHONE=
COMPANY_EMAIL=
COMPANY_TAX_ID=
COMPANY_WEBSITE=
COMPANY_LOGO_PATH=
# Outbound mail (Amazon SES — the channel the office already uses for bulk
# notification, see docs/MASS_EMAIL_NOTIFICATIONS.md). Without all four
# vars the API still boots; in dev the MailService logs sends to stdout,
# in production every send throws ServiceUnavailableException.
SES_REGION=
SES_ACCESS_KEY=
SES_SECRET_KEY=
SES_FROM=mail@jorgecuadros.com
SES_FROM_NAME=Information Server
# Optional — bounce/complaint event publishing configuration set.
SES_CONFIGURATION_SET=
# Comma-separated addresses that receive the per-job admin summary email
# (one summary per address, JSON body, sent after every sweep). Defaults to
# the legacy pair if unset.
NOTIFICATION_ADMIN_EMAILS=rmancinas@freakma.net,mpulido@freakma.net
+20 -1
View File
@@ -14,7 +14,16 @@
# ARG/ENV (APP_VERSION / GIT_SHA / BUILD_DATE) and as OCI labels, so a running # ARG/ENV (APP_VERSION / GIT_SHA / BUILD_DATE) and as OCI labels, so a running
# container can report exactly what is deployed. # container can report exactly what is deployed.
# #
# Release flow: git tag v1.2.0 && git push origin v1.2.0 -> versioned images. # Release flow: git tag v1.2.0 && git push origin v1.2.0 -> versioned images
# -> deploy-on-tag.yml waits for this run to go green and then
# dispatches deploy-galactus.yml.
#
# This workflow BUILDS ONLY — it never deploys. The deploy chain used to be a
# job here, gated to tag refs, but Gitea draws every job of a workflow into the
# run graph before it evaluates the job's `if`: a routine master build showed a
# pending "Deploy to galactus" and looked like prod was about to be redeployed
# off an unreleased commit. Keeping the deploy in a `on: push: tags` workflow of
# its own makes that structurally impossible.
name: Build and Push Images name: Build and Push Images
@@ -37,6 +46,16 @@ env:
jobs: jobs:
build: build:
name: Build ${{ matrix.image }} name: Build ${{ matrix.image }}
# release.yml pushes the release commit and its tag in a single `git push`,
# so Gitea creates two runs for the same commit: one for master, one for the
# tag. Only the tag run matters — it is the one that emits the X.Y.Z / X.Y
# image tags, and it publishes `latest` and `sha-<short>` too, since it is
# the same commit. Skip the branch run rather than racing or cancelling it.
# Ordinary pushes to master (any message but `chore(release):`) still build.
if: >-
github.event_name != 'push' ||
startsWith(github.ref, 'refs/tags/') ||
!startsWith(github.event.head_commit.message, 'chore(release):')
runs-on: docker runs-on: docker
container: container:
image: docker:27-dind image: docker:27-dind
+388
View File
@@ -0,0 +1,388 @@
# Manual PROD deploy to galactus — the office server, Portainer endpoint 3.
#
# galactus is STANDALONE Docker (`swarm: inactive`), so this workflow applies
# the compose files under deploy/galactus/, NOT the Swarm files in deploy/.
# .gitea/workflows/deploy.yml is the cubex/Swarm equivalent; the two are kept
# separate on purpose because plain compose silently ignores Swarm's `deploy:`
# keys rather than failing on them.
#
# This does NOT build. build.yml already built + pushed both images from one
# matrix run, so api and web at the same tag are always in step.
#
# Order of operations, and why:
# 1. db + minio (scope=full only) — the API depends on both.
# 2. pre-migrate backup dumped INSIDE the still-running OLD api container,
# so the file lands in the volume the Operaciones
# restore screen reads. Must precede the migration.
# 3. prisma migrate deploy forward-only. Prisma has no down-migrations; see
# docs/DEPLOY_AND_MIGRATIONS.md — expand/contract is
# the rule, the backup is the emergency lever.
# Done HERE so the schema moves while the OLD code is
# still serving. The api container ALSO migrates at
# start (docker/api-entrypoint.sh); `migrate deploy`
# is idempotent, so the second run is a no-op and the
# container is what covers a restart that never goes
# through this workflow at all.
# 4. app (api + web) the new images.
# 5. verify ask the running API what it actually is.
#
# Rollback = re-dispatch with an older `tag`. That rolls back CODE only; the
# schema stays forward. This is exactly why every schema change must be
# backward-compatible with the previous release.
#
# Prereqs (once):
# - Gitea repo secrets, galactus-specific (suffix _GALACTUS so the cubex
# secrets keep working side by side):
# PORTAINER_URL_GALACTUS https://100.103.77.46:9443
# PORTAINER_API_KEY_GALACTUS Portainer access token for galactus
# PORTAINER_ENDPOINT_ID_GALACTUS 3
# PORTAINER_APP_STACK_NAME_GALACTUS e.g. jorgecuadros-prod-app
# PORTAINER_DB_STACK_NAME_GALACTUS e.g. jorgecuadros-prod-db
# PORTAINER_MINIO_STACK_NAME_GALACTUS e.g. jorgecuadros-prod-minio
# DATABASE_URL_GALACTUS mysql://jorgecuadros:<pass>@<galactus>:3306/jorgecuadros
# APP_API_ORIGIN_GALACTUS browser-facing API URL
# APP_WEB_ORIGIN_GALACTUS web public origin (API CORS)
# APP_S3_ENDPOINT_GALACTUS server-side minio URL
# SESSION_SECRET_GALACTUS 64-hex (openssl rand -hex 32)
# MINIO_ROOT_USER / MINIO_ROOT_PASSWORD
# MYSQL_PASSWORD / MYSQL_ROOT_PASSWORD
# Optional — outbound mail. Not needed to deploy; needed for
# /notificaciones to send anything at all (the image sets
# NODE_ENV=production, which disables MailService's stdout fallback, so
# a blank config fails every send loudly):
# SES_REGION e.g. us-west-2
# SES_FROM a VERIFIED SES sending identity
# SES_FROM_NAME display name, optional
# SES_ACCESS_KEY / SES_SECRET_KEY
# SES_CONFIGURATION_SET optional, for bounce/complaint events
# NOTIFICATION_ADMIN_EMAILS fallback only — the summary recipients
# are edited in the UI and stored in
# app_settings; this is what a deployment
# uses until somebody saves them there
# These are NOT galactus-specific (no _GALACTUS suffix) — one SES identity
# serves every deployment.
# - The runner (which lives on cubex) must be able to reach galactus:9443
# (Portainer). It should also reach galactus:3306 for step 3, but that is
# no longer load-bearing: dispatch with skip_migrate=true and the api
# container applies the migrations itself at start.
# - ONE-TIME, on a database that predates migration history (i.e. one built
# with `prisma db push`): baseline it before the first run, or step 3 fails
# with P3005 "database schema is not empty":
# npx prisma@5 migrate resolve --applied 0000_init \
# --schema packages/database/prisma/schema.prisma
name: Deploy to galactus
on:
workflow_dispatch:
inputs:
tag:
description: "Image tag to deploy (1.2.3 — no leading v — or sha-<short>, or latest)"
required: true
default: "latest"
scope:
description: "What to deploy"
type: choice
required: true
default: "app"
options:
- app
- full
bootstrap:
description: "First-ever deploy: allow the pre-migrate backup to be skipped when no API container exists yet"
type: boolean
required: false
default: false
skip_migrate:
description: "Skip the runner-side migrate step (safe: the api container migrates at start)"
type: boolean
required: false
default: false
env:
REGISTRY: git.mancinas.io
jobs:
deploy:
name: Deploy ${{ github.event.inputs.tag }} (${{ github.event.inputs.scope }})
runs-on: docker
container:
image: node:20-alpine
steps:
- name: Install tools
# openssl: prisma's migration engine picks its musl/openssl build at
# runtime and cannot resolve one without it.
run: apk add --no-cache openssl ca-certificates git
- uses: actions/checkout@v4
# An unset secret arrives as an empty string, and the deploy action then
# fails with "Input required and not supplied: token" — which names the
# action's input, not the secret you forgot. Check them up front and say
# exactly which ones are missing.
- name: Preflight — required secrets
env:
PORTAINER_URL_GALACTUS: ${{ secrets.PORTAINER_URL_GALACTUS }}
PORTAINER_API_KEY_GALACTUS: ${{ secrets.PORTAINER_API_KEY_GALACTUS }}
PORTAINER_ENDPOINT_ID_GALACTUS: ${{ secrets.PORTAINER_ENDPOINT_ID_GALACTUS }}
PORTAINER_APP_STACK_NAME_GALACTUS: ${{ secrets.PORTAINER_APP_STACK_NAME_GALACTUS }}
PORTAINER_DB_STACK_NAME_GALACTUS: ${{ secrets.PORTAINER_DB_STACK_NAME_GALACTUS }}
PORTAINER_MINIO_STACK_NAME_GALACTUS: ${{ secrets.PORTAINER_MINIO_STACK_NAME_GALACTUS }}
DATABASE_URL_GALACTUS: ${{ secrets.DATABASE_URL_GALACTUS }}
SESSION_SECRET_GALACTUS: ${{ secrets.SESSION_SECRET_GALACTUS }}
APP_API_ORIGIN_GALACTUS: ${{ secrets.APP_API_ORIGIN_GALACTUS }}
APP_WEB_ORIGIN_GALACTUS: ${{ secrets.APP_WEB_ORIGIN_GALACTUS }}
APP_S3_ENDPOINT_GALACTUS: ${{ secrets.APP_S3_ENDPOINT_GALACTUS }}
MINIO_ROOT_USER: ${{ secrets.MINIO_ROOT_USER }}
MINIO_ROOT_PASSWORD: ${{ secrets.MINIO_ROOT_PASSWORD }}
MYSQL_PASSWORD: ${{ secrets.MYSQL_PASSWORD }}
MYSQL_ROOT_PASSWORD: ${{ secrets.MYSQL_ROOT_PASSWORD }}
# Not required — the app boots fine without mail. Warned about below,
# because the failure mode is remote: everything looks healthy until
# someone clicks "Ejecutar" and every send fails.
SES_REGION: ${{ secrets.SES_REGION }}
SES_FROM: ${{ secrets.SES_FROM }}
SES_ACCESS_KEY: ${{ secrets.SES_ACCESS_KEY }}
SES_SECRET_KEY: ${{ secrets.SES_SECRET_KEY }}
SCOPE: ${{ github.event.inputs.scope }}
run: |
REQUIRED="PORTAINER_URL_GALACTUS PORTAINER_API_KEY_GALACTUS
PORTAINER_ENDPOINT_ID_GALACTUS PORTAINER_APP_STACK_NAME_GALACTUS
DATABASE_URL_GALACTUS SESSION_SECRET_GALACTUS
APP_API_ORIGIN_GALACTUS APP_WEB_ORIGIN_GALACTUS
APP_S3_ENDPOINT_GALACTUS MINIO_ROOT_USER MINIO_ROOT_PASSWORD
MYSQL_ROOT_PASSWORD"
if [ "$SCOPE" = "full" ]; then
REQUIRED="$REQUIRED PORTAINER_DB_STACK_NAME_GALACTUS
PORTAINER_MINIO_STACK_NAME_GALACTUS MYSQL_PASSWORD"
fi
missing=""
for name in $REQUIRED; do
eval "value=\${$name}"
[ -z "$value" ] && missing="$missing $name"
done
if [ -n "$missing" ]; then
echo "::error::missing repo secrets:$missing"
echo "::error::set them under Settings > Actions > Secrets"
exit 1
fi
echo "all required secrets present for scope=$SCOPE"
# Mail is optional to deploy but not optional to work. Say so loudly
# rather than letting /notificaciones fail one send at a time.
mail_missing=""
for name in SES_REGION SES_FROM SES_ACCESS_KEY SES_SECRET_KEY; do
eval "value=\${$name}"
[ -z "$value" ] && mail_missing="$mail_missing $name"
done
if [ -n "$mail_missing" ]; then
echo "::warning::outbound mail is NOT configured, missing:$mail_missing"
echo "::warning::the deploy will succeed, but every notification and"
echo "::warning::renewal aviso will fail with 'El envío de correo no"
echo "::warning::está configurado.' See docs/MASS_EMAIL_NOTIFICATIONS.md"
fi
# --- full only: database ---------------------------------------------
- name: Deploy database stack
if: ${{ github.event.inputs.scope == 'full' }}
uses: cssnr/portainer-stack-deploy-action@v1
with:
url: ${{ secrets.PORTAINER_URL_GALACTUS }}
token: ${{ secrets.PORTAINER_API_KEY_GALACTUS }}
name: ${{ secrets.PORTAINER_DB_STACK_NAME_GALACTUS }}
file: deploy/galactus/jorgecuadros-db.compose.yml
type: file
standalone: true
endpoint: ${{ secrets.PORTAINER_ENDPOINT_ID_GALACTUS }}
env_data: |
{
"MYSQL_SERVER_ID": "1",
"MYSQL_PORT": "3306",
"MYSQL_DATABASE": "jorgecuadros",
"MYSQL_USER": "jorgecuadros",
"MYSQL_PASSWORD": "${{ secrets.MYSQL_PASSWORD }}",
"MYSQL_ROOT_PASSWORD": "${{ secrets.MYSQL_ROOT_PASSWORD }}"
}
# --- full only: object storage ---------------------------------------
- name: Deploy minio stack
if: ${{ github.event.inputs.scope == 'full' }}
uses: cssnr/portainer-stack-deploy-action@v1
with:
url: ${{ secrets.PORTAINER_URL_GALACTUS }}
token: ${{ secrets.PORTAINER_API_KEY_GALACTUS }}
name: ${{ secrets.PORTAINER_MINIO_STACK_NAME_GALACTUS }}
file: deploy/galactus/jorgecuadros-minio.compose.yml
type: file
standalone: true
endpoint: ${{ secrets.PORTAINER_ENDPOINT_ID_GALACTUS }}
env_data: |
{
"MINIO_API_PORT": "9000",
"MINIO_CONSOLE_PORT": "9001",
"MINIO_ROOT_USER": "${{ secrets.MINIO_ROOT_USER }}",
"MINIO_ROOT_PASSWORD": "${{ secrets.MINIO_ROOT_PASSWORD }}"
}
# --- restore point, taken while the OLD api container is still up ------
- name: Pre-migrate backup
env:
PORTAINER_URL: ${{ secrets.PORTAINER_URL_GALACTUS }}
PORTAINER_API_KEY: ${{ secrets.PORTAINER_API_KEY_GALACTUS }}
PORTAINER_ENDPOINT_ID: ${{ secrets.PORTAINER_ENDPOINT_ID_GALACTUS }}
DATABASE_URL: ${{ secrets.DATABASE_URL_GALACTUS }}
# The dump runs as root: --single-transaction issues FLUSH TABLES,
# which needs the global RELOAD privilege the application user
# deliberately does not have.
MYSQL_ROOT_PASSWORD: ${{ secrets.MYSQL_ROOT_PASSWORD }}
BACKUP_TAG: ${{ github.event.inputs.tag }}
ALLOW_MISSING_CONTAINER: ${{ github.event.inputs.bootstrap }}
# Portainer serves a self-signed certificate. Scoped to this step
# only, which does nothing but talk to Portainer.
NODE_TLS_REJECT_UNAUTHORIZED: "0"
run: node deploy/scripts/pre-migrate-backup.mjs
# --- schema, forward-only ---------------------------------------------
# Belt to the container's braces: this runs while the OLD code is still
# serving, which is the order expand/contract is designed around. The
# api container repeats it at start for the paths this step cannot
# reach (skip_migrate, a host reboot, a stack re-applied by hand).
- name: Apply database migrations
if: ${{ github.event.inputs.skip_migrate != 'true' }}
env:
DATABASE_URL: ${{ secrets.DATABASE_URL_GALACTUS }}
run: |
set -e
SCHEMA=packages/database/prisma/schema.prisma
npx --yes prisma@5 migrate status --schema "$SCHEMA" || true
if ! npx --yes prisma@5 migrate deploy --schema "$SCHEMA"; then
echo "::error::migrate deploy failed. If this is P3005 (schema not empty),"
echo "::error::the database predates migration history — baseline it once with:"
echo "::error:: npx prisma@5 migrate resolve --applied 0000_init --schema $SCHEMA"
exit 1
fi
# --- make sure the host actually has the images ------------------------
# The deploy action's `pull: true` does not reliably refresh an already
# cached moving tag. Pull explicitly, or a "successful" deploy can leave
# the host serving an older build of the same tag.
- name: Pull images
env:
PORTAINER_URL: ${{ secrets.PORTAINER_URL_GALACTUS }}
PORTAINER_API_KEY: ${{ secrets.PORTAINER_API_KEY_GALACTUS }}
PORTAINER_ENDPOINT_ID: ${{ secrets.PORTAINER_ENDPOINT_ID_GALACTUS }}
REGISTRY: ${{ env.REGISTRY }}
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
IMAGES: ${{ github.repository_owner }}/jorgecuadros-api,${{ github.repository_owner }}/jorgecuadros-web
TAG: ${{ github.event.inputs.tag }}
NODE_TLS_REJECT_UNAUTHORIZED: "0"
run: node deploy/scripts/pull-images.mjs
# --- always: the app (web + api) -------------------------------------
- name: Deploy app stack
uses: cssnr/portainer-stack-deploy-action@v1
with:
url: ${{ secrets.PORTAINER_URL_GALACTUS }}
token: ${{ secrets.PORTAINER_API_KEY_GALACTUS }}
name: ${{ secrets.PORTAINER_APP_STACK_NAME_GALACTUS }}
file: deploy/galactus/jorgecuadros-app.compose.yml
type: file
standalone: true
pull: true
endpoint: ${{ secrets.PORTAINER_ENDPOINT_ID_GALACTUS }}
# NOTE: the block below is parsed as JSON — no comments inside it.
#
# API_ORIGIN is deliberately absent. The browser derives the API origin
# from the page it loaded (apps/web/src/lib/api.ts), so the deployment
# survives the box moving between the tailnet, the office LAN and a
# demo domain. Setting it here would pin it again and re-break an https
# front door with mixed active content. APP_API_ORIGIN_GALACTUS lives
# on only as the URL the verify step probes.
env_data: |
{
"APP_TAG": "${{ github.event.inputs.tag }}",
"API_PORT": "3001",
"WEB_PORT": "3000",
"S3_BUCKET": "jorgecuadros-documents",
"WEB_ORIGIN": "${{ secrets.APP_WEB_ORIGIN_GALACTUS }}",
"S3_ENDPOINT": "${{ secrets.APP_S3_ENDPOINT_GALACTUS }}",
"DATABASE_URL": "${{ secrets.DATABASE_URL_GALACTUS }}",
"SESSION_SECRET": "${{ secrets.SESSION_SECRET_GALACTUS }}",
"SESSION_COOKIE_SECURE": "false",
"OPS_DB_ADMIN_USER": "root",
"OPS_DB_ADMIN_PASSWORD": "${{ secrets.MYSQL_ROOT_PASSWORD }}",
"REPLICA_DB_HOST": "${{ secrets.REPLICA_DB_HOST }}",
"REPLICA_DB_USER": "${{ secrets.REPLICA_DB_USER }}",
"REPLICA_DB_PASS": "${{ secrets.REPLICA_DB_PASS }}",
"MINIO_ROOT_USER": "${{ secrets.MINIO_ROOT_USER }}",
"MINIO_ROOT_PASSWORD": "${{ secrets.MINIO_ROOT_PASSWORD }}",
"SES_REGION": "${{ secrets.SES_REGION }}",
"SES_FROM": "${{ secrets.SES_FROM }}",
"SES_FROM_NAME": "${{ secrets.SES_FROM_NAME }}",
"SES_ACCESS_KEY": "${{ secrets.SES_ACCESS_KEY }}",
"SES_SECRET_KEY": "${{ secrets.SES_SECRET_KEY }}",
"SES_CONFIGURATION_SET": "${{ secrets.SES_CONFIGURATION_SET }}",
"NOTIFICATION_ADMIN_EMAILS": "${{ secrets.NOTIFICATION_ADMIN_EMAILS }}"
}
# --- prove it ----------------------------------------------------------
- name: Verify running version
env:
API_ORIGIN: ${{ secrets.APP_API_ORIGIN_GALACTUS }}
WEB_ORIGIN: ${{ secrets.APP_WEB_ORIGIN_GALACTUS }}
WANT: ${{ github.event.inputs.tag }}
# A stack naming a tag is not proof the containers run it. Ask BOTH
# tiers what they are, and require them to be the same commit: api and
# web are built from one matrix run, so a difference can only mean one
# of them did not actually get replaced.
run: |
set -e
apk add --no-cache curl >/dev/null
# These secrets are CORS origin LISTS as far as the app is concerned
# (WEB_ORIGIN is comma-separated so one deployment can be reached by
# LAN IP, tailnet name and demo domain at once). A list is not a URL,
# so probe the FIRST entry — keep the runner-reachable origin first.
API_ORIGIN=${API_ORIGIN%%,*}
WEB_ORIGIN=${WEB_ORIGIN%%,*}
fetch_version() {
for i in $(seq 1 30); do
if curl -fsS "$1/version" > "$2"; then return 0; fi
echo "waiting for $1 ($i/30)..."
sleep 5
done
echo "::error::$1/version never answered"
return 1
}
fetch_version "$API_ORIGIN" /tmp/api.json
fetch_version "$WEB_ORIGIN" /tmp/web.json
cat /tmp/api.json; echo; cat /tmp/web.json; echo
API_SHA=$(node -e 'console.log(require("/tmp/api.json").gitSha)')
WEB_SHA=$(node -e 'console.log(require("/tmp/web.json").gitSha)')
API_VER=$(node -e 'console.log(require("/tmp/api.json").version)')
# Compare the COMMIT, not the version string: on a branch build both
# tiers report "master", so version equality proves nothing.
if [ "$API_SHA" != "$WEB_SHA" ]; then
echo "::error::api and web are different builds — api $API_SHA, web $WEB_SHA"
echo "::error::one of the images was not replaced; check the Pull images step"
exit 1
fi
echo "api and web agree: $API_SHA"
# A semver dispatch is additionally comparable to the tag itself:
# metadata-action's {{version}} turns tag v1.2.3 into image 1.2.3,
# while `latest` and `sha-*` report the branch or short sha instead.
case "$WANT" in
[0-9]*.[0-9]*.[0-9]*)
if [ "$API_VER" != "$WANT" ]; then
echo "::error::deployed $WANT but the API reports $API_VER"
exit 1
fi
echo "verified: running $API_VER"
;;
*)
echo "dispatched '$WANT'; tiers report '$API_VER' (not directly comparable)"
;;
esac
+199
View File
@@ -0,0 +1,199 @@
# Chain the PROD deploy onto a green tag build.
#
# This is a SEPARATE workflow, not a job in build.yml, and the trigger is the
# whole point: `on: push: tags` cannot fire on a push to master. When this was a
# `deploy` job inside build.yml gated by `if: startsWith(github.ref,
# 'refs/tags/v')`, Gitea still drew "Deploy to galactus" into the job graph of
# every ordinary master build — the `if` is not evaluated until `needs` resolve,
# so the job sits there looking like an imminent production deploy on a commit
# nobody released. That is indistinguishable from a real misfire, and the only
# safe reaction is to cancel the run, which kills the images with it.
#
# What it does NOT do is build. build.yml already builds and pushes both images
# from one run; this waits for that run to go green and then dispatches
# deploy-galactus.yml, which only pulls.
#
# Why wait for the build run rather than just dispatching: deploy-galactus.yml
# pulls api and web at the same tag, and a half-pushed pair is exactly the state
# that leaves prod running one new image and one old one. The build run turning
# green is the signal that both are in the registry.
#
# Why a dispatch and not a `workflow_run:` trigger, which Gitea does support as
# of 1.24: deploy-galactus.yml reads `github.event.inputs.*` in ten places (tag,
# scope, bootstrap, skip_migrate). Under workflow_run every one of them is the
# empty string, so the deploy would silently run with no tag and scope != 'full'.
# A dispatch keeps that workflow's contract intact and keeps it hand-runnable for
# rollbacks, which is the whole point of it.
#
# Kill switch: set the repo variable AUTO_DEPLOY_GALACTUS to `false` to cut the
# chain and go back to dispatching the deploy by hand. Anything else (including
# unset) deploys.
name: Deploy on tag
on:
push:
tags: ["v*"]
jobs:
deploy:
name: Deploy to galactus
runs-on: docker
container:
image: node:20-alpine
steps:
- name: Preflight — RELEASE_TOKEN
env:
RELEASE_TOKEN: ${{ secrets.RELEASE_TOKEN }}
run: |
set -eu
if [ -z "${RELEASE_TOKEN:-}" ]; then
echo "::error::Secret RELEASE_TOKEN is not set, so this cannot wait"
echo "::error::for the build or dispatch the deploy. Once build.yml"
V=${GITHUB_REF#refs/tags/}
echo "::error::is green, run 'Deploy to galactus' by hand with tag=${V#v}."
exit 1
fi
- name: Wait for the tag build, then dispatch deploy-galactus.yml
env:
RELEASE_TOKEN: ${{ secrets.RELEASE_TOKEN }}
AUTO_DEPLOY: ${{ vars.AUTO_DEPLOY_GALACTUS }}
TAG_REF: ${{ github.ref }}
BUILD_SHA: ${{ github.sha }}
run: |
node -e '
const base = `${process.env.GITHUB_SERVER_URL}/api/v1/repos/${process.env.GITHUB_REPOSITORY}`;
const headers = { Authorization: `token ${process.env.RELEASE_TOKEN}` };
// refs/tags/v1.2.3 — derived from github.ref rather than ref_name so
// it does not depend on how Gitea populates GITHUB_REF_NAME.
const tagRef = process.env.TAG_REF;
const tag = tagRef.replace(/^refs\/tags\//, "");
// The git tag carries the leading v; the image tag does not.
const version = tag.replace(/^v/, "");
const sha = process.env.BUILD_SHA;
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
const runs = async () => {
const r = await fetch(`${base}/actions/runs?limit=50`, { headers });
if (!r.ok) throw new Error(`runs query failed: HTTP ${r.status}`);
return (await r.json()).workflow_runs || [];
};
// The release commit and its tag are the SAME sha, and build.yml
// skips the master run by design — so a sha match alone can latch
// onto that skipped run and call the build green when no image was
// ever pushed. Require the tag ref when the API reports one.
const isTagRun = (r) => {
const ref = r.head_branch || r.ref || "";
return !ref || ref === tag || ref === tagRef;
};
const buildRun = async () =>
(await runs()).find(
(r) =>
r.head_sha === sha &&
String(r.path || "").includes("build.yml") &&
isTagRun(r),
);
// Gitea reports a run as `status` and mirrors it into `conclusion`;
// read whichever is populated rather than betting on one field.
const outcome = (r) => String(r.conclusion || r.status || "").toLowerCase();
const DONE = ["success", "failure", "cancelled", "canceled", "skipped"];
// Every deploy-galactus run id visible right now. A dispatch is only
// confirmed by an id that is NOT in here — a plain "is there a deploy
// run" check is satisfied by the PREVIOUS release run, and would
// report success for a dispatch that never took.
const deployRunIds = async () =>
new Set(
(await runs())
.filter((r) => String(r.path || "").includes("deploy-galactus.yml"))
.map((r) => r.id),
);
(async () => {
if (process.env.AUTO_DEPLOY === "false") {
console.log("AUTO_DEPLOY_GALACTUS=false — not deploying.");
console.log(`Deploy by hand with tag=${version} when ready.`);
return;
}
// ~20 min. A build is about 90s; the rest is queue time behind
// other runs on a single runner.
let run = null;
for (let i = 0; i < 80; i++) {
run = await buildRun();
if (run && DONE.includes(outcome(run))) break;
if (!run && i === 11) {
// Two minutes with no run at all. The post-receive hook drops
// runs silently when it errors (this cost v1.0.3 its images),
// so say so rather than timing out with no explanation.
console.log(`::warning::No build.yml run for ${tag} yet after 2 min.`);
console.log(`::warning::If the Gitea post-receive hook is broken, dispatch`);
console.log(`::warning::"Build and Push Images" by hand with ref=${tag}.`);
}
await sleep(15_000);
}
if (!run) {
console.log(`::error::No build.yml run for ${tag} (${sha}) after 20 min.`);
console.log(`::error::Dispatch "Build and Push Images" with ref=${tag} (the`);
console.log(`::error::tag, not master), then deploy by hand with tag=${version}.`);
process.exit(1);
}
const result = outcome(run);
if (result !== "success") {
console.log(`::error::build.yml for ${tag} ended as "${result}" — not deploying.`);
console.log(`::error::Fix the build, re-run it, then deploy by hand with tag=${version}.`);
process.exit(1);
}
console.log(`build.yml for ${tag} is green (run ${run.id}). Deploying ${version}.`);
const before = await deployRunIds();
// Dispatch against the TAG, not master: the deploy applies the
// compose files under deploy/galactus/ from whatever ref it runs
// on, and those must be the ones this release was cut with.
const res = await fetch(
`${base}/actions/workflows/deploy-galactus.yml/dispatches`,
{
method: "POST",
headers: { ...headers, "Content-Type": "application/json" },
body: JSON.stringify({
ref: tagRef,
inputs: {
tag: version,
scope: "app",
bootstrap: "false",
skip_migrate: "false",
},
}),
},
);
if (!res.ok) {
console.log(`::error::Dispatch returned HTTP ${res.status}: ${await res.text()}`);
console.log(`::error::Images for ${version} are published. Run`);
console.log(`::error::"Deploy to galactus" by hand with tag=${version}.`);
process.exit(1);
}
// A 204 only means Gitea accepted the request. Confirm a NEW run
// exists — an accepted call that creates no run is the failure mode
// that cost v1.0.3 its images.
for (let i = 0; i < 3; i++) {
await sleep(5_000);
const fresh = [...(await deployRunIds())].filter((id) => !before.has(id));
if (fresh.length) {
console.log(`Deploy of ${version} to galactus is running (run ${fresh[0]}).`);
return;
}
}
console.log(`::error::Dispatch was accepted but no deploy run appeared.`);
console.log(`::error::Run "Deploy to galactus" by hand with tag=${version}.`);
process.exit(1);
})();
'
+343
View File
@@ -0,0 +1,343 @@
# Manual PROD deploy to Portainer.
#
# This does NOT build — build.yml already builds + pushes the api/web images.
# This workflow (re)applies the deploy/*.stack.yml files to the Portainer Swarm.
# Trigger it by hand from the Actions tab ("Run workflow") and choose:
# - tag: which already-published image tag to ship (default: latest)
# - scope: how much to deploy
# app = web + api only (the usual app release) [default]
# full = db + minio + web + api (bring up / update the whole platform)
#
# The `tag` input carries NO leading `v`: metadata-action's {{version}} turns
# git tag v1.2.3 into image tag 1.2.3. Tag v1.2.3, dispatch 1.2.3.
#
# Order: db+minio (full only) -> pre-migrate backup -> prisma migrate deploy ->
# (the api container also migrates at start; see docker/api-entrypoint.sh)
# app -> verify the API reports the version you asked for. Rollback = dispatch
# an older tag; that rolls back CODE only, never the schema, which is why every
# schema change must be expand/contract. See docs/DEPLOY_AND_MIGRATIONS.md.
#
# galactus (the office server) is standalone Docker, not this Swarm — it has its
# own workflow, .gitea/workflows/deploy-galactus.yml.
#
# cssnr/portainer-stack-deploy-action creates each stack on first run and updates
# it on every run, so no manual stack pre-creation in the Portainer UI. On a
# `full` deploy the db + minio stacks are applied BEFORE the app (the API depends
# on them). db + minio are stateful + pinned to node label jorgecuadros_db=true
# (see their stack files) — re-applying them is idempotent and keeps their data.
#
# Prereqs (once):
# - one swarm node labelled jorgecuadros_db=true (db + minio + api pin there).
# - Gitea repo secrets set (Settings > Actions > Secrets):
# # Portainer
# PORTAINER_URL https://192.168.4.212:9443
# PORTAINER_API_KEY Portainer access token
# PORTAINER_ENDPOINT_ID 2 (the local Swarm endpoint)
# PORTAINER_APP_STACK_NAME e.g. jorgecuadros-prod-app
# PORTAINER_DB_STACK_NAME e.g. jorgecuadros-prod-db (full only)
# PORTAINER_MINIO_STACK_NAME e.g. jorgecuadros-prod-minio (full only)
# # App runtime
# DATABASE_URL mysql://jorgecuadros:<pass>@192.168.4.212:3306/jorgecuadros
# SESSION_SECRET 64-hex (openssl rand -hex 32)
# APP_API_ORIGIN http://192.168.4.212:3001 (browser-facing API URL)
# APP_WEB_ORIGIN http://192.168.4.212:3000 (web public origin, API CORS)
# APP_S3_ENDPOINT http://192.168.4.212:9000 (server-side minio URL)
# # Object storage (app + minio stack)
# MINIO_ROOT_USER minio access key
# MINIO_ROOT_PASSWORD minio secret key
# # Database stack (full only)
# MYSQL_PASSWORD app-user password (matches DATABASE_URL)
# MYSQL_ROOT_PASSWORD mysql root password
# - the runner must reach Portainer (9443). It should also reach MySQL (3306)
# for the migrate step, but that is no longer load-bearing: dispatch with
# skip_migrate=true and the api container applies the migrations itself at
# start (docker/api-entrypoint.sh). `migrate deploy` is idempotent, so the
# two never conflict.
# - ONE-TIME on a database built with `prisma db push` (i.e. every database
# that exists today): baseline it before the first run, or the migrate step
# fails with P3005 "database schema is not empty":
# npx prisma@5 migrate resolve --applied 0000_init \
# --schema packages/database/prisma/schema.prisma
name: Deploy to Portainer
on:
workflow_dispatch:
inputs:
tag:
description: "Image tag to deploy (latest, sha-<short>, or vX.Y.Z)"
required: true
default: "latest"
scope:
description: "What to deploy"
type: choice
required: true
default: "app"
options:
- app
- full
bootstrap:
description: "First-ever deploy: allow the pre-migrate backup to be skipped when no API container exists yet"
type: boolean
required: false
default: false
skip_migrate:
description: "Skip the runner-side migrate step (safe: the api container migrates at start)"
type: boolean
required: false
default: false
env:
REGISTRY: git.mancinas.io
jobs:
deploy:
name: Deploy (${{ github.event.inputs.scope }})
runs-on: docker
container:
image: node:20-alpine
steps:
- name: Install tools
# openssl: prisma's migration engine picks its musl/openssl build at
# runtime and cannot resolve one without it.
run: apk add --no-cache openssl ca-certificates git
- uses: actions/checkout@v4
# An unset secret arrives as an empty string, and the deploy action then
# fails with "Input required and not supplied: token" — which names the
# action's input, not the secret you forgot.
- name: Preflight — required secrets
env:
PORTAINER_URL: ${{ secrets.PORTAINER_URL }}
PORTAINER_API_KEY: ${{ secrets.PORTAINER_API_KEY }}
PORTAINER_ENDPOINT_ID: ${{ secrets.PORTAINER_ENDPOINT_ID }}
PORTAINER_APP_STACK_NAME: ${{ secrets.PORTAINER_APP_STACK_NAME }}
PORTAINER_DB_STACK_NAME: ${{ secrets.PORTAINER_DB_STACK_NAME }}
PORTAINER_MINIO_STACK_NAME: ${{ secrets.PORTAINER_MINIO_STACK_NAME }}
DATABASE_URL: ${{ secrets.DATABASE_URL }}
SESSION_SECRET: ${{ secrets.SESSION_SECRET }}
APP_API_ORIGIN: ${{ secrets.APP_API_ORIGIN }}
APP_WEB_ORIGIN: ${{ secrets.APP_WEB_ORIGIN }}
APP_S3_ENDPOINT: ${{ secrets.APP_S3_ENDPOINT }}
MINIO_ROOT_USER: ${{ secrets.MINIO_ROOT_USER }}
MINIO_ROOT_PASSWORD: ${{ secrets.MINIO_ROOT_PASSWORD }}
MYSQL_PASSWORD: ${{ secrets.MYSQL_PASSWORD }}
MYSQL_ROOT_PASSWORD: ${{ secrets.MYSQL_ROOT_PASSWORD }}
SCOPE: ${{ github.event.inputs.scope }}
run: |
REQUIRED="PORTAINER_URL PORTAINER_API_KEY PORTAINER_ENDPOINT_ID
PORTAINER_APP_STACK_NAME DATABASE_URL SESSION_SECRET
APP_API_ORIGIN APP_WEB_ORIGIN APP_S3_ENDPOINT
MINIO_ROOT_USER MINIO_ROOT_PASSWORD MYSQL_ROOT_PASSWORD"
if [ "$SCOPE" = "full" ]; then
REQUIRED="$REQUIRED PORTAINER_DB_STACK_NAME PORTAINER_MINIO_STACK_NAME
MYSQL_PASSWORD"
fi
missing=""
for name in $REQUIRED; do
eval "value=\${$name}"
[ -z "$value" ] && missing="$missing $name"
done
if [ -n "$missing" ]; then
echo "::error::missing repo secrets:$missing"
echo "::error::set them under Settings > Actions > Secrets"
exit 1
fi
echo "all required secrets present for scope=$SCOPE"
# --- full only: database ---------------------------------------------
- name: Deploy database stack
if: ${{ github.event.inputs.scope == 'full' }}
uses: cssnr/portainer-stack-deploy-action@v1
with:
url: ${{ secrets.PORTAINER_URL }}
token: ${{ secrets.PORTAINER_API_KEY }}
name: ${{ secrets.PORTAINER_DB_STACK_NAME }}
file: deploy/jorgecuadros-db.stack.yml
type: file
endpoint: ${{ secrets.PORTAINER_ENDPOINT_ID }}
env_data: |
{
"MYSQL_SERVER_ID": "1",
"MYSQL_PORT": "3306",
"MYSQL_DATABASE": "jorgecuadros",
"MYSQL_USER": "jorgecuadros",
"MYSQL_PASSWORD": "${{ secrets.MYSQL_PASSWORD }}",
"MYSQL_ROOT_PASSWORD": "${{ secrets.MYSQL_ROOT_PASSWORD }}"
}
# --- full only: object storage ---------------------------------------
- name: Deploy minio stack
if: ${{ github.event.inputs.scope == 'full' }}
uses: cssnr/portainer-stack-deploy-action@v1
with:
url: ${{ secrets.PORTAINER_URL }}
token: ${{ secrets.PORTAINER_API_KEY }}
name: ${{ secrets.PORTAINER_MINIO_STACK_NAME }}
file: deploy/jorgecuadros-minio.stack.yml
type: file
endpoint: ${{ secrets.PORTAINER_ENDPOINT_ID }}
env_data: |
{
"MINIO_API_PORT": "9000",
"MINIO_CONSOLE_PORT": "9001",
"MINIO_ROOT_USER": "${{ secrets.MINIO_ROOT_USER }}",
"MINIO_ROOT_PASSWORD": "${{ secrets.MINIO_ROOT_PASSWORD }}"
}
# --- restore point, taken while the OLD api container is still up ------
# Dumped INSIDE the running api container so the file lands in the volume
# the "Operaciones" restore screen reads — a dump on the runner would be
# unreachable by the only restore path this platform has.
- name: Pre-migrate backup
env:
PORTAINER_URL: ${{ secrets.PORTAINER_URL }}
PORTAINER_API_KEY: ${{ secrets.PORTAINER_API_KEY }}
PORTAINER_ENDPOINT_ID: ${{ secrets.PORTAINER_ENDPOINT_ID }}
DATABASE_URL: ${{ secrets.DATABASE_URL }}
# The dump runs as root: --single-transaction issues FLUSH TABLES,
# which needs the global RELOAD privilege the application user
# deliberately does not have.
MYSQL_ROOT_PASSWORD: ${{ secrets.MYSQL_ROOT_PASSWORD }}
BACKUP_TAG: ${{ github.event.inputs.tag }}
ALLOW_MISSING_CONTAINER: ${{ github.event.inputs.bootstrap }}
# Portainer serves a self-signed certificate. Scoped to this step
# only, which does nothing but talk to Portainer.
NODE_TLS_REJECT_UNAUTHORIZED: "0"
run: node deploy/scripts/pre-migrate-backup.mjs
# --- schema, forward-only ---------------------------------------------
# Prisma has no down-migrations: a code rollback does NOT roll the schema
# back. See docs/DEPLOY_AND_MIGRATIONS.md — every change must be
# expand/contract so the previous release still runs against the new
# schema. Run as a deploy STEP, never as the container CMD: N replicas
# would race each other applying the same migration.
- name: Apply database migrations
if: ${{ github.event.inputs.skip_migrate != 'true' }}
env:
DATABASE_URL: ${{ secrets.DATABASE_URL }}
run: |
set -e
SCHEMA=packages/database/prisma/schema.prisma
npx --yes prisma@5 migrate status --schema "$SCHEMA" || true
if ! npx --yes prisma@5 migrate deploy --schema "$SCHEMA"; then
echo "::error::migrate deploy failed. If this is P3005 (schema not empty),"
echo "::error::the database predates migration history — baseline it once with:"
echo "::error:: npx prisma@5 migrate resolve --applied 0000_init --schema $SCHEMA"
exit 1
fi
# --- make sure the host actually has the images ------------------------
# The deploy action's `pull: true` does not reliably refresh an already
# cached moving tag; without this a "successful" deploy can leave the host
# serving an older build of the same tag.
- name: Pull images
env:
PORTAINER_URL: ${{ secrets.PORTAINER_URL }}
PORTAINER_API_KEY: ${{ secrets.PORTAINER_API_KEY }}
PORTAINER_ENDPOINT_ID: ${{ secrets.PORTAINER_ENDPOINT_ID }}
REGISTRY: ${{ env.REGISTRY }}
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
IMAGES: ${{ github.repository_owner }}/jorgecuadros-api,${{ github.repository_owner }}/jorgecuadros-web
TAG: ${{ github.event.inputs.tag }}
NODE_TLS_REJECT_UNAUTHORIZED: "0"
run: node deploy/scripts/pull-images.mjs
# --- always: the app (web + api) -------------------------------------
- name: Deploy app stack
uses: cssnr/portainer-stack-deploy-action@v1
with:
url: ${{ secrets.PORTAINER_URL }}
token: ${{ secrets.PORTAINER_API_KEY }}
name: ${{ secrets.PORTAINER_APP_STACK_NAME }}
file: deploy/jorgecuadros-app.stack.yml
type: file
pull: true
endpoint: ${{ secrets.PORTAINER_ENDPOINT_ID }}
# NOTE: the block below is parsed as JSON — no comments inside it.
#
# API_ORIGIN is deliberately absent. The browser derives the API origin
# from the page it loaded (apps/web/src/lib/api.ts), so the deployment
# survives the host moving. Setting it here would pin it again and
# re-break an https front door with mixed active content. APP_API_ORIGIN
# lives on only as the URL the verify step probes.
env_data: |
{
"APP_TAG": "${{ github.event.inputs.tag }}",
"API_PORT": "3001",
"WEB_PORT": "3000",
"S3_BUCKET": "jorgecuadros-documents",
"WEB_ORIGIN": "${{ secrets.APP_WEB_ORIGIN }}",
"S3_ENDPOINT": "${{ secrets.APP_S3_ENDPOINT }}",
"DATABASE_URL": "${{ secrets.DATABASE_URL }}",
"SESSION_SECRET": "${{ secrets.SESSION_SECRET }}",
"OPS_DB_ADMIN_USER": "root",
"OPS_DB_ADMIN_PASSWORD": "${{ secrets.MYSQL_ROOT_PASSWORD }}",
"MINIO_ROOT_USER": "${{ secrets.MINIO_ROOT_USER }}",
"MINIO_ROOT_PASSWORD": "${{ secrets.MINIO_ROOT_PASSWORD }}",
"SES_REGION": "${{ secrets.SES_REGION }}",
"SES_FROM": "${{ secrets.SES_FROM }}",
"SES_FROM_NAME": "${{ secrets.SES_FROM_NAME }}",
"SES_ACCESS_KEY": "${{ secrets.SES_ACCESS_KEY }}",
"SES_SECRET_KEY": "${{ secrets.SES_SECRET_KEY }}",
"SES_CONFIGURATION_SET": "${{ secrets.SES_CONFIGURATION_SET }}",
"NOTIFICATION_ADMIN_EMAILS": "${{ secrets.NOTIFICATION_ADMIN_EMAILS }}"
}
# --- prove it ----------------------------------------------------------
# A stack naming a tag is not proof the container is running it — a
# skipped pull leaves the old code up. Ask the API what it actually is.
- name: Verify running version
env:
API_ORIGIN: ${{ secrets.APP_API_ORIGIN }}
WEB_ORIGIN: ${{ secrets.APP_WEB_ORIGIN }}
WANT: ${{ github.event.inputs.tag }}
run: |
set -e
apk add --no-cache curl >/dev/null
# These secrets are CORS origin LISTS as far as the app is concerned
# (WEB_ORIGIN is comma-separated so one deployment can be reached under
# several origins at once). A list is not a URL, so probe the FIRST
# entry — keep the runner-reachable origin first.
API_ORIGIN=${API_ORIGIN%%,*}
WEB_ORIGIN=${WEB_ORIGIN%%,*}
fetch_version() {
for i in $(seq 1 30); do
if curl -fsS "$1/version" > "$2"; then return 0; fi
echo "waiting for $1 ($i/30)..."
sleep 5
done
echo "::error::$1/version never answered"
return 1
}
fetch_version "$API_ORIGIN" /tmp/api.json
fetch_version "$WEB_ORIGIN" /tmp/web.json
cat /tmp/api.json; echo; cat /tmp/web.json; echo
API_SHA=$(node -e 'console.log(require("/tmp/api.json").gitSha)')
WEB_SHA=$(node -e 'console.log(require("/tmp/web.json").gitSha)')
API_VER=$(node -e 'console.log(require("/tmp/api.json").version)')
# Compare the COMMIT, not the version string: on a branch build both
# tiers report "master", so version equality proves nothing.
if [ "$API_SHA" != "$WEB_SHA" ]; then
echo "::error::api and web are different builds — api $API_SHA, web $WEB_SHA"
echo "::error::one of the images was not replaced; check the Pull images step"
exit 1
fi
echo "api and web agree: $API_SHA"
case "$WANT" in
[0-9]*.[0-9]*.[0-9]*)
if [ "$API_VER" != "$WANT" ]; then
echo "::error::deployed $WANT but the API reports $API_VER"
exit 1
fi
echo "verified: running $API_VER"
;;
*)
echo "dispatched '$WANT'; tiers report '$API_VER' (not directly comparable)"
;;
esac
+280
View File
@@ -0,0 +1,280 @@
# Cut a release: stamp the version across every package.json, commit, tag, push.
#
# This does NOT build and does NOT deploy itself. Pushing the `vX.Y.Z` tag is
# what triggers both build.yml, which publishes the `X.Y.Z`, `X.Y`,
# `sha-<short>` and `latest` image tags, and deploy-on-tag.yml, which waits for
# that build to go green and then dispatches deploy-galactus.yml with
# `tag=X.Y.Z scope=app` (no leading v — the git tag carries the `v`, the image
# tag does not). A tag is the only ref that starts either chain; pushing to
# master builds images and stops there.
#
# So cutting a release DOES reach prod. To cut a version without deploying it,
# set the repo variable AUTO_DEPLOY_GALACTUS=false first; deploy-on-tag.yml then
# prints the manual command instead of running it.
#
# Why a workflow instead of three local commands: the release commit is the one
# thing that must be identical every time, and cutting it from a laptop is how
# a manifest bump gets forgotten or a tag lands on an unpushed commit. Here the
# only input is the number.
#
# Prereqs (once):
# - Repo secret RELEASE_TOKEN: a Gitea personal access token with
# write:repository on this repo. The built-in Actions token is deliberately
# NOT used — whether a push made with it re-triggers build.yml depends on the
# Gitea version, and a release that silently publishes no images is worse
# than one that fails. A PAT push is an ordinary push and always triggers.
# If build.yml somehow does not start, it has workflow_dispatch: run it
# against the new tag by hand.
name: Cut release
on:
workflow_dispatch:
inputs:
bump:
description: "Which part to bump (choose 'explicit' to type the number)"
type: choice
required: true
default: "minor"
options:
- patch
- minor
- major
- explicit
version:
description: "Exact version when bump=explicit (x.y.z, no leading v)"
required: false
default: ""
jobs:
release:
name: Release
runs-on: docker
container:
image: node:20-alpine
steps:
- name: Install tools
run: apk add --no-cache git
- name: Preflight — RELEASE_TOKEN
env:
RELEASE_TOKEN: ${{ secrets.RELEASE_TOKEN }}
run: |
set -eu
if [ -z "${RELEASE_TOKEN:-}" ]; then
echo "::error::Secret RELEASE_TOKEN is not set. Create a Gitea PAT with"
echo "::error::write:repository and add it as a repo secret named RELEASE_TOKEN."
exit 1
fi
# Full history + tags: the duplicate-tag check below is meaningless
# against a shallow clone, which has none of them.
- uses: actions/checkout@v4
with:
fetch-depth: 0
ref: master
token: ${{ secrets.RELEASE_TOKEN }}
- name: Resolve the new version
id: ver
env:
BUMP: ${{ github.event.inputs.bump }}
EXPLICIT: ${{ github.event.inputs.version }}
run: |
set -eu
CURRENT=$(node -p "require('./package.json').version")
echo "current: $CURRENT"
if [ "$BUMP" = "explicit" ]; then
NEXT="$EXPLICIT"
if [ -z "$NEXT" ]; then
echo "::error::bump=explicit requires the version input."
exit 1
fi
else
NEXT=$(node -e '
const [cur, part] = process.argv.slice(1);
const m = /^(\d+)\.(\d+)\.(\d+)/.exec(cur);
if (!m) { console.error(`unparseable current version: ${cur}`); process.exit(1); }
let [maj, min, pat] = m.slice(1).map(Number);
if (part === "major") { maj += 1; min = 0; pat = 0; }
else if (part === "minor") { min += 1; pat = 0; }
else { pat += 1; }
process.stdout.write(`${maj}.${min}.${pat}`);
' "$CURRENT" "$BUMP")
fi
# set-version.mjs validates the shape too, but failing here keeps the
# working tree clean when the input is a typo.
case "$NEXT" in
v*) echo "::error::Version must not carry a leading 'v' (got $NEXT)."; exit 1 ;;
esac
if ! printf '%s' "$NEXT" | grep -Eq '^[0-9]+\.[0-9]+\.[0-9]+(-[0-9A-Za-z.-]+)?$'; then
echo "::error::Invalid version: $NEXT (expected x.y.z)."
exit 1
fi
if [ "$NEXT" = "$CURRENT" ]; then
echo "::error::$NEXT is already the current version."
exit 1
fi
if git rev-parse -q --verify "refs/tags/v$NEXT" >/dev/null; then
echo "::error::Tag v$NEXT already exists. Releases are immutable — pick a new number."
exit 1
fi
echo "next: $NEXT"
echo "version=$NEXT" >> "$GITHUB_OUTPUT"
- name: Stamp the version across every manifest
run: node scripts/set-version.mjs "${{ steps.ver.outputs.version }}"
# A release whose only content is the version bump means the dispatch was
# a mistake — set-version.mjs already refused a no-op above, so an empty
# diff here means the manifests were somehow already at this number.
- name: Commit, tag, push
env:
RELEASE_TOKEN: ${{ secrets.RELEASE_TOKEN }}
VERSION: ${{ steps.ver.outputs.version }}
ACTOR: ${{ github.actor }}
run: |
set -eu
if git diff --quiet; then
echo "::error::No manifest changed. Nothing to release."
exit 1
fi
git config user.name "gitea-actions"
git config user.email "actions@git.mancinas.io"
git commit -a \
-m "chore(release): v${VERSION}" \
-m "Cut by ${ACTOR} via the \"Cut release\" workflow. Pushing the tag triggers build.yml; deploy separately with tag=${VERSION}."
git tag -a "v${VERSION}" -m "v${VERSION}"
# Re-point at an authenticated remote. The token is a secret, so Gitea
# masks it in the log; nothing here echoes the URL regardless.
git remote set-url origin \
"$(printf '%s' "${GITHUB_SERVER_URL}" | sed "s#://#://x-access-token:${RELEASE_TOKEN}@#")/${GITHUB_REPOSITORY}.git"
# One push for both refs: a commit that lands without its tag builds
# nothing and looks like a successful release.
#
# The output is captured because a failing *post-receive* hook does not
# fail the push: git prints `remote: error: ...`, updates both refs and
# exits 0. That is how v1.0.3 was cut — the hook 500'd, so Gitea never
# created the build run, and this step went green anyway.
if ! git push origin "HEAD:master" "refs/tags/v${VERSION}" 2>push.log; then
cat push.log
echo "::error::Push failed. Nothing was released."
exit 1
fi
cat push.log
if grep -q '^remote: error' push.log; then
echo "::warning::The remote's post-receive hook errored. Both refs landed,"
echo "::warning::but Gitea most likely created no workflow run for them."
echo "::warning::The next step checks and dispatches build.yml if needed."
fi
echo "sha=$(git rev-parse HEAD)" >> "$GITHUB_OUTPUT"
id: push
# Gitea creates workflow runs from the post-receive hook, so a hook error
# silently costs you the build: the tag exists, no image is ever published,
# and the failure only surfaces later as a 404 when deploy pulls the image.
# Confirm the run exists; dispatch it if it does not; fail loudly if that
# does not work either.
- name: Verify build.yml started
env:
RELEASE_TOKEN: ${{ secrets.RELEASE_TOKEN }}
VERSION: ${{ steps.ver.outputs.version }}
SHA: ${{ steps.push.outputs.sha }}
run: |
node -e '
const base = `${process.env.GITHUB_SERVER_URL}/api/v1/repos/${process.env.GITHUB_REPOSITORY}`;
const headers = { Authorization: `token ${process.env.RELEASE_TOKEN}` };
const sha = process.env.SHA;
const tag = `v${process.env.VERSION}`;
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
// The master push and the tag push carry the SAME commit, so a sha
// match alone is not enough: build.yml skips the master run by
// design, and that skipped run would satisfy a sha-only check even
// if the tag run were never created. When the API reports a ref for
// the run, require it to be the tag; when it reports none, fall back
// to the sha match rather than failing a release over a field name.
const isTagRun = (r) => {
const ref = r.head_branch || r.ref || "";
return !ref || ref === tag || ref === `refs/tags/${tag}`;
};
const started = async () => {
const res = await fetch(`${base}/actions/runs?limit=30`, { headers });
if (!res.ok) throw new Error(`runs query failed: HTTP ${res.status}`);
const body = await res.json();
return (body.workflow_runs || []).some(
(r) =>
r.head_sha === sha &&
String(r.path || "").includes("build.yml") &&
isTagRun(r),
);
};
// The hook fires synchronously with the push, so a run that is coming
// is usually already there; the retries cover a busy instance.
const poll = async (attempts) => {
for (let i = 0; i < attempts; i++) {
if (await started()) return true;
await sleep(10_000);
}
return started();
};
(async () => {
if (await poll(3)) {
console.log(`build.yml is running for ${sha}.`);
return;
}
console.log(`No build.yml run for ${sha}. Dispatching against ${tag}.`);
const res = await fetch(
`${base}/actions/workflows/build.yml/dispatches`,
{
method: "POST",
headers: { ...headers, "Content-Type": "application/json" },
// Must be the tag, not master: metadata-action only emits the
// X.Y.Z and X.Y image tags when the ref is a semver tag. And
// it must be the fully qualified ref — Gitea 404s on `v1.0.3`.
body: JSON.stringify({ ref: `refs/tags/${tag}` }),
},
);
if (!res.ok) console.log(`Dispatch returned HTTP ${res.status}.`);
if (await poll(3)) {
console.log(`build.yml is running for ${sha}.`);
return;
}
console.log(`::error::${tag} is pushed but nothing is building it, and`);
console.log(`::error::the dispatch did not take. Run "Build and Push Images"`);
console.log(`::error::by hand with ref=${tag} (the tag, not master), then`);
console.log(`::error::deploy. Check the Gitea server log for the`);
console.log(`::error::post-receive error while you are at it.`);
process.exit(1);
})();
'
- name: Summary
env:
VERSION: ${{ steps.ver.outputs.version }}
run: |
set -eu
echo "Released v${VERSION}."
echo ""
echo "build.yml is now building git.mancinas.io/rmancinas/jorgecuadros-{api,web}:${VERSION}."
echo "deploy-on-tag.yml is watching that build; when it goes green it dispatches"
echo "'Deploy to galactus' with tag=${VERSION} scope=app bootstrap=false skip_migrate=false."
echo ""
echo "Watch that run. If it did not start (or AUTO_DEPLOY_GALACTUS=false),"
echo "dispatch 'Deploy to galactus' by hand with the same inputs."
echo "Rollback = re-dispatch it with an older tag."
+67
View File
@@ -1,5 +1,11 @@
# Unified Customer / Insurance / Utilities Platform — Migration & Rebuild Plan # Unified Customer / Insurance / Utilities Platform — Migration & Rebuild Plan
> **Looking for what is still outstanding?** → [`docs/BACKLOG.md`](docs/BACKLOG.md).
> This document is the plan and its running status; the backlog collects every
> open item — blocked-on-Jorge decisions, live data defects, unbuilt features
> and deploy blockers — in one list, checked against the code rather than
> against these notes.
## Context ## Context
Jorge Cuadros & Assoc. runs two lines of business — property/utility management (`UTILITIES.accdb`) and insurance brokerage (`SEGUROS 16.mdb` + its linked backend `SEGUROS 16_be.mdb`) — out of separate, decades-old MS Access databases, plus a third file (`SCOTHIA.mdb`) that's the office's own Scotiabank checking-account register ("chequera"). The same people are customers of both business lines, but today there's no shared customer record: a person's utility account and their insurance policies live in unrelated systems with independent, inconsistent copies of their name/address/contact info. The bank register is a fourth, disconnected source of truth for the money actually moving through the office's own account. Jorge Cuadros & Assoc. runs two lines of business — property/utility management (`UTILITIES.accdb`) and insurance brokerage (`SEGUROS 16.mdb` + its linked backend `SEGUROS 16_be.mdb`) — out of separate, decades-old MS Access databases, plus a third file (`SCOTHIA.mdb`) that's the office's own Scotiabank checking-account register ("chequera"). The same people are customers of both business lines, but today there's no shared customer record: a person's utility account and their insurance policies live in unrelated systems with independent, inconsistent copies of their name/address/contact info. The bank register is a fourth, disconnected source of truth for the money actually moving through the office's own account.
@@ -130,6 +136,33 @@ Given the amount of near-duplicate/overlapping data across snapshot tables (mult
8. VPS provisioning + Tailscale + MySQL replication setup. `utility_dbo`'s schema is now available (full dump on disk — 55 tables; see Status), so the exact replicated table/column set and inbox-table shape can be finalized against the real portal DB and the portal PHP code (`my-jorgecuadros-web`) that reads/writes it. 8. VPS provisioning + Tailscale + MySQL replication setup. `utility_dbo`'s schema is now available (full dump on disk — 55 tables; see Status), so the exact replicated table/column set and inbox-table shape can be finalized against the real portal DB and the portal PHP code (`my-jorgecuadros-web`) that reads/writes it.
9. Sync worker (push replicated tables' relevant subset, poll inbox tables for payment/propane submissions) — depends on step 8. **The separate Phase B Access additive sync is implemented:** `migration/run_all.py --sync` and the admin `SYNC` job upsert legacy-owned rows without truncating the database or touching manual rows. Portal write points confirmed present in `utility_dbo`: `peticion_gas` (propane requests), PayPal payment writes, `notifications_settings`, `verification_codes` — these define the VPS→internal inbox set. 9. Sync worker (push replicated tables' relevant subset, poll inbox tables for payment/propane submissions) — depends on step 8. **The separate Phase B Access additive sync is implemented:** `migration/run_all.py --sync` and the admin `SYNC` job upsert legacy-owned rows without truncating the database or touching manual rows. Portal write points confirmed present in `utility_dbo`: `peticion_gas` (propane requests), PayPal payment writes, `notifications_settings`, `verification_codes` — these define the VPS→internal inbox set.
10. Reports/email campaigns/admin — parity with old app's `reports.php`/`emailCampaigns.php` intent, rebuilt properly. 10. Reports/email campaigns/admin — parity with old app's `reports.php`/`emailCampaigns.php` intent, rebuilt properly.
11. **Receipt capture ("Editor") completion + three net-new ops features — NOT STARTED, spec written.** Full design in [`docs/RECEIPT_CAPTURE_SPEC.md`](docs/RECEIPT_CAPTURE_SPEC.md), from the 2026-07-25/26 meeting with Jorge:
- **Receipt capture module — DONE** (2026-07-27). The legacy "Editor" replacement, built on the single-movement capture from step 6. Wires up the previously-unused `Transaction.outstanding` (NOPAGO): capture flag on `POST /billing`, `?outstanding=` list filter, `POST /billing/:id/resolve-outstanding` (gated `ledger:create`, not `ledger:void` — resolving *completes* a capture), and exclusion from every balance aggregate exactly as the legacy `SALDOS ULTIMO 0`'s `HAVING NOPAGO = 0` did. Adds `POST /billing/batch` (one `$transaction`, check-level fields shared, per-line customer/amount) and `GET /billing/by-check`, plus the `cheque-count` report replacing `REPORTE CHEQUE COUNT` / `REPORTE POR CHEQUE` / `EDITA CHEQUE ALF|COUNT|NUM` — print/PDF/CSV/XLSX come free from the existing `/reportes/:slug` machinery. Web: `/estado-cuenta/lote` (the actual "Editor" screen, with live reconciliation against the physical check amount), plus an "Estado de pago" filter, a "sin fondos" row tag and a Resolver dialog on `/estado-cuenta`. No new abilities. Verified end-to-end against dev, API + browser.
**Two pre-existing bugs found and fixed while building it:** (a) `statement()` filtered `legacySourceTable: { notIn: [...] }`, which compiles to SQL `NOT IN` — and `NULL NOT IN (…)` is NULL, so **every app-captured movement was invisible on the customer statement** (438 rows in the movement browser vs 392 on the statement) while still appearing everywhere else. This would have made the whole receipt-capture feature look broken to staff. Now NULL-safe. (b) The balances *count* query omitted the void filter its own page query applied, so the row count disagreed with the rows.
**OCR seam:** `BillingService.createBatch(dto, opts)` is the single multi-row write path and carries three contract guarantees for the step-11 OCR module to post through — `items[i]` maps to `lines[i]` (so `StatementDocument.postedTransactionId` can be zipped back on), `opts.refs[i]` stamps `captureRef` with a duplicate-post guard that a *voided* row deliberately does not block, and `opts.source` is service-level only so an HTTP client cannot label hand-keyed rows as machine-captured. Backed by a new `TransactionCaptureSource` enum (MANUAL/BATCH/OCR) + `captureRef`, both nullable so the 40,136 migrated rows stay NULL rather than being mislabelled.
- **PDF/OCR auto-capture — DONE** (2026-08-01). As-built write-up in [`docs/STATEMENT_OCR.md`](docs/STATEMENT_OCR.md); the design and the measured evidence stay in the spec's §2. The ingest→split→OCR→match→review pipeline for the 300+/month/service-provider statements staff key in by hand, built in `apps/api/src/statements/` and posting through §1.2's `createBatch` seam with `source: "OCR"` and a per-document `captureRef`. Web: `/recibos` + `/recibos/:id`. Abilities `statement:ingest`/`statement:review` (STAFF — the review step is what makes machine capture safe at that tier). OCR is self-hosted **Tesseract** behind a swappable `OcrProvider` interface; `tesseract-ocr`, `tesseract-ocr-data-spa` and `poppler-utils` were added to the API image.
**Every decision was driven by 10 real scans (46 pages).** Shipped-parser results on them: provider 46/46, account ref 43/46, amount 42/46, due date 44/46 — and against the dev database **39/46 (85%) exact auto-match, 40/46 (87%) identified**, the rest genuine review cases. The scans are pure images (no text layer), so OCR is mandatory, and they arrive **bundled one customer per page**.
**The three gaps are closed, and two of them were mis-stated in the spec.** (a) `TELEPHONE` now exists and is backfilled from `Property.phone1` only — coverage is 534/18/1 across phone1/2/3, so phone is one billed line per property, not three. (b) **Clave catastral ≠ predial**: `DATMEX.clave` (934 rows, `KA903009`) is what CESPT and predial bills actually print, while `predial` — what `PROPERTY_TAX.accountNumber` holds — has only 663 distinct values across 1135 rows and appears on no statement; the clave now lives on `Property.cadastralKey` as the matcher's secondary key and predial is left untouched. (c) Gas was **not** a dead end: 160 of the 334 `DATMEX.gas` values are real account numbers (the rest are `ESTACIONARIO`/`CILINDRO` descriptors), all recovered into `GAS.meterNumber`.
**Matching is scoped per service kind and never reads the customer name** — a CESPT receipt prints `ARNAIZ ROSAS ELSA AURORA` for an account this office holds under `CATT, RANDY`, because the name on a utility bill is the registrant, not the current owner. Normalisation is per provider: CFE strips leading zeros off `NO. DE SERVICIO`, Telnor strips the 664 LADA down to the stored local 7 digits. Where a provider prints a payment barcode it is preferred over the printed label (one CFE label OCR'd a digit too many while its barcode was correct) and the two are cross-checked, with disagreement forcing review. Confirming a document whose service had no reference writes it back, so gas and any other cold start is a one-time cost.
- **Policy OCR capture — DONE** (2026-08-01), **unplanned — it came out of building the bullet above.** Full write-up in [`docs/POLICY_OCR.md`](docs/POLICY_OCR.md). Once the receipt pipeline existed it was obvious the same render→OCR→parse→match→review shape fits the *other* stack of paper this office keys in by hand: the carrier policy PDFs behind every `Policy` row. Built in `apps/api/src/policy-ocr/` with a GMX parser, `policy_ocr_batches`/`policy_ocr_documents`, and abilities `policy:ingest`/`policy:ocr-review` (STAFF, same trust tier and same reason). Web: `/polizas/captura` is the "automática" tab of the policy-creation screen (`/polizas/nuevo` is the manual one, both render `PolicyCaptura.tsx`) with the review queue at `/polizas/captura/[id]`. The `OcrProvider` seam was **extracted out of `StatementsModule` into its own `OcrModule`** to make this possible — that was blocking, not cosmetic; `StatementsModule` now imports it and binds nothing.
**The statement pipeline's core assumption inverts here.** Utility statements arrive bundled *one customer per page*, so there a page is a document; a GMX certificate is one policy across two pages (header on 1, coverage table on 2), so the pipeline concatenates the pages and runs the parser and matcher **once per file**. `PolicyOcrDocument.pageNumber` is therefore the file ordinal in the batch, and `storageKey` points at the **source PDF** (the review screen embeds the exact artifact the office received) rather than at a page image. Matching is on `Policy.policyNumber` alone and never the printed insured name — the same registrant-vs-owner drift that rules names out on the utility side. Zero hits means a new policy and confirm creates it; more than one is surfaced, never auto-picked.
**The GMX certificate carries no premium at all** — the figure lives on a separate `recibo` PDF — so the premium fields stay null with a note saying why, confirm never overwrites an existing premium with null, and the optional ledger write is gated on staff ticking `postPremium` *and* a premium actually parsing. 8/8 parser tests against one real document (`HC_Folio_000767_Traduccion.pdf`). GMX is the only carrier implemented; the dispatcher is a pattern table, so a second one is a parser function and two entries.
- **Multi-bank chequera — DONE** (2026-07-27). `Bank`/`BankAccount` models so Seguros (US bank) and Utilities (Mexican bank, currently SCOTHIA) can each have their own register. `bank_transactions` gained a **required** `bankAccountId` (plus an `(bankAccountId, transactionDate)` index, since every read is now filtered by account and ordered by date), and all 22,669 existing rows were backfilled onto a seeded "Utilities — Scotiabank (MXN)" account by `migration/backfill_bank_accounts.py` — a standalone step because `prisma db push` cannot add a required column to a populated table. It is idempotent and now runs inside `run_all.py` (both normal and `--sync`) ahead of `transform_bank.py`, which fails fast if the account is missing. Every read path in `bank.service.ts` is account-scoped, including `facets()` (which had no filter at all) and *both* raw-SQL rollups in `summary()`. API: `?bankAccountId=` is required on `list`/`stats`/`facets`/`summary`**not** optional-with-an-all-accounts-default, since summing an MXN and a USD register repeats exactly the currency-collapsing mistake the billing module exists to prevent — plus a new `bank/accounts` + `bank/banks` sub-resource under a MANAGER `bank:manage-accounts` ability. Web: `/banco` gained an account picker (remembered per browser) and reads every figure in the selected account's currency, `/banco/cuentas` manages banks and accounts, and `/inicio`'s chequera card names the account it is showing instead of implying one register. An account's `currency` is immutable after creation by design — its booked movements are denominated in it. Verified against dev + browser: a second USD account showed full read/write isolation from the MXN register, whose totals were unchanged.
- **Customer-number recycling** — promotes the legacy `NUM id` (currently only inside `customer_legacy_refs`) into a first-class, reusable `Customer.customerNumber`, automates *finding* candidates for reuse (cancelled / 1-year-inactive), and auto-assigns the lowest free number at creation — the search is automated, the release/reuse decision stays a human action. Backfill needs care: ~140 utilities rows and all insurance-only customers have no real legacy number (synthetic `rownum_N`/`insrow_N` placeholders in `transform_customers.py`, not real `NUM id`s).
Several open questions block parts of this (OCR provider/budget, the Seguros bank's identity, the clave-catastral-vs-predial mismatch, exact recycling triggers, and whether "recycling" should ever mean true data purge vs. archive-and-reuse-the-number) — see the spec's collected open-questions section.
12. **Insurance features — one of four built, rest spec'd.** Full design in [`docs/INSURANCE_FEATURES_SPEC.md`](docs/INSURANCE_FEATURES_SPEC.md), the insurance half of the same 2026-07-25/26 meeting with Jorge that produced step 11:
- **Renewal notification emails — DONE** (2026-08-01, extended 08-02). A sweep that mails the customer 30 days before expiry, 15 days before, and 7 days after, mapping onto `RenewalNotice.generation` 1/2/3 with **no schema change**. Sending is **Amazon SES** (`@aws-sdk/client-sesv2`, mirroring `StorageService`'s optional-client/degrade-don't-crash pattern). The letter body is the *existing* `aviso-renovacion` report; `@@unique([policyId, generation])` is already-in-place idempotency, so a re-run cannot double-send. Volume ≈260 mails/month, and **815 of the 893 policyholders (91%) have an email**.
**Three things came out differently from the spec.** (a) The manual mark-as-sent mutation was **dropped on purpose** — a button that marks a notice sent without sending anything lets the list claim a customer was told when they were not. `POST /renewals/send` replaced it: sending from the list *is* the marking, and the report's `enviadas` total becomes real the same way. (b) The send history is **not renewal-specific** — every attempt, including the failures and no-email skips a `RenewalNotice` row cannot represent, also writes `email_notification_log` as `RENEWAL_NOTICE`/`POLICIES`, shared with the four bulk jobs from [`docs/MASS_EMAIL_NOTIFICATIONS.md`](docs/MASS_EMAIL_NOTIFICATIONS.md). `RenewalNotice` stays *gating* state; the log is *history*. (c) The `@Cron("0 6 * * *")` literal the spec called for lasted one day: both this sweep and the servicios jobs now take their cadence from `NotificationScheduleService`, stored in `app_settings` and reinstalled on save — no redeploy. Defaults preserve the old behaviour (pólizas 06:00 daily, servicios off).
**Both halves live on one screen.** `/notificaciones` has Servicios and Pólizas tabs over the one log; `/renovaciones` is an alias onto the Pólizas tab. The send flags (`debug` in particular) sit in the shell above the tabs and govern both — before that there was no way to test a renewal aviso without mailing a real customer. A debug send diverts the mail, skips the `RenewalNotice` upsert **and** does not advance the sweep's `lastSuccessfulAt`; all three are needed together, or a test run silently narrows tomorrow's window and drops the letters it only pretended to send.
**Production status:** the `SES_*` Gitea secrets were created 2026-08-02, clearing the last blocker — but the feature has not shipped yet (master is well past the newest tag) and nothing has confirmed that `SES_FROM` is a verified SES identity or that the account is out of the sandbox. Run the first sweep with `debug` on. See [`docs/BACKLOG.md`](docs/BACKLOG.md) §0.
- **Liquidación batch workflow** — ~70% already built (`liquidated`/`liquidationNumber`/`liquidationDate` are wired through DTOs, list filter, stats, form and detail page); only the *batch* print-and-mark step is missing, against a live pending set of 226 policies. Adds a ramo-parameterized pending report plus `POST /policies/liquidate-batch` under a new MANAGER `policy:liquidate` ability. Parameterized by ramo, not MULT-only — legacy `TABLA LIQUIDA MF` served `MULT`, `INCENDIO` and `M EMPR` alike.
- **Certificate / "Solicitud Atlas"** — renders from the same `format: "letter"` machinery `aviso-renovacion` uses, then reaches customers as an extension of the step-8/9 replication (PDF generated here, pushed to MinIO, pointer replicated), **not** as a new public surface in this repo. Half-blocked: "Solicitud" has zero referent in the legacy system and normally means an *application form*, a different artifact from a certificate.
- **Carrier API integration (ANA Seguros + GMX)** — shape only (`CarrierConnector` + an import-review queue rather than direct `Policy` writes, matching how step 11's OCR results are routed). Carrier research done 2026-07-27: **the two carriers are one company** — both belong to **Grupo Valore** (ANA writes autos, GMX writes daños, which is exactly this database's `AUTO`/`LICENCIAS` vs `MULT`/`INCENDIO`/`M_EMPR` split), so it is one commercial relationship, not two. **ANA has a real live SOAP service** (`server.anaseguros.com.mx/ananetws/service.asmx`, ASP.NET `.asmx`) with a published operation list — catalogs, `CalculaValor`/`CalculaMSI`, `ValidaSerie`, `RecuperaCotizacion`, `Transaccion`. **GMX publishes no machine interface at all**, only human agent portals. ⚠️ **Critical mismatch:** every ANA operation serves *new-business quoting/issuance*, not "list the policies where I am agent of record" — so if the ask is inbound portfolio sync, no evidence exists that either carrier sells it. Blocked on one phone call to Grupo Valore ((55) 5480-4000) for credentials + a direction answer, not on further research. ("GDMX" in the meeting notes was a typo for `GMX` — confirmed 2026-07-27.)
**Two pre-existing defects were found while verifying this spec and should be fixed as part of the liquidación work:** (a) `policy_types` is missing its `INCENDIO` and `M_EMPR` rows and, because `policies_policyTypeId_fkey` is `ON DELETE SET NULL`, 5 `m_empr` policies silently lost their ramo — 4 of them are pending liquidación and are invisible to every ramo-filtered query; (b) the legacy settlement slots don't match what the target model assumed — `MULT`/`INCENDIO` carry two and `M EMPR` carries four, while `Policy` collapses to one, so ≤41 MULT second settlements were dropped in migration. Spec recommends moving settlement onto `PolicyPaymentInstallment` rather than adding a second slot.
**One long-standing open question is closed by this spec:** `DATGRAL.[NUM UTIL]` is authoritative for Utilities↔Seguros reconciliation and **`UTILSEG` must not be used** — its numbers resolve to unrelated people under every reading tested (name match 58/1,024 vs. 298/563 for `NUM UTIL`), and where the two sources overlap they contradict each other on 170 of 218 shared ids. This matters to step 11's customer-number recycling, which touches the same identity space.
## Status ## Status
@@ -139,6 +172,16 @@ Repo scaffolded at `jorgecuadros-platform/`: npm workspaces, NestJS API with a r
**Portal live DB now in hand.** `utility_dbo.sql` (1.3 GB, 55 tables) and the portal codebase `my-jorgecuadros-web` (PHP/`mysqli`, Gitea repo, themed classic/modern, ~397 PHP files, core in `scripts/functions.php`) are both on disk — resolving the long-standing "`utility_dbo` schema unknown" blocker. Sync-relevant tables identified: statements/money (`utility_bills`, `accounting`, `email_alert_log`), customer/property (`home_owners`, `home_index`, `condominium`, `management`, `hoa_management`, `trust_assist`), portal-facing policy views (`fm2`/`fm3`/`fmt`, `full_coverage`, `mx_liability`, `usa_liability`), and portal write points (`peticion_gas`, PayPal payments, `notifications_settings`, `verification_codes`). A second dump, `jorgecuadros.sql` (38 MB, 11 tables — `pagos`/`pagosemail`/`PROPANO`/`TRUSTVENCE`/etc.), appears to be an older/partial export, not the portal live DB. **Portal live DB now in hand.** `utility_dbo.sql` (1.3 GB, 55 tables) and the portal codebase `my-jorgecuadros-web` (PHP/`mysqli`, Gitea repo, themed classic/modern, ~397 PHP files, core in `scripts/functions.php`) are both on disk — resolving the long-standing "`utility_dbo` schema unknown" blocker. Sync-relevant tables identified: statements/money (`utility_bills`, `accounting`, `email_alert_log`), customer/property (`home_owners`, `home_index`, `condominium`, `management`, `hoa_management`, `trust_assist`), portal-facing policy views (`fm2`/`fm3`/`fmt`, `full_coverage`, `mx_liability`, `usa_liability`), and portal write points (`peticion_gas`, PayPal payments, `notifications_settings`, `verification_codes`). A second dump, `jorgecuadros.sql` (38 MB, 11 tables — `pagos`/`pagosemail`/`PROPANO`/`TRUSTVENCE`/etc.), appears to be an older/partial export, not the portal live DB.
**Step 11 is now three-quarters built.** Receipt capture, the multi-bank chequera and PDF/OCR auto-capture are all done and verified; only customer-number recycling remains unbuilt. `docs/RECEIPT_CAPTURE_SPEC.md` carries a BUILT note per section recording what shipped and, for §2, the four things real scanned statements proved the spec had wrong or unknown.
Each of the two OCR intakes now has an as-built doc separate from its spec — `docs/STATEMENT_OCR.md` and `docs/POLICY_OCR.md`. The specs record what was designed and why; those record what is in the code. They share one `OcrProvider` seam (`apps/api/src/ocr/`), so the Tesseract-vs-managed-API decision is one line for both.
**It also produced a feature nobody planned.** The statement OCR pipeline generalised: the same render→OCR→parse→match→review shape reads **carrier policy PDFs** into `Policy` rows, which is `docs/POLICY_OCR.md` (built 2026-08-01, GMX only so far). It belongs to step 12's subject matter but to step 11's lineage, and it is in no spec — worth knowing before reading `INSURANCE_FEATURES_SPEC.md`, which does not mention it. It also partly overlaps what §4's carrier API was wanted for, and unlike that section it is not blocked on a phone call.
**Step 12 is one-quarter built.** `docs/INSURANCE_FEATURES_SPEC.md` covers the insurance half of the same meeting (renewal emails, liquidación batch, certificate + portal delivery, carrier APIs) — see Build sequencing step 12 above. Verified the same way, plus a live query of the dev DB for the counts it quotes (email coverage, pending liquidación, installment fill rates) and of the staged Parquet for the legacy settlement-slot usage. **§1 renewal emails is done** (2026-08-01/02) and carries a BUILT note recording the three places the build diverged from the spec; §2 liquidación is still the smallest remaining piece, since the per-policy fields are already wired end to end.
**Notifications are one screen, not two features.** The four legacy mass-email jobs (`docs/MASS_EMAIL_NOTIFICATIONS.md`) and the insurance renewal avisos both mean "tell a customer something by email", so they are tabs of `/notificaciones` over one `email_notification_log`, with one shared flags panel and one schedule editor. `app_settings` + `SettingsService` (db → env → default) is the operator-config seam they introduced: summary recipients and both sweep cadences live there, so changing any of them is a save, not a redeploy. Credentials stay in the environment.
## Decisions (locked) ## Decisions (locked)
- **Stack:** Next.js + NestJS + Prisma + **MySQL** (locked earlier — see engine rationale above). - **Stack:** Next.js + NestJS + Prisma + **MySQL** (locked earlier — see engine rationale above).
@@ -155,6 +198,30 @@ Repo scaffolded at `jorgecuadros-platform/`: npm workspaces, NestJS API with a r
- **VPS provisioning:** provider (Hetzner vs DigitalOcean), size, and Tailscale + MySQL replica setup on it — an ops task, still pending. Design is settled; only the box is missing. - **VPS provisioning:** provider (Hetzner vs DigitalOcean), size, and Tailscale + MySQL replica setup on it — an ops task, still pending. Design is settled; only the box is missing.
- **Old external-DB credential** (hardcoded plaintext MySQL password in the old repo's `dbConnection.php`, in git history) — rotate it regardless, since it's already exposed. - **Old external-DB credential** (hardcoded plaintext MySQL password in the old repo's `dbConnection.php`, in git history) — rotate it regardless, since it's already exposed.
## Open design questions (steps 11 & 12 — need Jorge before/while building)
Unlike the ops items above, these block design decisions, not just infrastructure. Full detail in each section of `docs/RECEIPT_CAPTURE_SPEC.md` (step 11) and `docs/INSURANCE_FEATURES_SPEC.md` (step 12):
**Step 11 — utilities/ops side:**
- ~~OCR provider/budget~~ — **CLOSED**: self-hosted Tesseract, chosen on measured accuracy against real scans, so there is no per-page cost to approve.
- ~~Whether `PROPERTY_TAX.accountNumber` (from `DATMEX.PREDIAL`) is the same number as "Clave Catastral" (`DATMEX.CLAVE`)~~ — **CLOSED**: they are different numbers. Answered from real CESPT bills plus the staged data; the clave is now migrated separately and predial was left alone.
- Whether the CFE figure to charge is the rounded headline/barcode amount (`$268` — what is actually paid at the window) or the exact breakdown `Total` (`$268.88`). The parser takes the barcode amount; one confirmation from Jorge would settle it.
- The actual bank name/currency/details for the Seguros USD account, and whether any historical Seguros bank register exists to migrate. (Multi-bank support itself is **built** — this is now only the missing content: staff can open the account in `/banco/cuentas` the moment the answer arrives, and it starts empty unless a historical register turns up.)
- The exact "1 year inactivity" / "cancelled" triggers for customer-number recycling eligibility.
- Whether customer-number recycling should ever include true PII purge (matching the office's paper-world habit) or archive-and-reuse-the-number is sufficient — recommended default is archive-only, consistent with this project's existing never-hard-delete convention.
**Step 12 — insurance side:**
- Which SES region + verified sending identity/configuration set the renewal mail goes out under, and whether it reuses the existing IAM credentials or gets its own scoped `ses:SendEmail` user. (Provider and budget are *not* open — SES is settled.)
- What to do with the 78 policyholders who have no email on file: skip silently, or produce a print worklist? Recommended: the worklist, since `aviso-renovacion` already renders exactly those letters.
- Whether renewal notices go out in Spanish or English — `Customer` carries no language preference.
- What "garantías" refers to — it has zero referent in the legacy data, and it blocks the liquidación batch's exclusion filter.
- Whether policy settlement should move onto `PolicyPaymentInstallment` (recommended) or gain a second slot on `Policy`, and whether to backfill the ≤41 MULT second settlements lost in migration.
- Whether batch liquidación warrants a new MANAGER-level `policy:liquidate` ability (recommended) or should reuse the existing STAFF-level `policy:update`.
- **What "Solicitud Atlas" actually is** — an application form or a certificate. These are different artifacts with different data and timing; this blocks the whole certificate feature.
- **Carrier integration direction** — outbound quote/issue (which ANA's SOAP service supports today) or inbound sync of the office's existing book (which nothing found suggests either carrier offers)? This decides whether the feature is buildable at all. Bundle with the other three carrier questions into one call to Grupo Valore ((55) 5480-4000): WSDL + credentials for the ANA service, whether a cartera/portfolio download exists for an agent's own book, whether GMX daños has any machine interface, and whether one credential spans both carriers. ("GDMX" is resolved — it was a typo for `GMX`.)
## Verification ## Verification
- Migration: automated row-count/sum reconciliation between `staging` and final schema per table group (see step 5 above), run as part of the migration script, not a manual spot-check. - Migration: automated row-count/sum reconciliation between `staging` and final schema per table group (see step 5 above), run as part of the migration script, not a manual spot-check.
+47 -4
View File
@@ -4,7 +4,8 @@ Internal platform for a Baja California insurance brokerage and property-service
firm: a single expedient joining each client's **properties/services**, firm: a single expedient joining each client's **properties/services**,
**insurance policies**, **account statement**, and the firm's **checkbook**. **insurance policies**, **account statement**, and the firm's **checkbook**.
It replaces a legacy PHP/Access app (see `RESUME.md` and `PLAN.md` for the full It replaces a legacy PHP/Access app (see `RESUME.md` and `PLAN.md` for the full
history and rebuild rationale). history and rebuild rationale, and [`docs/BACKLOG.md`](docs/BACKLOG.md) for
everything still outstanding).
The UI is Spanish-first; the codebase and this document are in English. The UI is Spanish-first; the codebase and this document are in English.
@@ -40,8 +41,21 @@ docker-compose.yml mysql + api + web
``` ```
API feature modules: `auth`, `users`, `customers`, `policies`, `properties`, API feature modules: `auth`, `users`, `customers`, `policies`, `properties`,
`billing`, `bank`. Web routes: `/clientes`, `/polizas`, `/servicios`, `billing`, `bank`, `reports`, `notifications`, `renewals`, `mail`, `statements`,
`/estado-cuenta`, `/banco` (chequera), `/catalogos`, `/usuarios`, `/login`. `policy-ocr`, `ocr`, `storage`, `settings`, `ops`.
Web routes: `/inicio`, `/clientes`, `/polizas` (+ `/polizas/captura`, policy
PDF OCR capture), `/servicios`, `/estado-cuenta`, `/banco` (chequera),
`/recibos` (utility statement OCR capture), `/notificaciones` (mass email +
renewal avisos; `/renovaciones` is an alias onto its Pólizas tab), `/reportes`,
`/catalogos`, `/operaciones` (DB ingest/backup, ADMIN), `/usuarios`, `/login`.
Two OCR intakes share one `OcrProvider` seam (`src/ocr/`, Tesseract today):
utility statements → ledger rows ([`docs/STATEMENT_OCR.md`](docs/STATEMENT_OCR.md))
and carrier policy PDFs → `Policy` rows ([`docs/POLICY_OCR.md`](docs/POLICY_OCR.md)).
Both need `tesseract-ocr`, `tesseract-ocr-data-spa`, `poppler-utils` and object
storage; each reports its own availability and disables only itself if either
is missing.
--- ---
@@ -82,7 +96,14 @@ NEXT_PUBLIC_API_ORIGIN=http://localhost:3001
``` ```
The API loads `DATABASE_URL`, `SESSION_SECRET`, `WEB_ORIGIN`, and optional The API loads `DATABASE_URL`, `SESSION_SECRET`, `WEB_ORIGIN`, and optional
`PORT` (default `3001`). The web app only needs `NEXT_PUBLIC_API_ORIGIN`. `PORT` (default `3001`). `WEB_ORIGIN` is comma-separated — list every origin the
app is reached under, or credentialed fetches from the missing ones fail CORS.
The web app needs no API URL of its own: the browser derives it from the page it
loaded (same host on port `3001` over plain HTTP, or the same-origin `/api` path
behind a TLS proxy). Set `NEXT_PUBLIC_API_ORIGIN` (dev) or `API_ORIGIN` (deploy,
read at request time) only to override that — for instance when running the API
on a non-default port.
### 3. Start MySQL ### 3. Start MySQL
@@ -195,6 +216,28 @@ python migration/run_all.py
--- ---
## Scheduled jobs
The API runs two automatic email sweeps. Neither cadence is in the source:
both are stored in `app_settings` and edited at `/notificaciones`
"Programación de envíos" (ADMIN, `setting:manage`), taking effect immediately
without a restart. Shipped defaults:
| Job | Default | What it does |
| --- | ------- | ------------ |
| Pólizas | **on**, 06:00 daily (America/Tijuana) | Renewal avisos at 30/15 days before expiry and 7 days after. |
| Servicios | **off** | All four mass-email jobs in order, same as "Ejecutar todos". |
A scheduled run never uses the UI's send flags — in particular it ignores
`debug`, so a forgotten test toggle cannot silently stop customer mail. Full
detail in [`docs/MASS_EMAIL_NOTIFICATIONS.md`](docs/MASS_EMAIL_NOTIFICATIONS.md).
Sending needs `SES_*` in the environment. Without it the API still boots and
logs mail to stdout in dev; in production every send fails loudly and is
recorded as `FAILED` rather than quietly going nowhere.
---
## Production notes ## Production notes
- Use `pnpm --filter @jorgecuadros/database exec prisma migrate deploy` if/when - Use `pnpm --filter @jorgecuadros/database exec prisma migrate deploy` if/when
+324 -26
View File
@@ -127,8 +127,14 @@ To rerun (from `migration/`, venv at `migration/.venv`):
```bash ```bash
./.venv/bin/python load_staging.py --output-dir ./output # re-extract from Access (needs mdbtools + the source files) ./.venv/bin/python load_staging.py --output-dir ./output # re-extract from Access (needs mdbtools + the source files)
./.venv/bin/python run_all.py --env dev # full transform+load; add --stage to re-extract first ./.venv/bin/python run_all.py --env dev # full transform+load; add --stage to re-extract first
./.venv/bin/python run_all.py --env dev --sync # additive sync: upsert legacy by provenance, keep manual rows, prune legacy empties
``` ```
`--sync` mode (Phase B) upserts legacy-owned rows by their provenance keys and preserves
manual rows (`legacyId IS NULL`); every transform reuses each row's existing PK, rebuilds
legacy-owned children by scoped delete + reinsert, and drops legacy rows gone from source.
Verified end-to-end against dev 2026-07-24 — see §6 item 6.
## 5. Infrastructure & sync architecture (designed, not yet built) ## 5. Infrastructure & sync architecture (designed, not yet built)
- **Internal server** — on-prem, private IP `192.168.1.xx`, no inbound internet exposure. Runs the platform + canonical MySQL (source of truth). - **Internal server** — on-prem, private IP `192.168.1.xx`, no inbound internet exposure. Runs the platform + canonical MySQL (source of truth).
@@ -162,17 +168,30 @@ the reconciliation pass (done, then corrected) are all closed. See §3 and §8.
5. **`TRASPASOS PAYPAL` is a clearing account, not a customer** — carries -7.03M MXN over 5. **`TRASPASOS PAYPAL` is a clearing account, not a customer** — carries -7.03M MXN over
309 movements and therefore tops the adeudo worklist. Deliberately not special-cased in 309 movements and therefore tops the adeudo worklist. Deliberately not special-cased in
code; needs a business decision on how to model it. code; needs a business decision on how to model it.
6. **DB Operations — Phase B (additive sync) — IMPLEMENTED, verification pending.** Phase A provides 6. **DB Operations — Phase B (additive sync) — VERIFIED END-TO-END against dev DB 2026-07-24.**
the admin-only `/operaciones` page + `ops` API module (ability `db:manage`, ADMIN), ingest Phase A provides the admin-only `/operaciones` page + `ops` API module (ability `db:manage`,
folder, backup, restore, and destructive re-import. Phase B now enables `SYNC`: `OpsService` ADMIN), ingest folder, backup, restore, and destructive re-import. Phase B enables `SYNC`:
creates a safety backup and runs `run_all.py --sync`; transforms upsert legacy-owned rows by `OpsService` creates a safety backup and runs `run_all.py --sync`; transforms upsert
provenance keys while preserving existing PKs and rows whose `legacyId IS NULL` (manual). legacy-owned rows by provenance keys while preserving existing PKs and rows whose
Prisma now enforces provenance uniqueness for properties, policies, transactions, vehicles, `legacyId IS NULL` (manual). Prisma enforces provenance uniqueness for properties, policies,
and bank transactions. Sync intentionally skips prune/blob steps so manual customers and transactions, and bank transactions (the vehicle unique was **removed** — one legacy policy
document pointers are not removed. Python compilation plus API/web production builds pass; row carries up to 3 vehicles that share a `legacyId`, so provenance is not unique per
still required before production use: push updated Prisma schema and run an end-to-end sync vehicle; vehicles are rebuilt by scoped delete + reinsert). Sync skips blob extraction, and
against a disposable/dev DB proving stable PKs, manual-row preservation, changed-row updates, runs a **manual-safe prune** (`prune_empty_customers.py --sync` — only prunes empties that
and legacy-delete handling. carry a legacy ref, never manually-added customers) because the customer upsert otherwise
re-creates every previously-pruned empty from Parquet.
**The as-written sync was broken and had never been run; a batch of bugs were fixed on
2026-07-24 before it passed** (fresh-uuid child FKs in policies/properties, unconditional
child inserts, a `zip(customers, refs)` mispairing in transform_customers, invalid vehicle
unique, lookup tables built with fresh uuids but never upserted, a `updatedAt=NOW()` on a
table with no such column, and report crashes on NULL `legacySourceTable` for manual rows).
Verified with `migration/` `verify_sync.py`-style harness: two consecutive `run_all.py --sync`
runs both exit 0 and pass 32/32 assertions (stable PKs, manual-row preservation, changed-row
updates, legacy-delete, no child duplication, zero FK orphans), idempotent (customers stable
at 1537). Schema pushed to dev, Prisma client regenerated, API build clean. Migration/web
changes uncommitted as of this update. Still open before production: run the same sync from
the `/operaciones` UI (OpsService path) and against a prod-shaped DB.
## 7. Environment notes (current macOS machine) ## 7. Environment notes (current macOS machine)
@@ -181,9 +200,12 @@ the reconciliation pass (done, then corrected) are all closed. See §3 and §8.
- **`npm` is pnpm-aliased**, and pnpm ignores the `workspaces` field. Consequences: - **`npm` is pnpm-aliased**, and pnpm ignores the `workspaces` field. Consequences:
- there is **no root `node_modules/.bin`**. Binaries live per-app: `apps/api/node_modules/.bin/nest`, `apps/web/node_modules/.bin/next`. - there is **no root `node_modules/.bin`**. Binaries live per-app: `apps/api/node_modules/.bin/nest`, `apps/web/node_modules/.bin/next`.
- Prisma CLI is run as `npx prisma@5`. - Prisma CLI is run as `npx prisma@5`.
- **Dev servers** (both must be up to use the UI): - **Dev servers** (both must be up to use the UI). ⚠️ **Ports come from the env files, not the
- API `cd apps/api && ./node_modules/.bin/nest start --watch``:3001` framework defaults** — `apps/api/.env` sets `PORT=4501` and `WEB_ORIGIN=http://localhost:4500`,
- Web `cd apps/web && ./node_modules/.bin/next dev``:3000` and `apps/web/.env.local` points at `NEXT_PUBLIC_API_ORIGIN=http://localhost:4501`. This doc
said `:3001`/`:3000` until 2026-07-27; that was wrong and cost a debugging detour.
- API `cd apps/api && ./node_modules/.bin/nest start --watch`**`:4501`**
- Web `cd apps/web && ./node_modules/.bin/next dev -p 4500`**`:4500`**
- Dev login: `admin@jorgecuadros.local`, password from `apps/api/scripts/seed-user.mjs` (`SEED_PASSWORD` env overrides the default). - Dev login: `admin@jorgecuadros.local`, password from `apps/api/scripts/seed-user.mjs` (`SEED_PASSWORD` env overrides the default).
- **Dev DB**: `192.168.4.212:3307` (cubex Swarm stack `jorgecuadros-dev-db`). Credentials in gitignored `deploy/.env.dev`. **MinIO** for documents: `192.168.4.212:9100`, bucket `jorgecuadros-documents`. - **Dev DB**: `192.168.4.212:3307` (cubex Swarm stack `jorgecuadros-dev-db`). Credentials in gitignored `deploy/.env.dev`. **MinIO** for documents: `192.168.4.212:9100`, bucket `jorgecuadros-documents`.
@@ -359,9 +381,18 @@ for what's actually next.
insurance/servicios/fideicomiso split the migration comment implied. A insurance/servicios/fideicomiso split the migration comment implied. A
classifier would invent data, so `categoryId` stays null and the module does classifier would invent data, so `categoryId` stays null and the module does
not filter on it. Register is browsable by date/payee/amount/cheque instead. not filter on it. Register is browsable by date/payee/amount/cheque instead.
(c) **Single currency (MXN).** `bank_transactions` has no currency column and (c) ~~**Single currency (MXN).**~~ **SUPERSEDED 2026-07-27 by the multi-bank
every `amountInWords` is spelled out in PESOS — so, unlike the customer chequera** (step 11, `docs/RECEIPT_CAPTURE_SPEC.md` §3). The office keeps
ledger, everything here is one currency and not split per-currency. more than one register, so `bank_transactions` now carries a **required**
`bankAccountId` and every read in the module is scoped to exactly one
`BankAccount`, whose `currency` the movements inherit — there is still no
currency column on the movement itself, because a real bank account doesn't
mix currencies. All 22,669 migrated rows are the Utilities/Scotiabank MXN
account (backfilled by `migration/backfill_bank_accounts.py`, which
`run_all.py` runs before `transform_bank.py`), which is why every
`amountInWords` is still spelled out in PESOS. There is deliberately no
"all accounts" option: summing an MXN and a USD register would repeat the
currency-collapsing mistake the billing module warns against.
(d) **The "acumulado" is net movement since the register opened, not a bank (d) **The "acumulado" is net movement since the register opened, not a bank
balance** — SCOTHIA carries no opening balance (its `ban` table holds only the balance** — SCOTHIA carries no opening balance (its `ban` table holds only the
bank's name), so the running total starts at 0 in 2013. Labelled as such in bank's name), so the running total starts at 0 in 2013. Labelled as such in
@@ -369,9 +400,18 @@ for what's actually next.
(e) Sign convention (from `transform_bank.py`): positive = ingreso, (e) Sign convention (from `transform_bank.py`): positive = ingreso,
negative = egreso, exactly zero = a cancelled/void cheque (787 of 791 say negative = egreso, exactly zero = a cancelled/void cheque (787 of 791 say
CANCELADO/VOID) — voids are excluded from both the income and expense sides. CANCELADO/VOID) — voids are excluded from both the income and expense sides.
(f) **Multi-account since 2026-07-27.** `/banco` opens on an account picker
(the last account is remembered per browser) and reads every figure in that
account's currency; `/banco/cuentas` manages banks and accounts under a new
MANAGER `bank:manage-accounts` ability. Accounts are never deleted — the
`bankAccountId` FK is required, so a used account can only be *closed*
(`active: false`), which hides it from new captures but keeps its history
readable. An account's currency is immutable after creation, since its
booked movements are denominated in it.
- Full pipeline reproducible in one command: `run_all.py --env <env>` runs customers → - Full pipeline reproducible in one command: `run_all.py --env <env>` runs customers →
properties → policies → transactions → prune → bank → blobs in order (all idempotent); properties → policies → transactions → prune → bank accounts → bank → blobs in order
add `--stage` to re-extract from the Access files first. Verified end-to-end against dev. (all idempotent); add `--stage` to re-extract from the Access files first. Verified
end-to-end against dev.
5. **Infra****DONE.** Dev MySQL deployed to the cubex Swarm via the Portainer API as stack 5. **Infra****DONE.** Dev MySQL deployed to the cubex Swarm via the Portainer API as stack
`jorgecuadros-dev-db` (MySQL 8.4, `192.168.4.212:3307`, node `cubex` labeled `jorgecuadros-dev-db` (MySQL 8.4, `192.168.4.212:3307`, node `cubex` labeled
@@ -385,15 +425,273 @@ for what's actually next.
--- ---
- **Sync implementation — DONE, validation pending.** `run_all.py --sync` performs the - **Sync implementation — DONE + VALIDATED end-to-end against dev 2026-07-24.** `run_all.py
non-destructive legacy upsert path for customers, properties, policies, transactions, and --sync` performs the non-destructive legacy upsert path for customers, properties, policies,
bank rows. It preserves manual rows and stable legacy-owned primary keys; the admin SYNC job transactions, and bank rows, plus a manual-safe empty-customer prune. It preserves manual rows
automatically creates a pre-sync backup. Next validation: apply schema changes, then exercise and stable legacy-owned primary keys; the admin SYNC job auto-creates a pre-sync backup. The
sync against a disposable DB with added, changed, removed, and manually-created rows. as-written code was broken and had never been run — a batch of bugs was fixed before it passed
(see §6 item 6). Two consecutive syncs both exit 0 and pass 32/32 assertions (added, changed,
removed, and manually-created rows), idempotent. Remaining: exercise the same path from the
`/operaciones` admin UI and against a prod-shaped DB.
- **Plan step 9: portal sync worker** remains separate and blocked on VPS provisioning. This - **Plan step 9: portal sync worker** remains separate and blocked on VPS provisioning. This
Phase B feature synchronizes Access source files into the internal platform; it does not yet Phase B feature synchronizes Access source files into the internal platform; it does not yet
poll `utility_dbo` inbox tables or replicate portal-facing data to a VPS. poll `utility_dbo` inbox tables or replicate portal-facing data to a VPS.
- **Small / open:** (a) `TRASPASOS PAYPAL` clearing account still tops the adeudo worklist - **Small / open:** (a) `TRASPASOS PAYPAL` clearing account still tops the adeudo worklist
(§6.4d) — a business modelling call, not code. (b) Credential rotation on the old repo's (§6.4d) — a business modelling call, not code. (b) Credential rotation on the old repo's
exposed MySQL password. (c) The `/estado-cuenta` browser visual pass`/banco` was verified exposed MySQL password. (c) ~~The `/estado-cuenta` browser visual pass.~~ **DONE 2026-07-24** —
in-browser this session; `/estado-cuenta` still worth a look. verified vs dev: Anular buttons admin-gated, voided rows struck + excluded from totals,
clicking Anular voids end-to-end (note: it uses a blocking `window.confirm`). Customer-detail
mini tx list now also strikes voided rows ("(anulado)" tag) — was the last void-UI gap.
---
## Statement OCR intake (`/recibos`) — DONE 2026-08-01
> As-built reference: **`docs/STATEMENT_OCR.md`** (written 2026-08-02) — the
> parsers, the matcher's scoped-field rules, confirm/learning semantics and the
> API surface. `docs/RECEIPT_CAPTURE_SPEC.md` §2 stays the design and the
> measured evidence. This section is the session record of building it.
Plan step 11 §2 (`docs/RECEIPT_CAPTURE_SPEC.md` §2). The last big utilities
feature: staff scan the month's utility bills and the machine proposes customer
+ amount per page, instead of keying 300+ statements per company by hand. Built
in `apps/api/src/statements/` and `apps/web/src/app/recibos/`, posting through
step 11 §1.2's `BillingService.createBatch` seam (`source: "OCR"`, per-document
`captureRef`) so machine and hand capture share one write path and one audit
trail. Abilities `statement:ingest` / `statement:review`, both STAFF.
**Verified end to end against the live dev API + MinIO**, not just built: real
CFE and Telnor scans uploaded over HTTP, OCR'd, matched, confirmed against a
check, and the resulting rows checked in MySQL — negative (charge) amounts,
`captureSource = OCR`, concept auto-derived from the batch's service kind,
`captureRef` linking each transaction back to its page. Re-confirming a posted
batch is refused. All test data was removed afterwards.
**Everything here was decided from 10 real scanned statements (46 pages), not
from the sample-free spec.** Shipped-parser results on them: provider 46/46,
account reference 43/46, amount 42/46, due date 44/46; matched against the dev
database, **39/46 (85%) exact auto-match, 40/46 (87%) identified**. The rest are
real review cases (one shared account number, three phones not on file, one
clave not in the book, one page too poor to read).
Findings that corrected the spec, each of which changed the build:
- **The scans have no text layer at all** — they are camera images of paper, so
OCR is mandatory rather than a convenience, and they arrive **bundled, one
customer per page**.
- **Clave catastral is not predial.** `DATMEX.clave` (934 rows, `KA903009`) is
what CESPT and predial bills print; `DATMEX.predial` — which
`PROPERTY_TAX.accountNumber` holds — has only 663 distinct values across 1135
rows and appears on no statement. The clave now lives on
`Property.cadastralKey` as the matcher's secondary key; predial was left
untouched. This is the question that had been blocking predial matching.
- **Gas was recoverable after all.** The spec said no legacy gas number existed;
in fact 160 of 334 `DATMEX.gas` values are real account numbers (the rest are
`ESTACIONARIO`/`CILINDRO` descriptors). Recovered into `GAS.meterNumber`.
- **Phone is one billed line per property** (534 / 18 / 1 across phone1/2/3), so
`TELEPHONE` — a new `ServiceKind` — backfills from `phone1` only.
- **Never match on the printed name.** A CESPT receipt for account `5365218`
reads `ARNAIZ ROSAS ELSA AURORA`; the office's book, corroborated by the
clave, has `CATT, RANDY`. The name on a utility bill is the registrant, not
the current owner.
`migration/backfill_statement_match_fields.py` closes those three data gaps on
an existing database (idempotent, wired into `run_all.py` after
`transform_properties.py`, which now produces them directly on a full rebuild).
Applied to dev: 934 claves, 160 gas numbers, 534 TELEPHONE rows.
Implementation notes worth keeping:
- OCR is self-hosted **Tesseract** behind an `OcrProvider` interface — the
provider question is closed on measured accuracy, and a managed API stays a
one-line swap in `statements.module.ts`. `tesseract-ocr`,
`tesseract-ocr-data-spa` and `poppler-utils` were added to the API image; if
they are missing the module reports itself unavailable and only this feature
is disabled.
- **Payment barcodes beat printed labels.** One CFE label OCR'd a digit too
many while its barcode was correct, so the barcode is the source and the label
the cross-check; disagreement forces review.
- **Detect the provider by brand first, layout only as a fallback** — and never
interleave the two passes. A scanned CESPT header came back as `E BAJA ES
PAGO / EALIFORNIA`, which is why the layout fallback exists; a Telnor page
contains words a CFE layout rule would otherwise claim, which is why ordering
matters.
- **Parse amounts by separator position.** A real Telnor bill OCR'd as
`$ 649,00`; stripping commas as thousands separators turns that into $64,900.
- Two of the three layouts are line-oriented, but the CESPT "RECIBO" is a
**table** whose values sit under column headers — that one needs the word
boxes, which is why `OcrPage` carries geometry and not just text.
- Confirming a document whose matched service had no reference **writes the
reference back** (only into an empty field, and only when exactly one blank
service of that kind is a candidate), so gas and any other cold start is a
one-time cost rather than a permanent queue.
- Handwritten folder numbers on the bills (`9`, `405`) are **not** used for
matching — Tesseract read `405` as `205`.
**Open:** whether the CFE charge should be the rounded barcode/headline figure
(`$268`, what is paid at the window — what the parser uses today) or the exact
breakdown total (`$268.88`). One question for Jorge.
## Policy OCR capture (`/polizas/captura`) — DONE 2026-08-01, unplanned
**This feature was not in any spec.** It is what the statement OCR work above
turned into once the pipeline existed. Having built render → OCR → parse →
match → review for CFE/CESPT/Telnor receipts, the same shape obviously fits
the *other* stack of paper this office keys in by hand every week: the carrier
policy PDFs behind every `Policy` row. Full write-up in `docs/POLICY_OCR.md`.
The pipeline was reused rather than copied. `OcrModule` was **extracted out of
`StatementsModule`** in the same commit so `PolicyOcrModule` could inject
`OCR_PROVIDER` without taking on the statement pipeline — that extraction was
blocking, not tidying; the policy module could not resolve the provider at all
until it existed. `StatementsModule` imports it now and binds nothing itself,
so the Tesseract-vs-managed-API decision stays one line in one file for both
features.
Screens mirror Captura exactly: `/polizas/nuevo` is the manual tab,
`/polizas/captura` the automática one, both rendering `PolicyCaptura.tsx`, with
the batch review queue at `/polizas/captura/[id]`. Abilities `policy:ingest` /
`policy:ocr-review`, both STAFF — same trust tier as statement OCR, and for the
same reason: nothing reaches the books unconfirmed.
**The statement pipeline's central assumption inverts here, and that is the
thing to remember.** Utility statements arrive bundled *one customer per page*,
so there a page is a document and the parser runs per page. A policy PDF is the
opposite: the GMX certificate is one policy spread across two pages (contract
header on page 1, the per-coverage table on page 2). So every page's text is
concatenated and the parser and matcher run **once per file**. Consequences:
`PolicyOcrDocument.pageNumber` is repurposed as the file ordinal within the
batch (the `(batchId, pageNumber)` unique constraint still holds), `ocrConfidence`
is the mean across the file's pages, and a file that fails to parse yields
exactly one `OCR_FAILED` row.
`storageKey` points at the **source PDF**, not a rendered page image, so the
review screen embeds the exact artifact the office received and gets the
browser's native PDF scrolling, zoom and text selection for free. The page PNGs
are still written for future re-OCR, but nothing treats them as the document's
identity. (The statement side is the reverse, because there a page *is* the
document.)
Findings worth keeping:
- **The GMX certificate has no premium on it at all.** Not intermittently
missing — the figure lives on GMX's separate `recibo` PDF. The parser leaves
the premium fields null and pushes a note saying so, confirm never overwrites
an existing `Policy.netPremium` with null, and the optional ledger write is
gated on staff ticking `postPremium` *and* a premium actually parsing.
Without that second gate a premium-less certificate would book a $0 charge on
every confirm.
- **Match on `Policy.policyNumber`, never the printed insured name.** Same
registrant-vs-current-owner drift that rules names out on the utility side.
Zero hits means a new policy and confirm creates the row under a picked
customer; more than one hit is surfaced for a human, never auto-picked —
duplicate numbers across related parties do occur.
- Deductible and loss participation are stored as **strings** (`"5%"`,
`"USD 1,000"`): they are printed as a mix of percentages, amounts and free
text, and normalising them would lose the distinction.
- Carrier-portal PDFs are usually **born-digital**, so the text layer wins and
no OCR runs at all most of the time — same precedence rule as the statement
pipeline.
- The digit-confusion map and the amount-by-separator-position parser are
**duplicated on purpose** rather than imported, to keep the module
self-contained. Fix a bug in one, check the other.
8/8 parser tests, all against verbatim text from one real document
(`HC_Folio_000767_Traduccion.pdf`).
**Open:** GMX is the only carrier implemented — the dispatcher is a
`[provider, pattern]` table plus a parser map, so a second carrier is a
function and two entries, but no other layout has been seen. Reading the
premium off the separate `recibo` PDF and pairing it to its certificate is the
obvious next piece; it is what would let `postPremium` stop being a manual
tick. And nothing versions a re-issued policy — confirm updates the existing
row, so there is no record that this is the 2027 issue of that number.
## Notificaciones (`/notificaciones`) — DONE 2026-08-01 → 08-02
Two features that were spec'd separately turned out to be one screen. The four
legacy mass-email jobs (`docs/MASS_EMAIL_NOTIFICATIONS.md`, ported from
`email.notifications/send*.php`) and the insurance renewal avisos
(`docs/INSURANCE_FEATURES_SPEC.md` §1) both mean *tell a customer something by
email*, so they are **tabs of one screen over one log**, not two menu entries.
`/renovaciones` is an alias that lands on the Pólizas tab, the same pattern
Captura uses.
- **Servicios tab** — the four jobs (pagos pendientes, confirmación de pago,
estado de cuenta, fideicomiso), individually or "Ejecutar todos". Ability
`notification:send` (MANAGER); STAFF sees the log read-only.
- **Pólizas tab** — pending avisos at 30/15 days before expiry and 7 days
after, sent one at a time or as a sweep. Ability `renewal:send` (MANAGER).
**One send log for the whole platform.** `email_notification_log` is not
job-specific: renewals write it too (`RENEWAL_NOTICE` / `POLICIES`) through the
same `NotificationLogService`. That is what makes "Registro de envíos" complete
— the failures and no-email skips exist *only* there. `RenewalNotice` was not
made redundant by it: that row is **gating** state (one per policy+generation,
drives the pending list), the log is **history** (every attempt). `level` is
therefore per-type and unreadable without its `notificationType` — 0/1
yellow/red on `ACCOUNT_STATUS`, the aviso generation 1/2/3 on
`RENEWAL_NOTICE`.
**Manual mark-as-sent was dropped on purpose.** The spec called for it; a
button that marks a notice sent without sending anything is a button that lets
the list claim a customer was told when they were not. `POST /renewals/send`
replaced it — sending from the list *is* the marking.
**`app_settings` is the operator-config seam this work introduced.**
`SettingsService` resolves every key **db → env → default** and reports which
rung a value came from, so an existing deployment keeps behaving exactly as it
did until somebody saves in the UI. Three keys today: the summary recipients
(was `NOTIFICATION_ADMIN_EMAILS`, now a fallback) and the two sweep cadences.
Credentials deliberately stay in the environment — SES keys, `DATABASE_URL`
and S3 config are deployment identity, must exist before the app can reach its
own database, and a table only widens who can read them.
**The send flags are global, and that was a real bug fix (08-02).** The
`debug` / `ignoreDayRestriction` / `useEmailLimit` panel lived inside the
Servicios tab, so there was **no way to test a renewal aviso without mailing a
real customer**. It now lives in the shell above the tabs and both halves read
it. On the pólizas path `debug` does three things, and all three are required
together: it diverts the mail, it skips the `RenewalNotice` upsert, and it does
not advance the sweep's `lastSuccessfulAt`. Miss the third and `renewalWindow()`
narrows back to a single day on the next real run — a test send would silently
destroy the letters it only pretended to send. Flags are per-visit UI state and
are **never persisted**; a stored `debug` would survive a reload and swallow
real customer mail until somebody noticed.
**Both cadences are operator-editable (08-02).** The renewal sweep's
`@Cron("0 6 * * *")` literal lasted one day. `NotificationScheduleService` now
owns both: the owning services register a handler in `onModuleInit`, the
service compiles the stored `{hour, minute, weekdays}` to a cron expression and
installs it in `SchedulerRegistry`, and saving from the UI reinstalls the job —
no restart, which was the point. It lives in its own module for the same reason
as `NotificationLogModule`: `NotificationsModule` and `RenewalsModule` both need
it and neither may import the other. Defaults preserve prior behaviour exactly
(pólizas 06:00 daily, servicios **off** — a default that starts mailing 260
customers after a deploy is not a default, it's an incident). A scheduled run
never inherits the UI flags: no `debug`, and no `ignoreDayRestriction`, since an
automatic run on the operator's own cadence is precisely the case the
Mon/Wed/Fri gate was written for.
Implementation notes worth keeping:
- `cron` had to become a **direct dependency of `apps/api`**. It is a
transitive dep of `@nestjs/schedule`, but pnpm's strict layout does not hoist
it, so `import { CronJob } from "cron"` does not resolve without it.
- The pólizas sweep already had a DB lock (`scheduled_job_states`); the
servicios run-all does not, and relies on the deployment being
single-replica, which it is on galactus today.
- Wire shapes of the four jobs are byte-for-byte the legacy PHP responses,
quirks included (Job 1 reports `result`, not `request`).
**Open:** the `SES_*` Gitea secrets were created 2026-08-02, so the feature is
no longer blocked — but it has not shipped (master is well past the newest tag)
and two things nobody has checked decide whether mail leaves the building:
`SES_FROM` must be a verified identity in `SES_REGION`, and the AWS account
must be out of the SES sandbox, which otherwise restricts delivery to verified
recipients and would fail a real sweep while looking correctly configured. Run
the first sweep with `debug` on. Still open beyond that: the 78 policyholders
with no email are logged as `SKIPPED_NO_EMAIL` but have no printable worklist,
and the notice body is English-only (`Customer` carries no language
preference).
Binary file not shown.

After

Width:  |  Height:  |  Size: 110 KiB

+7
View File
@@ -0,0 +1,7 @@
/** @type {import('jest').Config} */
module.exports = {
rootDir: "src",
testEnvironment: "node",
testRegex: ".*\\.spec\\.ts$",
transform: { "^.+\\.ts$": "ts-jest" },
};
+2 -1
View File
@@ -3,6 +3,7 @@
"collection": "@nestjs/schematics", "collection": "@nestjs/schematics",
"sourceRoot": "src", "sourceRoot": "src",
"compilerOptions": { "compilerOptions": {
"deleteOutDir": true "deleteOutDir": true,
"tsConfigPath": "tsconfig.build.json"
} }
} }
+8 -1
View File
@@ -1,6 +1,6 @@
{ {
"name": "@jorgecuadros/api", "name": "@jorgecuadros/api",
"version": "0.1.0", "version": "1.0.22",
"private": true, "private": true,
"scripts": { "scripts": {
"build": "nest build", "build": "nest build",
@@ -11,18 +11,24 @@
"test": "jest" "test": "jest"
}, },
"dependencies": { "dependencies": {
"@aws-sdk/client-s3": "^3.665.0",
"@aws-sdk/client-sesv2": "^3.1101.0",
"@jorgecuadros/database": "workspace:*", "@jorgecuadros/database": "workspace:*",
"@nestjs/common": "^10.4.4", "@nestjs/common": "^10.4.4",
"@nestjs/config": "^3.3.0", "@nestjs/config": "^3.3.0",
"@nestjs/core": "^10.4.4", "@nestjs/core": "^10.4.4",
"@nestjs/passport": "^10.0.3", "@nestjs/passport": "^10.0.3",
"@nestjs/platform-express": "^10.4.4", "@nestjs/platform-express": "^10.4.4",
"@nestjs/schedule": "^4.1.2",
"argon2": "^0.41.1", "argon2": "^0.41.1",
"class-transformer": "^0.5.1", "class-transformer": "^0.5.1",
"class-validator": "^0.14.1", "class-validator": "^0.14.1",
"cron": "^3.2.1",
"exceljs": "^4.4.0",
"express-session": "^1.18.0", "express-session": "^1.18.0",
"passport": "^0.7.0", "passport": "^0.7.0",
"passport-local": "^1.0.0", "passport-local": "^1.0.0",
"pdfkit": "^0.15.1",
"reflect-metadata": "^0.2.2", "reflect-metadata": "^0.2.2",
"rxjs": "^7.8.1" "rxjs": "^7.8.1"
}, },
@@ -35,6 +41,7 @@
"@types/node": "^20.16.11", "@types/node": "^20.16.11",
"@types/passport": "^1.0.17", "@types/passport": "^1.0.17",
"@types/passport-local": "^1.0.38", "@types/passport-local": "^1.0.38",
"@types/pdfkit": "^0.13.5",
"jest": "^29.7.0", "jest": "^29.7.0",
"ts-jest": "^29.2.5", "ts-jest": "^29.2.5",
"ts-node": "^10.9.2", "ts-node": "^10.9.2",
+21
View File
@@ -6,4 +6,25 @@ export class AppController {
health() { health() {
return { status: "ok" }; return { status: "ok" };
} }
/**
* What is actually running. The three values are baked into the image at
* build time by .gitea/workflows/build.yml (see docker/api.Dockerfile) and
* are the only way to confirm a deploy — or a rollback — landed: the tag you
* dispatched and the code inside the container can disagree if a stack was
* applied without pulling, or if the app stack still names an older tag.
*
* Deliberately unauthenticated, same as /health: the deploy workflow has to
* read it with no session, and it exposes nothing an attacker could not
* already infer from the repo.
*/
@Get("version")
version() {
return {
service: "api",
version: process.env.APP_VERSION ?? "dev",
gitSha: process.env.GIT_SHA ?? "unknown",
buildDate: process.env.BUILD_DATE ?? "unknown",
};
}
} }
+16
View File
@@ -1,30 +1,46 @@
import { Module } from "@nestjs/common"; import { Module } from "@nestjs/common";
import { ConfigModule } from "@nestjs/config"; import { ConfigModule } from "@nestjs/config";
import { ScheduleModule } from "@nestjs/schedule";
import { PrismaModule } from "./prisma/prisma.module"; import { PrismaModule } from "./prisma/prisma.module";
import { StorageModule } from "./storage/storage.module";
import { CommonModule } from "./common/common.module"; import { CommonModule } from "./common/common.module";
import { MailModule } from "./mail/mail.module";
import { UsersModule } from "./users/users.module"; import { UsersModule } from "./users/users.module";
import { AuthModule } from "./auth/auth.module"; import { AuthModule } from "./auth/auth.module";
import { CustomersModule } from "./customers/customers.module"; import { CustomersModule } from "./customers/customers.module";
import { PoliciesModule } from "./policies/policies.module"; import { PoliciesModule } from "./policies/policies.module";
import { PropertiesModule } from "./properties/properties.module"; import { PropertiesModule } from "./properties/properties.module";
import { BillingModule } from "./billing/billing.module"; import { BillingModule } from "./billing/billing.module";
import { StatementsModule } from "./statements/statements.module";
import { PolicyOcrModule } from "./policy-ocr/policy-ocr.module";
import { BankModule } from "./bank/bank.module"; import { BankModule } from "./bank/bank.module";
import { OpsModule } from "./ops/ops.module"; import { OpsModule } from "./ops/ops.module";
import { ReportsModule } from "./reports/reports.module";
import { RenewalsModule } from "./renewals/renewals.module";
import { NotificationsModule } from "./notifications/notifications.module";
import { AppController } from "./app.controller"; import { AppController } from "./app.controller";
@Module({ @Module({
imports: [ imports: [
ConfigModule.forRoot({ isGlobal: true }), ConfigModule.forRoot({ isGlobal: true }),
ScheduleModule.forRoot(),
PrismaModule, PrismaModule,
StorageModule,
CommonModule, CommonModule,
MailModule,
UsersModule, UsersModule,
AuthModule, AuthModule,
CustomersModule, CustomersModule,
PoliciesModule, PoliciesModule,
PropertiesModule, PropertiesModule,
BillingModule, BillingModule,
StatementsModule,
PolicyOcrModule,
BankModule, BankModule,
OpsModule, OpsModule,
ReportsModule,
RenewalsModule,
NotificationsModule,
], ],
controllers: [AppController], controllers: [AppController],
}) })
+39 -1
View File
@@ -21,9 +21,13 @@ export type Ability =
| "customer:create" | "customer:create"
| "customer:update" | "customer:update"
| "customer:delete" | "customer:delete"
| "customer:portal-access"
| "policy:create" | "policy:create"
| "policy:update" | "policy:update"
| "policy:delete" | "policy:delete"
| "policy:ingest"
| "policy:ocr-review"
| "renewal:send"
| "property:create" | "property:create"
| "property:update" | "property:update"
| "property:delete" | "property:delete"
@@ -31,18 +35,34 @@ export type Ability =
| "ledger:void" | "ledger:void"
| "bank:create" | "bank:create"
| "bank:void" | "bank:void"
| "bank:manage-accounts"
| "statement:ingest"
| "statement:review"
| "lookup:manage" | "lookup:manage"
| "user:manage" | "user:manage"
| "db:manage"; | "db:manage"
| "notification:send"
| "setting:manage";
/** Minimum role required for each ability. */ /** Minimum role required for each ability. */
export const ABILITY_MIN: Record<Ability, Role> = { export const ABILITY_MIN: Record<Ability, Role> = {
"customer:create": "STAFF", "customer:create": "STAFF",
"customer:update": "STAFF", "customer:update": "STAFF",
"customer:delete": "ADMIN", "customer:delete": "ADMIN",
// Assigning a portal NUMid is granting someone the ability to log in to
// my.jorgecuadros.com and read an account, so it sits above customer:update:
// editing a phone number is the day job, handing out portal identity is not.
// It is also close to irreversible in practice — the id is what the customer
// then types at every login.
"customer:portal-access": "MANAGER",
"policy:create": "STAFF", "policy:create": "STAFF",
"policy:update": "STAFF", "policy:update": "STAFF",
"policy:delete": "MANAGER", "policy:delete": "MANAGER",
// Insurance OCR intake is the same trust tier as statement OCR: STAFF can
// upload + confirm, nothing reaches the books unconfirmed.
"policy:ingest": "STAFF",
"policy:ocr-review": "STAFF",
"renewal:send": "MANAGER",
"property:create": "STAFF", "property:create": "STAFF",
"property:update": "STAFF", "property:update": "STAFF",
"property:delete": "MANAGER", "property:delete": "MANAGER",
@@ -50,9 +70,27 @@ export const ABILITY_MIN: Record<Ability, Role> = {
"ledger:void": "MANAGER", "ledger:void": "MANAGER",
"bank:create": "STAFF", "bank:create": "STAFF",
"bank:void": "MANAGER", "bank:void": "MANAGER",
// Opening or renaming a chequera is rarer and higher-stakes than posting a
// movement into one — a wrong account silently mixes two sets of books.
"bank:manage-accounts": "MANAGER",
// Uploading a stack of scans and reviewing what the OCR read are both
// "capturing a receipt" — the same trust tier as ledger:create, since
// confirming a statement *is* capturing it. The review step is what makes
// this safe at STAFF level: nothing reaches the ledger unconfirmed.
"statement:ingest": "STAFF",
"statement:review": "STAFF",
"lookup:manage": "MANAGER", "lookup:manage": "MANAGER",
"user:manage": "ADMIN", "user:manage": "ADMIN",
"db:manage": "ADMIN", "db:manage": "ADMIN",
// Mass email notifications — fires mail to customers on the office's
// behalf, with no per-row review. Same trust tier as `renewal:send`:
// a STAFF user typing one customer receipt is fine; a STAFF user firing
// 260 mail merges on the customer base is not.
"notification:send": "MANAGER",
// Editing operator configuration. Above `notification:send` on purpose:
// firing a sweep is the day job, but changing WHERE the audit summaries
// land is how someone would quietly stop them being read.
"setting:manage": "ADMIN",
}; };
export const ALL_ABILITIES = Object.keys(ABILITY_MIN) as Ability[]; export const ALL_ABILITIES = Object.keys(ABILITY_MIN) as Ability[];
+28 -1
View File
@@ -1,9 +1,21 @@
import { Controller, Get, HttpCode, Post, Req, Res, UseGuards } from "@nestjs/common"; import {
Body,
Controller,
Get,
HttpCode,
Patch,
Post,
Req,
Res,
UseGuards,
} from "@nestjs/common";
import { Request, Response } from "express"; import { Request, Response } from "express";
import { LocalAuthGuard } from "./local-auth.guard"; import { LocalAuthGuard } from "./local-auth.guard";
import { AuthenticatedGuard } from "./authenticated.guard"; import { AuthenticatedGuard } from "./authenticated.guard";
import { LoginDto } from "./login.dto"; import { LoginDto } from "./login.dto";
import { UpdatePreferencesDto } from "./update-preferences.dto";
import { abilitiesFor, Role } from "./abilities"; import { abilitiesFor, Role } from "./abilities";
import { UsersService } from "../users/users.service";
/** Attach the resolved ability map so the web can gate its UI off one payload. */ /** Attach the resolved ability map so the web can gate its UI off one payload. */
function withAbilities(user: unknown) { function withAbilities(user: unknown) {
@@ -14,6 +26,8 @@ function withAbilities(user: unknown) {
@Controller("auth") @Controller("auth")
export class AuthController { export class AuthController {
constructor(private readonly users: UsersService) {}
// LoginDto is only used for request-shape documentation/validation here — // LoginDto is only used for request-shape documentation/validation here —
// the actual credential check happens inside LocalStrategy via Passport, // the actual credential check happens inside LocalStrategy via Passport,
// which populates req.user before this handler runs. // which populates req.user before this handler runs.
@@ -30,6 +44,19 @@ export class AuthController {
return withAbilities(req.user); return withAbilities(req.user);
} }
/**
* Update the caller's own UI preferences. Deliberately not on /users/:id —
* that controller is ADMIN-only, and this has to work for every role. The
* target is always the session's own user id, never a body parameter.
*/
@UseGuards(AuthenticatedGuard)
@Patch("preferences")
async updatePreferences(@Req() req: Request, @Body() dto: UpdatePreferencesDto) {
const id = (req.user as { id: string }).id;
const user = await this.users.updatePreferences(id, dto.uiScale);
return withAbilities(user);
}
@Post("logout") @Post("logout")
@HttpCode(200) @HttpCode(200)
logout(@Req() req: Request) { logout(@Req() req: Request) {
@@ -0,0 +1,17 @@
import { IsNumber, Max, Min } from "class-validator";
/**
* Self-service UI preferences — any authenticated user may set these on their
* own account, including VIEWER. No ability gate: it changes nothing but how
* the app looks to that one person.
*
* The bounds mirror MIN_UI_SCALE/MAX_UI_SCALE in apps/web/src/lib/ui-scale.ts;
* keep them in sync. The API clamps rather than trusting the client because
* this endpoint is reachable outside the UI.
*/
export class UpdatePreferencesDto {
@IsNumber()
@Min(0.9)
@Max(1.5)
uiScale!: number;
}
+47
View File
@@ -0,0 +1,47 @@
import {
IsBoolean,
IsIn,
IsOptional,
IsString,
MinLength,
} from "class-validator";
/** Mirrors the Prisma `Currency` enum; a chequera's is fixed at creation. */
export const BANK_CURRENCIES = ["MXN", "USD"] as const;
export type BankAccountCurrency = (typeof BANK_CURRENCIES)[number];
/** Mirrors `TransactionDomain`. A soft hint on the account, never enforced. */
export const BANK_BUSINESS_LINES = ["UTILITY", "INSURANCE", "TRUST"] as const;
export type BankBusinessLine = (typeof BANK_BUSINESS_LINES)[number];
export class CreateBankDto {
@IsString() @MinLength(1) name!: string;
/** "MX" | "US" — free text, informational only. */
@IsOptional() @IsString() country?: string;
}
export class UpdateBankDto {
@IsOptional() @IsString() @MinLength(1) name?: string;
@IsOptional() @IsString() country?: string;
}
export class CreateBankAccountDto {
@IsString() @MinLength(1) bankId!: string;
@IsString() @MinLength(1) label!: string;
/**
* Immutable after creation (no field for it on the update DTO): every
* movement already booked into the account is denominated in it, so
* changing it would silently re-denominate history.
*/
@IsIn(BANK_CURRENCIES) currency!: BankAccountCurrency;
@IsOptional() @IsIn(BANK_BUSINESS_LINES) businessLine?: BankBusinessLine;
@IsOptional() @IsBoolean() active?: boolean;
}
export class UpdateBankAccountDto {
@IsOptional() @IsString() @MinLength(1) bankId?: string;
@IsOptional() @IsString() @MinLength(1) label?: string;
@IsOptional() @IsIn(BANK_BUSINESS_LINES) businessLine?: BankBusinessLine;
/** Closing an account hides it from the picker; its movements stay readable. */
@IsOptional() @IsBoolean() active?: boolean;
}
+5 -1
View File
@@ -2,10 +2,14 @@ import { IsBoolean, IsNumber, IsOptional, IsString, MinLength } from "class-vali
/** /**
* A new bank-register movement. `amount` is signed: positive = ingreso, * A new bank-register movement. `amount` is signed: positive = ingreso,
* negative = egreso (the module's sign convention). Single currency (MXN). * negative = egreso (the module's sign convention). The currency is the
* account's, not the movement's — `bankAccountId` decides it.
* Booked rows are never edited — a mistake is corrected by voiding + re-capture. * Booked rows are never edited — a mistake is corrected by voiding + re-capture.
*/ */
export class CreateBankMovementDto { export class CreateBankMovementDto {
/** Which chequera this lands in. Required — see BankAccount in the schema. */
@IsString() @MinLength(1) bankAccountId!: string;
@IsNumber() amount!: number; @IsNumber() amount!: number;
@IsString() @MinLength(1) transactionDate!: string; @IsString() @MinLength(1) transactionDate!: string;
+96 -6
View File
@@ -3,6 +3,7 @@ import {
Controller, Controller,
Get, Get,
Param, Param,
Patch,
Post, Post,
Query, Query,
Req, Req,
@@ -20,6 +21,12 @@ import {
BankSort, BankSort,
} from "./bank.service"; } from "./bank.service";
import { CreateBankMovementDto } from "./bank-movement.dto"; import { CreateBankMovementDto } from "./bank-movement.dto";
import {
CreateBankAccountDto,
CreateBankDto,
UpdateBankAccountDto,
UpdateBankDto,
} from "./bank-account.dto";
const DIRECTIONS: BankDirection[] = ["income", "expense", "void"]; const DIRECTIONS: BankDirection[] = ["income", "expense", "void"];
const CLEARED: BankCleared[] = ["cleared", "pending"]; const CLEARED: BankCleared[] = ["cleared", "pending"];
@@ -54,28 +61,108 @@ export class BankController {
return (req.user as { id: string }).id; return (req.user as { id: string }).id;
} }
// --- accounts -------------------------------------------------------------
// Declared before the parameterised routes below so `/bank/accounts` can
// never be swallowed by a `:id`-shaped path.
/**
* The account picker. Readable by any authenticated user, VIEWER included —
* nothing else on this page can render until an account is chosen.
*/
@Get("accounts")
accounts() {
return this.bank.listAccounts();
}
@Get("banks")
banks() {
return this.bank.listBanks();
}
@Post("banks")
@RequireAbility("bank:manage-accounts")
async createBank(@Body() dto: CreateBankDto, @Req() req: Request) {
const row = await this.bank.createBank(dto);
void this.audit.log(this.actingId(req), "bank.bank.create", {
bankId: row.id,
name: row.name,
});
return row;
}
@Patch("banks/:id")
@RequireAbility("bank:manage-accounts")
async updateBank(
@Param("id") id: string,
@Body() dto: UpdateBankDto,
@Req() req: Request,
) {
const row = await this.bank.updateBank(id, dto);
void this.audit.log(this.actingId(req), "bank.bank.update", { bankId: id });
return row;
}
@Post("accounts")
@RequireAbility("bank:manage-accounts")
async createAccount(
@Body() dto: CreateBankAccountDto,
@Req() req: Request,
) {
const row = await this.bank.createAccount(dto);
void this.audit.log(this.actingId(req), "bank.account.create", {
bankAccountId: row.id,
label: row.label,
currency: row.currency,
});
return row;
}
@Patch("accounts/:id")
@RequireAbility("bank:manage-accounts")
async updateAccount(
@Param("id") id: string,
@Body() dto: UpdateBankAccountDto,
@Req() req: Request,
) {
const row = await this.bank.updateAccount(id, dto);
void this.audit.log(this.actingId(req), "bank.account.update", {
bankAccountId: id,
});
return row;
}
// --- register reads (all scoped to one account) ---------------------------
@Get("stats") @Get("stats")
stats() { async stats(@Query("bankAccountId") bankAccountId?: string) {
return this.bank.stats(); const account = await this.bank.requireAccount(bankAccountId);
return this.bank.stats(account.id);
} }
@Get("facets") @Get("facets")
facets() { async facets(@Query("bankAccountId") bankAccountId?: string) {
return this.bank.facets(); const account = await this.bank.requireAccount(bankAccountId);
return this.bank.facets(account.id);
} }
/** Year and month rollups with a running net-movement figure. */ /** Year and month rollups with a running net-movement figure. */
@Get("summary") @Get("summary")
summary(@Query("year") year?: string) { async summary(
@Query("bankAccountId") bankAccountId?: string,
@Query("year") year?: string,
) {
const account = await this.bank.requireAccount(bankAccountId);
const y = Number(year); const y = Number(year);
return this.bank.summary( return this.bank.summary(
account.id,
Number.isInteger(y) && y >= 1900 && y <= 2999 ? y : undefined, Number.isInteger(y) && y >= 1900 && y <= 2999 ? y : undefined,
); );
} }
/** The register browser. */ /** The register browser. */
@Get() @Get()
list( async list(
@Query("bankAccountId") bankAccountId?: string,
@Query("query") query?: string, @Query("query") query?: string,
@Query("page") page?: string, @Query("page") page?: string,
@Query("pageSize") pageSize?: string, @Query("pageSize") pageSize?: string,
@@ -85,7 +172,9 @@ export class BankController {
@Query("to") to?: string, @Query("to") to?: string,
@Query("sort") sort?: string, @Query("sort") sort?: string,
) { ) {
const account = await this.bank.requireAccount(bankAccountId);
return this.bank.list({ return this.bank.list({
bankAccountId: account.id,
query, query,
page: Math.max(1, Number(page) || 1), page: Math.max(1, Number(page) || 1),
pageSize: Math.min(100, Math.max(1, Number(pageSize) || 25)), pageSize: Math.min(100, Math.max(1, Number(pageSize) || 25)),
@@ -105,6 +194,7 @@ export class BankController {
const row = await this.bank.createMovement(dto); const row = await this.bank.createMovement(dto);
void this.audit.log(this.actingId(req), "bank.create", { void this.audit.log(this.actingId(req), "bank.create", {
bankTransactionId: row.id, bankTransactionId: row.id,
bankAccountId: row.bankAccountId,
amount: dto.amount, amount: dto.amount,
}); });
return row; return row;
+172 -18
View File
@@ -2,6 +2,12 @@ import { BadRequestException, Injectable, NotFoundException } from "@nestjs/comm
import { Prisma } from "@jorgecuadros/database"; import { Prisma } from "@jorgecuadros/database";
import { PrismaService } from "../prisma/prisma.service"; import { PrismaService } from "../prisma/prisma.service";
import { CreateBankMovementDto } from "./bank-movement.dto"; import { CreateBankMovementDto } from "./bank-movement.dto";
import {
CreateBankAccountDto,
CreateBankDto,
UpdateBankAccountDto,
UpdateBankDto,
} from "./bank-account.dto";
/** /**
* App-voided rows (voidedAt set) are reversed and must leave every * App-voided rows (voidedAt set) are reversed and must leave every
@@ -28,9 +34,18 @@ const NOT_VOIDED: Prisma.BankTransactionWhereInput = { voidedAt: null };
* expense and are excluded from both sides, the way the ~193 zero rows are * expense and are excluded from both sides, the way the ~193 zero rows are
* in the customer ledger. * in the customer ledger.
* *
* SINGLE CURRENCY. Unlike the customer ledger there is no currency column here: * ONE ACCOUNT AT A TIME, CURRENCY FROM THE ACCOUNT. The office now keeps more
* `bank_transactions` has none, and every `amountInWords` on the egreso side is * than one chequera (Utilities banks in MXN, Seguros in USD), so every read
* spelled out in PESOS. All figures in this module are MXN. * path here is scoped to exactly one `bankAccountId` — never "all accounts".
* There is deliberately no currency column on `bank_transactions`: a movement
* inherits its account's, the way a real bank account doesn't mix currencies.
* Callers must therefore pass an account id; an unscoped total would sum MXN
* and USD into a figure that never existed, the same mistake the billing
* module's per-currency rule exists to prevent.
*
* The 22,669 migrated rows are all SCOTHIA = the Utilities MXN account
* (backfilled by `migration/backfill_bank_accounts.py`), and their
* `amountInWords` on the egreso side is spelled out in PESOS accordingly.
* *
* NO CATEGORY DIMENSION. `bank_transactions.categoryId` is NULL on all 22,354 * NO CATEGORY DIMENSION. `bank_transactions.categoryId` is NULL on all 22,354
* rows and this module does not filter or group by it, because the data cannot * rows and this module does not filter or group by it, because the data cannot
@@ -64,6 +79,8 @@ export type BankSort =
| "reference"; | "reference";
export interface BankListParams { export interface BankListParams {
/** Which chequera to read. Required — see the module header. */
bankAccountId: string;
query?: string; query?: string;
page: number; page: number;
pageSize: number; pageSize: number;
@@ -98,7 +115,9 @@ export class BankService {
constructor(private readonly prisma: PrismaService) {} constructor(private readonly prisma: PrismaService) {}
private where(p: BankListParams): Prisma.BankTransactionWhereInput { private where(p: BankListParams): Prisma.BankTransactionWhereInput {
const and: Prisma.BankTransactionWhereInput[] = []; const and: Prisma.BankTransactionWhereInput[] = [
{ bankAccountId: p.bankAccountId },
];
if (p.query && p.query.trim()) { if (p.query && p.query.trim()) {
const q = p.query.trim(); const q = p.query.trim();
@@ -124,7 +143,9 @@ export class BankService {
}); });
} }
return and.length ? { AND: and } : {}; // Never empty: the account clause above is always present, so no read can
// accidentally span every chequera.
return { AND: and };
} }
private orderBy( private orderBy(
@@ -233,22 +254,25 @@ export class BankService {
}; };
} }
/** Top-line figures for the bank page header. */ /** Top-line figures for the bank page header, for one chequera. */
async stats() { async stats(bankAccountId: string) {
const account = { bankAccountId };
const [count, bounds, pending, transferred, totals] = await Promise.all([ const [count, bounds, pending, transferred, totals] = await Promise.all([
this.prisma.bankTransaction.count({ where: NOT_VOIDED }), this.prisma.bankTransaction.count({
where: { AND: [account, NOT_VOIDED] },
}),
this.prisma.bankTransaction.aggregate({ this.prisma.bankTransaction.aggregate({
where: NOT_VOIDED, where: { AND: [account, NOT_VOIDED] },
_min: { transactionDate: true }, _min: { transactionDate: true },
_max: { transactionDate: true }, _max: { transactionDate: true },
}), }),
this.prisma.bankTransaction.count({ this.prisma.bankTransaction.count({
where: { AND: [{ cleared: false }, NOT_VOIDED] }, where: { AND: [account, { cleared: false }, NOT_VOIDED] },
}), }),
this.prisma.bankTransaction.count({ this.prisma.bankTransaction.count({
where: { AND: [{ transferred: true }, NOT_VOIDED] }, where: { AND: [account, { transferred: true }, NOT_VOIDED] },
}), }),
this.totalsFor({}), this.totalsFor(account),
]); ]);
return { return {
@@ -261,14 +285,16 @@ export class BankService {
}; };
} }
/** Year list for the period filter, newest first. */ /** Year list for the period filter, newest first, for one chequera. */
async facets() { async facets(bankAccountId: string) {
// Tagged-template `$queryRaw`: the interpolation below is a bound
// parameter, not string concatenation.
const years = await this.prisma.$queryRaw< const years = await this.prisma.$queryRaw<
{ year: number; count: bigint | number | string }[] { year: number; count: bigint | number | string }[]
>` >`
SELECT YEAR(transactionDate) AS year, COUNT(*) AS count SELECT YEAR(transactionDate) AS year, COUNT(*) AS count
FROM bank_transactions FROM bank_transactions
WHERE voidedAt IS NULL WHERE voidedAt IS NULL AND bankAccountId = ${bankAccountId}
GROUP BY year GROUP BY year
ORDER BY year DESC ORDER BY year DESC
`; `;
@@ -287,8 +313,12 @@ export class BankService {
* `BAN` table holds only the bank's name), so the register starts at zero on * `BAN` table holds only the bank's name), so the register starts at zero on
* its first row in 2013 and the running figure is the net movement since * its first row in 2013 and the running figure is the net movement since
* then. Labelled as such in the UI so it is never read as a statement balance. * then. Labelled as such in the UI so it is never read as a statement balance.
*
* Both rollups take the SAME `bankAccountId`. Scoping only one of them would
* leave the year list and its month drill-down describing different books —
* wrong in a way that still looks right.
*/ */
async summary(year?: number) { async summary(bankAccountId: string, year?: number) {
const years = await this.prisma.$queryRaw<PeriodRow[]>` const years = await this.prisma.$queryRaw<PeriodRow[]>`
SELECT SELECT
YEAR(transactionDate) AS period, YEAR(transactionDate) AS period,
@@ -297,7 +327,7 @@ export class BankService {
SUM(CASE WHEN amount < 0 THEN amount ELSE 0 END) AS expense, SUM(CASE WHEN amount < 0 THEN amount ELSE 0 END) AS expense,
SUM(amount) AS net SUM(amount) AS net
FROM bank_transactions FROM bank_transactions
WHERE voidedAt IS NULL WHERE voidedAt IS NULL AND bankAccountId = ${bankAccountId}
GROUP BY period GROUP BY period
ORDER BY period ASC ORDER BY period ASC
`; `;
@@ -311,7 +341,9 @@ export class BankService {
SUM(CASE WHEN amount < 0 THEN amount ELSE 0 END) AS expense, SUM(CASE WHEN amount < 0 THEN amount ELSE 0 END) AS expense,
SUM(amount) AS net SUM(amount) AS net
FROM bank_transactions FROM bank_transactions
WHERE YEAR(transactionDate) = ${year} AND voidedAt IS NULL WHERE YEAR(transactionDate) = ${year}
AND voidedAt IS NULL
AND bankAccountId = ${bankAccountId}
GROUP BY period GROUP BY period
ORDER BY period ASC ORDER BY period ASC
` `
@@ -365,13 +397,135 @@ export class BankService {
}; };
} }
// --- accounts -------------------------------------------------------------
/**
* Every chequera, closed ones included — a closed account still has to be
* selectable to read its history, it just isn't offered for new captures.
*/
async listAccounts() {
const rows = await this.prisma.bankAccount.findMany({
orderBy: [{ active: "desc" }, { label: "asc" }],
select: {
id: true,
label: true,
currency: true,
businessLine: true,
active: true,
bank: { select: { id: true, name: true, country: true } },
},
});
return rows.map((a) => ({
id: a.id,
label: a.label,
currency: a.currency,
businessLine: a.businessLine,
active: a.active,
bankId: a.bank.id,
bankName: a.bank.name,
bankCountry: a.bank.country,
}));
}
async listBanks() {
return this.prisma.bank.findMany({
orderBy: { name: "asc" },
select: { id: true, name: true, country: true },
});
}
/**
* Resolve an account id from a request, or reject. Every read route funnels
* through this so a bad/missing id is a 400 rather than a silently empty
* register that reads as "this account has no movements".
*/
async requireAccount(bankAccountId: string | undefined) {
if (!bankAccountId || !bankAccountId.trim())
throw new BadRequestException("Falta la cuenta bancaria (bankAccountId)");
const account = await this.prisma.bankAccount.findUnique({
where: { id: bankAccountId },
select: { id: true, label: true, currency: true, active: true },
});
if (!account)
throw new NotFoundException(`Cuenta bancaria ${bankAccountId} no existe`);
return account;
}
async createBank(dto: CreateBankDto) {
return this.prisma.bank.create({
data: { name: dto.name.trim(), country: dto.country?.trim() || null },
});
}
async updateBank(id: string, dto: UpdateBankDto) {
await this.getBankOr404(id);
return this.prisma.bank.update({
where: { id },
data: {
...(dto.name !== undefined ? { name: dto.name.trim() } : {}),
...(dto.country !== undefined
? { country: dto.country.trim() || null }
: {}),
},
});
}
private async getBankOr404(id: string) {
const bank = await this.prisma.bank.findUnique({
where: { id },
select: { id: true },
});
if (!bank) throw new NotFoundException(`Banco ${id} no existe`);
return bank;
}
async createAccount(dto: CreateBankAccountDto) {
await this.getBankOr404(dto.bankId);
return this.prisma.bankAccount.create({
data: {
bankId: dto.bankId,
label: dto.label.trim(),
currency: dto.currency,
businessLine: dto.businessLine ?? null,
active: dto.active ?? true,
},
});
}
/**
* `currency` is intentionally absent from the update DTO: the movements
* already booked in this account are denominated in it, so changing it would
* silently re-denominate history rather than convert it.
*/
async updateAccount(id: string, dto: UpdateBankAccountDto) {
await this.requireAccount(id);
if (dto.bankId !== undefined) await this.getBankOr404(dto.bankId);
return this.prisma.bankAccount.update({
where: { id },
data: {
...(dto.bankId !== undefined ? { bankId: dto.bankId } : {}),
...(dto.label !== undefined ? { label: dto.label.trim() } : {}),
...(dto.businessLine !== undefined
? { businessLine: dto.businessLine }
: {}),
...(dto.active !== undefined ? { active: dto.active } : {}),
},
});
}
// --- writes (append + void) ----------------------------------------------- // --- writes (append + void) -----------------------------------------------
async createMovement(dto: CreateBankMovementDto) { async createMovement(dto: CreateBankMovementDto) {
const date = new Date(dto.transactionDate); const date = new Date(dto.transactionDate);
if (isNaN(date.getTime())) throw new BadRequestException("Fecha inválida"); if (isNaN(date.getTime())) throw new BadRequestException("Fecha inválida");
const account = await this.requireAccount(dto.bankAccountId);
if (!account.active)
throw new BadRequestException(
`La cuenta "${account.label}" está cerrada; no admite movimientos nuevos.`,
);
return this.prisma.bankTransaction.create({ return this.prisma.bankTransaction.create({
data: { data: {
bankAccountId: account.id,
amount: dto.amount, amount: dto.amount,
transactionDate: date, transactionDate: date,
concept: dto.concept, concept: dto.concept,
+180
View File
@@ -0,0 +1,180 @@
import { Prisma } from "@jorgecuadros/database";
import {
BALANCE_FLOOR_JOIN,
BALANCE_FORWARD_TYPE,
BillingService,
NOT_SUPERSEDED,
} from "./billing.service";
/**
* The balance floor drops rows a later BALANCE FORWARD already accounts for.
*
* It is worth testing because it fails silently: nothing throws, the numbers are
* just wrong, and they were wrong for years — the whole book read +20.6M MXN in
* credit because every customer's pre-cutover history was counted twice, once
* inside their opening balance and once as itself.
*/
describe("balance floor", () => {
describe("SQL fragments", () => {
it("binds the type name rather than interpolating it", () => {
// A literal would be a second place to edit if the label ever changes,
// and this string reaches SQL from a module constant.
expect(BALANCE_FLOOR_JOIN.values).toEqual([BALANCE_FORWARD_TYPE]);
});
it("keys the floor to the row's own customer", () => {
// Without this the derived table cross-joins and every customer inherits
// the earliest BALANCE FORWARD in the book.
expect(BALANCE_FLOOR_JOIN.sql).toContain(
"bfloor ON bfloor.customerId = t.customerId",
);
});
it("takes the most recent opening balance, not the first", () => {
// A customer accumulates one BALANCE FORWARD per year. MIN would floor at
// the oldest and leave every intervening year double-counted.
expect(BALANCE_FLOOR_JOIN.sql).toContain("MAX(bf.transactionDate)");
expect(BALANCE_FLOOR_JOIN.sql).not.toContain("MIN(bf.transactionDate)");
});
it("ignores voided opening balances when locating the floor", () => {
expect(BALANCE_FLOOR_JOIN.sql).toContain("bf.voidedAt IS NULL");
});
it("is inclusive of the opening balance row itself", () => {
// `>` instead of `>=` would drop the carried balance and understate every
// customer by exactly that amount.
expect(NOT_SUPERSEDED.sql).toContain("t.transactionDate >= bfloor.floorDate");
expect(NOT_SUPERSEDED.sql).not.toMatch(/transactionDate\s*>\s*bfloor/);
});
it("leaves customers with no opening balance untouched", () => {
// NULL comparisons are never true, so without the explicit IS NULL branch
// a customer who has no BALANCE FORWARD row loses their entire ledger.
expect(NOT_SUPERSEDED.sql).toContain("bfloor.floorDate IS NULL");
});
it("only ever references the alias the join defines", () => {
// The predicate is useless without the join; pairing them wrongly is a
// runtime "unknown column", so keep the alias identical in both.
const aliases = NOT_SUPERSEDED.sql.match(/bfloor\.\w+/g) ?? [];
expect(aliases.length).toBeGreaterThan(0);
for (const ref of aliases) {
expect(BALANCE_FLOOR_JOIN.sql).toContain(ref.split(".")[1]);
}
});
});
describe("statement()", () => {
/**
* One customer means one floor date, so the statement uses a scalar lookup
* instead of the join. Asserting on the `where` Prisma is handed is the only
* way to see it without a database.
*/
function serviceWith(floor: Date | null) {
const findMany = jest.fn().mockResolvedValue([]);
const prisma = {
customer: {
findUnique: jest.fn().mockResolvedValue({
id: "c1",
name: "CUADROS, JORGE H.",
preferredCurrency: "USD",
_count: { properties: 0, policies: 0 },
}),
},
transaction: {
findFirst: jest
.fn()
.mockResolvedValue(floor ? { transactionDate: floor } : null),
findMany,
},
};
return {
service: new BillingService(prisma as never),
prisma,
findMany,
};
}
it("looks the floor up from the customer's newest opening balance", async () => {
const { service, prisma } = serviceWith(new Date("2026-01-01T00:00:00Z"));
await service.statement("c1");
expect(prisma.transaction.findFirst).toHaveBeenCalledWith(
expect.objectContaining({
where: {
customerId: "c1",
voidedAt: null,
type: { nameEn: BALANCE_FORWARD_TYPE },
},
orderBy: { transactionDate: "desc" },
select: { transactionDate: true },
}),
);
});
it("bounds the statement at the floor, inclusive", async () => {
const floor = new Date("2026-01-01T00:00:00Z");
const { service, findMany } = serviceWith(floor);
await service.statement("c1");
expect(findMany.mock.calls[0][0].where).toMatchObject({
customerId: "c1",
transactionDate: { gte: floor },
});
});
it("applies no date bound when the customer has no opening balance", async () => {
const { service, findMany } = serviceWith(null);
await service.statement("c1");
expect(findMany.mock.calls[0][0].where).not.toHaveProperty(
"transactionDate",
);
});
it("keeps the source-table exclusion alongside the floor", async () => {
// The two guards answer different questions — one reproduces legacy's
// DATOS2-only materialization, the other drops superseded history — and
// dropping either one changes the customer's balance.
const { service, findMany } = serviceWith(new Date("2026-01-01T00:00:00Z"));
await service.statement("c1");
const where = findMany.mock.calls[0][0].where;
expect(where.OR).toEqual([
{ legacySourceTable: null },
{ legacySourceTable: { notIn: expect.arrayContaining(["EFECTIVO"]) } },
]);
});
});
describe("regression: NUMid 501", () => {
/**
* The arithmetic that exposed the bug, pinned so it cannot silently return.
* Figures measured against the live ledger on 2026-08-05.
*/
const openingBalance = new Prisma.Decimal("-6732.29");
const activitySinceOpening = new Prisma.Decimal("-7333.00");
const preCutoverCashAlreadyInOpening = new Prisma.Decimal("3596.00");
it("matches the legacy portal once superseded rows are dropped", () => {
expect(openingBalance.plus(activitySinceOpening).toFixed(2)).toBe(
"-14065.29",
);
});
it("reproduces the wrong figure when they are not", () => {
expect(
openingBalance
.plus(activitySinceOpening)
.plus(preCutoverCashAlreadyInOpening)
.toFixed(2),
).toBe("-10469.29");
});
});
});
+63 -1
View File
@@ -1,4 +1,5 @@
import { import {
BadRequestException,
Body, Body,
Controller, Controller,
Get, Get,
@@ -22,7 +23,11 @@ import {
LedgerDirection, LedgerDirection,
MovementSort, MovementSort,
} from "./billing.service"; } from "./billing.service";
import { CreateMovementDto } from "./movement.dto"; import {
BatchCreateDto,
CreateMovementDto,
ResolveOutstandingDto,
} from "./movement.dto";
const DOMAINS: TransactionDomain[] = ["UTILITY", "INSURANCE", "TRUST"]; const DOMAINS: TransactionDomain[] = ["UTILITY", "INSURANCE", "TRUST"];
const CURRENCIES: LedgerCurrency[] = ["MXN", "USD"]; const CURRENCIES: LedgerCurrency[] = ["MXN", "USD"];
@@ -46,6 +51,11 @@ function one<T>(allowed: T[], value: string | undefined): T | undefined {
return allowed.includes(value as T) ? (value as T) : undefined; return allowed.includes(value as T) ? (value as T) : undefined;
} }
/** Tri-state query flag: "true"/"false" filter, anything else means no filter. */
function flag(v: string | undefined): boolean | undefined {
return v === "true" ? true : v === "false" ? false : undefined;
}
/** A `YYYY-MM-DD` bound; anything unparseable is treated as absent. */ /** A `YYYY-MM-DD` bound; anything unparseable is treated as absent. */
function parseDate(v: string | undefined, endOfDay = false): Date | undefined { function parseDate(v: string | undefined, endOfDay = false): Date | undefined {
if (!v) return undefined; if (!v) return undefined;
@@ -97,6 +107,18 @@ export class BillingController {
}); });
} }
/**
* Every movement cut against one check, with its total — the reconciliation
* view replacing the legacy REPORTE CHEQUE COUNT. Declared before the
* `customers/:id` and `:id`-shaped routes so the literal path wins.
*/
@Get("by-check")
byCheck(@Query("checkNumber") checkNumber?: string) {
const n = checkNumber?.trim();
if (!n) throw new BadRequestException("checkNumber es obligatorio");
return this.billing.byCheck(n);
}
/** One customer's full statement across both business lines. */ /** One customer's full statement across both business lines. */
@Get("customers/:id") @Get("customers/:id")
statement(@Param("id") id: string) { statement(@Param("id") id: string) {
@@ -115,6 +137,8 @@ export class BillingController {
@Query("typeId") typeId?: string, @Query("typeId") typeId?: string,
@Query("source") source?: string, @Query("source") source?: string,
@Query("customerId") customerId?: string, @Query("customerId") customerId?: string,
@Query("outstanding") outstanding?: string,
@Query("checkNumber") checkNumber?: string,
@Query("from") from?: string, @Query("from") from?: string,
@Query("to") to?: string, @Query("to") to?: string,
@Query("sort") sort?: string, @Query("sort") sort?: string,
@@ -129,6 +153,8 @@ export class BillingController {
typeId: typeId || undefined, typeId: typeId || undefined,
source: source || undefined, source: source || undefined,
customerId: customerId || undefined, customerId: customerId || undefined,
outstanding: flag(outstanding),
checkNumber: checkNumber?.trim() || undefined,
from: parseDate(from), from: parseDate(from),
to: parseDate(to, true), to: parseDate(to, true),
sort: one(MOVEMENT_SORTS, sort) ?? "date_desc", sort: one(MOVEMENT_SORTS, sort) ?? "date_desc",
@@ -150,6 +176,42 @@ export class BillingController {
return tx; return tx;
} }
/**
* Batch capture: many customers' receipts against one physical check.
* Same ability as single capture — batching is still capturing.
*/
@Post("batch")
@RequireAbility("ledger:create")
async createBatch(@Body() dto: BatchCreateDto, @Req() req: Request) {
const result = await this.billing.createBatch(dto);
void this.audit.log(this.actingId(req), "ledger.batch", {
checkNumber: dto.checkNumber,
count: result.count,
total: result.total,
currency: result.currency,
});
return result;
}
/**
* Resolve an outstanding (NOPAGO) row — `ledger:create`, not `ledger:void`:
* resolving completes a capture, it doesn't reverse one.
*/
@Post(":id/resolve-outstanding")
@RequireAbility("ledger:create")
async resolveOutstanding(
@Param("id") id: string,
@Body() dto: ResolveOutstandingDto,
@Req() req: Request,
) {
const tx = await this.billing.resolveOutstanding(id, dto);
void this.audit.log(this.actingId(req), "ledger.resolve-outstanding", {
transactionId: id,
checkNumber: dto.checkNumber,
});
return tx;
}
@Post(":id/void") @Post(":id/void")
@RequireAbility("ledger:void") @RequireAbility("ledger:void")
async void(@Param("id") id: string, @Req() req: Request) { async void(@Param("id") id: string, @Req() req: Request) {
+3
View File
@@ -5,5 +5,8 @@ import { BillingService } from "./billing.service";
@Module({ @Module({
controllers: [BillingController], controllers: [BillingController],
providers: [BillingService], providers: [BillingService],
// The statements module posts confirmed OCR captures through
// BillingService.createBatch rather than writing Transaction rows itself.
exports: [BillingService],
}) })
export class BillingModule {} export class BillingModule {}
+467 -60
View File
@@ -1,7 +1,15 @@
import { BadRequestException, Injectable, NotFoundException } from "@nestjs/common"; import { BadRequestException, Injectable, NotFoundException } from "@nestjs/common";
import { Prisma, TransactionDomain } from "@jorgecuadros/database"; import {
Prisma,
TransactionCaptureSource,
TransactionDomain,
} from "@jorgecuadros/database";
import { PrismaService } from "../prisma/prisma.service"; import { PrismaService } from "../prisma/prisma.service";
import { CreateMovementDto } from "./movement.dto"; import {
BatchCreateDto,
CreateMovementDto,
ResolveOutstandingDto,
} from "./movement.dto";
/** /**
* Shared billing / statements module — plan step 6. * Shared billing / statements module — plan step 6.
@@ -54,12 +62,29 @@ export interface MovementParams {
typeId?: string; typeId?: string;
source?: string; source?: string;
customerId?: string; customerId?: string;
/** Restrict to captured-but-unpaid rows (the legacy NOPAGO worklist). */
outstanding?: boolean;
/** Groups a capture batch: every row cut against one physical check. */
checkNumber?: string;
/** Inclusive ISO date bounds on `transactionDate`. */ /** Inclusive ISO date bounds on `transactionDate`. */
from?: Date; from?: Date;
to?: Date; to?: Date;
sort: MovementSort; sort: MovementSort;
} }
/**
* Non-client-supplied options for a capture. Kept out of the DTO on purpose:
* these are set by the calling *module*, never by an HTTP body, so a client
* can't label its own rows as machine-captured or forge a capture ref.
* See `BillingService.createBatch` for the seam contract.
*/
export interface CaptureOptions {
/** Defaults to BATCH for the HTTP path; the OCR pipeline passes OCR. */
source?: TransactionCaptureSource;
/** Per-line artifact ids, positionally parallel to `dto.lines`. */
refs?: (string | undefined)[];
}
export interface BalanceParams { export interface BalanceParams {
query?: string; query?: string;
page: number; page: number;
@@ -79,24 +104,31 @@ interface BalanceRow {
nameMissing: number; nameMissing: number;
city: string | null; city: string | null;
state: string | null; state: string | null;
movements: bigint | number | string; movements: RawCount;
balanceMxn: Prisma.Decimal | null; balanceMxn: Prisma.Decimal | null;
balanceUsd: Prisma.Decimal | null; balanceUsd: Prisma.Decimal | null;
chargesMxn: Prisma.Decimal | null; chargesMxn: Prisma.Decimal | null;
creditsMxn: Prisma.Decimal | null; creditsMxn: Prisma.Decimal | null;
chargesUsd: Prisma.Decimal | null; chargesUsd: Prisma.Decimal | null;
creditsUsd: Prisma.Decimal | null; creditsUsd: Prisma.Decimal | null;
utilityMovements: bigint | number | string; utilityMovements: RawCount;
insuranceMovements: bigint | number | string; insuranceMovements: RawCount;
lastMovement: Date | null; lastMovement: Date | null;
} }
/** /**
* Raw-query counts come back in three shapes depending on the aggregate: * Every shape a raw-query count can arrive in. `COUNT(*)` is a bigint,
* `COUNT(*)` as bigint, `SUM(bool)` as a decimal *string*, and plain numbers. * `SUM(bool)` is a Prisma.Decimal, and plain numbers occur too — none of which
* Normalize all of them before they reach the client as JSON. * survive JSON serialization the way the client expects.
*/ */
function num(v: bigint | number | string | null | undefined): number { type RawCount = bigint | number | string | Prisma.Decimal;
/**
* Normalizes a raw-query count before it reaches the client as JSON. A bigint
* throws on JSON.stringify and a Decimal serializes to a *string*, so counts
* must not be passed through untouched.
*/
function num(v: RawCount | null | undefined): number {
if (v === null || v === undefined) return 0; if (v === null || v === undefined) return 0;
return typeof v === "number" ? v : Number(v); return typeof v === "number" ? v : Number(v);
} }
@@ -112,6 +144,91 @@ function dec(v: Prisma.Decimal | null | undefined): string {
*/ */
const NOT_VOIDED: Prisma.TransactionWhereInput = { voidedAt: null }; const NOT_VOIDED: Prisma.TransactionWhereInput = { voidedAt: null };
/**
* Outstanding ("NOPAGO") rows are captured but unpaid — the office recorded the
* bill without funds to cover it. They are excluded from every *balance*
* aggregate, exactly as the legacy `SALDOS ULTIMO 0` query did with its
* `HAVING NOPAGO = 0`: the office hasn't paid the bill, so it isn't yet owed by
* the customer. Resolving one (POST /billing/:id/resolve-outstanding) clears the
* flag and the amount starts counting.
*
* This is deliberately narrower than NOT_VOIDED. Voided rows are excluded
* everywhere; outstanding rows are excluded only from balances — the movement
* browser still totals them, because "how much water did we capture in April"
* means every captured row regardless of whether the check cleared.
*/
const NOT_OUTSTANDING: Prisma.TransactionWhereInput = { outstanding: false };
/**
* The legacy type name for a carried-forward opening balance.
*
* These rows are not movements. Access materialized one per customer per year,
* dated Jan 1, holding the closing balance of everything before it — that is
* what let the portal keep each year in its own table (`datosfreak` = current,
* `2025`, `2024`, ...) and still show a correct running balance from a single
* year's rows.
*/
export const BALANCE_FORWARD_TYPE = "BALANCE FORWARD";
/**
* Per-customer date of the most recent BALANCE FORWARD row.
*
* Joined rather than correlated: one small derived table (1,170 rows) beats a
* subquery evaluated per ledger row.
*/
export const BALANCE_FLOOR_JOIN = Prisma.sql`
LEFT JOIN (
SELECT bf.customerId, MAX(bf.transactionDate) AS floorDate
FROM transactions bf
JOIN type_transactions bft ON bft.id = bf.typeId
WHERE bft.nameEn = ${BALANCE_FORWARD_TYPE} AND bf.voidedAt IS NULL
GROUP BY bf.customerId
) bfloor ON bfloor.customerId = t.customerId`;
/**
* Excludes rows a later BALANCE FORWARD already accounts for.
*
* WHY THIS EXISTS. The platform holds both the synthetic BALANCE FORWARD rows
* and the real pre-cutover history they summarize, so summing a customer's
* whole ledger counts that history twice — once inside the opening balance,
* once as itself. NUMid 501 read -10,469.29 on the worklist against -14,065.29
* on the customer's own statement and on the legacy portal, the gap being two
* cash receipts from 2009 and 2012 that the 2026 opening balance had already
* absorbed.
*
* The scale is what settles it: summed the old way the entire book came to
* +20,605,447.86 MXN — the office owing its customers 20.6 million pesos.
* Floored, it is -56,855.90, a modest net receivable. A receivables ledger
* cannot be 20M in credit.
*
* Applies to BALANCES ONLY, in the same spirit as NOT_OUTSTANDING: the movement
* browser still totals every captured row, because "how much water did we
* capture in April" is a question about what was recorded, not about what is
* owed. Customers with no BALANCE FORWARD row (the floor is NULL) are
* unaffected.
*/
export const NOT_SUPERSEDED = Prisma.sql`(bfloor.floorDate IS NULL OR t.transactionDate >= bfloor.floorDate)`;
/**
* Source tables excluded from the customer-facing statement.
*
* The legacy portal's `datosfreak` table was materialized from DATOS2 only
* (`objects.json:1358`), so the customer's "current balance" never saw
* EFECTIVO / EFECTIVO FM3 / CHEQUE FM3 / EFECTIVO_BACKUP cash receipts, nor
* the IVA 2015 snapshot. The unified `transactions` table has all of them, so
* the statement must drop them to match the legacy number the customer has
* been quoted for years. The staff-facing balances worklist and movement
* browser keep them — they're real money, just tracked separately
* (FM3 = visa fee stream, EFECTIVO = cash receipt stream).
*/
const STATEMENT_EXCLUDED_SOURCE_TABLES: readonly string[] = [
"EFECTIVO",
"EFECTIVO_BACKUP",
"EFECTIVO FM3",
"CHEQUE FM3",
"IVA 2015",
];
@Injectable() @Injectable()
export class BillingService { export class BillingService {
constructor(private readonly prisma: PrismaService) {} constructor(private readonly prisma: PrismaService) {}
@@ -140,6 +257,10 @@ export class BillingService {
if (p.typeId) and.push({ typeId: p.typeId }); if (p.typeId) and.push({ typeId: p.typeId });
if (p.source) and.push({ legacySourceTable: p.source }); if (p.source) and.push({ legacySourceTable: p.source });
if (p.customerId) and.push({ customerId: p.customerId }); if (p.customerId) and.push({ customerId: p.customerId });
if (p.outstanding !== undefined) and.push({ outstanding: p.outstanding });
// Exact match, not `contains`: this is the by-check reconciliation lookup,
// where "1234" must not drag in "51234".
if (p.checkNumber) and.push({ checkNumber: p.checkNumber });
if (p.from || p.to) { if (p.from || p.to) {
and.push({ and.push({
transactionDate: { transactionDate: {
@@ -196,6 +317,7 @@ export class BillingService {
message: true, message: true,
legacySourceTable: true, legacySourceTable: true,
voidedAt: true, voidedAt: true,
outstanding: true,
type: { select: { nameEn: true, nameEs: true } }, type: { select: { nameEn: true, nameEs: true } },
customer: { customer: {
select: { id: true, name: true, nameSource: true, city: true }, select: { id: true, name: true, nameSource: true, city: true },
@@ -243,6 +365,7 @@ export class BillingService {
source: r.legacySourceTable, source: r.legacySourceTable,
type: r.type, type: r.type,
voided: r.voidedAt != null, voided: r.voidedAt != null,
outstanding: r.outstanding,
customerId: r.customer.id, customerId: r.customer.id,
customerName: r.customer.name, customerName: r.customer.name,
customerNameSource: r.customer.nameSource, customerNameSource: r.customer.nameSource,
@@ -336,19 +459,24 @@ export class BillingService {
MAX(t.transactionDate) AS lastMovement MAX(t.transactionDate) AS lastMovement
FROM customers c FROM customers c
JOIN transactions t ON t.customerId = c.id JOIN transactions t ON t.customerId = c.id
WHERE t.voidedAt IS NULL ${nameFilter} ${txFilter} ${BALANCE_FLOOR_JOIN}
WHERE t.voidedAt IS NULL AND t.outstanding = 0 AND ${NOT_SUPERSEDED} ${nameFilter} ${txFilter}
GROUP BY c.id, c.name, c.nameSource, c.nameMissing, c.city, c.state GROUP BY c.id, c.name, c.nameSource, c.nameMissing, c.city, c.state
${having} ${having}
${orderBy} ${orderBy}
LIMIT ${pageSize} OFFSET ${(page - 1) * pageSize} LIMIT ${pageSize} OFFSET ${(page - 1) * pageSize}
`; `;
const counted = await this.prisma.$queryRaw<{ total: bigint | number | string }[]>` const counted = await this.prisma.$queryRaw<{ total: RawCount }[]>`
SELECT COUNT(*) AS total FROM ( SELECT COUNT(*) AS total FROM (
SELECT c.id SELECT c.id
FROM customers c FROM customers c
JOIN transactions t ON t.customerId = c.id JOIN transactions t ON t.customerId = c.id
WHERE 1 = 1 ${nameFilter} ${txFilter} ${BALANCE_FLOOR_JOIN}
-- Must match the page query's filters exactly, or the total disagrees
-- with the rows. (The void exclusion was missing here before the
-- outstanding work; a voided-only customer inflated the count.)
WHERE t.voidedAt IS NULL AND t.outstanding = 0 AND ${NOT_SUPERSEDED} ${nameFilter} ${txFilter}
GROUP BY c.id GROUP BY c.id
${having} ${having}
) x ) x
@@ -389,9 +517,18 @@ export class BillingService {
}; };
} }
/** Top-line figures for the billing page header. */ /**
* Top-line figures for the billing page header.
*
* Two different questions live here and they use different row sets.
* `movements`, `ledgerCustomers`, `crossLineCustomers` and the date range are
* INVENTORY — what is stored — and count everything not voided. Everything
* under `byCurrency` / `byDomain` is a BALANCE, so it applies NOT_SUPERSEDED
* and drops rows an opening balance already accounts for. The four aggregates
* moved from Prisma groupBy to raw SQL to express that join; groupBy cannot.
*/
async stats() { async stats() {
const [movements, ledgerCustomers, byCurrency, byDomain] = await Promise.all([ const [movements, ledgerCustomers] = await Promise.all([
this.prisma.transaction.count({ where: NOT_VOIDED }), this.prisma.transaction.count({ where: NOT_VOIDED }),
this.prisma.transaction this.prisma.transaction
.findMany({ .findMany({
@@ -400,34 +537,47 @@ export class BillingService {
select: { customerId: true }, select: { customerId: true },
}) })
.then((r) => r.length), .then((r) => r.length),
this.prisma.transaction.groupBy({
by: ["currency"],
where: NOT_VOIDED,
_sum: { amount: true },
_count: { _all: true },
}),
this.prisma.transaction.groupBy({
by: ["domain", "currency"],
where: NOT_VOIDED,
_sum: { amount: true },
_count: { _all: true },
}),
]); ]);
const charges = await this.prisma.transaction.groupBy({ const byCurrency = await this.prisma.$queryRaw<
by: ["currency"], {
where: { AND: [{ amount: { lt: 0 } }, NOT_VOIDED] }, currency: string;
_sum: { amount: true }, net: Prisma.Decimal | null;
_count: { _all: true }, count: RawCount;
}); charges: Prisma.Decimal | null;
const credits = await this.prisma.transaction.groupBy({ chargeCount: RawCount;
by: ["currency"], credits: Prisma.Decimal | null;
where: { AND: [{ amount: { gt: 0 } }, NOT_VOIDED] }, creditCount: RawCount;
_sum: { amount: true }, }[]
_count: { _all: true }, >`
}); SELECT t.currency AS currency,
const chargeMap = new Map(charges.map((c) => [c.currency, c])); SUM(t.amount) AS net,
const creditMap = new Map(credits.map((c) => [c.currency, c])); COUNT(*) AS count,
SUM(CASE WHEN t.amount < 0 THEN t.amount ELSE 0 END) AS charges,
SUM(t.amount < 0) AS chargeCount,
SUM(CASE WHEN t.amount > 0 THEN t.amount ELSE 0 END) AS credits,
SUM(t.amount > 0) AS creditCount
FROM transactions t
${BALANCE_FLOOR_JOIN}
WHERE t.voidedAt IS NULL AND ${NOT_SUPERSEDED}
GROUP BY t.currency
`;
const byDomain = await this.prisma.$queryRaw<
{
domain: string;
currency: string;
net: Prisma.Decimal | null;
count: RawCount;
}[]
>`
SELECT t.domain AS domain, t.currency AS currency,
SUM(t.amount) AS net, COUNT(*) AS count
FROM transactions t
${BALANCE_FLOOR_JOIN}
WHERE t.voidedAt IS NULL AND ${NOT_SUPERSEDED}
GROUP BY t.domain, t.currency
`;
// How many customers sit on each side of the line, per currency — the // How many customers sit on each side of the line, per currency — the
// headline for a receivables view. Counted in SQL; a customer can be // headline for a receivables view. Counted in SQL; a customer can be
@@ -435,16 +585,19 @@ export class BillingService {
const sides = await this.prisma.$queryRaw< const sides = await this.prisma.$queryRaw<
{ {
currency: string; currency: string;
owing: bigint | number | string; owing: RawCount;
inCredit: bigint | number | string; inCredit: RawCount;
}[] }[]
>` >`
SELECT currency, SELECT currency,
SUM(bal < -0.005) AS owing, SUM(bal < -0.005) AS owing,
SUM(bal > 0.005) AS inCredit SUM(bal > 0.005) AS inCredit
FROM ( FROM (
SELECT customerId, currency, SUM(amount) AS bal SELECT t.customerId, t.currency, SUM(t.amount) AS bal
FROM transactions WHERE voidedAt IS NULL GROUP BY customerId, currency FROM transactions t
${BALANCE_FLOOR_JOIN}
WHERE t.voidedAt IS NULL AND ${NOT_SUPERSEDED}
GROUP BY t.customerId, t.currency
) x ) x
GROUP BY currency GROUP BY currency
`; `;
@@ -465,7 +618,7 @@ export class BillingService {
// Customers whose ledger spans both business lines — the whole reason this // Customers whose ledger spans both business lines — the whole reason this
// module is one view instead of two. // module is one view instead of two.
const crossLine = await this.prisma.$queryRaw<{ n: bigint | number | string }[]>` const crossLine = await this.prisma.$queryRaw<{ n: RawCount }[]>`
SELECT COUNT(*) AS n FROM ( SELECT COUNT(*) AS n FROM (
SELECT customerId FROM transactions WHERE voidedAt IS NULL SELECT customerId FROM transactions WHERE voidedAt IS NULL
GROUP BY customerId HAVING COUNT(DISTINCT domain) > 1 GROUP BY customerId HAVING COUNT(DISTINCT domain) > 1
@@ -480,20 +633,20 @@ export class BillingService {
lastMovement: lastRow?.transactionDate ?? null, lastMovement: lastRow?.transactionDate ?? null,
byCurrency: byCurrency.map((c) => ({ byCurrency: byCurrency.map((c) => ({
currency: c.currency, currency: c.currency,
net: c._sum.amount, net: c.net,
count: c._count._all, count: num(c.count),
charges: chargeMap.get(c.currency)?._sum.amount ?? null, charges: c.charges,
chargeCount: chargeMap.get(c.currency)?._count._all ?? 0, chargeCount: num(c.chargeCount),
credits: creditMap.get(c.currency)?._sum.amount ?? null, credits: c.credits,
creditCount: creditMap.get(c.currency)?._count._all ?? 0, creditCount: num(c.creditCount),
owing: num(sideMap.get(c.currency)?.owing), owing: num(sideMap.get(c.currency)?.owing),
inCredit: num(sideMap.get(c.currency)?.inCredit), inCredit: num(sideMap.get(c.currency)?.inCredit),
})), })),
byDomain: byDomain.map((d) => ({ byDomain: byDomain.map((d) => ({
domain: d.domain, domain: d.domain,
currency: d.currency, currency: d.currency,
net: d._sum.amount, net: d.net,
count: d._count._all, count: num(d.count),
})), })),
}; };
} }
@@ -520,7 +673,7 @@ export class BillingService {
}); });
const years = await this.prisma.$queryRaw< const years = await this.prisma.$queryRaw<
{ year: number; count: bigint | number | string }[] { year: number; count: RawCount }[]
>` >`
SELECT YEAR(transactionDate) AS year, COUNT(*) AS count SELECT YEAR(transactionDate) AS year, COUNT(*) AS count
FROM transactions WHERE voidedAt IS NULL GROUP BY year ORDER BY year DESC FROM transactions WHERE voidedAt IS NULL GROUP BY year ORDER BY year DESC
@@ -578,8 +731,47 @@ export class BillingService {
throw new NotFoundException(`Customer ${customerId} not found`); throw new NotFoundException(`Customer ${customerId} not found`);
} }
// One customer, so the balance floor is a single date rather than the
// derived table the aggregate queries join. See NOT_SUPERSEDED: rows before
// the opening balance are already inside it, and showing them would both
// double the total and make every balanceAfter below wrong.
//
// This is also what stops FEE ANUAL and fee15 leaking in. They are not in
// STATEMENT_EXCLUDED_SOURCE_TABLES — that list exists to reproduce legacy's
// DATOS2-only `datosfreak`, and it was letting 2,092 pre-cutover fee rows
// across 1,062 customers through, skewing the statement by -5,129,764
// against the number those customers have been quoted for years. Dating
// rather than source is the right test: a FEE ANUAL row *after* the opening
// balance is a real charge and still counts.
const floor = await this.prisma.transaction.findFirst({
where: {
customerId,
voidedAt: null,
type: { nameEn: BALANCE_FORWARD_TYPE },
},
orderBy: { transactionDate: "desc" },
select: { transactionDate: true },
});
const rows = await this.prisma.transaction.findMany({ const rows = await this.prisma.transaction.findMany({
where: { customerId }, where: {
customerId,
...(floor ? { transactionDate: { gte: floor.transactionDate } } : {}),
// NULL-safe exclusion. `notIn` alone compiles to SQL `NOT IN`, and
// `NULL NOT IN (...)` is NULL, not true — so every app-captured row
// (which has no legacySourceTable) silently vanished from the
// statement while still showing in the movement browser. Rows the app
// books must appear on the customer's statement, so the null case is
// spelled out.
OR: [
{ legacySourceTable: null },
{
legacySourceTable: {
notIn: STATEMENT_EXCLUDED_SOURCE_TABLES as string[],
},
},
],
},
orderBy: [{ transactionDate: "asc" }, { id: "asc" }], orderBy: [{ transactionDate: "asc" }, { id: "asc" }],
select: { select: {
id: true, id: true,
@@ -593,6 +785,7 @@ export class BillingService {
message: true, message: true,
legacySourceTable: true, legacySourceTable: true,
voidedAt: true, voidedAt: true,
outstanding: true,
type: { select: { nameEn: true, nameEs: true } }, type: { select: { nameEn: true, nameEs: true } },
}, },
}); });
@@ -601,9 +794,10 @@ export class BillingService {
const movements = rows.map((r) => { const movements = rows.map((r) => {
const voided = r.voidedAt != null; const voided = r.voidedAt != null;
const prev = running.get(r.currency) ?? new Prisma.Decimal(0); const prev = running.get(r.currency) ?? new Prisma.Decimal(0);
// A voided row does not move the running balance — it shows struck-through // Neither a voided row nor an outstanding (unpaid) one moves the running
// with the balance unchanged from the previous live movement. // balance — both show tagged, with the balance unchanged from the previous
const next = voided ? prev : prev.plus(r.amount); // live movement. Outstanding rows start counting once resolved.
const next = voided || r.outstanding ? prev : prev.plus(r.amount);
running.set(r.currency, next); running.set(r.currency, next);
return { return {
id: r.id, id: r.id,
@@ -619,6 +813,7 @@ export class BillingService {
source: r.legacySourceTable, source: r.legacySourceTable,
type: r.type, type: r.type,
voided, voided,
outstanding: r.outstanding,
/** Balance in this row's currency after applying it. */ /** Balance in this row's currency after applying it. */
balanceAfter: next.toFixed(2), balanceAfter: next.toFixed(2),
}; };
@@ -652,7 +847,9 @@ export class BillingService {
>(); >();
for (const r of rows) { for (const r of rows) {
if (r.voidedAt != null) continue; // voided rows never enter a total // Voided rows never enter a total; outstanding rows don't either until
// they're resolved (legacy SALDOS ULTIMO 0's `HAVING NOPAGO = 0`).
if (r.voidedAt != null || r.outstanding) continue;
const c = const c =
perCurrency.get(r.currency) ?? perCurrency.get(r.currency) ??
{ {
@@ -700,7 +897,7 @@ export class BillingService {
{ name: string; currency: string; total: Prisma.Decimal; count: number } { name: string; currency: string; total: Prisma.Decimal; count: number }
>(); >();
for (const r of rows) { for (const r of rows) {
if (r.voidedAt != null) continue; if (r.voidedAt != null || r.outstanding) continue;
if (!r.amount.lessThan(0)) continue; if (!r.amount.lessThan(0)) continue;
const name = r.type?.nameEs || r.type?.nameEn || "Sin clasificar"; const name = r.type?.nameEs || r.type?.nameEn || "Sin clasificar";
const key = `${name}|${r.currency}`; const key = `${name}|${r.currency}`;
@@ -772,10 +969,220 @@ export class BillingService {
reference: dto.reference, reference: dto.reference,
checkNumber: dto.checkNumber, checkNumber: dto.checkNumber,
message: dto.message, message: dto.message,
outstanding: dto.outstanding ?? false,
captureSource: "MANUAL",
}, },
}); });
} }
/**
* Batch capture by check — many customers' receipts against one physical
* check. One `$transaction`, so a bad line rejects the whole batch rather
* than leaving a half-captured check that reconciles against nothing.
*
* Returns the check-level total alongside the rows so the UI can show it
* against the physical check amount, which is the entire point of the legacy
* flow this replaces (`CAPTURA *` feeding `EDITA CHEQUE COUNT`).
*
* ── Integration seam for OCR auto-capture (RECEIPT_CAPTURE_SPEC §2) ────────
* This method is the SINGLE write path for multi-row capture, and the OCR
* pipeline is required to post through it rather than writing `Transaction`
* rows itself — one validation path, one audit trail. Three guarantees exist
* for that caller specifically, and must not be broken:
*
* 1. `items[i]` corresponds to `dto.lines[i]`. Prisma's array
* `$transaction` preserves order, so the caller can zip the result back
* onto its own records — which is how `StatementDocument.postedTransactionId`
* gets set after a confirmed batch posts.
* 2. `opts.refs[i]` stamps `captureRef` on row `i` (a `StatementDocument.id`).
* Re-posting a ref that already has a live row is rejected, so a
* double-clicked "confirm" or a retried job cannot double-charge a
* customer. Voided rows don't block a re-post — a corrected statement
* must be re-postable after its bad row is voided.
* 3. `opts.source` records the capture path; it is NOT accepted over HTTP,
* so a client cannot label its hand-keyed rows as machine-captured.
*
* Everything the OCR module adds on top (batches, per-document status, the
* review queue) lives in its own module; nothing about it needs to change
* this signature.
*/
async createBatch(dto: BatchCreateDto, opts: CaptureOptions = {}) {
const date = new Date(dto.transactionDate);
if (isNaN(date.getTime())) throw new BadRequestException("Fecha inválida");
// Validate every customer up front, in one query — a per-line lookup inside
// the transaction would be N round-trips and would fail halfway through.
const ids = [...new Set(dto.lines.map((l) => l.customerId))];
const found = await this.prisma.customer.findMany({
where: { id: { in: ids } },
select: { id: true },
});
if (found.length !== ids.length) {
const known = new Set(found.map((c) => c.id));
const missing = ids.filter((id) => !known.has(id));
throw new BadRequestException(
`Cliente(s) no encontrado(s): ${missing.join(", ")}`,
);
}
// Duplicate-post guard (seam guarantee 2). Only live rows block: a voided
// row means the earlier post was reversed, so the corrected statement must
// be allowed through.
const refs = (opts.refs ?? []).filter((r): r is string => !!r);
if (refs.length) {
const clash = await this.prisma.transaction.findMany({
where: { captureRef: { in: refs }, voidedAt: null },
select: { captureRef: true },
});
if (clash.length) {
const dupes = [...new Set(clash.map((c) => c.captureRef))];
throw new BadRequestException(
`Ya existen movimientos para: ${dupes.join(", ")}`,
);
}
}
const currency = dto.currency ?? "MXN";
const source = opts.source ?? "BATCH";
const created = await this.prisma.$transaction(
dto.lines.map((line, i) =>
this.prisma.transaction.create({
data: {
customerId: line.customerId,
domain: dto.domain,
amount: line.amount,
transactionDate: date,
currency,
typeId: dto.typeId,
checkNumber: dto.checkNumber,
period: line.period,
reference: line.reference,
message: line.message,
outstanding: line.outstanding ?? false,
captureSource: source,
captureRef: opts.refs?.[i],
},
}),
),
);
// Outstanding lines are captured but unfunded, so they don't belong in the
// figure staff reconcile against the physical check.
const total = created.reduce(
(sum, t) => (t.outstanding ? sum : sum.plus(t.amount)),
new Prisma.Decimal(0),
);
return {
/** Parallel to `dto.lines` — see seam guarantee 1. */
items: created,
checkNumber: dto.checkNumber,
currency,
source,
count: created.length,
outstandingCount: created.filter((t) => t.outstanding).length,
total: total.toFixed(2),
};
}
/**
* Resolve an outstanding row: the check was finally cut. Takes the resolution
* date and check number and clears the flag, so the amount starts counting
* toward the balance. Legacy: "se actualiza registro con fecha del día y el
* cheque a pagar y quitas outstanding".
*/
async resolveOutstanding(id: string, dto: ResolveOutstandingDto) {
const tx = await this.prisma.transaction.findUnique({
where: { id },
select: { id: true, voidedAt: true, outstanding: true },
});
if (!tx) throw new NotFoundException(`Transaction ${id} not found`);
if (tx.voidedAt) {
throw new BadRequestException("El movimiento está anulado");
}
if (!tx.outstanding) {
throw new BadRequestException("El movimiento no está pendiente de pago");
}
const date = new Date(dto.resolvedDate);
if (isNaN(date.getTime())) throw new BadRequestException("Fecha inválida");
return this.prisma.transaction.update({
where: { id },
data: {
outstanding: false,
checkNumber: dto.checkNumber,
transactionDate: date,
},
});
}
/**
* Every live movement cut against one check, plus its total — the
* reconciliation view replacing `EDITA CHEQUE ALF/COUNT/NUM` and
* `REPORTE POR CHEQUE`. Voided rows are dropped entirely (they reconcile
* against nothing); outstanding rows are listed but excluded from the total,
* since the check didn't fund them.
*/
async byCheck(checkNumber: string) {
const rows = await this.prisma.transaction.findMany({
where: { checkNumber, voidedAt: null },
orderBy: [{ transactionDate: "asc" }, { id: "asc" }],
select: {
id: true,
transactionDate: true,
domain: true,
amount: true,
currency: true,
reference: true,
period: true,
message: true,
outstanding: true,
type: { select: { nameEn: true, nameEs: true } },
customer: { select: { id: true, name: true, nameSource: true } },
},
});
// Per currency: a check is one currency in practice, but the ledger has
// both and this module never sums across them.
const totals = new Map<string, { currency: string; total: Prisma.Decimal; count: number }>();
for (const r of rows) {
if (r.outstanding) continue;
const e =
totals.get(r.currency) ??
{ currency: r.currency, total: new Prisma.Decimal(0), count: 0 };
e.total = e.total.plus(r.amount);
e.count += 1;
totals.set(r.currency, e);
}
return {
checkNumber,
items: rows.map((r) => ({
id: r.id,
transactionDate: r.transactionDate,
domain: r.domain,
amount: r.amount,
currency: r.currency,
direction: r.amount.lessThan(0) ? "charge" : "credit",
reference: r.reference,
period: r.period,
message: r.message,
outstanding: r.outstanding,
type: r.type,
customerId: r.customer.id,
customerName: r.customer.name,
customerNameSource: r.customer.nameSource,
})),
count: rows.length,
outstandingCount: rows.filter((r) => r.outstanding).length,
totals: [...totals.values()].map((t) => ({
currency: t.currency,
total: t.total.toFixed(2),
count: t.count,
})),
};
}
/** Reverse a movement by marking it voided; it stops counting toward totals. */ /** Reverse a movement by marking it voided; it stops counting toward totals. */
async voidMovement(id: string, userId: string) { async voidMovement(id: string, userId: string) {
const tx = await this.prisma.transaction.findUnique({ const tx = await this.prisma.transaction.findUnique({
+58
View File
@@ -1,10 +1,16 @@
import { import {
ArrayMaxSize,
ArrayMinSize,
IsArray,
IsBoolean,
IsEnum, IsEnum,
IsNumber, IsNumber,
IsOptional, IsOptional,
IsString, IsString,
MinLength, MinLength,
ValidateNested,
} from "class-validator"; } from "class-validator";
import { Type } from "class-transformer";
import { Currency, TransactionDomain } from "@jorgecuadros/database"; import { Currency, TransactionDomain } from "@jorgecuadros/database";
/** /**
@@ -24,4 +30,56 @@ export class CreateMovementDto {
@IsOptional() @IsString() reference?: string; @IsOptional() @IsString() reference?: string;
@IsOptional() @IsString() checkNumber?: string; @IsOptional() @IsString() checkNumber?: string;
@IsOptional() @IsString() message?: string; @IsOptional() @IsString() message?: string;
/**
* Legacy "NOPAGO": the bill was captured but not actually paid (no funds).
* The row posts normally and stays visible, but is kept out of every balance
* aggregate until resolved — see BillingService's NOT_OUTSTANDING.
*/
@IsOptional() @IsBoolean() outstanding?: boolean;
}
/**
* Resolving an outstanding row: the check finally got cut, so the movement
* takes the resolution date and check number and starts counting toward the
* balance. Legacy behavior: "se actualiza registro con fecha del día y el
* cheque a pagar y quitas outstanding".
*/
export class ResolveOutstandingDto {
@IsString() @MinLength(1) checkNumber!: string;
@IsString() @MinLength(1) resolvedDate!: string;
}
/** One customer's line within a batch; check-level fields live on the parent. */
export class BatchLineDto {
@IsString() @MinLength(1) customerId!: string;
@IsNumber() amount!: number;
@IsOptional() @IsString() reference?: string;
@IsOptional() @IsString() period?: string;
@IsOptional() @IsString() message?: string;
@IsOptional() @IsBoolean() outstanding?: boolean;
}
/**
* Batch capture by check — the legacy "Editor" flow: key many customers'
* receipts against one check, then reconcile the captured total against the
* physical check. Deliberately NOT a persisted batch entity: `checkNumber` is
* already a column, and grouping by it answers every legacy by-check query.
*/
export class BatchCreateDto {
@IsEnum(TransactionDomain) domain!: TransactionDomain;
@IsString() @MinLength(1) transactionDate!: string;
@IsString() @MinLength(1) checkNumber!: string;
@IsOptional() @IsEnum(Currency) currency?: Currency;
@IsOptional() @IsString() typeId?: string;
// Capped so one request can't open a transaction over an unbounded row set;
// a physical check batch is tens of lines, not thousands.
@IsArray()
@ArrayMinSize(1)
@ArrayMaxSize(500)
@ValidateNested({ each: true })
@Type(() => BatchLineDto)
lines!: BatchLineDto[];
} }
@@ -29,6 +29,7 @@ export class CreateCustomerDto {
@IsOptional() @IsString() mobile?: string; @IsOptional() @IsString() mobile?: string;
@IsOptional() @IsString() fax?: string; @IsOptional() @IsString() fax?: string;
@IsOptional() @IsEmail() email?: string; @IsOptional() @IsEmail() email?: string;
@IsOptional() @IsBoolean() emailOptOut?: boolean;
@IsOptional() @IsString() notes?: string; @IsOptional() @IsString() notes?: string;
@IsOptional() @IsString() identificationType?: string; @IsOptional() @IsString() identificationType?: string;
@IsOptional() @IsString() identificationNumber?: string; @IsOptional() @IsString() identificationNumber?: string;
@@ -16,6 +16,7 @@ import { AbilityGuard } from "../auth/ability.guard";
import { RequireAbility } from "../auth/require-ability.decorator"; import { RequireAbility } from "../auth/require-ability.decorator";
import { AuditService } from "../common/audit.service"; import { AuditService } from "../common/audit.service";
import { CustomersService } from "./customers.service"; import { CustomersService } from "./customers.service";
import { NumidService } from "./numid.service";
import { CreateCustomerDto } from "./create-customer.dto"; import { CreateCustomerDto } from "./create-customer.dto";
import { UpdateCustomerDto } from "./update-customer.dto"; import { UpdateCustomerDto } from "./update-customer.dto";
@@ -24,6 +25,7 @@ import { UpdateCustomerDto } from "./update-customer.dto";
export class CustomersController { export class CustomersController {
constructor( constructor(
private readonly customers: CustomersService, private readonly customers: CustomersService,
private readonly numids: NumidService,
private readonly audit: AuditService, private readonly audit: AuditService,
) {} ) {}
@@ -36,6 +38,13 @@ export class CustomersController {
return this.customers.stats(); return this.customers.stats();
} }
/** Reusable portal ids, lowest first. Declared above `:id` so the literal
* path is not swallowed by the wildcard route. */
@Get("numid/candidates")
async numidCandidates() {
return { candidates: await this.numids.emptyCandidates() };
}
@Get() @Get()
list( list(
@Query("query") query?: string, @Query("query") query?: string,
@@ -95,4 +104,27 @@ export class CustomersController {
void this.audit.log(this.actingId(req), "customer.restore", { customerId: id }); void this.audit.log(this.actingId(req), "customer.restore", { customerId: id });
return c; return c;
} }
/**
* Give this customer a portal NUMid so they can log in to
* my.jorgecuadros.com. Idempotent — a customer who already has one gets it
* back rather than a second identity.
*/
@Post(":id/portal-access")
@RequireAbility("customer:portal-access")
async portalAccess(@Param("id") id: string, @Req() req: Request) {
const allocation = await this.numids.allocate(id);
if (allocation.origin !== "existing") {
// Logged with the origin and the previous holder: a recycled id is the one
// case where reading this record later has to answer "whose number was
// this before, and was it taken or minted".
void this.audit.log(this.actingId(req), "customer.portal-access", {
customerId: id,
numid: allocation.numid,
origin: allocation.origin,
previousCustomerId: allocation.previousCustomerId,
});
}
return allocation;
}
} }
+5 -1
View File
@@ -1,9 +1,13 @@
import { Module } from "@nestjs/common"; import { Module } from "@nestjs/common";
import { SettingsModule } from "../settings/settings.module";
import { CustomersController } from "./customers.controller"; import { CustomersController } from "./customers.controller";
import { CustomersService } from "./customers.service"; import { CustomersService } from "./customers.service";
import { NumidService } from "./numid.service";
@Module({ @Module({
imports: [SettingsModule],
controllers: [CustomersController], controllers: [CustomersController],
providers: [CustomersService], providers: [CustomersService, NumidService],
exports: [NumidService],
}) })
export class CustomersModule {} export class CustomersModule {}
@@ -0,0 +1,206 @@
import { ConflictException, NotFoundException } from "@nestjs/common";
import { Prisma } from "@jorgecuadros/database";
import { NumidService } from "./numid.service";
/**
* What matters about the allocator is the two things it must never do: hand the
* same id to two customers, and hand out a recycled id while Access can still
* take it back. Both are tested here; the emptiness SQL itself is exercised
* against real data by scripts/numid-audit.mjs.
*/
interface Options {
existingRef?: { legacyId: string } | null;
archived?: boolean;
missing?: boolean;
recycle?: boolean;
empty?: { numid: string; refId: string; customerId: string }[];
max?: number | null;
/** Make the first N create() calls fail the unique key, as a race would. */
createConflicts?: number;
/** Make updateMany report "nothing matched", as a lost recycle race would. */
recycleMisses?: number;
}
function build(opts: Options = {}) {
const created: { legacyId: string }[] = [];
let conflictsLeft = opts.createConflicts ?? 0;
let missesLeft = opts.recycleMisses ?? 0;
const prisma = {
customer: {
findUnique: jest.fn().mockResolvedValue(
opts.missing ? null : { id: "cust-new", archivedAt: opts.archived ? new Date() : null },
),
},
customerLegacyRef: {
findFirst: jest.fn().mockResolvedValue(opts.existingRef ?? null),
updateMany: jest.fn().mockImplementation(() => {
if (missesLeft > 0) {
missesLeft -= 1;
return Promise.resolve({ count: 0 });
}
return Promise.resolve({ count: 1 });
}),
create: jest.fn().mockImplementation(({ data }: { data: { legacyId: string } }) => {
if (conflictsLeft > 0) {
conflictsLeft -= 1;
return Promise.reject(
new Prisma.PrismaClientKnownRequestError("dup", {
code: "P2002",
clientVersion: "5",
}),
);
}
created.push(data);
return Promise.resolve(data);
}),
},
// Two different raw queries share one mock: the MAX lookup returns a single
// {max} row, everything else is the empty-candidate list.
$queryRaw: jest.fn().mockImplementation((sql: { strings?: string[]; sql?: string }) => {
const text = String((sql as unknown as { sql?: string }).sql ?? "");
if (text.includes("MAX(")) return Promise.resolve([{ max: opts.max ?? null }]);
return Promise.resolve(opts.empty ?? []);
}),
};
const settings = {
numidRecycleEmpty: jest
.fn()
.mockResolvedValue({ value: opts.recycle ?? false, source: "default" }),
};
return {
service: new NumidService(prisma as never, settings as never),
prisma,
created,
};
}
describe("NUMid allocation", () => {
it("returns the id a customer already holds instead of minting a second one", async () => {
// A double-clicked button must not fork the customer's portal identity.
const { service, prisma } = build({ existingRef: { legacyId: "501" } });
await expect(service.allocate("cust-new")).resolves.toEqual({
numid: "501",
origin: "existing",
});
expect(prisma.customerLegacyRef.create).not.toHaveBeenCalled();
});
it("allocates one past the highest id in the pool", async () => {
const { service, created } = build({ max: 1171 });
await expect(service.allocate("cust-new")).resolves.toEqual({
numid: "1172",
origin: "new",
});
expect(created[0]).toMatchObject({
sourceSystem: "utilities",
sourceTable: "DATGRAL",
legacyId: "1172",
});
});
it("starts at 1 when the pool is empty", async () => {
const { service } = build({ max: null });
await expect(service.allocate("cust-new")).resolves.toMatchObject({ numid: "1" });
});
it("does NOT recycle while the setting is off, even with candidates free", async () => {
// The default has to be the safe one: every reusable id still exists in
// Access, and a --sync run reassigns it back to its Access owner.
const { service, prisma } = build({
max: 1171,
empty: [{ numid: "1089", refId: "ref-1089", customerId: "cust-old" }],
});
await expect(service.allocate("cust-new")).resolves.toMatchObject({
numid: "1172",
origin: "new",
});
expect(prisma.customerLegacyRef.updateMany).not.toHaveBeenCalled();
});
it("takes the lowest empty id once recycling is switched on", async () => {
const { service, prisma } = build({
recycle: true,
max: 1171,
empty: [
{ numid: "1089", refId: "ref-1089", customerId: "cust-old" },
{ numid: "1094", refId: "ref-1094", customerId: "cust-other" },
],
});
await expect(service.allocate("cust-new")).resolves.toEqual({
numid: "1089",
origin: "recycled",
previousCustomerId: "cust-old",
});
// Guarded on the owner read a moment ago, so a ref that moved underneath us
// matches nothing rather than being stolen.
expect(prisma.customerLegacyRef.updateMany).toHaveBeenCalledWith({
where: { id: "ref-1089", customerId: "cust-old" },
data: { customerId: "cust-new" },
});
});
it("skips a candidate that someone else took first", async () => {
const { service } = build({
recycle: true,
max: 1171,
recycleMisses: 1,
empty: [
{ numid: "1089", refId: "ref-1089", customerId: "cust-old" },
{ numid: "1094", refId: "ref-1094", customerId: "cust-other" },
],
});
await expect(service.allocate("cust-new")).resolves.toMatchObject({
numid: "1094",
origin: "recycled",
});
});
it("falls back to a new id when recycling is on but nothing is free", async () => {
const { service } = build({ recycle: true, max: 1171, empty: [] });
await expect(service.allocate("cust-new")).resolves.toMatchObject({
numid: "1172",
origin: "new",
});
});
it("retries when two writers pick the same id", async () => {
// The unique key on (sourceSystem, sourceTable, legacyId) is what decides
// the winner; the loser must retry, never overwrite.
const { service, prisma } = build({ max: 1171, createConflicts: 1 });
await expect(service.allocate("cust-new")).resolves.toMatchObject({
numid: "1172",
origin: "new",
});
expect(prisma.customerLegacyRef.create).toHaveBeenCalledTimes(2);
});
it("gives up loudly rather than looping forever", async () => {
const { service } = build({ max: 1171, createConflicts: 99 });
await expect(service.allocate("cust-new")).rejects.toBeInstanceOf(ConflictException);
});
it("refuses an archived customer", async () => {
const { service } = build({ archived: true });
await expect(service.allocate("cust-new")).rejects.toBeInstanceOf(ConflictException);
});
it("refuses a customer that does not exist", async () => {
const { service } = build({ missing: true });
await expect(service.allocate("nope")).rejects.toBeInstanceOf(NotFoundException);
});
});
+255
View File
@@ -0,0 +1,255 @@
import {
ConflictException,
Injectable,
Logger,
NotFoundException,
} from "@nestjs/common";
import { Prisma } from "@jorgecuadros/database";
import { PrismaService } from "../prisma/prisma.service";
import { SettingsService } from "../settings/settings.service";
/**
* Allocation of the portal NUMid — the "Security Number" my.jorgecuadros.com
* asks for at login.
*
* The NUMid is not a column on `Customer`. It is a `CustomerLegacyRef` row with
* (sourceSystem='utilities', sourceTable='DATGRAL'), and `CustomersService.create`
* deliberately writes none: a natively created customer has no legacy provenance.
* The consequence is that every customer created in the staff UI is invisible to
* the portal until this service gives them an id.
*
* WHY THIS IS NOT DONE AT CREATE TIME. Insurance is expected to move to the
* platform before utilities, and an insurance-only customer has no reason to hold
* a portal identity. Allocating on every create would spend utilities ids — and
* the handful of reusable ones — on people who will never log in. So this is an
* explicit staff action instead.
*/
/** The pair that identifies a portal NUMid. */
export const UTILITIES_SYSTEM = "utilities";
export const UTILITIES_TABLE = "DATGRAL";
/**
* insurance/DATGRAL is a SEPARATE id space that reuses the same sourceTable name
* and runs past 4,000. It must never be read as a NUMid, and never allocated
* from: the portal cannot resolve those ids. Every query here filters on BOTH
* columns for that reason, never on sourceTable alone. A customer can also hold
* more than one insurance ref — 16 of them do, where several insurance rows
* folded into one customer — so those are tested with EXISTS rather than joined.
*/
const POOL = {
sourceSystem: UTILITIES_SYSTEM,
sourceTable: UTILITIES_TABLE,
} as const;
export type AllocationOrigin = "existing" | "new" | "recycled";
export interface Allocation {
numid: string;
origin: AllocationOrigin;
/** Set only on a recycle — the customer the id was taken from. */
previousCustomerId?: string;
}
/**
* NUMids that were created and never used, safe for an allocator to take.
*
* THE TWO OBVIOUS RULES BOTH FIND NOTHING, which is why this one looks the way
* it does. "Owns no rows" matches nobody: migration gave all 1,171 NUMids a
* property and a transaction. "No transaction in N years" also matches nobody:
* every customer carries a synthetic Jan-1 opening-balance row, so everyone
* looks active in the current year. That row has to be subtracted before any
* activity test means anything, which is what `bf` does below.
*
* The balance-forward row is matched in two shapes on purpose.
* transform_transactions.py:120 mints a type literally named 'BALANCE FORWARD';
* databases loaded before that change carry the same rows with typeId NULL,
* dated Jan 1, legacySourceTable='datos2'. Matching only the type name floors
* nothing on such a database and turns the balance test into a raw lifetime sum
* — the double-count that read the whole book as +20.6M MXN in credit before
* d173c9e, and which here would mark live customers as empty.
*
* Services are tested as "any service" rather than "any ACTIVE service": a
* deactivated water account is still a record of somebody having lived behind
* this id.
*
* Kept in step with scripts/numid-audit.sql, which reports the same tier for a
* human. That script is the reporting copy of this rule; change both together.
*/
const EMPTY_NUMID_SQL = Prisma.sql`
WITH bf AS (
SELECT t.id, t.customerId
FROM transactions t
LEFT JOIN type_transactions tt ON tt.id = t.typeId
WHERE t.voidedAt IS NULL
AND (
tt.nameEn = 'BALANCE FORWARD'
OR (t.typeId IS NULL AND MONTH(t.transactionDate) = 1 AND DAY(t.transactionDate) = 1
AND t.legacySourceTable = 'datos2')
)
)
SELECT r.legacyId AS numid, r.id AS refId, r.customerId AS customerId
FROM customer_legacy_refs r
JOIN customers c ON c.id = r.customerId
WHERE r.sourceSystem = ${UTILITIES_SYSTEM} AND r.sourceTable = ${UTILITIES_TABLE}
AND (c.email IS NULL OR c.email = '')
AND NOT EXISTS (SELECT 1 FROM transactions t
WHERE t.customerId = c.id AND t.voidedAt IS NULL
AND t.id NOT IN (SELECT id FROM bf))
AND NOT EXISTS (SELECT 1 FROM transactions t
WHERE t.customerId = c.id AND t.voidedAt IS NULL AND t.outstanding = 1)
AND NOT EXISTS (SELECT 1 FROM property_services ps
JOIN properties p ON p.id = ps.propertyId WHERE p.customerId = c.id)
AND NOT EXISTS (SELECT 1 FROM policies p WHERE p.customerId = c.id)
AND NOT EXISTS (SELECT 1 FROM vehicles v WHERE v.customerId = c.id)
AND NOT EXISTS (SELECT 1 FROM trust_accounts ta
JOIN properties p ON p.id = ta.propertyId WHERE p.customerId = c.id)
AND NOT EXISTS (SELECT 1 FROM statement_documents s WHERE s.matchedCustomerId = c.id)
AND NOT EXISTS (SELECT 1 FROM policy_ocr_documents o WHERE o.matchedCustomerId = c.id)
AND NOT EXISTS (SELECT 1 FROM email_notification_log e WHERE e.customerId = c.id)
AND NOT EXISTS (SELECT 1 FROM email_log e WHERE e.customerId = c.id)
AND NOT EXISTS (SELECT 1 FROM account_status_history a WHERE a.customerId = c.id)
AND NOT EXISTS (SELECT 1 FROM customer_legacy_refs i
WHERE i.customerId = c.id AND i.sourceSystem = 'insurance')
ORDER BY CAST(r.legacyId AS UNSIGNED)`;
interface EmptyRow {
numid: string;
refId: string;
customerId: string;
}
@Injectable()
export class NumidService {
private readonly logger = new Logger(NumidService.name);
constructor(
private readonly prisma: PrismaService,
private readonly settings: SettingsService,
) {}
/** The customer's portal id, or null if they have none. */
async current(customerId: string): Promise<string | null> {
const ref = await this.prisma.customerLegacyRef.findFirst({
where: { customerId, ...POOL },
select: { legacyId: true },
});
return ref?.legacyId ?? null;
}
/** Reusable ids, lowest first. Empty unless recycling is switched on. */
async emptyCandidates(): Promise<string[]> {
const rows = await this.prisma.$queryRaw<EmptyRow[]>(EMPTY_NUMID_SQL);
return rows.map((r) => r.numid);
}
/**
* Give a customer a portal NUMid.
*
* Idempotent: a customer who already holds one gets it back rather than a
* second id, so a double-clicked button cannot fork an identity.
*/
async allocate(customerId: string): Promise<Allocation> {
const customer = await this.prisma.customer.findUnique({
where: { id: customerId },
select: { id: true, archivedAt: true },
});
if (!customer) throw new NotFoundException(`Customer ${customerId} not found`);
if (customer.archivedAt) {
throw new ConflictException(
"No se puede asignar un número de portal a un cliente archivado",
);
}
const existing = await this.current(customerId);
if (existing) return { numid: existing, origin: "existing" };
const recycle = await this.recycleEnabled();
// Two writers can pick the same id between the read and the write. The
// unique key on (sourceSystem, sourceTable, legacyId) is what actually
// decides the winner; the loser retries and takes the next id rather than
// silently overwriting. Bounded so a genuinely wedged pool fails loudly.
for (let attempt = 0; attempt < 5; attempt++) {
try {
if (recycle) {
const recycled = await this.tryRecycle(customerId);
if (recycled) return recycled;
}
return await this.allocateNext(customerId);
} catch (error) {
if (!isUniqueViolation(error)) throw error;
this.logger.warn(
`NUMid allocation for ${customerId} lost a race (attempt ${attempt + 1}), retrying`,
);
}
}
throw new ConflictException(
"No se pudo asignar un número de portal; intente de nuevo",
);
}
/**
* Whether the recycle tier is live.
*
* Off by default, and that default is the safe one while Access is still the
* utilities master. Every id in the pool ALSO exists in Access DATGRAL, and a
* `--sync` migration run upserts refs with ON DUPLICATE KEY UPDATE customerId
* (transform_customers.py:327) — so an id recycled today is silently handed
* back to its Access owner on the next sync, and the customer who was given it
* loses their portal identity. Turn this on once utilities has cut over, or
* for ids that have been deleted at the source.
*/
private async recycleEnabled(): Promise<boolean> {
const { value } = await this.settings.numidRecycleEmpty();
return value;
}
/** Re-point the lowest empty id at this customer. Null when none is free. */
private async tryRecycle(customerId: string): Promise<Allocation | null> {
const rows = await this.prisma.$queryRaw<EmptyRow[]>(EMPTY_NUMID_SQL);
for (const row of rows) {
// Guarded by the owner we just read: if anything moved the ref in the
// meantime the update matches nothing and we fall through to the next
// candidate rather than stealing an id that is no longer empty.
const moved = await this.prisma.customerLegacyRef.updateMany({
where: { id: row.refId, customerId: row.customerId },
data: { customerId },
});
if (moved.count === 1) {
this.logger.log(
`NUMid ${row.numid} recycled from ${row.customerId} to ${customerId}`,
);
return {
numid: row.numid,
origin: "recycled",
previousCustomerId: row.customerId,
};
}
}
return null;
}
/** One past the highest id in the pool. */
private async allocateNext(customerId: string): Promise<Allocation> {
const [{ max }] = await this.prisma.$queryRaw<{ max: number | null }[]>(
// MAX over a CAST, not over the string: legacyId is VARCHAR, so a plain
// MAX returns '999' as the highest of 1,171 rows and the allocator hands
// out an id that is already taken.
Prisma.sql`SELECT MAX(CAST(legacyId AS UNSIGNED)) AS max
FROM customer_legacy_refs
WHERE sourceSystem = ${UTILITIES_SYSTEM} AND sourceTable = ${UTILITIES_TABLE}`,
);
const numid = String(Number(max ?? 0) + 1);
await this.prisma.customerLegacyRef.create({
data: { customerId, ...POOL, legacyId: numid },
});
return { numid, origin: "new" };
}
}
function isUniqueViolation(error: unknown): boolean {
return (
error instanceof Prisma.PrismaClientKnownRequestError && error.code === "P2002"
);
}
@@ -22,6 +22,7 @@ export class UpdateCustomerDto {
@IsOptional() @IsString() mobile?: string; @IsOptional() @IsString() mobile?: string;
@IsOptional() @IsString() fax?: string; @IsOptional() @IsString() fax?: string;
@IsOptional() @IsEmail() email?: string; @IsOptional() @IsEmail() email?: string;
@IsOptional() @IsBoolean() emailOptOut?: boolean;
@IsOptional() @IsString() notes?: string; @IsOptional() @IsString() notes?: string;
@IsOptional() @IsString() identificationType?: string; @IsOptional() @IsString() identificationType?: string;
@IsOptional() @IsString() identificationNumber?: string; @IsOptional() @IsString() identificationNumber?: string;
+13
View File
@@ -0,0 +1,13 @@
import { Global, Module } from "@nestjs/common";
import { MailService } from "./mail.service";
/** Global so any feature module can inject MailService without re-importing.
* Matches the StorageService pattern: env-driven, null when unconfigured,
* and never blocks API boot. Notifications use it; renewals reuse it.
* ConfigService comes from the global ConfigModule in AppModule. */
@Global()
@Module({
providers: [MailService],
exports: [MailService],
})
export class MailModule {}
+189
View File
@@ -0,0 +1,189 @@
import {
Injectable,
Logger,
ServiceUnavailableException,
} from "@nestjs/common";
import { ConfigService } from "@nestjs/config";
import {
SESv2Client,
SendEmailCommand,
SendEmailCommandInput,
SendEmailCommandOutput,
} from "@aws-sdk/client-sesv2";
/**
* Outbound mail transport. Amazon SES — the channel the office already uses
* for bulk notification, per docs/INSURANCE_FEATURES_SPEC.md §1.3 (the
* renewal-notice spec settled on SES for the same reason: established sender
* reputation, existing IAM, negligible incremental cost at our volume).
*
* Mirrors `StorageService` exactly: env-driven config, null client when
* unconfigured, `ServiceUnavailableException` on use, never blocks API boot.
* When the env vars are missing AND we're in dev/test we fall back to a
* console-logging transport so the NotificationsService can be exercised
* end-to-end without SES credentials — a missing mail setup in production
* still throws, so a real deployment can't accidentally no-op its sends.
*
* Env:
* SES_REGION — required when client is configured
* SES_ACCESS_KEY / SES_SECRET_KEY — required
* SES_FROM — verified sending identity (e.g. mail@jorgecuadros.com)
* SES_FROM_NAME — display name, optional
* SES_CONFIGURATION_SET — optional, for bounce/complaint event publishing
*/
export interface SendArgs {
to: string;
/** Optional display name; SES will not display it for "to" but we keep it on
* the log row so customer-facing audit reads naturally. */
toName?: string;
subject: string;
/** HTML body. The four notification jobs all produce HTML. */
html: string;
/** Optional override of the configured From; rare but useful for the
* trust-payment test mail to a different identity. */
from?: string;
fromName?: string;
/** Marker header kept on every send so a downstream mail-log search for
* "X-Tracking: 1" surfaces only this app's outbound traffic. The legacy
* PHP sendEmail() always set it; we keep the convention. */
xTracking?: string;
}
export interface SendResult {
/** SES MessageId (or our mock prefix in dev). Stored verbatim on the
* notification log row so a SES bounce/complaint webhook can be matched
* back to the exact send. */
messageId: string;
/** Truncated SES response payload (or empty in dev). 4k cap matches the
* notification log column width. */
response: string;
}
@Injectable()
export class MailService {
private readonly logger = new Logger(MailService.name);
private readonly client: SESv2Client | null;
private readonly fromAddress: string | null;
private readonly fromName: string;
private readonly configurationSet: string | undefined;
private readonly devMode: boolean;
constructor(config: ConfigService) {
const region = config.get<string>("SES_REGION");
const accessKeyId = config.get<string>("SES_ACCESS_KEY");
const secretAccessKey = config.get<string>("SES_SECRET_KEY");
this.fromAddress =
config.get<string>("SES_FROM") ??
config.get<string>("MAIL_FROM") ??
null;
this.fromName =
config.get<string>("SES_FROM_NAME") ??
config.get<string>("MAIL_FROM_NAME") ??
"Information Server";
this.configurationSet = config.get<string>("SES_CONFIGURATION_SET");
// Dev fallback: when nothing is configured, log sends to stdout instead
// of throwing. Lets the API boot in a fresh checkout and lets the
// notifications UI show "0 sent" meaningfully on `debug=1`. Production
// (NODE_ENV !== development) still requires real config.
this.devMode = process.env.NODE_ENV !== "production";
if (!region || !accessKeyId || !secretAccessKey || !this.fromAddress) {
if (!this.devMode) {
this.logger.warn(
"SES not configured (SES_REGION / SES_ACCESS_KEY / SES_SECRET_KEY / SES_FROM). " +
"Outbound mail will throw ServiceUnavailableException.",
);
}
this.client = null;
return;
}
this.client = new SESv2Client({
region,
credentials: { accessKeyId, secretAccessKey },
});
this.logger.log(
`SES mail client configured (region=${region}, from=${this.fromAddress}).`,
);
}
/** Whether the deployment has a real mail transport. Callers use this to
* refuse work up front — a mass-notification job that throws on its
* first send half-completes and the log is unrecoverable, so we fail
* fast at the controller. */
get available(): boolean {
return this.client !== null || this.devMode;
}
/** True when the underlying transport is the dev console-log fallback. */
get isDevFallback(): boolean {
return this.client === null && this.devMode;
}
private require(): SESv2Client {
if (!this.client) {
throw new ServiceUnavailableException(
"El envío de correo no está configurado.",
);
}
return this.client;
}
/**
* Send a single HTML email. The dev fallback logs to stdout and returns a
* synthetic `dev-<timestamp>` message id; the real transport talks to SES
* and returns the SES MessageId.
*
* Throws `ServiceUnavailableException` when no transport is configured and
* we are not in dev — the caller (NotificationsService) catches and records
* it on the log row so a failed sweep produces a coherent audit trail
* instead of an aborted one.
*/
async send(args: SendArgs): Promise<SendResult> {
const from = `${args.fromName ?? this.fromName} <${
args.from ?? this.fromAddress ?? ""
}>`.trim();
if (!this.client) {
if (!this.devMode) this.require();
const fakeId = `dev-${Date.now().toString(36)}-${Math.random()
.toString(36)
.slice(2, 8)}`;
this.logger.log(
`[dev-mail] to=${args.to} subject="${args.subject}" id=${fakeId} ` +
`len=${args.html.length}`,
);
return { messageId: fakeId, response: "" };
}
const input: SendEmailCommandInput = {
FromEmailAddress: from,
Destination: { ToAddresses: [args.to] },
Content: {
Simple: {
Subject: { Data: args.subject, Charset: "UTF-8" },
Body: { Html: { Data: args.html, Charset: "UTF-8" } },
},
},
...(this.configurationSet
? { ConfigurationSetName: this.configurationSet }
: {}),
...(args.xTracking
? {
EmailTags: [
{ Name: "X-Tracking", Value: args.xTracking },
],
}
: {}),
};
const out: SendEmailCommandOutput = await this.client.send(
new SendEmailCommand(input),
);
return {
messageId: out.MessageId ?? "",
response: JSON.stringify({ MessageId: out.MessageId ?? null }).slice(0, 4096),
};
}
}
+34 -2
View File
@@ -25,6 +25,26 @@ async function bootstrap() {
throw new Error("SESSION_SECRET must be set (see .env.example)"); throw new Error("SESSION_SECRET must be set (see .env.example)");
} }
// Whether the session cookie carries the Secure flag. This CANNOT simply
// follow NODE_ENV: express-session silently declines to send a Secure cookie
// over a plain-HTTP connection, so a production image served over http://ial
// issues no cookie at all. Login then returns 200 with a user, no session is
// established, every later request 403s, and the UI loops back to /login —
// which is exactly what happened on the first galactus deploy.
//
// Leave it ON wherever the app is reached over TLS. Turn it OFF only for a
// deployment that is HTTP but reached over an already-encrypted transport
// (the galactus install is Tailscale-only, so WireGuard encrypts the wire).
// Behind a TLS-terminating proxy, set trust proxy instead of turning this off.
// An EMPTY value counts as unset, not as "false". Compose interpolation turns
// an absent `${SESSION_COOKIE_SECURE:-}` into the empty string, so testing
// `!== undefined` here would silently drop the Secure flag on any deployment
// that merely passes the variable through without setting it.
const cookieSecureRaw = process.env.SESSION_COOKIE_SECURE;
const cookieSecure = cookieSecureRaw
? cookieSecureRaw === "true"
: process.env.NODE_ENV === "production";
app.use( app.use(
session({ session({
secret: sessionSecret, secret: sessionSecret,
@@ -32,7 +52,7 @@ async function bootstrap() {
saveUninitialized: false, saveUninitialized: false,
cookie: { cookie: {
httpOnly: true, httpOnly: true,
secure: process.env.NODE_ENV === "production", secure: cookieSecure,
maxAge: 1000 * 60 * 60 * 8, // 8-hour session, matches a staff workday maxAge: 1000 * 60 * 60 * 8, // 8-hour session, matches a staff workday
}, },
}) })
@@ -40,7 +60,19 @@ async function bootstrap() {
app.use(passport.initialize()); app.use(passport.initialize());
app.use(passport.session()); app.use(passport.session());
app.enableCors({ credentials: true, origin: process.env.WEB_ORIGIN ?? "http://localhost:3000" }); // The same deployment is reached under several origins — the office LAN IP,
// the tailnet name, the demo domain — and the browser derives the API origin
// from whichever one served the page (apps/web/src/lib/api.ts). So WEB_ORIGIN
// is a comma-separated LIST, not a single value. A request whose Origin is
// not listed gets no CORS headers and the credentialed fetch fails, so add an
// entry when a new way of reaching the app is introduced. Same-origin setups
// (web and API behind one proxy) never hit CORS at all.
const webOrigins = (process.env.WEB_ORIGIN ?? "http://localhost:3000")
.split(",")
.map((o) => o.trim())
.filter(Boolean);
app.enableCors({ credentials: true, origin: webOrigins });
const port = process.env.PORT ? Number(process.env.PORT) : 3001; const port = process.env.PORT ? Number(process.env.PORT) : 3001;
await app.listen(port); await app.listen(port);
@@ -0,0 +1,14 @@
import { Module } from "@nestjs/common";
import { NotificationLogService } from "./notification-log.service";
/**
* Just the log writer, so a feature that sends mail can record it without
* importing `NotificationsModule` (which carries the four bulk-job pipelines
* and their controller). Imported by `NotificationsModule` and
* `RenewalsModule`.
*/
@Module({
providers: [NotificationLogService],
exports: [NotificationLogService],
})
export class NotificationLogModule {}
@@ -0,0 +1,74 @@
import { Injectable } from "@nestjs/common";
import {
EmailNotificationServicio,
EmailNotificationStatus,
EmailNotificationType,
} from "@jorgecuadros/database";
import { PrismaService } from "../prisma/prisma.service";
import { AttemptStatus } from "./notification.types";
/**
* The single writer for `email_notification_log`.
*
* Extracted out of `NotificationsService` so the renewal sweep can write the
* same rows as the four bulk jobs without pulling that service (and its four
* job pipelines) into `RenewalsModule`. Every outbound email the platform
* sends goes through here, which is what makes /notificaciones' "Registro de
* envíos" complete rather than per-feature.
*/
export interface NotificationLogEntry {
notificationType: EmailNotificationType;
servicio: EmailNotificationServicio;
/** Defaults to now(). Pass it when the row must line up exactly with
* another record of the same send (the renewal sweep pins it to
* `RenewalNotice.sentAt`). */
sendDate?: Date;
/** Type-dependent discriminator — see the `level` doc on the Prisma model.
* 0/1 for ACCOUNT_STATUS, the generation for RENEWAL_NOTICE. */
level?: number | null;
customerId: string | null;
customerName: string;
customerEmail: string;
subject: string;
bodySnapshot: string;
bodyRequestUrl?: string;
status: AttemptStatus;
debug: boolean;
providerMessageId?: string;
providerResponse?: string;
error?: string;
}
/** `providerResponse` is a VARCHAR(191); anything longer is a provider dump
* we only need the head of. Errors go to the TEXT `error` column and get
* the 4k cap the schema documents. */
const PROVIDER_RESPONSE_MAX = 180;
const ERROR_MAX = 4096;
@Injectable()
export class NotificationLogService {
constructor(private readonly prisma: PrismaService) {}
async record(entry: NotificationLogEntry): Promise<void> {
await this.prisma.emailNotificationLog.create({
data: {
notificationType: entry.notificationType,
servicio: entry.servicio,
...(entry.sendDate && { sendDate: entry.sendDate }),
level: entry.level ?? null,
customerId: entry.customerId,
customerName: entry.customerName,
customerEmail: entry.customerEmail,
subject: entry.subject,
bodySnapshot: entry.bodySnapshot,
bodyRequestUrl: entry.bodyRequestUrl ?? null,
debug: entry.debug,
providerMessageId: entry.providerMessageId ?? null,
providerResponse:
entry.providerResponse?.slice(0, PROVIDER_RESPONSE_MAX) ?? null,
status: entry.status as EmailNotificationStatus,
error: entry.error?.slice(0, ERROR_MAX) ?? null,
},
});
}
}
@@ -0,0 +1,15 @@
import { Module } from "@nestjs/common";
import { SettingsModule } from "../settings/settings.module";
import { NotificationScheduleService } from "./notification-schedule.service";
/**
* Just the cadence registry, split out for the same reason as
* `NotificationLogModule`: both `NotificationsModule` and `RenewalsModule`
* need it, and neither may import the other.
*/
@Module({
imports: [SettingsModule],
providers: [NotificationScheduleService],
exports: [NotificationScheduleService],
})
export class NotificationScheduleModule {}
@@ -0,0 +1,203 @@
import { Injectable, Logger } from "@nestjs/common";
import { SchedulerRegistry } from "@nestjs/schedule";
import { CronJob } from "cron";
import { SettingsService } from "../settings/settings.service";
import type { ResolvedSetting } from "../settings/settings.service";
/**
* When the two automatic envíos run.
*
* Both halves of /notificaciones used to be hardcoded: pólizas swept at 06:00
* from a `@Cron` decorator, servicios had no automatic run at all and had to
* be clicked. Neither could be changed without a redeploy. This service owns
* the cadence for both, stores it in `app_settings`, and re-installs the job
* the moment an operator saves — no restart.
*
* The owning services register their handler at boot rather than this service
* importing them: `NotificationsService` and `RenewalsService` would otherwise
* have to be injected here, and this file is imported by both.
*/
export const SCHEDULE_TIME_ZONE = "America/Tijuana";
export type ScheduleKind = "servicios" | "polizas";
export const SCHEDULE_KINDS: ScheduleKind[] = ["servicios", "polizas"];
export interface NotificationSchedule {
enabled: boolean;
/** Local hour/minute in `SCHEDULE_TIME_ZONE`, not UTC — the office thinks
* in Tijuana time and DST would otherwise drift the run by an hour. */
hour: number;
minute: number;
/** 0 = Sunday … 6 = Saturday. Empty means every day. */
weekdays: number[];
}
export interface ResolvedSchedule extends ResolvedSetting<NotificationSchedule> {
/** The cron expression the value compiles to, shown in the UI so the
* operator can see exactly what was installed. */
cron: string;
/** Next fire time, or null when disabled. */
nextRun: string | null;
}
/**
* Defaults preserve what each half did before this existed: pólizas keeps its
* 06:00 daily sweep, servicios stays OFF. Turning a mass send on is an
* operator decision — a default that starts mailing 260 customers on its own
* after a deploy is not a default, it's an incident.
*/
const DEFAULTS: Record<ScheduleKind, NotificationSchedule> = {
servicios: { enabled: false, hour: 7, minute: 0, weekdays: [1, 3, 5] },
polizas: { enabled: true, hour: 6, minute: 0, weekdays: [] },
};
/** Human label used in log lines and audit entries. */
export const SCHEDULE_LABELS: Record<ScheduleKind, string> = {
servicios: "envíos de servicios",
polizas: "avisos de renovación",
};
export function scheduleCron(schedule: NotificationSchedule): string {
const dow = schedule.weekdays.length
? [...new Set(schedule.weekdays)].sort((a, b) => a - b).join(",")
: "*";
return `${schedule.minute} ${schedule.hour} * * ${dow}`;
}
/** Reject anything that would compile to a cron we can't install. Returns the
* normalized value, or a message naming the offending field. */
export function parseSchedule(
raw: unknown,
): { ok: true; value: NotificationSchedule } | { ok: false; error: string } {
const v = raw as Partial<NotificationSchedule> | null;
if (!v || typeof v !== "object") return { ok: false, error: "Horario inválido." };
const hour = Number(v.hour);
const minute = Number(v.minute);
if (!Number.isInteger(hour) || hour < 0 || hour > 23) {
return { ok: false, error: "La hora debe estar entre 0 y 23." };
}
if (!Number.isInteger(minute) || minute < 0 || minute > 59) {
return { ok: false, error: "Los minutos deben estar entre 0 y 59." };
}
const weekdays = Array.isArray(v.weekdays) ? v.weekdays.map(Number) : [];
if (weekdays.some((d) => !Number.isInteger(d) || d < 0 || d > 6)) {
return { ok: false, error: "Los días deben estar entre 0 (domingo) y 6." };
}
return {
ok: true,
value: {
enabled: !!v.enabled,
hour,
minute,
weekdays: [...new Set(weekdays)].sort((a, b) => a - b),
},
};
}
@Injectable()
export class NotificationScheduleService {
private readonly logger = new Logger(NotificationScheduleService.name);
private readonly handlers = new Map<ScheduleKind, () => Promise<unknown>>();
constructor(
private readonly settings: SettingsService,
private readonly registry: SchedulerRegistry,
) {}
/**
* Called once per kind at boot by the service that owns the sweep. Installs
* the job immediately so a freshly started process honours the stored
* cadence without waiting for someone to open the UI.
*/
async register(kind: ScheduleKind, handler: () => Promise<unknown>) {
this.handlers.set(kind, handler);
await this.apply(kind);
}
async get(kind: ScheduleKind): Promise<ResolvedSchedule> {
const resolved = await this.settings.notificationSchedule(
kind,
DEFAULTS[kind],
);
const cron = scheduleCron(resolved.value);
return { ...resolved, cron, nextRun: this.nextRun(kind) };
}
async getAll(): Promise<Record<ScheduleKind, ResolvedSchedule>> {
const entries = await Promise.all(
SCHEDULE_KINDS.map(async (k) => [k, await this.get(k)] as const),
);
return Object.fromEntries(entries) as Record<ScheduleKind, ResolvedSchedule>;
}
async set(
kind: ScheduleKind,
schedule: NotificationSchedule,
userId: string,
): Promise<ResolvedSchedule> {
await this.settings.setNotificationSchedule(kind, schedule, userId);
await this.apply(kind);
return this.get(kind);
}
/** (Re)install the cron job for one kind from whatever is stored now. */
private async apply(kind: ScheduleKind): Promise<void> {
const handler = this.handlers.get(kind);
if (!handler) return;
this.remove(kind);
const { value } = await this.settings.notificationSchedule(
kind,
DEFAULTS[kind],
);
if (!value.enabled) {
this.logger.log(`Horario de ${SCHEDULE_LABELS[kind]}: desactivado.`);
return;
}
const cron = scheduleCron(value);
const job = new CronJob(
cron,
() => {
void handler().catch((error) =>
this.logger.error(
`Falló la corrida programada de ${SCHEDULE_LABELS[kind]}: ` +
`${(error as Error).message}`,
),
);
},
null,
false,
SCHEDULE_TIME_ZONE,
);
this.registry.addCronJob(this.jobName(kind), job);
job.start();
this.logger.log(
`Horario de ${SCHEDULE_LABELS[kind]}: ${cron} (${SCHEDULE_TIME_ZONE}).`,
);
}
private remove(kind: ScheduleKind): void {
const name = this.jobName(kind);
// `deleteCronJob` throws when the job was never installed, which is the
// normal case on first apply — presence check instead of try/catch so a
// real failure still surfaces.
if (!this.registry.doesExist("cron", name)) return;
this.registry.getCronJob(name).stop();
this.registry.deleteCronJob(name);
}
private nextRun(kind: ScheduleKind): string | null {
const name = this.jobName(kind);
if (!this.registry.doesExist("cron", name)) return null;
const next = this.registry.getCronJob(name).nextDate();
return next ? next.toJSDate().toISOString() : null;
}
private jobName(kind: ScheduleKind): string {
return `notification-schedule:${kind}`;
}
}
@@ -0,0 +1,57 @@
import { parseSchedule, scheduleCron } from "./notification-schedule.service";
/**
* The cadence editor's only sharp edge: a stored value compiles to a cron
* expression that the scheduler installs verbatim. A malformed one either
* throws at install time (taking the sweep down) or silently installs the
* wrong cadence, so validation happens before anything is written.
*/
describe("scheduleCron", () => {
it("compiles a daily schedule with no weekday filter", () => {
expect(
scheduleCron({ enabled: true, hour: 6, minute: 0, weekdays: [] }),
).toBe("0 6 * * *");
});
it("compiles the legacy Mon/Wed/Fri cadence, sorted and de-duplicated", () => {
expect(
scheduleCron({ enabled: true, hour: 7, minute: 30, weekdays: [5, 1, 3, 1] }),
).toBe("30 7 * * 1,3,5");
});
});
describe("parseSchedule", () => {
it("normalizes weekdays and coerces enabled to a boolean", () => {
const parsed = parseSchedule({
enabled: 1,
hour: 6,
minute: 0,
weekdays: [3, 1, 3],
});
expect(parsed).toEqual({
ok: true,
value: { enabled: true, hour: 6, minute: 0, weekdays: [1, 3] },
});
});
it("defaults a missing weekday list to every day", () => {
const parsed = parseSchedule({ enabled: true, hour: 0, minute: 0 });
expect(parsed.ok && parsed.value.weekdays).toEqual([]);
});
it.each([
[{ enabled: true, hour: 24, minute: 0 }, "hora"],
[{ enabled: true, hour: 6, minute: 60 }, "minutos"],
[{ enabled: true, hour: 6, minute: 0, weekdays: [7] }, "días"],
[{ enabled: true, hour: 6.5, minute: 0 }, "hora"],
])("rejects %p", (input, field) => {
const parsed = parseSchedule(input);
expect(parsed.ok).toBe(false);
expect(!parsed.ok && parsed.error.toLowerCase()).toContain(field);
});
it("rejects a non-object", () => {
expect(parseSchedule(null).ok).toBe(false);
});
});
@@ -0,0 +1,175 @@
import {
EmailNotificationServicio,
EmailNotificationType,
} from "@jorgecuadros/database";
import { IsBoolean, IsEnum, IsOptional } from "class-validator";
/**
* Where `debug` sends everything. The PHP used `rmancinas@freakma.net`;
* same here. Exported because the flag is platform-wide — the renewal
* notices honour it too, and two copies of this address would eventually
* disagree.
*/
export const DEBUG_RECIPIENT = "rmancinas@freakma.net";
/**
* Shared flags for every notification send — the four servicios jobs and
* the pólizas renewal notices alike. Every endpoint takes the same shape
* so the UI can offer one set of switches for the whole screen; each flag
* is documented inline so the per-job semantics are obvious in one place.
*
* `debug` — replace every recipient with `DEBUG_RECIPIENT` so a
* real customer never receives mail during a test run.
* Logged on every row. On the renewal side a debug send
* also does NOT write the `RenewalNotice` row, so a test
* can't gate the letter the customer is still owed.
* `ignoreDayRestriction` — Job 3 only: bypass the Mon/Wed/Fri (red) and
* Wed-only (yellow) day gates. Off by default so
* the on-demand sweep behaves like the legacy
* script.
* `useEmailLimit` — Job 3 only: pause the sweep 1 hour after 100
* sends (a vestigial SMTP-era throttling limit).
* Off by default; SES does not need it.
*/
export class NotificationFlagsDto {
@IsOptional()
@IsBoolean()
debug?: boolean;
@IsOptional()
@IsBoolean()
ignoreDayRestriction?: boolean;
@IsOptional()
@IsBoolean()
useEmailLimit?: boolean;
}
/**
* What we know at job-end and put on the wire. Field names match the
* legacy PHP scripts' `echo json_encode(...)` so a downstream log scraper
* that already parses `notificationType: "sendPaymentConfirmation"`
* keeps working — see `~/Documents/Claude-Memory/email-notifications-spec.md`
* for the verbatim PHP shapes. Specifically: Job 1 reports
* `notificationType: "sendPaymentConfirmation"` (the legacy literal), and
* uses field `result` instead of `request`; the other three use
* `notificationType` matching the script's purpose.
*
* Every variant carries `sent/skipped/failed/debug` for the audit log;
* the legacy fields stay where they were so the response shape is
* exactly backward-compatible.
*/
export type NotificationJobResponse =
| {
// Job 1
result: "success";
notificationType: "sendPaymentConfirmation";
reason: string;
statusCode: 200;
sent: number;
skipped: number;
failed: number;
debug: boolean;
type: "OUTSTANDING_PAYMENT";
}
| {
// Job 2
request: "success";
notificationType: "sendPaymentConfirmation";
confirmationSent: string;
statusCode: 200;
sent: number;
skipped: number;
failed: number;
debug: boolean;
type: "PAYMENT_CONFIRMATION";
}
| {
// Job 3 — sent/skipped/failed included so the audit log can record
// totals without depending on (red+yellow) alone.
request: "success";
notificationType: "sendAccountStatus";
statusSent: string;
statusReport: string;
statusCode: 200;
red: number;
yellow: number;
total: number;
sent: number;
skipped: number;
failed: number;
debug: boolean;
type: "ACCOUNT_STATUS";
}
| {
// Job 4
request: "success";
notificationType: "sendTrustPaymentConfirmation";
confirmationSent: string;
statusCode: 200;
sent: number;
skipped: number;
failed: number;
debug: boolean;
type: "TRUST_PAYMENT_CONFIRMATION";
};
/** The four jobs, in the order the "ejecutar todos" sweep runs them. */
export type NotificationJobKind =
| "outstanding"
| "payment"
| "account"
| "trust";
/**
* One entry of the run-all sweep. A job that throws does NOT abort the
* sweep — it is recorded with `ok: false` and the next job still runs, so a
* single bad query can't silently block the other three envíos.
*/
export interface NotificationRunAllJobResult {
kind: NotificationJobKind;
ok: boolean;
result?: NotificationJobResponse;
error?: string;
}
/**
* Aggregate response for `POST /notifications/run-all`. `sent/skipped/failed`
* are the sums across every job that completed; `jobs` keeps each job's own
* verbatim legacy response so the UI can still show per-job detail.
*/
export interface NotificationRunAllResponse {
request: "success";
notificationType: "runAllNotifications";
statusCode: 200;
debug: boolean;
sent: number;
skipped: number;
failed: number;
/** Jobs that threw — sweep continued past them. */
errors: number;
jobs: NotificationRunAllJobResult[];
type: "RUN_ALL";
}
/** Normalized record for a single send attempt, fed by all four jobs. */
export interface SendAttempt {
notificationType: EmailNotificationType;
servicio: EmailNotificationServicio;
customerId: string | null;
customerName: string;
customerEmail: string;
subject: string;
bodySnapshot: string;
bodyRequestUrl?: string;
/** Account-status-only — 0 yellow / 1 red. Null on the other three jobs. */
level?: 0 | 1;
/** Account-status-only — DEBAJO DEL TIPO / EN ROJO. */
historyTipo?: string;
historyBalance?: string;
historyTCambio?: string;
historySolicitado?: string;
}
/** Status enum values, mirrored from `EmailNotificationStatus`. */
export type AttemptStatus = "SENT" | "FAILED" | "SKIPPED_NO_EMAIL" | "SKIPPED_GATE";
@@ -0,0 +1,331 @@
import {
BadRequestException,
Body,
Controller,
Get,
HttpCode,
Param,
Post,
Put,
Query,
Req,
UseGuards,
} from "@nestjs/common";
import { Request } from "express";
import {
EmailNotificationServicio,
EmailNotificationStatus,
EmailNotificationType,
} from "@jorgecuadros/database";
import { Transform, Type } from "class-transformer";
import {
ArrayMaxSize,
IsArray,
IsBoolean,
IsEnum,
IsInt,
IsOptional,
IsString,
Max,
Min,
} from "class-validator";
import { AuthenticatedGuard } from "../auth/authenticated.guard";
import { AbilityGuard } from "../auth/ability.guard";
import { RequireAbility } from "../auth/require-ability.decorator";
import { AuditService } from "../common/audit.service";
import { invalidEmails, SettingsService } from "../settings/settings.service";
import {
NotificationScheduleService,
parseSchedule,
SCHEDULE_KINDS,
ScheduleKind,
} from "./notification-schedule.service";
import { NotificationFlagsDto } from "./notification.types";
import { NotificationsService } from "./notifications.service";
/** Same flags for every job, query-string OR body (the PHP scripts took
* both via STDIN vs HTTP-CGI — we accept either for parity). */
class RunJobDto extends NotificationFlagsDto {}
class ListLogDto {
@IsOptional() @Type(() => Number) @IsInt() @Min(1) page?: number;
@IsOptional() @Type(() => Number) @IsInt() @Min(1) @Max(200) pageSize?: number;
@IsOptional() @IsEnum(EmailNotificationType) type?: EmailNotificationType;
/** One or more servicios, comma-separated. The /notificaciones tabs each
* read their own slice of the one log: Servicios passes
* `CUSTOMERS,TRUST`, Pólizas passes `POLICIES`. Omitted = every servicio. */
@IsOptional()
@Transform(({ value }) =>
typeof value === "string"
? value.split(",").map((s) => s.trim()).filter(Boolean)
: value,
)
@IsEnum(EmailNotificationServicio, { each: true })
servicio?: EmailNotificationServicio[];
@IsOptional() @IsEnum(EmailNotificationStatus) status?: EmailNotificationStatus;
@IsOptional() @IsEnum(["sent", "failed", "skipped", "all"]) view?: "sent" | "failed" | "skipped" | "all";
}
/** An empty array is valid and means "send no summaries" — the cap only
* exists so a paste accident can't write an unbounded blob. */
class AdminEmailsDto {
@IsArray()
@ArrayMaxSize(50)
@IsString({ each: true })
emails!: string[];
}
/** Cadence of one automatic envío. Ranges are re-checked by `parseSchedule`,
* which is also what the scheduler itself uses — the decorators here only
* reject wrong *types* so a bad payload fails at the edge. */
class ScheduleDto {
@IsBoolean() enabled!: boolean;
@IsInt() @Min(0) @Max(23) hour!: number;
@IsInt() @Min(0) @Max(59) minute!: number;
@IsOptional() @IsArray() @IsInt({ each: true }) weekdays?: number[];
}
function actingId(req: Request): string {
return (req.user as { id: string }).id;
}
/**
* HTTP surface for the mass-notification jobs. Four trigger endpoints +
* two read endpoints (list log, stats). All mutations gated by the
* `notification:send` ability so a STAFF user can't accidentally fire a
* 260-mail sweep.
*/
@UseGuards(AuthenticatedGuard, AbilityGuard)
@Controller("notifications")
export class NotificationsController {
constructor(
private readonly svc: NotificationsService,
private readonly audit: AuditService,
private readonly settings: SettingsService,
private readonly schedule: NotificationScheduleService,
) {}
/* -------------------------------------------------------------- triggers */
@Post("outstanding-payments")
@RequireAbility("notification:send")
@HttpCode(200)
async runOutstanding(
@Body() body: RunJobDto,
@Query() query: RunJobDto,
@Req() req: Request,
) {
const flags = { ...query, ...body };
const result = await this.svc.runOutstandingPayments(flags);
void this.audit.log(actingId(req), "notification.outstanding.run", {
debug: !!flags.debug,
sent: result.sent,
skipped: result.skipped,
failed: result.failed,
});
return result;
}
@Post("payment-confirmation")
@RequireAbility("notification:send")
@HttpCode(200)
async runPaymentConfirm(
@Body() body: RunJobDto,
@Query() query: RunJobDto,
@Req() req: Request,
) {
const flags = { ...query, ...body };
const result = await this.svc.runPaymentConfirmation(flags);
void this.audit.log(actingId(req), "notification.payment-confirm.run", {
debug: !!flags.debug,
sent: result.sent,
skipped: result.skipped,
failed: result.failed,
});
return result;
}
@Post("account-status")
@RequireAbility("notification:send")
@HttpCode(200)
async runAccountStatus(
@Body() body: RunJobDto,
@Query() query: RunJobDto,
@Req() req: Request,
) {
const flags = { ...query, ...body };
const result = await this.svc.runAccountStatus(flags);
// Narrow the discriminated union to the ACCOUNT_STATUS variant before
// pulling red/yellow/total — TS can't follow this through `await` alone.
if (result.type === "ACCOUNT_STATUS") {
void this.audit.log(actingId(req), "notification.account-status.run", {
debug: !!flags.debug,
red: result.red,
yellow: result.yellow,
total: result.total,
sent: result.sent,
skipped: result.skipped,
failed: result.failed,
});
}
return result;
}
@Post("trust-payment-confirmation")
@RequireAbility("notification:send")
@HttpCode(200)
async runTrustConfirm(
@Body() body: RunJobDto,
@Query() query: RunJobDto,
@Req() req: Request,
) {
const flags = { ...query, ...body };
const result = await this.svc.runTrustConfirmation(flags);
void this.audit.log(actingId(req), "notification.trust-confirm.run", {
debug: !!flags.debug,
sent: result.sent,
skipped: result.skipped,
failed: result.failed,
});
return result;
}
/**
* Run all four jobs sequentially with one set of flags. Audited as a
* single `notification.run-all.run` entry carrying the aggregate totals
* plus each job's outcome — the per-job endpoints are NOT re-audited, so
* the log has exactly one row per staff click.
*/
@Post("run-all")
@RequireAbility("notification:send")
@HttpCode(200)
async runAll(
@Body() body: RunJobDto,
@Query() query: RunJobDto,
@Req() req: Request,
) {
const flags = { ...query, ...body };
const result = await this.svc.runAll(flags);
void this.audit.log(actingId(req), "notification.run-all.run", {
debug: !!flags.debug,
ignoreDayRestriction: !!flags.ignoreDayRestriction,
useEmailLimit: !!flags.useEmailLimit,
sent: result.sent,
skipped: result.skipped,
failed: result.failed,
errors: result.errors,
jobs: result.jobs.map((j) => ({ kind: j.kind, ok: j.ok })),
});
return result;
}
/* ----------------------------------------------------------- read views */
@Get("log")
listLog(@Query() q: ListLogDto) {
const page = q.page ?? 1;
const pageSize = q.pageSize ?? 50;
return this.svc.listLog({
page,
pageSize,
type: q.type,
servicio: q.servicio,
status: this.mapViewStatus(q.view, q.status),
customerId: undefined,
});
}
@Get("stats")
stats(@Query() q: ListLogDto) {
return this.svc.stats(q.servicio);
}
/* -------------------------------------------------------------- settings */
/** Who receives the per-job summary email. Readable by any logged-in user
* so the UI can show the current list; editing needs `setting:manage`. */
@Get("settings/admin-emails")
adminEmails() {
return this.settings.notificationAdminEmails();
}
@Put("settings/admin-emails")
@RequireAbility("setting:manage")
async setAdminEmails(@Body() dto: AdminEmailsDto, @Req() req: Request) {
const emails = dto.emails.map((e) => e.trim()).filter(Boolean);
const bad = invalidEmails(emails);
if (bad.length) {
throw new BadRequestException(
`Correo inválido: ${bad.join(", ")}`,
);
}
const result = await this.settings.setNotificationAdminEmails(
emails,
actingId(req),
);
void this.audit.log(actingId(req), "notification.settings.admin-emails", {
emails,
});
return result;
}
/* -------------------------------------------------------------- schedule */
/**
* Cadence of both automatic envíos. Readable by any logged-in user so the
* screen can show "próxima corrida" without needing edit rights; changing
* it needs `setting:manage`, same as the summary recipients.
*/
@Get("settings/schedule")
schedules() {
return this.schedule.getAll();
}
@Put("settings/schedule/:kind")
@RequireAbility("setting:manage")
async setSchedule(
@Param("kind") kind: string,
@Body() dto: ScheduleDto,
@Req() req: Request,
) {
if (!SCHEDULE_KINDS.includes(kind as ScheduleKind)) {
throw new BadRequestException(
`Horario desconocido: ${kind}. Use ${SCHEDULE_KINDS.join(" o ")}.`,
);
}
const parsed = parseSchedule({ ...dto, weekdays: dto.weekdays ?? [] });
if (!parsed.ok) throw new BadRequestException(parsed.error);
const result = await this.schedule.set(
kind as ScheduleKind,
parsed.value,
actingId(req),
);
void this.audit.log(actingId(req), "notification.settings.schedule", {
kind,
...parsed.value,
cron: result.cron,
});
return result;
}
/** Resolve the UI's coarse view tabs to concrete statuses. An explicit
* `status` wins. "Omitidos" covers both SKIPPED_* variants, which is why
* this returns a list rather than a single value. */
private mapViewStatus(
view: ListLogDto["view"],
status: ListLogDto["status"],
): EmailNotificationStatus[] | undefined {
if (status) return [status];
if (!view || view === "all") return undefined;
if (view === "sent") return [EmailNotificationStatus.SENT];
if (view === "failed") return [EmailNotificationStatus.FAILED];
if (view === "skipped") {
return [
EmailNotificationStatus.SKIPPED_NO_EMAIL,
EmailNotificationStatus.SKIPPED_GATE,
];
}
return undefined;
}
}
@@ -0,0 +1,22 @@
import { Module } from "@nestjs/common";
import { NotificationLogModule } from "./notification-log.module";
import { NotificationScheduleModule } from "./notification-schedule.module";
import { SettingsModule } from "../settings/settings.module";
import { NotificationsController } from "./notifications.controller";
import { NotificationsService } from "./notifications.service";
/**
* Mass email notifications. MailModule is global (registered in AppModule),
* so this module needs no MailService import — it picks it up by injection.
*
* The automatic sweep is registered by `NotificationsService` against
* `NotificationScheduleService`, which owns the cadence for both halves of
* /notificaciones and stores it in `app_settings`.
*/
@Module({
imports: [NotificationLogModule, NotificationScheduleModule, SettingsModule],
controllers: [NotificationsController],
providers: [NotificationsService],
exports: [NotificationsService],
})
export class NotificationsModule {}
File diff suppressed because it is too large Load Diff
+125
View File
@@ -0,0 +1,125 @@
import {
renderAccountStatus,
renderOutstanding,
renderPaymentConfirm,
renderTrustConfirm,
} from "./render";
/**
* Render-level tests. The legacy PHP scripts fetched these bodies by URL;
* we render server-side and inline. The tests assert the *shape* of each
* body — account id, name, subject, balance/tipo, color band — because
* the customer base has been seeing these letters for years and a visual
* regression costs trust faster than any backend change does.
*/
describe("renderOutstanding", () => {
it("includes the customer id, name, total, and per-row table", () => {
const html = renderOutstanding({
customerId: "C-001",
customerName: "Acme & Co.",
total: "1234.50",
rows: [
{
date: "2026-07-01",
reference: "INV-1",
period: "Jul-26",
type: "CHECK",
amount: "-500.00",
balance: "-500.00",
},
{
date: "2026-07-15",
reference: "INV-2",
period: "Jul-26",
type: "CASH",
amount: "-734.50",
balance: "-1234.50",
},
],
year: 2026,
});
expect(html).toContain("Acme &amp; Co.");
expect(html).toContain("ACCOUNT #C-001");
expect(html).toContain("$ 1,234.50");
expect(html).toContain("INV-1");
expect(html).toContain("CHECK");
expect(html).toContain("IF YOU ALREADY SENT THE CHECK");
});
it("escapes HTML in the customer name", () => {
const html = renderOutstanding({
customerId: "x",
customerName: "<script>alert(1)</script>",
total: "0.00",
rows: [],
year: 2026,
});
expect(html).not.toContain("<script>alert(1)</script>");
expect(html).toContain("&lt;script&gt;alert(1)&lt;/script&gt;");
});
});
describe("renderPaymentConfirm", () => {
it("uses the transaction type in the heading and the amount in the body", () => {
const html = renderPaymentConfirm({
customerId: "C-002",
customerName: "Bob",
typeOfTrx: "CHECK DEPOSIT",
reference: "DEP-99",
amount: "500.00",
year: 2026,
});
expect(html).toContain("CHECK DEPOSIT CONFIRMATION");
expect(html).toContain("HI, Bob");
expect(html).toContain("REFER# DEP-99");
expect(html).toContain("$ 500.00");
});
});
describe("renderAccountStatus", () => {
it("uses the yellow band and the under-minimum phrasing for level=0", () => {
const html = renderAccountStatus({
customerId: "C-003",
customerName: "Carol",
level: 0,
balance: "10.00",
tipo: "40.00",
year: 2026,
});
expect(html).toContain("#88D5EE");
expect(html).toContain("under our minimum");
expect(html).toContain("Carol");
expect(html).toContain("$ 10.00");
expect(html).toContain("$ 40.00");
});
it("uses the red band and the rush phrasing for level=1", () => {
const html = renderAccountStatus({
customerId: "C-003",
customerName: "Carol",
level: 1,
balance: "-25.50",
tipo: "25.50",
year: 2026,
});
expect(html).toContain("#FF8D71");
expect(html).toContain("overdrawn");
expect(html).toContain("reactivate your payments");
expect(html).toContain("$ 25.50");
});
});
describe("renderTrustConfirm", () => {
it("labels the trust annual fee and quotes the amount", () => {
const html = renderTrustConfirm({
customerId: "C-004",
customerName: "Dan",
amount: "350.00",
year: 2026,
});
expect(html).toContain("Annual Bank Fee Payment Confirmation");
expect(html).toContain("$ 350.00");
expect(html).toContain("Most banks always request");
});
});
+254
View File
@@ -0,0 +1,254 @@
/**
* HTML body renderers for the four notification jobs. These are the modern
* in-process equivalent of the legacy `getXxxForEmail.php` files the PHP
* scripts `fetch()`ed by URL. Rendering server-side and inlining the body
* in the response keeps a single SES MessageId tied to one frozen HTML
* snapshot (vs. the legacy flow, where the URL kept re-rendering with
* whatever the database looked like at click time).
*
* The visual style mirrors the legacy PHP templates where it makes sense
* (the office's customer base has been seeing these letters for years;
* gratuitous redesign costs trust). The body shell, table layout and the
* canonical contact block are preserved verbatim. English copy because the
* legacy letters were English; switching to Spanish is a future decision
* (see INSURANCE_FEATURES_SPEC §1.6 "Spanish or English body?").
*/
const HEAD = `<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Transitional//EN" "http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd">
<html xmlns="http://www.w3.org/1999/xhtml">
<head>
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
<title>{title}</title>
</head>`;
const FOOT_CONTACT = `<p>If you have any questions regarding this notice please contact us at:
Tel. 011 52 (661) 612 - 1295 &nbsp; Fax. (661) 612 - 1285 &nbsp;
For any type of a 24 Hrs. emergencies: please dial 52 (664) 304 - 7778 |
<a href="mailto:jorge@jorgecuadros.com">jorge@jorgecuadros.com</a> |
<a href="https://www.jorgecuadros.com/contactus.php">Contact Us Form</a></p>`;
const SIGNED = (year: number) => `<center><span class="small">This message has been generated by the Jorge Cuadros &amp; Assoc. Information Server.<br />Copyright ${year}&nbsp;<a href="http://www.freakma.net/">Developed by FreaKmA.Net</a></span></center>`;
const esc = (s: string | null | undefined): string =>
String(s ?? "")
.replace(/&/g, "&amp;")
.replace(/</g, "&lt;")
.replace(/>/g, "&gt;")
.replace(/"/g, "&quot;");
const usd = (n: number | string | null | undefined): string => {
if (n === null || n === undefined) return "$ 0.00";
const v = typeof n === "string" ? Number(n) : n;
if (!isFinite(v)) return "$ 0.00";
return `$ ${v.toLocaleString("en-US", {
minimumFractionDigits: 2,
maximumFractionDigits: 2,
})}`;
};
/** Shared shell: a 2-column table that matches the PHP output layout. */
function shell(opts: {
title: string;
bg: string;
heading: string;
accountId: string | number;
accountName: string;
body: string;
note?: string;
statementLink?: string;
year: number;
}): string {
const { title, bg, heading, accountId, accountName, body, note, statementLink, year } = opts;
const stmt = statementLink ?? "https://my.jorgecuadros.com/";
return `${HEAD.replace("{title}", esc(title))}
<body style="background-color:${bg};color:#333;font-family:'Courier New', Courier, monospace;">
<table width="100%" border="0" cellspacing="0" cellpadding="0">
<tr>
<td width="43%" style="font-size:20px;font-weight:bold;">${esc(heading)}</td>
<td width="57%" style="font-size:12px;">Please do not reply to this message. For any Jorge Cuadros &amp; Assoc. customer service inquiries, visit: <a href="https://www.jorgecuadros.com/contactus.php">Customer Support</a></td>
</tr>
<tr>
<td><strong>${esc(accountName)}<br />ACCOUNT #${esc(String(accountId))}</strong></td>
<td><div align="center"><a href="${esc(stmt)}" target="_blank" style="color:#006699;font-weight:bold">Click Here to View Your Account Statement</a></div></td>
</tr>
<tr><td colspan="2">&nbsp;</td></tr>
<tr><td colspan="2">${body}</td></tr>
<tr><td colspan="2">&nbsp;</td></tr>
${
note
? `<tr><td colspan="2"><h4>${esc(note)}</h4>${FOOT_CONTACT}</td></tr>`
: `<tr><td colspan="2">${FOOT_CONTACT}</td></tr>`
}
<tr><td colspan="2">&nbsp;</td></tr>
<tr><td colspan="2">${SIGNED(year)}</td></tr>
</table>
</body>
</html>`;
}
/* -------------------------------------------------------------------------- */
/* Outstanding payments — Job 1 */
/* -------------------------------------------------------------------------- */
export interface OutstandingRow {
date: Date | string;
reference: string | null;
period: string | null;
type: string | null;
/** Signed amount (negative for charges). */
amount: number | string;
/** Running balance in the customer's currency, after this row. */
balance: number | string;
}
export function renderOutstanding(args: {
customerId: string;
customerName: string;
total: number | string;
rows: OutstandingRow[];
year: number;
}): string {
const rows = args.rows
.map(
(r) => `<tr>
<td>${esc(String(r.date))}</td>
<td>${esc(r.reference ?? "")}</td>
<td>${esc(r.period ?? "")}</td>
<td>${esc(r.type ?? "")}</td>
<td align="right">${esc(usd(r.amount))}</td>
<td align="right">${esc(usd(r.balance))}</td>
</tr>`,
)
.join("\n");
const body = `<p>This needs your prompt attention in order to avoid any disruption(s):</p>
<p align="center"><strong><font color="#FF0000">TOTAL OF OUTSTANDING BILLS: ${esc(
usd(args.total),
)} PESOS.</font></strong></p>
<table width="100%" border="0" cellpadding="0" cellspacing="0">
<tr><th>DATE</th><th>REFER</th><th>PERIOD</th><th>TYPEOFTRX</th><th>CHARGECREDIT</th><th>BALANCE</th></tr>
${rows}
</table>`;
return shell({
title: "Outstanding Payments",
bg: "#9CC",
heading: "Outstanding Payments",
accountId: args.customerId,
accountName: args.customerName,
body,
note: "NOTE : IF YOU ALREADY SENT THE CHECK, PLEASE DISREGARD THIS EMAIL",
year: args.year,
});
}
/* -------------------------------------------------------------------------- */
/* Payment confirmation — Job 2 */
/* -------------------------------------------------------------------------- */
export function renderPaymentConfirm(args: {
customerId: string;
customerName: string;
typeOfTrx: string;
reference: string | null;
/** The deposited amount (positive number — credits are positive in the
* unified ledger). */
amount: number | string;
year: number;
}): string {
const body = `<table width="100%" border="0" cellspacing="0" cellpadding="0">
<tr>
<td width="48%" style="font-size:20px;font-weight:bold;">${esc(
args.typeOfTrx,
)} CONFIRMATION</td>
<td width="52%" style="font-size:12px;">Please do not reply to this message. For any Jorge Cuadros &amp; Assoc. customer service inquiries, visit: <a href="https://www.jorgecuadros.com/contactus.php" target="_blank">Customer Support</a></td>
</tr>
<tr>
<td>
<strong>HI, ${esc(args.customerName)}</strong><br/>
<strong>ACCOUNT #${esc(args.customerId)}</strong><br/>
<strong>REFER# ${esc(args.reference ?? "")}</strong>
</td>
<td>
<div align="center" style="padding:20px;">
<a href="https://my.jorgecuadros.com/" target="_blank" style="color:#006699;font-weight:bold"><em>Click Here to View Your Account Statement</em></a>
</div>
</td>
</tr>
<tr><td colspan="2">&nbsp;</td></tr>
<tr><td colspan="2">
<p>Your account is now current to keep paying your future obligations. If for any reason your next bill is more than what's available; our system will email you our automatic alert requesting more funds. Thank You,</p>
<p align="center" style="color:#006600;font-weight:bold;">Your deposit was for ${esc(
usd(args.amount),
)} PESOS.</p>
</td></tr>
<tr><td colspan="2">&nbsp;</td></tr>
<tr><td colspan="2"><h4>NOTE : IF YOU ALREADY SENT THE CHECK, PLEASE DISREGARD THIS EMAIL</h4>${FOOT_CONTACT}</td></tr>
<tr><td colspan="2">&nbsp;</td></tr>
<tr><td colspan="2">${SIGNED(args.year)}</td></tr>
</table>`;
return `${HEAD.replace("{title}", "Payment Confirmation")}<body>${body}</body></html>`;
}
/* -------------------------------------------------------------------------- */
/* Account status — Job 3 (yellow + red) */
/* -------------------------------------------------------------------------- */
export function renderAccountStatus(args: {
customerId: string;
customerName: string;
level: 0 | 1; // 0 = yellow (DEBAJO DEL TIPO), 1 = red (EN ROJO)
balance: number | string;
/** Amount the customer needs to deposit to clear the threshold. */
tipo: number | string;
year: number;
}): string {
const isYellow = args.level === 0;
const body = isYellow
? `<p>In order to avoid any disruptions please mail or bring ${esc(
usd(args.tipo),
)} USD ASAP. As your current Balance ${esc(
usd(args.balance),
)} is under our minimum required to run this account.</p>`
: `<p>Sorry Account is overdrawn and all utility bills are on hold please rush ${esc(
usd(args.tipo),
)} USD these funds must be on hand ASAP to reactivate your payments.</p>`;
return shell({
title: "Account Alert",
bg: isYellow ? "#88D5EE" : "#FF8D71",
heading: "Account Alert",
accountId: args.customerId,
accountName: args.customerName,
body,
note: "NOTE : PLEASE MAKE YOUR CHECK PAYABLE TO UMC AND ASSOCIATES. IF YOU ALREADY SENT THE CHECK, PLEASE DISREGARD THIS EMAIL.",
year: args.year,
});
}
/* -------------------------------------------------------------------------- */
/* Trust payment confirmation — Job 4 */
/* -------------------------------------------------------------------------- */
export function renderTrustConfirm(args: {
customerId: string;
customerName: string;
/** Annual fee amount posted (positive, in MXN per the PHP). */
amount: number | string;
year: number;
}): string {
const body = `<p>This automatic notice is to confirm, that your Annual Bank Fee has been paid by, and posted in your account. Thank You,</p>
<p align="center"><strong>The annual fee was posted for the amount of <font color="#FF0000">${esc(
usd(args.amount),
)} PESOS.</font></strong></p>`;
return shell({
title: "Trust Payment Confirmation",
bg: "#C0BEA0",
heading: "Annual Bank Fee Payment Confirmation",
accountId: args.customerId,
accountName: args.customerName,
body,
note: "NOTE : Most banks always request to make such payment in advance.",
statementLink: "https://my.jorgecuadros.com/",
year: args.year,
});
}
+18
View File
@@ -0,0 +1,18 @@
import { Module } from "@nestjs/common";
import { OCR_PROVIDER } from "../statements/ocr/ocr.provider";
import { TesseractOcrProvider } from "../statements/ocr/tesseract.provider";
/**
* Lifts the OCR seam out of StatementsModule so other modules (today:
* PolicyOcrModule) can inject OCR_PROVIDER without taking on the rest of
* the statement intake. StatementsModule itself imports this and gets the
* provider the same way.
*
* The concrete engine is still bound here — Tesseract today, a managed
* extraction API later is a one-line change in this file.
*/
@Module({
providers: [{ provide: OCR_PROVIDER, useClass: TesseractOcrProvider }],
exports: [OCR_PROVIDER],
})
export class OcrModule {}
+76
View File
@@ -0,0 +1,76 @@
import { jobProgress } from "./ops.service";
/** Shape run_all.py emits, with the shell trace lines it interleaves. */
const line = (i: number, n: number, name: string) =>
`[paso ${i}/${n}] ${name}\n+ /repo/migration/.venv/bin/python /repo/migration/${name} --env prod\n[${name}] target env: prod\n validation: OK`;
describe("jobProgress", () => {
it("returns null before any step marker appears", () => {
// The safety backup runs before run_all.py, so this is the real state for
// the first stretch of every REIMPORT.
expect(jobProgress("== Respaldo de seguridad previo ==\ntablas capturadas: 39", "RUNNING")).toBeNull();
});
it("returns null for jobs that have no steps at all", () => {
// BACKUP/RESTORE are a single mysqldump; a fabricated percentage would be
// worse than none.
expect(jobProgress("mysqldump ... done", "SUCCESS")).toBeNull();
});
it("tracks the most recent marker, not the first", () => {
const log = [line(1, 9, "transform_customers.py"), line(2, 9, "transform_properties.py")].join("\n");
const p = jobProgress(log, "RUNNING");
expect(p).toMatchObject({ step: 2, total: 9, name: "transform_properties.py" });
});
/**
* The point of the whole feature. While RUNNING, step i is IN PROGRESS, so
* only i-1 are done. Counting i as complete would show 100% while the final
* and slowest step (blob_extract) is still working.
*/
it("does not claim a running step is finished", () => {
expect(jobProgress(line(1, 9, "transform_customers.py"), "RUNNING")?.percent).toBe(0);
expect(jobProgress(line(9, 9, "blob_extract.py"), "RUNNING")?.percent).toBe(88);
});
it("reaches 100 only once the job is no longer running", () => {
expect(jobProgress(line(9, 9, "blob_extract.py"), "SUCCESS")?.percent).toBe(100);
});
/** A job that died mid-way must report where it died, not 100%. */
it("reports the failed step rather than completion", () => {
const p = jobProgress(line(5, 9, "transform_transactions.py"), "FAILED");
expect(p).toMatchObject({ step: 5, total: 9 });
expect(p!.percent).toBe(55);
});
it("handles the 8-step SYNC list as well as the 9-step REIMPORT one", () => {
expect(jobProgress(line(8, 8, "transform_bank.py"), "SUCCESS")?.percent).toBe(100);
expect(jobProgress(line(4, 8, "transform_policies.py"), "RUNNING")?.percent).toBe(37);
});
/**
* Captured verbatim from `run_all.run(..., step=8, total=9)`. This is the
* contract between the Python and this parser; if run_all.py's format
* changes, this fails rather than the panel silently showing no progress.
*/
it("parses the exact line run_all.py emits", () => {
const real =
"[paso 8/9] transform_bank.py\n+ /repo/migration/.venv/bin/python /repo/migration/transform_bank.py --env prod";
expect(jobProgress(real, "RUNNING")).toMatchObject({
step: 8,
total: 9,
name: "transform_bank.py",
percent: 77,
});
});
it("ignores a malformed marker instead of reporting NaN", () => {
expect(jobProgress("[paso 3/0] x.py", "RUNNING")).toBeNull();
});
/** The marker must be at line start so log text quoting it cannot spoof it. */
it("does not match a marker embedded mid-line", () => {
expect(jobProgress("some output mentioning [paso 4/9] fake.py", "RUNNING")).toBeNull();
});
});
+34 -2
View File
@@ -19,6 +19,7 @@ import { AbilityGuard } from "../auth/ability.guard";
import { RequireAbility } from "../auth/require-ability.decorator"; import { RequireAbility } from "../auth/require-ability.decorator";
import { AuditService } from "../common/audit.service"; import { AuditService } from "../common/audit.service";
import { OpsService } from "./ops.service"; import { OpsService } from "./ops.service";
import { ReplicationService } from "./replication.service";
import { StartJobDto } from "./start-job.dto"; import { StartJobDto } from "./start-job.dto";
/** Every route is ADMIN-only (ability "db:manage"). */ /** Every route is ADMIN-only (ability "db:manage"). */
@@ -28,6 +29,7 @@ import { StartJobDto } from "./start-job.dto";
export class OpsController { export class OpsController {
constructor( constructor(
private readonly ops: OpsService, private readonly ops: OpsService,
private readonly replication: ReplicationService,
private readonly audit: AuditService, private readonly audit: AuditService,
) {} ) {}
@@ -44,7 +46,7 @@ export class OpsController {
@Post("ingest/:name") @Post("ingest/:name")
@UseInterceptors( @UseInterceptors(
FileInterceptor("file", { limits: { fileSize: 500 * 1024 * 1024 } }), FileInterceptor("file", { limits: { fileSize: 2 * 1024 * 1024 * 1024 } }),
) )
async uploadIngest( async uploadIngest(
@Param("name") name: string, @Param("name") name: string,
@@ -96,6 +98,30 @@ export class OpsController {
/* --------------------------------------------------------------- jobs */ /* --------------------------------------------------------------- jobs */
/** Health of the my.jorgecuadros.com read replica. Read-only, no audit entry. */
@Get("replication")
replicationStatus() {
return this.replication.status();
}
/**
* Full row-by-row comparison of the customer-visible tables against the master.
*
* POST rather than GET despite reading nothing: it is a full scan of both
* servers and must not be something a browser prefetch, a retry, or a refresh
* can set off. Audited for the same reason — it is a deliberate, costly act,
* and "who ran this while the site was slow" is a question worth answering.
*/
@Post("replication/verify")
async verifyReplication(@Req() req: Request) {
const result = await this.replication.verify();
void this.audit.log(this.actingId(req), "ops.replication.verify", {
identical: result.identical,
elapsedMs: result.elapsedMs,
});
return result;
}
@Get("jobs") @Get("jobs")
listJobs() { listJobs() {
return this.ops.listJobs(); return this.ops.listJobs();
@@ -109,11 +135,17 @@ export class OpsController {
@Post("jobs") @Post("jobs")
async startJob(@Body() dto: StartJobDto, @Req() req: Request) { async startJob(@Body() dto: StartJobDto, @Req() req: Request) {
const userId = this.actingId(req); const userId = this.actingId(req);
const job = await this.ops.startJob(dto.kind, { file: dto.file }, userId); const job = await this.ops.startJob(
dto.kind,
{ file: dto.file, forceFull: dto.forceFull },
userId,
);
void this.audit.log(userId, "ops.job.start", { void this.audit.log(userId, "ops.job.start", {
jobId: job.id, jobId: job.id,
kind: dto.kind, kind: dto.kind,
file: dto.file, file: dto.file,
// Recorded because this is the flag that authorised deleting native rows.
forceFull: dto.forceFull,
}); });
return job; return job;
} }
+2 -1
View File
@@ -1,9 +1,10 @@
import { Module } from "@nestjs/common"; import { Module } from "@nestjs/common";
import { OpsController } from "./ops.controller"; import { OpsController } from "./ops.controller";
import { OpsService } from "./ops.service"; import { OpsService } from "./ops.service";
import { ReplicationService } from "./replication.service";
@Module({ @Module({
controllers: [OpsController], controllers: [OpsController],
providers: [OpsService], providers: [OpsService, ReplicationService],
}) })
export class OpsModule {} export class OpsModule {}
+230 -12
View File
@@ -31,6 +31,14 @@ export const INGEST_FILES = [
] as const; ] as const;
export type IngestName = (typeof INGEST_FILES)[number]; export type IngestName = (typeof INGEST_FILES)[number];
/**
* Prefix for every command containing a pipe. Without it the exit status of
* `mysqldump | gzip` is gzip's, so a dump that failed immediately still looks
* like a successful job. Both Alpine's busybox ash (the API image) and macOS
* `sh` (dev) support it; POSIX does not require it, so `sh -c` is the contract.
*/
const PIPEFAIL = "set -o pipefail; ";
interface MysqlConn { interface MysqlConn {
host: string; host: string;
port: string; port: string;
@@ -43,8 +51,11 @@ interface MysqlConn {
export class OpsService implements OnModuleInit { export class OpsService implements OnModuleInit {
private readonly logger = new Logger(OpsService.name); private readonly logger = new Logger(OpsService.name);
// Resolve from this source file so it works regardless of process.cwd()
// (the API runs from apps/api/, but the Python ETL lives at repo-root migration/).
private readonly migrationDir = private readonly migrationDir =
process.env.MIGRATION_DIR ?? path.resolve(process.cwd(), "migration"); process.env.MIGRATION_DIR ??
path.resolve(__dirname, "..", "..", "..", "..", "migration");
private readonly ingestDir = private readonly ingestDir =
process.env.INGEST_DIR ?? path.join(this.migrationDir, "ingest"); process.env.INGEST_DIR ?? path.join(this.migrationDir, "ingest");
private readonly backupDir = private readonly backupDir =
@@ -56,6 +67,55 @@ export class OpsService implements OnModuleInit {
async onModuleInit(): Promise<void> { async onModuleInit(): Promise<void> {
await fs.mkdir(this.ingestDir, { recursive: true }); await fs.mkdir(this.ingestDir, { recursive: true });
await fs.mkdir(this.backupDir, { recursive: true }); await fs.mkdir(this.backupDir, { recursive: true });
await this.reconcileOrphanedJobs();
}
/**
* Fail any job still marked RUNNING at startup.
*
* Jobs run as a child of THIS process, so no job can outlive it: if a row says
* RUNNING while we are booting, its process died with the previous instance
* and nothing will ever finalize it. Since startJob() refuses to start while
* any RUNNING row exists, one interrupted job wedges the panel permanently
* with no way out from the UI — it took a manual UPDATE against production to
* recover the first time this happened, when a deploy landed 110 seconds into
* a REIMPORT.
*
* Deliberately unconditional rather than filtered on age: "started recently"
* does not mean "still alive" here, and a fresh boot is proof enough that
* nothing survived.
*/
private async reconcileOrphanedJobs(): Promise<void> {
try {
// Read then write one by one rather than updateMany: the log needs the
// reason APPENDED, and a job whose log just stops mid-step with no
// explanation is what made the first occurrence hard to diagnose.
const orphans = await this.prisma.opsJob.findMany({
where: { status: "RUNNING" },
select: { id: true, kind: true, log: true },
});
for (const job of orphans) {
await this.prisma.opsJob.update({
where: { id: job.id },
data: {
status: "FAILED",
finishedAt: new Date(),
log: {
set:
job.log +
"\n[interrumpido: el contenedor se reinició mientras el trabajo corría; " +
"el proceso hijo no sobrevive a un redespliegue. " +
"Vuelva a ejecutar la operación desde el principio.]\n",
},
},
});
this.logger.warn(`trabajo ${job.kind} ${job.id} quedó huérfano; marcado FAILED`);
}
} catch (e) {
// Never block startup on this. A failed reconcile leaves the panel
// wedged, which is bad, but an API that will not boot is worse.
this.logger.error(`no se pudieron reconciliar trabajos huérfanos: ${String(e)}`);
}
} }
/* -------------------------------------------------------------- ingest */ /* -------------------------------------------------------------- ingest */
@@ -152,7 +212,9 @@ export class OpsService implements OnModuleInit {
async getJob(id: string) { async getJob(id: string) {
const job = await this.prisma.opsJob.findUnique({ where: { id } }); const job = await this.prisma.opsJob.findUnique({ where: { id } });
if (!job) throw new NotFoundException("Trabajo no encontrado."); if (!job) throw new NotFoundException("Trabajo no encontrado.");
return job; // Derived, never stored: the log is the single source of truth for how far
// a job got, so progress cannot drift out of sync with it.
return { ...job, progress: jobProgress(job.log, job.status) };
} }
/** /**
@@ -172,7 +234,7 @@ export class OpsService implements OnModuleInit {
); );
} }
const conn = this.parseDbUrl(); const conn = this.opsConn();
const { cmd, resolvedParams } = await this.buildCommand(kind, params, conn); const { cmd, resolvedParams } = await this.buildCommand(kind, params, conn);
const job = await this.prisma.opsJob.create({ const job = await this.prisma.opsJob.create({
@@ -204,6 +266,37 @@ export class OpsService implements OnModuleInit {
}; };
} }
/**
* The credentials mysqldump/mysql run as — deliberately NOT the application
* user. `--single-transaction` issues FLUSH TABLES, which needs the global
* RELOAD privilege, and the app user is granted only `ALL ON jorgecuadros.*`
* plus `USAGE ON *.*`; `--skip-lock-tables` does not avoid it. A restore of a
* dump taken before --set-gtid-purged=OFF likewise needs SUPER to replay its
* SET @@GLOBAL.GTID_PURGED. So an admin credential is supplied out of band
* rather than elevating the runtime user for the sake of one admin screen —
* the same choice deploy/scripts/pre-migrate-backup.mjs makes.
*
* Host, port and database always come from DATABASE_URL: the ops user is a
* different login on the SAME server, never a way to point at another one.
*
* With the vars unset this falls back to the DATABASE_URL credentials, which
* is what local development wants — a dev MySQL grants the app user far more.
*/
private opsConn(): MysqlConn {
const conn = this.parseDbUrl();
const user = process.env.OPS_DB_ADMIN_USER;
const password = process.env.OPS_DB_ADMIN_PASSWORD;
if (!user || !password) {
this.logger.warn(
"OPS_DB_ADMIN_USER/OPS_DB_ADMIN_PASSWORD no configuradas; " +
`usando el usuario de la aplicación (${conn.user}) para mysqldump. ` +
"En producción esto falla por falta del privilegio RELOAD.",
);
return conn;
}
return { ...conn, user, password };
}
/** mysql/mysqldump connection flags. The password goes through MYSQL_PWD in /** mysql/mysqldump connection flags. The password goes through MYSQL_PWD in
* the child env, never on the command line (which would leak via `ps`). */ * the child env, never on the command line (which would leak via `ps`). */
private connFlags(c: MysqlConn): string { private connFlags(c: MysqlConn): string {
@@ -214,6 +307,58 @@ export class OpsService implements OnModuleInit {
return new Date().toISOString().replace(/[:.]/g, "-").replace("T", "_").slice(0, 19); return new Date().toISOString().replace(/[:.]/g, "-").replace("T", "_").slice(0, 19);
} }
/**
* One hardened mysqldump, shared by BACKUP and by the safety backups SYNC and
* REIMPORT take first. Kept byte-for-byte in spirit with the dump in
* deploy/scripts/pre-migrate-backup.mjs — the two write into the same volume
* and both are listed as restore points by this same screen.
*
* The dumper is probed at runtime rather than assumed. This command runs
* inside the API image, whose `mysql-client` is Alpine's — i.e. MariaDB's —
* where `mysqldump` is a deprecation-warning shim over `mariadb-dump` that
* rejects --set-gtid-purged outright:
* mysqldump: unknown variable 'set-gtid-purged=OFF'
* which failed every backup, including the safety backups SYNC and REIMPORT
* take first. MariaDB's dumper emits no GTID state unless asked (--gtid), so
* there is nothing to suppress there; the flag is passed only when the dumper
* on PATH advertises it, and the real binary is called directly only in the
* MariaDB case (calling `mariadb-dump` whenever it merely exists would pick
* it over a MySQL `mysqldump` earlier in PATH on a host carrying both).
*
* The probe is a command substitution, not `--help | grep -q`: PIPEFAIL is in
* effect and grep closing the pipe early would make a supported flag look
* unsupported.
*
* --set-gtid-purged=OFF (MySQL only): the production server is the
* replication SOURCE with GTID on, so without it every dump embeds
* SET @@GLOBAL.GTID_PURGED and is unrestorable onto the very server it came
* from.
*
* The table-count assertion is not belt-and-braces: `gzip -t` passes on the
* ~372-byte output of a mysqldump that died on its first statement, so a
* failed dump would otherwise be recorded as a successful backup. (`set -o
* pipefail` is set by the caller for the same reason — without it the exit
* status of the pipeline is gzip's, and gzip succeeded.)
*
* A failed attempt deletes its own output, so a truncated file never appears
* in the restore list looking like an ordinary restore point.
*/
private dumpCommand(flags: string, db: string, out: string): string {
return (
`DUMP=mysqldump; GTID=; ` +
`case "$(mysqldump --help 2>/dev/null || true)" in ` +
`*set-gtid-purged*) GTID=--set-gtid-purged=OFF;; ` +
`*) command -v mariadb-dump >/dev/null 2>&1 && DUMP=mariadb-dump;; esac; ` +
`( $DUMP ${flags} --single-transaction --routines --triggers ` +
`--no-tablespaces $GTID ${db} | gzip -c > ${out} && ` +
`gzip -t ${out} && ` +
`TABLAS=$(gunzip -c ${out} | grep -c 'CREATE TABLE') && ` +
`echo "tablas capturadas: $TABLAS" && ` +
`[ "$TABLAS" -ge 1 ] ) || ` +
`{ rm -f ${out}; echo 'respaldo incompleto eliminado'; exit 1; }`
);
}
private async buildCommand( private async buildCommand(
kind: OpsJobKind, kind: OpsJobKind,
params: Record<string, unknown>, params: Record<string, unknown>,
@@ -226,7 +371,7 @@ export class OpsService implements OnModuleInit {
const file = `backup-${this.migrationEnv}-${this.timestamp()}.sql.gz`; const file = `backup-${this.migrationEnv}-${this.timestamp()}.sql.gz`;
const out = shq(path.join(this.backupDir, file)); const out = shq(path.join(this.backupDir, file));
return { return {
cmd: `mysqldump ${flags} --single-transaction --routines --triggers --no-tablespaces ${db} | gzip -c > ${out}`, cmd: `${PIPEFAIL}${this.dumpCommand(flags, db, out)}`,
resolvedParams: { file }, resolvedParams: { file },
}; };
} }
@@ -238,7 +383,10 @@ export class OpsService implements OnModuleInit {
throw new NotFoundException(`Respaldo no encontrado: ${name}`); throw new NotFoundException(`Respaldo no encontrado: ${name}`);
}); });
return { return {
cmd: `gunzip -c ${shq(full)} | mysql ${flags} ${db}`, // pipefail matters here too: a corrupt archive makes gunzip fail while
// mysql, fed a truncated stream, can still exit 0 — a restore that
// reported success having replayed only part of the dump.
cmd: `${PIPEFAIL}gunzip -c ${shq(full)} | mysql ${flags} ${db}`,
resolvedParams: { file: name }, resolvedParams: { file: name },
}; };
} }
@@ -249,10 +397,15 @@ export class OpsService implements OnModuleInit {
const py = await this.pythonBin(); const py = await this.pythonBin();
const runAll = shq(path.join(this.migrationDir, "run_all.py")); const runAll = shq(path.join(this.migrationDir, "run_all.py"));
const cmd = const cmd =
`echo '== Respaldo de seguridad previo ==' && ` + `${PIPEFAIL}echo '== Respaldo de seguridad previo ==' && ` +
`mysqldump ${flags} --single-transaction --routines --triggers --no-tablespaces ${db} | gzip -c > ${out} && ` + `${this.dumpCommand(flags, db, out)} && ` +
`echo '== Sincronización aditiva desde carpeta de ingesta ==' && ` + `echo '== Sincronización aditiva desde carpeta de ingesta ==' && ` +
`${shq(py)} ${runAll} --env ${shq(this.migrationEnv)} --sync`; // --stage is not optional here. The staged Parquet lives in the image
// at migration/output, NOT on a volume, so every redeploy wipes it and
// a sync without --stage dies on a missing stg_*/*.parquet. Re-staging
// is also the only thing that makes "desde carpeta de ingesta" true:
// stale Parquet would sync the previous upload, not the current one.
`${shq(py)} ${runAll} --env ${shq(this.migrationEnv)} --stage --sync`;
return { cmd, resolvedParams: { safetyBackup: file } }; return { cmd, resolvedParams: { safetyBackup: file } };
} }
@@ -262,12 +415,19 @@ export class OpsService implements OnModuleInit {
const out = shq(path.join(this.backupDir, file)); const out = shq(path.join(this.backupDir, file));
const py = await this.pythonBin(); const py = await this.pythonBin();
const runAll = shq(path.join(this.migrationDir, "run_all.py")); const runAll = shq(path.join(this.migrationDir, "run_all.py"));
// run_all.py runs native_guard.py before it truncates anything and exits
// without touching the database when the target holds rows that only
// exist here — allocated portal NUMids, app-created customers, OCR
// captures. --force-full is what the operator ticks to delete them
// anyway; without it the job fails with the list.
const force = params.forceFull === true;
const cmd = const cmd =
`echo '== Respaldo de seguridad previo ==' && ` + `${PIPEFAIL}echo '== Respaldo de seguridad previo ==' && ` +
`mysqldump ${flags} --single-transaction --routines --triggers --no-tablespaces ${db} | gzip -c > ${out} && ` + `${this.dumpCommand(flags, db, out)} && ` +
`echo '== Reimportación desde carpeta de ingesta ==' && ` + `echo '== Reimportación desde carpeta de ingesta ==' && ` +
`${shq(py)} ${runAll} --env ${shq(this.migrationEnv)} --stage`; `${shq(py)} ${runAll} --env ${shq(this.migrationEnv)} --stage` +
return { cmd, resolvedParams: { safetyBackup: file } }; (force ? " --force-full" : "");
return { cmd, resolvedParams: { safetyBackup: file, forceFull: force } };
} }
throw new BadRequestException(`Operación no soportada: ${kind}`); throw new BadRequestException(`Operación no soportada: ${kind}`);
@@ -284,6 +444,15 @@ export class OpsService implements OnModuleInit {
} }
} }
/**
* `password` is the ops credential from opsConn(), exported as MYSQL_PWD so it
* never reaches argv (which `ps` exposes to every process on the host).
*
* It does not leak into the Python ETL that SYNC and REIMPORT go on to run:
* migration/dbenv.py connects with pymysql using the credentials inside
* DATABASE_URL and never consults MYSQL_PWD. The ETL keeps running as the
* application user, which is what it should be doing.
*/
private run(jobId: string, cmd: string, password: string): void { private run(jobId: string, cmd: string, password: string): void {
const child = spawn("sh", ["-c", cmd], { const child = spawn("sh", ["-c", cmd], {
cwd: this.migrationDir, cwd: this.migrationDir,
@@ -359,3 +528,52 @@ export class OpsService implements OnModuleInit {
function shq(v: string): string { function shq(v: string): string {
return `'${v.replace(/'/g, `'\\''`)}'`; return `'${v.replace(/'/g, `'\\''`)}'`;
} }
/** Progress derived from a job's log. Null when the job reports no steps. */
export interface JobProgress {
/** 1-based index of the step currently running (or last reached). */
step: number;
total: number;
/** Script name, e.g. "transform_bank.py". */
name: string;
/** 0..100, floored. 100 only once the job is no longer RUNNING. */
percent: number;
}
/**
* Parse the "[paso i/N] name" markers migration/run_all.py emits.
*
* Progress is DERIVED from the log rather than tracked in a column: the log is
* already the record of what happened, and a separate counter could disagree
* with it — which is exactly the confusion a progress display is supposed to
* remove. run_all.py owns the step count, so adding a step cannot desync this.
*
* BACKUP and RESTORE are a single mysqldump with no steps, so they return null
* and the UI shows an indeterminate spinner. Reporting a fabricated percentage
* for them would be worse than showing none.
*/
export function jobProgress(
log: string,
status: string,
): JobProgress | null {
// Last marker wins: the log grows, and the newest line is the current step.
const matches = [...log.matchAll(/^\[paso (\d+)\/(\d+)\] (\S+)/gm)];
const last = matches[matches.length - 1];
if (!last) return null;
const step = Number(last[1]);
const total = Number(last[2]);
if (!Number.isFinite(step) || !Number.isFinite(total) || total <= 0) return null;
// While RUNNING, step i means i is IN PROGRESS, not finished — so report
// (i-1) completed. Claiming 100% while the last step is still working is the
// classic progress-bar lie, and here the last step (blob_extract) is also the
// slowest, so it would sit at "100%" for the longest stretch of the job.
const done = status === "RUNNING" ? step - 1 : step;
return {
step,
total,
name: last[3],
percent: Math.max(0, Math.min(100, Math.floor((done / total) * 100))),
};
}
+639
View File
@@ -0,0 +1,639 @@
import { Injectable, Logger } from "@nestjs/common";
import { execFile } from "node:child_process";
import { promisify } from "node:util";
const exec = promisify(execFile);
/**
* How far the SQL thread is behind the I/O thread, in source binlog bytes.
*
* This is a different question from `secondsBehind`, and it answers the case
* that lag hides: while the SQL thread grinds through one huge transaction,
* `Seconds_Behind_Source` can sit still or even read 0, but the relay backlog
* is plainly shrinking (or not). It costs nothing extra — every field here
* comes out of the same `SHOW REPLICA STATUS` the panel already runs.
*
* Both positions are coordinates in the SOURCE's binlog, so they are only
* comparable while both threads are working on the SAME source file. When they
* are not, the replica is whole files behind and the byte delta is meaningless
* (positions restart at ~4 in each new file), so `backlogBytes` and `percent`
* are null and `sameFile` says why.
*/
export interface ApplyProgress {
/** Source binlog file the I/O thread is currently reading. */
sourceLogFile: string | null;
/** Position in `sourceLogFile` that the I/O thread has fetched up to. */
readPos: number;
/** Source binlog file the SQL thread is currently applying. */
relayLogFile: string | null;
/** Position in `relayLogFile` that the SQL thread has applied up to. */
execPos: number;
/** True while both threads are on the same source file. */
sameFile: boolean;
/** Fetched-but-not-yet-applied bytes. Null when the files differ. */
backlogBytes: number | null;
/**
* `execPos / readPos` as a percentage, null when the files differ.
*
* Deliberately never rounded up to 100 while any backlog remains: binlog
* positions are large, so a real backlog of a few KB is 99.99% of the file
* and would render as "caught up" when it is not. Read `backlogBytes === 0`
* for actually caught up.
*/
percent: number | null;
}
/**
* How far the replica's executed history is from the master's, in transactions.
*
* This is the check `SHOW REPLICA STATUS` cannot give you, and it is stronger
* than everything else on the card for one specific reason: every other field is
* self-reported by the replica. `Seconds_Behind_Source` reads 0 both when there
* is genuinely nothing to apply AND when the I/O thread is disconnected — with
* no incoming event there is nothing to measure staleness against, so a dead
* link reports as perfectly current. `GTID_SUBTRACT(master, replica)` asks the
* master what it has done and the replica what it has applied, so a silent
* disconnect shows up immediately as a growing number.
*/
export interface GtidDrift {
/** Transactions the master executed that the replica has not. 0 = identical. */
missingTransactions: number;
/** The missing GTID set verbatim. Null when nothing is missing. */
missingGtidSet: string | null;
/**
* Transactions in the replica's `gtid_executed` under its OWN server UUID —
* writes that happened here and exist nowhere on the master.
*
* Reported, never alarmed on. A non-zero count is the expected residue of the
* seed load: restoring a dump executes its statements locally, and they take
* GTIDs from this server's UUID. They never propagate (`log_replica_updates`
* is off and nothing sources from this node), so they are harmless — right up
* until someone tries to promote this box, where they become a real divergence.
*/
localTransactions: number;
}
/** One table's row count and content fingerprint, on one side of the link. */
export interface TableFingerprint {
table: string;
masterRows: number;
replicaRows: number;
/** Order-independent checksum over every column of every row. */
masterChecksum: string;
replicaChecksum: string;
matches: boolean;
}
export interface VerifyResult {
/** True only when every table matched on both count and checksum. */
identical: boolean;
tables: TableFingerprint[];
/** Set instead of `tables` when the comparison could not be run at all. */
problem: string | null;
checkedAt: string;
/** Wall-clock cost, because this is a full scan and the caller should see it. */
elapsedMs: number;
}
export interface ReplicationStatus {
/** false when the replica is not configured for this environment at all. */
configured: boolean;
/** true only when both threads run, no error is set, and lag is within bounds. */
healthy: boolean;
host: string | null;
ioRunning: string | null;
sqlRunning: string | null;
/** null when MySQL reports NULL, which it does whenever a thread is down. */
secondsBehind: number | null;
lastIoError: string | null;
lastSqlError: string | null;
sourceHost: string | null;
/** Relay-log apply progress. Null when the status output has no positions. */
apply: ApplyProgress | null;
/** GTID comparison against the master. Null when the master was unreachable. */
drift: GtidDrift | null;
/** Human-readable reason when healthy is false. */
problem: string | null;
checkedAt: string;
}
/**
* The tables `my.jorgecuadros.com` reads through the `web_reader` grant.
*
* This list is the verification surface, not the replication surface — the
* replica carries the whole schema. These are the eight whose divergence would
* actually be visible to a customer, so they are the ones worth a full scan.
*/
export const REPLICATED_TABLES = [
"transactions",
"customers",
"customer_legacy_refs",
"type_transactions",
"exchange_rates",
"properties",
"property_services",
"trust_accounts",
] as const;
/**
* Reports whether the my.jorgecuadros.com read replica is still replicating.
*
* The replica is what the public site reads once the platformDataSource flag is
* on, and a replica that has silently stopped applying serves stale balances
* rather than erroring — the failure is invisible from the site itself, which is
* why it needs a panel.
*
* Shells out to the mysql client for the same reason the rest of OpsService
* does: there is no MySQL driver in this API's dependencies, and the image
* already ships one.
*/
@Injectable()
export class ReplicationService {
private readonly logger = new Logger(ReplicationService.name);
/** Lag above this many seconds is reported as unhealthy. */
private readonly maxLagSeconds = Number(process.env.REPLICA_MAX_LAG ?? 60);
async status(): Promise<ReplicationStatus> {
const host = process.env.REPLICA_DB_HOST;
const user = process.env.REPLICA_DB_USER;
const password = process.env.REPLICA_DB_PASS;
const now = new Date().toISOString();
const empty: ReplicationStatus = {
configured: false,
healthy: false,
host: host ?? null,
ioRunning: null,
sqlRunning: null,
secondsBehind: null,
lastIoError: null,
lastSqlError: null,
sourceHost: null,
apply: null,
drift: null,
problem: null,
checkedAt: now,
};
if (!host || !user || !password) {
return { ...empty, problem: "REPLICA_DB_* no configuradas" };
}
let raw: string;
try {
raw = await this.onReplica("SHOW REPLICA STATUS\\G");
} catch (e) {
const msg = e instanceof Error ? e.message : String(e);
this.logger.warn(`no se pudo consultar la réplica: ${msg}`);
return { ...empty, configured: true, problem: `No se pudo conectar: ${msg}` };
}
const field = (name: string): string | null => replicaField(raw, name);
// An empty result set means the server is not configured as a replica at
// all — distinct from "configured but broken", and worth saying plainly.
if (!raw.includes("Replica_IO_Running")) {
return {
...empty,
configured: true,
problem: "El servidor no está configurado como réplica",
};
}
const ioRunning = field("Replica_IO_Running");
const sqlRunning = field("Replica_SQL_Running");
const lagRaw = field("Seconds_Behind_Source");
const secondsBehind =
lagRaw === null || lagRaw === "NULL" ? null : Number(lagRaw);
const lastIoError = field("Last_IO_Error");
const lastSqlError = field("Last_SQL_Error");
// Order matters: report the most specific cause first. Checking lag before
// the threads would blame "sin dato de retraso" for what is really a
// stopped thread, because MySQL reports NULL lag whenever either is down.
let problem: string | null = null;
if (ioRunning !== "Yes") problem = "El hilo de E/S no está corriendo";
else if (sqlRunning !== "Yes") problem = "El hilo SQL no está corriendo";
else if (lastSqlError) problem = `Error SQL: ${lastSqlError}`;
else if (lastIoError) problem = `Error de E/S: ${lastIoError}`;
else if (secondsBehind === null) problem = "Sin dato de retraso";
else if (secondsBehind > this.maxLagSeconds)
problem = `Retraso de ${secondsBehind}s (máximo ${this.maxLagSeconds}s)`;
return {
configured: true,
healthy: problem === null,
host,
ioRunning,
sqlRunning,
secondsBehind,
lastIoError,
lastSqlError,
sourceHost: field("Source_Host"),
// Reported, never folded into `healthy`: a non-zero backlog is the normal
// state of a working replica for the instant between fetch and apply, so
// alarming on it would cry wolf. It is here to answer "is it moving?"
// when the lag counter is stuck.
apply: applyProgress(raw),
// Also reported rather than alarmed on, for the same reason: a busy master
// is always a few transactions ahead for the instant they are in flight.
// Null rather than zero when the master could not be reached — "unknown"
// and "identical" must not render the same.
drift: await this.gtidDrift(),
problem,
checkedAt: now,
};
}
/**
* Compare executed history between master and replica.
*
* Two round trips: ask the master what it has executed, then ask the replica
* to subtract its own history from that. The subtraction runs on the replica
* rather than in TypeScript because `GTID_SUBTRACT` already implements the
* interval algebra correctly, and reimplementing set subtraction over binlog
* ranges is exactly the kind of thing that looks right and is wrong at the
* boundaries.
*
* @returns null on any failure — a broken drift check must never be mistaken
* for a healthy zero.
*/
private async gtidDrift(): Promise<GtidDrift | null> {
try {
const masterGtid = (await this.onMaster("SELECT @@gtid_executed")).trim();
// GTID sets are UUIDs, digits, colons, commas, hyphens, whitespace and
// (since 8.4) alphanumeric tags. Nothing else is legal, so rejecting
// anything outside that alphabet is a whitelist, not a blacklist: with no
// quote and no backslash able to survive it, the value cannot escape the
// string literal it is interpolated into below.
if (masterGtid && !/^[0-9a-fA-F:,\s_-]+$/.test(masterGtid)) {
this.logger.warn("gtid_executed del maestro con formato inesperado");
return null;
}
// An empty set means the master has GTID mode off, and there is nothing
// meaningful to compare.
if (!masterGtid) return null;
const flat = masterGtid.replace(/\s+/g, "");
// Every GTID set is flattened with REPLACE before it leaves the server.
// MySQL wraps `gtid_executed` across lines once it holds more than one
// source UUID, and this is read back as tab-separated columns — an
// embedded newline would split one row into two and silently truncate the
// set at the first UUID.
const out = await this.onReplica(
"SELECT REPLACE(GTID_SUBTRACT(" +
`'${flat}', @@gtid_executed), '\\n', ''), ` +
"@@server_uuid, REPLACE(@@gtid_executed, '\\n', '')",
["-N"],
);
// Trailing newline only — never `.trim()`. When nothing is missing the
// first column is the empty string, so the line begins with a tab, and
// trimming it would shift every column one position left and report the
// replica's own UUID as the missing GTID set.
const [missingSet = "", serverUuid = "", executed = ""] = out
.replace(/\r?\n+$/, "")
.split("\t");
return {
missingTransactions: countGtids(missingSet),
missingGtidSet: missingSet || null,
localTransactions: countGtids(gtidsForUuid(executed, serverUuid)),
};
} catch (e) {
const msg = e instanceof Error ? e.message : String(e);
this.logger.warn(`no se pudo comparar GTIDs con el maestro: ${msg}`);
return null;
}
}
/**
* Full-scan comparison of the customer-visible tables on both sides.
*
* Deliberately NOT part of `status()`: this reads every row of every table in
* `REPLICATED_TABLES` on both servers, so it belongs behind a button, not a
* 30-second poll.
*
* It answers the one question GTID drift cannot. GTIDs prove the replica
* applied every transaction the master produced; they say nothing about rows
* changed on the replica by some other route. A local write is invisible to
* every other field on the card and shows up here as a checksum mismatch.
*/
async verify(): Promise<VerifyResult> {
const started = Date.now();
const base: VerifyResult = {
identical: false,
tables: [],
problem: null,
checkedAt: new Date().toISOString(),
elapsedMs: 0,
};
let sql: string;
try {
sql = await this.fingerprintSql();
} catch (e) {
const msg = e instanceof Error ? e.message : String(e);
return { ...base, problem: `No se pudo leer el esquema: ${msg}`, elapsedMs: Date.now() - started };
}
let masterOut: string;
let replicaOut: string;
try {
// Sequential, not parallel. Running both at once would have the master
// scan under the replica's own read load only sometimes, which makes a
// slow run hard to attribute; and the boxes are small enough that two
// concurrent full scans is a real memory event on the 946MB replica.
masterOut = await this.onMaster(sql);
replicaOut = await this.onReplica(sql);
} catch (e) {
const msg = e instanceof Error ? e.message : String(e);
return { ...base, problem: `No se pudo comparar: ${msg}`, elapsedMs: Date.now() - started };
}
const master = parseFingerprints(masterOut);
const replica = parseFingerprints(replicaOut);
const tables: TableFingerprint[] = REPLICATED_TABLES.map((table) => {
const m = master.get(table);
const r = replica.get(table);
return {
table,
masterRows: m?.rows ?? -1,
replicaRows: r?.rows ?? -1,
masterChecksum: m?.checksum ?? "?",
replicaChecksum: r?.checksum ?? "?",
// Both sides must have answered. A missing row on either side is a
// mismatch, never a pass — `undefined === undefined` would otherwise
// report two failed reads as agreement.
matches:
m !== undefined && r !== undefined && m.rows === r.rows && m.checksum === r.checksum,
};
});
return {
identical: tables.every((t) => t.matches),
tables,
problem: null,
checkedAt: base.checkedAt,
elapsedMs: Date.now() - started,
};
}
/**
* Build the count+checksum query from the live column list.
*
* The columns come from `information_schema` on the master rather than being
* hardcoded, so the check keeps covering the whole row after a migration adds
* one. Reading the schema from the master is safe by construction: if the two
* schemas had diverged, replication would already be broken.
*/
private async fingerprintSql(): Promise<string> {
const list = REPLICATED_TABLES.map((t) => `'${t}'`).join(",");
const raw = await this.onMaster(
"SELECT CONCAT(TABLE_NAME, '\\t', COLUMN_NAME) FROM information_schema.COLUMNS " +
`WHERE TABLE_SCHEMA = DATABASE() AND TABLE_NAME IN (${list}) ` +
"ORDER BY TABLE_NAME, ORDINAL_POSITION",
["-N"],
);
const cols = new Map<string, string[]>();
for (const line of raw.split("\n")) {
const [table, column] = line.trim().split("\t");
if (!table || !column) continue;
cols.set(table, [...(cols.get(table) ?? []), column]);
}
const selects = REPLICATED_TABLES.map((table) => {
const columns = cols.get(table);
if (!columns?.length) throw new Error(`tabla ${table} sin columnas`);
// CONVERT(... USING binary), never CAST(... AS CHAR).
//
// CAST to CHAR transcodes into the *connection* character set, which is
// not the same on the two servers: the mysql client inside the master's
// container negotiates latin1, while the replica's negotiates utf8mb4.
// Every accented character in a Mexican name, street or note therefore
// hashes to different bytes on each side, and the comparison reports a
// permanent mismatch on exactly the tables that hold free text — a
// verification tool that always cries wolf, which is worse than none.
// Comparing the stored bytes sidesteps the session entirely. (Verified
// 2026-08-06: with CAST, `customers.name` gave 3344437324815 vs
// 3339150372121; with CONVERT both give 3339150372121.)
//
// 0x1f (unit separator) joins the columns and 0x1e (record separator)
// stands in for NULL. Both matter: CONCAT_WS *skips* NULLs rather than
// emitting an empty field, so without a placeholder the rows
// ('a', NULL, 'b') and ('a', 'b', NULL) produce the same string and a
// column-shifting bug would checksum as identical.
const expr = columns
.map((c) => `IFNULL(CONVERT(\`${c}\` USING binary), 0x1e)`)
.join(", 0x1f, ");
// SUM, not a running hash: addition is commutative, so the result does not
// depend on the order rows come back in. The two servers have no reason to
// scan in the same order and are not asked to.
return (
`SELECT '${table}' AS t, COUNT(*) AS n, ` +
`IFNULL(SUM(CRC32(CONCAT_WS(0x1f, ${expr}))), 0) AS c FROM \`${table}\``
);
});
return selects.join(" UNION ALL ");
}
/* ------------------------------------------------------------ plumbing */
/**
* Run a statement on the replica.
*
* --ssl is required: the replica sets require_secure_transport=ON.
*
* --ssl-verify-server-cert=0 is deliberate and is NOT the same trade-off the
* website makes. This hop never leaves Tailscale — the replica is reached on
* its CGNAT tailnet address and the tailnet ACL admits only this host — so
* WireGuard already authenticates the peer. The DreamHost leg crosses the
* public internet and therefore pins the CA instead. The client here is
* MariaDB's, which rejects our self-signed CA outright unless it is handed the
* CA file, which would mean shipping a cert into this image for a link that is
* already authenticated.
*/
private async onReplica(sql: string, extra: string[] = []): Promise<string> {
const host = process.env.REPLICA_DB_HOST!;
const user = process.env.REPLICA_DB_USER!;
const password = process.env.REPLICA_DB_PASS!;
const { stdout } = await exec(
"mysql",
[
`--host=${host}`,
`--user=${user}`,
"--ssl",
"--ssl-verify-server-cert=0",
"--connect-timeout=5",
...extra,
"-e",
sql,
],
{ env: { ...process.env, MYSQL_PWD: password }, timeout: VERIFY_TIMEOUT_MS },
);
return stdout;
}
/**
* Run a statement on the master, using the application's own DATABASE_URL.
*
* The app credential is enough here on purpose — everything this class sends
* to the master is a SELECT against `information_schema` or a system variable.
* Reaching for OPS_DB_ADMIN_* the way OpsService does would hand a monitoring
* read path a credential that can also restore a dump.
*/
private async onMaster(sql: string, extra: string[] = []): Promise<string> {
const raw = process.env.DATABASE_URL;
if (!raw) throw new Error("DATABASE_URL no está configurada");
const u = new URL(raw);
const { stdout } = await exec(
"mysql",
[
`--host=${u.hostname}`,
`--port=${u.port || "3306"}`,
`--user=${decodeURIComponent(u.username)}`,
"--connect-timeout=5",
...extra,
"-N",
"-e",
sql,
u.pathname.replace(/^\//, ""),
],
{
env: { ...process.env, MYSQL_PWD: decodeURIComponent(u.password) },
timeout: VERIFY_TIMEOUT_MS,
},
);
return stdout;
}
}
/** Full scans on a 1-vCPU replica are not fast; 15s would cut them off. */
const VERIFY_TIMEOUT_MS = 120_000;
/** Parse the `t\tn\tc` rows the fingerprint query emits under `mysql -N`. */
function parseFingerprints(raw: string): Map<string, { rows: number; checksum: string }> {
const out = new Map<string, { rows: number; checksum: string }>();
for (const line of raw.split("\n")) {
const [table, n, c] = line.trim().split("\t");
if (!table || n === undefined || c === undefined) continue;
const rows = Number(n);
if (!Number.isFinite(rows)) continue;
// The checksum stays a string. Sums of CRC32 over 40k rows exceed 2^53, so
// parsing it as a number would round and make distinct tables compare equal.
out.set(table, { rows, checksum: c });
}
return out;
}
/**
* Count the transactions in a GTID set.
*
* Exported for testing. The format is `uuid[:tag]:interval[:interval]...`,
* comma-separated, where an interval is `N` or `N-M` inclusive at both ends —
* so `1-5` is five transactions, not four.
*
* MySQL 8.4 added an optional alphanumeric tag between the UUID and the first
* interval. It is skipped rather than parsed: any segment that is not a number
* or a number range is not an interval, whatever else it may be.
*/
export function countGtids(set: string): number {
if (!set.trim()) return 0;
let total = 0;
for (const group of set.split(",")) {
for (const part of group.trim().split(":").slice(1)) {
const m = /^(\d+)(?:-(\d+))?$/.exec(part.trim());
if (!m) continue;
const from = Number(m[1]);
const to = m[2] === undefined ? from : Number(m[2]);
if (Number.isFinite(from) && Number.isFinite(to) && to >= from) total += to - from + 1;
}
}
return total;
}
/**
* Narrow a GTID set to the intervals belonging to one server UUID.
*
* Exported for testing. Used to isolate the replica's own writes from the
* history it replicated, which are interleaved in the same `gtid_executed`.
*/
export function gtidsForUuid(set: string, uuid: string): string {
if (!uuid.trim()) return "";
const wanted = uuid.trim().toLowerCase();
return set
.split(",")
.map((g) => g.trim())
.filter((g) => g.toLowerCase().startsWith(`${wanted}:`))
.join(",");
}
/**
* Derive relay-apply progress from `SHOW REPLICA STATUS\G` output.
*
* Exported for testing. Free in query terms — it re-reads four more fields from
* the output the caller already has, with no second round trip to the replica
* and no connection to the source.
*
* @returns null when either position is missing or unparseable, which is what
* happens on a server that is not a replica at all.
*/
export function applyProgress(raw: string): ApplyProgress | null {
const num = (name: string): number | null => {
const v = replicaField(raw, name);
if (v === null || v === "NULL") return null;
const n = Number(v);
return Number.isFinite(n) ? n : null;
};
const readPos = num("Read_Source_Log_Pos");
const execPos = num("Exec_Source_Log_Pos");
if (readPos === null || execPos === null) return null;
const sourceLogFile = replicaField(raw, "Source_Log_File");
const relayLogFile = replicaField(raw, "Relay_Source_Log_File");
const sameFile =
sourceLogFile !== null && relayLogFile !== null && sourceLogFile === relayLogFile;
// Clamped at 0: the SQL thread cannot be ahead of the I/O thread, but the two
// fields are sampled independently, so a rotation racing this read can print
// a momentarily negative delta. Zero is the honest floor, not a bug.
const backlogBytes = sameFile ? Math.max(0, readPos - execPos) : null;
let percent: number | null = null;
if (backlogBytes !== null && readPos > 0) {
// Truncate rather than round, and hold short of 100 while bytes remain —
// see the doc on ApplyProgress.percent.
const p = Math.floor((execPos / readPos) * 10_000) / 100;
percent = backlogBytes === 0 ? 100 : Math.min(p, 99.99);
}
return { sourceLogFile, readPos, relayLogFile, execPos, sameFile, backlogBytes, percent };
}
/**
* Read one field out of `SHOW REPLICA STATUS\G` output.
*
* Exported for testing, and worth testing: the obvious regex is wrong.
* `\s` matches newlines in JavaScript, so `^\s*NAME:\s*(.*)$` lets the `\s*`
* after the colon swallow the line break of an EMPTY field and capture the
* following line instead. Last_SQL_Error is empty on a healthy replica, so that
* version reported the next line ("Replicate_Ignore_Server_Ids:") as a SQL
* error and rendered a perfectly healthy replica as broken.
*
* Hence `[^\S\n]` — horizontal whitespace only — on both sides of the name.
*
* @returns the trimmed value, or null when the field is absent OR empty. Empty
* and absent mean the same thing to every caller here: MySQL prints
* error fields as blank rather than omitting them.
*/
export function replicaField(raw: string, name: string): string | null {
const m = raw.match(new RegExp(`^[^\\S\\n]*${name}:[^\\S\\n]*(.*)$`, "m"));
const v = m?.[1]?.trim();
return v === undefined || v === "" ? null : v;
}
+254
View File
@@ -0,0 +1,254 @@
import {
applyProgress,
countGtids,
gtidsForUuid,
replicaField,
} from "./replication.service";
/**
* Verbatim shape of `SHOW REPLICA STATUS\G` from the live replica, trimmed to
* the fields the panel reads plus the neighbours that matter.
*
* The empty `Last_SQL_Error:` immediately followed by
* `Replicate_Ignore_Server_Ids:` is the whole point of the fixture — that exact
* adjacency is what the first implementation misread.
*/
const HEALTHY = [
"*************************** 1. row ***************************",
" Replica_IO_State: Waiting for source to send event",
" Source_Host: 100.103.77.46",
" Source_User: repl",
" Source_Log_File: binlog.000042",
" Read_Source_Log_Pos: 194884231",
" Relay_Source_Log_File: binlog.000042",
" Exec_Source_Log_Pos: 194884231",
" Replica_IO_Running: Yes",
" Replica_SQL_Running: Yes",
" Replicate_Do_DB: ",
" Last_Errno: 0",
" Last_Error: ",
" Seconds_Behind_Source: 0",
" Last_IO_Errno: 0",
" Last_IO_Error: ",
" Last_SQL_Errno: 0",
" Last_SQL_Error: ",
" Replicate_Ignore_Server_Ids: ",
" Source_Server_Id: 1",
].join("\n");
const BROKEN = [
" Replica_IO_Running: Yes",
" Replica_SQL_Running: No",
" Seconds_Behind_Source: NULL",
" Last_IO_Error: ",
" Last_SQL_Error: Could not execute Write_rows event on table jorgecuadros.customers",
" Replicate_Ignore_Server_Ids: ",
].join("\n");
describe("replicaField", () => {
it("reads plain values", () => {
expect(replicaField(HEALTHY, "Replica_IO_Running")).toBe("Yes");
expect(replicaField(HEALTHY, "Replica_SQL_Running")).toBe("Yes");
expect(replicaField(HEALTHY, "Source_Host")).toBe("100.103.77.46");
expect(replicaField(HEALTHY, "Seconds_Behind_Source")).toBe("0");
});
/**
* The regression this file exists for. `\s` matches newlines in JavaScript,
* so `^\s*NAME:\s*(.*)$` walks past an empty field's line break and captures
* the NEXT line — turning a healthy replica into
* "Error SQL: Replicate_Ignore_Server_Ids:" in the admin panel.
*/
it("returns null for an empty field instead of the following line", () => {
expect(replicaField(HEALTHY, "Last_SQL_Error")).toBeNull();
expect(replicaField(HEALTHY, "Last_IO_Error")).toBeNull();
expect(replicaField(HEALTHY, "Last_Error")).toBeNull();
expect(replicaField(HEALTHY, "Replicate_Do_DB")).toBeNull();
expect(replicaField(HEALTHY, "Replicate_Ignore_Server_Ids")).toBeNull();
});
it("still reads a real error when there is one", () => {
expect(replicaField(BROKEN, "Last_SQL_Error")).toBe(
"Could not execute Write_rows event on table jorgecuadros.customers",
);
expect(replicaField(BROKEN, "Replica_SQL_Running")).toBe("No");
});
/** NULL is a distinct state from empty and must survive as the literal. */
it("preserves the literal NULL that MySQL prints for unknown lag", () => {
expect(replicaField(BROKEN, "Seconds_Behind_Source")).toBe("NULL");
});
it("returns null for a field that is not present at all", () => {
expect(replicaField(HEALTHY, "Nonexistent_Field")).toBeNull();
});
/**
* Field names are matched at the start of a line. Without the line anchor,
* "Last_Error" would also match inside "Last_SQL_Error" and read the wrong
* value — the two carry different things and both feed the panel.
*/
it("does not match a field name that is a suffix of another", () => {
const raw = " Last_SQL_Error: boom\n Last_Error: ";
expect(replicaField(raw, "Last_Error")).toBeNull();
expect(replicaField(raw, "Last_SQL_Error")).toBe("boom");
});
});
/** Builds the four position fields the apply-progress reader cares about. */
function positions(
sourceFile: string,
readPos: number | string,
relayFile: string,
execPos: number | string,
): string {
return [
` Source_Log_File: ${sourceFile}`,
` Read_Source_Log_Pos: ${readPos}`,
` Relay_Source_Log_File: ${relayFile}`,
` Exec_Source_Log_Pos: ${execPos}`,
].join("\n");
}
describe("applyProgress", () => {
it("reports zero backlog and 100% when both positions match", () => {
const p = applyProgress(HEALTHY)!;
expect(p.sameFile).toBe(true);
expect(p.sourceLogFile).toBe("binlog.000042");
expect(p.readPos).toBe(194884231);
expect(p.execPos).toBe(194884231);
expect(p.backlogBytes).toBe(0);
expect(p.percent).toBe(100);
});
it("reports the byte delta when the SQL thread trails inside one file", () => {
const p = applyProgress(positions("binlog.000042", 2_000_000, "binlog.000042", 1_500_000))!;
expect(p.backlogBytes).toBe(500_000);
expect(p.percent).toBe(75);
});
/**
* The reason the byte delta exists at all. `Seconds_Behind_Source` holds at 0
* while the SQL thread is mid-transaction, so the backlog is the only field
* that moves — and the only one that says the replica is not caught up.
*/
it("shows a backlog even when the lag counter reads zero", () => {
const raw = [
" Seconds_Behind_Source: 0",
positions("binlog.000042", 900, "binlog.000042", 400),
].join("\n");
expect(replicaField(raw, "Seconds_Behind_Source")).toBe("0");
expect(applyProgress(raw)!.backlogBytes).toBe(500);
});
/**
* Positions restart near 4 in every new binlog file, so subtracting across
* files produces a number that is not a backlog — here it would be a large
* NEGATIVE one, which would render as "ahead of the source".
*/
it("refuses to compare positions across different binlog files", () => {
const p = applyProgress(positions("binlog.000043", 500, "binlog.000042", 194_000_000))!;
expect(p.sameFile).toBe(false);
expect(p.backlogBytes).toBeNull();
expect(p.percent).toBeNull();
expect(p.sourceLogFile).toBe("binlog.000043");
expect(p.relayLogFile).toBe("binlog.000042");
});
/**
* Percent must not round up to 100 while bytes remain: binlog positions are
* large, so a genuine backlog is a rounding error away from the whole file
* and would otherwise render as "caught up" on a replica that is not.
*/
it("stops short of 100% while any backlog remains", () => {
const p = applyProgress(positions("binlog.000042", 194_884_231, "binlog.000042", 194_884_230))!;
expect(p.backlogBytes).toBe(1);
expect(p.percent).toBe(99.99);
});
/** Sampled independently, so a rotation racing the read can invert them. */
it("clamps a momentarily negative delta to zero", () => {
const p = applyProgress(positions("binlog.000042", 400, "binlog.000042", 500))!;
expect(p.backlogBytes).toBe(0);
expect(p.percent).toBe(100);
});
it("returns null when the server is not a replica and prints no positions", () => {
expect(applyProgress("")).toBeNull();
expect(applyProgress(BROKEN)).toBeNull();
});
/** A stopped thread makes MySQL print NULL, which is not a position. */
it("returns null when a position is NULL", () => {
expect(applyProgress(positions("binlog.000042", "NULL", "binlog.000042", 400))).toBeNull();
});
});
/**
* Real GTID sets from the live pair, captured 2026-08-06. The replica's own
* server UUID (3b103283…) carries the transactions the seed dump load executed
* locally; the master's UUID (defc34e2…) carries the replicated history.
*/
const REPLICA_EXECUTED =
"3b103283-8f15-11f1-a52b-020017027b33:1-513," +
"defc34e2-8c5d-11f1-8e58-52c4c853bce8:1-525";
const REPLICA_UUID = "3b103283-8f15-11f1-a52b-020017027b33";
describe("countGtids", () => {
it("counts an inclusive range at both ends", () => {
// 1-5 is five transactions. Off-by-one here understates the gap, which is
// the direction that hides a problem.
expect(countGtids("defc34e2-8c5d-11f1-8e58-52c4c853bce8:1-5")).toBe(5);
});
it("counts a bare single transaction", () => {
expect(countGtids("defc34e2-8c5d-11f1-8e58-52c4c853bce8:7")).toBe(1);
});
it("sums several intervals under one UUID", () => {
expect(countGtids("defc34e2-8c5d-11f1-8e58-52c4c853bce8:1-5:8:10-12")).toBe(9);
});
it("sums across UUIDs, including the wrapped form MySQL prints", () => {
expect(countGtids(REPLICA_EXECUTED)).toBe(513 + 525);
// `gtid_executed` comes back wrapped once it holds more than one UUID.
expect(countGtids(REPLICA_EXECUTED.replace(",", ",\n"))).toBe(513 + 525);
});
/** An empty subtraction result is the caught-up case and must be zero. */
it("returns 0 for an empty or blank set", () => {
expect(countGtids("")).toBe(0);
expect(countGtids(" \n ")).toBe(0);
});
/**
* MySQL 8.4 allows an alphanumeric tag between the UUID and the intervals.
* It is not an interval and must not be counted as one.
*/
it("skips a tag without counting it", () => {
expect(countGtids("defc34e2-8c5d-11f1-8e58-52c4c853bce8:mytag:1-3")).toBe(3);
});
});
describe("gtidsForUuid", () => {
it("isolates the replica's own transactions from the replicated history", () => {
expect(countGtids(gtidsForUuid(REPLICA_EXECUTED, REPLICA_UUID))).toBe(513);
});
it("returns nothing for a UUID that is not in the set", () => {
expect(gtidsForUuid(REPLICA_EXECUTED, "00000000-0000-0000-0000-000000000000")).toBe("");
});
/**
* The colon matters. Without it a UUID prefix would match a longer UUID that
* merely starts the same way, and the replica's local writes would be
* over-reported.
*/
it("does not match on a bare prefix", () => {
expect(gtidsForUuid(REPLICA_EXECUTED, "3b103283")).toBe("");
});
it("returns nothing when the UUID is blank", () => {
expect(gtidsForUuid(REPLICA_EXECUTED, "")).toBe("");
});
});
+10 -1
View File
@@ -1,4 +1,4 @@
import { IsEnum, IsOptional, IsString } from "class-validator"; import { IsBoolean, IsEnum, IsOptional, IsString } from "class-validator";
import { OpsJobKind } from "@jorgecuadros/database"; import { OpsJobKind } from "@jorgecuadros/database";
export class StartJobDto { export class StartJobDto {
@@ -9,4 +9,13 @@ export class StartJobDto {
@IsOptional() @IsOptional()
@IsString() @IsString()
file?: string; file?: string;
/**
* REIMPORT only: proceed even though the rebuild deletes rows that exist only
* in the platform. Off by default, so the guard in run_all.py stops the job
* and lists what would be lost rather than the operator finding out after.
*/
@IsOptional()
@IsBoolean()
forceFull?: boolean;
} }
@@ -0,0 +1,94 @@
import { BadRequestException } from "@nestjs/common";
import { PoliciesService } from "./policies.service";
/**
* Deleting a lookup row that policies still reference used to succeed and
* silently blank the field on every one of them, because both FKs are
* `ON DELETE SET NULL` (`0000_init`). That is not a hypothetical: it is how
* the `M_EMPR` policy type disappeared from the dev database and left 5
* policies with a null `policyTypeId`, found only by querying months later.
*
* These tests pin the refusal. They drive the service with a stub client
* rather than a database because what is being asserted is the guard, not
* Prisma — and a test that needed a live MySQL would not run in CI.
*/
function serviceWith(counts: {
policies?: number;
claims?: number;
}): { service: PoliciesService; deleted: string[] } {
const deleted: string[] = [];
const prisma = {
policy: { count: async () => counts.policies ?? 0 },
claim: { count: async () => counts.claims ?? 0 },
insuranceProvider: {
findUnique: async () => ({ id: "p1", name: "ANA SEGUROS" }),
delete: async () => {
deleted.push("provider");
return { id: "p1" };
},
},
policyType: {
findUnique: async () => ({ id: "t1", name: "M_EMPR" }),
delete: async () => {
deleted.push("policyType");
return { id: "t1" };
},
},
adjuster: {
findUnique: async () => ({ id: "a1", name: "JUAN PEREZ" }),
delete: async () => {
deleted.push("adjuster");
return { id: "a1" };
},
},
};
const storage = {} as never;
return {
service: new PoliciesService(prisma as never, storage),
deleted,
};
}
describe("lookup deletes refuse while the row is in use", () => {
it("refuses a policy type that policies still carry, and names the count", () => {
const { service, deleted } = serviceWith({ policies: 5 });
return service.removePolicyType("t1").then(
() => {
throw new Error("expected the delete to be refused");
},
(err: unknown) => {
expect(err).toBeInstanceOf(BadRequestException);
// The operator has to be told WHICH row and HOW MANY, or the message
// is not actionable.
expect((err as Error).message).toContain("M_EMPR");
expect((err as Error).message).toContain("5");
expect(deleted).toEqual([]);
},
);
});
it("refuses a carrier that policies still carry", async () => {
const { service, deleted } = serviceWith({ policies: 738 });
await expect(service.removeProvider("p1")).rejects.toBeInstanceOf(
BadRequestException,
);
expect(deleted).toEqual([]);
});
it("refuses an adjuster still assigned to claims", async () => {
// Same `ON DELETE SET NULL` trap, on `claims.adjusterId`.
const { service, deleted } = serviceWith({ claims: 2 });
await expect(service.removeAdjuster("a1")).rejects.toBeInstanceOf(
BadRequestException,
);
expect(deleted).toEqual([]);
});
it("allows the delete once nothing references the row", async () => {
const { service, deleted } = serviceWith({ policies: 0, claims: 0 });
await service.removePolicyType("t1");
await service.removeProvider("p1");
await service.removeAdjuster("a1");
expect(deleted).toEqual(["policyType", "provider", "adjuster"]);
});
});
+43 -1
View File
@@ -8,9 +8,15 @@ import {
Post, Post,
Query, Query,
Req, Req,
Res,
StreamableFile,
UploadedFile,
UseGuards, UseGuards,
UseInterceptors,
} from "@nestjs/common"; } from "@nestjs/common";
import { Request } from "express"; import { FileInterceptor } from "@nestjs/platform-express";
import { Request, Response } from "express";
import { downloadName, type UploadedFileLike } from "../storage/upload-file";
import { AuthenticatedGuard } from "../auth/authenticated.guard"; import { AuthenticatedGuard } from "../auth/authenticated.guard";
import { AbilityGuard } from "../auth/ability.guard"; import { AbilityGuard } from "../auth/ability.guard";
import { RequireAbility } from "../auth/require-ability.decorator"; import { RequireAbility } from "../auth/require-ability.decorator";
@@ -240,4 +246,40 @@ export class PoliciesController {
removeClaim(@Param("id") id: string, @Param("childId") childId: string) { removeClaim(@Param("id") id: string, @Param("childId") childId: string) {
return this.policies.removeClaim(id, childId); return this.policies.removeClaim(id, childId);
} }
// --- documents ------------------------------------------------------------
@Post(":id/documents")
@RequireAbility("policy:update")
@UseInterceptors(
FileInterceptor("file", { limits: { fileSize: 50 * 1024 * 1024 } }),
)
addDocument(
@Param("id") id: string,
@UploadedFile() file: UploadedFileLike | undefined,
@Query("type") type: string | undefined,
) {
if (!file) throw new Error("No se recibió ningún archivo.");
return this.policies.addDocument(id, file, type);
}
@Get(":id/documents/:childId/download")
async downloadDocument(
@Param("id") id: string,
@Param("childId") childId: string,
@Res({ passthrough: true }) res: Response,
): Promise<StreamableFile> {
const { row, stream, contentType } = await this.policies.getDocument(id, childId);
res.set({
"Content-Type": contentType ?? "application/octet-stream",
"Content-Disposition": `attachment; filename="${downloadName(row.storageKey, row.documentType)}"`,
});
return new StreamableFile(stream);
}
@Delete(":id/documents/:childId")
@RequireAbility("policy:update")
removeDocument(@Param("id") id: string, @Param("childId") childId: string) {
return this.policies.removeDocument(id, childId);
}
} }
+97 -5
View File
@@ -1,6 +1,9 @@
import { Injectable, NotFoundException } from "@nestjs/common"; import { BadRequestException, Injectable, NotFoundException } from "@nestjs/common";
import { randomUUID } from "node:crypto";
import { Prisma } from "@jorgecuadros/database"; import { Prisma } from "@jorgecuadros/database";
import { PrismaService } from "../prisma/prisma.service"; import { PrismaService } from "../prisma/prisma.service";
import { StorageService } from "../storage/storage.service";
import { extForUpload, type UploadedFileLike } from "../storage/upload-file";
import { toDate } from "../common/coerce"; import { toDate } from "../common/coerce";
import { CreatePolicyDto, UpdatePolicyDto } from "./policy.dto"; import { CreatePolicyDto, UpdatePolicyDto } from "./policy.dto";
import { import {
@@ -77,7 +80,10 @@ function daysUntil(policyTo: Date | null, from: Date): number | null {
@Injectable() @Injectable()
export class PoliciesService { export class PoliciesService {
constructor(private readonly prisma: PrismaService) {} constructor(
private readonly prisma: PrismaService,
private readonly storage: StorageService,
) {}
private statusWhere( private statusWhere(
status: PolicyStatus | undefined, status: PolicyStatus | undefined,
@@ -352,6 +358,7 @@ export class PoliciesService {
return this.prisma.policy.update({ where: { id }, data: { archivedAt: null } }); return this.prisma.policy.update({ where: { id }, data: { archivedAt: null } });
} }
private async ensurePolicy(id: string) { private async ensurePolicy(id: string) {
const found = await this.prisma.policy.findUnique({ const found = await this.prisma.policy.findUnique({
where: { id }, where: { id },
@@ -479,6 +486,47 @@ export class PoliciesService {
}; };
} }
// --- documents ------------------------------------------------------------
// Blob in object storage under `policy/<policyId>/…`; row is the pointer.
async addDocument(
policyId: string,
file: UploadedFileLike,
documentType?: string,
) {
await this.ensurePolicy(policyId);
const key = `policy/${policyId}/${randomUUID()}${extForUpload(file)}`;
await this.storage.put(key, file.buffer, file.mimetype);
return this.prisma.policyDocument.create({
data: {
policyId,
documentType: documentType?.trim() || "DOCUMENT",
storageKey: key,
},
});
}
async getDocument(policyId: string, id: string) {
const row = await this.prisma.policyDocument.findFirst({
where: { id, policyId },
});
if (!row) throw new NotFoundException(`Document ${id} not found on policy ${policyId}`);
const blob = await this.storage.getStream(row.storageKey);
return { row, ...blob };
}
async removeDocument(policyId: string, id: string) {
await this.ensurePolicy(policyId);
const row = await this.prisma.policyDocument.findFirst({
where: { id, policyId },
select: { id: true, storageKey: true },
});
if (!row) throw new NotFoundException(`Document ${id} not found on policy ${policyId}`);
const deleted = await this.prisma.policyDocument.delete({ where: { id } });
await this.storage.delete(row.storageKey);
return deleted;
}
// --- lookups (providers / policy types / adjusters) ----------------------- // --- lookups (providers / policy types / adjusters) -----------------------
listLookups() { listLookups() {
@@ -506,7 +554,40 @@ export class PoliciesService {
updateProvider(id: string, dto: UpdateProviderDto) { updateProvider(id: string, dto: UpdateProviderDto) {
return this.prisma.insuranceProvider.update({ where: { id }, data: dto }); return this.prisma.insuranceProvider.update({ where: { id }, data: dto });
} }
removeProvider(id: string) { /**
* Deleting a lookup row that policies still point at is silent data loss.
*
* Both FKs are `ON DELETE SET NULL` (see `0000_init`), so the delete
* succeeds, returns 200, and blanks the field on every policy that used it
* — with no error and nothing in the UI to suggest anything happened. That
* is how the `M_EMPR` policy type disappeared and left 5 policies with a
* null `policyTypeId`, only found later by querying.
*
* Refusing is the whole fix. There is no "are you sure": the operator
* reassigns those policies first, which is work the app cannot do for them
* because only they know which type is correct.
*/
private async assertLookupUnused(
kind: "provider" | "policyType",
id: string,
): Promise<void> {
const where = kind === "provider" ? { insuranceProviderId: id } : { policyTypeId: id };
const count = await this.prisma.policy.count({ where });
if (count === 0) return;
const label =
kind === "provider"
? (await this.prisma.insuranceProvider.findUnique({ where: { id } }))?.name
: (await this.prisma.policyType.findUnique({ where: { id } }))?.name;
const noun = kind === "provider" ? "La aseguradora" : "El tipo de póliza";
throw new BadRequestException(
`${noun} «${label ?? id}» está en uso por ${count} póliza(s). ` +
"Reasígnelas antes de eliminarlo.",
);
}
async removeProvider(id: string) {
await this.assertLookupUnused("provider", id);
return this.prisma.insuranceProvider.delete({ where: { id } }); return this.prisma.insuranceProvider.delete({ where: { id } });
} }
@@ -516,7 +597,8 @@ export class PoliciesService {
updatePolicyType(id: string, dto: UpdatePolicyTypeDto) { updatePolicyType(id: string, dto: UpdatePolicyTypeDto) {
return this.prisma.policyType.update({ where: { id }, data: dto }); return this.prisma.policyType.update({ where: { id }, data: dto });
} }
removePolicyType(id: string) { async removePolicyType(id: string) {
await this.assertLookupUnused("policyType", id);
return this.prisma.policyType.delete({ where: { id } }); return this.prisma.policyType.delete({ where: { id } });
} }
@@ -526,7 +608,17 @@ export class PoliciesService {
updateAdjuster(id: string, dto: UpdateAdjusterDto) { updateAdjuster(id: string, dto: UpdateAdjusterDto) {
return this.prisma.adjuster.update({ where: { id }, data: dto }); return this.prisma.adjuster.update({ where: { id }, data: dto });
} }
removeAdjuster(id: string) { /** Same `ON DELETE SET NULL` trap as the two above, on `claims.adjusterId`:
* deleting a busy adjuster would quietly strip them off their claims. */
async removeAdjuster(id: string) {
const count = await this.prisma.claim.count({ where: { adjusterId: id } });
if (count > 0) {
const row = await this.prisma.adjuster.findUnique({ where: { id } });
throw new BadRequestException(
`El ajustador «${row?.name ?? id}» está asignado a ${count} siniestro(s). ` +
"Reasígnelos antes de eliminarlo.",
);
}
return this.prisma.adjuster.delete({ where: { id } }); return this.prisma.adjuster.delete({ where: { id } });
} }
} }
@@ -0,0 +1,161 @@
import {
nameTokens,
suggestCustomersByName,
suggestionNote,
type CustomerNameRow,
} from "./name-matcher";
/**
* Every row here is a real name out of the customer book (1536 rows, dev
* mirror of production), chosen because it is one of the shapes that breaks
* naive matching: surname-first ordering, a middle initial, a Spanish double
* surname, a joint account, a missing comma, and the `(SIN NOMBRE)`
* placeholder the migration left for customers whose DATGRAL row had no name.
*/
const BOOK: CustomerNameRow[] = [
{ id: "c1", name: "WAGONER, PAMELA" },
{ id: "c2", name: "MCWILLIAMS, BRIAN MICHAEL" },
{ id: "c3", name: "MCWILLIAMS, BRIAN" },
{ id: "c4", name: "WEAKLAND, RICHARD E." },
{ id: "c5", name: "ESTRADA, JERRY & MARILYN" },
{ id: "c6", name: "CABALLERO PRIETO, GUILLERMO" },
{ id: "c7", name: "GREENE STEPHANIE" },
{ id: "c8", name: "(SIN NOMBRE)" },
{ id: "c9", name: "MUÑOZ, LUIS ALBERTO" },
{ id: "c10", name: "SMITH, DANIEL" },
{ id: "c11", name: "SMITH, JOHN" },
];
describe("nameTokens", () => {
it("makes the two orderings the same set", () => {
expect(nameTokens("PAMELA WAGONER").sort()).toEqual(
nameTokens("WAGONER, PAMELA").sort(),
);
});
it("drops initials, particles and corporate suffixes", () => {
expect(nameTokens("WEAKLAND, RICHARD E.")).toEqual(["WEAKLAND", "RICHARD"]);
expect(nameTokens("GARCIA DE LA TORRE, ANA")).toEqual(["GARCIA", "TORRE", "ANA"]);
expect(nameTokens("CONSTRUCTORA BAJA S.A. DE C.V.")).toEqual([
"CONSTRUCTORA",
"BAJA",
]);
});
it("folds accents so OCR's MUNOZ reaches the book's MUÑOZ", () => {
expect(nameTokens("MUÑOZ")).toEqual(["MUNOZ"]);
});
it("drops the phone number ANA prints against the insured name", () => {
// Observed verbatim from the ANA automobile face.
expect(nameTokens("MARIA GARCIA Ph.3102001538")).toEqual([
"MARIA",
"GARCIA",
"PH",
]);
});
});
describe("suggestCustomersByName", () => {
it("matches the reversed name exactly", () => {
const [top] = suggestCustomersByName("PAMELA WAGONER", BOOK);
expect(top).toMatchObject({ customerId: "c1", tier: "EXACT", score: 1 });
});
it("treats a printed middle name the book lacks as a partial hit", () => {
const hits = suggestCustomersByName("PAMELA DENISE WAGONER", BOOK);
expect(hits[0]).toMatchObject({ customerId: "c1", tier: "PARTIAL" });
expect(hits[0].score).toBeCloseTo(2 / 3);
});
it("ranks the exact row above the row that merely contains it", () => {
// Both MCWILLIAMS rows are reachable from this name; the one that holds
// the middle name is the exact set and must come first.
const hits = suggestCustomersByName("BRIAN MICHAEL MCWILLIAMS", BOOK);
expect(hits.map((h) => h.customerId)).toEqual(["c2", "c3"]);
expect(hits[0].tier).toBe("EXACT");
expect(hits[1].tier).toBe("PARTIAL");
});
it("reaches a joint account from the one spouse the carrier printed", () => {
const hits = suggestCustomersByName("JERRY ESTRADA", BOOK);
expect(hits[0]).toMatchObject({ customerId: "c5", tier: "PARTIAL" });
});
it("will not reach a joint account on given names alone", () => {
// No surname printed: `JERRY MARILYN` overlaps ESTRADA, JERRY & MARILYN
// on two tokens, and matching on that would book a stranger's policy.
expect(suggestCustomersByName("JERRY MARILYN", BOOK)).toEqual([]);
});
it("matches a Spanish double surname regardless of where the comma fell", () => {
const [top] = suggestCustomersByName("GUILLERMO CABALLERO PRIETO", BOOK);
expect(top).toMatchObject({ customerId: "c6", tier: "EXACT" });
});
it("still matches a book row that has no comma", () => {
const [top] = suggestCustomersByName("STEPHANIE GREENE", BOOK);
expect(top).toMatchObject({ customerId: "c7", tier: "EXACT" });
});
it("never suggests the (SIN NOMBRE) placeholder", () => {
expect(suggestCustomersByName("SIN NOMBRE", BOOK)).toEqual([]);
expect(suggestCustomersByName("NOMBRE DEL ASEGURADO", BOOK)).toEqual([]);
});
it("returns nothing on a shared surname alone", () => {
// 185 surnames are shared by 524 customers; one token is not evidence.
expect(suggestCustomersByName("SMITH", BOOK)).toEqual([]);
});
it("returns nothing for a different person with the same surname", () => {
expect(suggestCustomersByName("ROBERT SMITH", BOOK)).toEqual([]);
});
it("refuses a page-sized blob", () => {
// GMX's especificación has no field labels and the parser has handed its
// whole first page over as the insured name.
const blob =
"ESPECIFICACION DE LA POLIZA DE SEGURO DE RESPONSABILIDAD CIVIL " +
"EXPEDIDA A FAVOR DE PAMELA WAGONER CON VIGENCIA DEL 01 DE ENERO";
expect(suggestCustomersByName(blob, BOOK)).toEqual([]);
});
it("caps the list", () => {
expect(suggestCustomersByName("BRIAN MICHAEL MCWILLIAMS", BOOK, 1)).toHaveLength(1);
});
it("handles a null insured name", () => {
expect(suggestCustomersByName(null, BOOK)).toEqual([]);
});
});
describe("suggestionNote", () => {
it("says nothing when there is nothing", () => {
expect(suggestionNote([])).toBeNull();
});
it("names a single exact hit", () => {
expect(suggestionNote(suggestCustomersByName("PAMELA WAGONER", BOOK))).toBe(
"posible cliente por nombre: WAGONER, PAMELA",
);
});
it("reports a tie rather than picking one", () => {
// The book really does hold EMERY, LAURA twice and KIRCHHOFF, CINDY
// three times.
const dupes: CustomerNameRow[] = [
{ id: "d1", name: "EMERY, LAURA" },
{ id: "d2", name: "EMERY, LAURA" },
];
expect(suggestionNote(suggestCustomersByName("LAURA EMERY", dupes))).toBe(
"2 clientes tienen ese mismo nombre; elija cuál",
);
});
it("lists partial hits", () => {
expect(suggestionNote(suggestCustomersByName("PAMELA DENISE WAGONER", BOOK))).toBe(
"posibles clientes por nombre: WAGONER, PAMELA",
);
});
});
+199
View File
@@ -0,0 +1,199 @@
/**
* Suggests which existing customer a printed insured name belongs to.
*
* The office books customers surname-first ("WAGONER, PAMELA") and carriers
* print them given-name-first ("PAMELA DENISE WAGONER"), so a string compare
* never hits. Comparing *token sets* does, and it is order-insensitive by
* construction — which is the whole trick.
*
* **These are suggestions, never matches.** Nothing here sets
* `matchedCustomerId` or `confident`; the review screen offers the ranked
* names and a human picks. That line is not caution, it is what the book
* measures out to: of 1536 customers, 1487 have a distinct normalized token
* set — but loosen the rule to surname + first given name only and 131 of
* them (8.5%) collide, because the book holds `MCWILLIAMS, BRIAN MICHAEL`
* *and* `MCWILLIAMS, BRIAN`, and `CUADROS, JORGE JR` alongside three
* `CUADROS, JORGE H.`. 185 surnames are shared by 524 customers, so a
* surname alone carries no information at all.
*
* The two tiers below are drawn at the two places that measurement puts a
* cliff: full token-set equality, where cross-person collisions are
* effectively zero, and strict containment, where they are common enough
* that the result can only ever be a hint.
*/
/** A customer row as the matcher needs it — id and the book's name. */
export interface CustomerNameRow {
id: string;
name: string;
}
export type NameMatchTier = "EXACT" | "PARTIAL";
export interface CustomerNameSuggestion {
customerId: string;
customerName: string;
/**
* `EXACT` — the two names carry the same tokens, in any order.
* `PARTIAL` — one name's tokens are all present in the other's, plus the
* surname. A printed middle name the book does not hold, or a joint
* account where the carrier named one spouse, both land here.
*/
tier: NameMatchTier;
/** Shared tokens over the longer name's token count, 0..1. */
score: number;
}
/**
* Words that carry no identity. Spanish particles and the ampersand joining
* a couple are noise; the corporate suffixes are dropped so `S.A. DE C.V.`
* does not make every company look alike.
*/
const NOISE = new Set([
"DE", "DEL", "LA", "LAS", "LOS", "Y", "AND", "VDA",
"JR", "SR", "II", "III", "IV",
"SA", "CV", "SAPI", "SRL", "RL", "SC", "INC", "LLC", "LTD", "CORP", "CO",
]);
/**
* Placeholder rows the migration left behind. Fourteen customers are named
* literally `(SIN NOMBRE)`; without this they would be one 14-way tie on
* every unreadable name.
*/
const PLACEHOLDER = new Set(["SIN NOMBRE", "NOMBRE SIN"]);
/**
* A name blob longer than this is not a name. GMX's PVL especificación has
* no field labels, and the parser has been seen handing its entire first
* page over as `insuredName`; matching that against the book would find
* a surname somewhere in the prose and suggest a stranger.
*/
const MAX_TOKENS = 8;
const MAX_CHARS = 80;
/**
* Splits a name into comparable tokens.
*
* Accents go first, and deliberately in both directions: the book holds
* `MUÑOZ` where OCR routinely reads `MUNOZ`, and folding both to the same
* ASCII makes that a hit rather than a miss.
*
* Tokens containing digits are dropped outright. ANA's automobile face
* prints the phone number hard against the insured name — the parser has
* emitted `MARIA GARCIA Ph.3102001538` — and the digits would otherwise
* be an extra token forever blocking `EXACT`.
*
* Single letters are dropped as initials: the book is full of
* `WEAKLAND, RICHARD E.`, and a carrier that prints the middle name in
* full should still match the row that abbreviates it.
*/
export function nameTokens(raw: string): string[] {
const cleaned = raw
.normalize("NFD")
.replace(/[\u0300-\u036f]/g, "")
.toUpperCase()
.replace(/[^A-Z0-9]+/g, " ")
.trim();
const tokens = cleaned
.split(" ")
.filter((t) => t.length > 1 && !/\d/.test(t) && !NOISE.has(t));
return [...new Set(tokens)];
}
/** The surname tokens — everything before the comma the book writes. */
function surnameTokens(bookName: string): string[] {
const comma = bookName.indexOf(",");
// 54 of 1536 rows have no comma at all ("GREENE STEPHANIE",
// "FAROOQ VAKIL"), and which half is the surname is unknowable. Requiring
// a surname we cannot identify would silently exclude those rows, so they
// fall back to requiring nothing beyond the containment rule.
if (comma < 0) return [];
return nameTokens(bookName.slice(0, comma));
}
function isPlaceholder(tokens: string[]): boolean {
return tokens.length === 0 || PLACEHOLDER.has([...tokens].sort().join(" "));
}
function containsAll(haystack: Set<string>, needles: string[]): boolean {
return needles.every((n) => haystack.has(n));
}
/**
* Ranks the book against one printed name.
*
* Returns at most `limit` suggestions, `EXACT` before `PARTIAL` and higher
* score first. An empty array means the printed name was unusable (too
* long, too few real tokens) or nothing in the book came close — both of
* which leave the review screen exactly as it is today.
*/
export function suggestCustomersByName(
printedName: string | null | undefined,
customers: CustomerNameRow[],
limit = 3,
): CustomerNameSuggestion[] {
if (!printedName || printedName.length > MAX_CHARS) return [];
const printed = nameTokens(printedName);
// One usable token is a surname or a given name on its own, and 34% of the
// book shares a surname with someone. Nothing useful can come of it.
if (printed.length < 2 || printed.length > MAX_TOKENS) return [];
const printedSet = new Set(printed);
const out: CustomerNameSuggestion[] = [];
for (const c of customers) {
const book = nameTokens(c.name);
if (isPlaceholder(book) || book.length < 2) continue;
const bookSet = new Set(book);
const overlap = printed.filter((t) => bookSet.has(t)).length;
// Two shared tokens is the floor: one is a bare surname collision.
if (overlap < 2) continue;
const bookInPrinted = containsAll(printedSet, book);
const printedInBook = containsAll(bookSet, printed);
if (!bookInPrinted && !printedInBook) continue;
// When the book's name is the shorter one, containment already proves
// the surname was printed. When the printed name is shorter — the book
// holds a middle name or a second spouse the carrier omitted — the
// surname must be there explicitly, or `JERRY MARILYN` would match
// `ESTRADA, JERRY & MARILYN` on given names alone.
if (!bookInPrinted && !containsAll(printedSet, surnameTokens(c.name))) continue;
out.push({
customerId: c.id,
customerName: c.name,
tier: bookInPrinted && printedInBook ? "EXACT" : "PARTIAL",
score: overlap / Math.max(book.length, printed.length),
});
}
out.sort((a, b) => {
if (a.tier !== b.tier) return a.tier === "EXACT" ? -1 : 1;
if (b.score !== a.score) return b.score - a.score;
return a.customerName.localeCompare(b.customerName);
});
return out.slice(0, limit);
}
/** Review-queue wording for what the suggestions amount to. */
export function suggestionNote(suggestions: CustomerNameSuggestion[]): string | null {
if (suggestions.length === 0) return null;
const exact = suggestions.filter((s) => s.tier === "EXACT");
// More than one exact hit is the duplicate-customer case the book really
// has (`EMERY, LAURA` twice, `KIRCHHOFF, CINDY` three times). Saying so is
// more useful than naming whichever one sorted first.
if (exact.length > 1) {
return `${exact.length} clientes tienen ese mismo nombre; elija cuál`;
}
if (exact.length === 1) {
return `posible cliente por nombre: ${exact[0].customerName}`;
}
return `posibles clientes por nombre: ${suggestions.map((s) => s.customerName).join(", ")}`;
}
@@ -0,0 +1,919 @@
import type { OcrPage } from "../../statements/ocr/ocr.provider";
import {
detectPolicyProvider,
parsePolicy,
type ParsedCoverage,
} from "./policy-parser";
/**
* Verbatim excerpts of what the GMX portal's translation PDF actually
* rendered through pdftotext — same convention as the statement parser
* tests, where invented-clean input would test nothing because clean input
* is not the failure mode.
*/
function page(text: string): OcrPage {
return { text, words: [], confidence: 0.95 };
}
/** Coverages keyed by their risk label, so an assertion names the coverage
* it is about instead of an array index that shifts when one is added. */
const byRisk = (p: ReturnType<typeof parsePolicy>): Record<string, ParsedCoverage> =>
Object.fromEntries(p.coverages.map((c) => [c.risk, c]));
describe("detectPolicyProvider", () => {
it("claims GMX from the brand wordmark on the letterhead", () => {
expect(
detectPolicyProvider(
"Grupo Mexicano de Seguros, S.A. de C.V.\nTecoyotitla 412, Edificio GMX",
),
).toBe("GMX");
});
it("claims GMX from the 'gmx.com.mx' footer URL", () => {
expect(detectPolicyProvider("JUNTOS EL RIESGO ES MENOR\nwww.gmx.com.mx")).toBe("GMX");
});
});
describe("parsePolicy / GMX", () => {
// Verbatim text extracted from ~/Downloads/HC_Folio_000767_Traduccion.pdf via
// `pdftotext -layout`. Two pages joined by "\n\n".
const GMX_FULL = page(
"Multiple Policy\nHome\n" +
"Policy 007-037-07005947-0000-02 in accordance with the enclosed clauses, to insurance:\n" +
"Insured JON ASHLEY STRABALA\n" +
"Additional insured VIVIAN\n" +
"Legal address BONAMPACK No. EXT26 No.INT 0 COL. Punta Bandera, Tijuana, Baja California, C.P. 22550\n" +
"ZIP 22550 Income Tax No. XEXX-010101-000\n" +
"Broker (1176) Jorge Humberto Cuadros\n" +
"Term 12 months\n" +
"From 19/07/2026\n" +
"To 19/07/2027 at twelve hours (noon) Mexico City time.\n" +
"Currency DOLARES Premium payment CONTADO\n" +
"Free translation from the Spanish Insurance contract. The English text is just copy given by courtesy. In case of a dispute, the Spanish will prevail over the English version.\n" +
"Agreed clauses:\n" +
"•The insured and GMX Hereby declared...\n" +
"From the above, the present contract shall not be considered under the condition mentioned within article 36-B from the Insurance Companies General Law. Therefore it shall not be required its registration before the Comision National de Seguros y Fianzas.\n" +
"July 23, 2026\n" +
"Authority sign.\n" +
"Grupo Mexicano de Seguros, S.A. de C.V.\n" +
"Tecoyotitla 412, Edificio GMX\n" +
"JUNTOS EL RIESGO ES MENOR\n" +
"www.gmx.com.mx\n\n" +
"Risk Insured Amount Deductible Loss Participation\n" +
"Building $350,000.00 Not applies Not applies\n" +
"Contents $60,000.00 Not applies Not applies\n" +
"ADDITIONAL RISK\n" +
"Risk Insured Amount Deductible Loss Participation\n" +
"Debris removal Building $35,000.00 Not applies Not applies\n" +
"Debris removal Contents $6,000.00 Not applies Not applies\n" +
"Outdoors Constructions $10,000.00 5% 10%\n" +
"Coverage Extention Covered Not applies Not applies\n" +
"All Risk Covered Not applies Not applies\n" +
"Earthquake and/or volcanic eruption Covered 2% of the sum insured for each damage structure 20%\n" +
"Extra Expenses $41,000.00 Not applies Not applies\n" +
"Robbery with violence $10,000.00 Not applies Not applies\n" +
"Jewerly $3,900.00 Not applies Not applies\n" +
"Electronic Equipment $10,000.00 Not applies Not applies\n" +
"Glasses $10,000.00 Not applies Not applies\n" +
"Tenant $200,000.00 Not applies Not applies\n" +
"Family $200,000.00 Not applies Not applies\n" +
"Family $200,000.00 Not applies Not applies\n" +
"Domestic workers $7,010.00 Not applies Not applies\n" +
"VALUES ADDED, HOME GMX",
);
it("extracts the policy number, insured name, broker, dates, and currency", () => {
const p = parsePolicy(GMX_FULL);
expect(p.provider).toBe("GMX");
expect(p.policyNumber).toBe("007-037-07005947-0000-02");
expect(p.insuredName).toBe("JON ASHLEY STRABALA");
expect(p.additionalInsured).toBe("VIVIAN");
expect(p.agentName).toBe("Jorge Humberto Cuadros");
expect(p.policyFrom?.toISOString().slice(0, 10)).toBe("2026-07-19");
expect(p.policyTo?.toISOString().slice(0, 10)).toBe("2027-07-19");
expect(p.policyDate?.toISOString().slice(0, 10)).toBe("2026-07-23");
expect(p.currency).toBe("USD");
expect(p.zip).toBe("22550");
expect(p.legalAddress).toContain("BONAMPACK");
expect(p.premiumPayment).toBe("CONTADO");
});
it("extracts every coverage row off the second page table", () => {
const p = parsePolicy(GMX_FULL);
const byName = Object.fromEntries(p.coverages.map((c) => [c.risk, c]));
expect(byName.Building?.insuredAmount).toBe(350000);
expect(byName.Contents?.insuredAmount).toBe(60000);
expect(byName["Debris removal Building"]?.insuredAmount).toBe(35000);
expect(byName["Outdoors Constructions"]?.insuredAmount).toBe(10000);
expect(byName["Outdoors Constructions"]?.deductible).toBe("5%");
expect(byName["Outdoors Constructions"]?.lossParticipation).toBe("10%");
// Free-text coverage cells kept verbatim (the policy form surfaces them
// as observations, not as numbers).
expect(byName["Earthquake and/or volcanic eruption"]?.insuredAmount).toBeNull();
expect(byName["Earthquake and/or volcanic eruption"]?.deductible).toContain("2%");
expect(byName["Earthquake and/or volcanic eruption"]?.lossParticipation).toBe("20%");
expect(byName["All Risk"]?.insuredAmount).toBeNull();
expect(p.coverages.length).toBeGreaterThan(10);
});
it("names the product MULT for confirm to resolve", () => {
// The caratula's own header reads "Multiple Policy / Home". MULT is the
// legacy discriminator for that multi-line home policy; INCENDIO is
// fire-only and no policy in the book has ever used it.
expect(parsePolicy(GMX_FULL).policyTypeName).toBe("MULT");
});
it("leaves premium fields null on the certificate page and notes it", () => {
const p = parsePolicy(GMX_FULL);
expect(p.netPremium).toBeNull();
expect(p.total).toBeNull();
expect(p.policyFee).toBeNull();
expect(p.notes.join(" ")).toMatch(/prima/i);
});
it("still parses when the broker parens are missing", () => {
const p = parsePolicy(
page(
"Insured JON ASHLEY STRABALA\nBroker Jorge Humberto Cuadros\n" +
"From 19/07/2026\nTo 19/07/2027\nCurrency DOLARES\n" +
"Grupo Mexicano de Seguros",
),
);
expect(p.agentName).toBe("Jorge Humberto Cuadros");
});
it("rejects a page that carries no GMX signal at all", () => {
const p = parsePolicy(page("Random unrelated document with no policy data."));
expect(p.provider).toBe("");
expect(p.notes.join(" ")).toContain("no se reconoció el proveedor");
});
it("captures the deductible / loss-participation columns verbatim as strings", () => {
const p = parsePolicy(GMX_FULL);
const eq = p.coverages.find((c) => c.risk === "Earthquake and/or volcanic eruption");
expect(eq).toBeDefined();
const eqTyped = eq as ParsedCoverage;
expect(eqTyped.deductible).toContain("sum insured");
expect(eqTyped.lossParticipation).toBe("20%");
});
});
/**
* The second GMX document family: the Spanish PVL "especificación" the office
* receives as `…-CondicionesParticulares.pdf`. Verbatim excerpts from
* `007_LGS-HGMX_07006957_01_0-CondicionesParticulares.pdf` through
* `pdftotext -layout`, indentation included — the column positions and the
* blank lines between blocks are what the parser reads, so a cleaned-up
* fixture would test nothing.
*/
describe("parsePolicy / GMX especificación (PVL Hogar)", () => {
const HEADER =
" ESPECIFICACIÓN QUE SE ADHIERE Y FORMA PARTE INTEGRANTE DE LA PÓLIZA\n" +
" 07-037-07006957-00000-01\n" +
"\n";
const GMX_ESPEC = page(
HEADER +
"\n" +
" Nombre del asegurado EMMER . KATHLEEN\n" +
"\n" +
" Tipo Persona Asegurada Propietario\n" +
"\n" +
" Ubicación del riesgo LOS PELICANOS ESTE NO. 98 Col. LAS GAVIOTAS PLAYAS\n" +
" DE ROSARITO BAJA CALIFORNIA 22713\n" +
"\n" +
" Características del Inmueble Casa Tipo constructivo Combinado: Macizo y Madera.\n" +
" Consta de 2 pisos incluyendo sótanos y planta baja.\n" +
"\n" +
" -500 mts.cuerpo agua SI\n" +
"\n" +
" Asegurado Adicional\n" +
"\n" +
"PVL Hogar - GMX Seguros Página: 1 de 10\n" +
HEADER +
"\n" +
" SECCIÓN INCENDIO EDIFICIO Y CONTENIDOS\n" +
"\n" +
" EDIFICIO\n" +
"\n" +
" Límite Máximo de Responsabilidad:\n" +
" $200,000.00 USD\n" +
"\n" +
" Quedan amparados los muros de contención y bardas, así como puertas y portones, hasta un sublimite de $ 50,000.00 M.N. o su\n" +
" equivalente en dólares americanos, o hasta el 10% de la suma asegurada de la sección de Edificio, lo que resulte menor.\n" +
"\n" +
"\n" +
" CONTENIDOS\n" +
"\n" +
" Límite Máximo de Responsabilidad:\n" +
" $20,000.00 USD\n" +
"\n" +
" 2. Terremoto o erupción volcánica: Sección Edificio EXCLUIDO, Sección Contenidos EXCLUIDO\n" +
"\n" +
" 3. Fenómenos hidrometeorológicos: Sección Edificio $200,000.00 USD, Sección Contenidos $20,000.00 USD\n" +
"\n" +
" Riesgos adicionales.\n" +
"\n" +
" Remoción de escombros\n" +
"\n" +
" Límite Máximo de Responsabilidad:\n" +
" Edificio\n" +
" $20,000.00 USD\n" +
" Contenidos\n" +
" $2,000.00 USD\n" +
"\n" +
" Gastos extraordinarios para casa habitación\n" +
"\n" +
" En caso de siniestro por los riesgos cubiertos en esta póliza, GMX Seguros pagará la renta de casa o departamento, casa de\n" +
" huéspedes u hotel cuando se asegure el inmueble, así como los gastos de mudanza, seguro de transporte del menaje de casa y\n" +
" efectuados.\n" +
"\n" +
" Límite Máximo de Responsabilidad:\n" +
" $22,000.00 USD\n" +
" Periodo de indemnización: 4 meses.\n" +
"\n" +
" Bienes a la Intemperie:\n" +
"\n" +
"\n" +
" 5 POR CIENTO SOBRE SUMA ASEGURADA, 20 PORCIENTO DE PARTICIPACIÓN A CARGO DEL ASEGURADO DE TODA\n" +
" Y CADA PÉRDIDA.\n" +
"\n" +
"\n" +
" Límite Máximo de Responsabilidad: $10,000.00 USD\n" +
"\n" +
" DEDUCIBLES:\n" +
"\n" +
" El procedimiento que se seguirá para la aplicación de deducibles en caso de que la póliza cuente con cláusula inflacionaria en todas\n" +
" y/o en algunas de sus coberturas será como sigue:\n" +
"\n" +
" Fenómenos hidrometeorológicos\n" +
" Zona: A2\n" +
" Deducible\n" +
" Edificio: 1 POR CIENTO SOBRE SUMA ASEGURADA\n" +
" Coaseguro:\n" +
" Zona 1: (INTERIOR) Participación a cargo del asegurado del 10% de toda y cada pérdida.\n" +
" Zona 2: Participación a cargo del asegurado del 10% de toda y cada pérdida.\n" +
"\n" +
" Deducible\n" +
" Contenidos: 1 POR CIENTO SOBRE SUMA ASEGURADA\n" +
"\n" +
" II.- SECCIÓN DIVERSOS MISCELÁNEOS\n" +
"\n" +
" ROBO DE CONTENIDOS\n" +
"\n" +
" Límite de Responsabilidad:\n" +
" $4,000.00 USD\n" +
"\n" +
"\n" +
" Deducible:\n" +
" Sin deducible\n" +
"\n" +
"\n" +
" Sublímites:\n" +
" Joyas, artículos de oro y plata, armas, relojes, pieles, piedras preciosas montadas, colecciones, obras de arte y demás que por su\n" +
"\n" +
"PVL Hogar - GMX Seguros Página: 7 de 10\n" +
HEADER +
"\n" +
" naturaleza se consideran como objetos de difícil o imposible reposición\n" +
"\n" +
"\n" +
" Límite de Responsabilidad:\n" +
"\n" +
" $2,000.00 USD\n" +
"\n" +
"\n" +
" Deducible:\n" +
" Sin deducible\n" +
"\n" +
" Las condiciones generales que forman parte de la presente póliza son las identificadas bajo el nombre:\n" +
" W_HogarGMX_12.11.2025.pdf\n" +
"\n" +
"PVL Hogar - GMX Seguros Página: 10 de 10\n",
);
it("reads a policy number whose groups are not the caratula's widths", () => {
// 2-3-8-5-2 here vs 3-3-8-4-2 on the English caratula. Pinning the widths
// reads one family and returns null on the other.
expect(parsePolicy(GMX_ESPEC).policyNumber).toBe("07-037-07006957-00000-01");
});
it("reads the insured, the risk location across its wrapped line, and the ZIP", () => {
const p = parsePolicy(GMX_ESPEC);
expect(p.provider).toBe("GMX");
expect(p.insuredName).toBe("EMMER . KATHLEEN");
expect(p.legalAddress).toBe(
"LOS PELICANOS ESTE NO. 98 Col. LAS GAVIOTAS PLAYAS DE ROSARITO BAJA CALIFORNIA 22713",
);
expect(p.zip).toBe("22713");
// The cell is printed but empty on this policy — an empty label must not
// capture the next line of the form.
expect(p.additionalInsured).toBeNull();
});
it("leaves the fields this document does not carry null, and says so", () => {
const p = parsePolicy(GMX_ESPEC);
expect(p.policyFrom).toBeNull();
expect(p.policyTo).toBeNull();
expect(p.policyDate).toBeNull();
expect(p.agentName).toBeNull();
expect(p.netPremium).toBeNull();
expect(p.total).toBeNull();
// The note must tell the reviewer to key them in — those three are
// captured by hand on this layout — and must say what silently breaks if
// the vigencia is left empty.
const notes = p.notes.join(" ");
expect(notes).toMatch(/no trae vigencia, agente ni prima/i);
expect(notes).toMatch(/captúrelos a mano/i);
expect(notes).toMatch(/avisos de renovación/i);
});
it("takes the currency from the printed limits, not from the M.N. sublimits", () => {
// The body prose quotes sublimits in pesos ("$ 50,000.00 M.N."); every
// limit is in USD, and only the limits vote.
expect(parsePolicy(GMX_ESPEC).currency).toBe("USD");
});
it("reads each coverage under its own heading", () => {
const c = byRisk(parsePolicy(GMX_ESPEC));
expect(c.EDIFICIO?.insuredAmount).toBe(200000);
expect(c.CONTENIDOS?.insuredAmount).toBe(20000);
expect(c["ROBO DE CONTENIDOS"]?.insuredAmount).toBe(4000);
expect(c["ROBO DE CONTENIDOS"]?.deductible).toBe("Sin deducible");
});
it("splits a limit printed under Edificio / Contenidos sub-labels", () => {
const c = byRisk(parsePolicy(GMX_ESPEC));
expect(c["Remoción de escombros — Edificio"]?.insuredAmount).toBe(20000);
expect(c["Remoción de escombros — Contenidos"]?.insuredAmount).toBe(2000);
});
it("names a coverage after its heading, not after the wrapped tail of the prose above it", () => {
// Walking back from the limit hits "efectuados." — short, and the only
// thing separating it from a heading is that it is not preceded by a
// blank line.
const c = byRisk(parsePolicy(GMX_ESPEC));
expect(c["Gastos extraordinarios para casa habitación"]?.insuredAmount).toBe(22000);
expect(c["efectuados."]).toBeUndefined();
});
it("reads a limit printed on the label's own line", () => {
const c = byRisk(parsePolicy(GMX_ESPEC));
expect(c["Bienes a la Intemperie"]?.insuredAmount).toBe(10000);
});
it("reads a deductible stated as a sentence above the limit", () => {
const c = byRisk(parsePolicy(GMX_ESPEC));
expect(c["Bienes a la Intemperie"]?.deductible).toBe(
"5 POR CIENTO SOBRE SUMA ASEGURADA, 20 PORCIENTO DE PARTICIPACIÓN A CARGO DEL ASEGURADO DE TODA Y CADA PÉRDIDA.",
);
});
it("never borrows a neighbouring coverage's prose as a deductible", () => {
// "…o hasta el 10% de la suma asegurada de la sección de Edificio" is a
// sublimit rule for EDIFICIO, printed two paragraphs above CONTENIDOS.
const c = byRisk(parsePolicy(GMX_ESPEC));
expect(c.CONTENIDOS?.deductible).toBeNull();
expect(c.EDIFICIO?.deductible).toBeNull();
});
it("does not read the page-level DEDUCIBLES paragraph as a deductible", () => {
const p = parsePolicy(GMX_ESPEC);
expect(
p.coverages.some((c) => (c.deductible ?? "").includes("cláusula inflacionaria")),
).toBe(false);
});
it("reads a sublimit block as a sublimit OF the coverage above it", () => {
// The amount sits after a blank line AND a page break, and the block's
// own heading ("Sublímites:") names no risk.
const c = byRisk(parsePolicy(GMX_ESPEC));
expect(c["ROBO DE CONTENIDOS — sublímite"]?.insuredAmount).toBe(2000);
});
it("records an excluded catastrophic risk as excluded, never as zero", () => {
const p = parsePolicy(GMX_ESPEC);
const quake = p.coverages.filter((c) => /Terremoto/i.test(c.risk));
expect(quake).toHaveLength(2);
for (const c of quake) {
expect(c.risk).toMatch(/EXCLUIDO/);
// A coverage insured for $0 and an excluded coverage are the same
// number and very different facts.
expect(c.insuredAmount).toBeNull();
}
});
it("attaches the hydrometeorological deductible and coinsurance from its own block", () => {
const c = byRisk(parsePolicy(GMX_ESPEC));
const building = c["Fenómenos hidrometeorológicos — Sección Edificio"];
expect(building?.insuredAmount).toBe(200000);
expect(building?.deductible).toBe("1 POR CIENTO SOBRE SUMA ASEGURADA");
expect(building?.lossParticipation).toBe("10%");
expect(c["Fenómenos hidrometeorológicos — Sección Contenidos"]?.insuredAmount).toBe(20000);
});
it("names the same product as the caratula — one policy, two artifacts", () => {
expect(parsePolicy(GMX_ESPEC).policyTypeName).toBe("MULT");
});
it("carries the underwriting context the fields have no home for", () => {
const notes = parsePolicy(GMX_ESPEC).notes.join(" | ");
expect(notes).toMatch(/tipo de persona asegurada: Propietario/);
expect(notes).toMatch(/características del inmueble: Casa/);
expect(notes).toMatch(/cuerpo de agua/);
expect(notes).toMatch(/zona catastrófica declarada: A2/);
expect(notes).toMatch(/W_HogarGMX_12\.11\.2025\.pdf/);
});
});
/* ------------------------------------------------------------------ ANA */
/**
* Verbatim `pdftotext -layout` output of the PDFs A.N.A.'s portal produced
* for three real policies, cut at the end of the risk table (the legal
* boilerplate and the repeated AGENT COPY below it are not parsed, and the
* repeats are covered by their own test).
*
* The column padding is load-bearing on the driver's policy, which
* distinguishes SUM INSURED from PREMIUM by horizontal position alone — do
* not reflow these strings.
*/
const ANA_AUTO_AMPLIA = page(`A.N.A. COMPAÑIA DE SEGUROS SA DE CV
LUIS CABRERA #2033 INT. 201, Col. ZONA URBANA RIO TIJUANA
C.P. 22010 MUNICIPIO DE TIJUANA, BAJA CALIFORNIA
www.anaseguros.com.mx
AUTOMOBILE
ALL CLAIMS MUST BE REPORTED BEFORE LEAVING MEXICO
U.S. CELL PHONES TRY + 011-52-55-5322-82-66 MEXICAN CELL PHONES 800-911-911-9 SPECIAL POLICY FOR TOURISTS
TOLL-FREE FROM THE U.S.A. 888-335-7072 BELIZE CELL PHONES 00-52-55-5322-8266
WHATSAPP + 52-55-80-50-3633
No. 700489651
ISSUED BY: DATE ISSUED TERM OF INSURANCE
DAYS
JORGE HUMBERTO CUADROS DAY MONTH YEAR DAY MONTH YEAR TIME
BENITO JUAREZ 25 No.50 INT 38 CENTRO
04 08 2026 FROM 07 08 2026 12:01
365
ROSARITO, BAJA CALIFORNIA 22710
. 70175 TO 07 08 2027 12:01
DISCOUNT PREMIUM POLICY FEE TAX LOCAL TAX TOTAL
- 298.61 30.00 26.29 0.00 354.90
INSURED RAY DEAN II AND SUSAN ROCKHOLD
LICENSE P0066762
ADDRESS 10308 DONNA AVE EMAIL PROLABSALE@AOL.COM
CITY & STATE NORTHRIDGE, CA 91326 TELEPHONE 8184453524
PAYMENT DEADLINE
INSURANCE COMPANY LIEN HOLDER
IMMEDIATE
ITEM YEAR MAKE BODY SERIAL No. PLATES
VEHICLE 2017 CHRYSLER PACIFICA 2C4RC1DG7HR654698 8BPX206
TRAILER . .
TOWING . .
*** VALUE STATED MUST NOT EXCEED MARKET VALUE ***
***VEHICLES THAT HAVE BEEN ACQUIRED AS SALVAGE, REBUILT, OR HAVE BEEN USED PREVIOUSLY AS A TAXI WILL BE CONSIDERED WITH A REDUCED VALUE OF 35% (thirty-five percent), TAKING
AS A BASE THE VALUE OF A SIMILAR NORMAL VEHICLE, THAT IS, ONE THAT HAS NOT BEEN ACQUIRED AS SALVAGE AND ITS PREVIOUS USE HAS NOT BEEN AS A TAXI OR REBUILT. IT WILL BE THE
SOLE OBLIGATION AND RESPONSIBILITY OF THE INSURED TO DECLARATE THIS WHEN ACQUIRING THE POLICY.
SECTION SPECIFICATION OF RISKS LIMIT OF LIABILITY
MATERIAL DAMAGE WITH MANDATORY DEDUCTIBLE COVERED/EXCLUDED VEHICLE 8,000.00 DLLS.
1 DEDUCTIBLE: WITH MINIMUM OF $500.00 ON AUTOS
TRAILER
(SEDANS, COUPES, CONVERTIBLES AND STATION WAGONS) COVERED
AND $500.00 ON ALL OTHERS (PICK UPS, VANS, SUV´s AND MOTOR HOMES). 0.00 DLLS.
TOTAL THEFT WITH MANDATORY DEDUCTIBLE COVERED/EXCLUDED TOWING
2 DEDUCTIBLE: WITH MINIMUM OF $1,000.00 ON AUTOS 0.00 DLLS.
(SEDANS, COUPES, CONVERTIBLES AND STATION WAGONS) COVERED
AND $1,000.00 ON ALL OTHERS (PICK UPS, VANS, SUV´s AND MOTOR HOMES).
LIABILITY FOR PROPERTY DAMAGE TO THIRD PARTIES
3 100,000.00 DLLS.
BODILY INJURY LIABILITY PER PER
4 PERSON 100,000.00 ACCIDENT 200,000.00 DLLS.
MEDICAL EXPENSES PER PER
5 PERSON 5,000.00 ACCIDENT 25,000.00 DLLS.
COVERED/EXCLUDED PREMIUM
6 A.N.A.'s LEGAL AID
COVERED 40.00
COVERED/EXCLUDED PREMIUM
7 A.N.A.'s ROADSIDE ASSISTANCE
COVERED 40.00
CATASTROPHIC LIABILITY FOR DEATH OF THIRD PREMIUM
8 EXCLUDED
PARTIES DLLS. 0.00
ELITE OR ELITE PLUS WITH MANDATORY DEDUCTIBLE COVERED/EXCLUDED
9 PARTIAL THEFT (LIMIT 0.00 DLLS.WITH DEDUCTIBLE: 0.00 DLLS. PER EVENT) 0.00
VANDALISM (LIMIT 0.00 DLLS.WITH DEDUCTIBLE: 0.00 DLLS. PER EVENT) EXCLUDED
ISSUED ONLINE`);
const ANA_AUTO_RC_DIAS = page(`A.N.A. COMPAÑIA DE SEGUROS SA DE CV
LUIS CABRERA #2033 INT. 201, Col. ZONA URBANA RIO TIJUANA
C.P. 22010 MUNICIPIO DE TIJUANA, BAJA CALIFORNIA
www.anaseguros.com.mx
AUTOMOBILE
ALL CLAIMS MUST BE REPORTED BEFORE LEAVING MEXICO
U.S. CELL PHONES TRY + 011-52-55-5322-82-66 MEXICAN CELL PHONES 800-911-911-9 SPECIAL POLICY FOR TOURISTS
TOLL-FREE FROM THE U.S.A. 888-335-7072 BELIZE CELL PHONES 00-52-55-5322-8266
WHATSAPP + 52-55-80-50-3633
No. 700487807
ISSUED BY: DATE ISSUED TERM OF INSURANCE
DAYS
JORGE HUMBERTO CUADROS DIARIA DAY MONTH YEAR DAY MONTH YEAR TIME
BENITO JUAREZ 25 NO50 INT 38 COL CENTRO
22 07 2026 FROM 23 07 2026 12:01
3
ROSARITO BAJA CALIFORNIA 22710
(661) 612 12 55 70175 TO 26 07 2026 12:01
DISCOUNT PREMIUM POLICY FEE TAX LOCAL TAX TOTAL
- 10.77 25.00 2.86 0.00 38.63
INSURED STEPHEN RUPAN SHATAFIAN
LICENSE C1394198
ADDRESS 13181 CROSSROADS PARKWAY NORTH STE 300 EMAIL sshatafian@lee-associates.com
CITY & STATE CITY OF INDUSTRY, CA 91746 TELEPHONE 7143221072
PAYMENT DEADLINE
INSURANCE COMPANY LIEN HOLDER
IMMEDIATE
ITEM YEAR MAKE BODY SERIAL No. PLATES
VEHICLE 2022 FORD TRANSIT 1FBAX2CG3NKA69091 EC46T99
TRAILER . .
TOWING . .
*** VALUE STATED MUST NOT EXCEED MARKET VALUE ***
***VEHICLES THAT HAVE BEEN ACQUIRED AS SALVAGE, REBUILT, OR HAVE BEEN USED PREVIOUSLY AS A TAXI WILL BE CONSIDERED WITH A REDUCED VALUE OF 35% (thirty-five percent), TAKING
AS A BASE THE VALUE OF A SIMILAR NORMAL VEHICLE, THAT IS, ONE THAT HAS NOT BEEN ACQUIRED AS SALVAGE AND ITS PREVIOUS USE HAS NOT BEEN AS A TAXI OR REBUILT. IT WILL BE THE
SOLE OBLIGATION AND RESPONSIBILITY OF THE INSURED TO DECLARATE THIS WHEN ACQUIRING THE POLICY.
SECTION SPECIFICATION OF RISKS LIMIT OF LIABILITY
MATERIAL DAMAGE WITH MANDATORY DEDUCTIBLE COVERED/EXCLUDED VEHICLE 0.00 DLLS.
1 DEDUCTIBLE: ON AUTOS (SEDANS, COUPES, CONVERTIBLES AND
TRAILER
STATION WAGONS) AND OTHERS (PICK UPS, VANS, EXCLUDED
SUV´s AND MOTOR HOMES). 0.00 DLLS.
TOTAL THEFT WITH MANDATORY DEDUCTIBLE COVERED/EXCLUDED TOWING
2 DEDUCTIBLE: ON AUTOS (SEDANS, COUPES, CONVERTIBLES AND 0.00 DLLS.
STATION WAGONS) AND OTHERS (PICK UPS, VANS, EXCLUDED
SUV´s AND MOTOR HOMES).
LIABILITY FOR PROPERTY DAMAGE TO THIRD PARTIES
3 100,000.00 DLLS.
BODILY INJURY LIABILITY PER PER
4 PERSON 100,000.00 ACCIDENT 200,000.00 DLLS.
MEDICAL EXPENSES PER PER
5 PERSON 5,000.00 ACCIDENT 25,000.00 DLLS.
COVERED/EXCLUDED PREMIUM
6 A.N.A.'s LEGAL AID
COVERED 2.25
COVERED/EXCLUDED PREMIUM
7 A.N.A.'s ROADSIDE ASSISTANCE
COVERED 2.25
CATASTROPHIC LIABILITY FOR DEATH OF THIRD PREMIUM
8 EXCLUDED
PARTIES DLLS. 0.00
ELITE OR ELITE PLUS WITH MANDATORY DEDUCTIBLE COVERED/EXCLUDED
9 PARTIAL THEFT (LIMIT 0.00 DLLS.WITH DEDUCTIBLE: 0.00 DLLS. PER EVENT) 0.00
VANDALISM (LIMIT 0.00 DLLS.WITH DEDUCTIBLE: 0.00 DLLS. PER EVENT) EXCLUDED
ISSUED ONLINE`);
const ANA_LICENCIA = page(`A.N.A. COMPAÑIA DE SEGUROS SA DE CV
LUIS CABRERA #2033 INT. 201, Col.4 ZONA URBANA RIO TIJUANA
C.P. 22010 MUNICIPIO DE TIJUANA, BAJA CALIFORNIA
www.anaseguros.com.mx
DRIVER´S POLICY FOR AUTOMOBILE
ALL CLAIMS MUST BE REPORTED BEFORE LEAVING MEXICO
U.S. CELL PHONES TRY + 011-52-55-5322-82-66 MEXICAN CELL PHONES 800-911-911-9
SPECIAL POLICY FOR TOURISTS
TOLL-FREE FROM THE U.S.A. 888-335-7072 BELIZE CELL PHONES 00-52-55-5322-8266
WHATSAPP + 52-55-80-50-3633 No. 700489616
ISSUED BY: DATE ISSUED & TIME TERM OF INSURANCE
JORGE HUMBERTO CUADROS
DAYS
DAY MONTH YEAR DAY MONTH YEAR TIME
BENITO JUAREZ 25 No.50 INT 38 CENTRO 04 08 2026 FROM 06 08 2026 12:01
365
ROSARITO, BAJA CALIFORNIA 22710 TO 06 08 2027 12:01
. 70175
DISCOUNT PREMIUM POLICY FEE TAX LOCAL TAX TOTAL
- 142.78 30.00 13.82 0.00 186.60
LICENSE N0017668 EMAIL PWAGONER49@AOL.COM TELEPHONE 3102001538
POLICY HOLDER
1. NAME : PAMELA DENISE WAGONER Ph.3102001538
ADDRESS : 49305 HIGHWAY 74 SPC 10, PALM DESERT, CA, 92260,
DRIVER LICENSE : N0017668
2. NAME :
ADDRESS :
DRIVER LICENSE :
NONE
3. NAME :
ADDRESS :
DRIVER LICENSE :
NONE
4. NAME :
ADDRESS :
DRIVER LICENSE : NONE
5. NAME :
ADDRESS :
DRIVER LICENSE : NONE
SPECIFICATION OF RISKS SUM INSURED PREMIUM
LIABILITY FOR PROPERTY DAMAGE TO THIRD PARTIES 100,000.00 usd. 18.70 usd.
BODILY INJURY LIABILITY ( EXCLUDING OCCUPANTS OF THE VEHICLE ) 100,000.00 usd. Per Person
54.27 usd.
200,000.00 usd. Per Accident
CATASTROPHIC LIABILITY FOR DEATH OF THIRD PARTIES 0.00 usd. 0.00 usd.
MEDICAL EXPENSES 4,000.00 usd. Per Person
9.81 usd.
20,000.00 usd. Per Accident
COVERED/EXCLUDED PREMIUM
LEGAL AID
COVERED 30.00 usd.
COVERED/EXCLUDED PREMIUM
AUTOMOBILE ASSISTANCE
COVERED 30.00 usd.
The following risks are excluded Collision, overtuning and glass breakage, fire, total theft and natural disasters, partial theft and vandalism.`);
describe("detectPolicyProvider / ANA", () => {
it("claims ANA from the letterhead", () => {
expect(
detectPolicyProvider("A.N.A. COMPAÑIA DE SEGUROS SA DE CV\nwww.anaseguros.com.mx"),
).toBe("ANA");
});
it("does not let GMX's layout rules claim an ANA page", () => {
// Both books print "MATERIAL DAMAGE"-ish headings; the brand pass runs
// before any layout rule precisely so this can't go the other way.
expect(detectPolicyProvider(ANA_AUTO_AMPLIA.text)).toBe("ANA");
expect(detectPolicyProvider(ANA_LICENCIA.text)).toBe("ANA");
});
});
describe("parsePolicy / ANA automobile", () => {
const p = parsePolicy(ANA_AUTO_AMPLIA);
it("reads the header band", () => {
expect(p.provider).toBe("ANA");
expect(p.policyNumber).toBe("700489651");
expect(p.insuredName).toBe("RAY DEAN II AND SUSAN ROCKHOLD");
expect(p.agentName).toBe("JORGE HUMBERTO CUADROS");
expect(p.legalAddress).toBe("10308 DONNA AVE, NORTHRIDGE, CA 91326");
expect(p.zip).toBe("91326");
expect(p.currency).toBe("USD");
expect(p.premiumPayment).toBe("IMMEDIATE");
});
it("reads DD MM YYYY out of the three date column cells", () => {
expect(p.policyDate?.toISOString().slice(0, 10)).toBe("2026-08-04");
expect(p.policyFrom?.toISOString().slice(0, 10)).toBe("2026-08-07");
expect(p.policyTo?.toISOString().slice(0, 10)).toBe("2027-08-07");
});
it("maps the six money cells positionally, not by finding six amounts", () => {
// DISCOUNT prints as a bare "-" here. A "take the amounts in order"
// reading would shift every value one column left.
expect(p.netPremium).toBe(298.61);
expect(p.policyFee).toBe(30);
expect(p.total).toBe(354.9);
expect(p.notes.join(" | ")).toMatch(/impuesto: 26\.29/);
});
it("reads the vehicle by token role, not by column", () => {
expect(p.vehicles).toHaveLength(1);
expect(p.vehicles[0]).toEqual({
item: "VEHICLE",
modelYear: "2017",
make: "CHRYSLER",
bodyType: "PACIFICA",
vinNumber: "2C4RC1DG7HR654698",
licensePlate: "8BPX206",
});
});
it("reads a two-word BODY cell without losing the VIN", () => {
// "GENESIS SEDAN" is two tokens where "PACIFICA" is one — the VIN shape
// is the anchor, not the token count.
const v = parsePolicy(ANA_AUTO_RC_DIAS).vehicles[0];
expect(v.make).toBe("FORD");
expect(v.vinNumber).toBe("1FBAX2CG3NKA69091");
expect(v.licensePlate).toBe("EC46T99");
});
it("skips the empty TRAILER and TOWING slots", () => {
// Both print a "." per cell rather than being absent.
expect(p.vehicles.map((v) => v.item)).toEqual(["VEHICLE"]);
});
it("records the insured as a named driver with their licence", () => {
expect(p.drivers).toHaveLength(1);
expect(p.drivers[0].fullName).toBe("RAY DEAN II AND SUSAN ROCKHOLD");
expect(p.drivers[0].licenseNumber).toBe("P0066762");
expect(p.drivers[0].email).toBe("PROLABSALE@AOL.COM");
});
it("does not read the agent's own street number as the policy number", () => {
// "BENITO JUAREZ 25 No.50 INT 38" sits three lines above the No. cell.
expect(p.policyNumber).not.toBe("50");
expect(p.notes.join(" | ")).not.toMatch(/formas/);
});
it("reads the agent clave without picking up their postal code", () => {
// "ROSARITO, BAJA CALIFORNIA 22710" is five digits in the same band.
expect(p.notes.join(" | ")).toMatch(/clave de agente: 70175/);
expect(p.notes.join(" | ")).not.toMatch(/22710/);
});
it("labels the declared values by their printed item slot", () => {
const c = byRisk(p);
expect(c["MATERIAL DAMAGE — VEHICLE"]?.insuredAmount).toBe(8000);
expect(c["MATERIAL DAMAGE — TRAILER"]?.insuredAmount).toBe(0);
expect(c["TOTAL THEFT — TOWING"]?.insuredAmount).toBe(0);
});
it("keeps the deductible sentence out of the value columns", () => {
const c = byRisk(p);
expect(c["MATERIAL DAMAGE — VEHICLE"]?.deductible).toBe(
"WITH MINIMUM OF $500.00 ON AUTOS (SEDANS, COUPES, CONVERTIBLES AND " +
"STATION WAGONS) AND $500.00 ON ALL OTHERS (PICK UPS, VANS, SUV´s AND " +
"MOTOR HOMES).",
);
});
it("does not mistake the $500.00 inside the deductible for a sum insured", () => {
// It is the one amount in the block not suffixed "DLLS.".
const amounts = p.coverages.map((c) => c.insuredAmount);
expect(amounts).not.toContain(500);
});
it("splits the per-person and per-accident limits", () => {
const c = byRisk(p);
expect(c["BODILY INJURY LIABILITY — POR PERSONA"]?.insuredAmount).toBe(100000);
expect(c["BODILY INJURY LIABILITY — POR EVENTO"]?.insuredAmount).toBe(200000);
expect(c["MEDICAL EXPENSES — POR PERSONA"]?.insuredAmount).toBe(5000);
expect(c["MEDICAL EXPENSES — POR EVENTO"]?.insuredAmount).toBe(25000);
});
it("records an add-on's figure as a premium, never as a sum insured", () => {
// $40 is what legal aid COST. As `insuredAmount` it would read on the
// review screen as a $40 liability limit.
const c = byRisk(p);
expect(c["LEGAL AID"]?.premium).toBe(40);
expect(c["LEGAL AID"]?.insuredAmount).toBeNull();
expect(c["ROADSIDE ASSISTANCE"]?.premium).toBe(40);
});
it("unpacks section 9's parenthesised limit and deductible", () => {
const c = byRisk(p);
const theft = c["ELITE / ELITE PLUS — PARTIAL THEFT: EXCLUDED"];
expect(theft?.insuredAmount).toBe(0);
expect(theft?.deductible).toBe("0.00 DLLS. POR EVENTO");
expect(c["ELITE / ELITE PLUS — VANDALISM: EXCLUDED"]).toBeDefined();
});
it("emits each coverage once even though the PDF prints the face twice", () => {
// The real upload is ORIGINAL + AGENT COPY + receipt + three travel
// cards, all concatenated into one string before parsing.
const doubled = page(ANA_AUTO_AMPLIA.text + "\n\n" + ANA_AUTO_AMPLIA.text);
expect(parsePolicy(doubled).coverages).toHaveLength(p.coverages.length);
expect(parsePolicy(doubled).vehicles).toHaveLength(1);
});
});
describe("policy type, as a name for confirm to resolve", () => {
it("names ANA's two faces after the legacy tables they belong to", () => {
expect(parsePolicy(ANA_AUTO_AMPLIA).policyTypeName).toBe("AUTO");
expect(parsePolicy(ANA_AUTO_RC_DIAS).policyTypeName).toBe("AUTO");
expect(parsePolicy(ANA_LICENCIA).policyTypeName).toBe("LICENCIAS");
});
it("emits a NAME, never an id — the parser must not need a database", () => {
// Anything id-shaped here would mean the parser had reached for the DB.
for (const p of [ANA_AUTO_AMPLIA, ANA_AUTO_RC_DIAS, ANA_LICENCIA]) {
expect(parsePolicy(p).policyTypeName).toMatch(/^[A-Z_]+$/);
}
});
it("leaves the type unnamed when no parser claimed the page", () => {
expect(parsePolicy(page("a laundry receipt")).policyTypeName).toBeNull();
});
});
describe("parsePolicy / ANA responsabilidad civil por días", () => {
const p = parsePolicy(ANA_AUTO_RC_DIAS);
it("reads a by-the-day term rather than defaulting to a year", () => {
// Left at the schema's 365 default this weekend policy would sit in the
// renewals window a year out.
expect(p.policyFrom?.toISOString().slice(0, 10)).toBe("2026-07-23");
expect(p.policyTo?.toISOString().slice(0, 10)).toBe("2026-07-26");
expect(p.coveragePeriodDays).toBe(3);
});
it("reads the clave when the agent's phone occupies the left cell", () => {
// The by-the-day products print "(661) 612 12 55" ahead of the clave, so
// it is no longer the first thing on its line.
expect(p.notes.join(" | ")).toMatch(/clave de agente: 70175/);
});
it("marks the excluded sections as excluded, not as insured for zero", () => {
const risks = p.coverages.map((c) => c.risk);
expect(risks).toContain("MATERIAL DAMAGE — VEHICLE: EXCLUDED");
expect(risks).toContain("TOTAL THEFT — TOWING: EXCLUDED");
// The liability sections are what this product actually sells, and they
// are NOT excluded.
expect(risks).toContain("LIABILITY FOR PROPERTY DAMAGE TO THIRD PARTIES");
});
});
describe("parsePolicy / ANA driver's policy (licencia)", () => {
const p = parsePolicy(ANA_LICENCIA);
it("reads the holder off the numbered POLICY HOLDER list", () => {
expect(p.policyNumber).toBe("700489616");
expect(p.insuredName).toBe("PAMELA DENISE WAGONER");
expect(p.legalAddress).toBe("49305 HIGHWAY 74 SPC 10, PALM DESERT, CA, 92260");
expect(p.zip).toBe("92260");
});
it("lists one driver, not one per printed copy of the page", () => {
// The face renders three times in the real PDF; an unbounded walk
// returns the same person three times, which reads as a three-driver
// policy rather than as a parse bug.
const tripled = page([ANA_LICENCIA.text, ANA_LICENCIA.text, ANA_LICENCIA.text].join("\n\n"));
expect(p.drivers).toHaveLength(1);
expect(parsePolicy(tripled).drivers).toHaveLength(1);
});
it("splits the phone off the name even without the printed column gap", () => {
// The phone shares the name cell, and the only thing marking it off is
// white space — which the OCR seam is free to collapse. Depending on the
// gap surviving is what put "PAMELA DENISE WAGONER Ph.3102001538" in the
// insured field, where it matched no customer.
const collapsed = page(ANA_LICENCIA.text.replace(/ {2,}/g, " "));
expect(parsePolicy(collapsed).insuredName).toBe("PAMELA DENISE WAGONER");
});
it("drops the four empty driver slots", () => {
// Slots 2-5 print an empty NAME and a bare "NONE" licence.
expect(p.drivers.map((d) => d.fullName)).toEqual(["PAMELA DENISE WAGONER"]);
expect(p.drivers[0].licenseNumber).toBe("N0017668");
expect(p.drivers[0].phone).toBe("3102001538");
});
it("insures no vehicle", () => {
expect(p.vehicles).toEqual([]);
expect(p.notes.join(" | ")).toMatch(/no ampara un veh[íi]culo determinado/);
});
it("separates the SUM INSURED and PREMIUM columns by position", () => {
// Both columns print the same shape ("100,000.00 usd." / "18.70 usd.")
// and neither is labelled per row — only the offset tells them apart.
const c = byRisk(p);
const pd = c["LIABILITY FOR PROPERTY DAMAGE TO THIRD PARTIES"];
expect(pd?.insuredAmount).toBe(100000);
expect(pd?.premium).toBe(18.7);
});
it("reads the trailing Per Person / Per Accident labels on this layout", () => {
// They FOLLOW their amount here and PRECEDE it on the automobile face.
const c = byRisk(p);
expect(c["BODILY INJURY LIABILITY — POR PERSONA"]?.insuredAmount).toBe(100000);
expect(c["BODILY INJURY LIABILITY — POR EVENTO"]?.insuredAmount).toBe(200000);
expect(c["MEDICAL EXPENSES — POR PERSONA"]?.insuredAmount).toBe(4000);
expect(c["MEDICAL EXPENSES — POR EVENTO"]?.insuredAmount).toBe(20000);
});
it("charges a section's premium once, not once per limit", () => {
const c = byRisk(p);
expect(c["BODILY INJURY LIABILITY — POR PERSONA"]?.premium).toBe(54.27);
expect(c["BODILY INJURY LIABILITY — POR EVENTO"]?.premium).toBeNull();
});
it("handles the section order this layout uses", () => {
// CATASTROPHIC LIABILITY prints ABOVE MEDICAL EXPENSES here and below it
// on the automobile face; blocks are keyed by where the labels land.
const c = byRisk(p);
expect(c["CATASTROPHIC LIABILITY FOR DEATH OF THIRD PARTIES"]?.insuredAmount).toBe(0);
expect(c["LEGAL AID"]?.premium).toBe(30);
expect(c["ROADSIDE ASSISTANCE"]?.premium).toBe(30);
});
it("carries the excluded-risk sentence that defines the product", () => {
expect(p.notes.join(" | ")).toMatch(/riesgos excluidos: Collision, overtuning/);
});
});
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,100 @@
import { PolicyMatcherService } from "./policy-matcher.service";
import type { PrismaService } from "../prisma/prisma.service";
import type { ParsedPolicy } from "./parsers/policy-parser";
function parsed(over: Partial<ParsedPolicy> = {}): ParsedPolicy {
return {
provider: "GMX",
policyNumber: null,
insuredName: null,
notes: [],
coverages: [],
vehicles: [],
drivers: [],
...over,
} as unknown as ParsedPolicy;
}
function prismaStub(policies: unknown[], customers: { id: string; name: string }[]) {
const findManyPolicy = jest.fn().mockResolvedValue(policies);
const findManyCustomer = jest.fn().mockResolvedValue(customers);
return {
prisma: {
policy: { findMany: findManyPolicy },
customer: { findMany: findManyCustomer },
} as unknown as PrismaService,
findManyPolicy,
findManyCustomer,
};
}
const BOOK = [
{ id: "cust-1", name: "WAGONER, PAMELA" },
{ id: "cust-2", name: "SMITH, JOHN" },
];
describe("PolicyMatcherService name suggestions", () => {
it("suggests a customer when the policy number is new", async () => {
const { prisma } = prismaStub([], BOOK);
const svc = new PolicyMatcherService(prisma);
const r = await svc.match(
parsed({ policyNumber: "P-999", insuredName: "PAMELA DENISE WAGONER" } as never),
);
expect(r.customerSuggestions).toEqual([
expect.objectContaining({ customerId: "cust-1", tier: "PARTIAL" }),
]);
// The suggestion is surfaced, never applied.
expect(r.customerId).toBeNull();
expect(r.confident).toBe(false);
expect(r.note).toContain("posibles clientes por nombre: WAGONER, PAMELA");
});
it("suggests when the policy number could not be read at all", async () => {
const { prisma } = prismaStub([], BOOK);
const svc = new PolicyMatcherService(prisma);
const r = await svc.match(parsed({ insuredName: "PAMELA WAGONER" } as never));
expect(r.customerSuggestions[0]).toMatchObject({ customerId: "cust-1", tier: "EXACT" });
expect(r.customerId).toBeNull();
expect(r.note).toBe(
"no se pudo leer el número de póliza; posible cliente por nombre: WAGONER, PAMELA",
);
});
it("does not touch the book when the policy number hits", async () => {
const { prisma, findManyCustomer } = prismaStub(
[
{
id: "pol-1",
policyNumber: "P-1",
customerId: "cust-2",
customer: { name: "SMITH, JOHN" },
},
],
BOOK,
);
const svc = new PolicyMatcherService(prisma);
const r = await svc.match(
parsed({ policyNumber: "P-1", insuredName: "PAMELA WAGONER" } as never),
);
expect(r.confident).toBe(true);
expect(r.customerId).toBe("cust-2");
expect(r.customerSuggestions).toEqual([]);
expect(findManyCustomer).not.toHaveBeenCalled();
});
it("reads the customer book once across a batch", async () => {
const { prisma, findManyCustomer } = prismaStub([], BOOK);
const svc = new PolicyMatcherService(prisma);
await svc.match(parsed({ policyNumber: "A", insuredName: "PAMELA WAGONER" } as never));
await svc.match(parsed({ policyNumber: "B", insuredName: "JOHN SMITH" } as never));
expect(findManyCustomer).toHaveBeenCalledTimes(1);
});
});
@@ -0,0 +1,180 @@
import { Injectable } from "@nestjs/common";
import { PrismaService } from "../prisma/prisma.service";
import type { ParsedPolicy } from "./parsers/policy-parser";
import {
suggestCustomersByName,
suggestionNote,
type CustomerNameRow,
type CustomerNameSuggestion,
} from "./name-matcher";
export interface MatchResult {
policyId: string | null;
customerId: string | null;
/** Why it landed here — shown in the review queue verbatim. */
note: string;
/** True only for an unambiguous hit on `Policy.policyNumber`. */
confident: boolean;
/**
* Every policy that carries the parsed number, with its customer. >1 means
* the policy number is shared across customers and a human must pick.
*/
candidates: { policyId: string; customerId: string; customerName: string; policyNumber: string }[];
/**
* Customers whose name resembles the printed insured name. Populated only
* when the policy number resolved to nothing, and never used to set
* `customerId` or `confident` — see the class comment.
*/
customerSuggestions: CustomerNameSuggestion[];
}
/**
* How long the customer book is reused across documents in a batch.
*
* A twenty-page batch would otherwise read all 1536 rows twenty times. The
* only cost of the staleness is that a customer created in the last minute
* is not suggested — the picker still finds them, so nothing is lost that a
* reviewer cannot do in one click.
*/
const BOOK_TTL_MS = 60_000;
/**
* Resolves a parsed policy page to an existing Policy (and its customer) the
* office already holds.
*
* **Match on `Policy.policyNumber` alone, never on the printed insured name.**
* The certificate's "Insured" line is the account's registrant, which drifts
* from the current owner — the same problem the statement matcher cites for
* utility bills ("ARNAIZ ROSAS ELSA AURORA" on a CESPT receipt for a
* customer this office holds as "CATT, RANDY"). Names are surfaced for the
* reviewer to sanity-check and never feed matching.
*
* A policy number that matches zero rows means the policy is new: the
* review screen then offers a customer picker and the confirm step creates
* the row. Multiple hits are surfaced rather than auto-picked — duplicate
* policy numbers across customers do occur (same group policy bound by two
* related parties), and picking one arbitrarily would silently book the
* wrong coverage.
*
* On that zero-hit path only, the printed name is used to *rank the picker*
* — see `name-matcher.ts`. That is not a walk-back of the rule above: the
* suggestion never reaches `customerId` or `confident`, a human still picks,
* and the ranking exists because the office writes names surname-first
* ("WAGONER, PAMELA") while carriers print them given-name-first ("PAMELA
* DENISE WAGONER"), so the reviewer is retyping a name the machine could
* have offered.
*/
@Injectable()
export class PolicyMatcherService {
private book: { rows: CustomerNameRow[]; loadedAt: number } | null = null;
constructor(private readonly prisma: PrismaService) {}
async match(parsed: ParsedPolicy): Promise<MatchResult> {
if (!parsed.policyNumber) {
// No number to search on, so the page goes to review with a picker —
// the same place the name suggestions help.
return this.unmatched(
"no se pudo leer el número de póliza",
await this.suggestByName(parsed.insuredName),
);
}
const rows = await this.prisma.policy.findMany({
where: { policyNumber: parsed.policyNumber },
select: {
id: true,
policyNumber: true,
customerId: true,
customer: { select: { name: true } },
},
});
const candidates = rows.map((r) => ({
policyId: r.id,
customerId: r.customerId,
customerName: r.customer.name,
policyNumber: r.policyNumber,
}));
if (rows.length === 0) {
const suggestions = await this.suggestByName(parsed.insuredName);
const hint = suggestionNote(suggestions);
return {
policyId: null,
customerId: null,
note: [
`no se encontró ninguna póliza con el número ${parsed.policyNumber}`,
hint,
]
.filter(Boolean)
.join("; "),
confident: false,
candidates: [],
customerSuggestions: suggestions,
};
}
if (rows.length > 1) {
// The policy number did find rows; the reviewer picks among those, and
// adding name guesses on top would only add noise.
return {
policyId: null,
customerId: null,
note: `${rows.length} pólizas comparten el número ${parsed.policyNumber}`,
confident: false,
candidates,
customerSuggestions: [],
};
}
return {
policyId: candidates[0].policyId,
customerId: candidates[0].customerId,
note: `coincidencia exacta por número de póliza ${parsed.policyNumber}`,
confident: true,
candidates,
customerSuggestions: [],
};
}
private async suggestByName(
insuredName: string | null | undefined,
): Promise<CustomerNameSuggestion[]> {
if (!insuredName) return [];
return suggestCustomersByName(insuredName, await this.customerBook());
}
/**
* The whole customer book, held briefly. 1536 rows of `{id, name}` is a
* few hundred kilobytes and the comparison is pure token-set work, so
* scanning it beats any SQL approximation — and a `LIKE` search would in
* any case have to guess which token is the surname, which is the one
* thing the office's own data does not agree on.
*/
private async customerBook(): Promise<CustomerNameRow[]> {
if (this.book && Date.now() - this.book.loadedAt < BOOK_TTL_MS) {
return this.book.rows;
}
const rows = await this.prisma.customer.findMany({
select: { id: true, name: true },
});
this.book = { rows, loadedAt: Date.now() };
return rows;
}
private unmatched(
note: string,
customerSuggestions: CustomerNameSuggestion[] = [],
): MatchResult {
const hint = suggestionNote(customerSuggestions);
return {
policyId: null,
customerId: null,
note: [note, hint].filter(Boolean).join("; "),
confident: false,
candidates: [],
customerSuggestions,
};
}
}
@@ -0,0 +1,173 @@
import {
Body,
Controller,
Get,
Param,
Patch,
Post,
Query,
Req,
Res,
StreamableFile,
UploadedFiles,
UseGuards,
UseInterceptors,
} from "@nestjs/common";
import { FilesInterceptor } from "@nestjs/platform-express";
import type { Request, Response } from "express";
import { AuthenticatedGuard } from "../auth/authenticated.guard";
import { AbilityGuard } from "../auth/ability.guard";
import { RequireAbility } from "../auth/require-ability.decorator";
import { AuditService } from "../common/audit.service";
import type { UploadedFileLike } from "../storage/upload-file";
import { PolicyOcrService } from "./policy-ocr.service";
import {
ConfirmPolicyBatchDto,
CreatePolicyOcrBatchDto,
ReviewPolicyDocumentDto,
} from "./policy-ocr.dto";
/**
* Insurance OCR intake (policy_ocr_intake).
*
* Mirrors StatementsController shape: one batch = one upload session of
* policy PDFs from a provider portal (GMX today), one document per page.
* Confirming a batch delegates nothing to a separate billing path —
* everything goes through `Policy` (and optionally a Transaction for the
* premium), the same tables the manual `PolicyForm` writes.
*/
@Controller("policy-ocr")
@UseGuards(AuthenticatedGuard, AbilityGuard)
export class PolicyOcrController {
constructor(
private readonly policyOcr: PolicyOcrService,
private readonly audit: AuditService,
) {}
private actingId(req: Request): string {
return (req.user as { id: string } | undefined)?.id ?? "";
}
@Get("status")
async status() {
return {
ocrAvailable: await this.policyOcr.ocrAvailable(),
storageAvailable: this.policyOcr.storageAvailable(),
};
}
@Get("batches")
listBatches(@Query("page") page?: string, @Query("pageSize") pageSize?: string) {
return this.policyOcr.listBatches(
Math.max(1, Number(page) || 1),
Math.min(100, Math.max(1, Number(pageSize) || 25)),
);
}
@Get("batches/:id")
getBatch(@Param("id") id: string) {
return this.policyOcr.getBatch(id);
}
@Get("batches/:id/documents")
listDocuments(@Param("id") id: string) {
return this.policyOcr.listDocuments(id);
}
/**
* The source PDF for a parsed policy document. One PDF = one parsed policy,
* so this returns the entire upload (typically multi-page for insurance
* certificates). The review screen embeds it in an iframe.
*/
@Get("documents/:id/page")
async pageImage(
@Param("id") id: string,
@Res({ passthrough: true }) res: Response,
) {
const { stream, contentType, contentLength } = await this.policyOcr.pageImage(id);
res.set({
// The doc row stores the source PDF, not a rendered page image.
"Content-Type": contentType ?? "application/pdf",
...(contentLength ? { "Content-Length": String(contentLength) } : {}),
});
return new StreamableFile(stream);
}
// --- writes ---------------------------------------------------------------
@Post("batches")
@RequireAbility("policy:ingest")
@UseInterceptors(
FilesInterceptor("files", 25, { limits: { fileSize: 50 * 1024 * 1024 } }),
)
async createBatch(
@UploadedFiles() files: UploadedFileLike[] | undefined,
@Body() _dto: CreatePolicyOcrBatchDto,
@Query("label") label: string | undefined,
@Req() req: Request,
) {
const batch = await this.policyOcr.createBatch(
files ?? [],
this.actingId(req),
label ?? _dto.label,
);
void this.audit.log(this.actingId(req), "policyOcr.batch.create", {
batchId: batch.id,
fileCount: batch.fileCount,
});
return batch;
}
@Patch("documents/:id")
@RequireAbility("policy:ocr-review")
async review(
@Param("id") id: string,
@Body() dto: ReviewPolicyDocumentDto,
@Req() req: Request,
) {
const doc = await this.policyOcr.review(id, dto, this.actingId(req));
void this.audit.log(this.actingId(req), "policyOcr.document.review", {
documentId: id,
status: doc.status,
});
return doc;
}
@Post("documents/:id/reject")
@RequireAbility("policy:ocr-review")
async reject(@Param("id") id: string, @Req() req: Request) {
const doc = await this.policyOcr.reject(id, this.actingId(req));
void this.audit.log(this.actingId(req), "policyOcr.document.reject", {
documentId: id,
});
return doc;
}
/** Abandon a batch pending review — rejects every unapplied page. */
@Post("batches/:id/discard")
@RequireAbility("policy:ocr-review")
async discard(@Param("id") id: string, @Req() req: Request) {
const result = await this.policyOcr.discardBatch(id, this.actingId(req));
void this.audit.log(this.actingId(req), "policyOcr.batch.discard", {
batchId: id,
rejected: result.rejected,
});
return result;
}
@Post("batches/:id/confirm")
@RequireAbility("policy:ocr-review")
async confirm(
@Param("id") id: string,
@Body() dto: ConfirmPolicyBatchDto,
@Req() req: Request,
) {
const result = await this.policyOcr.confirmBatch(id, dto, this.actingId(req));
void this.audit.log(this.actingId(req), "policyOcr.batch.confirm", {
batchId: id,
applied: result.applied,
postedTransactions: result.postedTransactions,
});
return result;
}
}
+97
View File
@@ -0,0 +1,97 @@
import { Type } from "class-transformer";
import {
IsArray,
IsDateString,
IsEnum,
IsInt,
IsNumber,
IsObject,
IsOptional,
IsString,
Max,
Min,
MinLength,
ValidateNested,
} from "class-validator";
/** One document's confirmed-after-review state. The service reads these
* fields and writes them onto either a matched Policy or a freshly created
* one. Anything null here is not written. */
export class ConfirmPolicyDocumentDto {
@IsString() documentId!: string;
/** Required when creating a new Policy; ignored if `policyId` is set. */
@IsOptional() @IsString() customerId?: string;
/** Reviewer's explicit lookup picks. Both beat the parsed name; omitted,
* the service resolves `policy_types` / `insurance_providers` by name and
* leaves the FK null when there is no such row. */
@IsOptional() @IsString() policyTypeId?: string;
@IsOptional() @IsString() insuranceProviderId?: string;
/** Set when the document matched an existing Policy. */
@IsOptional() @IsString() policyId?: string;
@IsOptional() @IsString() policyNumber?: string;
@IsOptional() @IsString() insuredName?: string;
@IsOptional() @IsString() additionalInsured?: string;
@IsOptional() @IsString() agentName?: string;
@IsOptional() @IsString() legalAddress?: string;
@IsOptional() @IsString() zip?: string;
@IsOptional() @IsDateString() policyFrom?: string;
@IsOptional() @IsDateString() policyTo?: string;
@IsOptional() @IsDateString() policyDate?: string;
@IsOptional() @IsEnum(["MXN", "USD", "EUR"]) currency?: "MXN" | "USD" | "EUR";
@IsOptional() @IsNumber() netPremium?: number;
@IsOptional() @IsNumber() policyFee?: number;
@IsOptional() @IsNumber() brokerFee?: number;
@IsOptional() @IsNumber() total?: number;
@IsOptional() @IsString() premiumPayment?: string;
/** Printed term in days. Omitted leaves the parsed value (or the schema's
* 365 default) in place; ANA sells 3- and 4-day tourist policies. */
@IsOptional() @IsInt() @Min(1) @Max(3660) coveragePeriodDays?: number;
/** Coverages parsed off the PDF, passed through verbatim to Policy.coveragesJson. */
@IsOptional() @IsObject() coveragesJson?: unknown;
/** When true, write a Transaction(domain=INSURANCE, amount=-netPremium)
* in addition to creating/updating the Policy. Skipped if netPremium is
* null or zero. */
@IsOptional() postPremium?: boolean;
}
export class ConfirmPolicyBatchDto {
@IsArray()
@ValidateNested({ each: true })
@Type(() => ConfirmPolicyDocumentDto)
documents!: ConfirmPolicyDocumentDto[];
}
/** Staff correction of one document's extracted fields or its match. */
export class ReviewPolicyDocumentDto {
@IsOptional() @IsString() policyNumber?: string;
@IsOptional() @IsString() insuredName?: string;
@IsOptional() @IsString() additionalInsured?: string;
@IsOptional() @IsString() agentName?: string;
@IsOptional() @IsString() legalAddress?: string;
@IsOptional() @IsString() zip?: string;
@IsOptional() @IsDateString() policyFrom?: string;
@IsOptional() @IsDateString() policyTo?: string;
@IsOptional() @IsDateString() policyDate?: string;
@IsOptional() @IsString() currency?: string;
@IsOptional() @IsNumber() netPremium?: number;
@IsOptional() @IsNumber() policyFee?: number;
@IsOptional() @IsNumber() brokerFee?: number;
@IsOptional() @IsNumber() total?: number;
@IsOptional() @IsString() premiumPayment?: string;
@IsOptional() @IsInt() @Min(1) @Max(3660) coveragePeriodDays?: number;
@IsOptional() @IsObject() coveragesJson?: unknown;
/** Set by the reviewer when the document matched an existing Policy. */
@IsOptional() @IsString() matchedPolicyId?: string;
/** Set by the reviewer when creating a new Policy. */
@IsOptional() @IsString() matchedCustomerId?: string;
/** Force-confirm a doc even when the matcher left it ambiguous. */
@IsOptional() forceConfirm?: boolean;
}
export class CreatePolicyOcrBatchDto {
@IsOptional() @IsString() @MinLength(1) label?: string;
}
@@ -0,0 +1,18 @@
import { Module } from "@nestjs/common";
import { OcrModule } from "../ocr/ocr.module";
import { PolicyOcrController } from "./policy-ocr.controller";
import { PolicyOcrService } from "./policy-ocr.service";
import { PolicyMatcherService } from "./policy-matcher.service";
/**
* Reuses the OCR seam from OcrModule unchanged: the Tesseract provider is
* bound there and `OcrProvider` is the only thing the parsers touch. This
* module registers its own controller + service + matcher; nothing about
* utility ingestion needs to know about it.
*/
@Module({
imports: [OcrModule],
controllers: [PolicyOcrController],
providers: [PolicyOcrService, PolicyMatcherService],
})
export class PolicyOcrModule {}
@@ -0,0 +1,997 @@
import {
BadRequestException,
Inject,
Injectable,
Logger,
NotFoundException,
} from "@nestjs/common";
import { Currency, Prisma } from "@jorgecuadros/database";
import { PrismaService } from "../prisma/prisma.service";
import { StorageService } from "../storage/storage.service";
import type { UploadedFileLike } from "../storage/upload-file";
import { OCR_PROVIDER, type OcrPage, type OcrProvider } from "../statements/ocr/ocr.provider";
import { parsePolicy, type ParsedDriver, type ParsedVehicle } from "./parsers/policy-parser";
import { PolicyMatcherService } from "./policy-matcher.service";
import type {
ConfirmPolicyBatchDto,
ConfirmPolicyDocumentDto,
ReviewPolicyDocumentDto,
} from "./policy-ocr.dto";
/**
* Insurance OCR intake — mirrors the statement pipeline at
* `apps/api/src/statements/statements.service.ts`. Reuses the OCR seam and
* Tesseract binding unchanged; the parsers and matcher are policy-specific.
*
* Why a parallel pipeline rather than a column on StatementDocument: the
* matcher keys on `Policy.policyNumber`, the confirm step writes to a
* different table (`Policy`, not `Transaction`), and the review UI shows
* different fields. Sharing one queue would either bloat the row with null
* columns or force the review screen to branch on a discriminator — both
* worse than a thin second table.
*/
@Injectable()
export class PolicyOcrService {
private readonly logger = new Logger(PolicyOcrService.name);
constructor(
private readonly prisma: PrismaService,
private readonly storage: StorageService,
private readonly matcher: PolicyMatcherService,
@Inject(OCR_PROVIDER) private readonly ocr: OcrProvider,
) {}
ocrAvailable(): Promise<boolean> {
return this.ocr.available();
}
storageAvailable(): boolean {
return this.storage.available;
}
// --- ingest ---------------------------------------------------------------
async createBatch(
files: UploadedFileLike[],
uploadedById: string,
label?: string,
) {
if (!files?.length) throw new BadRequestException("No se recibió ningún archivo.");
if (!(await this.ocr.available())) {
throw new BadRequestException(
"El servidor no tiene OCR instalado; no se pueden leer PDFs de pólizas.",
);
}
if (!this.storage.available) {
throw new BadRequestException(
"El almacenamiento de documentos no está configurado; no se pueden " +
"guardar los PDFs escaneados.",
);
}
// The provider is not asked of the uploader and not assumed: `process`
// sets it from what the parsers actually claimed, so the batch label can
// never contradict its own documents. Until then it says so.
const batch = await this.prisma.policyOcrBatch.create({
data: { provider: "por detectar", uploadedById, label, fileCount: files.length },
});
const copies = files.map((f) => ({ buffer: f.buffer, name: f.originalname }));
void this.process(batch.id, copies).catch(async (err) => {
this.logger.error(`Policy OCR batch ${batch.id} failed: ${(err as Error).message}`);
await this.prisma.policyOcrBatch.update({
where: { id: batch.id },
data: { status: "FAILED", error: (err as Error).message },
});
});
return batch;
}
/**
* Render → text → parse → match, **one PolicyOcrDocument row per uploaded
* file**. The GMX certificate is a 2-page PDF where page 1 carries the
* contract header and page 2 carries the per-coverage table — both pages
* describe the SAME policy, so the parser concatenates them and the
* matcher runs once. `pageNumber` on the row is repurposed as the file
* ordinal within the batch (1, 2, 3…) — the unique constraint
* `(batchId, pageNumber)` still holds and lets a single batch carry many
* policies.
*
* The doc's `storageKey` is the SOURCE PDF (`policy-ocr/{batchId}/source-N.pdf`)
* rather than a rendered page image, so the review screen can embed the
* exact artifact the office received. The rendered page PNGs are still
* stored under `policy-ocr/{batchId}/page-M.png` for any future re-OCR or
* image-based audit, but they aren't used as `storageKey` for the document.
*/
private async process(
batchId: string,
files: { buffer: Buffer; name?: string }[],
) {
await this.prisma.policyOcrBatch.update({
where: { id: batchId },
data: { status: "PROCESSING" },
});
let fileOrdinal = 0;
let globalPageOrdinal = 0;
const providersSeen = new Set<string>();
for (const file of files) {
fileOrdinal += 1;
const sourceKey = `policy-ocr/${batchId}/source-${fileOrdinal}.pdf`;
await this.storage.put(sourceKey, file.buffer, "application/pdf");
const pages = await this.ocr.renderPages(file.buffer);
const textLayer = await this.ocr.textPages(file.buffer).catch(() => []);
// One OcrPage per rendered page: text-layer wins when present (cheap,
// exact), OCR the rendered image when it isn't. Same precedence rule
// as the statement OCR pipeline.
const perPageOcr: OcrPage[] = [];
for (const [index, image] of pages.entries()) {
globalPageOrdinal += 1;
const pageStorageKey = `policy-ocr/${batchId}/page-${globalPageOrdinal}.png`;
await this.storage.put(pageStorageKey, image, "image/png");
const embedded = textLayer[index] ?? null;
const pageOcr = embedded ?? (await this.ocr.recognize(image));
perPageOcr.push(pageOcr);
}
// Concatenate every page's text with a blank line between pages so the
// parser's anchored regexes (^From$, ^Currency\s+...) still work
// across page boundaries — pdftotext -bbox-layout produces newline-
// separated text per page already, the `\n\n` just preserves a clear
// boundary in ocrRawText for debugging.
const mergedText = perPageOcr.map((p) => p.text).join("\n\n");
const avgConfidence =
perPageOcr.length === 0
? 0
: perPageOcr.reduce((s, p) => s + p.confidence, 0) / perPageOcr.length;
const synthetic: OcrPage = {
text: mergedText,
words: [],
confidence: avgConfidence,
};
try {
const parsed = parsePolicy(synthetic);
if (parsed.provider === "") {
throw new Error("no se reconoció el proveedor");
}
providersSeen.add(parsed.provider);
const match = await this.matcher.match(parsed);
const notes = [...parsed.notes, match.note].filter(Boolean);
// Confident when exactly one Policy carries the printed number —
// the only unambiguous hit we trust. A new policy (no match) still
// needs a customer pick, so it stays in review.
const trusted = match.confident && parsed.policyNumber != null;
await this.prisma.policyOcrDocument.create({
data: {
batchId,
pageNumber: fileOrdinal,
storageKey: sourceKey,
status: trusted ? "MATCHED" : "NEEDS_REVIEW",
ocrRawText: mergedText,
ocrConfidence: new Prisma.Decimal(avgConfidence.toFixed(3)),
provider: parsed.provider,
extractedPolicyNumber: parsed.policyNumber,
extractedInsuredName: parsed.insuredName,
extractedAdditionalInsured: parsed.additionalInsured,
extractedAgentName: parsed.agentName,
extractedLegalAddress: parsed.legalAddress,
extractedZip: parsed.zip,
extractedPolicyFrom: parsed.policyFrom,
extractedPolicyTo: parsed.policyTo,
extractedPolicyDate: parsed.policyDate,
extractedCurrency: parsed.currency,
extractedNetPremium:
parsed.netPremium != null ? new Prisma.Decimal(parsed.netPremium) : null,
extractedPolicyFee:
parsed.policyFee != null ? new Prisma.Decimal(parsed.policyFee) : null,
extractedBrokerFee:
parsed.brokerFee != null ? new Prisma.Decimal(parsed.brokerFee) : null,
extractedTotal:
parsed.total != null ? new Prisma.Decimal(parsed.total) : null,
extractedCoveragesJson: parsed.coverages.length
? (parsed.coverages as unknown as Prisma.InputJsonValue)
: Prisma.DbNull,
extractedPremiumPayment: parsed.premiumPayment,
extractedCoveragePeriodDays: parsed.coveragePeriodDays,
extractedVehiclesJson: parsed.vehicles.length
? (parsed.vehicles as unknown as Prisma.InputJsonValue)
: Prisma.DbNull,
extractedDriversJson: parsed.drivers.length
? (parsed.drivers as unknown as Prisma.InputJsonValue)
: Prisma.DbNull,
extractedPolicyTypeName: parsed.policyTypeName,
matchedPolicyId: match.policyId,
matchedCustomerId: match.customerId,
matchCandidates: match.candidates.length
? (match.candidates as unknown as Prisma.InputJsonValue)
: Prisma.DbNull,
customerSuggestions: match.customerSuggestions.length
? (match.customerSuggestions as unknown as Prisma.InputJsonValue)
: Prisma.DbNull,
matchNote: notes.join("; "),
},
});
} catch (err) {
// The file as a whole failed to parse (no provider, parse exception).
// One OCR_FAILED row per file is the right granularity — the page
// images are still on disk for a re-run after a parser fix.
await this.prisma.policyOcrDocument.create({
data: {
batchId,
pageNumber: fileOrdinal,
storageKey: sourceKey,
status: "OCR_FAILED",
matchNote: (err as Error).message,
},
});
}
}
await this.prisma.policyOcrBatch.update({
where: { id: batchId },
data: {
status: "READY_FOR_REVIEW",
// Whatever the parsers claimed. A mixed upload is labelled as mixed
// rather than as whichever provider happened to come first — the
// review header is the only place staff see what they dropped in.
provider: [...providersSeen].sort().join(" + ") || "desconocido",
},
});
}
// --- reads ----------------------------------------------------------------
async listBatches(page: number, pageSize: number) {
const [total, items] = await this.prisma.$transaction([
this.prisma.policyOcrBatch.count(),
this.prisma.policyOcrBatch.findMany({
orderBy: { createdAt: "desc" },
skip: (page - 1) * pageSize,
take: pageSize,
include: {
uploadedBy: { select: { name: true } },
_count: { select: { documents: true } },
},
}),
]);
return { items, total, page, pageSize, pageCount: Math.ceil(total / pageSize) };
}
async getBatch(id: string) {
const batch = await this.prisma.policyOcrBatch.findUnique({
where: { id },
include: { uploadedBy: { select: { name: true } } },
});
if (!batch) throw new NotFoundException("Lote no encontrado.");
const counts = await this.prisma.policyOcrDocument.groupBy({
by: ["status"],
where: { batchId: id },
_count: { _all: true },
});
return {
...batch,
byStatus: Object.fromEntries(counts.map((c) => [c.status, c._count._all])),
};
}
async listDocuments(batchId: string) {
return this.prisma.policyOcrDocument.findMany({
where: { batchId },
orderBy: { pageNumber: "asc" },
include: {
matchedCustomer: { select: { id: true, name: true } },
matchedPolicy: {
select: {
id: true,
policyNumber: true,
customerId: true,
customer: { select: { name: true } },
},
},
},
});
}
/**
* The source PDF for the document, so the review screen can show the
* exact artifact the office uploaded (the browser's PDF viewer handles
* scrolling, zoom, and selection natively). The rendered page PNGs
* remain on disk under `policy-ocr/{batchId}/page-N.png` for any
* future re-OCR, but the doc row points here at the source.
*/
async pageImage(documentId: string) {
const doc = await this.prisma.policyOcrDocument.findUnique({
where: { id: documentId },
select: { storageKey: true },
});
if (!doc) throw new NotFoundException("Documento no encontrado.");
return this.storage.getStream(doc.storageKey);
}
// --- review ---------------------------------------------------------------
async review(id: string, dto: ReviewPolicyDocumentDto, reviewedById: string) {
const doc = await this.prisma.policyOcrDocument.findUnique({ where: { id } });
if (!doc) throw new NotFoundException("Documento no encontrado.");
if (doc.status === "POSTED") {
throw new BadRequestException("Este documento ya fue aplicado.");
}
// Trusting a customer-supplied pair (policyId, customerId) without
// cross-check is how a document lands on the wrong customer's ledger;
// pin them here from the DB.
let matchedPolicyId = dto.matchedPolicyId ?? doc.matchedPolicyId;
let matchedCustomerId = doc.matchedCustomerId;
if (matchedPolicyId) {
const p = await this.prisma.policy.findUnique({
where: { id: matchedPolicyId },
select: { customerId: true },
});
if (!p) throw new BadRequestException("Póliza no encontrada.");
matchedCustomerId = p.customerId;
} else if (dto.matchedCustomerId) {
const c = await this.prisma.customer.findUnique({
where: { id: dto.matchedCustomerId },
select: { id: true },
});
if (!c) throw new BadRequestException("Cliente no encontrado.");
matchedCustomerId = c.id;
}
return this.prisma.policyOcrDocument.update({
where: { id },
data: {
extractedPolicyNumber: dto.policyNumber ?? undefined,
extractedInsuredName: dto.insuredName ?? undefined,
extractedAdditionalInsured: dto.additionalInsured ?? undefined,
extractedAgentName: dto.agentName ?? undefined,
extractedLegalAddress: dto.legalAddress ?? undefined,
extractedZip: dto.zip ?? undefined,
extractedPolicyFrom: dto.policyFrom ? new Date(dto.policyFrom) : undefined,
extractedPolicyTo: dto.policyTo ? new Date(dto.policyTo) : undefined,
extractedPolicyDate: dto.policyDate ? new Date(dto.policyDate) : undefined,
extractedCurrency: dto.currency ?? undefined,
extractedNetPremium:
dto.netPremium != null ? new Prisma.Decimal(dto.netPremium) : undefined,
extractedPolicyFee:
dto.policyFee != null ? new Prisma.Decimal(dto.policyFee) : undefined,
extractedBrokerFee:
dto.brokerFee != null ? new Prisma.Decimal(dto.brokerFee) : undefined,
extractedTotal:
dto.total != null ? new Prisma.Decimal(dto.total) : undefined,
extractedCoveragesJson: dto.coveragesJson
? (dto.coveragesJson as Prisma.InputJsonValue)
: undefined,
extractedPremiumPayment: dto.premiumPayment ?? undefined,
extractedCoveragePeriodDays: dto.coveragePeriodDays ?? undefined,
matchedPolicyId,
matchedCustomerId,
status: dto.forceConfirm ? "CONFIRMED" : "MATCHED",
reviewedById,
reviewedAt: new Date(),
},
});
}
async reject(id: string, reviewedById: string) {
const doc = await this.prisma.policyOcrDocument.findUnique({ where: { id } });
if (!doc) throw new NotFoundException("Documento no encontrado.");
if (doc.status === "POSTED") {
throw new BadRequestException("Este documento ya fue aplicado.");
}
const updated = await this.prisma.policyOcrDocument.update({
where: { id },
data: { status: "REJECTED", reviewedById, reviewedAt: new Date() },
});
// Rejecting the last open page settles the batch just as confirming it
// would — without this, a fully-rejected batch sat in READY_FOR_REVIEW
// forever because only confirmBatch() ever closed one.
await this.closeIfDone(doc.batchId);
return updated;
}
/**
* Throw away a whole batch that is pending review: every page that has not
* been applied is marked REJECTED and the batch itself becomes DISCARDED.
*
* Refuses once any page is POSTED — a partly-applied batch has already
* written Policy (and possibly Transaction) rows, and hiding the paperwork
* behind a "discarded" label would leave those rows unexplained. Reject the
* remaining pages individually instead.
*/
async discardBatch(batchId: string, reviewedById: string) {
const batch = await this.prisma.policyOcrBatch.findUnique({
where: { id: batchId },
});
if (!batch) throw new NotFoundException("Lote no encontrado.");
if (batch.status === "DISCARDED") {
throw new BadRequestException("Este lote ya fue descartado.");
}
const posted = await this.prisma.policyOcrDocument.count({
where: { batchId, status: "POSTED" },
});
if (posted > 0) {
throw new BadRequestException(
`No se puede descartar: ${posted} página(s) ya se aplicaron a una póliza.`,
);
}
const { count } = await this.prisma.policyOcrDocument.updateMany({
where: { batchId, status: { notIn: ["POSTED", "REJECTED"] } },
data: { status: "REJECTED", reviewedById, reviewedAt: new Date() },
});
await this.prisma.policyOcrBatch.update({
where: { id: batchId },
data: { status: "DISCARDED", completedAt: new Date() },
});
return { batchId, rejected: count };
}
// --- confirm --------------------------------------------------------------
/**
* Apply every confirmed document: create or update the Policy, attach the
* source PDF as a PolicyDocument, and (when staff asked + premium parses)
* write a Transaction row. Each step is guarded by status checks so a
* double-confirm cannot re-apply a document.
*/
async confirmBatch(batchId: string, dto: ConfirmPolicyBatchDto, reviewedById: string) {
const batch = await this.prisma.policyOcrBatch.findUnique({ where: { id: batchId } });
if (!batch) throw new NotFoundException("Lote no encontrado.");
const results: { documentId: string; policyId: string; postedTransactionId: string | null }[] = [];
for (const item of dto.documents) {
const doc = await this.prisma.policyOcrDocument.findUnique({
where: { id: item.documentId },
});
if (!doc) {
throw new BadRequestException(`Documento ${item.documentId} no encontrado.`);
}
if (doc.status === "POSTED") {
throw new BadRequestException(
`El documento página ${doc.pageNumber} ya fue aplicado.`,
);
}
if (!item.policyId && !item.customerId) {
throw new BadRequestException(
`Documento página ${doc.pageNumber}: falta póliza destino o cliente.`,
);
}
// 1. Resolve the lookup rows the parser can only name. The reviewer's
// explicit pick always wins; the parsed name is the fallback.
const lookups = await this.resolveLookups(item, doc);
// 2. Resolve target Policy (create or update). Field selection: every
// non-null `extracted*` on the doc (post-review) is written. Null is
// preserved — never overwrite an existing Policy's `netPremium` with
// null because the certificate page didn't carry one.
let policyId = item.policyId ?? null;
if (policyId) {
const updateData = buildPolicyUpdateFromDoc(item, doc, lookups);
await this.prisma.policy.update({
where: { id: policyId },
data: updateData,
});
} else {
// Create under the picked customer. `policyNumber` is the only field
// that must be present.
if (!item.policyNumber && !doc.extractedPolicyNumber) {
throw new BadRequestException(
`Documento página ${doc.pageNumber}: falta número de póliza.`,
);
}
const createData = buildPolicyCreateFromDoc(item, doc, item.customerId!, lookups);
const created = await this.prisma.policy.create({
data: createData,
});
policyId = created.id;
}
// 3. Vehicles and named drivers, for the providers whose face carries
// them (ANA's automobile and driver's policies; never GMX Hogar).
await this.applyVehiclesAndDrivers(doc, policyId);
// 4. Attach the source PDF as a PolicyDocument. `doc.storageKey`
// already points at the exact upload (`policy-ocr/{batchId}/source-N.pdf`)
// so the attach is just a stream copy into the policy's namespace —
// the previous per-page "which file did this page come from" walk is
// gone because one PDF = one doc now.
await this.attachSourcePdf(doc.storageKey, policyId, doc.provider);
// 5. Optionally post the premium to the ledger. Only when staff
// explicitly asked (`postPremium` true) and netPremium parses — without
// that gate a missing premium would silently book $0.
let postedTransactionId: string | null = null;
const premium =
item.netPremium != null
? item.netPremium
: doc.extractedNetPremium != null
? Number(doc.extractedNetPremium)
: null;
if (item.postPremium && premium && premium > 0) {
const tx = await this.prisma.transaction.create({
data: {
customerId: (await this.policyCustomerId(policyId))!,
domain: "INSURANCE",
amount: new Prisma.Decimal(-Math.abs(premium)),
transactionDate: doc.extractedPolicyDate ?? doc.extractedPolicyFrom ?? new Date(),
currency: (item.currency ??
doc.extractedCurrency ??
"MXN") as Currency,
reference: item.policyNumber ?? doc.extractedPolicyNumber ?? null,
period: null,
captureSource: "OCR",
captureRef: doc.id,
message: `Prima de póliza ${item.policyNumber ?? doc.extractedPolicyNumber ?? ""}`,
},
});
postedTransactionId = tx.id;
}
await this.prisma.policyOcrDocument.update({
where: { id: doc.id },
data: {
status: "POSTED",
matchedPolicyId: policyId,
reviewedById,
reviewedAt: new Date(),
createdPolicyId: item.policyId ? null : policyId,
postedTransactionId,
},
});
results.push({
documentId: doc.id,
policyId,
postedTransactionId,
});
}
await this.closeIfDone(batchId);
return {
applied: results.length,
policies: results.map((r) => r.policyId),
postedTransactions: results.filter((r) => r.postedTransactionId).length,
};
}
/**
* Turn the two things the parser can only NAME into foreign keys.
*
* The parser is a pure function over text and never touches the database,
* so it emits `policyTypeName` ("AUTO") and `provider` ("ANA"). Resolving
* them here keeps that boundary and means a renamed lookup row is a data
* change rather than a parser change.
*
* **Resolve, never create.** A missing `policy_types` row is a signal that
* a human deleted it (that is exactly how M_EMPR disappeared), and silently
* recreating it would undo that decision with no record. The field stays
* null and the reviewer can add the row through the lookups screen.
*
* An explicit pick from the reviewer always beats the parsed name.
*/
private async resolveLookups(
item: ConfirmPolicyDocumentDto,
doc: { extractedPolicyTypeName: string | null; provider: string | null },
): Promise<{ policyTypeId?: string; insuranceProviderId?: string }> {
const out: { policyTypeId?: string; insuranceProviderId?: string } = {};
if (item.policyTypeId) {
out.policyTypeId = item.policyTypeId;
} else if (doc.extractedPolicyTypeName) {
const row = await this.prisma.policyType.findUnique({
where: { name: doc.extractedPolicyTypeName },
select: { id: true },
});
if (row) out.policyTypeId = row.id;
}
if (item.insuranceProviderId) {
out.insuranceProviderId = item.insuranceProviderId;
} else if (doc.provider) {
const name = PROVIDER_ROW_NAME[doc.provider] ?? doc.provider;
const row = await this.prisma.insuranceProvider.findFirst({
where: { name },
select: { id: true },
});
if (row) out.insuranceProviderId = row.id;
}
return out;
}
/**
* Write the parsed `Vehicle` and `InsuredDriver` rows onto the policy.
*
* Both inserts are skipped when an equivalent row is already on the policy.
* The reason is `confirmBatch` applying to an EXISTING policy: the office
* uploads a renewal for a car already on file, and a blind insert would
* leave the customer with the same VIN listed twice with no way to tell
* which row the renewal belongs to. Matching is on the identifier the
* document actually prints — the VIN for a vehicle (falling back to the
* plate, since ANA's TRAILER/TOWING slots have no VIN), the licence number
* for a driver (falling back to the name).
*
* Nothing is ever updated or deleted here. A vehicle whose plate changed
* lands as a second row for a human to reconcile, which is the safe half
* of the mistake: an over-write would destroy the only record of what was
* insured last term.
*/
private async applyVehiclesAndDrivers(
doc: { extractedVehiclesJson: Prisma.JsonValue | null; extractedDriversJson: Prisma.JsonValue | null },
policyId: string,
): Promise<void> {
const vehicles = asArray<ParsedVehicle>(doc.extractedVehiclesJson);
const drivers = asArray<ParsedDriver>(doc.extractedDriversJson);
if (vehicles.length === 0 && drivers.length === 0) return;
const policy = await this.prisma.policy.findUnique({
where: { id: policyId },
select: { customerId: true },
});
if (!policy) return;
if (vehicles.length) {
const existing = await this.prisma.vehicle.findMany({
where: { policyId },
select: { vinNumber: true, licensePlate: true },
});
const seen = new Set(
existing.flatMap((v) =>
[v.vinNumber, v.licensePlate].filter((k): k is string => !!k).map(norm),
),
);
for (const v of vehicles) {
const key = norm(v.vinNumber ?? v.licensePlate ?? "");
if (!key || seen.has(key)) continue;
seen.add(key);
await this.prisma.vehicle.create({
data: {
policyId,
customerId: policy.customerId,
make: v.make,
// ANA prints one BODY cell, not separate model/body columns, so
// it lands on `bodyType`; `model` stays null rather than being
// guessed out of the same string.
bodyType: v.bodyType,
modelYear: v.modelYear,
vinNumber: v.vinNumber,
licensePlate: v.licensePlate,
// "VEHICLE" / "TRAILER" / "TOWING" — the printed slot, which is
// the difference between the insured car and the trailer behind
// it and has no column of its own.
notes: v.item && v.item !== "VEHICLE" ? v.item : null,
},
});
}
}
if (drivers.length) {
const existing = await this.prisma.insuredDriver.findMany({
where: { policyId },
select: { licenseNumber: true, fullName: true },
});
const seen = new Set(
existing.flatMap((d) =>
[d.licenseNumber, d.fullName].filter((k): k is string => !!k).map(norm),
),
);
for (const d of drivers) {
const key = norm(d.licenseNumber ?? d.fullName ?? "");
if (!key || seen.has(key)) continue;
seen.add(key);
await this.prisma.insuredDriver.create({
data: { policyId, fullName: d.fullName, licenseNumber: d.licenseNumber },
});
}
}
}
/**
* Stream the source PDF (`sourceKey`, set by `process` on the doc row)
* into the policy's storage namespace and create a `PolicyDocument`
* pointer. Trivial now that the doc row holds the exact source key —
* the old per-page "which file did this page come from" walk is gone.
*/
private async attachSourcePdf(
sourceKey: string,
policyId: string,
provider: string | null,
): Promise<void> {
const got = await this.storage.getStream(sourceKey);
const chunks: Buffer[] = [];
for await (const c of got.stream) chunks.push(c as Buffer);
const buf = Buffer.concat(chunks);
const newKey = `policy/${policyId}/${Date.now()}-${crypto.randomUUID()}.pdf`;
await this.storage.put(newKey, buf, "application/pdf");
await this.prisma.policyDocument.create({
data: {
policyId,
// Named after whichever parser claimed the page. Was hardcoded
// `GMX_POLICY`, which mislabelled every ANA upload as a GMX
// document in the policy's file list.
documentType: `${provider ?? "OCR"}_POLICY`,
storageKey: newKey,
},
});
}
private async policyCustomerId(policyId: string): Promise<string | null> {
const p = await this.prisma.policy.findUnique({
where: { id: policyId },
select: { customerId: true },
});
return p?.customerId ?? null;
}
private async closeIfDone(batchId: string) {
const open = await this.prisma.policyOcrDocument.count({
where: {
batchId,
status: { in: ["PENDING_OCR", "NEEDS_REVIEW", "MATCHED", "CONFIRMED"] },
},
});
if (open === 0) {
await this.prisma.policyOcrBatch.updateMany({
// `updateMany` + a status filter so a discarded batch is never quietly
// relabelled COMPLETED by a late reject on one of its pages.
where: { id: batchId, status: { not: "DISCARDED" } },
data: { status: "COMPLETED", completedAt: new Date() },
});
}
}
}
/**
* The parser's provider code is not the carrier's row name in
* `insurance_providers`, and the two namespaces are allowed to differ.
*
* ANA is the case that forces this: the office's book is filed under
* "ANA SEGUROS" (738 policies). A bare "ANA" row also existed with 1 policy
* and is merged away by `20260815160000_policy_type_repair`, so an exact-name
* lookup on the parser's "ANA" would find nothing at all after that migration.
*
* Anything not listed resolves by its own name.
*/
const PROVIDER_ROW_NAME: Record<string, string> = {
ANA: "ANA SEGUROS",
};
/** The lookup FKs resolved for one document, absent when unresolvable. */
interface ResolvedLookups {
policyTypeId?: string;
insuranceProviderId?: string;
}
/** Map a (post-review) doc + final confirmed fields onto a `Policy.update`
* payload. Every field that is null in both inputs is omitted so we never
* write null over a value the Policy already carries (the GMX certificate
* has no premium — we must not blank the existing Policy.netPremium). */
function buildPolicyUpdateFromDoc(
item: ConfirmPolicyDocumentDto,
doc: {
extractedPolicyNumber: string | null;
extractedInsuredName: string | null;
extractedAdditionalInsured: string | null;
extractedAgentName: string | null;
extractedLegalAddress: string | null;
extractedZip: string | null;
extractedPolicyFrom: Date | null;
extractedPolicyTo: Date | null;
extractedPolicyDate: Date | null;
extractedCurrency: string | null;
extractedNetPremium: Prisma.Decimal | null;
extractedPolicyFee: Prisma.Decimal | null;
extractedBrokerFee: Prisma.Decimal | null;
extractedTotal: Prisma.Decimal | null;
extractedCoveragesJson: Prisma.JsonValue | null;
extractedPremiumPayment: string | null;
extractedCoveragePeriodDays: number | null;
},
lookups: ResolvedLookups,
): Prisma.PolicyUpdateInput {
const numOrUndef = (a: number | undefined, b: Prisma.Decimal | null): Prisma.Decimal | undefined => {
if (a != null) return new Prisma.Decimal(a);
if (b != null) return b;
return undefined;
};
const dateOrUndef = (a: string | undefined, b: Date | null): Date | undefined => {
if (a) return new Date(a);
if (b) return b;
return undefined;
};
const strOrUndef = (a: string | undefined, b: string | null): string | undefined => {
if (a != null && a !== "") return a;
if (b != null && b !== "") return b;
return undefined;
};
return {
policyNumber: strOrUndef(item.policyNumber, doc.extractedPolicyNumber),
// `connect` rather than a raw id: this is the CHECKED update input. Left
// undefined when unresolved, so an existing Policy never loses a type or
// carrier it already had because this document could not name one.
policyType: lookups.policyTypeId ? { connect: { id: lookups.policyTypeId } } : undefined,
insuranceProvider: lookups.insuranceProviderId
? { connect: { id: lookups.insuranceProviderId } }
: undefined,
agentName: strOrUndef(item.agentName, doc.extractedAgentName),
policyFrom: dateOrUndef(item.policyFrom, doc.extractedPolicyFrom),
policyTo: dateOrUndef(item.policyTo, doc.extractedPolicyTo),
policyDate: dateOrUndef(item.policyDate, doc.extractedPolicyDate),
// Left undefined when the document didn't print a term, so the schema
// default (365) stands for GMX. ANA's by-the-day policies DO print one,
// and the default would otherwise turn a 4-day tourist policy into an
// annual one on the renewals screen.
coveragePeriodDays:
item.coveragePeriodDays ?? doc.extractedCoveragePeriodDays ?? undefined,
currency: strOrUndef(item.currency, doc.extractedCurrency) as Currency | undefined,
netPremium: numOrUndef(item.netPremium, doc.extractedNetPremium),
policyFee: numOrUndef(item.policyFee, doc.extractedPolicyFee),
brokerFee: numOrUndef(item.brokerFee, doc.extractedBrokerFee),
total: numOrUndef(item.total, doc.extractedTotal),
// coveragesJson / observations: freeform, keep the GMX data when present.
coveragesJson:
item.coveragesJson !== undefined
? (item.coveragesJson as Prisma.InputJsonValue)
: doc.extractedCoveragesJson != null
? (doc.extractedCoveragesJson as Prisma.InputJsonValue)
: undefined,
// Premium payment cadence ("CONTADO") and insured-name fields land in
// `observations` so the PolicyForm's edits stay the source of truth for
// structured fields. The reviewer can move them by hand if needed.
observations: joinObservations(
doc.extractedInsuredName,
doc.extractedAdditionalInsured,
doc.extractedLegalAddress,
doc.extractedZip,
doc.extractedPremiumPayment,
item,
),
};
}
/** Same shape as `buildPolicyUpdateFromDoc`, but for `Policy.create`. The
* `customerId` is supplied separately and `policyNumber` is required (a
* Policy without a number can't be re-matched by the OCR pipeline). */
function buildPolicyCreateFromDoc(
item: ConfirmPolicyDocumentDto,
doc: {
extractedPolicyNumber: string | null;
extractedInsuredName: string | null;
extractedAdditionalInsured: string | null;
extractedAgentName: string | null;
extractedLegalAddress: string | null;
extractedZip: string | null;
extractedPolicyFrom: Date | null;
extractedPolicyTo: Date | null;
extractedPolicyDate: Date | null;
extractedCurrency: string | null;
extractedNetPremium: Prisma.Decimal | null;
extractedPolicyFee: Prisma.Decimal | null;
extractedBrokerFee: Prisma.Decimal | null;
extractedTotal: Prisma.Decimal | null;
extractedCoveragesJson: Prisma.JsonValue | null;
extractedPremiumPayment: string | null;
extractedCoveragePeriodDays: number | null;
},
customerId: string,
lookups: ResolvedLookups,
): Prisma.PolicyUncheckedCreateInput {
const numOrUndef = (a: number | undefined, b: Prisma.Decimal | null): Prisma.Decimal | undefined => {
if (a != null) return new Prisma.Decimal(a);
if (b != null) return b;
return undefined;
};
const dateOrUndef = (a: string | undefined, b: Date | null): Date | undefined => {
if (a) return new Date(a);
if (b) return b;
return undefined;
};
const strOrUndef = (a: string | undefined, b: string | null): string | undefined => {
if (a != null && a !== "") return a;
if (b != null && b !== "") return b;
return undefined;
};
const policyNumber =
strOrUndef(item.policyNumber, doc.extractedPolicyNumber);
if (!policyNumber) {
// Caller already guards this; the throw is a type-narrowing aid.
throw new Error("policyNumber required for create");
}
return {
policyNumber,
customerId,
policyTypeId: lookups.policyTypeId,
insuranceProviderId: lookups.insuranceProviderId,
agentName: strOrUndef(item.agentName, doc.extractedAgentName),
policyFrom: dateOrUndef(item.policyFrom, doc.extractedPolicyFrom),
policyTo: dateOrUndef(item.policyTo, doc.extractedPolicyTo),
policyDate: dateOrUndef(item.policyDate, doc.extractedPolicyDate),
// Left undefined when the document didn't print a term, so the schema
// default (365) stands for GMX. ANA's by-the-day policies DO print one,
// and the default would otherwise turn a 4-day tourist policy into an
// annual one on the renewals screen.
coveragePeriodDays:
item.coveragePeriodDays ?? doc.extractedCoveragePeriodDays ?? undefined,
currency: strOrUndef(item.currency, doc.extractedCurrency) as Currency | undefined,
netPremium: numOrUndef(item.netPremium, doc.extractedNetPremium),
policyFee: numOrUndef(item.policyFee, doc.extractedPolicyFee),
brokerFee: numOrUndef(item.brokerFee, doc.extractedBrokerFee),
total: numOrUndef(item.total, doc.extractedTotal),
coveragesJson:
item.coveragesJson !== undefined
? (item.coveragesJson as Prisma.InputJsonValue)
: doc.extractedCoveragesJson != null
? (doc.extractedCoveragesJson as Prisma.InputJsonValue)
: undefined,
observations: joinObservations(
doc.extractedInsuredName,
doc.extractedAdditionalInsured,
doc.extractedLegalAddress,
doc.extractedZip,
doc.extractedPremiumPayment,
item,
),
};
}
function joinObservations(
insured: string | null,
additional: string | null,
address: string | null,
zip: string | null,
premiumPayment: string | null,
item: ConfirmPolicyDocumentDto,
): string | undefined {
const lines: string[] = [];
const insuredName = strOrUndefDb(item.insuredName, insured);
if (insuredName) lines.push(`Asegurado: ${insuredName}`);
const additionalInsured = strOrUndefDb(item.additionalInsured, additional);
if (additionalInsured) lines.push(`Asegurado adicional: ${additionalInsured}`);
const legalAddress = strOrUndefDb(item.legalAddress, address);
if (legalAddress) lines.push(`Dirección: ${legalAddress}`);
const zipVal = strOrUndefDb(item.zip, zip);
if (zipVal) lines.push(`C.P.: ${zipVal}`);
const cadence = strOrUndefDb(item.premiumPayment, premiumPayment);
if (cadence) lines.push(`Pago de prima: ${cadence}`);
return lines.length ? lines.join("\n") : undefined;
}
function strOrUndefDb(a: string | undefined, b: string | null): string | undefined {
if (a != null && a !== "") return a;
if (b != null && b !== "") return b;
return undefined;
}
/** A JSON column the parser wrote as an array, read back as one. Anything
* else (null, DbNull, a legacy object shape) is an empty list rather than a
* crash — these columns are only ever populated by the parser, so a
* surprise shape means old data, not a caller to reject. */
function asArray<T>(value: Prisma.JsonValue | null): T[] {
return Array.isArray(value) ? (value as unknown as T[]) : [];
}
/** Compare identifiers the way a person would: case- and space-insensitive.
* VINs and plates are printed inconsistently ("8BPX206" vs "8BPX 206"). */
function norm(s: string): string {
return s.replace(/\s+/g, "").toUpperCase();
}
@@ -9,10 +9,16 @@ import {
Put, Put,
Query, Query,
Req, Req,
Res,
StreamableFile,
UploadedFile,
UseGuards, UseGuards,
UseInterceptors,
} from "@nestjs/common"; } from "@nestjs/common";
import { FileInterceptor } from "@nestjs/platform-express";
import { ServiceKind } from "@jorgecuadros/database"; import { ServiceKind } from "@jorgecuadros/database";
import { Request } from "express"; import { Request, Response } from "express";
import { downloadName, type UploadedFileLike } from "../storage/upload-file";
import { AuthenticatedGuard } from "../auth/authenticated.guard"; import { AuthenticatedGuard } from "../auth/authenticated.guard";
import { AbilityGuard } from "../auth/ability.guard"; import { AbilityGuard } from "../auth/ability.guard";
import { RequireAbility } from "../auth/require-ability.decorator"; import { RequireAbility } from "../auth/require-ability.decorator";
@@ -196,7 +202,35 @@ export class PropertiesController {
return this.properties.removeTrust(id); return this.properties.removeTrust(id);
} }
// --- documents (remove pointer only) -------------------------------------- // --- documents ------------------------------------------------------------
@Post(":id/documents")
@RequireAbility("property:update")
@UseInterceptors(
FileInterceptor("file", { limits: { fileSize: 50 * 1024 * 1024 } }),
)
addDocument(
@Param("id") id: string,
@UploadedFile() file: UploadedFileLike | undefined,
@Query("type") type: string | undefined,
) {
if (!file) throw new Error("No se recibió ningún archivo.");
return this.properties.addDocument(id, file, type);
}
@Get(":id/documents/:childId/download")
async downloadDocument(
@Param("id") id: string,
@Param("childId") childId: string,
@Res({ passthrough: true }) res: Response,
): Promise<StreamableFile> {
const { row, stream, contentType } = await this.properties.getDocument(id, childId);
res.set({
"Content-Type": contentType ?? "application/octet-stream",
"Content-Disposition": `attachment; filename="${downloadName(row.storageKey, row.documentType)}"`,
});
return new StreamableFile(stream);
}
@Delete(":id/documents/:childId") @Delete(":id/documents/:childId")
@RequireAbility("property:update") @RequireAbility("property:update")
+40 -5
View File
@@ -1,6 +1,9 @@
import { Injectable, NotFoundException } from "@nestjs/common"; import { Injectable, NotFoundException } from "@nestjs/common";
import { randomUUID } from "node:crypto";
import { Prisma, ServiceKind } from "@jorgecuadros/database"; import { Prisma, ServiceKind } from "@jorgecuadros/database";
import { PrismaService } from "../prisma/prisma.service"; import { PrismaService } from "../prisma/prisma.service";
import { StorageService } from "../storage/storage.service";
import { extForUpload } from "../storage/upload-file";
import { toDate } from "../common/coerce"; import { toDate } from "../common/coerce";
import { import {
CreatePropertyDto, CreatePropertyDto,
@@ -83,7 +86,10 @@ function daysUntil(dueDate: Date | null | undefined, from: Date): number | null
@Injectable() @Injectable()
export class PropertiesService { export class PropertiesService {
constructor(private readonly prisma: PrismaService) {} constructor(
private readonly prisma: PrismaService,
private readonly storage: StorageService,
) {}
private trustWhere( private trustWhere(
trust: TrustFilter | undefined, trust: TrustFilter | undefined,
@@ -543,16 +549,45 @@ export class PropertiesService {
} }
// --- documents ------------------------------------------------------------ // --- documents ------------------------------------------------------------
// Removing a pointer row only; uploading files needs the object-storage // The blob lives in object storage (MinIO); the row is just the pointer. Keys
// client wired into the API (today only the migration writes to MinIO). // stay under the `service/<propertyId>/…` prefix the migration established.
async addDocument(
propertyId: string,
file: { buffer: Buffer; originalname?: string; mimetype?: string },
documentType?: string,
) {
await this.ensureProperty(propertyId);
const ext = extForUpload(file);
const key = `service/${propertyId}/${randomUUID()}${ext}`;
await this.storage.put(key, file.buffer, file.mimetype);
return this.prisma.serviceDocument.create({
data: {
propertyId,
documentType: documentType?.trim() || "DOCUMENT",
storageKey: key,
},
});
}
async getDocument(propertyId: string, id: string) {
const row = await this.prisma.serviceDocument.findFirst({
where: { id, propertyId },
});
if (!row) throw new NotFoundException(`Document ${id} not found on property ${propertyId}`);
const blob = await this.storage.getStream(row.storageKey);
return { row, ...blob };
}
async removeDocument(propertyId: string, id: string) { async removeDocument(propertyId: string, id: string) {
await this.ensureProperty(propertyId); await this.ensureProperty(propertyId);
const row = await this.prisma.serviceDocument.findFirst({ const row = await this.prisma.serviceDocument.findFirst({
where: { id, propertyId }, where: { id, propertyId },
select: { id: true }, select: { id: true, storageKey: true },
}); });
if (!row) throw new NotFoundException(`Document ${id} not found on property ${propertyId}`); if (!row) throw new NotFoundException(`Document ${id} not found on property ${propertyId}`);
return this.prisma.serviceDocument.delete({ where: { id } }); const deleted = await this.prisma.serviceDocument.delete({ where: { id } });
await this.storage.delete(row.storageKey);
return deleted;
} }
} }
@@ -0,0 +1,64 @@
import type { RenewalLetterRow } from "../reports/renewal-letter";
import { renderRenewalEmail } from "./renewal-email";
function letter(overrides: Partial<RenewalLetterRow> = {}): RenewalLetterRow {
return {
__kind: "letter",
policyId: "policy-1",
policyNumber: "POL-123",
policyType: "AUTO",
customerName: "Ana Pérez",
customerEmail: "ana@example.com",
customerPhone: "664-111-2222",
customerMobile: null,
customerAddress: ["Calle Uno 123", "Tijuana, BC, 22000"],
provider: "Aseguradora Uno",
policyTo: "2026-09-01",
netPremium: "1200.00",
policyFee: null,
total: "1392.00",
currency: "MXN",
coverageDays: null,
cslLimit: null,
medicalCoverage: null,
propertyDamage: null,
perPersonLiability: null,
additionalService: null,
vehicle: null,
generation: 1,
sentAt: null,
...overrides,
};
}
describe("renderRenewalEmail", () => {
it("includes policy, premium, expiration, type, and customer information", () => {
const result = renderRenewalEmail(letter());
expect(result.subject).toContain("POL-123");
expect(result.html).toContain("primer aviso");
expect(result.html).toContain("AUTO");
expect(result.html).toContain("01/09/2026");
expect(result.html).toContain("1,392.00");
expect(result.html).toContain("Ana Pérez");
expect(result.html).toContain("ana@example.com");
expect(result.html).toContain("664-111-2222");
expect(result.html).toContain("Calle Uno 123");
});
it("uses overdue wording for generation three", () => {
const result = renderRenewalEmail(letter({ generation: 3 }));
expect(result.subject).toContain("Póliza vencida");
expect(result.html).toContain("está vencida");
});
it("escapes customer-provided HTML", () => {
const result = renderRenewalEmail(
letter({ customerName: '<img src=x onerror="alert(1)">' }),
);
expect(result.html).not.toContain("<img");
expect(result.html).toContain("&lt;img");
});
});
+65
View File
@@ -0,0 +1,65 @@
import type { RenewalLetterRow } from "../reports/renewal-letter";
const GENERATION_TEXT: Record<number, string> = {
1: "Le enviamos el primer aviso para renovar su póliza.",
2: "Le enviamos el segundo aviso para renovar su póliza.",
3: "Le informamos que su póliza está vencida.",
};
function escapeHtml(value: unknown): string {
return String(value ?? "")
.replaceAll("&", "&amp;")
.replaceAll("<", "&lt;")
.replaceAll(">", "&gt;")
.replaceAll('"', "&quot;")
.replaceAll("'", "&#039;");
}
function displayDate(value: string): string {
if (value === "—") return value;
const [year, month, day] = value.split("-");
return `${day}/${month}/${year}`;
}
function money(value: string | null, currency: string): string {
if (!value) return "No disponible";
return new Intl.NumberFormat("es-MX", {
style: "currency",
currency,
minimumFractionDigits: 2,
}).format(Number(value));
}
function row(label: string, value: string): string {
return `<tr><th style="padding:8px 12px;text-align:left;background:#f4f4f4;border:1px solid #ddd">${escapeHtml(label)}</th><td style="padding:8px 12px;border:1px solid #ddd">${escapeHtml(value)}</td></tr>`;
}
export function renderRenewalEmail(letter: RenewalLetterRow): {
subject: string;
html: string;
} {
const expired = letter.generation === 3;
const subject = expired
? `Póliza vencida: ${letter.policyNumber}`
: `Aviso de renovación: póliza ${letter.policyNumber}`;
const phone = letter.customerMobile ?? letter.customerPhone ?? "No disponible";
const address = letter.customerAddress.join(", ") || "No disponible";
const premium = letter.total ?? letter.netPremium;
const details = [
row("Número de póliza", letter.policyNumber),
row("Tipo de póliza", letter.policyType),
row("Aseguradora", letter.provider),
row("Fecha de vencimiento", displayDate(letter.policyTo)),
row("Prima", money(premium, letter.currency)),
row("Cliente", letter.customerName),
row("Correo", letter.customerEmail ?? "No disponible"),
row("Teléfono", phone),
row("Dirección", address),
].join("");
return {
subject,
html: `<div style="font-family:Arial,sans-serif;color:#222;line-height:1.5"><p>Estimado(a) ${escapeHtml(letter.customerName)}:</p><p>${escapeHtml(GENERATION_TEXT[letter.generation] ?? "Le enviamos un aviso sobre la renovación de su póliza.")}</p><table style="border-collapse:collapse;width:100%;max-width:680px">${details}</table><p>Por favor, comuníquese con Jorge Cuadros &amp; Asociados para revisar su renovación.</p><p>Atentamente,<br>Jorge Cuadros &amp; Asociados</p></div>`,
};
}
+197
View File
@@ -0,0 +1,197 @@
import { RenewalsService } from "./renewals.service";
/**
* The renewal sweep's half of the unified notification log.
*
* `RenewalNotice` only records that a policy WAS notified — it has no way to
* say a send failed or that a customer had no address. Those rows exist only
* in `email_notification_log`, so they are what these tests pin down.
*/
const POLICY_ID = "policy-1";
const CUSTOMER_ID = "cust-1";
function makePolicy(email: string | null) {
return {
id: POLICY_ID,
policyNumber: "700442181",
policyTo: new Date("2026-09-01T00:00:00.000Z"),
netPremium: null,
policyFee: null,
total: null,
currency: "MXN",
coveragesJson: null,
customer: {
id: CUSTOMER_ID,
name: "ACME SA DE CV",
nameMissing: false,
email,
phone: null,
mobile: null,
addressLine1: null,
addressLine2: null,
city: null,
state: null,
zipCode: null,
country: null,
},
policyType: { name: "AUTO" },
insuranceProvider: { name: "GMX" },
vehicles: [],
renewalNotices: [],
};
}
function build(overrides: {
policies?: ReturnType<typeof makePolicy>[];
sendImpl?: () => Promise<{ messageId: string; response: string }>;
}) {
const policies = overrides.policies ?? [makePolicy("cliente@example.com")];
const record = jest.fn().mockResolvedValue(undefined);
const send =
overrides.sendImpl ??
jest.fn().mockResolvedValue({ messageId: "ses-1", response: "{}" });
const prisma = {
// Only generation 1 has a candidate; the other two cadences return none,
// so a sweep produces exactly one outcome to assert on.
policy: {
findMany: jest
.fn()
.mockResolvedValueOnce(policies)
.mockResolvedValue([]),
findFirst: jest.fn().mockResolvedValue(policies[0]),
},
renewalNotice: { upsert: jest.fn().mockResolvedValue({}) },
scheduledJobState: {
upsert: jest.fn().mockResolvedValue({}),
updateMany: jest.fn().mockResolvedValue({ count: 1 }),
findUniqueOrThrow: jest.fn().mockResolvedValue({ lastSuccessfulAt: null }),
update: jest.fn().mockResolvedValue({}),
},
};
// `register` is a no-op here: these tests drive the sweep directly, so no
// cron job is ever installed.
const schedule = { register: jest.fn().mockResolvedValue(undefined) };
const service = new RenewalsService(
prisma as never,
{ available: true, send } as never,
{ log: jest.fn() } as never,
{ record } as never,
schedule as never,
);
return { service, record, send, prisma };
}
describe("renewal notices write the shared notification log", () => {
it("records a SENT row tagged RENEWAL_NOTICE / POLICIES", async () => {
const { service, record, prisma } = build({});
await service.sweep("user-1");
expect(record).toHaveBeenCalledTimes(1);
const row = record.mock.calls[0][0];
expect(row).toMatchObject({
notificationType: "RENEWAL_NOTICE",
servicio: "POLICIES",
status: "SENT",
customerId: CUSTOMER_ID,
customerEmail: "cliente@example.com",
providerMessageId: "ses-1",
debug: false,
});
// `level` carries the aviso generation, not an alert colour.
expect(row.level).toBe(1);
expect(row.subject).toContain("700442181");
expect(row.bodySnapshot).toContain("ACME SA DE CV");
// The gating row is still written — the log does not replace it.
expect(prisma.renewalNotice.upsert).toHaveBeenCalledTimes(1);
});
it("records a FAILED row and no gating row when the send throws", async () => {
const { service, record, prisma } = build({
sendImpl: jest.fn().mockRejectedValue(new Error("SES rejected")),
});
const result = await service.sweep("user-1");
expect(result.sent).toBe(0);
expect(result.failed).toBe(1);
expect(record).toHaveBeenCalledTimes(1);
expect(record.mock.calls[0][0]).toMatchObject({
status: "FAILED",
error: "SES rejected",
notificationType: "RENEWAL_NOTICE",
});
// Nothing was delivered, so nothing may gate tomorrow's retry.
expect(prisma.renewalNotice.upsert).not.toHaveBeenCalled();
});
it("records SKIPPED_NO_EMAIL for a candidate with no address", async () => {
const { service, record, send, prisma } = build({
policies: [makePolicy(" ")],
});
const result = await service.sweep("user-1");
expect(result.skipped).toBe(1);
expect(send).not.toHaveBeenCalled();
expect(prisma.renewalNotice.upsert).not.toHaveBeenCalled();
expect(record.mock.calls[0][0]).toMatchObject({
status: "SKIPPED_NO_EMAIL",
customerEmail: "",
});
});
it("diverts a debug sweep and leaves the notice pending", async () => {
const { service, record, send, prisma } = build({});
const result = await service.sweep("user-1", { debug: true });
expect(result.sent).toBe(1);
expect(result.debug).toBe(true);
// The customer's own address is never contacted.
expect(jest.mocked(send).mock.calls[0][0]).toMatchObject({
to: "rmancinas@freakma.net",
xTracking: "debug",
});
expect(record.mock.calls[0][0]).toMatchObject({
status: "SENT",
customerEmail: "rmancinas@freakma.net",
debug: true,
});
// The letter is still owed, so nothing may gate it: no RenewalNotice row,
// and `lastSuccessfulAt` must not advance past the days we only tested.
expect(prisma.renewalNotice.upsert).not.toHaveBeenCalled();
const release = prisma.scheduledJobState.update.mock.calls.at(-1)?.[0];
expect(release.data.lastSuccessfulAt).toBeUndefined();
});
it("sends one notice on demand in debug without marking it sent", async () => {
const { service, send, prisma } = build({});
const result = await service.sendOne(POLICY_ID, 1, "user-1", {
debug: true,
});
expect(result.debug).toBe(true);
expect(result.to).toBe("rmancinas@freakma.net");
expect(send).toHaveBeenCalledTimes(1);
expect(prisma.renewalNotice.upsert).not.toHaveBeenCalled();
});
it("does not fail a delivered notice when the log write throws", async () => {
const { service, record } = build({});
record.mockRejectedValue(new Error("log table gone"));
const result = await service.sweep("user-1");
// The mail went out and the gating row was written; a lost audit row must
// not report that as a failure, which would re-send tomorrow.
expect(result.sent).toBe(1);
expect(result.failed).toBe(0);
});
});
@@ -0,0 +1,72 @@
import {
Body,
Controller,
Get,
HttpCode,
Post,
Query,
Req,
UseGuards,
} from "@nestjs/common";
import { Request } from "express";
import { Type } from "class-transformer";
import { IsBoolean, IsInt, IsOptional, IsString, Max, Min } from "class-validator";
import { AbilityGuard } from "../auth/ability.guard";
import { AuthenticatedGuard } from "../auth/authenticated.guard";
import { RequireAbility } from "../auth/require-ability.decorator";
import { RenewalsService } from "./renewals.service";
/** The pólizas half of the shared "Flags del envío" panel. Only `debug`
* means anything here — the day gate and the send limit are estado-de-cuenta
* concepts — so the other two are simply not accepted. */
class RenewalFlagsDto {
@IsOptional()
@IsBoolean()
debug?: boolean;
}
class SendRenewalDto extends RenewalFlagsDto {
@IsString()
policyId!: string;
/** 1 = 30 días antes, 2 = 15 días antes, 3 = 7 días después. */
@Type(() => Number)
@IsInt()
@Min(1)
@Max(3)
generation!: number;
}
@UseGuards(AuthenticatedGuard, AbilityGuard)
@Controller("renewals")
export class RenewalsController {
constructor(private readonly renewals: RenewalsService) {}
@Get("pending")
pending(@Query("days") days?: string) {
return this.renewals.pending(
Math.min(365, Math.max(1, Number(days) || 30)),
);
}
@Post("sweep")
@RequireAbility("renewal:send")
sweep(@Body() dto: RenewalFlagsDto, @Req() req: Request) {
return this.renewals.sweep((req.user as { id: string }).id, {
debug: dto?.debug,
});
}
/** Send a single pending notice from the /notificaciones list. */
@Post("send")
@RequireAbility("renewal:send")
@HttpCode(200)
send(@Body() dto: SendRenewalDto, @Req() req: Request) {
return this.renewals.sendOne(
dto.policyId,
dto.generation,
(req.user as { id: string }).id,
{ debug: dto.debug },
);
}
}
+15
View File
@@ -0,0 +1,15 @@
import { Module } from "@nestjs/common";
import { NotificationLogModule } from "../notifications/notification-log.module";
import { NotificationScheduleModule } from "../notifications/notification-schedule.module";
import { RenewalsController } from "./renewals.controller";
import { RenewalsService } from "./renewals.service";
@Module({
// Renewal sends write to the same `email_notification_log` the four bulk
// jobs write, so /notificaciones has one send history across both tabs, and
// take their cadence from the same operator-editable schedule.
imports: [NotificationLogModule, NotificationScheduleModule],
controllers: [RenewalsController],
providers: [RenewalsService],
})
export class RenewalsModule {}
@@ -0,0 +1,42 @@
import {
addUtcDays,
dateInTimeZone,
renewalWindow,
RENEWAL_CADENCE,
} from "./renewals.service";
describe("renewal scheduling dates", () => {
it("uses the America/Tijuana calendar date", () => {
expect(dateInTimeZone(new Date("2026-08-01T05:00:00.000Z"))).toEqual(
new Date("2026-07-31T00:00:00.000Z"),
);
});
it("maps generations to 30 days, 15 days, and 7 days overdue", () => {
const today = new Date("2026-08-01T00:00:00.000Z");
expect(
RENEWAL_CADENCE.map(({ generation, offsetDays }) => ({
generation,
target: addUtcDays(today, offsetDays).toISOString().slice(0, 10),
})),
).toEqual([
{ generation: 1, target: "2026-08-31" },
{ generation: 2, target: "2026-08-16" },
{ generation: 3, target: "2026-07-25" },
]);
});
it("uses an inclusive catch-up window after a missed run", () => {
const window = renewalWindow(
new Date("2026-08-10T00:00:00.000Z"),
30,
new Date("2026-08-07T18:00:00.000Z"),
);
expect(window).toEqual({
from: new Date("2026-09-07T00:00:00.000Z"),
to: new Date("2026-09-09T00:00:00.000Z"),
});
});
});
+449
View File
@@ -0,0 +1,449 @@
import {
BadRequestException,
ConflictException,
Injectable,
Logger,
NotFoundException,
OnModuleInit,
ServiceUnavailableException,
} from "@nestjs/common";
import { AuditService } from "../common/audit.service";
import { MailService } from "../mail/mail.service";
import { NotificationLogService } from "../notifications/notification-log.service";
import {
NotificationScheduleService,
SCHEDULE_TIME_ZONE,
} from "../notifications/notification-schedule.service";
import { DEBUG_RECIPIENT } from "../notifications/notification.types";
import { PrismaService } from "../prisma/prisma.service";
import {
RenewalLetterPolicy,
renewalLetterSelect,
toRenewalLetterRow,
} from "../reports/renewal-letter";
import { renderRenewalEmail } from "./renewal-email";
export const RENEWAL_CADENCE = [
{ generation: 1, offsetDays: 30 },
{ generation: 2, offsetDays: 15 },
{ generation: 3, offsetDays: -7 },
] as const;
const JOB_NAME = "renewal-email-sweep";
/** The window maths runs in office time; the cadence itself is owned by
* `NotificationScheduleService`, which uses the same zone. */
const TIME_ZONE = SCHEDULE_TIME_ZONE;
const DAY_MS = 86400000;
export function dateInTimeZone(now: Date, timeZone = TIME_ZONE): Date {
const parts = new Intl.DateTimeFormat("en-US", {
timeZone,
year: "numeric",
month: "2-digit",
day: "2-digit",
}).formatToParts(now);
const value = (type: Intl.DateTimeFormatPartTypes) =>
Number(parts.find((part) => part.type === type)?.value);
return new Date(Date.UTC(value("year"), value("month") - 1, value("day")));
}
export function addUtcDays(date: Date, days: number): Date {
return new Date(date.getTime() + days * DAY_MS);
}
export function renewalWindow(
today: Date,
offsetDays: number,
lastSuccessfulAt?: Date | null,
): { from: Date; to: Date } {
const to = addUtcDays(today, offsetDays);
if (!lastSuccessfulAt) return { from: to, to };
const previousDay = dateInTimeZone(lastSuccessfulAt);
if (previousDay >= today) return { from: to, to };
return { from: addUtcDays(previousDay, offsetDays + 1), to };
}
@Injectable()
export class RenewalsService implements OnModuleInit {
private readonly logger = new Logger(RenewalsService.name);
constructor(
private readonly prisma: PrismaService,
private readonly mail: MailService,
private readonly audit: AuditService,
private readonly notificationLog: NotificationLogService,
private readonly schedule: NotificationScheduleService,
) {}
/** The cadence used to be a `@Cron("0 6 * * *")` literal here; it is now
* operator-editable, and the stored value defaults to that same 06:00
* daily run. */
async onModuleInit(): Promise<void> {
await this.schedule.register("polizas", () => this.scheduledSweep());
}
/** The unattended run always sends for real: `debug` is a per-click switch
* in the UI, never persisted, so the schedule cannot inherit a forgotten
* test toggle and silently stop mailing customers. */
async scheduledSweep(): Promise<void> {
try {
await this.sweep();
} catch (error) {
this.logger.error(
`Falló el barrido de renovaciones: ${(error as Error).message}`,
);
}
}
async pending(days = 30) {
const today = dateInTimeZone(new Date());
const state = await this.prisma.scheduledJobState.findUnique({
where: { name: JOB_NAME },
select: { lastSuccessfulAt: true },
});
const cadence = RENEWAL_CADENCE.filter(
(item) => item.offsetDays < 0 || item.offsetDays <= days,
);
const groups = await Promise.all(
cadence.map(async (item) => ({
generation: item.generation,
rows: await this.findCandidates(
item,
today,
state?.lastSuccessfulAt ?? null,
),
})),
);
return groups.flatMap(({ generation, rows }) =>
rows
.filter((policy) => Boolean(policy.customer.email?.trim()))
.map((policy) => toRenewalLetterRow(policy, generation)),
);
}
async sweep(userId?: string, flags: { debug?: boolean } = {}) {
const debug = !!flags.debug;
const now = new Date();
const state = await this.acquireLock(now);
try {
if (!this.mail.available) {
throw new ServiceUnavailableException(
"El servicio de correo no está configurado.",
);
}
const today = dateInTimeZone(now);
let eligible = 0;
let sent = 0;
let skipped = 0;
const failures: Array<{ policyId: string; generation: number; error: string }> = [];
for (const cadence of RENEWAL_CADENCE) {
const policies = await this.findCandidates(
cadence,
today,
state.lastSuccessfulAt,
);
eligible += policies.length;
for (const policy of policies) {
const to = policy.customer.email?.trim();
if (!to) {
// Logged rather than silently counted: "we had nobody to mail"
// is a finding the office acts on, and only the log survives the
// HTTP response.
await this.recordLog(policy, cadence.generation, "", {
status: "SKIPPED_NO_EMAIL",
debug,
});
skipped++;
continue;
}
try {
await this.deliver(policy, cadence.generation, to, userId, debug);
sent++;
} catch (error) {
failures.push({
policyId: policy.id,
generation: cadence.generation,
error: (error as Error).message,
});
}
}
}
const result = {
eligible,
sent,
skipped,
failed: failures.length,
failures,
debug,
};
// A debug run must not advance `lastSuccessfulAt`: it wrote no
// RenewalNotice rows, so the days it "covered" are still owed, and
// narrowing tomorrow's window back to a single day would drop them.
await this.releaseLock(!debug && failures.length === 0 ? now : null);
void this.audit.log(userId, "renewalNotice.sweep", result);
return result;
} catch (error) {
await this.releaseLock(null);
throw error;
}
}
/**
* Send one pending renewal notice on demand, from the /notificaciones
* list. Same path the sweep takes — render, send, then record the notice —
* so a letter sent by hand is marked exactly like a swept one and drops
* off the pending list. Refuses a generation already sent so a double
* click can't mail the customer twice.
*
* Under `debug` the notice is NOT marked as sent, so the row stays in the
* pending list — the customer has still not been told anything.
*/
async sendOne(
policyId: string,
generation: number,
userId?: string,
flags: { debug?: boolean } = {},
) {
const debug = !!flags.debug;
if (!this.mail.available) {
throw new ServiceUnavailableException(
"El servicio de correo no está configurado.",
);
}
const policy = await this.prisma.policy.findFirst({
where: { id: policyId, archivedAt: null },
select: renewalLetterSelect(generation),
});
if (!policy) {
throw new NotFoundException("Póliza no encontrada.");
}
if (policy.renewalNotices.some((notice) => notice.sentAt)) {
throw new ConflictException("Este aviso ya fue enviado.");
}
const to = policy.customer.email?.trim();
if (!to) {
throw new BadRequestException("El cliente no tiene correo registrado.");
}
const { sentAt, providerMessageId, addressedTo } = await this.deliver(
policy,
generation,
to,
userId,
debug,
);
return {
policyId,
generation,
// The address the mail actually went to — under debug that is the
// override inbox, and the UI says so rather than claiming the customer
// was notified.
to: addressedTo,
debug,
sentAt: sentAt.toISOString(),
providerMessageId,
};
}
/** Render + send + record one notice. Shared by the sweep and `sendOne`.
*
* Two records come out of a send: the `RenewalNotice` row, which gates the
* pending list, and an `email_notification_log` row, which is the send
* history the /notificaciones "Registro de envíos" reads. A failed send
* writes only the second — there is no notice to gate on — and rethrows so
* the sweep counts it as a failure.
*
* Under `debug` the mail is diverted to `DEBUG_RECIPIENT` and the
* `RenewalNotice` row is deliberately skipped: the customer was not
* notified, so nothing may gate the letter they are still owed. Only the
* log row is written, flagged `debug`. */
private async deliver(
policy: RenewalLetterPolicy,
generation: number,
to: string,
userId?: string,
debug = false,
) {
const letter = toRenewalLetterRow(policy, generation);
const message = renderRenewalEmail(letter);
const addressedTo = debug ? DEBUG_RECIPIENT : to;
let result: Awaited<ReturnType<MailService["send"]>>;
try {
result = await this.mail.send({
to: addressedTo,
toName: letter.customerName,
subject: message.subject,
html: message.html,
xTracking: debug ? "debug" : "renewals",
});
} catch (error) {
const detail = error instanceof Error ? error.message : String(error);
await this.recordLog(policy, generation, addressedTo, {
status: "FAILED",
error: detail,
debug,
});
throw error;
}
const sentAt = new Date();
if (!debug) {
await this.prisma.renewalNotice.upsert({
where: {
policyId_generation: { policyId: policy.id, generation },
},
create: {
policyId: policy.id,
generation,
channel: "EMAIL",
sentAt,
sentById: userId,
providerMessageId: result.messageId,
},
update: {
channel: "EMAIL",
sentAt,
sentById: userId,
providerMessageId: result.messageId,
},
});
}
await this.recordLog(policy, generation, addressedTo, {
status: "SENT",
providerMessageId: result.messageId || undefined,
providerResponse: result.response || undefined,
sendDate: sentAt,
debug,
});
void this.audit.log(userId, "renewalNotice.send", {
policyId: policy.id,
generation,
debug,
providerMessageId: result.messageId,
});
return { sentAt, providerMessageId: result.messageId, addressedTo };
}
/**
* Write one row to the shared notification log.
*
* Never throws: the mail is already gone (or already failed) by the time we
* get here, and losing the audit row must not turn a delivered notice into
* a reported failure — which on the SENT path would also strand the
* `RenewalNotice` we just wrote and re-send tomorrow.
*/
private async recordLog(
policy: RenewalLetterPolicy,
generation: number,
/** Recipient as addressed. Empty on the SKIPPED_NO_EMAIL path — that
* emptiness IS the reason the row exists. */
to: string,
outcome: {
status: "SENT" | "FAILED" | "SKIPPED_NO_EMAIL";
providerMessageId?: string;
providerResponse?: string;
error?: string;
sendDate?: Date;
debug?: boolean;
},
): Promise<void> {
const letter = toRenewalLetterRow(policy, generation);
const message = renderRenewalEmail(letter);
try {
await this.notificationLog.record({
notificationType: "RENEWAL_NOTICE",
servicio: "POLICIES",
sendDate: outcome.sendDate,
// `level` carries the aviso generation for RENEWAL_NOTICE rows — see
// the column doc on the Prisma model.
level: generation,
customerId: policy.customer.id,
customerName: letter.customerName,
customerEmail: to,
subject: message.subject,
bodySnapshot: message.html,
status: outcome.status,
debug: !!outcome.debug,
providerMessageId: outcome.providerMessageId,
providerResponse: outcome.providerResponse,
error: outcome.error,
});
} catch (error) {
this.logger.warn(
`No se pudo registrar el aviso de renovación en el log ` +
`(póliza ${policy.id}, aviso ${generation}): ` +
`${(error as Error).message}`,
);
}
}
private findCandidates(
cadence: (typeof RENEWAL_CADENCE)[number],
today: Date,
lastSuccessfulAt: Date | null,
) {
const window = renewalWindow(today, cadence.offsetDays, lastSuccessfulAt);
return this.prisma.policy.findMany({
where: {
archivedAt: null,
policyTo: { gte: window.from, lte: window.to },
customer: {
archivedAt: null,
emailOptOut: false,
email: { not: "" },
},
renewalNotices: {
none: { generation: cadence.generation, sentAt: { not: null } },
},
},
orderBy: [{ policyTo: "asc" }, { policyNumber: "asc" }],
select: renewalLetterSelect(cadence.generation),
});
}
private async acquireLock(now: Date) {
await this.prisma.scheduledJobState.upsert({
where: { name: JOB_NAME },
create: { name: JOB_NAME },
update: { updatedAt: now },
});
const acquired = await this.prisma.scheduledJobState.updateMany({
where: {
name: JOB_NAME,
OR: [{ lockedUntil: null }, { lockedUntil: { lte: now } }],
},
data: { lockedUntil: new Date(now.getTime() + 2 * 60 * 60 * 1000) },
});
if (acquired.count !== 1) {
throw new ConflictException(
"Ya hay un barrido de renovaciones en curso.",
);
}
return this.prisma.scheduledJobState.findUniqueOrThrow({
where: { name: JOB_NAME },
});
}
private async releaseLock(lastSuccessfulAt: Date | null): Promise<void> {
await this.prisma.scheduledJobState.update({
where: { name: JOB_NAME },
data: {
lockedUntil: null,
...(lastSuccessfulAt && { lastSuccessfulAt }),
},
});
}
}
+103
View File
@@ -0,0 +1,103 @@
/**
* Company info used on every report header (PDF + print). Read from
* the environment so the office can edit it without a code change —
* the .env.example file lists the keys; defaults below are placeholders
* the office should override for production.
*
* Single source of truth: the API renders the header. The web header
* (login + AppShell) still reads the static "Jorge Cuadros & Asociados"
* strings for now — those are visual brand, the API's COMPANY_INFO
* block is the legal/locator block on printed documents.
*/
import * as fs from "node:fs";
import * as path from "node:path";
export interface CompanyInfo {
name: string;
/** Street address — line 1. */
addressLine1: string;
/** Street address — line 2 (suite, floor, etc.). Optional. */
addressLine2: string;
/** "City, State, ZIP, Country" — single line. */
cityState: string;
phone: string;
email: string;
/** Mexican tax ID ("RFC"). Optional. */
taxId: string;
website: string;
/** Absolute path to the logo PNG. Null when missing — renderers fall
* back to a text mark. */
logoPath: string | null;
/** Logo buffer + intrinsic size, eagerly loaded so the PDF renderer
* doesn't do a sync read on every report. Null when no logo. */
logo: { buffer: Buffer; width: number; height: number } | null;
}
function envOr(key: string, fallback: string): string {
const v = process.env[key];
return v && v.trim() ? v : fallback;
}
function resolveLogoPath(): string | null {
const explicit = process.env.COMPANY_LOGO_PATH;
if (explicit) {
return fs.existsSync(explicit) ? explicit : null;
}
// Default: look in apps/api/assets/company_logo.png (copied from
// apps/web/public/images/company_logo.png — single canonical image
// kept in lock-step; see .env.example for the override path).
const candidates = [
path.resolve(__dirname, "..", "..", "assets", "company_logo.png"),
path.resolve(__dirname, "..", "..", "..", "web", "public", "images", "company_logo.png"),
];
for (const c of candidates) {
if (fs.existsSync(c)) return c;
}
return null;
}
let cached: CompanyInfo | null = null;
export function getCompanyInfo(): CompanyInfo {
if (cached) return cached;
const logoPath = resolveLogoPath();
let logo: CompanyInfo["logo"] = null;
if (logoPath) {
try {
const buf = fs.readFileSync(logoPath);
// Intrinsic PNG size: read IHDR (bytes 16-23 of the file).
// Width = BE uint32 at offset 16, height = BE uint32 at offset 20.
const w =
logoPath.endsWith(".png") && buf.length >= 24
? buf.readUInt32BE(16)
: 0;
const h =
logoPath.endsWith(".png") && buf.length >= 24
? buf.readUInt32BE(20)
: 0;
logo = { buffer: buf, width: w, height: h };
} catch {
logo = null;
}
}
cached = {
name: envOr("COMPANY_NAME", "Jorge Cuadros & Asociados"),
addressLine1: envOr(
"COMPANY_ADDRESS_LINE1",
"Av. Revolución 1234, Int. 5",
),
addressLine2: envOr("COMPANY_ADDRESS_LINE2", ""),
cityState: envOr(
"COMPANY_CITY_STATE",
"Tijuana, Baja California 22000, México",
),
phone: envOr("COMPANY_PHONE", "(664) 000-0000"),
email: envOr("COMPANY_EMAIL", "contacto@jorgecuadros.local"),
taxId: envOr("COMPANY_TAX_ID", ""),
website: envOr("COMPANY_WEBSITE", "jorgecuadros.local"),
logoPath,
logo,
};
return cached;
}
+433
View File
@@ -0,0 +1,433 @@
/**
* Output renderers for the reports module.
*
* Every report's `run` returns `{ columns, rows, totals?, subtitle? }`.
* CSV/XLSX/PDF all derive from the same shape so adding a report = one
* registry entry, no per-format template.
*
* PDF uses pdfkit. The statement format (edo-cuenta-datos) uses a
* different layout than the tabular one — handled inline.
*/
// pdfkit exports its constructor via `module.exports = PDFDocument`, so a
// namespace import gets the type, and `import = require()` gets the value.
import PDFDocument = require("pdfkit");
import * as ExcelJS from "exceljs";
import type { ColumnDef, ReportResult } from "./reports.types";
import { getCompanyInfo } from "./company";
type Doc = PDFKit.PDFDocument;
/* ----------------------------------------------------------------- CSV */
function csvCell(v: unknown): string {
if (v === null || v === undefined) return "";
const s = String(v);
if (s.includes(",") || s.includes('"') || s.includes("\n") || s.includes("\r")) {
return `"${s.replace(/"/g, '""')}"`;
}
return s;
}
export function renderCsv(columns: ColumnDef[], result: ReportResult): string {
const headers = columns.map((c) => csvCell(c.label)).join(",");
const lines = result.rows.map((r) =>
columns
.map((c) => {
const v = r[c.key];
if (typeof v === "number") return v;
return csvCell(v);
})
.join(","),
);
const totals: string[] = [];
if (result.totals) {
for (const [k, v] of Object.entries(result.totals)) {
totals.push(csvCell(k), csvCell(v));
}
}
return [headers, ...lines, ...(totals.length ? [totals.join(",")] : [])].join(
"\n",
);
}
/* ----------------------------------------------------------------- XLSX */
export async function renderXlsx(
columns: ColumnDef[],
result: ReportResult,
): Promise<Buffer> {
const wb = new ExcelJS.Workbook();
wb.creator = "Jorge Cuadros & Asociados";
const ws = wb.addWorksheet("Reporte", {
views: [{ state: "frozen", ySplit: 1 }],
});
ws.columns = columns.map((c) => ({
header: c.label,
key: c.key,
width: Math.max(10, Math.min(40, (c.label.length + 2) * 1.2)),
}));
ws.getRow(1).font = { bold: true };
ws.getRow(1).fill = {
type: "pattern",
pattern: "solid",
fgColor: { argb: "FFE2EDE9" }, // brand-tint
};
for (const row of result.rows) {
ws.addRow(row);
}
// Number formatting for money columns.
for (const col of columns) {
if (col.type === "money" || col.type === "number") {
ws.getColumn(col.key).numFmt =
col.type === "money" ? "#,##0.00" : "#,##0";
ws.getColumn(col.key).alignment = { horizontal: "right" };
}
}
if (result.totals) {
const last = ws.addRow({});
let i = 1;
for (const [k, v] of Object.entries(result.totals)) {
const cell = ws.getCell(last.number, i);
cell.value = `${k}: ${v}`;
cell.font = { bold: true };
i++;
}
}
const buf = await wb.xlsx.writeBuffer();
return Buffer.from(buf);
}
/* ----------------------------------------------------------------- PDF */
const BRAND = "#0c322d";
const ACCENT = "#bf5a34";
const MUTED = "#756c5c";
const LINE = "#e4dccb";
function fmtMoney(v: unknown): string {
if (v === null || v === undefined || v === "") return "";
const n = Number(v);
if (!Number.isFinite(n)) return String(v);
return n.toLocaleString("es-MX", { minimumFractionDigits: 2, maximumFractionDigits: 2 });
}
function pdfRow(
doc: Doc,
y: number,
cols: Array<{ label: string; width: number; align?: "left" | "right" }>,
values: Array<{ text: string; align?: "left" | "right" }>,
x: number,
): number {
let cx = x;
for (let i = 0; i < cols.length; i++) {
const c = cols[i];
const v = values[i] ?? { text: "" };
const align = v.align ?? c.align ?? "left";
const w = c.width;
doc
.font("Helvetica")
.fontSize(9)
.fillColor("#211d17")
.text(v.text, cx, y, {
width: w - 4,
align,
ellipsis: true,
lineBreak: false,
height: 16,
});
cx += w;
}
return y + 18;
}
export function renderPdf(
columns: ColumnDef[],
result: ReportResult,
title: string,
): Promise<Buffer> {
return new Promise((resolve, reject) => {
const doc = new PDFDocument({
size: "LETTER",
layout: "landscape",
margins: { top: 96, bottom: 56, left: 48, right: 48 },
bufferPages: true,
info: {
Title: title,
Author: "Jorge Cuadros & Asociados",
Subject: "Reporte",
Creator: "Jorge Cuadros Platform — Reports module",
},
});
const chunks: Buffer[] = [];
doc.on("data", (c: Buffer) => chunks.push(c));
doc.on("end", () => resolve(Buffer.concat(chunks)));
doc.on("error", reject);
const company = getCompanyInfo();
const pageW = doc.page.width - 96;
/** The header is repeated on every page (via addPage + manual draw). */
const drawHeader = () => {
// Background bar (brand pine) for the masthead.
doc.rect(0, 0, doc.page.width, 60).fill(BRAND);
// Logo, fitted to a 40px box, with 8px padding.
let textX = 48;
if (company.logo) {
const targetH = 40;
const scale = targetH / company.logo.height;
const w = company.logo.width * scale;
doc.image(company.logo.buffer, 48, 10, { height: targetH });
textX = 48 + w + 14;
}
// Company name (large) + "Reporte" tag below it.
doc
.fillColor("#f5f1e8")
.font("Helvetica-Bold")
.fontSize(15)
.text(company.name, textX, 14, { width: pageW - (textX - 48), lineBreak: false });
doc
.font("Helvetica")
.fontSize(8)
.fillColor("#cde0db")
.text("Reporte", textX, 36, { lineBreak: false });
// Right-aligned company locator (address + phone + email).
const rightLines = [
company.addressLine1,
company.addressLine2,
[company.cityState].filter(Boolean).join(" · "),
[company.phone, company.email].filter(Boolean).join(" · "),
company.taxId ? `RFC: ${company.taxId}` : "",
].filter(Boolean);
doc.font("Helvetica").fontSize(8).fillColor("#cde0db");
let ry = 12;
for (const line of rightLines) {
doc.text(line, 48, ry, {
width: pageW,
align: "right",
lineBreak: false,
ellipsis: true,
});
ry += 10;
}
// Thin accent line under the masthead.
doc.rect(0, 60, doc.page.width, 2).fill(ACCENT);
// Title + subtitle + printed-at.
doc
.font("Helvetica-Bold")
.fontSize(15)
.fillColor(BRAND)
.text(title, 48, 72, { lineBreak: false });
let metaY = 92;
if (result.subtitle) {
doc
.font("Helvetica")
.fontSize(9)
.fillColor(MUTED)
.text(result.subtitle, 48, metaY, { lineBreak: false });
metaY += 12;
}
const printedAt = new Date().toLocaleString("es-MX");
doc
.font("Helvetica")
.fontSize(8)
.fillColor(MUTED)
.text(`Impreso: ${printedAt}`, 48, metaY, { lineBreak: false });
};
drawHeader();
// Column widths: distribute page width minus margins, weighted.
const totalW = columns.reduce((s, c) => s + (c.width ?? 12), 0);
const cols = columns.map((c) => ({
label: c.label,
width: ((c.width ?? 12) / totalW) * pageW,
align: c.align,
}));
let y = 130;
const drawTableHeader = () => {
doc.rect(48, y, pageW, 18).fill("#faf6ee");
y = pdfRow(
doc,
y + 4,
cols,
cols.map((c) => ({ text: c.label, align: c.align })),
48,
);
doc
.moveTo(48, y)
.lineTo(48 + pageW, y)
.strokeColor(LINE)
.lineWidth(0.5)
.stroke();
};
drawTableHeader();
// Body rows.
for (const r of result.rows) {
if (y > doc.page.height - 64) {
doc.addPage({ layout: "landscape", margins: { top: 96, bottom: 56, left: 48, right: 48 } });
drawHeader();
y = 130;
drawTableHeader();
}
const vals = columns.map((c) => {
const v = r[c.key];
const text = c.type === "money" ? fmtMoney(v) : v == null ? "" : String(v);
return { text, align: c.align };
});
y = pdfRow(doc, y + 4, cols, vals, 48);
doc
.moveTo(48, y)
.lineTo(48 + pageW, y)
.strokeColor("#e4dccb")
.lineWidth(0.4)
.stroke();
}
// Totals.
if (result.totals) {
y += 6;
doc.rect(48, y, pageW, 18).fill(ACCENT);
doc
.font("Helvetica-Bold")
.fontSize(9)
.fillColor("#f5f1e8")
.text(
Object.entries(result.totals)
.map(([k, v]) => `${k}: ${v}`)
.join(" · "),
52,
y + 5,
{ width: pageW - 8, align: "left" },
);
}
doc.end();
});
}
/* ----------------------------------------------------------------- print (HTML) */
/**
* Print-stylesheet-friendly HTML. The web app's print stylesheet hides
* nav, but otherwise this is a plain table the browser paginates itself.
*
* The header carries the company info (logo + name + locator) so printed
* pages stand alone — staff can hand one to a customer and the office
* identification is on every sheet, not buried in the cover page.
*/
export function renderPrintHtml(
columns: ColumnDef[],
result: ReportResult,
title: string,
): string {
const company = getCompanyInfo();
const head = (label: string, align?: "left" | "right") =>
`<th style="text-align:${align ?? "left"};padding:6px 8px;border-bottom:2px solid #0c322d;background:#faf6ee;font-size:11px">${escapeHtml(label)}</th>`;
const cell = (v: unknown, c: ColumnDef) => {
const text = c.type === "money" ? fmtMoney(v) : v == null ? "" : String(v);
const align = c.align ?? "left";
return `<td style="text-align:${align};padding:4px 8px;border-bottom:1px solid #e4dccb;font-size:11px;${c.type === "money" ? "font-variant-numeric:tabular-nums" : ""}">${escapeHtml(text)}</td>`;
};
const rows = result.rows
.map(
(r) =>
`<tr>${columns
.map((c) => cell(r[c.key], c))
.join("")}</tr>`,
)
.join("");
const totals = result.totals
? `<tr><td colspan="${columns.length}" style="padding:8px;background:#bf5a34;color:#f5f1e8;font-weight:600;font-size:11px">${Object.entries(
result.totals,
)
.map(([k, v]) => `${escapeHtml(k)}: ${escapeHtml(String(v))}`)
.join(" &nbsp;·&nbsp; ")}</td></tr>`
: "";
// Logo embedded as base64 data URL — the print page is opened as a
// new tab and printed standalone, so a relative path to the web app
// wouldn't resolve when launched outside the web's origin.
const logoDataUrl = company.logo
? `data:image/png;base64,${company.logo.buffer.toString("base64")}`
: null;
const locatorLines = [
company.addressLine1,
company.addressLine2,
company.cityState,
[company.phone, company.email].filter(Boolean).join(" · "),
company.taxId ? `RFC: ${company.taxId}` : "",
].filter(Boolean);
return `<!doctype html>
<html lang="es"><head>
<meta charset="utf-8" />
<title>${escapeHtml(title)}${escapeHtml(company.name)}</title>
<style>
@page { size: letter landscape; margin: 0.5in; }
body { font-family: -apple-system, "Helvetica Neue", Helvetica, Arial, sans-serif; color: #211d17; margin: 0; }
.masthead { display: flex; align-items: flex-start; gap: 16px; padding: 12px 16px; background: #0c322d; color: #f5f1e8; border-radius: 6px 6px 0 0; }
.masthead-logo { flex: 0 0 auto; }
.masthead-logo img { display: block; height: 56px; width: auto; }
.masthead-text { flex: 1; min-width: 0; }
.masthead-name { font-family: Georgia, "Times New Roman", serif; font-size: 20px; font-weight: 600; line-height: 1.1; }
.masthead-tag { font-size: 11px; color: #cde0db; margin-top: 2px; text-transform: uppercase; letter-spacing: 0.06em; }
.masthead-locator { font-size: 10px; color: #cde0db; text-align: right; line-height: 1.35; white-space: nowrap; }
.accent { height: 3px; background: #bf5a34; }
.head { padding: 12px 4px 8px; }
.head h1 { font-family: Georgia, "Times New Roman", serif; font-size: 18px; margin: 0; color: #0c322d; }
.head p { font-size: 11px; color: #756c5c; margin: 2px 0 0; }
table { width: 100%; border-collapse: collapse; }
@media print {
.noprint { display: none; }
.masthead { border-radius: 0; }
}
.noprint { padding: 8px 0; }
.noprint button { padding: 6px 12px; background: #0c322d; color: #f5f1e8; border: 0; border-radius: 4px; cursor: pointer; font-size: 12px; }
.footer { margin-top: 16px; font-size: 9px; color: #756c5c; border-top: 1px solid #e4dccb; padding-top: 6px; display: flex; justify-content: space-between; }
</style>
</head><body>
<div class="noprint"><button onclick="window.print()">Imprimir / Guardar PDF</button></div>
<div class="masthead">
${logoDataUrl ? `<div class="masthead-logo"><img src="${logoDataUrl}" alt="" /></div>` : ""}
<div class="masthead-text">
<div class="masthead-name">${escapeHtml(company.name)}</div>
<div class="masthead-tag">Reporte</div>
</div>
<div class="masthead-locator">
${locatorLines.map((l) => escapeHtml(l)).join("<br/>")}
${company.website ? `<br/>${escapeHtml(company.website)}` : ""}
</div>
</div>
<div class="accent"></div>
<div class="head">
<h1>${escapeHtml(title)}</h1>
${result.subtitle ? `<p>${escapeHtml(result.subtitle)}</p>` : ""}
<p>Impreso: ${new Date().toLocaleString("es-MX")}</p>
</div>
<table>
<thead><tr>${columns.map((c) => head(c.label, c.align)).join("")}</tr></thead>
<tbody>${rows}${totals}</tbody>
</table>
<div class="footer">
<span>${escapeHtml(company.name)} · ${escapeHtml(company.phone)} · ${escapeHtml(company.email)}</span>
<span>${escapeHtml(title)}</span>
</div>
</body></html>`;
}
function escapeHtml(s: string): string {
return s
.replace(/&/g, "&amp;")
.replace(/</g, "&lt;")
.replace(/>/g, "&gt;")
.replace(/"/g, "&quot;");
}
+141
View File
@@ -0,0 +1,141 @@
import { Prisma } from "@jorgecuadros/database";
export function renewalLetterSelect(generation: number) {
return Prisma.validator<Prisma.PolicySelect>()({
id: true,
policyNumber: true,
policyTo: true,
netPremium: true,
policyFee: true,
total: true,
currency: true,
coveragesJson: true,
customer: {
select: {
// Needed by the notification log's customerId FK, not by the letter.
id: true,
name: true,
nameMissing: true,
email: true,
phone: true,
mobile: true,
addressLine1: true,
addressLine2: true,
city: true,
state: true,
zipCode: true,
country: true,
},
},
policyType: { select: { name: true } },
insuranceProvider: { select: { name: true } },
vehicles: {
take: 1,
select: {
make: true,
model: true,
modelYear: true,
bodyType: true,
engineNumber: true,
licensePlate: true,
},
},
renewalNotices: {
where: { generation },
select: { sentAt: true, channel: true },
},
});
}
export type RenewalLetterPolicy = Prisma.PolicyGetPayload<{
select: ReturnType<typeof renewalLetterSelect>;
}>;
export interface RenewalLetterRow extends Record<string, unknown> {
__kind: "letter";
policyId: string;
policyNumber: string;
policyType: string;
customerName: string;
customerEmail: string | null;
customerPhone: string | null;
customerMobile: string | null;
customerAddress: string[];
provider: string;
policyTo: string;
netPremium: string | null;
policyFee: string | null;
total: string | null;
currency: string;
coverageDays: unknown;
cslLimit: unknown;
medicalCoverage: unknown;
propertyDamage: unknown;
perPersonLiability: unknown;
additionalService: unknown;
vehicle: {
make: string | null;
model: string | null;
modelYear: string | null;
bodyType: string | null;
engineNumber: string | null;
licensePlate: string | null;
} | null;
generation: number;
sentAt: string | null;
}
export function toRenewalLetterRow(
policy: RenewalLetterPolicy,
generation: number,
): RenewalLetterRow {
const notice = policy.renewalNotices[0];
const coverage = (policy.coveragesJson ?? {}) as Record<string, unknown>;
const address = [
policy.customer.addressLine1,
policy.customer.addressLine2,
[policy.customer.city, policy.customer.state, policy.customer.zipCode]
.filter(Boolean)
.join(", "),
policy.customer.country,
].filter((part): part is string => Boolean(part));
return {
__kind: "letter",
policyId: policy.id,
policyNumber: policy.policyNumber,
policyType: policy.policyType?.name ?? "—",
customerName: policy.customer.nameMissing ? "(sin nombre)" : policy.customer.name,
customerEmail: policy.customer.email,
customerPhone: policy.customer.phone,
customerMobile: policy.customer.mobile,
customerAddress: address,
provider: policy.insuranceProvider?.name ?? "—",
policyTo: policy.policyTo ? policy.policyTo.toISOString().slice(0, 10) : "—",
netPremium: policy.netPremium ? policy.netPremium.toFixed(2) : null,
policyFee: policy.policyFee ? policy.policyFee.toFixed(2) : null,
total: policy.total ? policy.total.toFixed(2) : null,
currency: policy.currency,
coverageDays: coverage.cobertura ?? null,
cslLimit: coverage.csl_limite ?? null,
medicalCoverage: coverage.gastos_medico ?? null,
propertyDamage: coverage.propiedades ?? null,
perPersonLiability: coverage.personas ?? null,
additionalService:
coverage.servicio_adicional ?? coverage.servicio_adiconal ?? null,
vehicle: policy.vehicles[0]
? {
make: policy.vehicles[0].make,
model: policy.vehicles[0].model,
modelYear: policy.vehicles[0].modelYear,
bodyType: policy.vehicles[0].bodyType,
engineNumber: policy.vehicles[0].engineNumber,
licensePlate: policy.vehicles[0].licensePlate,
}
: null,
generation,
sentAt: notice?.sentAt
? notice.sentAt.toISOString().slice(0, 10)
: null,
};
}
+130
View File
@@ -0,0 +1,130 @@
import {
Body,
Controller,
Get,
Header,
Param,
Post,
Query,
Res,
UseGuards,
} from "@nestjs/common";
import type { Response } from "express";
import { AuthenticatedGuard } from "../auth/authenticated.guard";
import { ReportsService } from "./reports.service";
import {
renderCsv,
renderPdf,
renderPrintHtml,
renderXlsx,
} from "./outputs";
import { findReport } from "./reports.registry";
/**
* Reports routes. Every report is dispatched by slug; outputs are
* differentiated by `?format=...` (default `json`). Reads only — gated
* by AuthenticatedGuard alone, like every other read in the app.
*/
@UseGuards(AuthenticatedGuard)
@Controller("reports")
export class ReportsController {
constructor(private readonly reports: ReportsService) {}
/** Catalog of all registered reports (the /reportes index). */
@Get()
catalog() {
return { items: this.reports.catalog() };
}
/** Run a report and return the JSON result (rows + totals + the def's columns). */
@Get(":slug")
async runJson(
@Param("slug") slug: string,
@Query() query: Record<string, string | undefined>,
) {
const def = findReport(slug);
const result = await this.reports.run(slug, query);
return { ...result, columns: def?.columns ?? [] };
}
/** CSV download. */
@Get(":slug/csv")
@Header("Content-Type", "text/csv; charset=utf-8")
async runCsv(
@Param("slug") slug: string,
@Query() query: Record<string, string | undefined>,
@Res() res: Response,
) {
const def = findReport(slug);
const result = await this.reports.run(slug, query);
const filename = `${def?.title ?? slug}-${new Date().toISOString().slice(0, 10)}.csv`;
res.setHeader(
"Content-Disposition",
`attachment; filename="${filename.replace(/[^\wÀ-ſ .-]/g, "_")}"`,
);
res.send(renderCsv(def?.columns ?? [], result));
}
/** XLSX download. */
@Get(":slug/xlsx")
async runXlsx(
@Param("slug") slug: string,
@Query() query: Record<string, string | undefined>,
@Res() res: Response,
) {
const def = findReport(slug);
const result = await this.reports.run(slug, query);
const filename = `${def?.title ?? slug}-${new Date().toISOString().slice(0, 10)}.xlsx`;
const buf = await renderXlsx(def?.columns ?? [], result);
res.setHeader(
"Content-Type",
"application/vnd.openxmlformats-officedocument.spreadsheetml.sheet",
);
res.setHeader(
"Content-Disposition",
`attachment; filename="${filename.replace(/[^\wÀ-ſ .-]/g, "_")}"`,
);
res.send(buf);
}
/** PDF download. */
@Get(":slug/pdf")
async runPdf(
@Param("slug") slug: string,
@Query() query: Record<string, string | undefined>,
@Res() res: Response,
) {
const def = findReport(slug);
const result = await this.reports.run(slug, query);
const buf = await renderPdf(def?.columns ?? [], result, def?.title ?? slug);
const filename = `${def?.title ?? slug}-${new Date().toISOString().slice(0, 10)}.pdf`;
res.setHeader("Content-Type", "application/pdf");
res.setHeader(
"Content-Disposition",
`attachment; filename="${filename.replace(/[^\wÀ-ſ .-]/g, "_")}"`,
);
res.send(buf);
}
/** Browser-printable HTML view (the user hits Print → Save as PDF). */
@Get(":slug/print")
@Header("Content-Type", "text/html; charset=utf-8")
async runPrint(
@Param("slug") slug: string,
@Query() query: Record<string, string | undefined>,
) {
const def = findReport(slug);
const result = await this.reports.run(slug, query);
return renderPrintHtml(def?.columns ?? [], result, def?.title ?? slug);
}
/** POST a customer-picker-driven report (statement). Mirrors GET to keep
* the param contract simple: same body shape, same response. */
@Post(":slug")
async runPost(
@Param("slug") slug: string,
@Body() body: Record<string, string | undefined>,
) {
return this.reports.run(slug, body);
}
}
+9
View File
@@ -0,0 +1,9 @@
import { Module } from "@nestjs/common";
import { ReportsController } from "./reports.controller";
import { ReportsService } from "./reports.service";
@Module({
controllers: [ReportsController],
providers: [ReportsService],
})
export class ReportsModule {}
File diff suppressed because it is too large Load Diff
+45
View File
@@ -0,0 +1,45 @@
import { Injectable, NotFoundException } from "@nestjs/common";
import { PrismaService } from "../prisma/prisma.service";
import { findReport, REPORTS } from "./reports.registry";
import type { ReportDef } from "./reports.types";
/**
* The reports service. Two responsibilities:
* 1. Run a report by slug with the given params — just dispatch.
* 2. Return the catalog for the /reportes index page.
*
* Output rendering (CSV/XLSX/PDF/HTML print) lives in `outputs.ts`; this
* service is data only. The controller maps URLs to (slug, format) and
* hands the result to outputs.
*/
@Injectable()
export class ReportsService {
constructor(private readonly prisma: PrismaService) {}
/** List every registered report, in display order. */
catalog(): Array<{
slug: string;
title: string;
description: string;
domain: string;
legacyName: string | null;
format: string;
params: ReportDef["params"];
}> {
return REPORTS.map((r) => ({
slug: r.slug,
title: r.title,
description: r.description,
domain: r.domain,
legacyName: r.legacyName,
format: r.format,
params: r.params,
}));
}
async run(slug: string, params: Record<string, string | undefined>) {
const def = findReport(slug);
if (!def) throw new NotFoundException(`Reporte "${slug}" no encontrado`);
return def.run(this.prisma, params);
}
}
+142
View File
@@ -0,0 +1,142 @@
/**
* The reports module — plan step 10.
*
* Each "report" is one entry in `reports.registry.ts`. The entry declares
* its slug (URL id), title, what filters it accepts, what columns it
* returns, and a `run` function that produces the data from Prisma. The
* service dispatches on slug; the controller exposes JSON + CSV + XLSX +
* PDF + HTML print; the catalog endpoint exposes the registry itself so
* the `/reportes` page can render the same data.
*
* Output philosophy: a report returns a uniform shape — `columns` (typed
* schema) + `rows` (any[] of values matching the column types) + `totals`
* (record of column key → summary value). All three output formats
* (CSV/XLSX/PDF/print) derive from this same shape so adding a new
* report is one entry, never a per-format template.
*/
import { Prisma } from "@jorgecuadros/database";
import { PrismaService } from "../prisma/prisma.service";
/** Top-level grouping for the catalog page; matches the existing nav. */
export type ReportDomain =
| "clientes"
| "polizas"
| "servicios"
| "estado-cuenta"
| "chequera";
/** How the runner should render rows: a grid, a per-customer statement, or
* one printable letter per row (e.g. renewal notices — see `format:
* "letter"` reports for the `__kind: "letter"` row shape they emit). */
export type ReportFormat = "tabular" | "statement" | "letter";
/** Filter controls the report's UI should render. */
export type ParamDef =
| {
key: string;
label: string;
kind: "text" | "number";
placeholder?: string;
defaultValue?: string;
}
| {
key: string;
label: string;
kind: "date";
/** Inclusive bound, true for `to`, false for `from`. */
endOfDay?: boolean;
defaultValue?: string;
}
| {
key: string;
label: string;
kind: "select";
options: { value: string; label: string }[];
defaultValue?: string;
}
| {
key: string;
label: string;
kind: "customer-picker";
};
/** One column of the output table. */
export interface ColumnDef {
key: string;
label: string;
/** Render hint for the on-screen + print table. */
type: "text" | "number" | "money" | "date";
/** Right-align numbers/money; default false (left). */
align?: "left" | "right";
/** Used for column-width hints in the print/PDF layout. */
width?: number;
}
/** Shape every report's `run` resolves to. Columns come from the def. */
export interface ReportResult {
rows: Array<Record<string, unknown>>;
totals?: Record<string, string | number>;
/** Optional free-form subtitle for print/PDF (e.g. date range, scope). */
subtitle?: string;
}
/** A report's static declaration. */
export interface ReportDef {
slug: string;
title: string;
description: string;
domain: ReportDomain;
/** The original Access report name (per docs/LEGACY_DATABASES_OBJECTS.md)
* for traceability. Null when this is a new report with no legacy equiv. */
legacyName: string | null;
format: ReportFormat;
params: ParamDef[];
columns: ColumnDef[];
/**
* Run the report. Receives the Prisma client and the validated params
* record (keys are the `key` from ParamDef, values are the strings the
* runner collected; numeric/date params arrive as strings — the report
* parses them). Must apply the same NOT_VOIDED filter on transactions as
* the billing module so totals match.
*/
run: (
prisma: PrismaService,
params: Record<string, string | undefined>,
) => Promise<ReportResult>;
}
/** A typed bag of helpers for the report functions. */
export interface ReportCtx {
prisma: PrismaService;
params: Record<string, string | undefined>;
}
/** Helper: a `YYYY-MM-DD` bound; unparseable is undefined. */
export function parseDate(
v: string | undefined,
endOfDay = false,
): Date | undefined {
if (!v) return undefined;
const d = new Date(endOfDay ? `${v}T23:59:59.999Z` : `${v}T00:00:00.000Z`);
return Number.isNaN(d.getTime()) ? undefined : d;
}
/** Helper: integer param with default. */
export function intParam(
p: Record<string, string | undefined>,
key: string,
def: number,
min = 1,
max = 1000,
): number {
const n = Number(p[key]);
if (!Number.isFinite(n)) return def;
return Math.min(max, Math.max(min, Math.round(n)));
}
/** Helper: not-voided filter, shared with billing.service. */
export const NOT_VOIDED: Prisma.TransactionWhereInput = { voidedAt: null };
export const NOT_VOIDED_BANK: Prisma.BankTransactionWhereInput = {
voidedAt: null,
};
+14
View File
@@ -0,0 +1,14 @@
import { Module } from "@nestjs/common";
import { SettingsService } from "./settings.service";
/**
* Operator-editable configuration. No controller of its own — each setting is
* exposed by the feature that owns it (summary recipients live under
* /notifications), so the validation and the permission live next to the
* thing they protect rather than behind a generic key/value endpoint.
*/
@Module({
providers: [SettingsService],
exports: [SettingsService],
})
export class SettingsModule {}
@@ -0,0 +1,89 @@
import { SettingsService, invalidEmails, parseEmailList } from "./settings.service";
/**
* The db → env → default ladder is the whole contract of this service: it is
* what lets the setting move out of the environment without changing how any
* existing deployment behaves.
*/
function build(row: { value: string } | null, env?: string) {
const prisma = {
appSetting: {
findUnique: jest.fn().mockResolvedValue(
row ? { key: "k", updatedAt: new Date("2026-08-02"), updatedById: "u1", ...row } : null,
),
upsert: jest.fn().mockResolvedValue({}),
},
};
const config = { get: jest.fn().mockReturnValue(env) };
return {
service: new SettingsService(prisma as never, config as never),
prisma,
};
}
describe("notification admin emails resolve db > env > default", () => {
it("prefers the stored row", async () => {
const { service } = build({ value: "a@x.com,b@x.com" }, "env@x.com");
await expect(service.notificationAdminEmails()).resolves.toMatchObject({
value: ["a@x.com", "b@x.com"],
source: "db",
updatedById: "u1",
});
});
it("falls back to the environment when nothing is stored", async () => {
const { service } = build(null, "env@x.com, other@x.com");
await expect(service.notificationAdminEmails()).resolves.toMatchObject({
value: ["env@x.com", "other@x.com"],
source: "env",
});
});
it("falls back to the built-in defaults when neither is set", async () => {
const { service } = build(null, undefined);
const resolved = await service.notificationAdminEmails();
expect(resolved.source).toBe("default");
expect(resolved.value).toHaveLength(2);
});
it("treats a stored empty list as 'nobody', not as unset", async () => {
// The regression this guards: falling through to env/defaults here would
// keep mailing people who were deliberately removed.
const { service } = build({ value: "" }, "env@x.com");
await expect(service.notificationAdminEmails()).resolves.toMatchObject({
value: [],
source: "db",
});
});
it("writes the list back as CSV", async () => {
const { service, prisma } = build({ value: "" });
await service.setNotificationAdminEmails(["a@x.com", "b@x.com"], "user-9");
expect(prisma.appSetting.upsert).toHaveBeenCalledWith(
expect.objectContaining({
create: expect.objectContaining({ value: "a@x.com,b@x.com", updatedById: "user-9" }),
update: expect.objectContaining({ value: "a@x.com,b@x.com", updatedById: "user-9" }),
}),
);
});
});
describe("email list parsing", () => {
it("trims and drops blanks", () => {
expect(parseEmailList(" a@x.com , ,b@x.com ")).toEqual(["a@x.com", "b@x.com"]);
});
it("rejects entries that are not addresses at all", () => {
expect(invalidEmails(["ok@x.com", "nope", "also@bad"])).toEqual([
"nope",
"also@bad",
]);
});
});
+218
View File
@@ -0,0 +1,218 @@
import { Injectable, Logger } from "@nestjs/common";
import { ConfigService } from "@nestjs/config";
import { PrismaService } from "../prisma/prisma.service";
/**
* Reader/writer for `app_settings` — the configuration staff can change
* without a redeploy.
*
* Every setting resolves through the same three-step ladder: the database row
* if an operator has set one, else the environment variable it used to live
* in, else a hardcoded default. That ordering is what makes this migration
* safe — an existing deployment keeps behaving exactly as it did until
* somebody edits the value in the UI, and `source` tells the UI which of the
* three it is looking at so "this came from the env, editing it here will
* take over" is visible rather than surprising.
*/
export const SETTING_KEYS = {
/** Comma-separated recipients of the per-job notification summary. */
notificationAdminEmails: "notification.adminEmails",
/** JSON cadence of the automatic servicios sweep. */
scheduleServicios: "notification.schedule.servicios",
/** JSON cadence of the automatic pólizas renewal sweep. */
schedulePolizas: "notification.schedule.polizas",
/** Whether the NUMid allocator may reuse empty portal ids. */
numidRecycleEmpty: "numid.recycleEmpty",
} as const;
/** Where a resolved value came from. Shown in the UI. */
export type SettingSource = "db" | "env" | "default";
export interface ResolvedSetting<T> {
value: T;
source: SettingSource;
updatedAt: Date | null;
updatedById: string | null;
}
/** Last resort when neither the database nor the environment says otherwise.
* Matches what `NotificationsService` hardcoded before this table existed. */
const DEFAULT_ADMIN_EMAILS = ["rmancinas@freakma.net", "mpulido@freakma.net"];
/** Deliberately permissive — this rejects "not an address at all", not
* "not deliverable". Only SES can tell us the latter, and a validator strict
* enough to argue with is a validator that blocks a legitimate address. */
const EMAIL_RE = /^[^\s@,]+@[^\s@,]+\.[^\s@,]+$/;
export function parseEmailList(raw: string): string[] {
return raw
.split(",")
.map((s) => s.trim())
.filter(Boolean);
}
export function invalidEmails(list: string[]): string[] {
return list.filter((e) => !EMAIL_RE.test(e));
}
@Injectable()
export class SettingsService {
private readonly logger = new Logger(SettingsService.name);
constructor(
private readonly prisma: PrismaService,
private readonly config: ConfigService,
) {}
/**
* Recipients of the per-job summary email.
*
* Read on every send rather than cached at boot: the point of moving this
* out of the environment was that it changes while the app is running, and
* a cache would reintroduce exactly the restart-to-apply behaviour we are
* removing. It is one indexed primary-key lookup per sweep, not per email.
*/
async notificationAdminEmails(): Promise<ResolvedSetting<string[]>> {
const row = await this.read(SETTING_KEYS.notificationAdminEmails);
if (row) {
const parsed = parseEmailList(row.value);
// An empty stored value is a legitimate choice — "send no summaries" —
// and must not silently fall through to the env or the defaults, or an
// operator who cleared the field would keep receiving mail.
return {
value: parsed,
source: "db",
updatedAt: row.updatedAt,
updatedById: row.updatedById,
};
}
const env = this.config.get<string>("NOTIFICATION_ADMIN_EMAILS");
if (env && env.trim()) {
return {
value: parseEmailList(env),
source: "env",
updatedAt: null,
updatedById: null,
};
}
return {
value: [...DEFAULT_ADMIN_EMAILS],
source: "default",
updatedAt: null,
updatedById: null,
};
}
/** Persist the summary recipients. An empty list is stored as an empty
* string and means "nobody" — see the read path above. */
async setNotificationAdminEmails(
emails: string[],
userId: string,
): Promise<ResolvedSetting<string[]>> {
await this.write(
SETTING_KEYS.notificationAdminEmails,
emails.join(","),
userId,
);
return this.notificationAdminEmails();
}
/**
* Cadence of one automatic envío, stored as JSON.
*
* No env rung on this ladder: a schedule was never an environment variable
* (it was a `@Cron` literal in the source), so the only two sources are the
* operator's row and the caller's default — which is the previous hardcoded
* behaviour. A row that fails to parse is treated as absent and logged
* rather than thrown: a bad JSON blob must not take the scheduler down with
* it, and falling back to the shipped cadence is the safe reading.
*/
async notificationSchedule<T>(
kind: "servicios" | "polizas",
fallback: T,
): Promise<ResolvedSetting<T>> {
const key =
kind === "servicios"
? SETTING_KEYS.scheduleServicios
: SETTING_KEYS.schedulePolizas;
const row = await this.read(key);
if (row) {
try {
return {
value: { ...fallback, ...(JSON.parse(row.value) as T) },
source: "db",
updatedAt: row.updatedAt,
updatedById: row.updatedById,
};
} catch (error) {
this.logger.warn(
`Setting ${key} is not valid JSON, using the default: ` +
`${(error as Error).message}`,
);
}
}
return { value: fallback, source: "default", updatedAt: null, updatedById: null };
}
async setNotificationSchedule(
kind: "servicios" | "polizas",
schedule: unknown,
userId: string,
): Promise<void> {
await this.write(
kind === "servicios"
? SETTING_KEYS.scheduleServicios
: SETTING_KEYS.schedulePolizas,
JSON.stringify(schedule),
userId,
);
}
/**
* Whether the NUMid allocator may reuse empty portal ids instead of only
* issuing new ones.
*
* Defaults to OFF, and the default is the safety property rather than a
* preference: while Access remains the utilities master, every reusable id
* still exists in DATGRAL, and a `--sync` migration run reassigns the ref back
* to its Access owner (transform_customers.py:327). Recycling before utilities
* cuts over therefore hands out ids that quietly stop working. No env rung —
* this has never been an environment variable and should be flipped
* deliberately, in the UI, by someone who knows the cutover happened.
*/
async numidRecycleEmpty(): Promise<ResolvedSetting<boolean>> {
const row = await this.read(SETTING_KEYS.numidRecycleEmpty);
if (row) {
return {
value: row.value === "true",
source: "db",
updatedAt: row.updatedAt,
updatedById: row.updatedById,
};
}
return { value: false, source: "default", updatedAt: null, updatedById: null };
}
async setNumidRecycleEmpty(
enabled: boolean,
userId: string,
): Promise<ResolvedSetting<boolean>> {
await this.write(SETTING_KEYS.numidRecycleEmpty, String(enabled), userId);
return this.numidRecycleEmpty();
}
private read(key: string) {
return this.prisma.appSetting.findUnique({ where: { key } });
}
private async write(key: string, value: string, userId: string) {
await this.prisma.appSetting.upsert({
where: { key },
create: { key, value, updatedById: userId },
update: { value, updatedById: userId },
});
}
}
@@ -0,0 +1,71 @@
/**
* The OCR seam. Everything above this interface works in terms of page text and
* word boxes, so the concrete engine is swappable without touching the parsers,
* the matcher, or the schema.
*
* The shipped implementation is self-hosted Tesseract (see tesseract.provider).
* That choice is evidence-based rather than assumed: run against 46 pages of
* real scanned CFE, CESPT and Telnor statements, it identified the provider on
* 46/46 and extracted a usable account reference on 43/46, and on a later
* corpus of 19 scanned municipal predial receipts it read the provider on
* 19/19 and an identifier on 18/19 — well past the bar for a queue whose whole
* point is that a human confirms every row. A
* managed document-extraction API (Textract, Document Intelligence, Document
* AI) fits behind this same interface if per-page accuracy ever proves
* insufficient, with no schema change — but at 300+ pages/month/company it
* would carry a real recurring cost for accuracy that is not currently the
* bottleneck.
*/
/** One OCR'd word, with where it sits on the page. */
export interface OcrWord {
text: string;
/** Pixel box in the rendered page image. */
left: number;
top: number;
width: number;
height: number;
/** Engine confidence for this word, 0..1. */
confidence: number;
}
export interface OcrPage {
/** Full page text, reading order, newline-separated. */
text: string;
/**
* Word boxes. Needed because several of the real layouts are *tables* — the
* CESPT "RECIBO" prints `No. DE CUENTA` as a column header with the value in
* the row beneath it, which line-oriented text cannot associate. Parsers fall
* back to geometry for exactly those fields.
*/
words: OcrWord[];
/** Mean word confidence across the page, 0..1. */
confidence: number;
}
export interface OcrProvider {
/** True when the engine is actually usable in this deployment. */
available(): Promise<boolean>;
/** Split a PDF into one rendered page image per page. */
renderPages(pdf: Buffer): Promise<Buffer[]>;
/** OCR a single rendered page image. */
recognize(pageImage: Buffer): Promise<OcrPage>;
/**
* Read a PDF's own text layer, one entry per page, `null` where the page has
* none worth using.
*
* Not every statement is a scan. The gas company e-mails born-digital CFDI
* invoices whose text is already exact and already positioned — running those
* through a rasteriser and a character recogniser can only lose information
* (one sample turned `MEDIDOR: VM01014426` into `ar (LTR): 014420`) while
* costing about a minute of CPU per page for the privilege. Where the layer
* exists it is strictly better input for the same parsers, so it is tried
* first and OCR remains the fallback for genuine scans.
*
* Positions are reported in the same pixel space `recognize` uses, so the
* geometric helpers in the parsers work unchanged on either source.
*/
textPages(pdf: Buffer): Promise<(OcrPage | null)[]>;
}
export const OCR_PROVIDER = Symbol("OCR_PROVIDER");
@@ -0,0 +1,106 @@
import { parseBboxLayout } from "./tesseract.provider";
/**
* Shaped like real `pdftotext -bbox-layout` output: the gas invoice lays its
* header out as two columns of independent text flows, so poppler puts a label
* and the value printed beside it in *different* `<line>` elements. Trusting
* that grouping is what left `PERIODO FACTURADO` with no value next to it and
* every period field empty on a batch whose text was perfectly readable.
*/
/**
* Boxes are sized from the text, at 6 units a character: the reassembler now
* reads the space BETWEEN two boxes, so a fixed width would put a fabricated
* gap after every short word and every row would come back column-padded.
*/
function word(x: number, y: number, text: string): string {
return `<word xMin="${x}" yMin="${y}" xMax="${x + text.length * 6}" yMax="${y + 8}">${text}</word>`;
}
function doc(...lines: string[]): string {
return `<doc><page width="612" height="792">${lines
.map((l) => `<flow><block><line>${l}</line></block></flow>`)
.join("")}</page></doc>`;
}
/** Enough words on the page to clear the "is this a real text layer" floor. */
function padding(): string {
return Array.from({ length: 50 }, (_, i) => word(10, 400 + i * 10, `w${i}`)).join("");
}
describe("parseBboxLayout", () => {
it("rejoins a label with the value printed beside it in another flow", () => {
const [page] = parseBboxLayout(
doc(
word(20, 100, "PERIODO") + word(68, 100, "FACTURADO:"),
word(300, 100.4, "20260630-20260630"),
padding(),
),
1,
);
expect(page).not.toBeNull();
expect(page!.text).toMatch(/PERIODO FACTURADO:\s+20260630-20260630/);
});
it("keeps genuinely separate lines apart", () => {
const [page] = parseBboxLayout(
doc(word(20, 100, "Cuenta:") + word(68, 100, "0900003463"), word(20, 130, "Nombre:"), padding()),
1,
);
const lines = page!.text.split("\n").map((l) => l.trim());
expect(lines).toContain("Cuenta: 0900003463");
expect(lines).toContain("Nombre:");
});
/**
* The layout is data. A borderless table separates its cells with nothing
* but white space, so the parsers read a run of spaces as a cell boundary
* (`INSURED\s{2,}`) and a column offset as a column (`SUM INSURED` vs
* `PREMIUM`). Both regressed to nothing when this collapsed every gap to a
* single space, and the fixtures — taken from `pdftotext -layout`, which
* prints the gaps — could not see it.
*/
it("preserves the gap between two cells of a borderless table", () => {
const [page] = parseBboxLayout(
doc(word(20, 100, "INSURED") + word(300, 100, "PAMELA") + word(340, 100, "WAGONER"), padding()),
1,
);
const line = page!.text.split("\n").find((l) => l.includes("INSURED"))!;
expect(line).toMatch(/INSURED\s{2,}PAMELA WAGONER/);
});
it("preserves the blank line between two blocks", () => {
const [page] = parseBboxLayout(
doc(word(20, 100, "Insured"), word(20, 112, "wraps"), word(20, 200, "Next"), padding()),
1,
);
const lines = page!.text.split("\n").map((l) => l.trim());
// The wrapped continuation stays attached; the next block is cut off from
// it, which is what stops a "join until the cell ends" walk running away.
expect(lines.slice(lines.indexOf("Insured"), lines.indexOf("Next") + 1)).toEqual([
"Insured",
"wraps",
"",
"Next",
]);
});
it("scales point coordinates into the render's pixel space", () => {
// Word boxes have to land in the same coordinate space tesseract reports,
// or the geometric helpers the parsers share silently stop finding values.
const [page] = parseBboxLayout(doc(word(72, 144, "X") + padding()), 300 / 72);
const x = page!.words.find((w) => w.text === "X")!;
expect(x.left).toBeCloseTo(300);
expect(x.top).toBeCloseTo(600);
});
it("reports no text layer for a scan carrying a few stray glyphs", () => {
expect(parseBboxLayout(doc(word(10, 10, "3") + word(40, 10, "of") + word(60, 10, "5")), 1)).toEqual([
null,
]);
});
it("decodes the entities poppler escapes", () => {
const [page] = parseBboxLayout(doc(word(10, 10, "A&amp;B") + padding()), 1);
expect(page!.text).toContain("A&B");
});
});
@@ -0,0 +1,438 @@
import { Injectable, Logger, ServiceUnavailableException } from "@nestjs/common";
import { ConfigService } from "@nestjs/config";
import { execFile } from "node:child_process";
import { mkdtemp, readFile, readdir, rm, writeFile } from "node:fs/promises";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { promisify } from "node:util";
import type { OcrPage, OcrProvider, OcrWord } from "./ocr.provider";
const run = promisify(execFile);
/**
* Self-hosted OCR: `pdftoppm` (poppler) to rasterise, `tesseract` to read.
*
* Both are external binaries rather than a native npm addon, which keeps the
* pnpm workspace free of a compiled dependency and makes the alpine runtime
* image a two-package change (see docker/api.Dockerfile). Like StorageService,
* a missing binary degrades rather than crashes the API: the module reports
* itself unavailable and statement ingest returns 503, while every other
* feature keeps working.
*
* The settings below are not arbitrary — they were measured against the real
* scanned samples:
* - 300 DPI grayscale. The source scans are phone photos of paper at ~5MB a
* page; below 300 the small print (RMU, clave catastral) stops resolving,
* above it costs time for no additional fields.
* - `--psm 6` ("assume a single uniform block of text"). The default page
* segmentation splits these dense forms into columns and interleaves them,
* which destroys the label-then-value adjacency every parser depends on.
* - Spanish traineddata, with a graceful fall back to English if the language
* pack is absent — an accented label reads worse but the digits, which are
* what actually gets matched, are unaffected.
*/
@Injectable()
export class TesseractOcrProvider implements OcrProvider {
private readonly logger = new Logger(TesseractOcrProvider.name);
private readonly dpi: number;
private readonly lang: string;
private probe: Promise<boolean> | null = null;
constructor(config: ConfigService) {
this.dpi = Number(config.get("OCR_DPI") ?? 300);
this.lang = config.get<string>("OCR_LANG") ?? "spa";
}
/** Cached — the binaries do not appear or vanish while the process runs. */
available(): Promise<boolean> {
if (!this.probe) {
this.probe = (async () => {
try {
await Promise.all([
run("tesseract", ["--version"]),
run("pdftoppm", ["-v"]),
]);
return true;
} catch {
this.logger.warn(
"OCR unavailable: `tesseract` and/or `pdftoppm` not found on PATH. " +
"Statement ingest is disabled; every other feature is unaffected.",
);
return false;
}
})();
}
return this.probe;
}
private async require(): Promise<void> {
if (!(await this.available())) {
throw new ServiceUnavailableException(
"El servicio de OCR no está disponible en este servidor.",
);
}
}
private async scratch<T>(fn: (dir: string) => Promise<T>): Promise<T> {
const dir = await mkdtemp(join(tmpdir(), "stmt-ocr-"));
try {
return await fn(dir);
} finally {
await rm(dir, { recursive: true, force: true });
}
}
async renderPages(pdf: Buffer): Promise<Buffer[]> {
await this.require();
return this.scratch(async (dir) => {
const src = join(dir, "in.pdf");
await writeFile(src, pdf);
// -gray: these are grayscale scans already; colour triples the bytes
// handed to tesseract for no gain in character recognition.
await run("pdftoppm", [
"-r",
String(this.dpi),
"-gray",
"-png",
src,
join(dir, "page"),
]);
const files = (await readdir(dir))
.filter((f) => f.startsWith("page") && f.endsWith(".png"))
// pdftoppm zero-pads its page numbers, so lexical order is page order.
.sort();
return Promise.all(files.map((f) => readFile(join(dir, f))));
});
}
/**
* `pdftotext -bbox-layout` — the same poppler package `pdftoppm` comes from,
* so this costs no extra dependency in the runtime image.
*
* A page is only accepted when it carries a real text layer. Scanned PDFs
* frequently contain a handful of stray glyphs (a scanner watermark, a page
* number stamped by the MFP), and treating those as the page's text would
* hand every parser an almost-empty string and silently take OCR out of the
* loop — so a floor of MIN_TEXT_WORDS words has to be present before the
* layer is believed.
*/
async textPages(pdf: Buffer): Promise<(OcrPage | null)[]> {
await this.require();
return this.scratch(async (dir) => {
const src = join(dir, "in.pdf");
await writeFile(src, pdf);
const out = join(dir, "out.html");
try {
await run("pdftotext", ["-bbox-layout", src, out]);
} catch (err) {
this.logger.warn(
`pdftotext failed; falling back to OCR for this file: ${(err as Error).message}`,
);
return [];
}
// Points to pixels at the render DPI, so word boxes from either source
// land in one coordinate space and `valueUnder`'s thresholds hold.
return parseBboxLayout(await readFile(out, "utf8"), this.dpi / 72);
});
}
async recognize(pageImage: Buffer): Promise<OcrPage> {
await this.require();
return this.scratch(async (dir) => {
const img = join(dir, "page.png");
await writeFile(img, pageImage);
// One tesseract invocation produces both outputs; TSV carries the word
// boxes and per-word confidence, and its text can be reassembled into
// reading order, so there is no need to run the engine twice.
const out = join(dir, "out");
try {
await run("tesseract", [img, out, "-l", this.lang, "--psm", "6", "tsv"]);
} catch (err) {
if (this.lang !== "eng") {
this.logger.warn(
`Tesseract failed with lang "${this.lang}", retrying with "eng": ${
(err as Error).message
}`,
);
await run("tesseract", [img, out, "-l", "eng", "--psm", "6", "tsv"]);
} else {
throw err;
}
}
const tsv = await readFile(`${out}.tsv`, "utf8");
return parseTsv(tsv);
});
}
}
/**
* Below this many words a "text layer" is scanner debris, not a document.
* The real born-digital samples carry 400+ words a page; the scanned ones
* carry none at all, so the exact threshold is not delicate.
*/
const MIN_TEXT_WORDS = 40;
const ENTITIES: Record<string, string> = {
amp: "&",
lt: "<",
gt: ">",
quot: '"',
apos: "'",
};
function decodeEntities(s: string): string {
return s.replace(/&(#x?[0-9a-fA-F]+|[a-z]+);/g, (whole, body: string) => {
if (body[0] === "#") {
const code =
body[1] === "x" || body[1] === "X"
? parseInt(body.slice(2), 16)
: parseInt(body.slice(1), 10);
return Number.isFinite(code) ? String.fromCodePoint(code) : whole;
}
return ENTITIES[body] ?? whole;
});
}
/**
* Turn `pdftotext -bbox-layout`'s XHTML into one OcrPage per PDF page.
*
* Parsed with regexes rather than an XML library on purpose: the output is
* machine-generated by poppler with a fixed element shape (`page` > `flow` >
* `block` > `line` > `word`), and the alternative is a parser dependency in
* the API for one file format read in one place. Only `page` and `word` are
* consulted — see below for why poppler's own `line` grouping is discarded.
*
* `confidence` is 1 for every word: these are the document's own characters,
* not a recognition guess.
*/
export function parseBboxLayout(xhtml: string, scale: number): (OcrPage | null)[] {
const pages: (OcrPage | null)[] = [];
for (const pageMatch of xhtml.matchAll(/<page\b[^>]*>([\s\S]*?)<\/page>/g)) {
const words: OcrWord[] = [];
for (const w of pageMatch[1].matchAll(
/<word\s+xMin="([\d.eE+-]+)"\s+yMin="([\d.eE+-]+)"\s+xMax="([\d.eE+-]+)"\s+yMax="([\d.eE+-]+)"\s*>([\s\S]*?)<\/word>/g,
)) {
const text = decodeEntities(w[5]).trim();
if (!text) continue;
const left = Number(w[1]) * scale;
const top = Number(w[2]) * scale;
words.push({
text,
left,
top,
width: Number(w[3]) * scale - left,
height: Number(w[4]) * scale - top,
confidence: 1,
});
}
pages.push(
words.length >= MIN_TEXT_WORDS
? { text: toVisualRows(words), words, confidence: 1 }
: null,
);
}
return pages;
}
/**
* Reassemble words into the rows a reader sees, left to right.
*
* Poppler's own `<line>` grouping cannot be used for this. It groups by text
* flow, and these invoices lay their fields out as two columns of independent
* flows — so `PERIODO FACTURADO:` and the `20260630-20260630` printed beside
* it end up in different `<line>` elements, and every label-then-value pattern
* in the parsers misses a value that is plainly there on the page. Regrouping
* by vertical position restores the adjacency, and matches what tesseract
* hands back for the scanned version of the same layout.
*
* Rows are cut when a word's vertical centre leaves the band established by
* the row's first word, which tolerates the sub-pixel baseline differences
* between fonts on one line without merging two genuinely separate lines.
*
* Vertical WHITE SPACE is preserved as a blank line. Rows alone are not the
* whole layout: on a form, the blank between two blocks is what says where a
* cell's wrapped value stops, and dropping it leaves parsers that walk a
* block ("keep joining until the cell ends") running to the end of the page.
* That is not hypothetical — the GMX PVL especificación read its whole first
* page as the insured's name, because the fixtures were taken from
* `pdftotext -layout` (which prints the blanks) while the runtime fed it this
* function's output (which did not).
*
* Horizontal white space is preserved the same way, by padding each word out
* to its own column. The same fixture mismatch bit here: a run of spaces is
* the ONLY thing separating two cells of a borderless table, so ANA's
* `INSURED\s{2,}` label matches and its `SUM INSURED` / `PREMIUM` column
* split (taken from `head.search()` offsets) both need real offsets. Joining
* on one space put every driver's-policy premium in the sum-insured column
* and left the phone glued to the insured's name.
*/
function toVisualRows(words: OcrWord[]): string {
const centre = (w: OcrWord) => w.top + w.height / 2;
const sorted = [...words].sort((a, b) => centre(a) - centre(b) || a.left - b.left);
const rows: OcrWord[][] = [];
let current: OcrWord[] = [];
let band = 0;
for (const w of sorted) {
if (!current.length) {
current = [w];
band = centre(w);
continue;
}
// Half the word's own height: tall headings and body text both sit within
// their own line's band, and neither reaches into the next one.
if (Math.abs(centre(w) - band) <= Math.max(w.height, current[0].height) / 2) {
current.push(w);
} else {
rows.push(current);
current = [w];
band = centre(w);
}
}
if (current.length) rows.push(current);
const charWidth = estimateCharWidth(words);
const out: string[] = [];
rows.forEach((r, i) => {
if (i > 0 && isBlankBetween(rows[i - 1], r)) out.push("");
out.push(layoutRow(r, charWidth));
});
return out.join("\n");
}
/**
* One row rendered at its printed column offsets.
*
* Words that merely follow one another inside the same cell are separated by
* exactly one space, whatever the column arithmetic says: one `charWidth` for
* a page that mixes fonts leaves a rounding error on every word, and letting
* that accumulate sprinkles `\s{2,}` runs through ordinary prose — which is
* the very thing the parsers read as a cell boundary. Only a gap wide enough
* to be deliberate (more than one blank character) is rendered as one, and
* only there is the word re-anchored to its true column, so the offsets a
* column split depends on stay honest while values stay clean.
*/
function layoutRow(row: OcrWord[], charWidth: number): string {
let line = "";
let right = 0;
for (const w of [...row].sort((a, b) => a.left - b.left)) {
const col = Math.round(w.left / charWidth);
if (!line.length) {
line = " ".repeat(Math.max(0, col));
} else if (w.left - right > charWidth * 1.5) {
line += " ".repeat(Math.max(2, col - line.length));
} else {
line += " ";
}
line += w.text;
right = w.left + w.width;
}
return line.trimEnd();
}
/**
* Width of one character, in the same units the word boxes use.
*
* The median of each word's own width-per-character: robust to the handful of
* oversized headings and to the wide-tracked letterhead, both of which would
* drag a mean. Only words of 3+ characters vote, since a one-character box is
* mostly side bearing. Falls back to a value derived from line height when a
* page has nothing long enough to measure.
*/
function estimateCharWidth(words: OcrWord[]): number {
const samples = words
.filter((w) => w.text.length >= 3 && w.width > 0)
.map((w) => w.width / w.text.length)
.sort((a, b) => a - b);
if (samples.length) return samples[Math.floor(samples.length / 2)];
const heights = words.map((w) => w.height).filter((h) => h > 0);
return heights.length ? Math.max(...heights) / 2 : 1;
}
/**
* Does the space between two consecutive rows read as an empty line?
*
* Measured against the taller of the two rows so a heading and its body text
* are judged on their own scale. On the real documents the two populations do
* not overlap: consecutive lines of one paragraph sit at 0.31.1 line heights
* apart, and anything the reader sees as blank-separated starts at 2.1. The
* threshold is placed in that empty middle, biased high — a missed blank only
* restores today's behaviour, while a spurious one would cut a wrapped value
* short.
*/
function isBlankBetween(prev: OcrWord[], row: OcrWord[]): boolean {
const bottom = Math.max(...prev.map((w) => w.top + w.height));
const top = Math.min(...row.map((w) => w.top));
const unit = Math.max(
...prev.map((w) => w.height),
...row.map((w) => w.height),
);
return unit > 0 && top - bottom > unit * 1.6;
}
/**
* Turn tesseract's TSV into words plus reassembled text.
*
* Columns are: level, page_num, block_num, par_num, line_num, word_num, left,
* top, width, height, conf, text. Rows with level < 5 are structural (page,
* block, paragraph, line) and carry no text; only level 5 is a word. A conf of
* -1 marks a structural row, so those are dropped rather than averaged in —
* including them would drag every page's confidence toward zero.
*/
export function parseTsv(tsv: string): OcrPage {
const lines = tsv.split("\n");
const header = lines[0]?.split("\t") ?? [];
const col = (name: string) => header.indexOf(name);
const iLeft = col("left");
const iTop = col("top");
const iWidth = col("width");
const iHeight = col("height");
const iConf = col("conf");
const iText = col("text");
const iLine = col("line_num");
const iBlock = col("block_num");
const words: OcrWord[] = [];
// Keyed by block+line so the reassembled text preserves the engine's own
// reading order instead of sorting words by raw y, which interleaves columns.
const byLine = new Map<string, string[]>();
for (let i = 1; i < lines.length; i++) {
const f = lines[i].split("\t");
if (f.length <= iText) continue;
const text = f[iText]?.trim();
if (!text) continue;
const confidence = Number(f[iConf]);
if (!Number.isFinite(confidence) || confidence < 0) continue;
words.push({
text,
left: Number(f[iLeft]) || 0,
top: Number(f[iTop]) || 0,
width: Number(f[iWidth]) || 0,
height: Number(f[iHeight]) || 0,
confidence: confidence / 100,
});
const key = `${f[iBlock]}:${f[iLine]}`;
const bucket = byLine.get(key);
if (bucket) bucket.push(text);
else byLine.set(key, [text]);
}
const text = [...byLine.values()].map((w) => w.join(" ")).join("\n");
const confidence = words.length
? words.reduce((sum, w) => sum + w.confidence, 0) / words.length
: 0;
return { text, words, confidence };
}
@@ -0,0 +1,270 @@
import type { OcrPage } from "../ocr/ocr.provider";
import {
detectProvider,
normalizeCadastralKey,
normalizeZofematKey,
parseStatement,
} from "./statement-parser";
/**
* Every string in this file is a verbatim excerpt of what the OCR engine
* actually returned for a real receipt — misreads, dropped spaces, mangled
* accents and all. That is the point: these are the specific ways these five
* layouts have been observed to fail, and the assertions pin down what the
* parser is supposed to do about each one. Inventing clean input here would
* test nothing, because clean input was never the problem.
*/
function page(text: string): OcrPage {
return { text, words: [], confidence: 0.9 };
}
describe("detectProvider", () => {
it("reads a Rosarito predial receipt as predial, not as a water bill", () => {
// "Clave Catastral" is also a CESPT structural marker, so a predial page
// whose header OCR'd badly must still not be claimed by the CESPT rule.
expect(
detectProvider(
"e | Clave Catastral. KP-128-105 IMPUESTO PREDIAL ea rita\n" +
"TASA | VALOR FISCAL | BIMESTRES | INCISO. | IMPUESTO",
),
).toBe("PREDIAL ROSARITO");
});
it("keeps telling the three municipalities apart by their RFC", () => {
expect(detectProvider("R.F.C. ATB-541201-KK2")).toBe("PREDIAL TIJUANA");
expect(detectProvider("R.F.C. AMP-981201-HJ4")).toBe("PREDIAL ROSARITO");
expect(detectProvider("MEN-540301-9J5")).toBe("PREDIAL ENSENADA");
});
it("does not let the CFE rule claim a gas bill over 'PERIODO FACTURADO'", () => {
expect(
detectProvider("Orden de Facturación: 000009801640\nPERIODO FACTURADO: 20260630-20260630"),
).toBe("GAS TIJUANA");
});
});
describe("normalizeCadastralKey", () => {
it("keeps a letter in the third position instead of digitising it", () => {
// `MMB01041` is a real key on file; mapping its B to 8 produced a key that
// matches no property at all.
expect(normalizeCadastralKey("MM-B01-041", [])).toBe("MMB01041");
});
it("repairs the spurious I tesseract inserts into the prefix", () => {
expect(normalizeCadastralKey("MIM-200-010", [])).toBe("MM200010");
});
it("digitises confusable glyphs from position four onward", () => {
expect(normalizeCadastralKey("KP-1O8-O45", [])).toBe("KP108045");
});
it("flags a prefix it had to truncate", () => {
const notes: string[] = [];
expect(normalizeCadastralKey("KPX-128-106", notes)).toBe("KP128106");
expect(notes).toHaveLength(1);
});
});
describe("parsePredialTijuana", () => {
const TIJUANA = page(
"Hats | AYUNTAMIENTO DE TIJUANA, BC $2,613.00 23/01/2026\n" +
"y) TELEFONO: 973-7000 R.F.C. ATB-541201-KK2\n" +
"ER AÑO VALOR FISCAL TASA IMPUESTO |CONCEPTO IMPORTE\n" +
"ED ca 2026 1,207,15778 246 2,969.61 1102 - IMPUESTO PREDIAL 2,969.61\n" +
"55164964310126000002613000054192\n" +
"se 0 O (54427 [a] | TOTALAPAGAR: 2,613.00\n" +
"Dc 1097 : FECHA VENCE : 31/ENE/2026",
);
it("splits the payment barcode into account, deadline and amount", () => {
const p = parseStatement(TIJUANA);
expect(p.provider).toBe("PREDIAL TIJUANA");
expect(p.serviceKind).toBe("PROPERTY_TAX");
expect(p.accountRef).toBe("55164964");
expect(p.amount).toBe(2613);
expect(p.dueDate?.toISOString().slice(0, 10)).toBe("2026-01-31");
expect(p.period).toBe("2026");
});
it("reads the printed total even when the space in the label is lost", () => {
// The real page OCR'd the label as "TOTALAPAGAR:", and it is that reading
// that cross-checks the barcode's amount.
expect(parseStatement(TIJUANA).crossChecked).toBe(true);
});
it("refuses to trust a barcode the printed total contradicts", () => {
const p = parseStatement(
page(
"R.F.C. ATB-541201-KK2\n" +
"55164964310126000002613000054192\n" +
"TOTAL A PAGAR: 9,613.00\nFECHA VENCE : 31/ENE/2026",
),
);
expect(p.crossChecked).toBe(false);
expect(p.notes.join(" ")).toContain("no coincide");
});
});
describe("parsePredialRosarito", () => {
it("takes the rounded Total, not the Sub Total printed above it", () => {
const p = parseStatement(
page(
"AYUNTAMIENTO MUNICIPAL DE PLAYAS DE ROSARITO, B.C.\n" +
"Ce Clave Catastral: + JR-400-008 7 | IMPUESTO PREDIAL\n" +
"SUPERFICIE: 228.31 ZONA 30025 “Redondeo IT049 -$0.39 Sub Total $5,409.39\n" +
"¿XTEMPORANEO DESPUES DE: 31/01/2026 Elaboro: MGLG\n" +
"Total | $5,409.00\n" +
"| Periodo por Pagar: 2026/1 2026/6",
),
);
expect(p.cadastralKey).toBe("JR400008");
expect(p.amount).toBe(5409);
expect(p.dueDate?.toISOString().slice(0, 10)).toBe("2026-01-31");
expect(p.period).toBe("2026");
});
it("is not fooled by the unspaced 'SubTotal' spelling", () => {
// This exact page read $9,624.85 off a receipt for $9,625.00 while the
// lookbehind still assumed a space.
const p = parseStatement(
page(
"AMP-981201-HJ4 IMPUESTO PREDIAL\n" +
"SUPERFICIE. 367.62 ZONA:30151 | Redondco 17049 $0.15 SubTotal $9,624.85\n" +
": Total | $9,625.00",
),
);
expect(p.amount).toBe(9625);
});
});
describe("parsePredialEnsenada", () => {
const totals = (tail: string) =>
page(
"IMPRESION MAQUINA REGISTRADORA ez | MUNICIPIO DE ENSENADA\n" +
"+7] DATOS. DEL.CAUSANTE alta A pe CLAVE MM-200-010 2 CUENTA\n" +
`ES g € S| TOTALES 12,744.47 0.00 0.00 324.56 0.00 13,069.03 ${tail} |`,
);
it("reads the paid total off the TOTALES row however the label OCR'd", () => {
expect(parseStatement(totals("TOTA LA A $5,797.00")).amount).toBe(5797);
expect(parseStatement(totals("orAL: M7 z] $14,414.00")).amount).toBe(14414);
expect(parseStatement(totals("| TOTAL: = $6 246.00")).amount).toBe(6246);
});
it("reports no amount rather than one whose $ was misread as an 8", () => {
// `TOTAL: A 82,203.00` is a $2,203.00 receipt. Posting $82,203 would look
// entirely ordinary in the ledger, so this page must go to review instead.
const p = parseStatement(totals("TOTAL: A 82,203.00"));
expect(p.amount).toBeNull();
expect(p.notes.join(" ")).toContain("capturarlo a mano");
});
it("never falls back to the assessed total on the same row", () => {
expect(parseStatement(totals("yo: se TE= 58/4690]")).amount).toBeNull();
});
});
describe("parseGas", () => {
const gas = (...cuentas: string[]) =>
page(
"GTI4608032K2 COMPAÑIA DE GAS DE TIJUANA\n" +
"Fecha de Vencimiento: 2026/08/08\n" +
cuentas.map((c) => `Cuenta: ${c}`).join("\n") +
"\nPERIODO FACTURADO: 20260630-20260630\nTOTAL A PAGAR: $275.82",
);
it("strips the printed leading zero to the stored account number", () => {
const p = parseStatement(gas("0900003463", "0900003463", "0900003463"));
expect(p.serviceKind).toBe("GAS");
expect(p.accountRef).toBe("900003463");
expect(p.amount).toBe(275.82);
expect(p.dueDate?.toISOString().slice(0, 10)).toBe("2026-08-08");
expect(p.period).toBe("2026-06");
expect(p.crossChecked).toBe(true);
});
it("takes the majority reading but still sends a disagreement to review", () => {
const p = parseStatement(gas("0900003463", "0900003463", "0900003468"));
expect(p.accountRef).toBe("900003463");
expect(p.crossChecked).toBe(false);
});
it("claims no cross-check from a single printing", () => {
expect(parseStatement(gas("0900003463")).crossChecked).toBeNull();
});
});
describe("parseZonaFederal", () => {
/**
* The Tijuana zona federal receipt, trimmed to the rows the parser reads.
* Verbatim from page 7 of the August 2026 batch, including the two ways the
* heading OCR'd: the clave line is struck through by the office's own
* highlighter, which is what cost two of eight pages their concession clave.
*/
const zf = (clave: string, body = "") =>
page(
"ESIZ <pYl Av. Independencia y Esq. Paseo del CentenaxiaiiArlhnto de Tijuana, B.C.\n" +
"Teléfono: 9737000 R.F.C. ATB-541201-BK2 0070000146 12:54 PM\n" +
"Zona Federal Marítimo Terrestre\n" +
`${clave} Nombre: DENNIS JOHN SEIN Concesión:\n` +
"Periodo Construcción Tasa Ornato Tasa Impuesto Actualiza. Recargo Multa Importe\n" +
"2026-2 / 2026-2 316.40 35.00 0.00 12.11 1,845.66 0.00 27.13 1,000.00 2,872.79\n" +
"SubTotal 1,845.66 0.00 27.13 1,000.00 2,872.79\n" +
"Concepto: Derechos de ocupación de Zona Federal Marítimo Terrestre\n" +
body,
);
it("is not claimed by the predial parser that shares its RFC and header", () => {
// Tijuana bills predial and zona federal from the same treasury, so
// "Ayuntamiento de Tijuana" and ATB-541201 identify neither on their own.
expect(detectProvider("R.F.C. ATB-541201-BK2\nZona Federal Marítimo Terrestre")).toBe(
"ZONA FEDERAL TIJUANA",
);
expect(parseStatement(zf("Clave: 14-D -014")).serviceKind).toBe("FEDERAL_ZONE");
});
it("still recognises the layout when the heading itself did not survive OCR", () => {
// Real: page 1 came back as "Zona Ledera) Maritimo Terrestre".
expect(
detectProvider("Zona Ledera) Maritimo Terrestre\nClave EJ -012% Nombre: STEFAN"),
).toBe("ZONA FEDERAL TIJUANA");
});
it("reads the clave through the loose spacing the receipt prints", () => {
expect(parseStatement(zf("Clave: 14-D -014")).accountRef).toBe("14D014");
expect(parseStatement(zf("Clave: 14-A-119")).accountRef).toBe("14A119");
});
it("keeps the letter instead of digitising it", () => {
// toDigits maps D to 0 and B to 8; a real 14-D -014 must not become 140014.
expect(normalizeZofematKey("14-D -014")).toBe("14D014");
expect(normalizeZofematKey("12-B -013")).toBe("12B013");
});
it("takes the payable amount from the SubTotal row, rounded to whole pesos", () => {
// The municipality rounds and prints the difference as "Ajuste Ley Hacienda
// Mpal"; 2,872.79 is charged as $2,873.00.
expect(parseStatement(zf("Clave: 14-D -014")).amount).toBe(2873);
});
it("prefers the printed total and cross-checks it against the subtotal", () => {
const p = parseStatement(zf("Clave: 14-D -014", "Total a pagar $2,873.00"));
expect(p.amount).toBe(2873);
expect(p.crossChecked).toBe(true);
});
it("sends a printed total that contradicts the subtotal to review", () => {
const p = parseStatement(zf("Clave: 14-D -014", "Total a pagar $2,973.00"));
expect(p.crossChecked).toBe(false);
});
it("translates the printed bimester into the ledger's own vocabulary", () => {
expect(parseStatement(zf("Clave: 14-D -014")).period).toBe("MAR/APR");
});
it("leaves the clave blank rather than guessing when the marker ate it", () => {
const p = parseStatement(zf("Clave EJ -012%"));
expect(p.accountRef).toBeNull();
expect(p.notes.join(" ")).toContain("clave");
});
});
@@ -0,0 +1,872 @@
import type { ServiceKind } from "@jorgecuadros/database";
import type { OcrPage, OcrWord } from "../ocr/ocr.provider";
/**
* What one parsed statement page yields. `accountRef` is already normalised to
* the form the migrated `PropertyService` columns hold, so the matcher compares
* like with like and never has to know about provider-specific formatting.
*/
export interface ParsedStatement {
/**
* "CFE" | "CESPT" | "TELNOR" | "GAS TIJUANA" | "PREDIAL TIJUANA" |
* "PREDIAL ROSARITO" | "PREDIAL ENSENADA" | "ZONA FEDERAL TIJUANA", or null
* when no parser claimed the page.
*/
provider: string | null;
serviceKind: ServiceKind | null;
accountRef: string | null;
/** Clave catastral, when printed — a second key to match on. */
cadastralKey: string | null;
amount: number | null;
dueDate: Date | null;
period: string | null;
/**
* Independent corroboration of `accountRef`. CFE and Telnor both print a
* payment barcode that repeats the account number (and the amount), so when
* the barcode and the label agree the extraction is near-certainly right;
* when they disagree, or only one is present, the page is worth a human
* glance. Null when the layout has no second source.
*/
crossChecked: boolean | null;
/** Human-readable trail of what was read, surfaced in the review queue. */
notes: string[];
}
// --- shared helpers ---------------------------------------------------------
/**
* Tesseract confuses these glyphs inside numeric runs with some regularity —
* a real clave catastral `KB078025` came back as `KBO78025`. Applied ONLY to
* fields known to be digits, never to free text, where it would corrupt words.
*/
const DIGIT_CONFUSIONS: Record<string, string> = {
O: "0",
o: "0",
D: "0",
I: "1",
l: "1",
"|": "1",
S: "5",
B: "8",
};
export function toDigits(s: string | null | undefined): string {
if (!s) return "";
return s
.split("")
.map((c) => DIGIT_CONFUSIONS[c] ?? c)
.join("")
.replace(/\D/g, "");
}
/**
* Parse a printed amount, treating `,` and `.` by position rather than by
* assumption. A real Telnor bill OCR'd as "$ 649,00" — blindly stripping commas
* as thousands separators turned $649.00 into $64,900, a hundredfold error that
* would post silently. Two trailing digits after a single separator are always
* cents here; a separator followed by three digits is a thousands group.
*/
function money(s: string | null | undefined): number | null {
if (!s) return null;
const cleaned = s.replace(/[\s$]/g, "");
// 1.234,56 or 1,234.56 — grouped thousands plus optional cents.
let m = cleaned.match(/^(\d{1,3}(?:[.,]\d{3})+)([.,]\d{1,2})?$/);
if (m) {
const whole = m[1].replace(/[.,]/g, "");
const cents = m[2] ? m[2].slice(1) : "";
return Number(cents ? `${whole}.${cents.padEnd(2, "0")}` : whole);
}
// 649,00 / 649.00 — a single separator with exactly two digits after it.
m = cleaned.match(/^(\d+)[.,](\d{2})$/);
if (m) return Number(`${m[1]}.${m[2]}`);
const n = Number(cleaned.replace(/[,.]/g, ""));
return Number.isFinite(n) ? n : null;
}
function firstMatch(text: string, patterns: RegExp[]): string | null {
for (const p of patterns) {
const m = text.match(p);
if (m?.[1]) return m[1].trim();
}
return null;
}
/** Every capture of `pattern` across the page, in order. */
function allMatches(text: string, pattern: RegExp): string[] {
const out: string[] = [];
const re = new RegExp(pattern.source, pattern.flags.includes("g") ? pattern.flags : `${pattern.flags}g`);
for (const m of text.matchAll(re)) {
if (m[1]) out.push(m[1].trim());
}
return out;
}
const MONTHS: Record<string, number> = {
ENE: 0, FEB: 1, MAR: 2, ABR: 3, MAY: 4, JUN: 5,
JUL: 6, AGO: 7, SEP: 8, OCT: 9, NOV: 10, DIC: 11,
};
/** Parses the three date shapes these statements actually print. */
export function parseDate(raw: string | null | undefined): Date | null {
if (!raw) return null;
const s = raw.trim().toUpperCase();
// 16/07/2026
let m = s.match(/^(\d{1,2})\/(\d{1,2})\/(\d{4})$/);
if (m) return utc(+m[3], +m[2] - 1, +m[1]);
// 22-JUL-2026 / 22 JUN 26 / 31/ENE/2026 (Tijuana predial)
m = s.match(/^(\d{1,2})[-\s/]([A-Z]{3})[A-Z]*[-\s/](\d{2,4})$/);
if (m && MONTHS[m[2]] !== undefined) {
const y = m[3].length === 2 ? 2000 + +m[3] : +m[3];
return utc(y, MONTHS[m[2]], +m[1]);
}
// 2026-07-22 (already normalised, e.g. decoded from a barcode) and the
// 2026/08/08 the gas bill prints — same field order, different separator.
m = s.match(/^(\d{4})[-/](\d{2})[-/](\d{2})$/);
if (m) return utc(+m[1], +m[2] - 1, +m[3]);
return null;
}
function utc(y: number, mo: number, d: number): Date | null {
const dt = new Date(Date.UTC(y, mo, d));
return Number.isNaN(dt.getTime()) ? null : dt;
}
/**
* Read the value printed *underneath* a column header.
*
* The CESPT "RECIBO" is a table: `No. DE CUENTA` is a header cell and its value
* sits in the row below it, so no amount of label-adjacent regex on line text
* can associate the two. This walks the word boxes instead — find the header
* word, then take the nearest word below it whose horizontal centre falls
* within the column.
*/
export function valueUnder(
page: OcrPage,
header: RegExp,
opts: { maxDy?: number; tolerance?: number; match?: RegExp } = {},
): string | null {
const { maxDy = 300, tolerance = 200, match } = opts;
const centre = (w: OcrWord) => ({
x: w.left + w.width / 2,
y: w.top + w.height / 2,
});
for (const h of page.words.filter((w) => header.test(w.text))) {
const hc = centre(h);
const below = page.words
.filter((w) => {
const c = centre(w);
return c.y > hc.y && c.y <= hc.y + maxDy && Math.abs(c.x - hc.x) <= tolerance;
})
.sort((a, b) => centre(a).y - centre(b).y);
for (const w of below) {
if (!match || match.test(w.text)) return w.text;
}
}
return null;
}
// --- provider detection -----------------------------------------------------
/**
* Brand wordmarks first, page structure only as a fallback — and the two passes
* must not be interleaved. Scanned logos OCR badly (one CESPT header came back
* as "E BAJA ES PAGO / EALIFORNIA", with neither "CESPT" nor "COMISIÓN ESTATAL"
* readable), so the structural pass is what rescues those pages. But a Telnor
* bill contains the words "Pagar antes de", which a CFE structural rule
* evaluated first will happily claim — running all brand checks before any
* structural check is what keeps that from happening.
*/
const BRAND: [string, RegExp][] = [
["CFE", /comisi[oó]n federal de electricidad|CFE.?contigo|Suministrador de Servicios/i],
["CESPT", /CESPT|COMISI[OÓ]N ESTATAL DE SERVICIOS/i],
["TELNOR", /TELNOR|TELEFONOS DEL NOROESTE/i],
["GAS TIJUANA", /COMPA[ÑN][IÍ]?A\s*DE\s*GAS\s*DE\s*TIJUANA|bajagas/i],
// Ahead of the predial rules on purpose. Tijuana's zona federal receipt is
// issued by the same treasury and carries the same header — "Ayuntamiento de
// Tijuana", the same address, the same `ATB-541201` RFC — so every predial
// discriminator matches it too, and whichever rule is asked first wins the
// page. What only the zona federal layout says is "Marítimo Terrestre", which
// survived OCR on all eight sample pages even where the heading above it came
// back as "Zona Ledera) Maritimo Terrestre" and the printed concession clave
// was lost under a highlighter mark.
["ZONA FEDERAL TIJUANA", /ZOFEMAT|Mar[ií]timo\s*Terrestre|ocupaci[oó]n\s*de\s*Zona\s*Federal/i],
// The municipal RFCs are the single most reliable discriminator on a predial
// receipt: they are printed in a clean monospaced run on every layout, they
// never change, and they say which of the three city treasuries issued the
// page — which the wordmarks alone do not, since a Tijuana receipt also
// carries "PLAYAS DE TIJUANA" and a Rosarito one "TIJUANA ENSENADA".
["PREDIAL TIJUANA", /AYUNTAMIENTO\s*DE\s*TIJUANA|ATB.?541201/i],
["PREDIAL ROSARITO", /AYUNTAMIENTO\s*MUNICIPAL\s*DE\s*PLAYAS\s*DE\s*ROSARITO|AMP.?981201|rosarito\.gob/i],
["PREDIAL ENSENADA", /MUNICIPIO\s*DE\s*ENSENADA|MEN.?540301/i],
];
/**
* The predial rules come first because a Rosarito receipt prints "Clave
* Catastral" as a boxed label — the very string the CESPT structural rule
* looks for — so a page whose municipal header failed to OCR would otherwise
* be claimed as a water bill and matched against the wrong column entirely.
* "IMPUESTO PREDIAL" appears on all three municipal layouts and on none of the
* utility ones, so it is the safe first question to ask.
*/
const LAYOUT: [string, RegExp][] = [
// Same reasoning as the brand pass, one rule earlier: the concept line
// "Derechos de ocupación de Zona Federal Marítimo Terrestre" is printed on
// the stub of every zona federal page and on no other layout, and it read
// cleanly on 8 of 8 samples — including the two whose heading did not.
["ZONA FEDERAL TIJUANA", /Derechos\s*de\s*ocupaci[oó]n/i],
["PREDIAL TIJUANA", /IMPUESTO\s*PREDIAL[\s\S]*?(?:CERTIFICACION\s*DE\s*CAJA|PASEO\s*DEL\s*CENTENARIO|PAGA\s*TU\s*PREDIAL)/i],
["PREDIAL ENSENADA", /(?:IMPUESTO\s*PREDIAL[\s\S]*?TRANSPENINSULAR)|(?:IMPRESION\s*MAQUINA\s*REGISTRADORA)/i],
["PREDIAL ROSARITO", /IMPUESTO\s*PREDIAL/i],
["GAS TIJUANA", /Orden\s*de\s*Facturaci[oó]n|FACTOR\s*DE\s*PRESI[OÓ]N|GAS\s*LP/i],
["CFE", /NO\.?\s*DE\s*SERVICIO|L[IÍ]MITE\s*DE\s*PAGO|PERIODO\s*FACTURADO/i],
["CESPT", /SALDO\s+CORRIENTE|CLAVE\s*CATASTRAL|No\.?\s*DE\s*CUENTA/i],
["TELNOR", /Mes\s*de\s*Facturaci[oó]n|Pagar\s*antes\s*de/i],
];
export function detectProvider(text: string): string | null {
for (const group of [BRAND, LAYOUT]) {
for (const [name, pattern] of group) {
if (pattern.test(text)) return name;
}
}
return null;
}
// --- CFE (electric) ---------------------------------------------------------
function parseCfe(page: OcrPage): ParsedStatement {
const text = page.text;
const notes: string[] = [];
// The payment barcode line repeats the service number, the due date (YYMMDD)
// and the amount in one fixed-width run, and reads far more reliably than the
// label: on one sample the label came back as "0059603001917" (a digit too
// many) while its barcode gave the correct "005960300191". So the barcode
// wins, and the label becomes the cross-check rather than the source.
const barcode = text.match(/\b01\s+([0-9OIlSBD]{12})\s+([0-9OIlSBD]{6})\s+([0-9OIlSBD]{9})\b/);
const label = firstMatch(text, [/NO\.?\s*DE\s*SERVICIO\s*[:;.]?\s*([0-9OIlSBD]{10,14})/i]);
let accountRef: string | null = null;
let amount: number | null = null;
let dueDate: Date | null = null;
let crossChecked: boolean | null = null;
if (barcode) {
// Leading zeros are print padding: DATMEX.rpu holds the bare 10 digits.
accountRef = toDigits(barcode[1]).replace(/^0+/, "");
amount = Number(toDigits(barcode[3]));
const d = toDigits(barcode[2]);
dueDate = parseDate(`20${d.slice(0, 2)}-${d.slice(2, 4)}-${d.slice(4, 6)}`);
notes.push("importe y vencimiento leídos del código de barras");
if (label) {
crossChecked = toDigits(label).replace(/^0+/, "") === accountRef;
if (!crossChecked) {
notes.push(
`el número impreso (${toDigits(label).replace(/^0+/, "")}) no coincide con el código de barras`,
);
}
}
} else if (label) {
accountRef = toDigits(label).replace(/^0+/, "");
notes.push("sin código de barras legible; número tomado de la etiqueta");
}
if (amount == null) {
amount = money(firstMatch(text, [/TOTAL\s*A\s*PAGAR\s*[:;.]?\s*\$?\s*([\d,]+\.?\d*)/i]));
}
if (!dueDate) {
dueDate = parseDate(
firstMatch(text, [/L[IÍ]MITE\s*DE\s*PAGO\s*[:;.]?\s*(\d{1,2}\s+\w{3}\s+\d{2,4})/i]),
);
}
return {
provider: "CFE",
serviceKind: "ELECTRIC",
accountRef: accountRef || null,
cadastralKey: null,
amount,
dueDate,
period: firstMatch(text, [
/PERIODO\s*FACTURADO\s*[:;.]?\s*(\d{1,2}\s+\w{3}\s+\d{2}\s*-\s*\d{1,2}\s+\w{3}\s+\d{2})/i,
]),
crossChecked,
notes,
};
}
// --- CESPT (water) ----------------------------------------------------------
/**
* Two different layouts arrive under the same brand:
* - the line-oriented "COMPROBANTE DE PAGO" (`Cuenta : 7604192`), and
* - the tabular "RECIBO", where `No. DE CUENTA` is a column header.
* Line patterns are tried first; anything they miss falls through to the
* geometric read, which is what the tabular layout needs.
*/
function parseCespt(page: OcrPage): ParsedStatement {
const text = page.text;
const notes: string[] = [];
let account = firstMatch(text, [/Cuenta\s*[:;.]?\s*([0-9OIlSBD]{5,9})/i]);
if (!account) {
account = valueUnder(page, /^CUENTA$/i, { match: /^[0-9OIlSBD]{5,9}$/ });
if (account) notes.push("número de cuenta leído de la columna del recibo");
}
let clave = firstMatch(text, [/Cve\.?\s*Cat\.?\s*[:;.]?\s*([A-Z]{2}\s?[0-9OIlSBD]{6})/i]);
if (!clave) {
clave = valueUnder(page, /^CATASTRAL$/i, { match: /^[A-Z]{2}[0-9OIlSBD]{6}$/i });
if (clave) notes.push("clave catastral leída de la columna del recibo");
}
let due = firstMatch(text, [/Fecha\s*Venc\s*[:;.]?\s*(\d{2}\/\d{2}\/\d{4})/i]);
if (!due) due = valueUnder(page, /^VENCIMIENTO$/i, { match: /^\d{2}\/\d{2}\/\d{4}$/ });
const amount = money(
firstMatch(text, [
/TOTAL\s*[:;.]?\s*\$?\s*([\d,]+\.\d{2})/i,
/SALDO\s+CORRIENTE[^\n]*?([\d,]+\.\d{2})/i,
]),
);
// Leading zeros are print padding here too: the RECIBO prints `0457341` for
// what DATMEX.agua holds as `457341`.
const accountRef = account ? toDigits(account).replace(/^0+/, "") : null;
const cadastralKey = clave
? clave.replace(/\s/g, "").slice(0, 2).toUpperCase() +
toDigits(clave.replace(/\s/g, "").slice(2))
: null;
return {
provider: "CESPT",
serviceKind: "WATER",
accountRef: accountRef || null,
cadastralKey: cadastralKey || null,
amount,
dueDate: parseDate(due),
period: null,
crossChecked: null,
notes,
};
}
// --- TELNOR (telephone) -----------------------------------------------------
function parseTelnor(page: OcrPage): ParsedStatement {
const text = page.text;
const notes: string[] = [];
const label = firstMatch(text, [
/Tel[eé]fono\s*[:;.]?\s*([0-9OIlSBD]{3}\s?[0-9OIlSBD]{3}\s?[0-9OIlSBD]{4})/i,
]);
// The payment stub prints phone (10 digits) + amount in cents (9) + a check
// digit: `6646093444 000099900 7` for a $999.00 bill. Reading the amount as
// 10 digits swallows the check digit and inflates the figure 100-fold.
const barcode = text.match(/\b(\d{10})(\d{9})\d\b/);
let accountRef: string | null = null;
let crossChecked: boolean | null = null;
// The bill prints the number with its 664 Tijuana LADA; DATMEX stores the
// bare local 7 digits, so the LADA is dropped rather than the stored value
// being padded — padding would guess at an area code for the 500+ existing
// rows that never recorded one.
if (label) accountRef = toDigits(label).slice(-7);
if (barcode) {
const fromBarcode = barcode[1].slice(-7);
if (accountRef) {
crossChecked = fromBarcode === accountRef;
if (!crossChecked) notes.push("el teléfono impreso no coincide con el código de barras");
} else {
accountRef = fromBarcode;
notes.push("teléfono leído del código de barras");
}
}
let amount = money(firstMatch(text, [/Total\s*a\s*Pagar\s*[:;.]?\s*\$?\s*([\d,]+\.?\d{0,2})/i]));
if (amount == null && barcode) {
amount = Number(barcode[2]) / 100;
notes.push("importe leído del código de barras");
}
return {
provider: "TELNOR",
serviceKind: "TELEPHONE",
accountRef: accountRef || null,
cadastralKey: null,
amount,
dueDate: parseDate(
firstMatch(text, [/Pagar\s*antes\s*de\s*[:;.]?\s*(\d{2}-\w{3}-\d{4})/i]),
),
period: firstMatch(text, [/Mes\s*de\s*Facturaci[oó]n\s*[:;.]?\s*(\w+)/i]),
crossChecked,
notes,
};
}
// --- GAS (Compañía de Gas de Tijuana / bajagas) ------------------------------
/**
* These arrive as born-digital CFDI PDFs rather than scans, so the text layer
* (see `TesseractOcrProvider.textPages`) usually reads them exactly and the
* patterns below only have to be tolerant enough for the scanned case.
*
* The account number is printed three times — supply address, fiscal data, and
* the payment stub at the foot — which is a free cross-check: three readings
* that agree are near-certainly right, and any disagreement means one of them
* was misread and the page deserves a human glance.
*
* `Cuenta` is what the matcher compares, not `Contrato`. The migration
* recovered gas references out of `PropertyService.notes` into `meterNumber`
* and what sat there is the 9-digit account (`900003463`), printed here with a
* leading zero as `0900003463`.
*/
function parseGas(page: OcrPage): ParsedStatement {
const text = page.text;
const notes: string[] = [];
const seen = allMatches(text, /Cuenta\s*[:;.]?\s*([0-9OIlSBD]{6,12})/i).map((s) =>
toDigits(s).replace(/^0+/, ""),
);
const distinct = [...new Set(seen.filter(Boolean))];
let accountRef: string | null = null;
let crossChecked: boolean | null = null;
if (distinct.length === 1) {
accountRef = distinct[0];
if (seen.length > 1) crossChecked = true;
} else if (distinct.length > 1) {
// Majority wins — the stub and the two address blocks print the same
// number, so a single divergent reading is the misread one. It still goes
// to review: `crossChecked: false` is what keeps the batch from
// auto-matching a number one of three readings disagreed with.
const tally = new Map<string, number>();
for (const s of seen) tally.set(s, (tally.get(s) ?? 0) + 1);
accountRef = [...tally.entries()].sort((a, b) => b[1] - a[1])[0][0];
crossChecked = false;
notes.push(`el número de cuenta se leyó de ${distinct.length} formas distintas (${distinct.join(", ")})`);
}
const amount = money(
firstMatch(text, [
/TOTAL\s*A\s*PAGAR\s*[:;.]?\s*\$\s*([\d,]+\.\d{2})/i,
/Total\s*a\s*pagar\s*[:;.]?\s*\$\s*([\d,]+\.\d{2})/i,
]),
);
// `20260630-20260630` — the range the bill was cut for. Both ends are the
// same reading date on every sample, so the period is reported as the ISO
// month rather than a range no ledger row would ever be searched by.
const facturado = firstMatch(text, [/PERIODO\s*FACTURADO\s*[:;.]?\s*(\d{8})\s*-\s*\d{8}/i]);
const period = facturado ? `${facturado.slice(0, 4)}-${facturado.slice(4, 6)}` : null;
return {
provider: "GAS TIJUANA",
serviceKind: "GAS",
accountRef: accountRef || null,
cadastralKey: null,
amount,
dueDate: parseDate(
firstMatch(text, [/Fecha\s*de\s*Vencimiento\s*[:;.]?\s*(\d{4}\s*\/\s*\d{2}\s*\/\s*\d{2})/i])?.replace(
/\s/g,
"",
),
),
period,
crossChecked,
notes,
};
}
// --- PREDIAL (municipal property tax) ---------------------------------------
/**
* Normalise a printed clave catastral to the eight-character form
* `Property.cadastralKey` holds. The municipalities print it grouped
* (`KP-128-106`, `MM-B01-041`); the stored value drops the separators
* (`KP128106`, `MMB01041`).
*
* The shape is *not* two letters and six digits, which is the assumption that
* has to be resisted here. Across the 932 distinct claves on file, characters
* four through eight are digits without exception, but the third is a digit in
* 917 of them and one of `A`, `B`, `H`, `T` in the other fifteen. Running the
* whole tail through `toDigits` — which maps `B` to `8` — is what turned a real
* `MMB01041` into a nonexistent `MM801041`, so only positions four onward get
* that treatment and a letter in the third position is kept as printed.
*
* That leaves a genuine ambiguity at that one position: a `B` there might be a
* misread `8`, and 34 stored claves do carry an `8` there against six with a
* `B`. It is left as read rather than guessed, because a page that fails to
* match lands in the review queue where a human fixes it in seconds, while a
* page that matches the wrong property posts a charge to the wrong customer.
*
* The two-letter prefix is the other fragile part. Tesseract inserts a spurious
* `I` into letter pairs with some regularity — a real `MM-200-010` came back as
* `MIM-200-010` — so a run longer than two letters has its `I`/`L` dropped
* first, which recovers exactly that case. Anything still not two letters is
* truncated and flagged, because a wrong prefix silently matches the wrong
* property or, more often, nothing at all.
*/
export function normalizeCadastralKey(
raw: string,
notes: string[],
): string | null {
const m = raw.match(/^([A-Za-z|]{2,5})[-\s]?([A-Za-z0-9|]{3})[-\s]?([0-9OIlSBD]{3})$/);
if (!m) return null;
let letters = m[1].toUpperCase().replace(/[^A-Z]/g, "");
if (letters.length > 2) {
const stripped = letters.replace(/[IL]/g, "");
if (stripped.length === 2) {
letters = stripped;
} else {
letters = letters.slice(0, 2);
notes.push(`la clave catastral se leyó como "${m[1]}"; se tomó "${letters}"`);
}
}
if (letters.length !== 2) return null;
const third = m[2][0].toUpperCase();
const tail =
(/[A-Z]/.test(third) ? third : toDigits(third)) +
toDigits(m[2].slice(1)) +
toDigits(m[3]);
return tail.length === 6 ? letters + tail : null;
}
/** The grouped clave as printed, anchored to its label when one survived OCR. */
const GROUPED_CLAVE = "[A-Z|]{2,5}-[A-Z0-9OIlSBD]{3}-[0-9OIlSBD]{3}";
function findCadastralKey(text: string, notes: string[]): string | null {
const labelled = firstMatch(text, [
new RegExp(`Clave\\s*Catastral\\s*[^A-Z0-9]{0,8}(${GROUPED_CLAVE})`, "i"),
new RegExp(`CLAVE\\s*[^A-Z0-9]{0,8}(${GROUPED_CLAVE})`, "i"),
]);
if (labelled) return normalizeCadastralKey(labelled, notes);
// Ensenada's label ("CLAVE") lands inside a table header that OCRs into
// noise more often than not, so the bare grouped shape is accepted as a
// fallback. It is distinctive enough — two letters and two three-character
// groups joined by hyphens appears nowhere else on these pages.
const bare = firstMatch(text, [new RegExp(`\\b(${GROUPED_CLAVE})\\b`)]);
return bare ? normalizeCadastralKey(bare, notes) : null;
}
/**
* Tijuana: a "CERTIFICACIÓN DE CAJA" whose payment barcode is one 32-digit run
* of `account(8) + due date(DDMMYY) + amount(9) + folio(9)`, verified against
* all five sample pages. Municipal totals are whole pesos (the receipt itself
* carries a "Redondeo" line), so the barcode amount needs no decimal point.
*
* No clave catastral is printed anywhere on this layout — the 8-digit
* municipal account is the only identifier, and it is not a number the legacy
* database ever held. Until a reviewer confirms one, every Tijuana page lands
* in review; confirming teaches the matcher (see `learnAccountRefs`) so the
* same property matches itself next year.
*/
function parsePredialTijuana(page: OcrPage): ParsedStatement {
const text = page.text;
const notes: string[] = [];
const barcode = text.match(/(?<![0-9OIlSBD])([0-9OIlSBD]{32})(?![0-9OIlSBD])/);
const printedTotal = money(
firstMatch(text, [/TOTAL\s*A?\s*PAGAR\s*[:;.]?\s*\$?\s*([\d,]+\.?\d{0,2})/i]),
);
let accountRef: string | null = null;
let amount: number | null = printedTotal;
let dueDate: Date | null = null;
let crossChecked: boolean | null = null;
if (barcode) {
const run = toDigits(barcode[1]);
const d = run.slice(8, 14);
const fromBarcode = Number(run.slice(14, 23));
accountRef = run.slice(0, 8);
dueDate = parseDate(`20${d.slice(4, 6)}-${d.slice(2, 4)}-${d.slice(0, 2)}`);
notes.push("cuenta, importe y vencimiento leídos del código de barras");
if (printedTotal != null) {
// Guarding the money, not the account number: the printed total is the
// figure a human would key, so when the two disagree one of them is a
// misread peso amount and nothing should post unreviewed.
crossChecked = Math.abs(printedTotal - fromBarcode) < 0.5;
if (!crossChecked) {
notes.push(
`el total impreso (${printedTotal}) no coincide con el código de barras (${fromBarcode})`,
);
}
}
if (amount == null) amount = fromBarcode;
}
if (!dueDate) {
dueDate = parseDate(
firstMatch(text, [/FECHA\s*VENCE\s*[:;.]?\s*(\d{1,2}\/\w{3}\/\d{4})/i]),
);
}
return {
provider: "PREDIAL TIJUANA",
serviceKind: "PROPERTY_TAX",
accountRef: accountRef || null,
cadastralKey: null,
amount,
dueDate,
// The fiscal year, which is what the legacy ledger's `period` holds for
// predial ("2026" is its single most common value). It is read from the
// assessment table's year column, and failing that from the deadline: a
// predial bill for year N falls due on 31 January of year N.
period:
firstMatch(text, [/VALOR\s*FISCAL[\s\S]{0,160}?\b(20\d{2})\b/i]) ??
(dueDate ? String(dueDate.getUTCFullYear()) : null),
crossChecked,
notes,
};
}
/**
* Rosarito: a wide "CERTIFICACIÓN DE CAJA" keyed by clave catastral, with no
* account number of its own — the clave is the identifier, which is exactly
* what `Property.cadastralKey` holds, so these match on the first pass.
*
* The total is read with a negative lookbehind on "Sub": the receipt prints
* `Sub Total $5,409.39` (before the peso rounding) directly above
* `Total $5,409.00`, and taking the first "Total" on the page books 39 cents
* that the municipality did not charge. The lookbehind allows zero spaces
* because the label prints both ways — `Sub Total` on one sample and
* `SubTotal` on the next, and the tight one is what slipped past a fixed
* `Sub\s` and read $9,624.85 off a receipt for $9,625.00.
*/
function parsePredialRosarito(page: OcrPage): ParsedStatement {
const notes: string[] = [];
const text = page.text;
return {
provider: "PREDIAL ROSARITO",
serviceKind: "PROPERTY_TAX",
accountRef: null,
cadastralKey: findCadastralKey(text, notes),
amount: money(firstMatch(text, [/(?<!Sub\s{0,3})Total\s*[|:;.]?\s*\$\s*([\d,]+\.\d{2})/i])),
// "EXTEMPORANEO DESPUES DE: 31/01/2026" — the leading E is regularly eaten
// by the box rule printed over it, so the anchor starts at "XTEMPORANEO".
dueDate: parseDate(
firstMatch(text, [/XTEMPOR[AÁ]NEO\s*DESPU[EÉ]S\s*DE\s*[:;.]?\s*(\d{2}\/\d{2}\/\d{4})/i]),
),
period: firstMatch(text, [/Periodo\s*por\s*Pagar\s*[:;.]?\s*(20\d{2})/i]),
crossChecked: null,
notes,
};
}
/**
* Ensenada: a dot-matrix "IMPRESION MAQUINA REGISTRADORA" statement, by some
* distance the worst-scanning of the three. Matching is by clave catastral.
*
* The amount is read positionally rather than by label, because the label does
* not survive: across five real pages the same word came back as `TOTAL:`,
* `TOTA LA A` and `orAL:`. What is stable is the row — the summary line that
* starts `TOTALES` carries the assessed figures across it and the amount
* actually paid last, at the right margin.
*
* That last figure must carry a literal `$`. On a real sample the paid total
* printed as `TOTAL: A $2,203.00` and OCR'd as `TOTAL: A 82,203.00` — the
* dollar sign read as an 8, a mistake that would post a $2,203 charge as
* $82,203 and look entirely ordinary in the ledger. Requiring the `$` costs
* that page its amount and sends it to review, which is the only acceptable
* failure here. The unprefixed figures earlier on the row are deliberately not
* a fallback: they are the tax assessed before the early-payment discount, not
* what was paid.
*/
function parsePredialEnsenada(page: OcrPage): ParsedStatement {
const notes: string[] = [];
const text = page.text;
const totalsRow = text.split("\n").find((l) => /TOTALES/i.test(l)) ?? "";
const figures = allMatches(totalsRow, /\$\s*(\d[\d,.\s]*\.\d{2})/);
const amount = figures.length ? money(figures[figures.length - 1]) : null;
if (amount == null) {
notes.push("no se pudo leer el importe con certeza; capturarlo a mano");
}
return {
provider: "PREDIAL ENSENADA",
serviceKind: "PROPERTY_TAX",
accountRef: null,
cadastralKey: findCadastralKey(text, notes),
amount,
// This layout prints no payment deadline at all — it is a receipt for a
// payment already made at the municipal window.
dueDate: null,
period: firstMatch(text, [/A[ÑN]O\s*[\s\S]{0,60}?\b(20\d{2})\b/i]),
crossChecked: null,
notes,
};
}
// --- ZONA FEDERAL (ZOFEMAT, Tijuana) ----------------------------------------
/**
* Normalise the concession clave the zona federal receipt is keyed by.
*
* It is printed grouped and loosely spaced — `12-T -012`, `14-A-119`,
* `14-K -031` — and is a different shape from the cadastral key entirely: two
* digits, one letter, three digits. The letter is kept as printed rather than
* digitised, for the same reason `normalizeCadastralKey` keeps its third
* character: `toDigits` maps `B` to `8` and `D` to `0`, and a real `14-D -014`
* run through it becomes `140014`, which is not a clave at all.
*
* Stored without separators, because nothing on file holds this value yet (see
* `parseZonaFederal`) so the canonical form is ours to pick, and a bare run
* cannot be broken by the hyphen the scan renders as a dash, a minus or
* nothing.
*/
export function normalizeZofematKey(raw: string): string | null {
const m = raw.match(/^([0-9OIlSBD]{2})\s*-\s*([A-Za-z])\s*-?\s*([0-9OIlSBD]{3})$/);
if (!m) return null;
const zone = toDigits(m[1]);
const lot = toDigits(m[3]);
if (zone.length !== 2 || lot.length !== 3) return null;
return `${zone}${m[2].toUpperCase()}${lot}`;
}
/**
* The bimester the receipt prints as `2026-2 / 2026-2`, rendered in the
* vocabulary the ledger already speaks.
*
* All 258 legacy FEDERAL ZONE transactions carry a period of `JAN/FEB`,
* `MAR/APR`, `MAY/JUN` or `NOV/DEC`, and their payment dates confirm the
* ordering — JAN/FEB was paid in March, MAR/APR in May, MAY/JUN in July,
* NOV/DEC in January, i.e. always the month after the bimester closes. The
* receipts agree: the two `2026-3` samples fall due 17/07/2026 with no
* surcharge, which is bimester three, May and June. Writing `2026-3` instead
* would leave the OCR-posted rows unsearchable alongside every hand-keyed one.
*/
const BIMESTERS = ["JAN/FEB", "MAR/APR", "MAY/JUN", "JUL/AUG", "SEP/OCT", "NOV/DEC"];
/**
* Tijuana's "Zona Federal Marítimo Terrestre" — the federal maritime-zone
* occupancy fee, billed by the municipality for beachfront lots.
*
* Nothing on file identifies these. `PropertyService.accountNumber` for
* FEDERAL_ZONE holds DATMEX.zfed, which is not a reference at all but an
* amount: its 77 values include `246.06`, `2369.09`, `22653.94` and a negative
* `-1679`, and the concession claves these receipts are keyed by appear nowhere
* in the database. So the clave goes to `meterNumber` (see `scopedRefField`),
* every page starts cold, and the first confirm teaches the match — the same
* arrangement Tijuana predial needed, for the same reason.
*
* The amount is taken from the SubTotal row rather than the "Total a pagar"
* box, which is printed on a grey fill and OCR'd on only 1 of 8 sample pages
* while the SubTotal row read on 8 of 8. The two differ by design: the
* municipality rounds to whole pesos and prints the difference on its own
* "Ajuste Ley Hacienda Mpal" line — `-$0.05` against a 591.05 subtotal, `$0.21`
* against 2,872.79 — so the payable figure is the rounded subtotal, and where
* the printed box did read, it agreed.
*/
function parseZonaFederal(page: OcrPage): ParsedStatement {
const text = page.text;
const notes: string[] = [];
// Printed twice, once on the receipt and once on the stub below it, which is
// a free second reading: on one sample the heading was struck through by the
// office's own highlighter and only the stub survived.
const claves = [
...new Set(
allMatches(text, /Clave\s*[:;.]?\s*([0-9OIlSBD]{2}\s*-\s*[A-Za-z]\s*-?\s*[0-9OIlSBD]{3})/i)
.map(normalizeZofematKey)
.filter((k): k is string => k != null),
),
];
const accountRef = claves[0] ?? null;
let crossChecked: boolean | null = null;
const subtotalRow = text.split("\n").find((l) => /SubTotal/i.test(l)) ?? "";
const figures = allMatches(subtotalRow, /(\d[\d,]*\.\d{2})/);
// Impuesto, Actualización, Recargo, Multa, Importe — the payable one is last.
const importe = figures.length ? money(figures[figures.length - 1]) : null;
const rounded = importe != null ? Math.round(importe) : null;
const printed = money(
firstMatch(text, [/Total\s*a\s*pagar\s*[:;.]?\s*\$?\s*([\d,]+\.\d{2})/i]),
);
if (printed != null && rounded != null) {
crossChecked = Math.abs(printed - rounded) < 0.5;
if (!crossChecked) {
notes.push(
`el total impreso (${printed}) no coincide con el subtotal redondeado (${rounded})`,
);
}
} else if (rounded != null) {
notes.push("importe tomado del subtotal, redondeado al peso");
} else if (printed == null) {
notes.push("no se pudo leer el importe con certeza; capturarlo a mano");
}
// A clave read two different ways means one of the two readings is wrong and
// there is no third to break the tie, so the page goes to a human even if the
// money cross-checked.
if (claves.length > 1) {
crossChecked = false;
notes.push(`la clave se leyó de ${claves.length} formas distintas (${claves.join(", ")})`);
}
if (!accountRef) notes.push("no se pudo leer la clave de la concesión");
const bimester = text.match(/\b(20\d{2})\s*-\s*([1-6])\s*\/\s*20\d{2}\s*-\s*[1-6]/);
return {
provider: "ZONA FEDERAL TIJUANA",
serviceKind: "FEDERAL_ZONE",
accountRef,
cadastralKey: null,
amount: printed ?? rounded,
dueDate: parseDate(
firstMatch(text, [/Vencimiento\s*[:;.]?\s*(\d{2}\/\d{2}\/\d{4})/i]),
),
period: bimester ? BIMESTERS[+bimester[2] - 1] : null,
crossChecked,
notes,
};
}
const PARSERS: Record<string, (page: OcrPage) => ParsedStatement> = {
CFE: parseCfe,
CESPT: parseCespt,
TELNOR: parseTelnor,
"GAS TIJUANA": parseGas,
"PREDIAL TIJUANA": parsePredialTijuana,
"PREDIAL ROSARITO": parsePredialRosarito,
"PREDIAL ENSENADA": parsePredialEnsenada,
"ZONA FEDERAL TIJUANA": parseZonaFederal,
};
const EMPTY: ParsedStatement = {
provider: null,
serviceKind: null,
accountRef: null,
cadastralKey: null,
amount: null,
dueDate: null,
period: null,
crossChecked: null,
notes: [],
};
/** Detect the provider and run its parser. */
export function parseStatement(page: OcrPage): ParsedStatement {
const provider = detectProvider(page.text);
if (!provider) return { ...EMPTY, notes: ["no se reconoció el proveedor"] };
return PARSERS[provider](page);
}
@@ -0,0 +1,236 @@
import { Injectable } from "@nestjs/common";
import type { ServiceKind } from "@jorgecuadros/database";
import { PrismaService } from "../prisma/prisma.service";
import type { ParsedStatement } from "./parsers/statement-parser";
export interface MatchResult {
propertyServiceId: string | null;
customerId: string | null;
/** Why it landed here — shown in the review queue verbatim. */
note: string;
/** True only for an unambiguous hit on the scoped field. */
confident: boolean;
/** Populated when more than one service claims the same number. */
candidates: { propertyServiceId: string; customerId: string; customerName: string }[];
}
/**
* Resolves a parsed statement to the customer who should be billed for it.
*
* Two rules govern everything here.
*
* **Match on one scoped field, never fuzzily across all identifiers.** Each
* service kind has exactly one column its statements print, and only that
* column is consulted. A blanket search over accountNumber/meterNumber/route
* would let a water account number collide with an unrelated phone number, and
* the resulting mis-post would look perfectly ordinary in the ledger.
*
* **Never match on the customer name.** The name on a utility bill is the
* account's registrant, which drifts from the current owner and is often years
* stale — one sample CESPT receipt is printed to "ARNAIZ ROSAS ELSA AURORA"
* for an account this office holds under "CATT, RANDY", who is not the same
* person. Names are displayed for the reviewer to sanity-check, and are never
* an input to matching.
*/
/**
* Which `PropertyService` column a given kind's statements actually print.
*
* Exported because the same answer governs three places that must agree: the
* lookup here, the blank-service fill on review, and the write-back on confirm.
* When they disagree, a reference gets learned into a column nothing searches,
* and the same page returns to the review queue every month forever.
*
* `meterNumber` is doing double duty for the three kinds whose printed
* reference DATMEX never held in `accountNumber`:
* - GAS, where the number lived in free-text notes,
* - PROPERTY_TAX, where `accountNumber` holds DATMEX.predial — a 3-4 digit
* office file number that is neither unique nor printed on any statement.
* The Tijuana municipal receipt prints an 8-digit account and no clave
* catastral at all, so it needs a column of its own; overwriting the legacy
* predial numbers to make room would destroy the only link back to the
* original records, and
* - FEDERAL_ZONE, where `accountNumber` holds DATMEX.zfed, which is not a
* reference of any kind but a peso amount: 3 of its 77 values carry cents
* (`246.06`, `2369.09`, `22653.94`) and one is negative. Searching it for
* the concession clave the receipt prints would never hit, and — worse —
* because every row already has a value, the `[field]: null` guards in
* `learnAccountRefs` and the blank-service fill would never fire either, so
* the same page would return to the review queue every bimester forever.
*/
export function scopedRefField(
kind: ServiceKind,
): "accountNumber" | "meterNumber" | null {
switch (kind) {
case "ELECTRIC": // CFE "NO. DE SERVICIO" -> DATMEX.rpu
case "WATER": // CESPT "Cuenta" / "No. DE CUENTA" -> DATMEX.agua
case "TELEPHONE": // Telnor "Teléfono" (LADA stripped) -> DATMEX.telefono
case "CABLE":
return "accountNumber";
case "GAS": // bajagas "Cuenta" -> recovered from notes into meterNumber
case "PROPERTY_TAX": // Tijuana's 8-digit municipal account
case "FEDERAL_ZONE": // ZOFEMAT concession clave, e.g. `12T012`
return "meterNumber";
default:
return null;
}
}
@Injectable()
export class StatementMatcherService {
constructor(private readonly prisma: PrismaService) {}
async match(parsed: ParsedStatement, expectedKind: ServiceKind): Promise<MatchResult> {
const kind = parsed.serviceKind ?? expectedKind;
// The uploader labels a batch with one service kind. If the parser reads a
// page as a different provider, that is a mis-sorted page, not a match —
// posting it would book a phone bill as a water charge.
if (parsed.serviceKind && parsed.serviceKind !== expectedKind) {
return this.unmatched(
`la página parece de ${parsed.provider} (${parsed.serviceKind}) pero el lote es de ${expectedKind}`,
);
}
const field = scopedRefField(kind);
if (field && parsed.accountRef) {
const hit = await this.byServiceField(kind, field, parsed.accountRef);
if (hit) return hit;
}
// The clave catastral is printed on CESPT bills as well as predial ones, so
// it rescues a page whose account number did not OCR — which happened on
// real samples, where the clave read cleanly and the account number did
// not. On the Rosarito and Ensenada predial layouts it is not a rescue at
// all but the only identifier the receipt carries, so a unique hit there is
// as good as any account-number match and is treated as one.
if (parsed.cadastralKey) {
const primary = kind === "PROPERTY_TAX" && !parsed.accountRef;
const hit = await this.byCadastralKey(kind, parsed.cadastralKey, primary);
if (hit) return hit;
}
if (!field && !parsed.cadastralKey) {
return this.unmatched(`no hay campo de búsqueda definido para ${kind}`);
}
if (!parsed.accountRef && !parsed.cadastralKey) {
return this.unmatched(
kind === "PROPERTY_TAX"
? "no se leyó ni la clave catastral ni la cuenta municipal"
: "no se pudo leer la referencia de la cuenta",
);
}
return this.unmatched(
parsed.accountRef
? `no se encontró ningún servicio de ${kind} con la referencia ${parsed.accountRef}`
: `no se encontró ninguna propiedad con la clave catastral ${parsed.cadastralKey}`,
);
}
private async byServiceField(
kind: ServiceKind,
field: "accountNumber" | "meterNumber",
ref: string,
): Promise<MatchResult | null> {
const rows = await this.prisma.propertyService.findMany({
where: { kind, [field]: ref },
select: {
id: true,
property: {
select: { customerId: true, customer: { select: { name: true } } },
},
},
});
if (rows.length === 0) return null;
const candidates = rows.map((r) => ({
propertyServiceId: r.id,
customerId: r.property.customerId,
customerName: r.property.customer.name,
}));
// Duplicate account numbers do occur in the legacy data (the office's own
// DUPLICADOS report existed for a reason), so every candidate is surfaced
// for the reviewer to choose rather than one being picked arbitrarily.
if (rows.length > 1) {
return {
propertyServiceId: null,
customerId: null,
note: `${rows.length} servicios comparten la referencia ${ref}`,
confident: false,
candidates,
};
}
return {
propertyServiceId: candidates[0].propertyServiceId,
customerId: candidates[0].customerId,
note: `coincidencia exacta por ${field === "accountNumber" ? "número de cuenta" : "medidor"} ${ref}`,
confident: true,
candidates,
};
}
private async byCadastralKey(
kind: ServiceKind,
key: string,
/** True when the clave is the identifier the statement was issued against. */
primary: boolean,
): Promise<MatchResult | null> {
const props = await this.prisma.property.findMany({
where: { cadastralKey: key },
select: {
customerId: true,
customer: { select: { name: true } },
services: { where: { kind }, select: { id: true } },
},
});
if (props.length === 0) return null;
const candidates = props.flatMap((p) =>
(p.services.length ? p.services.map((s) => s.id) : [null]).map((sid) => ({
propertyServiceId: sid as string,
customerId: p.customerId,
customerName: p.customer.name,
})),
);
if (candidates.length > 1) {
return {
propertyServiceId: null,
customerId: null,
note: `${candidates.length} propiedades comparten la clave catastral ${key}`,
confident: false,
candidates,
};
}
// When the clave is the *secondary* key — a utility bill that also happens
// to print it — the page is left for review, because the clave was not the
// number the statement was issued against and confirming is what teaches
// the matcher the account number for next month. When it is the primary key
// (Rosarito and Ensenada predial, which print nothing else), a unique hit
// is a real match and there is no second number to learn.
return {
propertyServiceId: candidates[0].propertyServiceId ?? null,
customerId: candidates[0].customerId,
note: primary
? `coincidencia exacta por clave catastral ${key}`
: `identificado por clave catastral ${key}; confirme para registrar también el número de cuenta`,
// A clave with no service row of the right kind behind it still needs a
// human: there is nothing to attach the posting to.
confident: primary && candidates[0].propertyServiceId != null,
candidates,
};
}
private unmatched(note: string): MatchResult {
return {
propertyServiceId: null,
customerId: null,
note,
confident: false,
candidates: [],
};
}
}
+52
View File
@@ -0,0 +1,52 @@
import {
IsBoolean,
IsEnum,
IsInt,
IsNumber,
IsOptional,
IsString,
MinLength,
} from "class-validator";
import { Currency, ServiceKind, StatementDocumentStatus } from "@jorgecuadros/database";
export class CreateStatementBatchDto {
@IsEnum(ServiceKind) serviceKind!: ServiceKind;
@IsOptional() @IsString() label?: string;
}
/** Staff correction of one document's extracted fields or its match. */
export class ReviewDocumentDto {
@IsOptional() @IsString() accountRef?: string;
@IsOptional() @IsNumber() amount?: number;
@IsOptional() @IsString() period?: string;
@IsOptional() @IsString() dueDate?: string;
@IsOptional() @IsString() matchedPropertyServiceId?: string;
@IsOptional() @IsString() matchedCustomerId?: string;
// Restricted to the review-reachable states: a client cannot declare a
// document POSTED, because only a successful ledger write may do that.
@IsOptional()
@IsEnum(StatementDocumentStatus)
status?: Extract<StatementDocumentStatus, "MATCHED" | "NEEDS_REVIEW" | "CONFIRMED">;
}
/**
* Post a batch's confirmed documents. The check-level fields are shared by
* every line, exactly as on the manual batch-capture screen — an OCR batch is
* still "these receipts, paid by this check".
*/
export class ConfirmBatchDto {
@IsString() @MinLength(1) checkNumber!: string;
@IsString() @MinLength(1) transactionDate!: string;
@IsOptional() @IsEnum(Currency) currency?: Currency;
/** Overrides the concept derived from the batch's service kind. */
@IsOptional() @IsString() typeId?: string;
/** Post as outstanding (sin fondos) — captured but not yet funded. */
@IsOptional() @IsBoolean() outstanding?: boolean;
/** Also post documents a reviewer explicitly marked CONFIRMED. */
@IsOptional() @IsBoolean() includeReviewed?: boolean;
}
export class ListBatchesQuery {
@IsOptional() @IsInt() page?: number;
@IsOptional() @IsInt() pageSize?: number;
}
@@ -0,0 +1,174 @@
import {
Body,
Controller,
Get,
Param,
Patch,
Post,
Query,
Req,
Res,
StreamableFile,
UploadedFiles,
UseGuards,
UseInterceptors,
} from "@nestjs/common";
import { FilesInterceptor } from "@nestjs/platform-express";
import type { ServiceKind, StatementDocumentStatus } from "@jorgecuadros/database";
import type { Request, Response } from "express";
import { AuthenticatedGuard } from "../auth/authenticated.guard";
import { AbilityGuard } from "../auth/ability.guard";
import { RequireAbility } from "../auth/require-ability.decorator";
import { AuditService } from "../common/audit.service";
import type { UploadedFileLike } from "../storage/upload-file";
import { StatementsService } from "./statements.service";
import { ConfirmBatchDto, ReviewDocumentDto } from "./statement.dto";
/**
* Statement OCR intake (RECEIPT_CAPTURE_SPEC §2).
*
* Nothing here writes to the ledger directly — confirming a batch delegates to
* BillingService, so an OCR-captured charge is indistinguishable from a
* hand-keyed one except for its `captureSource`.
*/
@Controller("statements")
@UseGuards(AuthenticatedGuard, AbilityGuard)
export class StatementsController {
constructor(
private readonly statements: StatementsService,
private readonly audit: AuditService,
) {}
private actingId(req: Request): string {
return (req.user as { id: string } | undefined)?.id ?? "";
}
/**
* Whether this deployment can ingest scans at all — the UI hides automatic
* capture without it. Both halves are needed: OCR to read the page, object
* storage to keep it.
*/
@Get("status")
async status() {
return {
ocrAvailable: await this.statements.ocrAvailable(),
storageAvailable: this.statements.storageAvailable(),
};
}
@Get("batches")
listBatches(@Query("page") page?: string, @Query("pageSize") pageSize?: string) {
return this.statements.listBatches(
Math.max(1, Number(page) || 1),
Math.min(100, Math.max(1, Number(pageSize) || 25)),
);
}
@Get("batches/:id")
getBatch(@Param("id") id: string) {
return this.statements.getBatch(id);
}
@Get("batches/:id/documents")
listDocuments(@Param("id") id: string, @Query("status") status?: string) {
return this.statements.listDocuments(
id,
(status || undefined) as StatementDocumentStatus | undefined,
);
}
/** The rendered page, so a reviewer can compare it against what was read. */
@Get("documents/:id/page")
async pageImage(@Param("id") id: string, @Res({ passthrough: true }) res: Response) {
const { stream, contentType, contentLength } = await this.statements.pageImage(id);
res.set({
"Content-Type": contentType ?? "image/png",
...(contentLength ? { "Content-Length": String(contentLength) } : {}),
});
return new StreamableFile(stream);
}
// --- writes ---------------------------------------------------------------
@Post("batches")
@RequireAbility("statement:ingest")
@UseInterceptors(
// A month of one company's statements is a handful of multi-page scans;
// 25 files at 50MB covers that with room to spare.
FilesInterceptor("files", 25, { limits: { fileSize: 50 * 1024 * 1024 } }),
)
async createBatch(
@UploadedFiles() files: UploadedFileLike[] | undefined,
@Query("serviceKind") serviceKind: ServiceKind,
@Query("label") label: string | undefined,
@Req() req: Request,
) {
const batch = await this.statements.createBatch(
files ?? [],
serviceKind,
this.actingId(req),
label,
);
void this.audit.log(this.actingId(req), "statement.batch.create", {
batchId: batch.id,
serviceKind,
fileCount: batch.fileCount,
});
return batch;
}
@Patch("documents/:id")
@RequireAbility("statement:review")
async review(
@Param("id") id: string,
@Body() dto: ReviewDocumentDto,
@Req() req: Request,
) {
const doc = await this.statements.review(id, dto, this.actingId(req));
void this.audit.log(this.actingId(req), "statement.document.review", {
documentId: id,
status: doc.status,
});
return doc;
}
@Post("documents/:id/reject")
@RequireAbility("statement:review")
async reject(@Param("id") id: string, @Req() req: Request) {
const doc = await this.statements.reject(id, this.actingId(req));
void this.audit.log(this.actingId(req), "statement.document.reject", {
documentId: id,
});
return doc;
}
/** Abandon a batch pending review — rejects every unposted page. */
@Post("batches/:id/discard")
@RequireAbility("statement:review")
async discard(@Param("id") id: string, @Req() req: Request) {
const result = await this.statements.discardBatch(id, this.actingId(req));
void this.audit.log(this.actingId(req), "statement.batch.discard", {
batchId: id,
rejected: result.rejected,
});
return result;
}
/** Post every matched document in the batch, against one check. */
@Post("batches/:id/confirm")
@RequireAbility("statement:review")
async confirm(
@Param("id") id: string,
@Body() dto: ConfirmBatchDto,
@Req() req: Request,
) {
const result = await this.statements.confirmBatch(id, dto, this.actingId(req));
void this.audit.log(this.actingId(req), "statement.batch.confirm", {
batchId: id,
posted: result.posted,
total: result.total,
checkNumber: dto.checkNumber,
});
return result;
}
}
@@ -0,0 +1,18 @@
import { Module } from "@nestjs/common";
import { BillingModule } from "../billing/billing.module";
import { OcrModule } from "../ocr/ocr.module";
import { StatementsController } from "./statements.controller";
import { StatementsService } from "./statements.service";
import { StatementMatcherService } from "./statement-matcher.service";
/**
* The concrete OCR engine is bound in OcrModule (see apps/api/src/ocr/) —
* everything downstream depends on the OcrProvider interface, so swapping
* Tesseract for a managed extraction API is a one-line change there.
*/
@Module({
imports: [BillingModule, OcrModule],
controllers: [StatementsController],
providers: [StatementsService, StatementMatcherService],
})
export class StatementsModule {}
@@ -0,0 +1,526 @@
import {
BadRequestException,
Inject,
Injectable,
Logger,
NotFoundException,
} from "@nestjs/common";
import {
Prisma,
type ServiceKind,
type StatementDocumentStatus,
} from "@jorgecuadros/database";
import { PrismaService } from "../prisma/prisma.service";
import { StorageService } from "../storage/storage.service";
import { BillingService } from "../billing/billing.service";
import type { UploadedFileLike } from "../storage/upload-file";
import { OCR_PROVIDER, type OcrProvider } from "./ocr/ocr.provider";
import { parseStatement } from "./parsers/statement-parser";
import { StatementMatcherService, scopedRefField } from "./statement-matcher.service";
import type { ConfirmBatchDto, ReviewDocumentDto } from "./statement.dto";
/**
* Default ledger concept per service kind. The names are the legacy
* `TYPE OF TRX` values already in `type_transactions`, resolved by name once
* per confirm rather than hard-coded as ids, which differ per environment.
*/
const CONCEPT_BY_KIND: Partial<Record<ServiceKind, string>> = {
ELECTRIC: "ELECTRIC",
WATER: "WATER",
TELEPHONE: "TELEPHONE",
GAS: "GAS BUTANO",
PROPERTY_TAX: "PROPERTY TAXES",
FEDERAL_ZONE: "FEDERAL ZONE",
CABLE: "CABLE",
};
/** Statuses a document can still be worked on from. */
const OPEN: StatementDocumentStatus[] = ["NEEDS_REVIEW", "MATCHED", "CONFIRMED"];
@Injectable()
export class StatementsService {
private readonly logger = new Logger(StatementsService.name);
constructor(
private readonly prisma: PrismaService,
private readonly storage: StorageService,
private readonly billing: BillingService,
private readonly matcher: StatementMatcherService,
@Inject(OCR_PROVIDER) private readonly ocr: OcrProvider,
) {}
ocrAvailable(): Promise<boolean> {
return this.ocr.available();
}
/** Scans are stored as blobs, so no object storage means no intake. */
storageAvailable(): boolean {
return this.storage.available;
}
// --- ingest ---------------------------------------------------------------
/**
* Accept a batch of scanned PDFs and start processing.
*
* Processing is kicked off but deliberately not awaited: 300 pages of OCR is
* minutes of CPU, far past any sane HTTP timeout. The caller gets the batch
* id immediately and polls its status, which is also what lets the review
* queue show partial progress.
*/
async createBatch(
files: UploadedFileLike[],
serviceKind: ServiceKind,
uploadedById: string,
label?: string,
) {
if (!files?.length) throw new BadRequestException("No se recibió ningún archivo.");
if (!(await this.ocr.available())) {
throw new BadRequestException(
"El servidor no tiene OCR instalado; no se pueden procesar recibos.",
);
}
// Checked here rather than at the first `put`, which would only surface as
// a FAILED batch minutes later.
if (!this.storage.available) {
throw new BadRequestException(
"El almacenamiento de documentos no está configurado; no se pueden " +
"guardar los recibos escaneados.",
);
}
const batch = await this.prisma.statementBatch.create({
data: { serviceKind, uploadedById, label, fileCount: files.length },
});
// Buffers are held for the background pass; the request's own copies would
// otherwise be garbage once the response is sent.
const copies = files.map((f) => ({ buffer: f.buffer, name: f.originalname }));
void this.process(batch.id, copies, serviceKind).catch(async (err) => {
this.logger.error(`Batch ${batch.id} failed: ${(err as Error).message}`);
await this.prisma.statementBatch.update({
where: { id: batch.id },
data: { status: "FAILED", error: (err as Error).message },
});
});
return batch;
}
/** Render → OCR → parse → match, one document row per page. */
private async process(
batchId: string,
files: { buffer: Buffer; name?: string }[],
serviceKind: ServiceKind,
) {
await this.prisma.statementBatch.update({
where: { id: batchId },
data: { status: "PROCESSING" },
});
let pageNumber = 0;
for (const file of files) {
// The source PDF is kept as well as the page images: it is the artifact
// the office actually received, and the only way to re-run a corrected
// parser over the original later.
const sourceKey = `statement/${batchId}/source-${pageNumber + 1}.pdf`;
await this.storage.put(sourceKey, file.buffer, "application/pdf");
const pages = await this.ocr.renderPages(file.buffer);
// Page images are still rendered and stored for every file, text layer or
// not: the review screen shows the reviewer the page, and "what the
// parser read" is only checkable against a picture of the paper.
const textLayer = await this.ocr.textPages(file.buffer).catch(() => []);
for (const [index, image] of pages.entries()) {
pageNumber += 1;
const storageKey = `statement/${batchId}/page-${pageNumber}.png`;
await this.storage.put(storageKey, image, "image/png");
try {
const embedded = textLayer[index] ?? null;
const ocr = embedded ?? (await this.ocr.recognize(image));
const parsed = parseStatement(ocr);
if (embedded) {
parsed.notes.unshift("texto leído del PDF original, sin OCR");
}
const match = await this.matcher.match(parsed, serviceKind);
const notes = [...parsed.notes, match.note].filter(Boolean);
// A confident field match is only trusted when nothing contradicts
// it: a barcode that disagrees with the printed number means one of
// the two was misread, and which one is a judgement call.
const trusted = match.confident && parsed.crossChecked !== false;
await this.prisma.statementDocument.create({
data: {
batchId,
pageNumber,
storageKey,
status: trusted ? "MATCHED" : "NEEDS_REVIEW",
ocrRawText: ocr.text,
ocrConfidence: new Prisma.Decimal(ocr.confidence.toFixed(3)),
provider: parsed.provider,
extractedAccountRef: parsed.accountRef,
extractedAmount:
parsed.amount != null ? new Prisma.Decimal(parsed.amount) : null,
extractedPeriod: parsed.period,
extractedDueDate: parsed.dueDate,
extractedCadastralKey: parsed.cadastralKey,
matchedPropertyServiceId: match.propertyServiceId,
matchedCustomerId: match.customerId,
matchNote: notes.join("; ").slice(0, 190),
},
});
} catch (err) {
// One unreadable page must not abandon the other 299.
await this.prisma.statementDocument.create({
data: {
batchId,
pageNumber,
storageKey,
status: "OCR_FAILED",
matchNote: (err as Error).message.slice(0, 190),
},
});
}
}
}
await this.prisma.statementBatch.update({
where: { id: batchId },
data: { status: "READY_FOR_REVIEW" },
});
}
// --- reads ----------------------------------------------------------------
async listBatches(page: number, pageSize: number) {
const [total, items] = await this.prisma.$transaction([
this.prisma.statementBatch.count(),
this.prisma.statementBatch.findMany({
orderBy: { createdAt: "desc" },
skip: (page - 1) * pageSize,
take: pageSize,
include: {
uploadedBy: { select: { name: true } },
_count: { select: { documents: true } },
},
}),
]);
return { items, total, page, pageSize, pageCount: Math.ceil(total / pageSize) };
}
async getBatch(id: string) {
const batch = await this.prisma.statementBatch.findUnique({
where: { id },
include: { uploadedBy: { select: { name: true } } },
});
if (!batch) throw new NotFoundException("Lote no encontrado.");
const counts = await this.prisma.statementDocument.groupBy({
by: ["status"],
where: { batchId: id },
_count: { _all: true },
});
const totals = await this.prisma.statementDocument.aggregate({
where: { batchId: id, status: { in: OPEN } },
_sum: { extractedAmount: true },
});
return {
...batch,
byStatus: Object.fromEntries(counts.map((c) => [c.status, c._count._all])),
pendingTotal: totals._sum.extractedAmount?.toFixed(2) ?? "0.00",
};
}
async listDocuments(batchId: string, status?: StatementDocumentStatus) {
return this.prisma.statementDocument.findMany({
where: { batchId, ...(status ? { status } : {}) },
orderBy: { pageNumber: "asc" },
include: {
matchedCustomer: { select: { id: true, name: true } },
matchedPropertyService: {
select: {
id: true,
kind: true,
accountNumber: true,
meterNumber: true,
property: { select: { id: true, addressLine1: true } },
},
},
},
});
}
/** The rendered page image, so a reviewer can read what the parser read. */
async pageImage(documentId: string) {
const doc = await this.prisma.statementDocument.findUnique({
where: { id: documentId },
select: { storageKey: true },
});
if (!doc) throw new NotFoundException("Documento no encontrado.");
return this.storage.getStream(doc.storageKey);
}
// --- review ---------------------------------------------------------------
/** Staff correction of an extracted field or of the match itself. */
async review(id: string, dto: ReviewDocumentDto, reviewedById: string) {
const doc = await this.prisma.statementDocument.findUnique({ where: { id } });
if (!doc) throw new NotFoundException("Documento no encontrado.");
if (doc.status === "POSTED") {
throw new BadRequestException("Este documento ya fue registrado.");
}
// Changing the service implies its owner; deriving the customer here rather
// than trusting a client-supplied pair is what stops a page being posted to
// one customer's ledger against another customer's service.
let matchedCustomerId = doc.matchedCustomerId;
let matchedPropertyServiceId = dto.matchedPropertyServiceId ?? undefined;
if (dto.matchedPropertyServiceId) {
const svc = await this.prisma.propertyService.findUnique({
where: { id: dto.matchedPropertyServiceId },
select: { property: { select: { customerId: true } } },
});
if (!svc) throw new BadRequestException("Servicio no encontrado.");
matchedCustomerId = svc.property.customerId;
} else if (dto.matchedCustomerId) {
matchedCustomerId = dto.matchedCustomerId;
// A reviewer picks a *customer*, not one of their service rows. Without
// a service the posting still works, but the confirmed reference has
// nowhere to be written back, so the same account would land in review
// again next month — which is exactly the behaviour that is supposed to
// make gas (whose numbers the migration never populated) a one-time cost.
// So: if the batch's service kind resolves to exactly one of that
// customer's services that has no reference yet, attach it. Exactly one
// — with two candidates there is no way to tell which meter or line the
// bill belongs to, and guessing would write a real number onto the wrong
// service.
const batch = await this.prisma.statementBatch.findUnique({
where: { id: doc.batchId },
select: { serviceKind: true },
});
const field = batch && scopedRefField(batch.serviceKind);
if (batch && field) {
const blank = await this.prisma.propertyService.findMany({
where: {
kind: batch.serviceKind,
[field]: null,
property: { customerId: matchedCustomerId },
},
select: { id: true },
take: 2,
});
if (blank.length === 1) matchedPropertyServiceId = blank[0].id;
}
}
return this.prisma.statementDocument.update({
where: { id },
data: {
extractedAccountRef: dto.accountRef ?? undefined,
extractedAmount:
dto.amount != null ? new Prisma.Decimal(dto.amount) : undefined,
extractedPeriod: dto.period ?? undefined,
extractedDueDate: dto.dueDate ? new Date(dto.dueDate) : undefined,
matchedPropertyServiceId,
matchedCustomerId,
status: dto.status ?? "MATCHED",
reviewedById,
reviewedAt: new Date(),
},
});
}
async reject(id: string, reviewedById: string) {
const doc = await this.prisma.statementDocument.findUnique({ where: { id } });
if (!doc) throw new NotFoundException("Documento no encontrado.");
if (doc.status === "POSTED") {
throw new BadRequestException("Este documento ya fue registrado.");
}
const updated = await this.prisma.statementDocument.update({
where: { id },
data: { status: "REJECTED", reviewedById, reviewedAt: new Date() },
});
// Rejecting the last open page settles the batch just as posting it would
// — without this, a fully-rejected batch sat in READY_FOR_REVIEW forever
// because only confirmBatch() ever closed one.
await this.closeIfDone(doc.batchId);
return updated;
}
/**
* Throw away a whole batch that is pending review: every page that has not
* been posted is marked REJECTED and the batch itself becomes DISCARDED.
*
* Refuses once any page is POSTED — those pages already wrote ledger rows
* against a check, and a "discarded" label on the batch would leave those
* charges unexplained. Reject the remaining pages individually instead.
*/
async discardBatch(batchId: string, reviewedById: string) {
const batch = await this.prisma.statementBatch.findUnique({
where: { id: batchId },
});
if (!batch) throw new NotFoundException("Lote no encontrado.");
if (batch.status === "DISCARDED") {
throw new BadRequestException("Este lote ya fue descartado.");
}
const posted = await this.prisma.statementDocument.count({
where: { batchId, status: "POSTED" },
});
if (posted > 0) {
throw new BadRequestException(
`No se puede descartar: ${posted} página(s) ya se registraron en el estado de cuenta.`,
);
}
const { count } = await this.prisma.statementDocument.updateMany({
where: { batchId, status: { notIn: ["POSTED", "REJECTED"] } },
data: { status: "REJECTED", reviewedById, reviewedAt: new Date() },
});
await this.prisma.statementBatch.update({
where: { id: batchId },
data: { status: "DISCARDED", completedAt: new Date() },
});
return { batchId, rejected: count };
}
// --- posting --------------------------------------------------------------
/**
* Post every confirmable document in a batch to the ledger.
*
* This goes through `BillingService.createBatch` — the same method the manual
* "Editor" screen uses — rather than writing `Transaction` rows directly, so
* OCR-sourced and hand-keyed receipts share one write path, one validation
* path and one audit trail. `source: "OCR"` and a per-line `captureRef` of
* the document id give the duplicate-post guard something to key on, so a
* batch confirmed twice cannot double-charge anyone.
*/
async confirmBatch(batchId: string, dto: ConfirmBatchDto, reviewedById: string) {
const batch = await this.prisma.statementBatch.findUnique({
where: { id: batchId },
});
if (!batch) throw new NotFoundException("Lote no encontrado.");
const docs = await this.prisma.statementDocument.findMany({
where: {
batchId,
status: { in: dto.includeReviewed ? ["MATCHED", "CONFIRMED"] : ["MATCHED"] },
matchedCustomerId: { not: null },
},
orderBy: { pageNumber: "asc" },
});
if (!docs.length) {
throw new BadRequestException("No hay documentos listos para registrar.");
}
const missing = docs.filter((d) => d.extractedAmount == null);
if (missing.length) {
throw new BadRequestException(
`Falta el importe en ${missing.length} documento(s): página(s) ` +
missing.map((d) => d.pageNumber).join(", "),
);
}
const typeId = dto.typeId ?? (await this.conceptFor(batch.serviceKind));
const result = await this.billing.createBatch(
{
domain: "UTILITY",
transactionDate: dto.transactionDate,
checkNumber: dto.checkNumber,
currency: dto.currency ?? "MXN",
typeId,
lines: docs.map((d) => ({
customerId: d.matchedCustomerId!,
// Charges are negative in this ledger: a negative amount is what the
// customer owes. The parser reads the printed (positive) figure, so
// the sign is applied here, at the single point where a statement
// becomes a ledger row.
amount: -Math.abs(Number(d.extractedAmount)),
reference: d.extractedAccountRef ?? undefined,
period: d.extractedPeriod ?? undefined,
outstanding: dto.outstanding ?? false,
})),
},
{ source: "OCR", refs: docs.map((d) => d.id) },
);
// `items[i]` is positionally parallel to `lines[i]` (seam guarantee 1), so
// the created rows zip straight back onto the documents that produced them.
await this.prisma.$transaction(
docs.map((d, i) =>
this.prisma.statementDocument.update({
where: { id: d.id },
data: {
status: "POSTED",
postedTransactionId: result.items[i].id,
reviewedById,
reviewedAt: new Date(),
},
}),
),
);
// Teach the matcher. When a document was matched by clave catastral or by
// hand because the scoped field was blank, writing the reference back means
// next month's statement for the same account matches on its own — this is
// what turns gas (whose numbers the migration never populated) from a
// permanent review queue into a one-time cost.
await this.learnAccountRefs(docs, batch.serviceKind);
await this.closeIfDone(batchId);
return { posted: result.count, total: result.total, checkNumber: dto.checkNumber };
}
/** Write a confirmed reference onto a service that had none. */
private async learnAccountRefs(
docs: { matchedPropertyServiceId: string | null; extractedAccountRef: string | null }[],
kind: ServiceKind,
) {
const field = scopedRefField(kind);
if (!field) return;
for (const d of docs) {
if (!d.matchedPropertyServiceId || !d.extractedAccountRef) continue;
await this.prisma.propertyService.updateMany({
// Only fills a hole — never overwrites a number already on file, which
// would let one misread page rewrite good reference data.
where: { id: d.matchedPropertyServiceId, [field]: null },
data: { [field]: d.extractedAccountRef },
});
}
}
private async closeIfDone(batchId: string) {
const open = await this.prisma.statementDocument.count({
where: { batchId, status: { in: OPEN } },
});
if (open === 0) {
await this.prisma.statementBatch.updateMany({
// `updateMany` + a status filter so a discarded batch is never quietly
// relabelled COMPLETED by a late reject on one of its pages.
where: { id: batchId, status: { not: "DISCARDED" } },
data: { status: "COMPLETED", completedAt: new Date() },
});
}
}
private async conceptFor(kind: ServiceKind): Promise<string | undefined> {
const name = CONCEPT_BY_KIND[kind];
if (!name) return undefined;
const row = await this.prisma.typeTransaction.findFirst({
where: { nameEn: name },
select: { id: true },
});
return row?.id;
}
}
+10
View File
@@ -0,0 +1,10 @@
import { Global, Module } from "@nestjs/common";
import { StorageService } from "./storage.service";
/** Global so any feature module can inject StorageService without re-importing. */
@Global()
@Module({
providers: [StorageService],
exports: [StorageService],
})
export class StorageModule {}
+132
View File
@@ -0,0 +1,132 @@
import {
Injectable,
Logger,
OnModuleInit,
ServiceUnavailableException,
} from "@nestjs/common";
import { ConfigService } from "@nestjs/config";
import {
CreateBucketCommand,
DeleteObjectCommand,
GetObjectCommand,
HeadBucketCommand,
PutObjectCommand,
S3Client,
} from "@aws-sdk/client-s3";
import type { Readable } from "node:stream";
/**
* S3 / MinIO object storage for document blobs. MySQL keeps only the pointer
* (`storageKey`) + metadata; the bytes live here. Same bucket the migration's
* `blob_extract.py` writes to, so keys stay under the `service/…` and
* `policy/…` prefixes it established.
*
* Env (see deploy/.env.dev): S3_ENDPOINT, S3_BUCKET, and creds — S3_ACCESS_KEY
* / S3_SECRET_KEY, falling back to MINIO_ROOT_USER / MINIO_ROOT_PASSWORD so a
* single MinIO credential set drives both the migration and the API.
*/
@Injectable()
export class StorageService implements OnModuleInit {
private readonly logger = new Logger(StorageService.name);
private readonly client: S3Client | null;
readonly bucket: string;
constructor(config: ConfigService) {
const endpoint = config.get<string>("S3_ENDPOINT");
this.bucket = config.get<string>("S3_BUCKET") ?? "jorgecuadros-documents";
const accessKeyId =
config.get<string>("S3_ACCESS_KEY") ?? config.get<string>("MINIO_ROOT_USER");
const secretAccessKey =
config.get<string>("S3_SECRET_KEY") ?? config.get<string>("MINIO_ROOT_PASSWORD");
if (!endpoint || !accessKeyId || !secretAccessKey) {
this.logger.warn(
"Object storage not configured (missing S3_ENDPOINT / credentials); " +
"document upload & download are disabled.",
);
this.client = null;
return;
}
this.client = new S3Client({
endpoint,
region: config.get<string>("S3_REGION") ?? "us-east-1",
credentials: { accessKeyId, secretAccessKey },
forcePathStyle: true, // MinIO needs path-style addressing
});
}
/** Best-effort bucket check on boot; never blocks API startup. */
async onModuleInit() {
if (!this.client) return;
try {
await this.client.send(new HeadBucketCommand({ Bucket: this.bucket }));
} catch {
try {
await this.client.send(new CreateBucketCommand({ Bucket: this.bucket }));
this.logger.log(`Created bucket "${this.bucket}".`);
} catch (err) {
this.logger.warn(
`Could not verify/create bucket "${this.bucket}": ${(err as Error).message}`,
);
}
}
}
/**
* Whether the deployment has object storage at all. Callers use this to
* refuse work up front instead of failing halfway through — a recibo batch
* that dies on its first `put` leaves a FAILED batch and no explanation the
* office can act on.
*/
get available(): boolean {
return this.client !== null;
}
private require(): S3Client {
if (!this.client) {
throw new ServiceUnavailableException(
"El almacenamiento de documentos no está configurado.",
);
}
return this.client;
}
async put(key: string, body: Buffer, contentType?: string): Promise<void> {
await this.require().send(
new PutObjectCommand({
Bucket: this.bucket,
Key: key,
Body: body,
ContentType: contentType,
}),
);
}
async getStream(key: string): Promise<{
stream: Readable;
contentType?: string;
contentLength?: number;
}> {
const out = await this.require().send(
new GetObjectCommand({ Bucket: this.bucket, Key: key }),
);
return {
stream: out.Body as Readable,
contentType: out.ContentType,
contentLength: out.ContentLength,
};
}
/** Best-effort blob delete; a missing object is not an error. */
async delete(key: string): Promise<void> {
if (!this.client) return;
try {
await this.client.send(
new DeleteObjectCommand({ Bucket: this.bucket, Key: key }),
);
} catch (err) {
this.logger.warn(`Failed to delete blob "${key}": ${(err as Error).message}`);
}
}
}

Some files were not shown because too many files have changed in this diff Show More