SYNC ran `run_all.py --sync` without `--stage`, so it depended on staged
Parquet under migration/output. That directory is part of the image, not a
volume, so any redeploy wiped it and the job died on the first transform:
FileNotFoundError: '/repo/migration/output/stg_utilities/datgral.parquet'
Re-staging is also what makes the job's own label true — without it a sync
would replay whatever upload staged last, not the files currently in the
ingest folder.
Staging now counts as a numbered step when it runs, so the Operaciones
progress bar moves during the slowest phase instead of sitting empty.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Seconds_Behind_Source cannot answer "is it moving?". While the SQL thread
works through one large transaction the lag counter holds still — often at
0 — even though the replica is not caught up. The relay backlog does move,
and it comes out of the SHOW REPLICA STATUS the panel already runs, so this
costs no extra query and no connection to the source.
Adds applyProgress(), which reads Source_Log_File / Read_Source_Log_Pos vs
Relay_Source_Log_File / Exec_Source_Log_Pos and reports the fetched-but-not-
applied byte delta plus a percentage. Both positions are source binlog
coordinates, so they are only comparable while the two threads are on the
same file; across files the delta is meaningless (positions restart at ~4 in
each new file) and is reported as null rather than as a huge negative number.
The percentage deliberately stops at 99.99 while any backlog remains —
binlog positions are large enough that a real backlog of a few KB rounds to
100% and would render a lagging replica as caught up.
Not folded into `healthy`: a non-zero backlog is the normal state of a
working replica between fetch and apply, so alarming on it would cry wolf.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A REIMPORT takes ~110 seconds and, until now, showed only a scrolling log —
there was no way to tell "halfway" from "wedged", which mattered the day one
actually did wedge.
run_all.py emits "[paso i/N] name" before each step and the API derives
progress from the job log. Emitting the marker from the Python rather than
having the UI count STEPS itself means the step count is stated in exactly
one place; adding a step cannot desync the display. Progress is derived, not
stored, for the same reason: the log is already the record of what happened,
and a separate counter could contradict it, which is precisely the confusion
a progress display exists to remove.
While RUNNING, step i is IN PROGRESS rather than finished, so only i-1 count
as done. Counting i would show 100% while the final step was still working —
and the final step (blob_extract) is the slowest, so the bar would sit at
"100%" for the longest stretch of the job.
BACKUP and RESTORE are a single mysqldump with no steps and deliberately
render no bar; a fabricated percentage would be worse than none. The safety
backup that precedes a REIMPORT is likewise named explicitly instead of
showing 0%, which reads as stuck.
Pinned by job-progress.spec.ts, including the literal line run_all.py emits,
so a change to the Python format fails a test rather than silently blanking
the panel.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The panel reported "Error SQL: Replicate_Ignore_Server_Ids:" against a
replica that was healthy — both threads running, zero lag.
`\s` matches newlines in JavaScript, so `^\s*NAME:\s*(.*)$` let the `\s*`
after the colon walk past an EMPTY field's line break and capture the
following line. Last_SQL_Error is blank on a healthy replica and
Replicate_Ignore_Server_Ids happens to be printed immediately after it, so
the blank error field returned the next field's name as its value. Every
empty field was affected; the visible damage was that a healthy replica
rendered as broken, which is the worst direction for a health panel to fail.
Fixed with `[^\S\n]` — horizontal whitespace only — on both sides of the
field name.
Extracted as replicaField() and pinned by replication.spec.ts against the
verbatim output of the live replica, keeping the empty Last_SQL_Error
adjacent to Replicate_Ignore_Server_Ids because that exact adjacency is what
broke. Also covers the literal "NULL" lag surviving as a distinct value from
empty, and a field name that is a suffix of another (Last_Error vs
Last_SQL_Error) not matching the wrong line.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Ops jobs run as a child of the API process, so no job can outlive it. When a
deploy landed 110 seconds into a REIMPORT, the child died and nothing was
left to finalize the row — it stayed RUNNING forever. Because startJob()
refuses to start while any RUNNING row exists, that one interrupted job
wedged the panel permanently with no way out from the UI; recovering it took
a manual UPDATE against the production database.
A fresh boot is proof that nothing survived, so this is unconditional rather
than filtered on age: "started recently" does not imply "still alive" here.
Rows are updated one at a time rather than with updateMany so the reason can
be APPENDED to the log. A job whose log simply stops mid-step with no
explanation is what made the first occurrence hard to diagnose.
Failure to reconcile is logged and swallowed: a wedged panel is bad, an API
that will not boot is worse.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
my.jorgecuadros.com serves customer balances from the Oracle VPS replica.
A replica whose SQL thread has stopped does not error — it keeps answering,
with data frozen at the moment it stopped — so nothing on the customer site
looks wrong and the only signal is a customer complaining about a stale
balance. This puts the failure somewhere a human sees it.
Deliberately does not trust the two fields an operator reaches for first.
Replica_IO_Running reports Yes while the SQL thread is stopped, because the
network thread keeps downloading binlog it will never apply; verified by
stopping SQL_THREAD and watching IO stay Yes. Seconds_Behind_Source reads
NULL whenever EITHER thread is down, so the card renders "sin dato" rather
than "0 s" — showing zero there would report an outage as perfect health.
The problem string is resolved most-specific-first for the same reason.
Shells out to the mysql client because the API has no MySQL driver and the
image already ships one. --ssl is required (the replica sets
require_secure_transport); --ssl-verify-server-cert=0 is deliberate and is
NOT the trade-off the website makes: this hop never leaves Tailscale and the
replica's firewall admits only this host, so WireGuard authenticates the
peer, whereas the DreamHost leg crosses the public internet and pins the CA.
The account behind it holds REPLICATION CLIENT and nothing else — it cannot
read a single row. REPLICA_DB_* unset is a supported state and renders "no
configurada", which is correct in dev and before cutover.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The "Flags del envío" panel lived inside the Servicios tab and only
governed the four bulk jobs. The pólizas half had no debug at all, so
there was no way to test a renewal notice without mailing a real
customer. The panel now lives in the /notificaciones shell above the
tabs and both halves read it.
`debug` on the renewal path diverts to the same override inbox as the
servicios jobs and deliberately does NOT write the `RenewalNotice` row
or advance the sweep's `lastSuccessfulAt` — the customer was not
notified, so nothing may gate the letter they are still owed.
`ignoreDayRestriction` and `useEmailLimit` stay estado-de-cuenta-only
and are labelled as such.
Both automatic sweeps are now operator-editable. The renewal cadence
was a `@Cron("0 6 * * *")` literal and servicios had no automatic run
at all; both now resolve through `NotificationScheduleService`, which
stores the cadence in `app_settings` and reinstalls the cron job on
save — no redeploy, no restart. Defaults preserve current behaviour:
pólizas 06:00 daily, servicios off. A scheduled run never inherits the
UI flags; it always sends for real.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
NOTIFICATION_ADMIN_EMAILS made "add Beto to the summaries" a redeploy —
the wrong unit of work for a list that changes when office staff change.
Adds `app_settings`, a key/value table for the configuration staff must
be able to change without a deploy, and `SettingsService`, which resolves
every key db -> env -> default and reports which of the three a value
came from. That ladder is what makes the move safe: a deployment behaves
exactly as before until somebody saves in the UI, and the screen can say
"this is still coming from the deployment" rather than implying somebody
chose it.
- new ability `setting:manage` (ADMIN) — deliberately above
`notification:send`, since redirecting the audit summaries is how
someone would quietly stop them being read
- GET/PUT /notifications/settings/admin-emails; read is open to any
logged-in user so the UI can display the list, write is gated
- resolved per job, not cached at boot, or we would reintroduce exactly
the restart-to-apply behaviour being removed
- a saved empty list means "nobody" and does NOT fall through to the env,
or clearing the field would keep mailing the people just removed
Credentials stay in env — see the model doc for where the line is drawn.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Renewal avisos left behind only a `RenewalNotice` row, whose sole job is
gating: a row with `sentAt` drops the policy off the pending list. It
cannot represent a failed send or a customer with no address, so the
Pólizas tab had no "Registro de envíos" to show and a sent notice simply
vanished from the list.
Renewals now write `email_notification_log` — the same table the four
bulk jobs write — as `RENEWAL_NOTICE` / `POLICIES`, with rows for
failures and no-email skips too. `RenewalNotice` keeps its gating role
unchanged; the two are complementary, not redundant.
- extend `EmailNotificationType` (+RENEWAL_NOTICE) and
`EmailNotificationServicio` (+POLICIES); `level` now carries the aviso
generation on renewal rows, so every reader must branch on the type
first (`notificationLevelLabel()` is the one place that lives)
- backfill emailed notices (`channel = 'EMAIL'`) into the log; MAIL-channel
rows are legacy printed letters and are deliberately left out
- extract `NotificationLogService`/`NotificationLogModule` as the single
writer, so a feature that sends mail records it without pulling the
bulk-job pipelines into its module
- `GET /notifications/log` and `/stats` take a comma-separated `servicio`
list; each tab reads its own slice. This also fixes the "Omitidos"
view, which mapped to no filter at all and showed every row
- share one `NotificationLogPanel` between both tabs
- pass SES_* / NOTIFICATION_ADMIN_EMAILS through the galactus compose,
which was missing them entirely — mail is runtime config, not a CI
secret, and the prod image sets NODE_ENV=production so a blank config
fails loudly instead of falling back to stdout
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The Pólizas tab now sends. Each pending row gets an "Enviar aviso" button
backed by POST /renewals/send, which renders, mails and records the notice
through the same path the daily sweep uses — so a hand-sent letter is
marked exactly like a swept one and drops off the pending list.
Sending is now the only way a notice gets marked as sent. Remove the
manual "Marcar impreso" / "Marcar EMAIL" buttons and the endpoint behind
them (POST /policies/:id/renewal-notices, PoliciesService.markRenewalNotice,
MarkRenewalNoticeDto): they wrote a sentAt with no mail behind it, which
let the list claim a customer was notified when nothing was sent.
sendOne refuses a generation that already has a sentAt (409) so a double
click cannot mail the customer twice, and 400s when the customer has no
email on file. Sweep and single send share the new deliver() helper.
Add POST /notifications/run-all: runs the four notification jobs
(outstanding, payment confirmation, account status, trust confirmation)
sequentially with one shared set of flags from "Flags del envío".
Sequential rather than parallel — the jobs share the SES transport and
account status can self-throttle via useEmailLimit. A job that throws is
captured and the sweep continues, so one bad query cannot swallow the
other three envíos; the aggregate response carries per-job results plus
summed sent/skipped/failed and an errors count.
Audited as a single notification.run-all.run entry so one staff click is
one audit row. UI adds the button to the flags card, with a confirm when
debug is off, and a per-job summary in "Última respuesta".
MailModule's provider used a `useFactory` with no `inject`, so the factory
received `undefined` and `new MailService(config)` threw on `config.get`,
taking the whole API down at boot. The module also wasn't actually
`@Global()` even though both NotificationsModule and RenewalsModule inject
MailService without importing it — that would have failed next. Replaced the
factory with a plain provider (ConfigModule is already `isGlobal`) and marked
the module global.
On the web side, mass email and renewal notices were two menu entries doing
the same job — telling a customer something by email. They are now two tabs
of `/notificaciones` (Servicios and Pólizas), following the Captura pattern:
`/renovaciones` still resolves, opening the same screen on its Pólizas tab so
existing bookmarks keep working.
The notifications page was also the last screen written in raw inline styles,
with blue buttons and filter pills that appear nowhere else in the app. It now
uses the shared design system: btn-primary/btn-outline, the seg segmented
control, card, tx-table, pager, and the servicios/fideicomiso badges.
Two supporting fixes found on the way: NOTIFICATION_STATUS_COLORS hardcoded
hex instead of the theme's positive/negative/muted vars, and `.small` was
referenced in 19 places across the app but never defined in globals.css.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Replaces the four legacy PHP scripts under email.notifications/send*.php
with a single NestJS module. Four jobs (outstanding payments, payment
confirmations, account-status alerts with day-of-week gates, trust
payment confirmations) share one MailService modelled on StorageService:
env-driven SES client, null fallback in dev with console logging, refuses
to send in production when unconfigured.
Schema adds email_notification_log (every attempt, sent/failed/skipped)
and account_status_history (one row per threshold hit, Job 3). Enums
encode the legacy wire shape so external log scrapers keep parsing
notificationType keys verbatim.
Web adds /notificaciones with four trigger cards, a flags panel, and a
paginated log browser. New notification:send ability gates all four
endpoints at MANAGER, matching the renewal:send trust tier.
INSURANCE_FEATURES_SPEC §1. The office printed and mailed renewal letters
from the legacy CONTROL <ramo> RENEW[2/3] paper log; 91% of policyholders
have an email on file, so send the notice instead and keep the paper log
as the fallback.
A daily cron (06:00 America/Tijuana) sweeps three generations off
policyTo — 30 and 15 days before expiry, 7 days after — sends each
through SES, and upserts RenewalNotice by [policyId, generation] so a
policy is never notified twice for the same milestone. RenewalNotice now
records providerMessageId, so a later bounce or complaint webhook can be
traced back to the row that sent it.
- customers.emailOptOut excludes a customer from every sweep; editable
from the customer form
- scheduled_job_states holds the sweep's lock and last successful run;
the window is widened to cover days the job did not run, so a weekend
outage does not silently drop a generation
- SES unconfigured is not an error outside production — messages are
logged and skipped, so dev and CI never send
- /renovaciones (renewal:send, MANAGER+) lists what is pending per
generation, runs the sweep by hand, and marks a notice sent by mail
for the customers with no email
- POST /policies/:id/renewal-notices records that manual mark
- the aviso-renovacion report and the emails now share one projection
(reports/renewal-letter.ts) instead of two copies of the mapping
A bad scan, the wrong PDFs or a duplicate upload used to leave a batch
sitting in READY_FOR_REVIEW forever, because the only exits were confirm
(posts to the books) or rejecting every page one at a time. Add a
DISCARDED terminal status to both OCR domains and a single endpoint per
domain that rejects every page still pending in one shot.
Discarding is refused once anything has landed: statements once a page is
POSTED, policies once a page is APPLIED. Those batches did real work and
have to be settled page by page.
- POST /statements/batches/:id/discard
- POST /policy-ocr/batches/:id/discard
- shared DiscardBatchCard on both review screens, gated the same way
Mirrors the utility statement intake on the insurance side: a policy_ocr
batch/document pair of tables, a GMX parser, a matcher keyed on
Policy.policyNumber, and a "Captura" screen under /polizas that proposes
policy -> customer for staff to confirm.
Lifts the OCR seam out of StatementsModule into its own OcrModule so
PolicyOcrModule can inject OCR_PROVIDER without taking on the rest of
the statement pipeline; StatementsModule now imports it and binds
nothing itself.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds the ZONA FEDERAL TIJUANA parser to the statement intake, measured
against 8 pages of real "Zona Federal Marítimo Terrestre" receipts — the
federal maritime-zone occupancy fee the municipality bills on beachfront
lots. Provider read on 8/8, amount on 8/8 (each verified against the
paper), concession clave on 6/8, period on 8/8, deadline on 2/8.
Four things the corpus forced:
- Tijuana bills predial and zona federal from the same treasury: same
header, same Paseo del Centenario address, same ATB-541201 RFC. Every
predial discriminator matches a zona federal page too, so whichever
rule is asked first wins it. The only words exclusive to this layout
are "Marítimo Terrestre", so its brand rule is asked ahead of all
three predial ones — and its structural rule, anchored on the stub's
"Derechos de ocupación", ahead of theirs.
- FEDERAL_ZONE.accountNumber is an amount, not a reference. It holds
DATMEX.zfed, whose 77 values include 246.06, 2369.09, 22653.94 and a
negative -1679, while the concession claves these receipts are keyed
by appear nowhere in the database. Matching on that column could never
hit — and because every row already has a value, the `[field]: null`
guards on learnAccountRefs and on the review blank-service fill would
never fire either, so every page would return to the queue every
bimester forever. The clave moves to meterNumber, joining gas and
Tijuana predial, and the first confirm teaches the match.
- The payable figure is not the printed subtotal. The municipality
rounds to whole pesos and prints the difference on its own "Ajuste Ley
Hacienda Mpal" line (-$0.05 against a 591.05 subtotal, $0.21 against
2,872.79). The "Total a pagar" box carrying the rounded figure sits on
a grey fill and OCR'd on 1 of 8 pages; the SubTotal row read on 8 of
8. So the amount is the rounded subtotal, cross-checked against the
printed box wherever it survives — where it did, it agreed.
- The clave is 2 digits, a letter and 3 digits (12-T -012), not the
cadastral shape, and the letter is kept as printed: toDigits maps D to
0, which turns a real 14-D -014 into 140014. It is printed twice,
which rescued a page whose heading was struck through by the office's
own highlighter — the failure mode behind both missing claves.
Deriving the deadline from the bimester is deliberately not attempted:
it is the 17th of the month after the bimester closes on a current bill,
but four of these eight are late (a $1,000 Multa) and print a
recalculated date, so a derived date would be wrong on exactly the pages
a human most wants to see.
Re-ran the earlier corpora (25 pages: predial Tijuana/Rosarito/Ensenada,
CFE, CESPT, Telnor) through detection to confirm the new rules steal
nothing — all 25 still read as their original provider, including the
five Tijuana predial pages that share the RFC.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds four parsers to the statement intake — GAS TIJUANA plus one per
municipality, because Tijuana, Rosarito and Ensenada issue three
completely different predial documents — and a text-layer fast path for
the born-digital invoices the gas company sends.
Measured against a new corpus of 14 documents / 29 pages: provider read
on 29/29, amount on 26/29, and 21/29 auto-matched against the dev
database (22/29 identified). The eight review cases are all legitimate.
Five things the corpus forced:
- Not every statement is a scan. The gas invoices are born-digital CFDIs
whose text layer is exact; rasterising them only loses information (one
sample turned `MEDIDOR: VM01014426` into `ar (LTR): 014420`). The new
`OcrProvider.textPages` reads the embedded layer via `pdftotext
-bbox-layout` — same poppler package as `pdftoppm`, so no new
dependency — and OCR stays the fallback for real scans. Poppler's own
`<line>` grouping follows text flow rather than the page, so words are
regrouped by vertical position; without that, a two-column header
leaves every label separated from the value printed beside it.
- The clave catastral is not two letters and six digits. Position three
is a letter in 15 of the 932 stored claves, and digitising the whole
tail mapped a real `MMB01041` to a nonexistent `MM801041`.
- Tijuana predial prints no clave at all. Its only identifier is an
8-digit municipal account carried in a 32-digit payment barcode, which
the legacy database never held, so it goes in `meterNumber` alongside
gas — `accountNumber` holds `DATMEX.predial`, which is not a
per-property key and must not be overwritten. Those pages start cold
and are taught by the first confirm.
- On Rosarito and Ensenada the clave is the primary key, not a fallback:
those receipts print nothing else, so a unique hit auto-matches. On a
utility bill that merely happens to print one it stays a review hint.
- A misread `$` is the dangerous failure. An Ensenada receipt for
$2,203.00 OCR'd as `82,203.00`, which would post a charge 37x too large
and look ordinary in the ledger. Predial amounts now require a literal
`$` and a page that cannot produce one goes to review.
The scoped match field is now one exported function rather than three
copies of `kind === "GAS" ? ... : ...`, since the lookup, the
blank-service fill and the confirm write-back have to agree or a
reference gets learned into a column nothing searches.
First tests in this package: 23 specs over the parsers and the text-layer
reader, every fixture a verbatim OCR excerpt from a real receipt. Adds
the jest config they need and a build tsconfig so they stay out of dist.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Backend already returns per-status counts via byStatus; render a real
progress bar (X% / N de M / en cola) while PENDING_OCR pages remain,
using the existing progress-track CSS. Falls back to indeterminate
when no docs have been reported yet.
Every backup on galactus died with:
mysqldump: unknown variable 'set-gtid-purged=OFF'
respaldo incompleto eliminado
Alpine's mysql-client is MariaDB's, so `mysqldump` inside the API
container is a shim over `mariadb-dump`, which has no --set-gtid-purged.
That took out BACKUP and, because they take a safety dump first, SYNC
and REIMPORT too.
Probe `mysqldump --help` and pass the flag only when it is advertised,
calling `mariadb-dump` directly otherwise — MariaDB writes no GTID state
unless asked with --gtid, so there is nothing to suppress. Testing
whether mariadb-dump merely exists would be wrong: on a host carrying
both clients it would shadow a perfectly good MySQL mysqldump.
The probe uses a command substitution rather than `--help | grep -q`
because PIPEFAIL is in effect for these commands and grep closing the
pipe early would report a supported flag as unsupported.
pre-migrate-backup.mjs is unaffected — it dumps from a real mysql:8.4
image, not from the API container.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The ingest upload used fetch(), which cannot report request-body
progress, so the only feedback was a static "Cargando…" label — no way
to tell a stalled 2 GB upload from a working one.
Switch uploadFile() to XMLHttpRequest and expose an optional onProgress
callback reporting loaded/total bytes, a smoothed transfer rate and a
remaining-time estimate. The Operaciones ingest table renders a progress
bar row under the file being uploaded. Once the bytes are all sent the
server still has to write the file, so that tail reads "Procesando en el
servidor…" rather than parking at 100%.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Scanning a stack of bills and keying them in are the same daily job, ending
in the same ledger path, so OCR intake becomes a mode of the capture screen
instead of a second menu entry:
- components/Captura.tsx holds the mode switch; the manual check form moves
verbatim to components/ManualCheckCapture.tsx and the OCR intake to
components/StatementIntake.tsx.
- /estado-cuenta/lote opens on manual, /recibos on automatic — both render
Captura, so batch-review links and old bookmarks still land right.
- Nav drops "Recibos (OCR)"; "Captura" covers both, with a NavLink.aliases
field so /recibos still highlights it.
Also fixes the "El almacenamiento de documentos no está configurado" failure
staff hit on upload. Uploading with no object storage configured used to
succeed, then die on the first put minutes later, leaving a FAILED batch
whose only explanation was that string. createBatch now refuses up front,
GET /statements/status reports storageAvailable alongside ocrAvailable, and
the intake tab explains the situation instead of offering an upload that
cannot work. S3_* documented in .env.example (deploy stacks already set it).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Staff key 300+ utility statements per company per month by hand. This adds
the ingest -> split -> OCR -> match -> review pipeline that proposes customer
and amount per page instead (RECEIPT_CAPTURE_SPEC §2), posting through the
existing BillingService.createBatch seam with source=OCR and a per-document
captureRef so machine and hand capture share one write path and audit trail.
Everything was designed against 10 real scanned statements (46 pages of CFE,
CESPT and Telnor bills) rather than from the sample-free spec. The scans have
no text layer at all — they are camera images — so OCR is mandatory, and they
arrive bundled one customer per page. Measured on those pages the parser
identifies the provider 46/46 and reads an account reference 43/46; against
the dev database that is 39/46 (85%) exact auto-match, 40/46 identified, with
the rest genuine review cases. That closes the OCR-provider question in favour
of self-hosted Tesseract: it clears the bar for a queue where a human confirms
every row, and OcrProvider keeps a managed API a one-line swap.
The samples corrected three things the spec had wrong or unknown:
- Clave catastral is NOT predial. DATMEX.clave (934 rows) is what CESPT and
predial bills print; DATMEX.predial, which PROPERTY_TAX.accountNumber holds,
has 663 distinct values across 1135 rows and appears on no statement. The
clave now lives on Property.cadastralKey as the matcher's secondary key;
predial is left untouched. This had been blocking predial matching.
- Gas was recoverable: 160 of 334 DATMEX.gas values are real account numbers
(the rest are ESTACIONARIO/CILINDRO descriptors), now in GAS.meterNumber.
- Phone is one billed line per property (534/18/1 across phone1/2/3), so the
new TELEPHONE ServiceKind backfills from phone1 only, not three rows.
Matching is scoped to one column per service kind and never reads the customer
name — a CESPT receipt prints ARNAIZ ROSAS ELSA AURORA for an account this
office holds under CATT, RANDY, because the printed name is the registrant,
not the current owner. Where a provider prints a payment barcode it beats the
printed label (one CFE label OCR'd a digit too many while its barcode was
correct) and the two cross-check, with disagreement forcing review.
Confirming a document whose service had no reference writes it back, so gas
and any other cold start is a one-time cost rather than a permanent queue.
Verified end to end against the live dev API and MinIO: real scans uploaded
over HTTP, matched, confirmed against a check, and the resulting rows checked
in MySQL (negative amounts, captureSource=OCR, concept derived from the batch
kind, captureRef linking back to each page). Re-confirming a posted batch is
refused. Test data was removed afterwards.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
v1.0.0's images were built from db2bd54, which predates the full-hash
footer. Deploying 1.0.0 would ship the abbreviated footer, so the version
that actually goes to galactus is 1.0.1.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The footer abbreviated to 7 characters, so the line read
"v master · 19f0319". That line exists to be pasted into `git show` or
compared against a registry tag, and an abbreviation makes both a manual
step — while the full 40-char value was already baked into the image
(build.yml passes `github.sha` whole, and /version returns it untouched).
`shortSha` had no other caller, so it goes with it.
The span gets `overflow-wrap: anywhere` and `min-width: 0`: hex offers no
break opportunity, and the api/web mismatch branch renders two of these
hashes side by side, which would otherwise push a phone into horizontal
scroll. Measured at a simulated 360px with both hashes present — the span
wraps, and documentElement.scrollWidth stays equal to clientWidth.
Verified in the browser against the dev database: footer renders
"v1.0.0 · db2bd54c0ffee1234567890abcdef0123456789a", hash length 40.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every manifest still read 0.1.0 while the deployed images were addressed by
the moving tag `latest`. That combination is what hid the stale-image bug:
a checkout could not be placed against a running container, and `latest`
silently kept serving two-commit-old web code through a green deploy.
Tagging v1.0.0 makes docker/metadata-action publish immutable `1.0.0` and
`1.0` image tags, so deploys can name a version instead of a moving target.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The Operaciones panel (backup, restore, sync, re-import) shelled out to
mysqldump as the application user, parsed straight out of DATABASE_URL.
`--single-transaction` issues FLUSH TABLES, which needs the global RELOAD
privilege, and the app user is granted only ALL ON jorgecuadros.* plus
USAGE ON *.*. BACKUP failed outright; SYNC and REIMPORT failed with it,
since both take a safety backup first.
An admin credential is now supplied out of band via OPS_DB_ADMIN_USER /
OPS_DB_ADMIN_PASSWORD, mirroring what deploy/scripts/pre-migrate-backup.mjs
already does, rather than permanently elevating the user the API serves
requests as. Host, port and database still come from DATABASE_URL, so the
override can only change who logs in, never which server. Unset, it falls
back to the DATABASE_URL credentials and warns — local development is
unaffected.
Two defects in the dumps themselves, both shared with the deploy backup
before it was rewritten:
- No --set-gtid-purged=OFF. The production server is the replication source
with GTID on, so every dump embedded SET @@GLOBAL.GTID_PURGED and was
unrestorable onto the server it came from — the one thing the restore
screen is for.
- The pipeline's exit status was gzip's, and gzip succeeded. A mysqldump
that died on its first statement left a small, perfectly valid archive
that the job recorded as SUCCESS and the restore screen listed as an
ordinary restore point. Dumps now run under `set -o pipefail`, assert a
CREATE TABLE count, and delete their own output on failure. Verified with
a stubbed mysqldump: a failing dump exits 1, surfaces the real error,
removes the partial file, and — critically — stops SYNC/REIMPORT before
the ETL touches anything.
Restores gained pipefail too: a corrupt archive made gunzip fail while
mysql, fed a truncated stream, could still exit 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Prod came up with nobody able to log in, in two separate ways.
1. No sign-in account exists. `prisma migrate deploy` creates tables, never
rows, and nothing in the deploy path seeds one — deliberately, since making
an administrator should not be a side effect of shipping code. But
apps/api/scripts was not in the runtime image either, so the only way to
create the first account was to run the script from a developer machine
against a production DATABASE_URL. Ship scripts/ in the image so it can be
run on the host with docker exec. Still never run automatically.
2. Login could not establish a session at all. cookie.secure followed NODE_ENV,
the image sets NODE_ENV=production, and the app is served over plain HTTP —
express-session then silently emits NO Set-Cookie header. POST /auth/login
still answered 200 with the full user object, no session was created, every
later request 403'd, and the UI would have looped back to /login. It reads
as an auth bug and is really a transport mismatch.
The flag is now driven by SESSION_COOKIE_SECURE, still defaulting to
NODE_ENV. An EMPTY value counts as unset rather than false, because compose
turns an absent `${SESSION_COOKIE_SECURE:-}` into the empty string and the
naive check would have quietly dropped Secure on any deployment that merely
passed the variable through.
galactus sets it to "false". That is acceptable ONLY because the host is
reachable exclusively over Tailscale, so WireGuard already encrypts the
wire. It must go back to "true" when the app is served over TLS or exposed
off-tailnet; behind a TLS-terminating proxy, set trust proxy instead.
Verified against live prod: seeded an admin, POST /auth/login returns 200 with
full ADMIN abilities, a wrong password is rejected with 401, and no Set-Cookie
was present before this change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first successful galactus deploy came up all-green while the web tier was
running a build from two commits earlier. The registry held web:latest from
3ff56e6; the host still had a web:latest cached from 4ee7ec7; the deploy
reported success and served the old one. The API was only current because it
had been pulled by hand during earlier debugging.
Two independent failures, both fixed here.
1. Images are not pulled. The deploy action's `pull: true` does not reliably
refresh an already-cached moving tag on a standalone endpoint. Added a
Pull images step (deploy/scripts/pull-images.mjs) that pulls each image
through Portainer's Docker API with registry credentials and fails the
deploy if a pull fails — note the endpoint answers 200 even when the pull
errored, so the stream body has to be inspected, not just the status.
2. The drift check could not see it. Both the verify step and the web footer
compared APP_VERSION, but on a branch build BOTH tiers report "master", so
equality proved nothing. They now compare gitSha, which is the only field
that differs between two builds of the same branch. api and web come from
one matrix run, so a difference can only mean an image was not replaced.
This needed a /version on the web tier too — previously its build identity
was only readable by scraping window.__APP_BUILD__ out of the HTML.
pull-images.mjs builds the X-Registry-Auth header as URL-safe base64 WITH
padding: Node's "base64url" omits the padding and Portainer's Go decoder
rejects it with "Illegal base64 data at input byte N".
Verified against galactus: pulls both images, and exits non-zero on a
nonexistent tag.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Closes the gap between "what tag did I deploy" and "what is actually running",
and gives the schema a history that can be reasoned about across releases.
Migrations
- Baseline the existing schema as 0000_init (migrate diff --from-empty). The
schema had only ever been applied with `prisma db push`, so no history
existed and schema state was disconnected from app version. Existing
databases must be baselined once with `migrate resolve --applied 0000_init`;
the workflows print this remedy on P3005.
- Run `prisma migrate deploy` as a deploy STEP, not the container CMD — as a
CMD, N replicas would race each other applying the same migration.
Version reporting
- GET /version on the API reports the APP_VERSION / GIT_SHA / BUILD_DATE that
build.yml already baked into both images but nothing ever read.
- The web footer shows the web build and flags an api/web mismatch. The two
cannot drift at build time (one matrix run) but can at deploy time.
- Both deploy workflows now fail if the running API does not report the tag
that was dispatched — a stack naming a tag is not proof of what is running.
- scripts/set-version.mjs stamps every package.json, which had all sat at
0.1.0 while real releases shipped as v1.x.
Pre-migrate backup
- deploy/scripts/pre-migrate-backup.mjs dumps the database from INSIDE the
still-running old API container over Portainer's Docker API, so the file
lands in the volume the Operaciones restore screen reads. A dump taken on
the CI runner would be unreachable by the only restore path we have.
Verifies the artefact with `gzip -t` before letting the migration proceed.
galactus
- deploy/galactus/*.compose.yml: standalone-Docker ports of the Swarm stacks.
Plain compose silently ignores `deploy:`, so restart_policy becomes
`restart: unless-stopped` — without it nothing returns after a host reboot.
- .gitea/workflows/deploy-galactus.yml drives endpoint 3 with its own secrets.
Fixes
- deploy.yml passed `endpoint_id` and `pull_image` to
cssnr/portainer-stack-deploy-action, which has no such inputs (they are
`endpoint` and `pull`). The endpoint was silently never set.
docs/DEPLOY_AND_MIGRATIONS.md documents expand/contract as the rule for schema
changes: Prisma has no down-migrations, so a code rollback never rolls the
schema back, and restoring the replication master from a dump diverges every
replica.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The office keeps more than one operating account (Utilities banks in MXN,
Seguros in USD), but bank_transactions was a single implicit MXN register by
design. Adds Bank/BankAccount and makes every read and write in the module
scoped to exactly one account.
Schema:
- Bank / BankAccount. Currency is fixed per account and BankTransaction has
no currency column of its own — a movement inherits its account's, the way
a real bank account doesn't mix currencies.
- BankTransaction.bankAccountId, required. A movement with no known account
isn't reconcilable against a statement.
- @@index([bankAccountId, transactionDate]): every read now filters by
account and orders/groups by date.
Migration:
- backfill_bank_accounts.py seeds Scotiabank + "Utilities — Scotiabank (MXN)"
and backfills all 22,669 existing rows onto it, then promotes the column to
NOT NULL and attaches the FK. Standalone because prisma db push cannot add
a required column to a populated table. Idempotent; re-running once a second
account exists does not re-point rows.
- run_all.py runs it (both modes) before transform_bank.py, which now resolves
the account by label and fails fast if it is missing.
API:
- ?bankAccountId= required on list/stats/facets/summary — not optional with an
"all accounts" default, since summing an MXN and a USD register repeats the
currency-collapsing mistake the billing module exists to prevent. Missing is
400, unknown is 404.
- facets() had no account clause at all and summary() has two raw-SQL rollups;
all three are now parameterised. Scoping only one of summary's queries would
leave the year list and its drill-down describing different books.
- New bank/accounts + bank/banks sub-resource under a MANAGER
bank:manage-accounts ability. currency is absent from the update DTO: booked
movements are denominated in it, so editing would re-denominate history.
Capture into a closed account is rejected.
Web:
- /banco gains an account picker (remembered per browser) and reads every
figure in the selected account's currency; the "single currency (MXN)"
doc-comment and the hardcoded MXN formatting are gone.
- New /banco/cuentas for banks and accounts. Accounts are closed, never
deleted — the FK is required, so deleting one would destroy its register.
- /inicio's chequera card names the account it is reading instead of implying
a single register.
Verified against dev + browser: a second USD account showed full read/write
isolation from the MXN register, whose totals were unchanged (22,669
movements, net 1,014,266.97).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two follow-ups to the text-size control.
Spacing now scales with the text. All padding, margin, gap and min-height
declarations in globals.css move from px to rem (263 declarations, converted
mechanically), so --ui-scale drives the whole layout rather than just the
glyphs. Deliberately left in px: border widths, which must stay hairlines;
box-shadow offsets; border-radius, which reads as bloated when scaled on large
cards; --shell-max, a container cap that must not outgrow the viewport; and
media-query breakpoints, which are conditions rather than declarations. With
spacing following along, the presets gain a 1.5 "Máximo" step and MAX_UI_SCALE
rises from 1.4.
The preference now lives on the account instead of only in one browser.
User.uiScale (Float, default 1) is added to the schema and to the safe select,
so it rides along on /auth/login and /auth/me. PATCH /auth/preferences writes
it, guarded by AuthenticatedGuard only — every role including VIEWER may set
their own, and the target is always the session's user id, never a body
parameter, so this cannot be used to touch another account. The global
ValidationPipe's whitelist rejects any extra field, so role cannot ride in
alongside uiScale.
localStorage stays, demoted to a pre-paint cache for the layout.tsx script;
AppShell reconciles it against the account once /auth/me answers, with the
account winning. FontScaleControl becomes a controlled component since the
same value is now edited from the appbar and the drawer.
Verified against the dev API: PATCH persists and is reflected by a subsequent
/auth/me, out-of-range values are rejected 400, and an extra "role" field in
the body is rejected 400.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The appbar had grown to 11 flat links with no responsive behaviour, and
overflowed below ~1100px.
Nav is now 7 top-level entries: Inicio, Clientes, Pólizas, Propiedades,
Reportes stay one click away, while the movement screens (Captura, Estado de
cuenta, Chequera) and the admin screens (Catálogos, Usuarios, Operaciones)
collapse into "Cobranza" and "Admin" dropdowns. Groups are ability-filtered
and disappear entirely when the user can see none of their items, so VIEWER
never renders an empty Admin menu. activeHref now scans the flattened link
list, and a group trigger highlights while one of its children is current.
Below 980px the nav collapses to a burger drawer that lists every group
expanded, closing on navigation and on Escape.
Text size is user-adjustable app-wide. Every font-size in globals.css is
converted from px to rem (mechanically, 133 declarations) and the root size
becomes calc(100% * var(--ui-scale)), so one variable on <html> rescales the
whole UI. The preference persists in localStorage and is applied by a
pre-hydration script in layout.tsx to avoid a flash at the default size; the
Aa control lives in the appbar and, as a segmented row, in the drawer.
Spacing stays in px by design, which is why 1.3 is the largest preset.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Implements docs/RECEIPT_CAPTURE_SPEC.md §1, the legacy "Editor"
replacement, on top of the single-movement capture from plan step 6.
No new abilities: batching and resolving are both capturing.
- outstanding (legacy NOPAGO): capture flag, ?outstanding= filter, and
POST /billing/:id/resolve-outstanding (gated ledger:create, not
ledger:void — resolving completes a capture rather than reversing
one). Outstanding rows are excluded from every balance aggregate,
matching the legacy SALDOS ULTIMO 0 query's HAVING NOPAGO = 0, but
still count in the movement browser's filtered totals.
- POST /billing/batch: many customers' receipts against one check, in
one $transaction. Deliberately not a persisted batch entity —
checkNumber is already a column and grouping by it answers every
legacy by-check query.
- GET /billing/by-check + a cheque-count report, replacing REPORTE
CHEQUE COUNT / REPORTE POR CHEQUE / EDITA CHEQUE ALF|COUNT|NUM. Print,
PDF, CSV and XLSX come free from the existing /reportes/:slug machinery.
- Web: /estado-cuenta/lote (the Editor screen, with live reconciliation
against the physical check amount), an "Estado de pago" filter, a
"sin fondos" row tag and a Resolver dialog, plus a top-level "Captura"
nav entry.
Integration seam for the OCR auto-capture module (spec §2), which is
required to post through createBatch rather than writing Transaction
rows itself: items[i] maps to lines[i] so postedTransactionId can be
zipped back on; opts.refs[i] stamps captureRef with a duplicate-post
guard that a voided row deliberately does not block; opts.source is
service-level only, so an HTTP client cannot label hand-keyed rows as
machine-captured. captureSource/captureRef are nullable so the 40,136
migrated rows stay NULL rather than being mislabelled.
Fixes two pre-existing bugs found while building this:
- statement() filtered legacySourceTable with `notIn`, which compiles to
SQL NOT IN — and `NULL NOT IN (...)` is NULL, so every app-captured
movement was invisible on the customer statement (438 rows in the
movement browser vs 392 on the statement) while showing everywhere
else. This would have made the whole capture feature look broken.
- The balances count query omitted the void filter its own page query
applied, so the total disagreed with the rows.
Nav highlighting now resolves by longest match; the previous
first-startsWith logic lit up both the parent and any nested entry.
Verified end-to-end against the dev DB, API and browser; all test rows
removed afterwards. Also corrects RESUME.md, which documented the dev
ports as :3001/:3000 — they are :4501/:4500, from the env files.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`r.netPremium` / `r.total` are strings, so an empty-string value leaked
""` into the JSX instead of rendering nothing. Wrap in Boolean() so the
guard is a real conditional.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Replaces ~40 legacy Access renewal-notice report clones (one per carrier
per coverage tier, e.g. AMPL/RC/LIC RENEW X MES/VENCE ATLAS 13/2013) with
one parameterized aviso-renovacion report driven by real Policy/Vehicle/
coveragesJson data instead of hand-typed label text per clone.
- schema.prisma: add RenewalNotice, replacing the legacy CONTROL <ramo>
RENEW[2/3] X MES paper log of which notice generation was sent
- reports: new "letter" ReportFormat + aviso-renovacion registry entry +
LetterLayout renderer in ReportRunner.tsx
- docs/RENEWAL_NOTICES.md + migration/legacy_report_defs/: extracted (via
Application.SaveAsText, since the VBA project wouldn't load) and
documented the legacy report/query chain this replaces
Coveragesjson key names and a mark-as-sent mutation are still unverified/
unbuilt — see caveats in docs/RENEWAL_NOTICES.md.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The --sync path had never been run and was broken in several ways. Fixed and
verified against the dev DB (two consecutive syncs, both exit 0, 32/32
assertions: stable PKs, manual-row preservation, changed-row updates,
legacy-delete, no child duplication, zero FK orphans; idempotent).
- policies/properties: reuse each legacy row's existing id (by provenance)
BEFORE building child rows, so children no longer point at a discarded fresh
uuid; rebuild legacy-owned children via scoped delete + reinsert.
- customers: replace zip(customers, refs) (mispaired almost every row) with a
ref-grouped id remap; names now restore and no spurious customers appear.
- drop the invalid Vehicle @@unique(legacySourceTable, legacyId) — one legacy
policy row carries up to 3 vehicles sharing a legacyId; handle via delete+reinsert.
- upsert lookup tables (policy_types, insurance_providers, type_transactions,
adjusters) by natural name and remap child FKs instead of inserting fresh
uuids that nothing points at.
- transactions: drop updatedAt=NOW() (no such column); guard report formatting
on NULL legacySourceTable (manual rows). Same report guard in bank.
- add manual-safe prune (prune_empty_customers.py --sync, in SYNC_STEPS): prune
only legacy-owned empties, never manually-added customers.
web: customer-detail mini tx list now strikes voided rows with an "(anulado)"
tag (was the last void-UI rendering gap; /estado-cuenta already handled it).
docs: RESUME.md updated — Phase B sync marked verified end-to-end, void-UI
browser pass recorded.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>