3 Commits
Author SHA1 Message Date
rmancinasandClaude Opus 5 d5ebb86cae fix(deploy): dump from a dedicated container as root, and prove the dump is real
The pre-migrate backup ran INSIDE the API container, which made it depend on
that image's toolchain — and deadlocked: the running image shipped a MySQL
client that could not authenticate, so the backup failed, which blocked the
very deploy that would have replaced the broken image. A backup must not depend
on the thing being deployed.

The dump now runs in a throwaway container built from mysql:8.4 with the API's
backup volume mounted. The volume name is discovered from the API container's
mounts, so the file still lands where the Operaciones restore screen looks. As
a container rather than an exec, its logs can simply be read — no more failures
reported as a bare exit code. The image is pulled if the host lacks it, since a
scope:app deploy never touches the db stack.

Three further defects found while verifying, none of which would have surfaced
without dumping against the real database:

- The dump now runs as root. mysqldump --single-transaction issues FLUSH
  TABLES, needing the global RELOAD privilege; the MySQL image grants the
  application user only ALL ON `<db>`.*, and --skip-lock-tables does not avoid
  it. Elevating the app's own runtime user would have been the worse trade.

- --set-gtid-purged=OFF. galactus is the replication SOURCE with GTID on, so a
  default dump embeds SET @@GLOBAL.GTID_PURGED and is unrestorable onto the
  server it came from. Verified: 0 GTID_PURGED lines in the output.

- Verification was too weak to be worth having. `test -s` plus `gzip -t` passes
  on a 372-byte gzip containing no tables, which is exactly what a dump that
  died on its first statement produces. It now asserts a CREATE TABLE count and
  logs it. A failed attempt also deletes its own output, so a truncated file
  never appears in the restore list.

Verified against live prod, both paths: success writes a 31-table dump the API
container can see; a wrong password fails with mysqldump's own error quoted and
leaves the volume empty.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 16:03:40 -07:00
rmancinasandClaude Opus 5 19f03198d6 fix(docker): install the MySQL 8.4 auth plugin; report why a dump fails
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m56s
Build and Push Images / Build jorgecuadros-api (push) Successful in 3m3s
The pre-migrate backup failed with "mysqldump exited 2" and nothing else.
Reproduced on the host with stderr captured:

  ERROR 1045: Plugin caching_sha2_password could not be loaded:
    /usr/lib/mariadb/plugin/caching_sha2_password.so: No such file or directory

Alpine's `mysql-client` is MariaDB's client and ships an EMPTY plugin
directory, so it cannot perform caching_sha2_password — MySQL 8.4's default and
effectively only auth method. `mariadb-connector-c` provides the plugin.

This was never about the deploy backup alone. Every mysqldump/mysql call from
the API container was broken, which means the whole Operaciones panel — backup,
restore, sync, re-import — could not work in a container. It went unnoticed
because that feature had only ever been run with the API on a developer
machine, where the Oracle client is installed. Verified after the fix: dump
exits 0, gzip valid, 31 CREATE TABLEs.

Also fixed, both found while chasing the above:

- The backup script reported an exit code and nothing else, because a detached
  exec captures no output — which is precisely why this needed a manual
  reproduction. mysqldump's stderr is now redirected to a file and read back
  through a short attached exec on failure, so the deploy log states the cause.
  Verified against live prod: the log now carries the 1045 line itself.

- Listing ONLY 100.100.100.100 as the containers' resolver costs them public
  DNS, since MagicDNS does not forward upstream unless the tailnet defines
  global nameservers. Nothing at runtime needed it, but `apk` inside the
  container stopped resolving, and anything outbound would have too. A public
  fallback resolver is now listed after MagicDNS.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 15:52:03 -07:00
rmancinasandClaude Opus 5 4ee7ec71f0 feat(deploy): prisma migration history, /version, galactus standalone deploy
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m49s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m2s
Closes the gap between "what tag did I deploy" and "what is actually running",
and gives the schema a history that can be reasoned about across releases.

Migrations
- Baseline the existing schema as 0000_init (migrate diff --from-empty). The
  schema had only ever been applied with `prisma db push`, so no history
  existed and schema state was disconnected from app version. Existing
  databases must be baselined once with `migrate resolve --applied 0000_init`;
  the workflows print this remedy on P3005.
- Run `prisma migrate deploy` as a deploy STEP, not the container CMD — as a
  CMD, N replicas would race each other applying the same migration.

Version reporting
- GET /version on the API reports the APP_VERSION / GIT_SHA / BUILD_DATE that
  build.yml already baked into both images but nothing ever read.
- The web footer shows the web build and flags an api/web mismatch. The two
  cannot drift at build time (one matrix run) but can at deploy time.
- Both deploy workflows now fail if the running API does not report the tag
  that was dispatched — a stack naming a tag is not proof of what is running.
- scripts/set-version.mjs stamps every package.json, which had all sat at
  0.1.0 while real releases shipped as v1.x.

Pre-migrate backup
- deploy/scripts/pre-migrate-backup.mjs dumps the database from INSIDE the
  still-running old API container over Portainer's Docker API, so the file
  lands in the volume the Operaciones restore screen reads. A dump taken on
  the CI runner would be unreachable by the only restore path we have.
  Verifies the artefact with `gzip -t` before letting the migration proceed.

galactus
- deploy/galactus/*.compose.yml: standalone-Docker ports of the Swarm stacks.
  Plain compose silently ignores `deploy:`, so restart_policy becomes
  `restart: unless-stopped` — without it nothing returns after a host reboot.
- .gitea/workflows/deploy-galactus.yml drives endpoint 3 with its own secrets.

Fixes
- deploy.yml passed `endpoint_id` and `pull_image` to
  cssnr/portainer-stack-deploy-action, which has no such inputs (they are
  `endpoint` and `pull`). The endpoint was silently never set.

docs/DEPLOY_AND_MIGRATIONS.md documents expand/contract as the rule for schema
changes: Prisma has no down-migrations, so a code rollback never rolls the
schema back, and restoring the replication master from a dump diverges every
replica.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 11:41:12 -07:00