The Operaciones panel (backup, restore, sync, re-import) shelled out to mysqldump as the application user, parsed straight out of DATABASE_URL. `--single-transaction` issues FLUSH TABLES, which needs the global RELOAD privilege, and the app user is granted only ALL ON jorgecuadros.* plus USAGE ON *.*. BACKUP failed outright; SYNC and REIMPORT failed with it, since both take a safety backup first. An admin credential is now supplied out of band via OPS_DB_ADMIN_USER / OPS_DB_ADMIN_PASSWORD, mirroring what deploy/scripts/pre-migrate-backup.mjs already does, rather than permanently elevating the user the API serves requests as. Host, port and database still come from DATABASE_URL, so the override can only change who logs in, never which server. Unset, it falls back to the DATABASE_URL credentials and warns — local development is unaffected. Two defects in the dumps themselves, both shared with the deploy backup before it was rewritten: - No --set-gtid-purged=OFF. The production server is the replication source with GTID on, so every dump embedded SET @@GLOBAL.GTID_PURGED and was unrestorable onto the server it came from — the one thing the restore screen is for. - The pipeline's exit status was gzip's, and gzip succeeded. A mysqldump that died on its first statement left a small, perfectly valid archive that the job recorded as SUCCESS and the restore screen listed as an ordinary restore point. Dumps now run under `set -o pipefail`, assert a CREATE TABLE count, and delete their own output on failure. Verified with a stubbed mysqldump: a failing dump exits 1, surfaces the real error, removes the partial file, and — critically — stops SYNC/REIMPORT before the ETL touches anything. Restores gained pipefail too: a corrupt archive made gunzip fail while mysql, fed a truncated stream, could still exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
305 lines
14 KiB
Markdown
305 lines
14 KiB
Markdown
# Releasing, deploying, and changing the schema
|
|
|
|
How a version gets from this repo onto a server, and the one rule that keeps
|
|
rollbacks possible.
|
|
|
|
## The short version
|
|
|
|
```bash
|
|
pnpm version:set 1.2.0 # stamp every package.json
|
|
git commit -am "chore(release): v1.2.0"
|
|
git tag v1.2.0 && git push origin master v1.2.0
|
|
```
|
|
|
|
That push triggers `.gitea/workflows/build.yml`, which builds **both** images in
|
|
one matrix run and publishes:
|
|
|
|
| tag pushed | image tags produced |
|
|
| --- | --- |
|
|
| `v1.2.0` | `1.2.0`, `1.2`, `sha-<short>` |
|
|
| push to `master` | `master`, `sha-<short>`, `latest` |
|
|
|
|
Then dispatch a deploy from the Actions tab:
|
|
|
|
- **galactus** (office server, standalone Docker) — *Deploy to galactus*
|
|
- **cubex** (3-node Swarm) — *Deploy to Portainer*
|
|
|
|
> **The `v` is not part of the image tag.** `docker/metadata-action`'s
|
|
> `{{version}}` strips it. Git tag `v1.2.0`, dispatch `1.2.0`. Dispatching
|
|
> `v1.2.0` deploys nothing that exists.
|
|
|
|
Because api and web are built from one matrix run, they cannot drift at build
|
|
time. They *can* drift at deploy time if a stack is applied with only one image
|
|
moved — the web footer shows both versions and flags a mismatch, and the deploy
|
|
workflow's last step fails if the API does not report the tag you dispatched.
|
|
|
|
## What a deploy actually does
|
|
|
|
1. **db + minio** — `scope: full` only. Idempotent; data lives on named volumes.
|
|
2. **Pre-migrate backup** — `deploy/scripts/pre-migrate-backup.mjs` runs
|
|
`mysqldump` *inside the still-running old API container*, via Portainer's
|
|
Docker API. The file lands in that container's `BACKUP_DIR` volume as
|
|
`pre-migrate-<tag>-<timestamp>.sql.gz`, which is exactly what the
|
|
**Operaciones** admin screen lists and can restore. A dump taken on the CI
|
|
runner would be unreachable by the only restore path the platform has.
|
|
3. **`prisma migrate deploy`** — as a workflow *step*, never the container
|
|
`CMD`. If it were the CMD, N replicas would race each other applying the
|
|
same migration.
|
|
4. **app** — the new api + web images.
|
|
5. **Verify** — `GET /version` on the running API must report the dispatched
|
|
tag.
|
|
|
|
Rollback is `tag: 1.1.9` re-dispatched. **That rolls back code only.** The
|
|
schema stays where it is. Which brings us to the rule.
|
|
|
|
## The rule: expand / contract
|
|
|
|
Prisma has no down-migrations. There is no `prisma migrate down`, and there
|
|
never will be. So a schema change that the *previous* release cannot tolerate
|
|
turns a 30-second rollback into a restore-from-backup outage.
|
|
|
|
**Every schema change must leave the previous release working.** Split anything
|
|
destructive across two releases:
|
|
|
|
| | Release N (expand) | Release N+1 (contract) |
|
|
| --- | --- | --- |
|
|
| Rename a column | add the new column, write to both, read the old | drop the old column |
|
|
| Drop a column | stop reading and writing it in code | drop it |
|
|
| Add a required column | add it nullable (or with a default), backfill | make it `NOT NULL` |
|
|
| Split a table | create the new table, dual-write | stop writing the old, drop it |
|
|
| Add an enum value | add the value; old code must not choke on unknowns | start emitting it |
|
|
|
|
Ship N, let it soak, *then* ship N+1. If N has to be rolled back you just
|
|
re-dispatch the old tag — the expanded schema still satisfies it.
|
|
|
|
Restoring from the pre-migrate dump is the **emergency lever, not the routine
|
|
path**, and on galactus it is worse than it sounds: galactus is the replication
|
|
master, DDL replicates through the binlog, and restoring the master from a dump
|
|
diverges every replica. GTIDs will not line up and each replica needs a full
|
|
re-seed. Assume a restore is a multi-hour, whole-topology event.
|
|
|
|
## Migration history
|
|
|
|
`packages/database/prisma/migrations/0000_init/` is a **baseline**. It is the
|
|
full schema as it stood on 2026-07-30, generated with:
|
|
|
|
```bash
|
|
prisma migrate diff --from-empty \
|
|
--to-schema-datamodel packages/database/prisma/schema.prisma --script
|
|
```
|
|
|
|
Until then the schema had only ever been applied with `prisma db push`, so no
|
|
history existed and the schema state was disconnected from the app version.
|
|
|
|
### One-time, on every database that already exists
|
|
|
|
`0000_init` describes tables those databases already have, so `migrate deploy`
|
|
would fail with **P3005 "the database schema is not empty"**. Mark it applied
|
|
instead of applying it — this writes a `_prisma_migrations` row and changes no
|
|
data:
|
|
|
|
```bash
|
|
DATABASE_URL=<the database> npx prisma@5 migrate resolve \
|
|
--applied 0000_init --schema packages/database/prisma/schema.prisma
|
|
```
|
|
|
|
Do this once per database (prod, dev, any local copy). Verify first that the
|
|
live schema really does match the baseline — this should print an empty
|
|
migration:
|
|
|
|
```bash
|
|
prisma migrate diff --from-url "$DATABASE_URL" \
|
|
--to-schema-datamodel packages/database/prisma/schema.prisma --script
|
|
```
|
|
|
|
If it prints actual statements, the live database has drifted from
|
|
`schema.prisma`. Reconcile *before* baselining, or the first real migration
|
|
will fail against a schema Prisma believes it already knows.
|
|
|
|
### From here on
|
|
|
|
```bash
|
|
# edit schema.prisma, then:
|
|
pnpm --filter @jorgecuadros/database exec prisma migrate dev --name add_foo
|
|
```
|
|
|
|
Commit the generated `migrations/<timestamp>_add_foo/` directory. `db push` is
|
|
now a local-scratch tool only — using it against a database with history
|
|
desynchronises it from `_prisma_migrations`.
|
|
|
|
## galactus vs cubex
|
|
|
|
`galactus` is standalone Docker (Portainer endpoint **3**), `cubex` is a 3-node
|
|
Swarm (endpoint **2**). They need different compose files because **plain
|
|
compose silently ignores Swarm's `deploy:` keys** rather than erroring:
|
|
|
|
| | Swarm (`deploy/*.stack.yml`) | standalone (`deploy/galactus/*.compose.yml`) |
|
|
| --- | --- | --- |
|
|
| restart | `deploy.restart_policy` | `restart: unless-stopped` — **without this nothing comes back after a host reboot** |
|
|
| placement | `node.labels.jorgecuadros_db == true` | dropped, one host |
|
|
| ports | `{mode: ingress}` long syntax | `"3306:3306"` |
|
|
| `depends_on` | ignored by Swarm | honoured, with `condition: service_healthy` |
|
|
| volumes | named | named (unchanged — the pinning hazard was a Swarm problem) |
|
|
|
|
Keep the two sets in sync when either changes.
|
|
|
|
On both hosts, cross-stack traffic goes over the **host address**, not compose
|
|
service DNS: db, minio and app are three separate stacks, so three separate
|
|
networks. `DATABASE_URL` and `S3_ENDPOINT` name the host and its published
|
|
port. Do not "simplify" them to `mysql:3306`.
|
|
|
|
### galactus is addressed by MagicDNS, and containers need help resolving it
|
|
|
|
galactus is Tailscale-only once it is installed in the office, so every URL
|
|
names `galactus.tail01aa2.ts.net`. Its LAN IP is a DHCP lease and has already
|
|
drifted once — never put a `192.168.4.x` address in a secret.
|
|
|
|
Containers on galactus cannot resolve that name by default. The host runs
|
|
systemd-resolved, whose `127.0.0.53` stub is unreachable from inside a
|
|
container, so Docker falls back to the upstream resolver listed in
|
|
`/run/systemd/resolve/resolv.conf` — the LAN router, which knows nothing about
|
|
the tailnet. Routing to `100.x` works fine; only the *lookup* fails, and the
|
|
symptom is Prisma **P1001 "can't reach database server"** on a container that
|
|
otherwise started cleanly.
|
|
|
|
`deploy/galactus/jorgecuadros-app.compose.yml` therefore pins the resolver:
|
|
|
|
```yaml
|
|
dns: [100.100.100.100] # Tailscale's fixed anycast MagicDNS address
|
|
dns_search: [tail01aa2.ts.net] # this tailnet's suffix
|
|
```
|
|
|
|
Both are overridable (`TAILSCALE_DNS`, `TAILNET_SUFFIX`) if the tailnet changes.
|
|
Browser-facing origins need none of this — those names resolve on the client.
|
|
|
|
## Replication
|
|
|
|
galactus's MySQL is the **master**; every other MySQL in the estate is a
|
|
replica. Consequences that bite:
|
|
|
|
- `server-id` must be unique across the whole topology (prod `1`, cubex dev
|
|
`11`). A duplicate breaks replication silently.
|
|
- GTID is on from first boot, so replicas attach with `SOURCE_AUTO_POSITION=1`.
|
|
- `binlog_expire_logs_seconds` is raised to 60 days in the galactus compose file
|
|
(`MYSQL_BINLOG_EXPIRE_SECONDS`). MySQL 8.4 defaults to 30 days; a replica
|
|
offline longer than the retention needs a full re-seed.
|
|
|
|
Still open, and **not** handled by anything in this repo:
|
|
|
|
- No replication user with `REPLICATION SLAVE` granted exists yet.
|
|
- Nothing sets `read_only` / `super_read_only` on the replicas, so a stray write
|
|
to a replica will diverge it.
|
|
- The channel to the VPS crosses the public internet. It needs a tunnel or TLS —
|
|
do not publish raw 3306.
|
|
|
|
## Seeding the first sign-in account
|
|
|
|
A freshly migrated database has a schema and **no users**, so nobody can log in.
|
|
`prisma migrate deploy` creates tables, never rows; nothing in the deploy path
|
|
seeds an account, by design — creating an administrator should be a deliberate
|
|
act, not a side effect of shipping code.
|
|
|
|
`apps/api/scripts/seed-user.mjs` ships inside the API image. On the target host:
|
|
|
|
```bash
|
|
docker exec -e SEED_PASSWORD='<a strong password>' \
|
|
<api-container> node apps/api/scripts/seed-user.mjs
|
|
```
|
|
|
|
Defaults are `admin@jorgecuadros.local` / `ChangeMe!2026` / role `ADMIN`,
|
|
overridable with `SEED_EMAIL`, `SEED_PASSWORD`, `SEED_NAME`. **Do not accept the
|
|
default password on anything but a dev database** — it is published in this
|
|
repo's README. The script upserts by email, so re-running is safe, but it also
|
|
**resets the password of an existing account**.
|
|
|
|
## The session cookie and TLS
|
|
|
|
`SESSION_COOKIE_SECURE` controls the `Secure` flag on the session cookie. It
|
|
defaults to on in production, and it must be explicitly `"false"` for a
|
|
deployment served over plain HTTP.
|
|
|
|
This is not cosmetic. express-session silently declines to emit a `Secure`
|
|
cookie over an unencrypted connection: no `Set-Cookie` header is sent at all,
|
|
`POST /auth/login` still answers `200` with the user object, no session is
|
|
established, every subsequent request gets `403`, and the UI bounces back to
|
|
`/login` in a loop. It looks like an auth bug and is really a transport
|
|
mismatch.
|
|
|
|
galactus runs with `SESSION_COOKIE_SECURE=false`, which is acceptable **only**
|
|
because it is reachable exclusively over Tailscale — WireGuard already encrypts
|
|
the wire, so the cookie never crosses an untrusted network. Turn it back on the
|
|
moment the app is served over TLS or reachable off-tailnet. Behind a
|
|
TLS-terminating reverse proxy, set `trust proxy` on the Nest app instead of
|
|
disabling the flag.
|
|
|
|
## The MySQL client inside the API image
|
|
|
|
Alpine's `mysql-client` package is **MariaDB's** client, and it installs an
|
|
empty `/usr/lib/mariadb/plugin`. It therefore cannot speak
|
|
`caching_sha2_password`, which is MySQL 8.4's default and effectively only auth
|
|
method, and every `mysqldump`/`mysql` call from the container fails with:
|
|
|
|
```
|
|
ERROR 1045: Plugin caching_sha2_password could not be loaded:
|
|
... /usr/lib/mariadb/plugin/caching_sha2_password.so: No such file or directory
|
|
```
|
|
|
|
`mariadb-connector-c` supplies that plugin and is installed in
|
|
`docker/api.Dockerfile` for exactly this reason — do not drop it as an unused
|
|
dependency. It affects far more than the deploy backup: the entire
|
|
**Operaciones** panel (backup, restore, sync, re-import) shells out to these
|
|
binaries, so without it none of those work in a container either. The feature
|
|
had only ever been exercised with the API running on a developer machine, where
|
|
the Oracle client is installed, which is why this went unnoticed until the
|
|
first containerised deploy.
|
|
|
|
## The Operaciones panel needs its own database login
|
|
|
|
The panel's four jobs all shell out to `mysqldump`/`mysql`, and they cannot do
|
|
so as the application user. `mysqldump --single-transaction` issues
|
|
`FLUSH TABLES`, which requires the **global** `RELOAD` privilege; the MySQL
|
|
image grants the app user only `ALL PRIVILEGES ON jorgecuadros.*` plus
|
|
`USAGE ON *.*`. `--skip-lock-tables` does not avoid it. BACKUP therefore failed
|
|
outright, and SYNC and RE-IMPORT with it, because both take a safety backup
|
|
first.
|
|
|
|
The API is given an admin login out of band rather than permanently elevating
|
|
the user it serves requests as:
|
|
|
|
```
|
|
OPS_DB_ADMIN_USER=root
|
|
OPS_DB_ADMIN_PASSWORD=<MYSQL_ROOT_PASSWORD>
|
|
```
|
|
|
|
Both deploy workflows pass these into the app stack from the existing
|
|
`MYSQL_ROOT_PASSWORD` secret. Host, port and database still come from
|
|
`DATABASE_URL` — the override changes *who logs in*, never *which server*. With
|
|
the pair unset the service falls back to the `DATABASE_URL` credentials and logs
|
|
a warning, which is what local development wants.
|
|
|
|
Two more things the panel's dumps now do, for the same reasons the pre-migrate
|
|
backup does them (see `deploy/scripts/pre-migrate-backup.mjs`):
|
|
|
|
- **`--set-gtid-purged=OFF`.** galactus is the replication *source* with GTID
|
|
on, so without this every dump embeds `SET @@GLOBAL.GTID_PURGED` and cannot be
|
|
restored onto the server it came from — which is precisely what the restore
|
|
screen exists to do.
|
|
- **`set -o pipefail` and a `CREATE TABLE` count.** `mysqldump | gzip` reports
|
|
gzip's exit status, and a `mysqldump` that dies on its first statement still
|
|
produces a ~372-byte perfectly valid archive that passes `gzip -t`. Without
|
|
both checks a failed backup was recorded as a successful one and listed as an
|
|
ordinary restore point. A dump that fails now deletes its own output.
|
|
|
|
## Known caveats in the deploy path
|
|
|
|
- The pre-migrate backup step sets `NODE_TLS_REJECT_UNAUTHORIZED=0` because
|
|
Portainer serves a self-signed certificate. It is scoped to that one step,
|
|
which talks to nothing but Portainer. Replacing the certificate and dropping
|
|
the flag is the real fix.
|
|
- The runner lives on cubex and must reach the target host's Portainer (9443)
|
|
**and** MySQL (3306). If it cannot reach 3306, run the migration by hand from
|
|
a host that can and dispatch with `skip_migrate: true`.
|
|
- `bootstrap: true` lets the pre-migrate backup be skipped when no API container
|
|
exists yet. Use it for a first-ever deploy only — it is the one switch that
|
|
lets a migration run with no restore point.
|