Files
jorgecuadros-platform/docs/DEPLOY_AND_MIGRATIONS.md
T
rmancinasandClaude Opus 5 b2cdcbe2cd
Build and Push Images / Build jorgecuadros-web (push) Successful in 1m41s
Build and Push Images / Build jorgecuadros-api (push) Successful in 2m5s
fix(api): session cookie never issued over HTTP; ship the seed script
Prod came up with nobody able to log in, in two separate ways.

1. No sign-in account exists. `prisma migrate deploy` creates tables, never
   rows, and nothing in the deploy path seeds one — deliberately, since making
   an administrator should not be a side effect of shipping code. But
   apps/api/scripts was not in the runtime image either, so the only way to
   create the first account was to run the script from a developer machine
   against a production DATABASE_URL. Ship scripts/ in the image so it can be
   run on the host with docker exec. Still never run automatically.

2. Login could not establish a session at all. cookie.secure followed NODE_ENV,
   the image sets NODE_ENV=production, and the app is served over plain HTTP —
   express-session then silently emits NO Set-Cookie header. POST /auth/login
   still answered 200 with the full user object, no session was created, every
   later request 403'd, and the UI would have looped back to /login. It reads
   as an auth bug and is really a transport mismatch.

   The flag is now driven by SESSION_COOKIE_SECURE, still defaulting to
   NODE_ENV. An EMPTY value counts as unset rather than false, because compose
   turns an absent `${SESSION_COOKIE_SECURE:-}` into the empty string and the
   naive check would have quietly dropped Secure on any deployment that merely
   passed the variable through.

   galactus sets it to "false". That is acceptable ONLY because the host is
   reachable exclusively over Tailscale, so WireGuard already encrypts the
   wire. It must go back to "true" when the app is served over TLS or exposed
   off-tailnet; behind a TLS-terminating proxy, set trust proxy instead.

Verified against live prod: seeded an admin, POST /auth/login returns 200 with
full ADMIN abilities, a wrong password is rejected with 401, and no Set-Cookie
was present before this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 15:41:55 -07:00

247 lines
11 KiB
Markdown

# Releasing, deploying, and changing the schema
How a version gets from this repo onto a server, and the one rule that keeps
rollbacks possible.
## The short version
```bash
pnpm version:set 1.2.0 # stamp every package.json
git commit -am "chore(release): v1.2.0"
git tag v1.2.0 && git push origin master v1.2.0
```
That push triggers `.gitea/workflows/build.yml`, which builds **both** images in
one matrix run and publishes:
| tag pushed | image tags produced |
| --- | --- |
| `v1.2.0` | `1.2.0`, `1.2`, `sha-<short>` |
| push to `master` | `master`, `sha-<short>`, `latest` |
Then dispatch a deploy from the Actions tab:
- **galactus** (office server, standalone Docker) — *Deploy to galactus*
- **cubex** (3-node Swarm) — *Deploy to Portainer*
> **The `v` is not part of the image tag.** `docker/metadata-action`'s
> `{{version}}` strips it. Git tag `v1.2.0`, dispatch `1.2.0`. Dispatching
> `v1.2.0` deploys nothing that exists.
Because api and web are built from one matrix run, they cannot drift at build
time. They *can* drift at deploy time if a stack is applied with only one image
moved — the web footer shows both versions and flags a mismatch, and the deploy
workflow's last step fails if the API does not report the tag you dispatched.
## What a deploy actually does
1. **db + minio**`scope: full` only. Idempotent; data lives on named volumes.
2. **Pre-migrate backup**`deploy/scripts/pre-migrate-backup.mjs` runs
`mysqldump` *inside the still-running old API container*, via Portainer's
Docker API. The file lands in that container's `BACKUP_DIR` volume as
`pre-migrate-<tag>-<timestamp>.sql.gz`, which is exactly what the
**Operaciones** admin screen lists and can restore. A dump taken on the CI
runner would be unreachable by the only restore path the platform has.
3. **`prisma migrate deploy`** — as a workflow *step*, never the container
`CMD`. If it were the CMD, N replicas would race each other applying the
same migration.
4. **app** — the new api + web images.
5. **Verify**`GET /version` on the running API must report the dispatched
tag.
Rollback is `tag: 1.1.9` re-dispatched. **That rolls back code only.** The
schema stays where it is. Which brings us to the rule.
## The rule: expand / contract
Prisma has no down-migrations. There is no `prisma migrate down`, and there
never will be. So a schema change that the *previous* release cannot tolerate
turns a 30-second rollback into a restore-from-backup outage.
**Every schema change must leave the previous release working.** Split anything
destructive across two releases:
| | Release N (expand) | Release N+1 (contract) |
| --- | --- | --- |
| Rename a column | add the new column, write to both, read the old | drop the old column |
| Drop a column | stop reading and writing it in code | drop it |
| Add a required column | add it nullable (or with a default), backfill | make it `NOT NULL` |
| Split a table | create the new table, dual-write | stop writing the old, drop it |
| Add an enum value | add the value; old code must not choke on unknowns | start emitting it |
Ship N, let it soak, *then* ship N+1. If N has to be rolled back you just
re-dispatch the old tag — the expanded schema still satisfies it.
Restoring from the pre-migrate dump is the **emergency lever, not the routine
path**, and on galactus it is worse than it sounds: galactus is the replication
master, DDL replicates through the binlog, and restoring the master from a dump
diverges every replica. GTIDs will not line up and each replica needs a full
re-seed. Assume a restore is a multi-hour, whole-topology event.
## Migration history
`packages/database/prisma/migrations/0000_init/` is a **baseline**. It is the
full schema as it stood on 2026-07-30, generated with:
```bash
prisma migrate diff --from-empty \
--to-schema-datamodel packages/database/prisma/schema.prisma --script
```
Until then the schema had only ever been applied with `prisma db push`, so no
history existed and the schema state was disconnected from the app version.
### One-time, on every database that already exists
`0000_init` describes tables those databases already have, so `migrate deploy`
would fail with **P3005 "the database schema is not empty"**. Mark it applied
instead of applying it — this writes a `_prisma_migrations` row and changes no
data:
```bash
DATABASE_URL=<the database> npx prisma@5 migrate resolve \
--applied 0000_init --schema packages/database/prisma/schema.prisma
```
Do this once per database (prod, dev, any local copy). Verify first that the
live schema really does match the baseline — this should print an empty
migration:
```bash
prisma migrate diff --from-url "$DATABASE_URL" \
--to-schema-datamodel packages/database/prisma/schema.prisma --script
```
If it prints actual statements, the live database has drifted from
`schema.prisma`. Reconcile *before* baselining, or the first real migration
will fail against a schema Prisma believes it already knows.
### From here on
```bash
# edit schema.prisma, then:
pnpm --filter @jorgecuadros/database exec prisma migrate dev --name add_foo
```
Commit the generated `migrations/<timestamp>_add_foo/` directory. `db push` is
now a local-scratch tool only — using it against a database with history
desynchronises it from `_prisma_migrations`.
## galactus vs cubex
`galactus` is standalone Docker (Portainer endpoint **3**), `cubex` is a 3-node
Swarm (endpoint **2**). They need different compose files because **plain
compose silently ignores Swarm's `deploy:` keys** rather than erroring:
| | Swarm (`deploy/*.stack.yml`) | standalone (`deploy/galactus/*.compose.yml`) |
| --- | --- | --- |
| restart | `deploy.restart_policy` | `restart: unless-stopped`**without this nothing comes back after a host reboot** |
| placement | `node.labels.jorgecuadros_db == true` | dropped, one host |
| ports | `{mode: ingress}` long syntax | `"3306:3306"` |
| `depends_on` | ignored by Swarm | honoured, with `condition: service_healthy` |
| volumes | named | named (unchanged — the pinning hazard was a Swarm problem) |
Keep the two sets in sync when either changes.
On both hosts, cross-stack traffic goes over the **host address**, not compose
service DNS: db, minio and app are three separate stacks, so three separate
networks. `DATABASE_URL` and `S3_ENDPOINT` name the host and its published
port. Do not "simplify" them to `mysql:3306`.
### galactus is addressed by MagicDNS, and containers need help resolving it
galactus is Tailscale-only once it is installed in the office, so every URL
names `galactus.tail01aa2.ts.net`. Its LAN IP is a DHCP lease and has already
drifted once — never put a `192.168.4.x` address in a secret.
Containers on galactus cannot resolve that name by default. The host runs
systemd-resolved, whose `127.0.0.53` stub is unreachable from inside a
container, so Docker falls back to the upstream resolver listed in
`/run/systemd/resolve/resolv.conf` — the LAN router, which knows nothing about
the tailnet. Routing to `100.x` works fine; only the *lookup* fails, and the
symptom is Prisma **P1001 "can't reach database server"** on a container that
otherwise started cleanly.
`deploy/galactus/jorgecuadros-app.compose.yml` therefore pins the resolver:
```yaml
dns: [100.100.100.100] # Tailscale's fixed anycast MagicDNS address
dns_search: [tail01aa2.ts.net] # this tailnet's suffix
```
Both are overridable (`TAILSCALE_DNS`, `TAILNET_SUFFIX`) if the tailnet changes.
Browser-facing origins need none of this — those names resolve on the client.
## Replication
galactus's MySQL is the **master**; every other MySQL in the estate is a
replica. Consequences that bite:
- `server-id` must be unique across the whole topology (prod `1`, cubex dev
`11`). A duplicate breaks replication silently.
- GTID is on from first boot, so replicas attach with `SOURCE_AUTO_POSITION=1`.
- `binlog_expire_logs_seconds` is raised to 60 days in the galactus compose file
(`MYSQL_BINLOG_EXPIRE_SECONDS`). MySQL 8.4 defaults to 30 days; a replica
offline longer than the retention needs a full re-seed.
Still open, and **not** handled by anything in this repo:
- No replication user with `REPLICATION SLAVE` granted exists yet.
- Nothing sets `read_only` / `super_read_only` on the replicas, so a stray write
to a replica will diverge it.
- The channel to the VPS crosses the public internet. It needs a tunnel or TLS —
do not publish raw 3306.
## Seeding the first sign-in account
A freshly migrated database has a schema and **no users**, so nobody can log in.
`prisma migrate deploy` creates tables, never rows; nothing in the deploy path
seeds an account, by design — creating an administrator should be a deliberate
act, not a side effect of shipping code.
`apps/api/scripts/seed-user.mjs` ships inside the API image. On the target host:
```bash
docker exec -e SEED_PASSWORD='<a strong password>' \
<api-container> node apps/api/scripts/seed-user.mjs
```
Defaults are `admin@jorgecuadros.local` / `ChangeMe!2026` / role `ADMIN`,
overridable with `SEED_EMAIL`, `SEED_PASSWORD`, `SEED_NAME`. **Do not accept the
default password on anything but a dev database** — it is published in this
repo's README. The script upserts by email, so re-running is safe, but it also
**resets the password of an existing account**.
## The session cookie and TLS
`SESSION_COOKIE_SECURE` controls the `Secure` flag on the session cookie. It
defaults to on in production, and it must be explicitly `"false"` for a
deployment served over plain HTTP.
This is not cosmetic. express-session silently declines to emit a `Secure`
cookie over an unencrypted connection: no `Set-Cookie` header is sent at all,
`POST /auth/login` still answers `200` with the user object, no session is
established, every subsequent request gets `403`, and the UI bounces back to
`/login` in a loop. It looks like an auth bug and is really a transport
mismatch.
galactus runs with `SESSION_COOKIE_SECURE=false`, which is acceptable **only**
because it is reachable exclusively over Tailscale — WireGuard already encrypts
the wire, so the cookie never crosses an untrusted network. Turn it back on the
moment the app is served over TLS or reachable off-tailnet. Behind a
TLS-terminating reverse proxy, set `trust proxy` on the Nest app instead of
disabling the flag.
## Known caveats in the deploy path
- The pre-migrate backup step sets `NODE_TLS_REJECT_UNAUTHORIZED=0` because
Portainer serves a self-signed certificate. It is scoped to that one step,
which talks to nothing but Portainer. Replacing the certificate and dropping
the flag is the real fix.
- The runner lives on cubex and must reach the target host's Portainer (9443)
**and** MySQL (3306). If it cannot reach 3306, run the migration by hand from
a host that can and dispatch with `skip_migrate: true`.
- `bootstrap: true` lets the pre-migrate backup be skipped when no API container
exists yet. Use it for a first-ever deploy only — it is the one switch that
lets a migration run with no restore point.