With the image fixed, the API got as far as connecting and then died with
Prisma P1001 "can't reach database server". The cause is DNS, not routing.
galactus runs systemd-resolved, whose 127.0.0.53 stub is unreachable from
inside a container, so Docker falls back to the upstream resolver in
/run/systemd/resolve/resolv.conf — the LAN router, which knows nothing about
the tailnet. Verified from a probe container on galactus: resolving
galactus.tail01aa2.ts.net fails outright, while `nc 100.103.77.46 3306` is
OPEN. Only the lookup was broken.
Pin the api and web services to Tailscale's own resolver (100.100.100.100,
the same anycast address on every tailnet) with this tailnet's search suffix.
Both are overridable via TAILSCALE_DNS / TAILNET_SUFFIX. db and minio need
nothing — they make no outbound calls.
Verified end to end: the published image, unmodified, with only these DNS
settings, boots on galactus against the real database and serves
/health {"status":"ok"}
/version {"service":"api","version":"master","gitSha":"3ff56e6b..."}
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Closes the gap between "what tag did I deploy" and "what is actually running",
and gives the schema a history that can be reasoned about across releases.
Migrations
- Baseline the existing schema as 0000_init (migrate diff --from-empty). The
schema had only ever been applied with `prisma db push`, so no history
existed and schema state was disconnected from app version. Existing
databases must be baselined once with `migrate resolve --applied 0000_init`;
the workflows print this remedy on P3005.
- Run `prisma migrate deploy` as a deploy STEP, not the container CMD — as a
CMD, N replicas would race each other applying the same migration.
Version reporting
- GET /version on the API reports the APP_VERSION / GIT_SHA / BUILD_DATE that
build.yml already baked into both images but nothing ever read.
- The web footer shows the web build and flags an api/web mismatch. The two
cannot drift at build time (one matrix run) but can at deploy time.
- Both deploy workflows now fail if the running API does not report the tag
that was dispatched — a stack naming a tag is not proof of what is running.
- scripts/set-version.mjs stamps every package.json, which had all sat at
0.1.0 while real releases shipped as v1.x.
Pre-migrate backup
- deploy/scripts/pre-migrate-backup.mjs dumps the database from INSIDE the
still-running old API container over Portainer's Docker API, so the file
lands in the volume the Operaciones restore screen reads. A dump taken on
the CI runner would be unreachable by the only restore path we have.
Verifies the artefact with `gzip -t` before letting the migration proceed.
galactus
- deploy/galactus/*.compose.yml: standalone-Docker ports of the Swarm stacks.
Plain compose silently ignores `deploy:`, so restart_policy becomes
`restart: unless-stopped` — without it nothing returns after a host reboot.
- .gitea/workflows/deploy-galactus.yml drives endpoint 3 with its own secrets.
Fixes
- deploy.yml passed `endpoint_id` and `pull_image` to
cssnr/portainer-stack-deploy-action, which has no such inputs (they are
`endpoint` and `pull`). The endpoint was silently never set.
docs/DEPLOY_AND_MIGRATIONS.md documents expand/contract as the rule for schema
changes: Prisma has no down-migrations, so a code rollback never rolls the
schema back, and restoring the replication master from a dump diverges every
replica.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>