fix(ops): fail orphaned RUNNING jobs at startup
Build and Push Images / Build jorgecuadros-api (push) Canceled after 1m21s
Build and Push Images / Build jorgecuadros-web (push) Canceled after 1m21s

Ops jobs run as a child of the API process, so no job can outlive it. When a
deploy landed 110 seconds into a REIMPORT, the child died and nothing was
left to finalize the row — it stayed RUNNING forever. Because startJob()
refuses to start while any RUNNING row exists, that one interrupted job
wedged the panel permanently with no way out from the UI; recovering it took
a manual UPDATE against the production database.

A fresh boot is proof that nothing survived, so this is unconditional rather
than filtered on age: "started recently" does not imply "still alive" here.

Rows are updated one at a time rather than with updateMany so the reason can
be APPENDED to the log. A job whose log simply stops mid-step with no
explanation is what made the first occurrence hard to diagnose.

Failure to reconcile is logged and swallowed: a wedged panel is bad, an API
that will not boot is worse.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-03 17:22:15 -07:00
co-authored by Claude Opus 5
parent dac1f1982f
commit 7797c45e9f
+49
View File
@@ -67,6 +67,55 @@ export class OpsService implements OnModuleInit {
async onModuleInit(): Promise<void> {
await fs.mkdir(this.ingestDir, { recursive: true });
await fs.mkdir(this.backupDir, { recursive: true });
await this.reconcileOrphanedJobs();
}
/**
* Fail any job still marked RUNNING at startup.
*
* Jobs run as a child of THIS process, so no job can outlive it: if a row says
* RUNNING while we are booting, its process died with the previous instance
* and nothing will ever finalize it. Since startJob() refuses to start while
* any RUNNING row exists, one interrupted job wedges the panel permanently
* with no way out from the UI — it took a manual UPDATE against production to
* recover the first time this happened, when a deploy landed 110 seconds into
* a REIMPORT.
*
* Deliberately unconditional rather than filtered on age: "started recently"
* does not mean "still alive" here, and a fresh boot is proof enough that
* nothing survived.
*/
private async reconcileOrphanedJobs(): Promise<void> {
try {
// Read then write one by one rather than updateMany: the log needs the
// reason APPENDED, and a job whose log just stops mid-step with no
// explanation is what made the first occurrence hard to diagnose.
const orphans = await this.prisma.opsJob.findMany({
where: { status: "RUNNING" },
select: { id: true, kind: true, log: true },
});
for (const job of orphans) {
await this.prisma.opsJob.update({
where: { id: job.id },
data: {
status: "FAILED",
finishedAt: new Date(),
log: {
set:
job.log +
"\n[interrumpido: el contenedor se reinició mientras el trabajo corría; " +
"el proceso hijo no sobrevive a un redespliegue. " +
"Vuelva a ejecutar la operación desde el principio.]\n",
},
},
});
this.logger.warn(`trabajo ${job.kind} ${job.id} quedó huérfano; marcado FAILED`);
}
} catch (e) {
// Never block startup on this. A failed reconcile leaves the panel
// wedged, which is bad, but an API that will not boot is worse.
this.logger.error(`no se pudieron reconciliar trabajos huérfanos: ${String(e)}`);
}
}
/* -------------------------------------------------------------- ingest */