You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
.github/workflows/deploy.yml runs on push to main and on pull_request. Nothing in it reacts
to its own failure. When a push-to-main run goes red, the only trace is a red check on the commit —
no issue, no ping. The web deploy job failed on 8 consecutive main pushes over 3 days (first red
commit 1e52c3c0, PR #8934, through 0cb9588f) and was found only because the founder asked why
main was red. The specific red is #9507; this issue is about the silence, not that bug.
The PR gates structurally cannot cover this. A preview stage has no custom domain, so the failing
step passed on every PR and failed only on the post-merge prod run. The same blind spot covers
anything else that only runs on push to main.
Triage note: I read deploy.yml at origin/main. Its jobs are changes, no-preview, app-roster and deploy; there is no failure-handling job and no if: failure() anywhere in the
file, so the gap is live.
The work — the founder's final ruled shape, 2026-09-20
This section is the one to build against. It supersedes every earlier ruling comment on this
issue, including the ones naming packages/ci-cli and "a registered verb under ci that
installs". Those comments stay on the record unedited; their homes are no longer the ones to build.
It is a generic main-went-red alarm, not a deploy tracker. Nothing in the decision is about
deploys: a push-to-default-branch run went red; open one tracking issue or comment on the open one;
close it on the next green. Name it accordingly — main-alarm, with the bin at packages/fabrika-cli/src/ci/main-alarm-bin.ts. PR #9522 currently calls it deploy-tracker; that
name goes.
Home is packages/fabrika-cli. No new package.
It is a no-install bare bin, deliberately. Follow the existing packages/fabrika-cli/src/ci/required-bin.ts precedent: the alarm job runs checkout + setup-node + node <bin> with no pnpm install, so the alarm still fires when the install is what broke.
Its import graph carries no effect and no installed dependency — one relative plain-TS import
chain is the whole graph.
The decision is a pure, unit-tested core. Given the run's conclusion and the currently open
alarm issue, if any: open / comment / close / nothing. Five cases, unchanged from the first ruling:
first red opens one issue; a further red comments on the open one; green with an open issue closes
it with a link to the green run; green with none does nothing; pull_request runs never fire.
Settings come from .fabrika.jsonc as a new sub-key under the existing ci key group: ci.mainAlarm, with mention (a list of handles) and label. Follow .patterns/fabrika-config-key-groups.md — one key module,
one registry line, a jsonSchema fragment, the four resolution arms, Unknown refuses and never
defaults. The bare bin reads the config through the existing zero-dependency parser
(packages/fabrika-cli/src/config/document.ts over packages/fabrika-cli/src/io/json.ts), never
the Effect file opener — that path would put effect back on the bin's graph.
Shipped defaults: label fabrika-main-alarm, mention empty. This repo declares "mention": ["@usirin"] in its own .fabrika.jsonc. No handle, repo name or workflow name is
hardcoded in the package; portability-guard is binding here.
One alarm issue per workflow. The open alarm is found again by that label plus a stable marker,
so deploy.yml and release-please.yml get separate issues. New alarm issues land carrying status:needs-triage, so normal intake picks them up.
Run facts come from the environment, not config. The workflow name, run URL, head commit and job
result arrive from the GitHub-provided env and a --result argument.
Wiring. A job in .github/workflows/deploy.yml on push-to-main only, if: always(), with issues: write scoped to that job alone while the workflow-level block stays contents: read. Also
replace the hand-rolled alarm job in .github/workflows/release-please.yml (lines ~388–441) with
the same bin: it is the known duplicate of this logic, it is a delete-and-call with no second
decision core, and leaving it would ship two alarms with different dedup rules.
Triage note: I judged that swap part of this unit rather than a follow-up. It adds no logic — the
existing job already runs without an install, which is exactly the shape the bin has — and the
per-workflow issue keying above is what makes the two coexist.
Docblocks under packages/fabrika-cli cite their ruling with an @ruling tag naming issue #9508 — never an ADR number and never a .decisions/ path; portability-guard reds both
(ADR 0394).
Spike residue. The proof runs on PR #9522 created a deploy-failure-tracker label on this
repository. It must not survive under that name.
Widening this issue to any of the three is a scope violation.
Pointers
packages/fabrika-cli/src/ci/required-bin.ts — the no-install bare-bin precedent, and its
docblock stating why it is not routed through fabrika's bin
packages/fabrika-cli/src/config/keys/ci.ts — the ci key group this adds a sub-key to; registry.ts, json-schema.ts, and fabrika config schema --write for the rest
packages/fabrika-cli/src/config/document.ts, packages/fabrika-cli/src/io/json.ts — the
zero-dependency read path
.github/workflows/deploy.yml — on: block, workflow-level permissions: contents: read, the deploy job's own permissions block
.github/workflows/release-please.yml — the hand-rolled alarm job (~388–441) this replaces
Triage note: I searched the board for an open epic or issue owning the post-merge-notification or
main-alarm surface. Nothing absorbed it — the nearest live hits fix one red in the same file
(#9507, merged) or own an adjacent CI-tooling question (#9516). Minted standalone in the standing
pipeline lane.
Engineering-led platform work under ADR 0078;
no product ruling is pending.
Acceptance criteria
The alarm's decision core and its bin live under packages/fabrika-cli/src/ci/, with the bin
at main-alarm-bin.ts. The diff creates no new workspace package, and no file, verb, label
default or identifier in it is named deploy-tracker or deploy-failure-tracker.
The bin is a no-install bare bin on the required-bin.ts precedent: its transitive import
graph reaches no effect import and no installed dependency, and the alarm jobs run
checkout + setup-node + node <bin> with no pnpm install step.
The open/comment/close/nothing decision is a pure function with unit tests covering all five
cases: first red opens one issue; a further red comments on the open one; green with an open
issue closes it with a link to the green run; green with no open issue does nothing;
a pull_request run never fires.
.fabrika.jsonc gains a ci.mainAlarm sub-key under the existing ci key group carrying mention (list of handles) and label, built per .patterns/fabrika-config-key-groups.md:
one key module, one register(...) line, a jsonSchema fragment, and the four resolution
arms with Unknown refusing rather than defaulting. The committed .fabrika.schema.json is
regenerated with fabrika config schema --write.
The bin reads that config through packages/fabrika-cli/src/config/document.ts and packages/fabrika-cli/src/io/json.ts, not through the Effect file opener.
The shipped defaults are label fabrika-main-alarm and an empty mention list, and this
repository declares "mention": ["@usirin"] in its own .fabrika.jsonc.
No handle, repo name, org name or workflow name is hardcoded anywhere in the package; fabrika guard portability-guard check is green on the diff.
The open alarm issue is matched by the configured label plus a stable marker — never by title
text — and the marker keys one issue per workflow, so deploy.yml and release-please.yml cannot collide on the same tracker.
A newly opened alarm issue carries the status:needs-triage label, names the failed workflow
and job, links the run URL, and names the head commit sha.
The workflow name, run URL, head commit and job result reach the bin from the GitHub-provided
environment and a --result argument. None of the four is read from .fabrika.jsonc.
.github/workflows/deploy.yml gains an alarm job that runs on push-to-main only, under if: always(), whose permissions: block grants issues: write scoped to that job while the
workflow-level permissions: stays contents: read. No new secret and no new third-party
action.
The hand-rolled alarm job in .github/workflows/release-please.yml is replaced by a job
calling the same bin, and no inline shell or github-script body deciding
open-vs-comment-vs-close survives in either workflow.
Nothing in the diff adds an outside uptime check on /api/health, and nothing pauses or gates
the merge queue. Both are ruled out of scope.
The changed workflows parse as YAML, and the alarm's trigger behaviour is confirmed against a
real run of the workflow after merge, or against a deliberately-failed run on a branch.
[evidence: a workflow run linked in the PR description]
deploy.yml's alarm job reports a cancelled run as cancelled, not as success. Its RUN_RESULT expression currently maps a need whose result is cancelled to 'success' when nothing else failed or skipped, so main-alarm.ts's cancelled arm is unreachable from that workflow and a cancelled deploy closes the standing alarm — the false all-clear that arm exists to prevent. release-please.yml already passes needs.release-please.result straight through and is correct; the two callers must agree.
Original report (verbatim)
Summary
When a workflow that only runs after merge goes red on main, nothing tells anyone. The web deploy job failed on 8 main pushes in a row over 3 days and was found only because the founder happened to ask why main was red.
What I was doing
Tracing why the main head showed a red check, which led to the deploy health gate failure now tracked in #9507.
What I observed
deploy.yml runs on push to main. Its job "deploy (web, @kampus/web, true)" failed on every main push from commit 1e52c3c (PR #8934, 2026-09-17 19:01Z) through 0cb9588 (2026-09-20 19:20Z), 8 of 8 runs checked. In that window no notification went out, no issue was filed, and merges kept landing on top of the red.
The PR gates could not have caught it. On PR #8934 the preview deploy passed, because a preview stage has no custom domain and the deploy log still printed a workers.dev URL. Prod has the custom domain and prints only that, so the failure exists only on the post-merge prod run.
The same holds for anything else that runs only on push to main, such as the prod health gate and release-please. I did not audit the full list.
Why it matters
A real prod break would look the same as a stale red and could sit just as long. The health gate from #1436 and ADR 0092 fails closed, but a closed gate nobody sees protects little. Merges continuing over a red prod deploy also make the first bad commit harder to find later. How often this happens is unknown; this is the one case I measured.
Pointers
.github/workflows/deploy.yml (push to main trigger, the deploy and verify steps)
On a failed push-to-main run of deploy.yml, open or update one issue, deduped so 8 reds make 1 issue, and ping the founder. Close it on the next green. Worth deciding separately whether a red prod deploy should pause the merge queue.
Filed by an agent · session 56fedc1c-913a-4bc4-baa5-2b748ea48b9b · branch main · 2026-09-20T19:34:32Z
What is wrong
.github/workflows/deploy.ymlruns onpushtomainand onpull_request. Nothing in it reactsto its own failure. When a push-to-main run goes red, the only trace is a red check on the commit —
no issue, no ping. The web deploy job failed on 8 consecutive main pushes over 3 days (first red
commit
1e52c3c0, PR #8934, through0cb9588f) and was found only because the founder asked whymain was red. The specific red is #9507; this issue is about the silence, not that bug.
The PR gates structurally cannot cover this. A preview stage has no custom domain, so the failing
step passed on every PR and failed only on the post-merge prod run. The same blind spot covers
anything else that only runs on push to main.
Triage note: I read
deploy.ymlatorigin/main. Its jobs arechanges,no-preview,app-rosteranddeploy; there is no failure-handling job and noif: failure()anywhere in thefile, so the gap is live.
The work — the founder's final ruled shape, 2026-09-20
This section is the one to build against. It supersedes every earlier ruling comment on this
issue, including the ones naming
packages/ci-cliand "a registered verb undercithatinstalls". Those comments stay on the record unedited; their homes are no longer the ones to build.
It is a generic main-went-red alarm, not a deploy tracker. Nothing in the decision is about
deploys: a push-to-default-branch run went red; open one tracking issue or comment on the open one;
close it on the next green. Name it accordingly —
main-alarm, with the bin atpackages/fabrika-cli/src/ci/main-alarm-bin.ts. PR #9522 currently calls itdeploy-tracker; thatname goes.
Home is
packages/fabrika-cli. No new package.It is a no-install bare bin, deliberately. Follow the existing
packages/fabrika-cli/src/ci/required-bin.tsprecedent: the alarm job runs checkout + setup-node +node <bin>with nopnpm install, so the alarm still fires when the install is what broke.Its import graph carries no
effectand no installed dependency — one relative plain-TS importchain is the whole graph.
The decision is a pure, unit-tested core. Given the run's conclusion and the currently open
alarm issue, if any: open / comment / close / nothing. Five cases, unchanged from the first ruling:
first red opens one issue; a further red comments on the open one; green with an open issue closes
it with a link to the green run; green with none does nothing;
pull_requestruns never fire.Settings come from
.fabrika.jsoncas a new sub-key under the existingcikey group:ci.mainAlarm, withmention(a list of handles) andlabel. Follow.patterns/fabrika-config-key-groups.md— one key module,one registry line, a
jsonSchemafragment, the four resolution arms,Unknownrefuses and neverdefaults. The bare bin reads the config through the existing zero-dependency parser
(
packages/fabrika-cli/src/config/document.tsoverpackages/fabrika-cli/src/io/json.ts), neverthe Effect file opener — that path would put
effectback on the bin's graph.Shipped defaults: label
fabrika-main-alarm, mention empty. This repo declares"mention": ["@usirin"]in its own.fabrika.jsonc. No handle, repo name or workflow name ishardcoded in the package;
portability-guardis binding here.One alarm issue per workflow. The open alarm is found again by that label plus a stable marker,
so
deploy.ymlandrelease-please.ymlget separate issues. New alarm issues land carryingstatus:needs-triage, so normal intake picks them up.Run facts come from the environment, not config. The workflow name, run URL, head commit and job
result arrive from the GitHub-provided env and a
--resultargument.Wiring. A job in
.github/workflows/deploy.ymlon push-to-main only,if: always(), withissues: writescoped to that job alone while the workflow-level block stayscontents: read. Alsoreplace the hand-rolled alarm job in
.github/workflows/release-please.yml(lines ~388–441) withthe same bin: it is the known duplicate of this logic, it is a delete-and-call with no second
decision core, and leaving it would ship two alarms with different dedup rules.
Triage note: I judged that swap part of this unit rather than a follow-up. It adds no logic — the
existing job already runs without an install, which is exactly the shape the bin has — and the
per-workflow issue keying above is what makes the two coexist.
Docblocks under
packages/fabrika-clicite their ruling with an@rulingtag naming issue#9508 — never an ADR number and never a
.decisions/path;portability-guardreds both(ADR 0394).
Spike residue. The proof runs on PR #9522 created a
deploy-failure-trackerlabel on thisrepository. It must not survive under that name.
Explicitly out of scope, unchanged
/api/health— tracked separately as Nothing probes prod /api/health outside a deploy run, so a prod outage between pushes is unseen #9512.have frozen merges for 3 days. Do not build it, do not add a hook for it.
packages/fabrika-cli/src/ci/main-alarm-bin.tsto pointnodeat. Out of scope here; filed as afollow-up as A repo consuming released fabrika cannot call its no-install bare bins at all #9539.
Widening this issue to any of the three is a scope violation.
Pointers
packages/fabrika-cli/src/ci/required-bin.ts— the no-install bare-bin precedent, and itsdocblock stating why it is not routed through
fabrika's binpackages/fabrika-cli/src/config/keys/ci.ts— thecikey group this adds a sub-key to;registry.ts,json-schema.ts, andfabrika config schema --writefor the restpackages/fabrika-cli/src/config/document.ts,packages/fabrika-cli/src/io/json.ts— thezero-dependency read path
.github/workflows/deploy.yml—on:block, workflow-levelpermissions: contents: read, thedeployjob's ownpermissionsblock.github/workflows/release-please.yml— the hand-rolledalarmjob (~388–441) this replacestoken-env follow-up), PR feat: A red post-merge run on main reaches nobody: 8 failed web deploys sat unseen for 3 days #9522 (the lane in flight)
Scope note
Triage note: I searched the board for an open epic or issue owning the post-merge-notification or
main-alarm surface. Nothing absorbed it — the nearest live hits fix one red in the same file
(#9507, merged) or own an adjacent CI-tooling question (#9516). Minted standalone in the standing
pipeline lane.
Engineering-led platform work under ADR 0078;
no product ruling is pending.
Acceptance criteria
packages/fabrika-cli/src/ci/, with the binat
main-alarm-bin.ts. The diff creates no new workspace package, and no file, verb, labeldefault or identifier in it is named
deploy-trackerordeploy-failure-tracker.required-bin.tsprecedent: its transitive importgraph reaches no
effectimport and no installed dependency, and the alarm jobs runcheckout + setup-node +
node <bin>with nopnpm installstep.cases: first red opens one issue; a further red comments on the open one; green with an open
issue closes it with a link to the green run; green with no open issue does nothing;
a
pull_requestrun never fires..fabrika.jsoncgains aci.mainAlarmsub-key under the existingcikey group carryingmention(list of handles) andlabel, built per.patterns/fabrika-config-key-groups.md:one key module, one
register(...)line, ajsonSchemafragment, and the four resolutionarms with
Unknownrefusing rather than defaulting. The committed.fabrika.schema.jsonisregenerated with
fabrika config schema --write.packages/fabrika-cli/src/config/document.tsandpackages/fabrika-cli/src/io/json.ts, not through the Effect file opener.fabrika-main-alarmand an emptymentionlist, and thisrepository declares
"mention": ["@usirin"]in its own.fabrika.jsonc.fabrika guard portability-guard checkis green on the diff.text — and the marker keys one issue per workflow, so
deploy.ymlandrelease-please.ymlcannot collide on the same tracker.status:needs-triagelabel, names the failed workflowand job, links the run URL, and names the head commit sha.
environment and a
--resultargument. None of the four is read from.fabrika.jsonc..github/workflows/deploy.ymlgains an alarm job that runs on push-to-main only, underif: always(), whosepermissions:block grantsissues: writescoped to that job while theworkflow-level
permissions:stayscontents: read. No new secret and no new third-partyaction.
alarmjob in.github/workflows/release-please.ymlis replaced by a jobcalling the same bin, and no inline shell or
github-scriptbody decidingopen-vs-comment-vs-close survives in either workflow.
packages/fabrika-cli/cites its ruling with an@rulingtag naming issue A red post-merge run on main reaches nobody: 8 failed web deploys sat unseen for 3 days #9508, and none of them names an ADR number or a.decisions/path.
deploy-failure-trackerlabel left on this repository by the PR feat: A red post-merge run on main reaches nobody: 8 failed web deploys sat unseen for 3 days #9522 proof runs nolonger exists under that name. [evidence: a
gh label listread linked or quoted in the PR description]/api/health, and nothing pauses or gatesthe merge queue. Both are ruled out of scope.
real run of the workflow after merge, or against a deliberately-failed run on a branch.
[evidence: a workflow run linked in the PR description]
deploy.yml's alarm job reports a cancelled run ascancelled, not assuccess. ItsRUN_RESULTexpression currently maps a need whose result iscancelledto'success'when nothing else failed or skipped, somain-alarm.ts'scancelledarm is unreachable from that workflow and a cancelleddeploycloses the standing alarm — the false all-clear that arm exists to prevent.release-please.ymlalready passesneeds.release-please.resultstraight through and is correct; the two callers must agree.Original report (verbatim)
Summary
When a workflow that only runs after merge goes red on main, nothing tells anyone. The web deploy job failed on 8 main pushes in a row over 3 days and was found only because the founder happened to ask why main was red.
What I was doing
Tracing why the main head showed a red check, which led to the deploy health gate failure now tracked in #9507.
What I observed
deploy.ymlruns on push to main. Its job "deploy (web, @kampus/web, true)" failed on every main push from commit 1e52c3c (PR #8934, 2026-09-17 19:01Z) through 0cb9588 (2026-09-20 19:20Z), 8 of 8 runs checked. In that window no notification went out, no issue was filed, and merges kept landing on top of the red.The PR gates could not have caught it. On PR #8934 the preview deploy passed, because a preview stage has no custom domain and the deploy log still printed a workers.dev URL. Prod has the custom domain and prints only that, so the failure exists only on the post-merge prod run.
The same holds for anything else that runs only on push to main, such as the prod health gate and release-please. I did not audit the full list.
Why it matters
A real prod break would look the same as a stale red and could sit just as long. The health gate from #1436 and ADR 0092 fails closed, but a closed gate nobody sees protects little. Merges continuing over a red prod deploy also make the first bad commit harder to find later. How often this happens is unknown; this is the one case I measured.
Pointers
.github/workflows/deploy.yml(push to main trigger, the deploy and verify steps)Suggested next step (non-binding)
On a failed push-to-main run of
deploy.yml, open or update one issue, deduped so 8 reds make 1 issue, and ping the founder. Close it on the next green. Worth deciding separately whether a red prod deploy should pause the merge queue.Filed by an agent · session
56fedc1c-913a-4bc4-baa5-2b748ea48b9b· branchmain· 2026-09-20T19:34:32Z