Skip to content

A red post-merge run on main reaches nobody: 8 failed web deploys sat unseen for 3 days #9508

Description

@usirin

What is wrong

.github/workflows/deploy.yml runs on push to main and on pull_request. Nothing in it reacts
to its own failure. When a push-to-main run goes red, the only trace is a red check on the commit —
no issue, no ping. The web deploy job failed on 8 consecutive main pushes over 3 days (first red
commit 1e52c3c0, PR #8934, through 0cb9588f) and was found only because the founder asked why
main was red. The specific red is #9507; this issue is about the silence, not that bug.

The PR gates structurally cannot cover this. A preview stage has no custom domain, so the failing
step passed on every PR and failed only on the post-merge prod run. The same blind spot covers
anything else that only runs on push to main.

Triage note: I read deploy.yml at origin/main. Its jobs are changes, no-preview,
app-roster and deploy; there is no failure-handling job and no if: failure() anywhere in the
file, so the gap is live.

The work — the founder's final ruled shape, 2026-09-20

This section is the one to build against. It supersedes every earlier ruling comment on this
issue
, including the ones naming packages/ci-cli and "a registered verb under ci that
installs". Those comments stay on the record unedited; their homes are no longer the ones to build.

It is a generic main-went-red alarm, not a deploy tracker. Nothing in the decision is about
deploys: a push-to-default-branch run went red; open one tracking issue or comment on the open one;
close it on the next green.
Name it accordingly — main-alarm, with the bin at
packages/fabrika-cli/src/ci/main-alarm-bin.ts. PR #9522 currently calls it deploy-tracker; that
name goes.

Home is packages/fabrika-cli. No new package.

It is a no-install bare bin, deliberately. Follow the existing
packages/fabrika-cli/src/ci/required-bin.ts precedent: the alarm job runs checkout + setup-node +
node <bin> with no pnpm install, so the alarm still fires when the install is what broke.
Its import graph carries no effect and no installed dependency — one relative plain-TS import
chain is the whole graph.

The decision is a pure, unit-tested core. Given the run's conclusion and the currently open
alarm issue, if any: open / comment / close / nothing. Five cases, unchanged from the first ruling:
first red opens one issue; a further red comments on the open one; green with an open issue closes
it with a link to the green run; green with none does nothing; pull_request runs never fire.

Settings come from .fabrika.jsonc as a new sub-key under the existing ci key group:
ci.mainAlarm, with mention (a list of handles) and label. Follow
.patterns/fabrika-config-key-groups.md — one key module,
one registry line, a jsonSchema fragment, the four resolution arms, Unknown refuses and never
defaults. The bare bin reads the config through the existing zero-dependency parser
(packages/fabrika-cli/src/config/document.ts over packages/fabrika-cli/src/io/json.ts), never
the Effect file opener — that path would put effect back on the bin's graph.

Shipped defaults: label fabrika-main-alarm, mention empty. This repo declares
"mention": ["@usirin"] in its own .fabrika.jsonc. No handle, repo name or workflow name is
hardcoded in the package; portability-guard is binding here.

One alarm issue per workflow. The open alarm is found again by that label plus a stable marker,
so deploy.yml and release-please.yml get separate issues. New alarm issues land carrying
status:needs-triage, so normal intake picks them up.

Run facts come from the environment, not config. The workflow name, run URL, head commit and job
result arrive from the GitHub-provided env and a --result argument.

Wiring. A job in .github/workflows/deploy.yml on push-to-main only, if: always(), with
issues: write scoped to that job alone while the workflow-level block stays contents: read. Also
replace the hand-rolled alarm job in .github/workflows/release-please.yml (lines ~388–441) with
the same bin: it is the known duplicate of this logic, it is a delete-and-call with no second
decision core, and leaving it would ship two alarms with different dedup rules.

Triage note: I judged that swap part of this unit rather than a follow-up. It adds no logic — the
existing job already runs without an install, which is exactly the shape the bin has — and the
per-workflow issue keying above is what makes the two coexist.

Docblocks under packages/fabrika-cli cite their ruling with an @ruling tag naming issue
#9508 — never an ADR number and never a .decisions/ path; portability-guard reds both
(ADR 0394).

Spike residue. The proof runs on PR #9522 created a deploy-failure-tracker label on this
repository. It must not survive under that name.

Explicitly out of scope, unchanged

Widening this issue to any of the three is a scope violation.

Pointers

Scope note

Triage note: I searched the board for an open epic or issue owning the post-merge-notification or
main-alarm surface. Nothing absorbed it — the nearest live hits fix one red in the same file
(#9507, merged) or own an adjacent CI-tooling question (#9516). Minted standalone in the standing
pipeline lane.

Engineering-led platform work under ADR 0078;
no product ruling is pending.

Acceptance criteria

  • The alarm's decision core and its bin live under packages/fabrika-cli/src/ci/, with the bin
    at main-alarm-bin.ts. The diff creates no new workspace package, and no file, verb, label
    default or identifier in it is named deploy-tracker or deploy-failure-tracker.
  • The bin is a no-install bare bin on the required-bin.ts precedent: its transitive import
    graph reaches no effect import and no installed dependency, and the alarm jobs run
    checkout + setup-node + node <bin> with no pnpm install step.
  • The open/comment/close/nothing decision is a pure function with unit tests covering all five
    cases: first red opens one issue; a further red comments on the open one; green with an open
    issue closes it with a link to the green run; green with no open issue does nothing;
    a pull_request run never fires.
  • .fabrika.jsonc gains a ci.mainAlarm sub-key under the existing ci key group carrying
    mention (list of handles) and label, built per .patterns/fabrika-config-key-groups.md:
    one key module, one register(...) line, a jsonSchema fragment, and the four resolution
    arms with Unknown refusing rather than defaulting. The committed .fabrika.schema.json is
    regenerated with fabrika config schema --write.
  • The bin reads that config through packages/fabrika-cli/src/config/document.ts and
    packages/fabrika-cli/src/io/json.ts, not through the Effect file opener.
  • The shipped defaults are label fabrika-main-alarm and an empty mention list, and this
    repository declares "mention": ["@usirin"] in its own .fabrika.jsonc.
  • No handle, repo name, org name or workflow name is hardcoded anywhere in the package;
    fabrika guard portability-guard check is green on the diff.
  • The open alarm issue is matched by the configured label plus a stable marker — never by title
    text — and the marker keys one issue per workflow, so deploy.yml and
    release-please.yml cannot collide on the same tracker.
  • A newly opened alarm issue carries the status:needs-triage label, names the failed workflow
    and job, links the run URL, and names the head commit sha.
  • The workflow name, run URL, head commit and job result reach the bin from the GitHub-provided
    environment and a --result argument. None of the four is read from .fabrika.jsonc.
  • .github/workflows/deploy.yml gains an alarm job that runs on push-to-main only, under
    if: always(), whose permissions: block grants issues: write scoped to that job while the
    workflow-level permissions: stays contents: read. No new secret and no new third-party
    action.
  • The hand-rolled alarm job in .github/workflows/release-please.yml is replaced by a job
    calling the same bin, and no inline shell or github-script body deciding
    open-vs-comment-vs-close survives in either workflow.
  • Every docblock the diff adds or touches under packages/fabrika-cli/ cites its ruling with an
    @ruling tag naming issue A red post-merge run on main reaches nobody: 8 failed web deploys sat unseen for 3 days #9508, and none of them names an ADR number or a .decisions/
    path.
  • The deploy-failure-tracker label left on this repository by the PR feat: A red post-merge run on main reaches nobody: 8 failed web deploys sat unseen for 3 days #9522 proof runs no
    longer exists under that name. [evidence: a gh label list read linked or quoted in the PR description]
  • Nothing in the diff adds an outside uptime check on /api/health, and nothing pauses or gates
    the merge queue. Both are ruled out of scope.
  • The changed workflows parse as YAML, and the alarm's trigger behaviour is confirmed against a
    real run of the workflow after merge, or against a deliberately-failed run on a branch.
    [evidence: a workflow run linked in the PR description]
  • deploy.yml's alarm job reports a cancelled run as cancelled, not as success. Its RUN_RESULT expression currently maps a need whose result is cancelled to 'success' when nothing else failed or skipped, so main-alarm.ts's cancelled arm is unreachable from that workflow and a cancelled deploy closes the standing alarm — the false all-clear that arm exists to prevent. release-please.yml already passes needs.release-please.result straight through and is correct; the two callers must agree.

Original report (verbatim)

Summary

When a workflow that only runs after merge goes red on main, nothing tells anyone. The web deploy job failed on 8 main pushes in a row over 3 days and was found only because the founder happened to ask why main was red.

What I was doing

Tracing why the main head showed a red check, which led to the deploy health gate failure now tracked in #9507.

What I observed

deploy.yml runs on push to main. Its job "deploy (web, @kampus/web, true)" failed on every main push from commit 1e52c3c (PR #8934, 2026-09-17 19:01Z) through 0cb9588 (2026-09-20 19:20Z), 8 of 8 runs checked. In that window no notification went out, no issue was filed, and merges kept landing on top of the red.

The PR gates could not have caught it. On PR #8934 the preview deploy passed, because a preview stage has no custom domain and the deploy log still printed a workers.dev URL. Prod has the custom domain and prints only that, so the failure exists only on the post-merge prod run.

The same holds for anything else that runs only on push to main, such as the prod health gate and release-please. I did not audit the full list.

Why it matters

A real prod break would look the same as a stale red and could sit just as long. The health gate from #1436 and ADR 0092 fails closed, but a closed gate nobody sees protects little. Merges continuing over a red prod deploy also make the first bad commit harder to find later. How often this happens is unknown; this is the one case I measured.

Pointers

Suggested next step (non-binding)

On a failed push-to-main run of deploy.yml, open or update one issue, deduped so 8 reds make 1 issue, and ping the founder. Close it on the next green. Worth deciding separately whether a red prod deploy should pause the merge queue.


Filed by an agent · session 56fedc1c-913a-4bc4-baa5-2b748ea48b9b · branch main · 2026-09-20T19:34:32Z

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    axis:pipeline-hardeningStanding cross-cutting axis: pipeline hardening (was milestone #1; go-forward label)class:codecreated by fabrika status bootstrap label-taxonomyp1Medium priorityready-for:agentAn execution engine may pick this up.status:triagedTriage signed off; ready for write-code to picktype:featureNew capability, directly implementable

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions