Skip to content

CI log reader stops at governance check while test logs remain available #9097

Description

@usirin

Fix unavailable CI logs hiding readable failures

heal-ci logs must retain available failure evidence when another selected context has no readable log, while clearly reporting that its evidence is incomplete. It currently stops the aggregate read and emits no log frames.

Verified cause and limits

Triage note: read at fresh origin/main commit ec425da. packages/fabrika-cli/src/heal-ci/logs-verb.ts matches workflow jobs by context name, buffers every frame, and returns immediately from the Expired branch before rendering those frames. packages/fabrika-cli/src/heal-ci/github.ts maps both HTTP 404 and HTTP 410 from a job-log endpoint to Expired.

The reported identifier is not merely a check absent from the Actions job list. GitHub's job endpoint identifies 103001620692 as governance floor at head, completed/success, attempt 1 of run 34515850336. That run has event dynamic, name PR #8934, and head fd967cf. Reading that job's log endpoint returned HTTP 404 during triage. The report independently records that the real integration logs in run 34515854007 remained readable without a rerun.

These reads establish the aggregate abort and an unsupported expiry claim. They do not establish that a retained log expired, that name matching selected the wrong job in this incident, or that the integration failures were transient. Preserve those distinctions in the repair.

Scope and home

This is a bug: the verb promises evidence for every failing context, and an unrelated unavailable log prevents it from serving available evidence. The delivery is one bounded CLI repair, with its tests and interface documentation. Home: axis:pipeline-hardening; priority p2.

Triage note: searched the exact observation and the wider aggregate-log owner. #9034 concerns governance re-fire detection; #8988 concerns review ordering. #7206 owns classification of a readable aggregator log, and #7208 owns setup-crash signatures. None owns unavailable-log retrieval, so folding this report into them would expand a different fix. Keep their classification and rerun rules intact.

Acceptance criteria

  • A regression reproduces a selected dynamic governance job whose metadata is readable but log endpoint returns 404, alongside a failed integration context whose log is readable. Use the reported job/run/head relationships as evidence for the fixture; do not require access to aging live CI records for the test.
  • The aggregate result retains each available selected context's log and explicitly identifies each unavailable context and its observed reason. Missing evidence cannot produce a complete-success claim, an empty-success result, or silent omission. Preserve complete enumeration and exact head binding.
  • Log availability distinguishes a context without log bytes, an unknown read, and proven expiry using source-backed evidence. A readable job followed by this HTTP 404 alone must not claim that retained logs expired or recommend a rerun as though expiry or transience were established. Verify the selected check's job/attempt relationship; do not invent a wrong-job explanation for this incident.
  • Tests cover mixed readable/unavailable results, actual expiry, transport or permission failure, and a context with no matching job, for both text and JSON outputs and their downstream reader. Preserve informational exclusions, truncation disclosure, and the classifier/rerun refusal to infer a transient failure from missing evidence.
  • Update the owning command help and contract for the explicit result and exit behavior. Required package tests and configured checks pass. Do not broaden this repair into the aggregator signature policy from issue 7206, governance re-fire changes, or automatic reruns.

Evidence


Original report (verbatim)

Summary

Fabrika's aggregate CI log read stopped at a governance check and returned no failing test logs. The same integration run's logs were still readable through GitHub CLI.

What I was doing

Diagnosing the failed integration check on PR 8934 at fd967cf using fabrika heal-ci diagnose 8934, then fabrika heal-ci logs 8934.

What I observed

Diagnosis returned stall red. The log command exited 15 with: heal-ci logs: run 103001620692's logs are expired — the platform no longer holds them; classify from the check-run summary or re-run to regenerate.

That identifier was the failing governance floor at head check. Immediately afterward, gh run view 34515854007 --repo kamp-us/phoenix --log-failed returned the real integration logs. They contained a Pano cache-refresh assertion failure and a separate connection-reset failure during live-posts setup. No rerun was needed to read this evidence.

Why it matters

One unreadable check prevented the aggregate reader from returning available failure evidence. I do not yet know whether the cause is job matching, transient log availability, or the handling of a check without ordinary job logs. The error's suggestion to rerun is not sufficient evidence that these test failures were transient.

Pointers

Suggested next step (non-binding)

Reproduce the selected-check/job/log mapping using these records. Preserve the refusal to claim a complete scan when evidence is missing, while making available test logs and each missing context distinguishable.


Filed by an agent · session 01a089a6-dd27-7003-8552-89997050082e · 2026-09-10T18:59:49Z

Activity

  1. usirin commented on Sep 10, 2026

    @usirin
    MemberAuthor
    No description provided.
  2. added
    axis:pipeline-hardeningStanding cross-cutting axis: pipeline hardening (was milestone #1; go-forward label)
    p2Lowest priority
    ready-for:agentAn execution engine may pick this up.
    status:triagedTriage signed off; ready for write-code to pick
    type:bugBehavior diverges from intent
    and removed on Sep 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    axis:pipeline-hardeningStanding cross-cutting axis: pipeline hardening (was milestone #1; go-forward label)p2Lowest priorityready-for:agentAn execution engine may pick this up.status:triagedTriage signed off; ready for write-code to picktype:bugBehavior diverges from intent

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions