Skip to content

Review FAIL text does not always request a local repair #9494

Description

@usirin

What this is

A write-up of spike #9481 and its three follow-ups (#9492, #9496, #9498, all closed). The spikes measured whether a model can read a review FAIL comment and tell "this class owes a local edit" from "this class is clean and only waits on another class's repair". Accuracy was 41/41 on a frozen sample; the cheapest variant parsed the structured PASS marker for 28 of 41 and sent only 13 comments to the model.

The open question this issue actually carries sits upstream of all of that: does a repair round in this pipeline waste work rereading a clean class today, and how much? Nothing here measured that.

Triage note — what I read at origin/main

packages/fabrika-cli/src/build/verdicts-verb.ts already emits polarity and current as structured fields per verdict row, and packages/fabrika-cli/src/build/range-verdicts.ts filters standing FAILs off that same parsed marker. So the distinction the spikes recovered from prose is already a field on the current verb for every verdict landed in the marker format. The issue's own second comment says the same. What remains unserved is the older prose-only records, and no issue names a cost for those.

Triage note — why this is not buildable as filed

  • No observed incident. The body states that no wasted builder run was observed and no production behaviour changed.
  • The suggested next step splits in two: a measurement (how often a repair round rereads a clean class) and a design choice (a reviewer-emitted repair-needed field vs a model helper). The choice cannot be made before the measurement, and the measurement has no owner.
  • Dedup on the surface found no open issue that owns it — the four closest matches are the spikes themselves, all closed. Minting standalone rather than folding.

Triage note — recommendation

This fails the value bar's hardening-with-no-incident clause, and the structured-field half is superseded by what build verdicts already returns. My recommendation is to close it rather than open a measurement lane. Left open and parked for a human call; the evidence stays readable on the four closed spike issues either way.


Original body preserved below.


Original report (verbatim)

Summary

A review's FAIL can mean that another class owes repairs while this class needs no edit. A Jev experiment extracted that distinction correctly from 41 new review comments, but the cost benefit and best place to store the distinction remain unmeasured.

What I was doing

Testing cheap classification opportunities in Fabrika with disposable spikes. The first review-intent pilot used 12 comments; the follow-up held its question fixed and used all code, doc and skill review records from 12 different older PRs.

What I observed

Jev matched 41/41 frozen intent labels, including ten requests for local repairs. A FAIL-header baseline matched 38/41 because three FAIL records explicitly needed no local repair. At a confidence cutoff of 0.90, 35 comments were accepted with no observed errors and six escalated. The run used 77,237 input tokens, estimated $0.003243954, with median latency 128 ms. The labels were agent judgments from reviewer statements, not independent human judgments. Cases share reviewer style and PRs.

Why it matters

A builder may need to distinguish a local edit request from a class that only needs another verdict posted after another class is repaired. This experiment establishes text extraction quality on a small sample, not an observed wasted builder run or permission to skip a review. A structured field emitted by the reviewer may be simpler and more reliable than an extra model request. No production behavior changed.

Pointers

Suggested next step (non-binding)

Measure how often repair shells reread clean classes and what that costs. Compare an explicit reviewer-emitted repair-needed field with a Jev helper for old prose records. Keep the existing verdict and required re-review rules. Test the helper in observation-only mode before allowing it to change work selection.


Filed by an agent · session 01a0bb3f-113c-7072-8566-0089336e6ff7 · branch main · 2026-09-19T22:54:25Z

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    axis:pipeline-hardeningStanding cross-cutting axis: pipeline hardening (was milestone #1; go-forward label)p2Lowest priorityready-for:humanA human picks this up.status:triagedTriage signed off; ready for write-code to picktype:investigationUnknown; output is knowledge

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions