What this is
A write-up of spike #9481 and its three follow-ups (#9492, #9496, #9498, all closed). The spikes measured whether a model can read a review FAIL comment and tell "this class owes a local edit" from "this class is clean and only waits on another class's repair". Accuracy was 41/41 on a frozen sample; the cheapest variant parsed the structured PASS marker for 28 of 41 and sent only 13 comments to the model.
The open question this issue actually carries sits upstream of all of that: does a repair round in this pipeline waste work rereading a clean class today, and how much? Nothing here measured that.
Triage note — what I read at origin/main
packages/fabrika-cli/src/build/verdicts-verb.ts already emits polarity and current as structured fields per verdict row, and packages/fabrika-cli/src/build/range-verdicts.ts filters standing FAILs off that same parsed marker. So the distinction the spikes recovered from prose is already a field on the current verb for every verdict landed in the marker format. The issue's own second comment says the same. What remains unserved is the older prose-only records, and no issue names a cost for those.
Triage note — why this is not buildable as filed
- No observed incident. The body states that no wasted builder run was observed and no production behaviour changed.
- The suggested next step splits in two: a measurement (how often a repair round rereads a clean class) and a design choice (a reviewer-emitted repair-needed field vs a model helper). The choice cannot be made before the measurement, and the measurement has no owner.
- Dedup on the surface found no open issue that owns it — the four closest matches are the spikes themselves, all closed. Minting standalone rather than folding.
Triage note — recommendation
This fails the value bar's hardening-with-no-incident clause, and the structured-field half is superseded by what build verdicts already returns. My recommendation is to close it rather than open a measurement lane. Left open and parked for a human call; the evidence stays readable on the four closed spike issues either way.
Original body preserved below.
Original report (verbatim)
Summary
A review's FAIL can mean that another class owes repairs while this class needs no edit. A Jev experiment extracted that distinction correctly from 41 new review comments, but the cost benefit and best place to store the distinction remain unmeasured.
What I was doing
Testing cheap classification opportunities in Fabrika with disposable spikes. The first review-intent pilot used 12 comments; the follow-up held its question fixed and used all code, doc and skill review records from 12 different older PRs.
What I observed
Jev matched 41/41 frozen intent labels, including ten requests for local repairs. A FAIL-header baseline matched 38/41 because three FAIL records explicitly needed no local repair. At a confidence cutoff of 0.90, 35 comments were accepted with no observed errors and six escalated. The run used 77,237 input tokens, estimated $0.003243954, with median latency 128 ms. The labels were agent judgments from reviewer statements, not independent human judgments. Cases share reviewer style and PRs.
Why it matters
A builder may need to distinguish a local edit request from a class that only needs another verdict posted after another class is repaired. This experiment establishes text extraction quality on a small sample, not an observed wasted builder run or permission to skip a review. A structured field emitted by the reviewer may be simpler and more reliable than an extra model request. No production behavior changed.
Pointers
Suggested next step (non-binding)
Measure how often repair shells reread clean classes and what that costs. Compare an explicit reviewer-emitted repair-needed field with a Jev helper for old prose records. Keep the existing verdict and required re-review rules. Test the helper in observation-only mode before allowing it to change work selection.
Filed by an agent · session 01a0bb3f-113c-7072-8566-0089336e6ff7 · branch main · 2026-09-19T22:54:25Z
What this is
A write-up of spike #9481 and its three follow-ups (#9492, #9496, #9498, all closed). The spikes measured whether a model can read a review FAIL comment and tell "this class owes a local edit" from "this class is clean and only waits on another class's repair". Accuracy was 41/41 on a frozen sample; the cheapest variant parsed the structured PASS marker for 28 of 41 and sent only 13 comments to the model.
The open question this issue actually carries sits upstream of all of that: does a repair round in this pipeline waste work rereading a clean class today, and how much? Nothing here measured that.
Triage note — what I read at
origin/mainpackages/fabrika-cli/src/build/verdicts-verb.tsalready emitspolarityandcurrentas structured fields per verdict row, andpackages/fabrika-cli/src/build/range-verdicts.tsfilters standing FAILs off that same parsed marker. So the distinction the spikes recovered from prose is already a field on the current verb for every verdict landed in the marker format. The issue's own second comment says the same. What remains unserved is the older prose-only records, and no issue names a cost for those.Triage note — why this is not buildable as filed
Triage note — recommendation
This fails the value bar's hardening-with-no-incident clause, and the structured-field half is superseded by what
build verdictsalready returns. My recommendation is to close it rather than open a measurement lane. Left open and parked for a human call; the evidence stays readable on the four closed spike issues either way.Original body preserved below.
Original report (verbatim)
Summary
A review's FAIL can mean that another class owes repairs while this class needs no edit. A Jev experiment extracted that distinction correctly from 41 new review comments, but the cost benefit and best place to store the distinction remain unmeasured.
What I was doing
Testing cheap classification opportunities in Fabrika with disposable spikes. The first review-intent pilot used 12 comments; the follow-up held its question fixed and used all code, doc and skill review records from 12 different older PRs.
What I observed
Jev matched 41/41 frozen intent labels, including ten requests for local repairs. A FAIL-header baseline matched 38/41 because three FAIL records explicitly needed no local repair. At a confidence cutoff of 0.90, 35 comments were accepted with no observed errors and six escalated. The run used 77,237 input tokens, estimated $0.003243954, with median latency 128 ms. The labels were agent judgments from reviewer statements, not independent human judgments. Cases share reviewer style and PRs.
Why it matters
A builder may need to distinguish a local edit request from a class that only needs another verdict posted after another class is repaired. This experiment establishes text extraction quality on a small sample, not an observed wasted builder run or permission to skip a review. A structured field emitted by the reviewer may be simpler and more reliable than an extra model request. No production behavior changed.
Pointers
claude-plugins/fabrika/skills/review/SKILL.mdpackages/fabrika-cli/src/review/Suggested next step (non-binding)
Measure how often repair shells reread clean classes and what that costs. Compare an explicit reviewer-emitted repair-needed field with a Jev helper for old prose records. Keep the existing verdict and required re-review rules. Test the helper in observation-only mode before allowing it to change work selection.
Filed by an agent · session
01a0bb3f-113c-7072-8566-0089336e6ff7· branchmain· 2026-09-19T22:54:25Z