Skip to content

feat(aidd-qa): create an implementation-independent QA plugin - #917

Merged
blafourcade merged 25 commits into
nextfrom
feat/aidd-qa-plugin
Sep 24, 2026
Merged

blafourcade merged 25 commits into
nextfrom
feat/aidd-qa-plugin

Conversation

@blafourcade

@blafourcade blafourcade commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

🎯 What & why

Acceptance QA lived in aidd-dev, the plugin that also implements the change: the author validated its own work. This adds aidd-qa, which owns implementation-independent acceptance validation and reviewer evidence, and moves Browser QA into it.

🛠️ How it works

  • New plugin plugins/aidd-qa/: first release 1.0.0 (pinned with release-as), Execution layer, recommended: false.
    • One skill, aidd-qa:01-acceptance-qa: acceptance criteria (issue, spec, story, or user) + reviewed candidate in, qa.md + qa/*.webm out. Browser is the only interface.
    • Actions: 01 load-scope → 02 prerequisites → 03 prepare-run → 04 run-scenarios.
  • Scope comes only from acceptance criteria. Never a plan, the diff, the code, or tests. A criterion, or a quoted part of one, with no browser-observable outcome is listed under Out of interface.
  • Safety: an entry is never started on storage not proven test-only when its reset deletes data; the skill asks once instead. Project memory (testing.md) decides the entry.
  • Report: per scenario Criteria, Expected, Actual, Verdict, Duration, Evidence; plus Out of interface and Rejected. Run verdict: any fail => fail; any blocked or rejected => blocked; skipped only when nothing is browser-observable; else pass.
  • Runner: run-code returns { step, expected, actual, ok } per step, so a product mismatch is a fail, and a throw is a tooling failure. Commands run from a temp directory outside the app repository. Video contract unchanged (vp9, 1280×720, ≤ 12 s, frame checks).
  • aidd-dev:11-browser-qa is removed: browser QA now lives in aidd-qa. aidd-dev:06-test is scoped to developer-side validation.
  • Registration: marketplace, release-please config and manifest, ci.yml build matrix, commitlint scopes, docs/ARCHITECTURE.md, docs/CATALOG.md, README, memory bank.

🧪 How to verify

  • pnpm exec lefthook run pre-commit: 503/503 pass; node scripts/check-architecture-rules.js: no violation.
  • claude plugin validate plugins/aidd-qa: passed. Codex translate and aidd plugin install aidd-qa --tool claude|codex: installed.
  • Real runs on a mobile-first PWA (7 runs):
    • Full run on an isolated Compose project: no question, 3 scenarios, 3 valid videos, machine left as found.
    • Memory pointing to a dev database: the skill stops before starting anything and asks.
    • A product gap (a "counter jumps" animation) surfaced as fail from the returned results.

⚠️ Heads-up

🔗 Linked issue

Closes #908

✅ I certify

  • I DO CERTIFY I READ EACH LINE OF THE PULL REQUEST BECAUSE I AM A SOFTWARE ENGINEER, NOT A AI PUPPY.

🤖 Generated with Claude Code

https://claude.ai/code/session_01WuxKN5bm96LFhU97rsgBUW

blafourcade and others added 16 commits September 24, 2026 05:58
Moves plugins/aidd-dev/skills/11-browser-qa to
plugins/aidd-qa/skills/01-acceptance-qa (git history preserved) and
rewrites it to derive scenarios only from acceptance criteria, never
the diff or the source code. Renames its Playwright reference to
interface-browser-playwright-cli.md, adds Criterion/Expected/Actual
columns to the QA report template, and pins the architecture-rules
sweep count to 49 skills-with-actions.

Refs #908

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
plugins/aidd-dev/skills/11-browser-qa now holds a single redirect
action that names the aidd-qa plugin and its install command, then
stops without loading a scope or recording evidence. Drops Browser QA
from the plugin description and README now that it moved.

Refs #908

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
Adds the aidd-qa entry to marketplace.json (recommended: false, off
the curated install path like aidd-ui and aidd-telemetry), versions it
in release-please-config.json and the manifest, adds it to the
build-plugin CI matrix, and adds aidd-qa/qa as commitlint scopes.
Updates docs/ARCHITECTURE.md, docs/CATALOG.md, docs/MAINTAINERS.md,
README.md and the project memory bank (architecture, deployment,
project-brief, testing) to name the 9th plugin and its new owner of
browser acceptance QA.

Refs #908

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
Adds the phased plan behind the aidd-qa plugin (issue #908) and the
gate, host-proof, and architecture-conformance evidence collected
while implementing it.

Refs #908

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
The previous version overstated what was observed: it claimed each
commit's hook run was gate evidence for that commit's own tree, when
every gate actually ran once against the final working tree. Records
the two intermediate trees are not independently clean, the blocked
push (pre-existing, machine-specific cli-test failure, isolated and
unrelated to this branch's changes), and deviations not yet noted.

Refs #908

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
The skills description did not exclude acceptance QA, so it could be
picked for reviewer evidence instead of the dedicated aidd-qa plugin.
Frame it as developer-side test validation during implementation and
add that exclusion; CATALOG.md is regenerated from the new description.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
…ract

01-load-scope admitted a plan as a criteria source outright; criteria now
come only from the issue, spec, or user story (or ones the user gives),
and a plans browser Test Scope edge case is admitted only when it maps
to one of those. The stale fewer-than-3 edge-case count check is
dropped, and the first process step now opens with a verb.

qa-report-template gains an Out of interface section so a criterion with
no browser-observable outcome is listed, never silently dropped, and the
header verdict cannot read pass without stating that list. Verdict
values are now explicit (pass | fail | blocked per scenario, plus
skipped at the header when nothing is browser-observable), and the
Duration column is restored. 03-run-scenarios is wired to fill both.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
architecture.mds fits-together diagram still read plugins/ dot 8 after
the aidd-qa plugin landed; the stack table already said 9. Align the
diagram with it.

The aidd-qa plugin tasks validation.md described the push as blocked on
a local node/codex PATH collision. It has since been resolved (node
binary copied into an isolated directory, prepended to PATH; a symlink
does not work because process.execPath resolves it) and cd048e6 was
pushed with the full pre-push gate passing, no --no-verify. Replace the
stale blocked account with what actually happened.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
Give run-scenarios an ordered rule for the header verdict (any fail
beats any blocked, beats no scenario ran producing skipped, else
pass), restoring the "mark the run failed" behavior review #908 found
missing after e8d3538. Move the zero-browser-observable-criteria
short circuit into load-scope itself: it now stops right after its
Filter step and reports skipped, so prerequisites and prepare-run
never run for a scope with no browser-observable criterion. Document
the exception in SKILL.md's transversal rules, since the router's
linear flow otherwise implies prerequisites always runs first. Also
merges a duplicate load-scope Test bullet review #908 flagged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
The 06-test SKILL.md description was narrowed to developer-side
validation with no acceptance evidence and no sibling-plugin address
in ba2a59a, but this hand-written README row still read like generic
"write and iterate on tests" coverage. Review #908 flagged the drift.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
docs/CATALOG.md still described 06-test as generic test coverage after
ba2a59a narrowed its scope; align it with the SKILL.md description
and the now-fixed aidd-dev README row.

The aidd-qa plugin's validation record had gone stale across three
commits landed after it was last corrected (ba2a59a, e8d3538,
cff70e6): its commit table stopped five commits short of HEAD and its
push section named only the first push's SHA. Rewrite both against the
current branch history, re-run the scripts suite and architecture
check on the final tree with one consistent number instead of the
earlier 503-pass/501-pass-plus-2-skip split, and describe the push
mechanism (an isolated copied-node PATH, never --no-verify) without
freezing it to a SHA the next push would make stale again.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
The prior transversal rule said prerequisites never runs when
load-scope finds no browser-observable criterion, but the router
still lists 00-prerequisites ahead of 01-load-scope, so an executor
following "read only the next action's file" would check and possibly
install ffmpeg/Playwright before load-scope ever got to decide. Move
the check itself ahead of 00: the router now says to test the
criteria before invoking prerequisites at all, and to run load-scope
alone (reporting skipped) when none is browser-observable.

Also corrects 8924333's commit body, which said this short circuit
happens "after its Filter step" -- load-scope actually stops one step
later, after Locate resolves the evidence folder the skip report
needs. That commit is already pushed and is not being amended;
recorded here instead.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
db49dea's fix copied the dispatch's own writing constraint ("no
sibling-plugin address") into the README row as if it described what
06-test does. Replace it with what a reader actually needs: 06-test is
not independent acceptance QA or reviewer evidence, mirroring the
SKILL.md description's own "Do NOT use for" clause.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
docs/CATALOG.md's 06-test row still ended with the addressing note
"no sibling-plugin address" that a6e0cda already dropped from the
README row for the same reason: it is this task's own writing
constraint, not a description of what the skill does. Replace it with
the same "not independent acceptance QA or reviewer evidence" wording.

Rewrite the validation record's "Commits" and "Push" sections, which
had gone stale across three commits (ba2a59a, e8d3538, cff70e6) and
now this repair's own nine, and describe the repair itself honestly:
what review #908 found, the routing-rule and docs-wording bugs the
first pass at fixing it introduced and a second pass corrected, and
the final gate re-run's numbers with exit codes captured directly
rather than through a wrapper that can mask them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
…ipped

A run with browser-observable criteria whose every scenario got rejected
(no executable teardown, etc.) fell through to header verdict skipped,
which the template and load-scope reserve for zero browser-observable
criteria. A rejected criterion then appeared in neither Scenarios nor Out
of interface, so a candidate that was never validated read as nothing to
validate.

Rejections are now carried forward with their reason from load-scope's
Validate and prepare-run's Reset into run-scenarios' report, which adds a
Rejected section and blocks the run whenever any exists. load-scope's own
Skip step (zero criteria survive the Filter) now writes and returns
<evidence-folder>/qa.md instead of leaving the destination to the agent.
SKILL.md's router row and mermaid are corrected to say "at most one" happy
path and to show the skip exit from load-scope to the report.

Review #908.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
…add install proof

Review #908's second follow-up pass on this record found "fourth defect"
where the load-scope-skips-prerequisites fix is really the second in its
own list, row 14's subject carrying a spurious "(final)" that never
appeared in 3c81dfb's real message, and "this repair's own three commits
(9-11, then 12-14)" undercounting a six-commit repair. All three are fixed
here; "defect four" describing the 06-test docs fix is left alone, since
that one really is the fourth defect in the list.

Also adds rows 15-16 for this round's own commits (row 16, this commit,
referred to generically since it cannot state its own SHA), a new "Second
repair re-run" gate table, and a Host proof entry for criterion 10: an
isolated-HOME sandbox install of aidd-qa for both --tool claude and
--tool codex, run to close the gap the review found between the criterion
and its recorded proof.

Review #908.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
blafourcade and others added 4 commits September 24, 2026 08:16
A first real run on a PWA wiped the developer's own database through the
documented test reset. Resets now require proof of test-only storage, or
one question.

- run load-scope before prerequisites; drop the special skip rule
- split partly observable criteria, quote the rest out of interface
- run-code returns per-step expected, actual, ok; a throw is tooling only
- run browser commands from a temp dir outside the application repository

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
- load-scope: 7 steps, 4 tests; edges only when a criterion names them,
  an unmapped plan edge is never a candidate nor a rejection
- prepare-run: prefer a test-only entry over the development one
- reference: tear down and verify baseline before closing the session;
  the no-substitution rule scopes to runner commands

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
…uard

- prepare-run: Reuse before Preflight; drop the generic test-only
  preference, project memory decides and the reset guard asks
- load-scope: teardown is checked in prepare-run only; Show names the
  evidence folder

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
The reset guard ran only at teardown, after the entry had already started
on possibly shared storage (and run its migrations). It now runs in Reuse,
before anything starts. The router row matches load-scope's edge rule.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
blafourcade and others added 2 commits September 24, 2026 10:08
Same behavior in fewer words: the verdict order is stated once, the
prepare-run return step folds into its output, and run-scenarios' input
names everything load-scope hands over. A blocked row reports evidence
`none`; only a failed take is kept.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
Browser QA now lives in the aidd-qa plugin, so aidd-dev owns no QA
surface. The catalog lists aidd-qa's renumbered actions, and the plan
records the removal.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
blafourcade and others added 3 commits September 24, 2026 11:30
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
Pin the first release with the package's `release-as`; the manifest
stays at 0.1.0 until release-please writes 1.0.0. A `Release-As` footer
would re-version every other path the squash commit touches.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WuxKN5bm96LFhU97rsgBUW
AIDD-Session-Id: 6d6ccc35-f6c1-417d-84a6-9a0b70474e6b
@blafourcade
blafourcade marked this pull request as ready for review September 24, 2026 10:21
@blafourcade
blafourcade requested a review from a team as a code owner September 24, 2026 10:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant