Skip to content

Grow the sequential test set from annotator imports, append-only - #36

Merged
Chouffe merged 18 commits into
mainfrom
docs/annotator-test-growth-design
Aug 14, 2026
Merged

Chouffe merged 18 commits into
mainfrom
docs/annotator-test-growth-design

Conversation

@Chouffe

@Chouffe Chouffe commented Aug 14, 2026

Copy link
Copy Markdown
Collaborator

What

Annotator imports can now assign sequences to all three splits (80/10/10), and sequential_test becomes append-only: every release's test set is a byte-identical superset of the previous one, so a model trained on an earlier dataset release stays comparable with one trained on a later release (re-evaluate the old model on the grown test set, or on the frozen old subset inside it).

Design: docs/specs/2026-08-14-annotator-test-growth-design.md (supersedes the deferred test-growth item and §4 of the 2026-08-13 import design).

Why the lockfile

The built test set was never a copy of the test split — it is a KMeans selection recomputed on every build. One new WF test sequence changes k and silently re-rolls all 151 historical negatives; test has only been stable because no dep of the stage changed. This PR replaces that accident with recorded state:

  • data/raw/sequential_test_lock.json — ordered, append-only, git-committed list of FP folders that is the FP half of sequential_test. Bootstrapped from the shipped test set (never recomputed), verified byte-identical through a full rebuild.
  • scripts/freeze_test_selection.py — manual command (same state-boundary rationale as plan_annotator_import.py): pins annotator test FPs up to quota in registry order, fills remaining slots via the existing two-stage selection, never removes or reorders.
  • build_sequential_dataset.py does no selection for test: it copies the lockfile verbatim and hard-errors on any mismatch (stale quota, unregistered folder, missing on disk).

Also in here

  • assignments_from_splits / add_data.py accept pre-assigned test splits; SPLIT_TARGETS opens to 80/10/10 (greedy deficit — expect the first ~15–20 new assignments to catch up into test).
  • Same-fire guard: new smoke alerts on one (camera, azimuth) view chaining within --same-fire-window (12 h, transitive) share one split, inheriting from previously planned smoke — a fire can no longer straddle train and test.
  • Leakage suite gains ledger split-pinning and lockfile↔quota consistency checks.
  • dvc.yaml: the build stage depends on the lockfile instead of the test embeddings, so test-embedding churn no longer rebuilds sequential_test.
  • Import workflow gains one step: plan → materialise → add_data ×2 → freeze_test_selection.py → dvc repro.

The detector's yolo_test deliberately keeps its current behaviour (see Non-goals in the spec).

This branch also carries the #35 recurring-object FP identity spec + plan (docs only).

Verification

  • 180 tests pass (uv run pytest tests/), including new suites for the lockfile, freeze command, same-fire grouping, and build errors.
  • Bootstrap fidelity: lockfile == shipped sequential_test/test/fp (151 folders), and dvc repro build_sequential_dataset reproduces the test split byte-for-byte (folder-list diff empty for both WF and FP).

Chouffe added 18 commits August 14, 2026 15:07
The lockfile lives at data/raw/, not data/raw/fp/: the fp directory is a
single DVC-tracked object whose gitignore would swallow the file and whose
hash every freeze would churn. Spec updated accordingly.
- The workflow (CLAUDE.md and the planner's Next output) now refreshes
  test embeddings before the freeze, as the spec always required —
  without it a fill would select from stale embeddings.
- CLAUDE.md notes that ANY ingest adding test WF sequences needs a
  freeze, not only annotator imports.
- test_data_leakage stage depends on the ledger, registries and
  lockfile, so hand-edits re-trigger the consistency checks.
- Same-fire chains touching planned folders in two splits now warn:
  that is the historical leak the guard exists for.
- Drop the stale 'pre-assigned to test' comment in add_data.py.
Every command with its real output, what to check at each step, the
commit checklist, and a recovery table. The README and CLAUDE.md point
to it. Also corrects add_data's misleading orphan-folder warning:
already-on-disk folders are skipped, not registered, so the remedy is
to delete them from the pool and re-run.
@Chouffe
Chouffe merged commit 851e98d into main Aug 14, 2026
3 checks passed
Chouffe added a commit that referenced this pull request Aug 14, 2026
The frozen-lockfile test branch and the pinning simplification compose:
test negatives copy the lockfile verbatim (no selection, no embeddings);
train/val keep per-source pinning with the shared helpers from
pyro_dataset.fp.selection. The pinning-helper tests stay in
tests/test_fp_pinning.py; tests/test_build_sequential_pinning.py keeps
only the frozen_test_selection coverage from main.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant