Skip to content

Phase I MVP: vision CIFAR continual scaffolding + smoke tests #1

Description

@icedac

Goal

Bring hal-sleep-wake from empty (README only) to Phase I MVP per L1 paper §5 Roadmap:

Phase I (Months 0–6): Proof of Concept — Implement basic sleep-wake cycle with existing LLMs, open-source repository with reproducible benchmarks, Target: 50% forgetting reduction on CIFAR sequence.

Paper: https://github.com/labforadvancedstudy/paper/blob/main/HAL/L1_Hierarchical_Abstraction_All_You_Need_v2.md

Decided scope

Vision-benchmark reproduction (Option A). §2.3 table's LLM stack is re-mapped onto a vision backbone to stay on Phase I target:

Layer Paper (LLM stack) This MVP (vision stack)
L1 Context window Task-batch buffer (working memory analog)
L2 LoRA weights LoRA adapter on convolutional layers
L3 Small model (7B) ResNet-18 base weights (long-term substrate)
L4 Large model (70B) Out of scope for Phase I

Dataset sequence (subset of paper §3.1): CIFAR-10 → CIFAR-100. Full 7-dataset sequence is a follow-up PR.

First PR boundary (this issue)

Deliverable: scaffolding complete + 2-task continual pipeline runnable + smoke tests pass + benchmark script callable in dry-run mode. Actual training results land in a follow-up PR because CI has no GPU.

File layout

hal-sleep-wake/
├── .github/workflows/ci.yml         # ruff + pytest (smoke only, no training)
├── .gitignore
├── pyproject.toml                   # deps: torch, torchvision, peft, pytest, ruff
├── README.md                        # updated: Phase I MVP status, run instructions
├── hal_sleep_wake/
│   ├── __init__.py
│   ├── config.py                    # dataclass: HyperParams (lr_wake, lr_nrem, sleep_every, lora_r, alpha)
│   ├── memory.py                    # L1 buffer + L2 LoRA adapter interface + L3 backbone interface
│   ├── model.py                     # ResNet-18 backbone + LoRA injection (conv+linear)
│   ├── wake.py                      # wake phase: train task_i with LoRA updates
│   ├── sleep.py                     # sleep phase: merge LoRA → base weights (α-scaled), reset LoRA
│   ├── metrics.py                   # per-task accuracy + forgetting rate (avg acc drop across tasks)
│   └── data.py                      # CIFAR-10 / CIFAR-100 loaders, class-incremental split helpers
├── scripts/
│   ├── train_cifar10_cifar100.py    # 2-task continual loop, supports --dry-run and --epochs-per-task
│   └── eval.py                      # loads saved state, emits accuracy matrix + forgetting rate
├── tests/
│   ├── test_memory.py               # LoRA merge preserves forward pass equivalence at α=0; differs at α>0
│   ├── test_metrics.py              # forgetting rate = avg_{i<T} (acc_i@i - acc_i@T)
│   ├── test_model.py                # ResNet-18 + LoRA: params trainable only on LoRA path
│   └── test_smoke.py                # 1-batch forward/backward on 2 mock classes (no dataset download)
└── results/
    └── README.md                    # placeholder for future benchmark results

Success criteria (this PR)

  • ruff check . clean
  • pytest passes (smoke only, no GPU, no dataset download)
  • python scripts/train_cifar10_cifar100.py --dry-run prints the planned loop (tasks, epochs, sleep schedule) without downloading data
  • README.md has run instructions + Phase I target statement + clear "results not yet measured" disclaimer
  • CI (GitHub Actions) green

Out of scope (follow-up PRs)

  • Actual GPU training and results/baseline_vs_hal.md with numbers
  • REM synthetic-dream phase (paper §2.4 line 99-101)
  • Entropy-aware context compression (paper §2.4 line 85-88)
  • L4 (large model) tier
  • Remaining 5 datasets in the paper §3.1 sequence

Technical notes for implementer

  • LoRA: use peft with LoraConfig(r=8, lora_alpha=16, target_modules=["conv1","fc"] or broader). For ResNet-18 we may need to attach LoRA to layer1..layer4 conv modules — verify peft supports conv layer targets; fallback: manual LoRA impl (< 40 LOC).
  • Sleep merge: LoRA delta W' = W + (α/r) · B @ A. After merge, zero out A,B and re-init. Keep optimizer state reset.
  • Classification heads: per-task heads (CIFAR-10 = 10 classes, CIFAR-100 = 100 classes). Save head weights per task; evaluation restores the matching head.
  • No GPU in CI: tests must use torch.device("cpu") and small tensors. Smoke test: torch.randn(2, 3, 32, 32) through backbone + LoRA.
  • Dry-run mode: --dry-run skips torchvision.datasets.CIFAR10(download=True) and skips training loops; just prints plan.

References

  • Paper §2.3 (memory hierarchy table)
  • Paper §2.4 (sleep-wake pseudocode)
  • Paper §3.1 (dataset sequence)
  • Paper §3.2 (baseline comparison table — our benchmark target)
  • Paper §5 Phase I (target: 50% forgetting reduction on CIFAR sequence)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions