Goal
Bring hal-sleep-wake from empty (README only) to Phase I MVP per L1 paper §5 Roadmap:
Phase I (Months 0–6): Proof of Concept — Implement basic sleep-wake cycle with existing LLMs, open-source repository with reproducible benchmarks, Target: 50% forgetting reduction on CIFAR sequence.
Paper: https://github.com/labforadvancedstudy/paper/blob/main/HAL/L1_Hierarchical_Abstraction_All_You_Need_v2.md
Decided scope
Vision-benchmark reproduction (Option A). §2.3 table's LLM stack is re-mapped onto a vision backbone to stay on Phase I target:
| Layer |
Paper (LLM stack) |
This MVP (vision stack) |
| L1 |
Context window |
Task-batch buffer (working memory analog) |
| L2 |
LoRA weights |
LoRA adapter on convolutional layers |
| L3 |
Small model (7B) |
ResNet-18 base weights (long-term substrate) |
| L4 |
Large model (70B) |
Out of scope for Phase I |
Dataset sequence (subset of paper §3.1): CIFAR-10 → CIFAR-100. Full 7-dataset sequence is a follow-up PR.
First PR boundary (this issue)
Deliverable: scaffolding complete + 2-task continual pipeline runnable + smoke tests pass + benchmark script callable in dry-run mode. Actual training results land in a follow-up PR because CI has no GPU.
File layout
hal-sleep-wake/
├── .github/workflows/ci.yml # ruff + pytest (smoke only, no training)
├── .gitignore
├── pyproject.toml # deps: torch, torchvision, peft, pytest, ruff
├── README.md # updated: Phase I MVP status, run instructions
├── hal_sleep_wake/
│ ├── __init__.py
│ ├── config.py # dataclass: HyperParams (lr_wake, lr_nrem, sleep_every, lora_r, alpha)
│ ├── memory.py # L1 buffer + L2 LoRA adapter interface + L3 backbone interface
│ ├── model.py # ResNet-18 backbone + LoRA injection (conv+linear)
│ ├── wake.py # wake phase: train task_i with LoRA updates
│ ├── sleep.py # sleep phase: merge LoRA → base weights (α-scaled), reset LoRA
│ ├── metrics.py # per-task accuracy + forgetting rate (avg acc drop across tasks)
│ └── data.py # CIFAR-10 / CIFAR-100 loaders, class-incremental split helpers
├── scripts/
│ ├── train_cifar10_cifar100.py # 2-task continual loop, supports --dry-run and --epochs-per-task
│ └── eval.py # loads saved state, emits accuracy matrix + forgetting rate
├── tests/
│ ├── test_memory.py # LoRA merge preserves forward pass equivalence at α=0; differs at α>0
│ ├── test_metrics.py # forgetting rate = avg_{i<T} (acc_i@i - acc_i@T)
│ ├── test_model.py # ResNet-18 + LoRA: params trainable only on LoRA path
│ └── test_smoke.py # 1-batch forward/backward on 2 mock classes (no dataset download)
└── results/
└── README.md # placeholder for future benchmark results
Success criteria (this PR)
ruff check . clean
pytest passes (smoke only, no GPU, no dataset download)
python scripts/train_cifar10_cifar100.py --dry-run prints the planned loop (tasks, epochs, sleep schedule) without downloading data
README.md has run instructions + Phase I target statement + clear "results not yet measured" disclaimer
- CI (GitHub Actions) green
Out of scope (follow-up PRs)
- Actual GPU training and
results/baseline_vs_hal.md with numbers
- REM synthetic-dream phase (paper §2.4 line 99-101)
- Entropy-aware context compression (paper §2.4 line 85-88)
- L4 (large model) tier
- Remaining 5 datasets in the paper §3.1 sequence
Technical notes for implementer
- LoRA: use
peft with LoraConfig(r=8, lora_alpha=16, target_modules=["conv1","fc"] or broader). For ResNet-18 we may need to attach LoRA to layer1..layer4 conv modules — verify peft supports conv layer targets; fallback: manual LoRA impl (< 40 LOC).
- Sleep merge: LoRA delta
W' = W + (α/r) · B @ A. After merge, zero out A,B and re-init. Keep optimizer state reset.
- Classification heads: per-task heads (CIFAR-10 = 10 classes, CIFAR-100 = 100 classes). Save head weights per task; evaluation restores the matching head.
- No GPU in CI: tests must use
torch.device("cpu") and small tensors. Smoke test: torch.randn(2, 3, 32, 32) through backbone + LoRA.
- Dry-run mode:
--dry-run skips torchvision.datasets.CIFAR10(download=True) and skips training loops; just prints plan.
References
- Paper §2.3 (memory hierarchy table)
- Paper §2.4 (sleep-wake pseudocode)
- Paper §3.1 (dataset sequence)
- Paper §3.2 (baseline comparison table — our benchmark target)
- Paper §5 Phase I (target: 50% forgetting reduction on CIFAR sequence)
Goal
Bring
hal-sleep-wakefrom empty (README only) to Phase I MVP per L1 paper §5 Roadmap:Paper: https://github.com/labforadvancedstudy/paper/blob/main/HAL/L1_Hierarchical_Abstraction_All_You_Need_v2.md
Decided scope
Vision-benchmark reproduction (Option A).
§2.3table's LLM stack is re-mapped onto a vision backbone to stay on Phase I target:Dataset sequence (subset of paper §3.1): CIFAR-10 → CIFAR-100. Full 7-dataset sequence is a follow-up PR.
First PR boundary (this issue)
Deliverable: scaffolding complete + 2-task continual pipeline runnable + smoke tests pass + benchmark script callable in dry-run mode. Actual training results land in a follow-up PR because CI has no GPU.
File layout
Success criteria (this PR)
ruff check .cleanpytestpasses (smoke only, no GPU, no dataset download)python scripts/train_cifar10_cifar100.py --dry-runprints the planned loop (tasks, epochs, sleep schedule) without downloading dataREADME.mdhas run instructions + Phase I target statement + clear "results not yet measured" disclaimerOut of scope (follow-up PRs)
results/baseline_vs_hal.mdwith numbersTechnical notes for implementer
peftwithLoraConfig(r=8, lora_alpha=16, target_modules=["conv1","fc"] or broader). For ResNet-18 we may need to attach LoRA tolayer1..layer4conv modules — verifypeftsupports conv layer targets; fallback: manual LoRA impl (< 40 LOC).W' = W + (α/r) · B @ A. After merge, zero out A,B and re-init. Keep optimizer state reset.torch.device("cpu")and small tensors. Smoke test:torch.randn(2, 3, 32, 32)through backbone + LoRA.--dry-runskipstorchvision.datasets.CIFAR10(download=True)and skips training loops; just prints plan.References