Skip to content

Latest commit

 

History

225 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ROSClaw-Memory

Evidence-Grounded Selective Memory for Physical AI

Memory conditioning is an intervention. ROSClaw-Memory governs when historical evidence is allowed to influence future physical action.

Relevance ≠ Utility ≠ Authority

Frozen evidence Result Status
Integrated Full16 Attempt 5 1288/2400 = 53.67% Verified same-host native execution
RoboMME PatternLock public official-test 140/150 = 93.33% Verified local task result
PatternLock same-host GPU Docker R2 142/150 = 94.67% Verified rehearsal; not external clean-room
PatternLock generated hard: correct 18/24 Qualification mechanism control
Same events, reversed order 0/24 Negative mechanism control
Wrong-session memory 0/24 Negative identity control

These are not organizer-heldout or official leaderboard scores. The selective Full16 projection remains a matched frozen composition; the separately executed Integrated Full16 result is 1288/2400. Its comparators are paired historical rows, not concurrent randomized controls. See CLAIMS.md before citing any number.

Memory can hurt

Generic Hybrid smoke arm Successes
Control 9/30
Memory rerank 5/30
Memory fallback 5/30

Each memory arm has one rescue and five harms on matched branches. This small experiment rejects the operational continuation gate; it does not estimate a universal population effect.

Authority can reject

Prospective VideoUnmask arm Successes
Auto8 candidate 66/90
Competitive MemER 66/90

With 18 rescues and 18 harms, the frozen decision is DO_NOT_QUALIFY. Registry v1 remains unchanged. No retuning on this cohort is authorized.

The scientific behavior is frozen. See submission readiness, reproduction, and organizer package.

Selective Memory Authority

                    ┌──────────── BYPASS ─────────→ Base Policy
                    │
Physical History
      ↓
Evidence Memory
      ↓
Retrieval
      ↓
Applicability
      ↓
Utility Evidence
      ↓
Authority Gate ─────┼──────────── AUGMENT ───────→ Specialist
                    │
                    └──────────── ABSTAIN ───────→ Safe Stop

Retrieval similarity does not authorize physical intervention. A specialist may produce policy conditioning, but it cannot issue motor commands. The Authority Plane grants AUGMENT only when task identity, memory identity, applicability, independently established utility, runtime safety and frozen specialist hashes all agree. The default is BYPASS.

Memory can hurt

The Authority Plane is motivated by a retained negative result, not by a post-hoc software story.

Generic Hybrid smoke Success Paired effect vs control
Control 9/30 = 30.0% —
Memory rerank 5/30 = 16.7% 1 rescue / 5 harms
Memory fallback 5/30 = 16.7% 1 rescue / 5 harms

Always-on generic memory injection was terminated. PatternLock is the only specialist qualified in AuthorityRegistryV1; the other 15 RoboMME tasks use the frozen base portfolio through BYPASS.

Evidence highlights

  • Memory content and order matter: on a generated hard PatternLock holdout, Correct Memory completed 18/24, while Reversed and Wrong-session Memory each completed 0/24.
  • The qualified specialist remains strong: PatternLock public official-test completed 140/150 across seeds 7, 42 and 123; episode-majority was 46/50 with Wilson 95% interval 81.16%–96.85%.
  • The frozen recipe now regenerates the task: a no-network, read-only, non-root same-host GPU Docker rehearsal completed 150/150 scheduled branches at 142/150; episode-majority was 49/50. It agreed with the original outcome on 146/150 paired branch keys. This is R2 evidence, not external clean-room.
  • Entity/spatial memory is a second system case: VideoUnmask official-val completed 138/180 with Correct Memory versus 0/180 with No Memory; its local public-test result was 117/150. In a later prospective competitive authority gate, Auto8 tied MemER at 66/90 with 18 rescues and 18 harms (one-sided p=.566), so it correctly remains outside registry v1.
  • Portfolio diversity matters: the frozen task-majority portfolio exceeded matched FrameSamp by 6.08 percentage points, clustered bootstrap 95% interval +3.29 to +8.92 points.
  • The end-to-end Authority system now runs across Full16: Attempt 5 completed all 2400 branches at 1288/2400 with zero retry, replacement, backfill, technical error, or prior-attempt import. It exceeds matched frozen FrameSamp by 8.71 points (cluster 95% interval +5.71 to +11.71) and is within 0.21 points of the pre-execution composition projection (interval -1.17 to +1.58). The latter comparison contains 142 gains and 137 harms, so aggregate calibration must not be mistaken for branch replay.

The original v1 values are recomputed by the CPU-only evidence audit. The additive Attempt 5 auditor independently verifies the complete matrix, denominators, hashes, and claim boundary against the frozen analysis implementation. The VideoUnmask qualification auditor independently recomputes all 180 terminals and binds the non-promotion decision. The original v1 ledger remains immutable; the Attempt 5 final audit and analysis bind the later integrated evidence.

Verify the science without a GPU

python -m pip install -e .
python -m rosclaw_memory.repro.audit_all

Expected terminal status:

"status": "PASS"

The canonical machine-readable experiment ledger is artifacts/evidence/experiment_ledger.json. The human-readable index is docs/evidence/EXPERIMENT_LEDGER.md.

Reproduce PatternLock

CPU evidence replay:

./repro/bootstrap.sh
./repro/run.sh patternlock --evidence-only

GPU execution additionally requires the exact RoboMME checkout and MME-VLA checkpoint described in release/evidence-v1/MANIFEST.json and assets/model_manifest.json:

export ROSCLAW_PATTERNLOCK_CHECKPOINT=/path/to/checkpoint/79999
make verify-assets
./repro/run.sh patternlock --execute

The completed canonical same-host R2 result and validation are under artifacts/repro/local_gpu_container_rehearsal_v1. The full technical report is ROSCLAW_MEMORY_R2_GPU_REHEARSAL_V1_20260825.html.

The current v0.1.0rc1 package does not claim a clean-room pass. A final v0.1.0 release requires a new clone, independent host/operator, exact code and asset hashes, 150 scheduled branches, zero retry/replacement, and a signed REPRODUCIBILITY_CERTIFICATE.json.

Package boundary

The distribution name is rosclaw-memory and its Python package is rosclaw_memory. PowerMem is an optional integration provider; ROSClaw-Memory does not publish the powermem namespace. Historical PowerMem-derived embodied infrastructure remains in this repository for license-preserving migration and is documented in NOTICE.

  • Core runtime research: evidence contracts, applicability, authority, specialists and PatchProof.
  • Supporting embodied infrastructure: spatial, temporal, object, trajectory and PowerMem integration.
  • Frozen reproduction track: exact evidence, source identities, protocols, hashes and execution adapters under release/ and repro/.

Documentation

License and citation

Apache-2.0. See LICENSE, NOTICE, and CITATION.cff.

About

Embodied intelligence memory for physical AI robots. Built on PowerMem + SeekDB. Zero ROS dependency. Brain-like architecture with spatial-temporal-causal graphs, DTW trajectory search, and gRPC.

Topics

Resources

Stars

7 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages