Repository navigation
spike: recovery Stage 3 - end-to-end experimental flow - #206
Closed
willemneal wants to merge 9 commits into
Closed
willemneal wants to merge 9 commits into
willemneal wants to merge 9 commits into
Conversation
Stage 3 of the staged recovery plan (firstmate/data/perch-zk-recovery-scout-p5/follow-up.md §8): contracts/recovery-controller (shared controller implementing guardian-only, ZK-only, and combined evidence paths against Stage 1's proposal model, completing via Stage 2's Variant A) and contracts/recovery-verifier (a new, fully constructorless UltraHonk verifier). circuits/zk_recovery_doc adapts the existing zk_recovery Noir circuit to bind a target-document hash instead of a raw pubkey, as a NEW, isolated circuit crate (not an in-place edit — the M1 circuits/zk_recovery module and its 33+ integration tests are byte-for-byte untouched and re-verified green). All three modes validated end-to-end against real contracts, including a real bb-proved UltraHonk proof verified on-chain (recovery_stage3_zk_only.rs, recovery_stage3_combined.rs) and the full guardian-only lifecycle through a real apply_doc completion (recovery_stage3_guardian_only.rs), matching Stage 2's Variant A call-ordering proof. See contracts/recovery-controller/src/lib.rs's crate doc comment for the architecture and the canonical "Known limits" list, and docs/recovery/stage3-measurements.md for proof generation/verification measurements. just test and just check both green. Client/SDK work is in progress in a follow-up commit.
The matrix wasn't kept in sync when these two Stage 3 crates were added (same maintenance gap the justfile's fmt-pkgs list already had for recovery-doc-completion, fixed in the previous commit).
packages/passkey-sdk/src/recoveryStage3/: enrollment config builders, target-document construction + diff (lost-key vs compromise, per follow-up.md §4.2), attempt/evidence builders for all three modes, read wrappers, and a JS reimplementation of the Rust contract's compute_doc_auth_hash (parity-tested against the same pinned zk.rs fixture, cross-validating the whole ZK path end to end). scripts/generate-recovery-proof.mjs: a Node CLI shelling out to nargo/bb for proof generation (no in-browser/mobile proving — a named limit, see the crate doc comment) — verified for real against the pinned circuits/zk_recovery_doc fixture: computed root/nullifier/auth_hash and VK sha256 both matched the committed Rust-side values exactly. packages/frontend/src/pages/security/recover-v3/: a plain, linear experimental page (enroll -> status -> begin attempt w/ diff preview -> evidence -> complete), reusing existing signing infrastructure (signAndSubmit, walletConnect.ts's kit) rather than inventing a new flow. Does not touch or modify the existing /security/recover (M1) page. Contract bindings generated for recovery-controller/recovery-verifier. Verified: tsc clean, 257/257 passkey-sdk tests, npm run build + astro check clean for the frontend.
|
Preview deployed! Account URLs use numeric preview suffixes, for example |
|
Example dApp preview deployed! https://example-pr-206.mysoroban.pages.dev The |
…_hash commitment Captain live-tested PR 206 and found a real testnet account's Enroll click was a no-op: `RecoveryController::enroll` only writes the controller's own storage — nothing ever checked or established the account's own `recovery_controller` field, so an account already wired to a different controller (or never wired at all) got orphaned, never-cross-called config. Adds `packages/passkey-sdk/src/recoveryStage3/accountWiring.ts` (checkAccountWiring/buildWireAccountTx) and wires it into the recover-v3 page: Enroll is disabled until wiring is confirmed, a "wire account" action handles the fresh-account case, and a different-controller mismatch blocks with an explanation instead of writing dead state. Also investigated follow-up.md §5.5's "reviewable configuration commitment in the doc" ask and confirmed it's unreachable today: perch's schema is strict, nido's lowering throws for the one principal shape that could fit, and the deployed/pinned perch-doc-compiler's wire-level CompiledRule has no field for an arbitrary policy address at all. Added RecoveryController::config_hash (sha256(xdr(RecoveryConfig)), on-chain, recomputable) as an equivalent, independently verifiable substitute, and documented the doc-embedding finding in the crate's Known Limits.
…-wired to M1 Added a live-testnet Playwright probe (tests/e2e/testnet/recover-v3-wiring.testnet.spec.ts) for the account-wiring fix. Running it surfaced something more specific than the wiring check was written to handle: a BRAND NEW account from the doc-only factory is not "fresh and unwired" — it already reports recovery_controller() == CAUZ6WFU... (the M1 nido-zk-recovery pool) at construction, per DEPLOYED.md's M2 genesis-insert behavior. So the captain's bug wasn't a one-off misconfiguration on his test account; it's the universal starting state for every account this factory has ever minted, and the only way into this Stage 3 controller today is the real 7-day rule-removal migration. Updated recovery-controller's crate doc comment and accountWiring.ts's module doc comment to state this as a confirmed, live-verified fact instead of a hypothetical case. The probe itself asserts the mismatch-detection path (the part that IS live-reachable): checkAccountWiring correctly reports 'wired-to-different-controller', Enroll stays disabled, and even a forced click is refused by the handler's own defensive re-check before any transaction is built.
Regenerating bindings (needed to add config_hash to the TS client) with a newer stellar-cli emitted a package.json missing the @nidohq scope (name: "recovery-controller" instead of "@nidohq/recovery-controller"), downgraded version 0.1.0 -> 0.0.0, and dropped publishConfig. npm workspaces no longer recognized it as satisfying passkey-sdk's "@nidohq/recovery-controller": "^0.1.0" dependency, so any fresh `npm install` (the PR-preview deploy jobs) tried the public registry and 404'd. Restored the fields to match every sibling bindings package.
…sh in measurements Adds two entries to stage3-measurements.md's limits/enrollment-data sections mirroring what the account-wiring fix and its live probe established: (1) account wiring is a separate, currently-unreachable-without-a-7-day-migration precondition from enroll, and (2) config_hash is the on-chain substitute for follow-up.md §5.5's reviewable-commitment ask, since literal doc embedding is blocked by the deployed perch-doc-compiler's wire protocol.
…egacy stub)
Second captain live-fail on PR 206: installing 1-of-1 friend recovery on the
REAL /security/ page (not the recover-v3 spike) failed with
"multisig-recovery.buildInstall: doc-only: the account has no rule
mutators; M-of-N friend recovery is not yet expressible as a policy
document" — a stale, inaccurate error from the 201 rework. buildInstall
threw unconditionally for every account regardless of state; the stub's
"doc v1 has all-signers principals only" claim was also stale (threshold
has existed since perch 0.2.0).
Rewrites multisigRecoveryModule to route through Stage 3's
RecoveryController (GuardianOnly mode) instead of the dead stub:
- buildInstall checks account wiring first (checkAccountWiring): wires
then enrolls a fresh account (two ops, one InvokeHostFunction each,
submitted sequentially), just enrolls an already-wired account, and
REFUSES with an accurate error (naming the real 7-day
initiate_recovery_rule_removal -> execute_recovery_rule_removal
constraint, not the false doc-schema claim) for an account already wired
to a different controller.
- Two real, documented implementer choices where the config has no
scriptable default: baseline_doc_hash is an inert sentinel hash (this
simplified form never exposes Compromise-mode recovery, the only case
that field is checked against); pending_activity_policy defaults to
Freeze, the more conservative of the two options TRANSITION_SPEC.md §10 /
follow-up.md §7 explicitly forbid a SPEC-level default for (this is an
implementer default at the UI layer, not a silently-chosen spec default).
- buildRevoke now explains the real constraint (no one-step revoke exists;
needs the same 7-day migration) instead of throwing the stale doc-only
message.
- fromChain recognizes BOTH the legacy multisig-policy on-chain-signers
shape (unchanged, so an already-installed legacy rule doesn't lose its
only Revoke path from the UI) and the new Stage 3 shape (guardians from
PolicyState, not on-chain signers — Stage 3's rule is zero-signer
CallContract(self)).
- policyChainFetch.ts's fetchPolicyState gets a branch for the Stage 3
controller address, reading RecoveryController::config and shaping it to
{guardians, threshold} for fromChain.
- New packages/passkey-sdk/src/recoveryStage3/deployment.ts holds the
canonical testnet controller/verifier addresses (not registry-resolvable
yet); recorded in DEPLOYED.md with provenance/verification notes.
Live-probed end to end against real testnet
(tests/e2e/testnet/security-recovery-install.testnet.spec.ts): a genuinely
fresh, unwired account (raw-deployed, bypassing the doc-only factory's
universal M1 pre-wiring found while fixing captain issue #1) completes
"Set up recovery" for 1 of 1 friend through the real production form, and
reloading /security/ renders "1 of 1 friend can rotate this account's
signers and rules" — confirming the full wire -> enroll -> render round
trip. Independently confirmed on-chain via `stellar contract invoke
... config` / `recovery_controller` (both match exactly).
just check (fmt + clippy -D pedantic) green; passkey-sdk: tsc clean,
263/263 vitest; astro check + npm run build clean.
Third captain live-fail on PR 206: enrolling ZK recovery first (via /security/'s "Add ZK recovery") wired the account directly to the M1 nido-zk-recovery pool -- a DIFFERENT controller from the guardian flow's Stage 3 RecoveryController -- so adding guardian recovery second always hit the wired-to-different-controller refusal shipped in the previous fix, and vice versa. On this stack ZK and guardians must be able to coexist on ONE controller (AuthMode::Combined) regardless of which is added first. A. New RecoveryController::reconfigure entry point (removes the "no reconfigure entry point" Known Limit -- lib.rs's crate doc comment explains what replaced it and its remaining bound). Accepts only two strictly-additive transitions (GuardianOnly -> Combined, ZkOnly -> Combined); every other field must match the stored config exactly or it refuses; blocked while has_pending; Profile::Loss needs only account.require_auth(), Profile::Protected + existing GuardianOnly needs the enrolled guardian quorum's nested require_auth_for_args in the same transaction, Profile::Protected + existing ZkOnly explicitly refuses (ReconfigureZkEvidenceUnsupported -- a real ZK reconfigure-evidence path needs a new circuit auth_hash domain, out of scope here; reusing an existing domain was considered and rejected as a cross-domain replay hazard). 12 new unit tests cover both transitions, all rejection cases, and both profiles. Deployed as a NEW testnet instance (v2) -- the constructorless controller has no upgrade/admin entry point at all, so the existing deployed instance could not gain reconfigure in place. See deployment.ts/DEPLOYED.md for the v1 -> v2 addresses and the "explicit upgrade = new immutable artifact" rationale. B. Killed the factory's genesis trap: contracts/factory::create_account/ create_account_v2 no longer unconditionally wire every new account to the M1 pool (recovery_controller: Some(M1), atomic genesis Merkle-leaf insert) -- a live probe already proved this made it impossible for any fresh account to ever reach Stage 3 without a real 7-day migration. Accounts now mint with recovery_controller: None; create_account_v2's commitment argument is ignored (kept for ABI compatibility). Redeployed the live factory in place (same address, upgrade + refresh_account_wasm_hash) and live-verified a freshly minted account reads recovery_controller() == null. C. Client flows route through Stage 3 consistently: security/index.astro's runZkEnrollment and new-account/index.astro's enrollRecoveryPostCreate both now check wiring (checkAccountWiring), wire if needed, and enroll-or-reconfigure based on RecoveryController::config(account)'s actual presence/mode -- instead of enroll_zk_recovery(M1 pool) directly. multisigRecoveryModule.buildInstall gained the mirror case: reconfigure (adding guardians) when ZK was already enrolled on the target controller. Shared enrollment defaults (baseline-doc sentinel, pending-activity policy, delay/expiry/max-cancels) extracted to defaults.ts so both flows agree on every field reconfigure requires to match exactly. D. Live-probed both orderings against the real production /security/ forms (tests/e2e/testnet/recovery-stage3-combined.testnet.spec.ts) -- caught and fixed a real bug in the process: policyChainFetch.ts's fetchRecoveryControllerState only returned guardian data for GuardianOnly mode, so a Combined-mode account's "N of M friends can rotate..." block silently vanished from the Security page even though the on-chain config was correct. Both accounts independently confirmed via RecoveryController::config() to reach identical Combined state (same guardians/threshold/verifier/zk_pool) regardless of order. just check / just test green (workspace); tsc/vitest/astro check/npm build green (client).
Contributor
Author
|
Superseded by #207, which merges this branch's work together with 200/201/202/204/205/206 into one reconciled, non-draft PR off main. This branch stays on origin for history. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Scope
Stage 3 of the staged recovery plan (
firstmate/data/perch-zk-recovery-scout-p5/follow-up.md§8, "Complete the experimental Nido flow"), building on Stage 1's transition
spec (#204) and Stage 2's completion-mechanism recommendation (#205 —
Variant A, no blocker recorded, adopted here per the captain's Stage 3
authorization).
This is the largest stage. Per the brief: kept honest rather than polished —
every simplification is named below and in
contracts/recovery-controller/src/lib.rs'scrate doc comment (the canonical, most detailed list). This is not
authorization to deploy Protected production accounts while follow-up.md §7
remains open.
What this ships
contracts/recovery-controller— a shared controller (one deployedinstance, many accounts) implementing guardian-only, ZK-only, and
combined evidence paths against the SAME proposal-commitment model
(
docs/recovery/TRANSITION_SPEC.md, Stage 1). Completion is Variant A —the controller's
Policy::enforcegates the account's EXISTINGapply_doc, zero smart-account code changes, same call-orderingcorrectness argument Stage 2 already proved.
GuardianOnlyenrollmentrequires no ZK machinery at all (enforced at
enroll, not justdocumented — follow-up.md §5.2 HARD requirement).
Combinedchecks bothfactors against the identical frozen commitment before promotion.
Cancellation uses its OWN action domain and per-attempt tally, separate
storage from initiation (follow-up.md §2.2).
contracts/recovery-verifier— a NEW, fully constructorless UltraHonkverifier: the VK is baked into the Wasm at compile time via
include_bytes!, no constructor, no admin, no upgrade entry point atall. A circuit/VK change means a new Wasm deploy (new address), never a
mutation — follow-up.md §5.4's "immutable until an explicit upgrade,"
interpreted as strongly as the constructorless-registry pattern allows.
circuits/zk_recovery_doc— a real, adapted Noir circuit. Starts fromcircuits/zk_recovery(the pre-existing M1 raw-signer-rotation circuit)and swaps its 5 raw-pubkey fields (
pk_prefix/pk_x_hi/lo/pk_y_hi/lo)for 5 new fields (
doc_hash_hi/lo,cfg_version,baseline_hi/lo) inthe
auth_hashbinding — same arity-15 Poseidon2 sponge, sameMerkle/nullifier logic untouched, real
bb provefixtures.This is a NEW, isolated circuit crate, not an in-place edit of
circuits/zk_recovery. The in-place edit was the first approach takenand reverted after discovering it would break the M1 module's own
still-referenced integration tests (
zk_recovery_lifecycle.rsand 5siblings, 33 tests total), which pin real
bb-proved fixtures againstthe OLD
auth_hashformula via sharedcrates/integration-tests/src/zk_fixture.rs. Isolating Stage 3's circuitinto its own crate (and its own Poseidon2 host-hash reconstruction in
contracts/recovery-controller/src/zk.rs, NOT depending onnido-zk-recoveryas a library) means this PR touches ZERO bytes of theM1 module or its tests —
circuits/zk_recovery/,contracts/zk-recovery/,and every M1 integration test are unchanged and re-verified green after
the revert.
Client/SDK (
packages/passkey-sdk/src/recoveryStage3/) — enrollmentconfig builders, target-document construction + diff (lost-key vs
compromise, per follow-up.md §4.2), attempt/evidence builders for all
three modes, read wrappers, and
docAuthHash.ts, a JS reimplementationof the Rust contract's
compute_doc_auth_hash— parity-tested againstthe SAME pinned
zk.rsfixture, cross-validating the JS↔Rust↔circuitpath.
scripts/generate-recovery-proof.mjsis a Node CLI shelling out tonargo/bbfor proof generation (no in-browser/mobile proving — a namedlimit). It was run for real against the pinned circuit fixture during
development: the computed root/nullifier/auth_hash AND the freshly
generated VK's sha256 both matched the committed Rust-side values
exactly — strong end-to-end correctness evidence spanning
client→CLI→circuit→contract.
packages/frontend/src/pages/security/recover-v3/is a plain, linear experimental page (enroll → status → begin attempt w/
diff preview → evidence → complete) reusing existing signing
infrastructure (
signAndSubmit,walletConnect.ts's kit) rather thaninventing a new flow; it does not touch the existing
/security/recover(M1) page. Verified:
tscclean, 257/257 passkey-sdk tests pass(including the new parity test),
npm run build+astro checkclean.docs/recovery/stage3-measurements.md— proof generation (~68msbb prove, 6,976-byte proof, on an Apple M5 Max dev machine — explicitlyNOT a realistic replacement device), on-chain verification (179.3M CPU
instructions, ~221M headroom under the 400M mainnet tx ceiling),
restoration behavior, enrollment-data availability, and every limit named
again with pointers.
Explicit limits (see
contracts/recovery-controller/src/lib.rsfor thecanonical, more detailed list — this is a summary)
reconfigureentry point. Enrollment is one-shot; baseline/mode/guardian-set changes require a fresh account.
replaced_credential_idsis client-declared, not on-chain-verifiedagainst the target document — the contract's only real cryptographic
guarantee is
target_doc_hashexactness; content correctness (does thetarget really equal baseline+replacements) depends on evidence providers
reviewing the client's target-document preview before approving.
has_pendinggatesapply_docandthe signer/rule/policy-removal entry points (the pre-existing guard set),
but NOT
execute()(arbitrary contract calls), confirmed by reading thatentry point directly (bare
require_auth+invoke_contract, no guard).add_context_rulehas its own "completion window" check(
has_pending() || completion_granted(), predates Stage 2/3) that admitsANY ordinarily-admin-authorized call — not just a doc-hash-bound
completion — to install an arbitrary new
ContextRulewhenever across-called controller's
has_pending()istrue. This controller'sFreeze-policyhas_pendingmakes that condition true exactly likeStage 2's controller already does. Variant A's own binding only
constrains completions through
apply_doc; it does nothing to close thisseparate, inherited vehicle. Flagged per follow-up.md §3.1's own warning
about exactly this class of gap — not fixed here (would need a
smart-account change, out of scope for "zero smart-account changes").
PendingActivityPolicy::Restrictexists in the type (mirroringTRANSITION_SPEC.md's three-way enum) but is refused atenroll—follow-up.md §7's "no default" made concrete as a real, tested refusal.
policyWriteConflictPolicy(Stage 1's separate invalidate-attempt-vs-blockaxis) is not implemented as its own axis at all.
this is a HARDER limit than first described. This bullet originally
assumed new accounts are created with
recovery_controller: Noneandjust need a post-hoc
enroll_zk_recovery(controller)call. A captainlive-test plus a live testnet probe proved that's wrong for every account
the doc-only factory actually mints today: they're ALREADY wired to the
M1
nido-zk-recoverypool at construction (DEPLOYED.md's M2genesis-insert behavior), so
enroll_zk_recoveryagainst THIS controlleris unreachable without first running the account's own real 7-day
initiate_recovery_rule_removal→execute_recovery_rule_removalmigration. Making this controller the factory's resolved default is a
deploy-time registry governance action requiring testnet/mainnet registry
author keys — named here per the brief's "report blocked naming exactly
what's needed," but NOT a blocker for this experiment (not needed to
complete Stage 3's scope).
recover-v3requirespasting controller/verifier/pool addresses and current/baseline doc JSON
by hand (no registry lookup or on-chain doc auto-fetch) — a time-boxed
scope cut; the underlying SDK functions are fully built and tested.
recoverStage3Actions.tssubmission helpers are classic-tx only — notwired to the Channels relayer, and assume the fee-payer is already funded.
— not a "realistic replacement device" claim.
Test plan
just test(full workspace, 119 passed / 0 failed / 3 pre-existingignored) green
just check(fmt + clippy -D pedantic) greennargo testin bothcircuits/zk_recovery(untouched, still 5/5green) and
circuits/zk_recovery_doc(new, 5/5 green)complete via real
apply_doc→ second attempt → expiry →cancellation (own action domain) → inactive-account restore →
atomicity on failed install → ordinary-admin-cannot-complete — all 8
tests in
recovery_stage3_guardian_only.rsbb proveUltraHonk proof verified on-chain throughpromotion (tampered-proof / wrong-target-doc-hash / unknown-root all
independently rejected) —
recovery_stage3_zk_only.rs, 4/4alone promotes — both required —
recovery_stage3_combined.rs, 2/2revert (
zk_recovery_*33/33,multisig_recovery3/3)tscclean, 257/257 vitest, including a JS↔Rustcompute_doc_auth_hashparity testroot/nullifier/auth_hash and VK sha256 both matched the committed
Rust-side fixture exactly
npm run build+astro checkcleanevery touched contract crate, Unit (TestAuthenticator), E2E (fast UI
+ CDP), and Dependency audit all green
Fix: account-wiring bug found by a captain live-test
The captain live-tested this PR: created a test account and reported the
contract at
CCI2HO73IVWPVYFPICNWMEJVF7ZZF36ZDI5HJ7WSEOWXZYIPCG2XOUH5(testnet) "doesn't seem to use a policy doc at all."
Investigation (read-only, direct contract introspection):
NidoSmartAccountfrom the current doc-onlyfactory (wasm hash
fe3b1878…, matches the expected build). Not a stalefactory, not a mistaken address.
get_applied_doc()/applied_doc_hash()are bothnull— correctand expected: the doc-only factory never auto-applies a Perch doc at
construction, only raw context rules. Not a bug.
recovery_controller()isCAUZ6WFUTTZCJQNNL5D3BNZSG7FYYGX46BDJE6G2XVVCGN76RKE5ESAR— the OLD M1
nido-zk-recoverypool, not this PR'sRecoveryController.RecoveryController::enrollonly ever writes the CONTROLLER's ownstorage; nothing checked or established the ACCOUNT's own
recovery_controllerfield before letting "Enroll" run. ClickingEnroll against this PR's controller silently wrote config nobody would
ever cross-call — a complete, silent no-op with respect to actual
on-chain protection. This was the real bug, not a doc/creation-path
issue.
(
tests/e2e/testnet/recover-v3-wiring.testnet.spec.ts) created aBRAND NEW account via the doc-only factory and found it was already
wired to that same M1 pool at construction — this is DEPLOYED.md's own
documented M2 genesis-insert behavior ("every account this factory
creates ... installs the recovery rule ... whether or not its owner ever
uses recovery"). So the captain's account wasn't a one-off
misconfiguration; it's the universal starting state of every account
this factory has ever minted. The only way into this controller today is
the account's own real 7-day
initiate_recovery_rule_removal→execute_recovery_rule_removalmigration — there is no faster path, bydesign (that delay is what keeps the removal safe against a
stolen-key attacker racing to strip recovery protection).
Fix:
packages/passkey-sdk/src/recoveryStage3/accountWiring.ts(new):checkAccountWiringreads the account's ownrecovery_controller()andclassifies it as
'wired-to-target'/'unwired'/'wired-to-different-controller';buildWireAccountTxhandles theone-shot fresh-account case via
enroll_zk_recovery.recover-v3now calls this before Enroll: the Enroll button startsdisabled, a "Wire account" action appears for the unwired case, and a
different-controller mismatch (the captain's exact scenario) blocks with
an explanation instead of a silent no-op. The Enroll click handler ALSO
re-checks wiring itself (defense in depth beyond the disabled attribute)
— proven live in the probe by force-enabling the button and clicking
anyway.
RecoveryController::config_hash(account) -> Option<BytesN<32>>(new,sha256(xdr(RecoveryConfig)), on-chain, deterministic) — investigatedwhether recovery config could instead be embedded in the account's own
Perch policy document (follow-up.md §5.5's "reviewable configuration ...
with an accurate commitment") and confirmed that's unreachable today, not
merely unimplemented: perch's schema is
.strict()(no extensionfields), nido's own doc-lowering throws for the one principal shape that
could fit, and — the decisive blocker — the deployed, pinned
perch-doc-compiler's wire-levelCompiledRuletype has no field for anarbitrary policy address at all.
config_hashis an equivalent,independently-verifiable substitute; surfaced in the recover-v3 Status
panel via
readConfigHash. Full reasoning incontracts/recovery-controller/src/lib.rs's "Known limits".packages/contract-bindings/recovery-controller/package.json: thebindings regen needed to add
config_hashto the TS client used a newerstellar-clithat emitted a package.json missing the@nidohqscope andpublishConfig— broke npm workspace resolution for every job that runsa fresh
npm install(PR preview deploys, TestAuthenticator unit tests,fast E2E). Restored to match every sibling bindings package.
Live probe evidence (
tests/e2e/testnet/recover-v3-wiring.testnet.spec.ts,real testnet, real relayer-sponsored passkey signing,
@testnettier):CB5YS4DZD6EZDB74DDHP6PBVT6KMVJILVCVZEPUB27CVK2FI2JG26ALBRecoveryControllerdeploy (this PR's contract, throwaway instance,deliberately not added to DEPLOYED.md):
CDHB5B3GI63EQKPLSQWB6OKBZKMBLRAA3SHOBCYAPTDZ6ZN3YLADTHL6checkAccountWiring→'wired-to-different-controller',currentControllerId = CAUZ6WFUTTZCJQNNL5D3BNZSG7FYYGX46BDJE6G2XVVCGN76RKE5ESAR(independently confirmed via
stellar contract invoke ... recovery_controller)anyway is refused by the handler's own re-check ("inert no-op") before any
transaction is built or signed
config()andconfig_hash()on the probe controller are both
nullfor this account — no orphanedstate was written
Fix 2: legacy friend-recovery stub still reachable (captain live-fail #2)
A second captain live-test on this PR: installing 1-of-1 friend recovery on
the REAL
/security/page (not therecover-v3spike) failed withmultisig-recovery.buildInstall: doc-only: the account has no rule mutators; M-of-N friend recovery is not yet expressible as a policy document (doc v1 has all-signers principals only)— a stale error from the201 rework.
buildInstallthrew unconditionally for every accountregardless of state; the "doc v1 has all-signers principals only" claim was
also stale (threshold has existed since perch 0.2.0 — the real historical
gap was stateful timelock, not principal expressiveness).
Fix:
multisigRecoveryModulenow routes through Stage 3'sRecoveryController(GuardianOnlymode) instead of the dead stub:buildInstallchecks account wiring first (checkAccountWiring): wiresthen enrolls a fresh account (two sequential ops — Soroban allows one
InvokeHostFunctionper tx), just enrolls an already-wired account, andrefuses with an accurate error (naming the real 7-day
initiate_recovery_rule_removal→execute_recovery_rule_removalconstraint) for an account already wired to a different controller —
instead of the old false doc-schema claim.
value but nothing mandates a default:
baseline_doc_hashis an inertsentinel hash (this simplified form never exposes
Compromise-moderecovery — the only case that field is checked against);
pending_activity_policydefaults toFreeze, the more conservative ofthe two options
TRANSITION_SPEC.md§10 / follow-up.md §7 explicitlyforbid a spec-level default for. This is an implementer default at the
UI layer (the type requires SOME explicit value), not a silently-chosen
spec default — documented prominently in
multisigRecovery.tsand opento being overridden by an explicit product decision.
buildRevokenow explains the real constraint (no one-step revokeexists) instead of throwing the stale doc-only message.
fromChainrecognizes BOTH the legacy multisig-policy on-chain-signersshape (unchanged — an already-installed legacy rule keeps its only
Revoke path in the UI) and the new Stage 3 shape (guardians read from
PolicyState, not on-chain signers, since Stage 3's rule is zero-signerCallContract(self)).policyChainFetch.ts'sfetchPolicyStategets a branch for the Stage 3controller address, reading
RecoveryController::configand shaping itto
{guardians, threshold}.packages/passkey-sdk/src/recoveryStage3/deployment.tsholds thecanonical testnet controller/verifier addresses (not registry-resolvable
yet — see "Known limits"); recorded in
DEPLOYED.mdwith provenancenotes (interface-checked against
contracts/recovery-controller's ownstellar contract info interfaceoutput before being recorded).Live-probed end to end against real testnet
(
tests/e2e/testnet/security-recovery-install.testnet.spec.ts): agenuinely fresh, unwired account (raw-deployed directly against the
smart-account wasm, bypassing the doc-only factory — which, per Fix 1's
finding, pre-wires every account to the M1 pool at genesis, so it cannot
produce an unwired account at all) completes "Set up recovery" for 1 of 1
friend through the real production form, and reloading
/security/renders"1 of 1 friend can rotate this account's signers and rules" — confirming
the full wire → enroll → render round trip. Independently confirmed
on-chain via
stellar contract invoke ... config/recovery_controller(both match exactly what the UI built).
Still true for essentially every REAL/existing account (per Fix 1's
universal-pre-wiring finding): reaching Stage 3 recovery requires the real
7-day rule-removal migration first. This fix makes the UI behave correctly
and honestly in that state (clear refusal, not a silent no-op or a false
error) rather than making it possible to skip that migration — no such
shortcut exists, by design.
Fix 3: ZK and guardian recovery couldn't coexist (captain live-fail #3)
A third captain live-test: enrolled ZK recovery first (via
/security/'s"Add ZK recovery"), then tried to add 1-of-1 friend recovery — refused with
the
wired-to-different-controllererror Fix 1 shipped. Diagnosis: ZKenrollment wired the account directly to the OLD M1
nido-zk-recoverypool, a different controller from the guardian flow's Stage 3
RecoveryController. On this stack ZK and guardians must be able tocoexist on ONE controller regardless of which is added first — that's
literally what
AuthMode::Combinedalready existed to express.A. New
RecoveryController::reconfigureentry point — removes the "noreconfigure entry point" Known Limit (
lib.rs's crate doc comment explainswhat replaced it and its remaining bound). Accepts only two strictly
additive transitions (
GuardianOnly -> Combined,ZkOnly -> Combined);every other field must match the stored config exactly or it refuses;
blocked while
has_pending. Authorization:Profile::Lossneeds onlyaccount.require_auth();Profile::Protected+ existingGuardianOnlyneeds the enrolled guardian quorum's nested
require_auth_for_argsin thesame transaction;
Profile::Protected+ existingZkOnlyexplicitlyrefuses (
ReconfigureZkEvidenceUnsupported) — a real ZKreconfigure-evidence path needs a new circuit
auth_hashdomain (out ofscope); reusing an existing domain was considered and rejected as a
cross-domain replay hazard. 12 new unit tests cover both transition
directions, every rejection case, and both profiles.
Deployed as a new testnet instance (v2) — the constructorless
controller has no upgrade/admin entry point at all, so the existing
deployed instance couldn't gain
reconfigurein place. Seedeployment.ts/DEPLOYED.mdfor the v1 → v2 addresses and the "explicitupgrade = a new immutable artifact, not a rewrite" rationale (the same
philosophy already documented for the verifier, exercised here for the
controller itself for the first time).
B. Killed the factory's genesis trap.
contracts/factory'screate_account/create_account_v2no longer unconditionally wire everynew account to the M1 pool (
recovery_controller: Some(M1), atomic genesisMerkle-leaf insert) — Fix 1's live probe already proved this made it
impossible for any fresh account to ever reach Stage 3 without a real
7-day migration; the captain's bug wasn't a misconfiguration, it was the
universal starting state. Accounts now mint with
recovery_controller: None;create_account_v2'scommitmentargument is ignored (kept for ABIcompatibility with existing callers). Redeployed the live factory in
place (same address
CCJFOM6U…,upgrade+refresh_account_wasm_hash)and live-verified: a freshly minted account reads
recovery_controller() == null.C. Client flows route through Stage 3 consistently.
security/index.astro'srunZkEnrollmentandnew-account/index.astro'senrollRecoveryPostCreateboth now check wiring (checkAccountWiring),wire if needed, and enroll-or-
reconfigurebased onRecoveryController::config(account)'s actual presence/mode — instead ofenroll_zk_recovery(M1 pool)directly.multisigRecoveryModule.buildInstallgained the mirror case:
reconfigure(adding guardians) when ZK wasalready enrolled on the target controller. Shared enrollment defaults
(baseline-doc sentinel, pending-activity policy, delay/expiry/max-cancels)
extracted to
defaults.tsso both flows agree on every fieldreconfigurerequires to match exactly — otherwise whichever factor enrolls second would
always fail the field-match check.
D. Live-probed both orderings against the real production
/security/forms (
tests/e2e/testnet/recovery-stage3-combined.testnet.spec.ts) —caught and fixed a real bug in the process:
policyChainFetch.ts'sfetchRecoveryControllerStateonly returned guardian data forGuardianOnlymode, so aCombined-mode account's "N of M friends canrotate…" block silently vanished from the Security page even though the
on-chain config was correct (
Combinedhas guardians too — onlyZkOnlyhas none). Fixed and re-verified. Both accounts independently confirmed via
RecoveryController::config()to reach identicalCombinedstate (sameguardians/guardian_threshold/verifier/zk_pool) regardless of order.Live-probe accounts (throwaway testnet test accounts, not real users):
CBY2E6AA7D5PTGB5JXASYAHDRJI6LY62JFEKWWQ45J3B76FNE4HIZI6BCAWNROKYENDO6W2FPF4VIDO7CBUH6GOKVG2D72TGWEM3HTKNCA6NIDFZBoth independently read back via
stellar contract invoke ... configagainst the new controller (
CBYSWPHNWAHYUBZO5TBTO5MCW2ZC45F2C3L4JSUZXYQFNMHTOBOCCHZU):mode: "Combined",guardians: ["GAMPJROH…"],guardian_threshold: 1,verifier/zk_poolboth populated — identical on both accounts.