Skip to content

[Bug] Host rescue-spawns its own bridge daemon during a systemd-daemon restart window — permanent port war, health stays green (2.0.20) #2467

Description

@kiwipaulrob

Environment

  • memos-local-plugin 2.0.20 (shared-bridge runtime; systemd unit memos-bridge.service runs dist/bridge.cjs --daemon)
  • Single Hermes instance; hosts = hermes-gateway (many sessions), hermes-dashboard, TUI — which normally spawn bridge.mjs --no-viewer clients against :18800

Summary
When the systemd-managed daemon is briefly unavailable (e.g. during a restart), a session host can rescue-spawn its own full daemon (bridge.mjs --agent=hermes --home=<plugin> --daemon). If that process binds :18800 before the systemd daemon finishes its ~8–10 s boot, the unit can never bind again: it crash-loops on port :18800 busy while the rogue process keeps serving. Critically, /api/v1/health stays green (answered by the rogue), so health-based monitoring reports the system as healthy. Unattended, this persists indefinitely — only stopping the spawning host clears it (killing the rogue alone respawns in ~3 s via the host's keepalive).

Observed (2026-10-07, this deployment)

  • 11:22:57 — watchdog restarted the unit after a legitimate down verdict (LLM-lane outage); recovery failed at 11:23:16.
  • 11:23:31 — the gateway process (the rogue's parent PID, cgroup hermes-gateway.service) spawned bridge.mjs ... --daemon; it bound :18800 in ~1 s while the systemd replacement was still initializing (it reached pipeline.ready at +8 s and only then saw the port busy).
  • Result: 413 unit start attempts in ~24 h, a restart attempt every ~26 s, port :18800 busy ×3,363 — while health answered ok:true throughout, and the watchdog's post-restart rogue sweep lost the race to the ~3 s respawn keepalive (it logged "recovery OK" every cycle).

Impact

  • 400+ restarts/day of churn (CPU, logs, journal); the process actually serving sessions runs outside systemd supervision; every loop tick interrupts in-flight recovery/backfill work; health + monitoring are actively misleading; clearing it requires stopping the spawning host first (operator footgun).

Related

Suggested directions (not prescriptive)

  • A host should not rescue-spawn a --daemon when a managed daemon exists or is mid-start: check systemctl is-active memos-bridge / the singleton pidfile / probe :18800 with a short retry window before spawning.
  • If rescue-spawning is intentional, make the rescued daemon yield to an in-progress systemd start (bind-retry with backoff) instead of racing for first-bind.
  • Optional: when the port is already owned by another bridge.* process, exit with a distinctive log line the supervisors can watch for (the unit currently logs port busy ... retrying (n/10) and then exits silently-ish at default log levels).

Verified local workaround (now standard ops here): stop unit → stop the spawning host (gateway) → sweep rogues → systemctl reset-failed memos-bridge; systemctl start memos-bridge (binds in ~2 s) → start host → verify the listener's cgroup is memos-bridge.service and no --daemon child reappears. Detection added locally: a 30-min guard alerts when the :18800 listener's cgroup ≠ memos-bridge.service or NRestarts climbs between runs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

area:pluginOpenClaw & Hermesstatus:needs-triageNeeds initial triage | 需要初步判断 & 问题复现

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions