Environment
- memos-local-plugin 2.0.20 (shared-bridge runtime; systemd unit
memos-bridge.service runs dist/bridge.cjs --daemon)
- Single Hermes instance; hosts = hermes-gateway (many sessions), hermes-dashboard, TUI — which normally spawn
bridge.mjs --no-viewer clients against :18800
Summary
When the systemd-managed daemon is briefly unavailable (e.g. during a restart), a session host can rescue-spawn its own full daemon (bridge.mjs --agent=hermes --home=<plugin> --daemon). If that process binds :18800 before the systemd daemon finishes its ~8–10 s boot, the unit can never bind again: it crash-loops on port :18800 busy while the rogue process keeps serving. Critically, /api/v1/health stays green (answered by the rogue), so health-based monitoring reports the system as healthy. Unattended, this persists indefinitely — only stopping the spawning host clears it (killing the rogue alone respawns in ~3 s via the host's keepalive).
Observed (2026-10-07, this deployment)
- 11:22:57 — watchdog restarted the unit after a legitimate
down verdict (LLM-lane outage); recovery failed at 11:23:16.
- 11:23:31 — the gateway process (the rogue's parent PID, cgroup
hermes-gateway.service) spawned bridge.mjs ... --daemon; it bound :18800 in ~1 s while the systemd replacement was still initializing (it reached pipeline.ready at +8 s and only then saw the port busy).
- Result: 413 unit start attempts in ~24 h, a restart attempt every ~26 s,
port :18800 busy ×3,363 — while health answered ok:true throughout, and the watchdog's post-restart rogue sweep lost the race to the ~3 s respawn keepalive (it logged "recovery OK" every cycle).
Impact
- 400+ restarts/day of churn (CPU, logs, journal); the process actually serving sessions runs outside systemd supervision; every loop tick interrupts in-flight recovery/backfill work; health + monitoring are actively misleading; clearing it requires stopping the spawning host first (operator footgun).
Related
Suggested directions (not prescriptive)
- A host should not rescue-spawn a
--daemon when a managed daemon exists or is mid-start: check systemctl is-active memos-bridge / the singleton pidfile / probe :18800 with a short retry window before spawning.
- If rescue-spawning is intentional, make the rescued daemon yield to an in-progress systemd start (bind-retry with backoff) instead of racing for first-bind.
- Optional: when the port is already owned by another
bridge.* process, exit with a distinctive log line the supervisors can watch for (the unit currently logs port busy ... retrying (n/10) and then exits silently-ish at default log levels).
Verified local workaround (now standard ops here): stop unit → stop the spawning host (gateway) → sweep rogues → systemctl reset-failed memos-bridge; systemctl start memos-bridge (binds in ~2 s) → start host → verify the listener's cgroup is memos-bridge.service and no --daemon child reappears. Detection added locally: a 30-min guard alerts when the :18800 listener's cgroup ≠ memos-bridge.service or NRestarts climbs between runs.
Environment
memos-bridge.servicerunsdist/bridge.cjs --daemon)bridge.mjs --no-viewerclients against :18800Summary
When the systemd-managed daemon is briefly unavailable (e.g. during a restart), a session host can rescue-spawn its own full daemon (
bridge.mjs --agent=hermes --home=<plugin> --daemon). If that process binds :18800 before the systemd daemon finishes its ~8–10 s boot, the unit can never bind again: it crash-loops onport :18800 busywhile the rogue process keeps serving. Critically,/api/v1/healthstays green (answered by the rogue), so health-based monitoring reports the system as healthy. Unattended, this persists indefinitely — only stopping the spawning host clears it (killing the rogue alone respawns in ~3 s via the host's keepalive).Observed (2026-10-07, this deployment)
downverdict (LLM-lane outage); recovery failed at 11:23:16.hermes-gateway.service) spawnedbridge.mjs ... --daemon; it bound :18800 in ~1 s while the systemd replacement was still initializing (it reachedpipeline.readyat +8 s and only then saw the port busy).port :18800 busy×3,363 — while health answeredok:truethroughout, and the watchdog's post-restart rogue sweep lost the race to the ~3 s respawn keepalive (it logged "recovery OK" every cycle).Impact
Related
initialize()spawn leak), fix(hermes-adapter): widen _ACTIVE_CLIENTS key with owner_id to prevent intra-process bridge fight #2292 (closed — superseded by shared-bridge mode; though 2.0.20 still exhibits host daemon-spawn behavior under this trigger), [Bug] dirty-closed rescan misses freshly-marked episodes for a full cursor lap; interrupted drains lose phase-2 accounting #2459 (dirty-rescan cursor accounting).Suggested directions (not prescriptive)
--daemonwhen a managed daemon exists or is mid-start: checksystemctl is-active memos-bridge/ the singleton pidfile / probe :18800 with a short retry window before spawning.bridge.*process, exit with a distinctive log line the supervisors can watch for (the unit currently logsport busy ... retrying (n/10)and then exits silently-ish at default log levels).Verified local workaround (now standard ops here): stop unit → stop the spawning host (gateway) → sweep rogues →
systemctl reset-failed memos-bridge; systemctl start memos-bridge(binds in ~2 s) → start host → verify the listener's cgroup ismemos-bridge.serviceand no--daemonchild reappears. Detection added locally: a 30-min guard alerts when the :18800 listener's cgroup ≠memos-bridge.serviceorNRestartsclimbs between runs.