Skip to content

WAL grows unbounded (115 GB in 4.5 h, fills system drive): checkpoint starvation from accumulated per-conversation servers + 11 concurrent index workers on one repo #1083

Description

@9981H

Environment

  • codebase-memory-mcp 0.9.0 (Windows 11, installed via install.ps1)
  • Multiple concurrent MCP clients against the same repo: Claude Code sessions (spawned by Cursor) + Codex sessions (ChatGPT desktop app). Each conversation spawns its own codebase-memory-mcp.exe; over one work day 30+ server processes accumulated (clients keep running after their conversations are closed, so servers are never reaped).

What happened

The main DB ~/.cache/codebase-memory-mcp/<repo>.db is only 60 MB, but its -wal file grew from 0 to 115 GB in ~4.5 hours (created 10:47, 115 GB by 15:30 local) and completely filled the system drive (free space went 11.3 GB -> 7.8 GB while we watched, ~0.5 GB per 2 minutes).

Process snapshot at 15:32 showed, besides 20+ long-lived server processes, 11 concurrent cli --index-worker index_repository workers all indexing the same repo path at the same time.

Diagnosis

Classic SQLite WAL checkpoint starvation:

  • dozens of long-lived connections (one per abandoned conversation) hold read marks on the WAL, so wal_checkpoint can never truncate;
  • each full re-index (clients run index_repository at the end of every task) appends the whole dataset to the WAL again;
  • log only grows, never shrinks.

Confirmation: at 15:39 a checkpoint finally succeeded and the WAL collapsed from 115 GB to a few MB — so the 115 GB was pure un-checkpointed log, not data volume.

Probably related: #914 (orphan cbm process), #937 (watcher re-index write amplification).

Suggestions

  1. Set PRAGMA journal_size_limit and force wal_checkpoint(TRUNCATE) after every index run (and/or periodically), logging a warning when the WAL exceeds a threshold.
  2. Serialize index workers per DB with a file lock: if an index_repository is already running for the same repo, queue or refuse instead of running 11 full indexes concurrently.
  3. Defense in depth: the server should exit when its client stdio disconnects (EOF), so abandoned conversations do not accumulate connection holders that block checkpoints.
  4. Consider one shared server per DB (with a supervisor/broker) instead of one server per conversation.

Activity

  1. added
    bugSomething isn't working
    stability/performanceServer crashes, OOM, hangs, high CPU/memory
    priority/highNeeds near-term maintainer attention; high-impact bug, regression, safety issue, or release blocker.
    on Jul 14, 2026
  2. added this to the 0.9.1-rc milestone on Jul 14, 2026
  3. DeusData commented on Jul 14, 2026

    @DeusData
    Owner

    Thanks for the unusually strong operational report. This is a high-priority 0.9.1-rc stability bug because unbounded WAL growth can exhaust the system disk. Investigation should cover per-repository single-flight indexing, stale MCP server lifetime, long-lived read transactions blocking checkpoints, and duplicate workers across clients. Please preserve the database/WAL evidence if possible rather than deleting it while this is being diagnosed.

  4. added a commit that references this issue on Jul 16, 2026
  5. DeusData commented on Aug 14, 2026

    @DeusData
    Owner

    @9981H — filling someone's system drive is about the worst thing a tool can do to their machine, and I am sorry it was ours that did it.

    Your report carried its own diagnosis: a 60 MB database against a 115 GB -wal, 11 concurrent workers on one repository, 30+ abandoned servers holding read marks, and then the WAL collapsing to a few MB the moment a checkpoint finally succeeded. That last observation is what proves it was pure un-checkpointed log rather than data volume, and it saved a lot of guessing.

    journal_size_limit and checkpoint-starvation warnings have shipped since, but the parts you named as the actual fix — per-repository single-flight indexing, and servers that exit when their client's stdio closes — are the ones that stop this recurring, and I am not going to claim those are done. Tonight's v0.10.5 is being cut for the issues that make cbm impossible to operate, and unbounded disk consumption is at the front of that queue.

  6. jameswilson commented on Aug 18, 2026

    @jameswilson

    Adding a macOS/Codex data point to this issue.

    Environment

    • codebase-memory-mcp: v0.9.0
    • MCP client: Codex Desktop 26.721.81911
    • macOS 12.7.6 (21H1320), x86_64
    • MacBookPro16,4
    • 8 physical / 16 logical CPUs
    • 32 GiB RAM
    • Binary SHA-256: 04ee3048810c19099502adc8bb83039423f02f2553d17677892a7f03b924e01f
    • Repository: private Git repository, 8,657 indexed files

    Observed behavior

    With auto_index=true and the default watcher behavior, Codex accumulated 23 MCP server processes. Thirteen simultaneous index_repository workers then started full indexes of the same repository.

    Each index reported:

    pipeline.mode mode=parallel workers=16 files=8657
    

    Observed resource use:

    • 13 simultaneous full index jobs. (!!!)
    • Approximately 592–850% aggregate CPU
    • Approximately 6.3–7.2 GiB aggregate RSS
    • macOS load averages exceeded 800
    • Approximately 108,000 extracted graph nodes per competing job

    The canonical SQLite database disappeared and was replaced by a 143 MiB .db.corrupt file. PRAGMA quick_check confirmed genuine B-tree damage, including out-of-order rowids.

    Several concurrent workers later produced a new 164 MiB database. It was also genuinely malformed:

    2nd reference to page ...
    Page ... is never used
    database disk image is malformed (11)
    

    Disabling auto_index alone was insufficient because existing MCP sessions continued to hold watchers/connections and spawned additional workers.

    A manual CLI rebuild completed analysis with:

    expected_nodes=115030
    expected_edges=129003
    status=degraded
    hint="Index database failed integrity check and was removed."
    

    After setting both:

    auto_index=false
    auto_watch=false
    

    six older MCP servers ignored SIGTERM and immediately spawned replacement index workers. They had to be force-stopped.

    With all MCP server processes gone, one isolated manual rebuild succeeded:

    nodes=115030
    edges=128898
    status=indexed
    

    Post-rebuild verification:

    PRAGMA quick_check;
    ok
    

    A subsequent no-change manual update completed in 0.55 seconds and retained the same node/edge counts.

    Expected behavior

    • Only one index worker should operate on a given repository/database.
    • Additional clients should queue or receive a clear “index already in progress” result.
    • Closing or terminating an MCP client should stop its server and watcher.
    • Concurrent clients must not create genuinely malformed SQLite databases.
    • Disabling automatic indexing/watching should prevent existing helpers from spawning replacement workers.

    Notes

    The corrupt databases and worker logs were preserved. They contain metadata and paths from a private repository, so I can provide sanitized excerpts rather than uploading the databases publicly.

    I understand that v0.10.3 fixed the transient-lock false-corruption behavior from #1206. I have not yet reproduced this on the current release; this report is primarily additional evidence for the still-open per-repository single-flight and stale-server lifecycle problems described here.

    For now, I'm upgrading to v0.10.6 which I understand confirms that auto-indexing now runs through a single shared per-user daemon, which should prevent the duplicate workers I saw.

  7. DeusData commented on Sep 25, 2026

    @DeusData
    Owner

    Thank you, @9981H, for a report that came with its own diagnosis, and @jameswilson for preserving the evidence and adding the macOS/Codex case!

    I need to correct my earlier comment: the fixes you both named are now in the code.

    • All MCP sessions talk to one shared per-user daemon. It merges identical index requests into a single job and serializes indexing of the same project with a per-project lock, so a dozen conversations can no longer start a dozen full indexes of one repository.
    • The per-session frontends exit when their client's stdin closes, so abandoned conversations no longer pile up holding the database open.
    • The WAL now has a size limit and a truncating checkpoint.

    I haven't reproduced your many-clients scenario end to end, so I'd rather hear it from you. Could you both run v0.11.0 in your normal setup for a day and tell us the largest -wal size you see, and whether more than one index worker ever runs for the same repo? If both hold up, we'll close this. Thanks!

  8. added
    awaiting-reporterWaiting on the reporter for info/repro; stale bot will warn then close
    on Sep 25, 2026
  9. jameswilson commented on Sep 26, 2026

    @jameswilson

    to be honest, since updating to 0.10.6, I haven't had any noticeable performance problems on my aging mac (see my prev comment for context). All the same, I'll be updating to 0.11 on Monday, but I'm expecting no negatives here. Will come back to report if something does come up tho.

    Thanks so much for your work on this project!

  10. github-actions commented on Sep 26, 2026

    @github-actions

    Thanks for the report! To move this forward we need a bit more so we can reproduce it ourselves:

    • the codebase-memory-mcp --version you're on
    • the exact steps or command you ran
    • a public repo (or a small dummy snippet) that shows the problem — please don't paste proprietary code

    Once that's here we'll pick it straight back up. Heads-up: issues left awaiting-reporter are automatically closed after a few weeks of silence, but a comment reopens the door anytime.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    awaiting-reporterWaiting on the reporter for info/repro; stale bot will warn then closebugSomething isn't workingeditor/integrationEditor compatibility and CLI integrationpriority/highNeeds near-term maintainer attention; high-impact bug, regression, safety issue, or release blocker.stability/performanceServer crashes, OOM, hangs, high CPU/memoryux/behaviorDisplay bugs, docs, adoption UXwindowsWindows-specific issues

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions