Repository navigation
WAL grows unbounded (115 GB in 4.5 h, fills system drive): checkpoint starvation from accumulated per-conversation servers + 11 concurrent index workers on one repo #1083
Description
Activity
- addededitor/integrationEditor compatibility and CLI integrationEditor compatibility and CLI integrationux/behaviorDisplay bugs, docs, adoption UXDisplay bugs, docs, adoption UXwindowsWindows-specific issuesWindows-specific issues
on Jul 14, 2026 - addedbugSomething isn't workingSomething isn't workingstability/performanceServer crashes, OOM, hangs, high CPU/memoryServer crashes, OOM, hangs, high CPU/memorypriority/highNeeds near-term maintainer attention; high-impact bug, regression, safety issue, or release blocker.Needs near-term maintainer attention; high-impact bug, regression, safety issue, or release blocker.
on Jul 14, 2026 Thanks for the unusually strong operational report. This is a high-priority 0.9.1-rc stability bug because unbounded WAL growth can exhaust the system disk. Investigation should cover per-repository single-flight indexing, stale MCP server lifetime, long-lived read transactions blocking checkpoints, and duplicate workers across clients. Please preserve the database/WAL evidence if possible rather than deleting it while this is being diagnosed.
- added a commit that references this issue
on Jul 16, 2026 - added a commit that references this issue
on Jul 16, 2026 - added a commit that references this issue
on Jul 28, 2026 @9981H — filling someone's system drive is about the worst thing a tool can do to their machine, and I am sorry it was ours that did it.
Your report carried its own diagnosis: a 60 MB database against a 115 GB
-wal, 11 concurrent workers on one repository, 30+ abandoned servers holding read marks, and then the WAL collapsing to a few MB the moment a checkpoint finally succeeded. That last observation is what proves it was pure un-checkpointed log rather than data volume, and it saved a lot of guessing.journal_size_limitand checkpoint-starvation warnings have shipped since, but the parts you named as the actual fix — per-repository single-flight indexing, and servers that exit when their client's stdio closes — are the ones that stop this recurring, and I am not going to claim those are done. Tonight's v0.10.5 is being cut for the issues that make cbm impossible to operate, and unbounded disk consumption is at the front of that queue.Adding a macOS/Codex data point to this issue.
Environment
- codebase-memory-mcp: v0.9.0
- MCP client: Codex Desktop 26.721.81911
- macOS 12.7.6 (21H1320), x86_64
- MacBookPro16,4
- 8 physical / 16 logical CPUs
- 32 GiB RAM
- Binary SHA-256:
04ee3048810c19099502adc8bb83039423f02f2553d17677892a7f03b924e01f - Repository: private Git repository, 8,657 indexed files
Observed behavior
With
auto_index=trueand the default watcher behavior, Codex accumulated 23 MCP server processes. Thirteen simultaneousindex_repositoryworkers then started full indexes of the same repository.Each index reported:
pipeline.mode mode=parallel workers=16 files=8657Observed resource use:
- 13 simultaneous full index jobs. (!!!)
- Approximately 592–850% aggregate CPU
- Approximately 6.3–7.2 GiB aggregate RSS
- macOS load averages exceeded 800
- Approximately 108,000 extracted graph nodes per competing job
The canonical SQLite database disappeared and was replaced by a 143 MiB
.db.corruptfile.PRAGMA quick_checkconfirmed genuine B-tree damage, including out-of-order rowids.Several concurrent workers later produced a new 164 MiB database. It was also genuinely malformed:
2nd reference to page ... Page ... is never used database disk image is malformed (11)Disabling
auto_indexalone was insufficient because existing MCP sessions continued to hold watchers/connections and spawned additional workers.A manual CLI rebuild completed analysis with:
expected_nodes=115030 expected_edges=129003 status=degraded hint="Index database failed integrity check and was removed."After setting both:
auto_index=false auto_watch=falsesix older MCP servers ignored
SIGTERMand immediately spawned replacement index workers. They had to be force-stopped.With all MCP server processes gone, one isolated manual rebuild succeeded:
nodes=115030 edges=128898 status=indexedPost-rebuild verification:
PRAGMA quick_check; okA subsequent no-change manual update completed in 0.55 seconds and retained the same node/edge counts.
Expected behavior
- Only one index worker should operate on a given repository/database.
- Additional clients should queue or receive a clear “index already in progress” result.
- Closing or terminating an MCP client should stop its server and watcher.
- Concurrent clients must not create genuinely malformed SQLite databases.
- Disabling automatic indexing/watching should prevent existing helpers from spawning replacement workers.
Notes
The corrupt databases and worker logs were preserved. They contain metadata and paths from a private repository, so I can provide sanitized excerpts rather than uploading the databases publicly.
I understand that v0.10.3 fixed the transient-lock false-corruption behavior from #1206. I have not yet reproduced this on the current release; this report is primarily additional evidence for the still-open per-repository single-flight and stale-server lifecycle problems described here.
For now, I'm upgrading to v0.10.6 which I understand confirms that auto-indexing now runs through a single shared per-user daemon, which should prevent the duplicate workers I saw.
- added 8 commits that reference this issue
on Sep 2, 2026 Thank you, @9981H, for a report that came with its own diagnosis, and @jameswilson for preserving the evidence and adding the macOS/Codex case!
I need to correct my earlier comment: the fixes you both named are now in the code.
- All MCP sessions talk to one shared per-user daemon. It merges identical index requests into a single job and serializes indexing of the same project with a per-project lock, so a dozen conversations can no longer start a dozen full indexes of one repository.
- The per-session frontends exit when their client's stdin closes, so abandoned conversations no longer pile up holding the database open.
- The WAL now has a size limit and a truncating checkpoint.
I haven't reproduced your many-clients scenario end to end, so I'd rather hear it from you. Could you both run v0.11.0 in your normal setup for a day and tell us the largest
-walsize you see, and whether more than one index worker ever runs for the same repo? If both hold up, we'll close this. Thanks!- addedawaiting-reporterWaiting on the reporter for info/repro; stale bot will warn then closeWaiting on the reporter for info/repro; stale bot will warn then close
on Sep 25, 2026 to be honest, since updating to 0.10.6, I haven't had any noticeable performance problems on my aging mac (see my prev comment for context). All the same, I'll be updating to 0.11 on Monday, but I'm expecting no negatives here. Will come back to report if something does come up tho.
Thanks so much for your work on this project!
Thanks for the report! To move this forward we need a bit more so we can reproduce it ourselves:
- the
codebase-memory-mcp --versionyou're on - the exact steps or command you ran
- a public repo (or a small dummy snippet) that shows the problem — please don't paste proprietary code
Once that's here we'll pick it straight back up. Heads-up: issues left
awaiting-reporterare automatically closed after a few weeks of silence, but a comment reopens the door anytime.- the
- added a commit that references this issue
on Sep 29, 2026
Environment
codebase-memory-mcp.exe; over one work day 30+ server processes accumulated (clients keep running after their conversations are closed, so servers are never reaped).What happened
The main DB
~/.cache/codebase-memory-mcp/<repo>.dbis only 60 MB, but its-walfile grew from 0 to 115 GB in ~4.5 hours (created 10:47, 115 GB by 15:30 local) and completely filled the system drive (free space went 11.3 GB -> 7.8 GB while we watched, ~0.5 GB per 2 minutes).Process snapshot at 15:32 showed, besides 20+ long-lived server processes, 11 concurrent
cli --index-worker index_repositoryworkers all indexing the same repo path at the same time.Diagnosis
Classic SQLite WAL checkpoint starvation:
wal_checkpointcan never truncate;index_repositoryat the end of every task) appends the whole dataset to the WAL again;Confirmation: at 15:39 a checkpoint finally succeeded and the WAL collapsed from 115 GB to a few MB — so the 115 GB was pure un-checkpointed log, not data volume.
Probably related: #914 (orphan cbm process), #937 (watcher re-index write amplification).
Suggestions
PRAGMA journal_size_limitand forcewal_checkpoint(TRUNCATE)after every index run (and/or periodically), logging a warning when the WAL exceeds a threshold.index_repositoryis already running for the same repo, queue or refuse instead of running 11 full indexes concurrently.