Skip to content

Windows BSOD (0x3B) in Ntfs.sys repeatedly triggered during codebase-memory-mcp file-name queries under multi-session load #2021

Description

@sxsorz

Version

codebase-memory-mcp 0.10.8

Platform

Windows (x64)

Install channel

GitHub release archive / install.sh / install.ps1

Binary variant

standard

What happened, and what did you expect?

codebase-memory-mcp 0.10.8 has repeatedly been the faulting process in Windows BSOD dumps. The three most recent crashes (2026-09-01, 2026-09-02, and 2026-09-03) have the same signature: SYSTEM_SERVICE_EXCEPTION (0x3B), exception 0xC0000005, faulting process codebase-memory-mcp.exe, and faulting location Ntfs!NtfsFsdDispatchSwitch+0x15f (module offset 0xDDA5F).

Expected: long-running indexing/query activity and multiple MCP sessions remain bounded and never trigger a kernel crash.

I understand that a user-mode process cannot directly corrupt NTFS and that the root kernel bug may be in Windows or a minifilter driver. However, codebase-memory-mcp is the repeatable triggering process, and the accumulated process/session objects may be relevant.

Reproduction

This is a workload-level reproduction; it is not tied to proprietary repository contents and can be attempted with an ordinary public source repository.

  1. Run codebase-memory-mcp 0.10.8 on Windows 11 x64 build 22631 with the standard binary.
  2. Connect multiple Codex/MCP client sessions to the shared daemon and allow them to index/query source repositories over many hours.
  3. Leave multiple stdio clients/sessions open while normal graph searches and indexing occur.
  4. Under additional system pressure, a codebase-memory-mcp object/file-name query eventually triggers the BSOD.

At the latest crash, the dump contained 64 codebase-memory-mcp process objects; 28 still had a live ObjectTable/handle table. After reboot, only five codebase-memory-mcp processes were present.

A VMware VM was active and using approximately 15 GB RAM. VMware drivers were loaded but none appeared in the faulting thread stack, so this may be a pressure/concurrency amplifier rather than the direct fault.

The issue is intermittent rather than deterministic, so I do not yet have a minimal one-command reproducer. Relevant potentially related reports: #581, #832, #1132, and #1569.

Logs

PROCESS_NAME:  codebase-memory-mcp.exe
BUGCHECK_CODE:  3b
BUGCHECK_P1:    c0000005

Failure.Exception.IP.Module: Ntfs
Failure.Exception.IP.Offset: 0xdda5f

Ntfs!NtfsFsdDispatchSwitch+0x15f
Ntfs!NtfsFsdDispatchWait+0x40
nt!IofCallDriver
FLTMGR!FltpGetFileName
...
luafv!LuafvGenerateFileName
...
gameflt+0x14e4c
...
bindflt!BfPreQueryInfo
...
nt!IopQueryNameInternal
nt!IopQueryName
nt!ObQueryNameStringMode
nt!NtQueryObject

The failing RIP landed in the middle of a valid NTFS jump instruction. WinDbg classified it as IP_MISALIGNED_GenuineIntel.sys, consistent with control-flow/return-address corruption rather than an ordinary NTFS error.

Additional checks:
- Ntfs.sys is Microsoft-signed and validated.
- Internal NVMe drives report healthy.
- NTFS blackbox showed zero slow-I/O and oplock-break timeouts.
- No WHEA hardware errors were recorded.
- No Resource-Exhaustion-Detector event 2004 was recorded around the crash.
- Microsoft gameflt.sys (GameInput minifilter) was directly in the faulting stack.
- A third-party security minifilter and VMware modules were loaded but were not in the faulting stack.

Questions:
1. Is 64 process objects / 28 live handle tables expected under the shared-daemon multi-session model?
2. Are file queries cancelled and workers/handles released when a client session disconnects?
3. Is there a recommended diagnostic build or CBM_DIAGNOSTICS configuration for session/process/handle lifecycle?
4. Would a Windows concurrency limit or stale-session cleanup be an appropriate mitigation?

Diagnostics trajectory (memory / performance / leak issues)

CBM_DIAGNOSTICS was not enabled before these crashes. From the kernel dump: 64 codebase-memory-mcp process objects existed and 28 had live handle tables. No Windows Resource-Exhaustion-Detector event 2004 was present. I can enable CBM_DIAGNOSTICS for a future reproduction and provide a sanitized trajectory; I am not uploading the full memory dump because it may contain private source paths/content.

Project scale (if relevant)

No response

Confirmations

  • I searched existing issues and this is not a duplicate.
  • My reproduction uses shareable code (a dummy snippet or a public OSS repository), not proprietary code.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    awaiting-reporterWaiting on the reporter for info/repro; stale bot will warn then closebugSomething isn't workingeditor/integrationEditor compatibility and CLI integrationpriority/highNeeds near-term maintainer attention; high-impact bug, regression, safety issue, or release blocker.stability/performanceServer crashes, OOM, hangs, high CPU/memorywindowsWindows-specific issues

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions