Logging this bug against myself so this information doesn't get lost. This came out of the JUCE investigations.
CMidiPorts lock contention stalls the WinMM MIDI send path during device changes
Component: wdmaud2.drv — MidiSrvPorts.cpp, MidiSrvPort.cpp, MidiSrvPort.h
Found: 2026-09-03, while analyzing JUCE's Windows MIDI back end (juce-framework/JUCE#1726). Not JUCE-specific — it affects every WinMM client.
Status: Partially fixed. One of three items addressed; two remain.
Summary
Every process that uses WinMM MIDI holds a single exclusive CMidiPorts::m_Lock. That lock is taken by the message send path, by all enumeration entry points, and by the PnP notification callback. During a device arrival or removal the notification thread holds it long enough to stall unrelated callers by tens of milliseconds, including midiOutShortMsg on an audio thread.
One cause — a full PnP walk per departing interface — has been fixed under Feature_Servicing_MIDI2WinMMInterfaceRemovalPerf. Two causes remain, plus a related staleness question.
Measurements
All on a machine with 33 MIDI OUT / 29 MIDI IN ports (39 / 35 device interfaces), unplugging a 4-out/3-in class-compliant controller. midiOutShortMsg sent to a software loopback port from a Highest-priority thread; getNumDevs polled from a second thread.
midiInGetNumDevs is the cleanest probe available: with Feature_Servicing_MIDI2NumDevsPerf enabled it takes the lock, reads a cached integer, and returns. No I/O, no RPC, no allocation. Any measurable duration is lock wait by construction.
Baseline, no device activity: midiOutShortMsg median 0.0155 ms, p99 0.049 ms, max 0.312 ms over 3,002 calls.
Before any fix, six removals:
| metric |
value |
getNumDevs pair, per removal |
14.5, 17.2, 18.3, 23.6, 25.2, 21.6 ms (mean 20.0) |
getNumDevs pair, per arrival |
~0.05 ms |
midiOutShortMsg max |
30.8 ms |
30.8 ms is 5.8 buffer periods at 48 kHz / 256 samples.
PnP walk cost (measured directly by replaying the RefreshPortsForFlow SetupDi sequence):
| flow |
interfaces |
median quiet |
median during device churn |
per interface |
| MIDI OUT |
39 |
2.05 ms |
2.19 ms |
0.045 ms |
| MIDI IN |
35 |
1.60 ms |
— |
0.046 ms |
The walk is not inflated by PnP activity (6% slower, and the storm maximum was lower than the quiet maximum). The stall was queueing, not an expensive walk.
ETW attribution. CMidiPorts is per-process, so refreshes must be counted per PID. Four unrelated processes each performed their own for the same device events — the test app, rtpMIDISvc, TotalMixFX_x64, NIHostIntegrationAgent — 320 refreshes total in 60 seconds. Per process:
| event |
refreshes |
span |
mean gap |
stall observed |
| removal |
4 (one per departing interface) |
~8 ms |
2.6 ms — contiguous |
14–25 ms |
| arrival |
10 |
~65 ms |
7.2 ms — interleavable |
~0.05 ms |
That asymmetry is the mechanism: on removal the walks are back to back, and wil::critical_section is not FIFO, so the notification thread releases and immediately re-acquires, starving waiters for 1.5–2.4× the total hold. On arrival the same work is spaced out and waiters slot in between.
Client-observed port creation span: ~435 ms from first activity to all ports visible, measured twice within 2 ms (434.7 and 432.9 ms), for a 4-out/3-in device. In-ports and out-ports became visible ~63 ms apart.
Every method that takes m_Lock
m_Lock is a wil::critical_section — exclusive only, no shared mode.
| line |
method |
hot path? |
| 89 |
MidiInterfaceChange |
held across the entire refresh |
| 1093 |
GetOpenedPort |
yes — every message send |
| 257 / 334 |
MidMessage / ModMessage GETNUMDEVS |
enumeration |
| 1041 |
GetDevCaps |
enumeration |
| 1232 |
GetDeviceInterface (DRV_QUERYDEVICEINTERFACE) |
enumeration |
| 1122 / 1204 |
Open / Close |
|
| 214 |
Shutdown |
|
| 396 |
GetMidiDeviceCount |
legacy, gated out |
ForwardModMessage (1335) calls GetOpenedPort (1349), which is why sending a MIDI message contends with PnP enumeration.
Item 1 — FIXED: removal performed a full PnP walk per departing interface
MidiInterfaceChange called RefreshPortsForFlow once per removal notification. Fixed under Feature_Servicing_MIDI2WinMMInterfaceRemovalPerf by net-new CMidiPorts::RemovePortForInterface, which erases the departed interface from m_MidiPortInfo[flow] and recomputes m_MidiPortCount[flow] from the ordered map — the same state a refresh would produce for a single departure, with no enumeration. Gated at the call site; the original path is preserved in the else.
Result:
|
before |
after |
getNumDevs per removal, mean |
20.0 ms |
6.6 / 7.6 ms (two runs) |
midiOutShortMsg max |
30.8 ms |
19.7 / 19.4 ms |
| loopback removal test |
3.246 ms |
0.23–0.65 ms |
| idle control |
0.352 ms |
0.324 ms — unchanged |
ETW confirms the removal path now performs zero refreshes and completes seven interface removals in under 1 ms.
Correctness verified independently: the full enumerable port table — every index, every name, both device counts — returns exactly to baseline after each of three create/remove cycles.
Item 2 — OPEN: NotifyInterfaceRemoval is called while holding the ports lock
MidiSrvPorts.cpp:102 iterates m_OpenPorts calling MidiSrvPort.h:33, which acquires the per-port lock — all while CMidiPorts::m_Lock is still held.
If any port's lock is held by an in-flight send (Item 3), the notification thread blocks on it while pinning the global ports lock, dragging in every other caller: getNumDevs, GetDevCaps, other ports' sends.
Fixable within wdmaud2. Snapshot the open-port com_ptrs under the lock, release it, then notify. Needs its own KIR; it changes lock scope, so it is more invasive than Item 1.
Helps enumeration callers and other ports. Does not help the thread whose own send is slow.
Item 3 — OPEN: the send path holds the port lock across a cross-process call
MidiSrvPort.cpp:849 takes the per-port m_Lock and holds it across m_MidisrvTransport->SendMidiMessage(...). While midisrv tears down a removed device that call runs ~19 ms.
This is the dominant remaining cost and is service-side. After Item 1, ETW shows the shape clearly at every hardware removal:
17:43:56.295 7 x RemovePortForInterface (two CM threads) — all within 1 ms
17:43:56.316 sender resumes
--> SENDER BLOCKED 21.27 ms, nothing logged by any thread during the block
Out of scope for a wdmaud2-only change.
Item 4 — LOWER PRIORITY: GETNUMDEVS can return a count that lags PnP
With Feature_Servicing_MIDI2NumDevsPerf enabled, GETNUMDEVS returns the cached m_MidiPortCount[flow], updated only from the CM callback. That change was necessary — callers exist that put midiInGetNumDevs() in a loop condition, which previously meant N+1 full enumerations per pass — but it removed an accidental self-healing property: a caller that raced the notification used to still get ground truth.
Two windows, only one of which is ours:
- Window A (ours): the interface is enabled and visible to
SetupDiGetClassDevs(DIGCF_PRESENT), but this process's CMEventCallback has not run, so the cache lags. Small.
- Window B (not fixable here): the service is still mid-creation, so the low count is the honest answer. ~435 ms measured. The caller must look again.
Window B dominates, which is why this is low priority.
If revisited: there is no PnP generation counter, so the shape is a cheap validity probe — CM_Get_Device_Interface_ListW with CM_GET_DEVICE_INTERFACE_LIST_PRESENT, hash the multi-sz, compare against the hash saved by the last full refresh, and only pay for RefreshPortsForFlow on mismatch. Add a 10–50 ms throttle so a looping caller cannot produce a PnP query per iteration.
Rejected: querying midisrv through m_MidisrvTransport for the count — adds an RPC to a hot path, creates a new failure mode when the service is down, and lets the service view diverge from what is actually enumerable.
Also considered and rejected
Converting m_Lock to an SRW lock with lock_shared() on the read paths. This is the most general fix — readers would never wait on a refresh at all. It cannot be KIR-gated cleanly: switching lock primitive at runtime requires routing all ten acquisition sites through helpers, and one missed site is a silent data race. Revisit only if Items 2 and 3 prove insufficient.
Repro and tests
midi-lock-probe.ps1 — sends Active Sensing to a software loopback out port at ~500/s from a Highest-priority thread, timing every call; second thread polls getNumDevs and timestamps count changes. Requires physical plug/unplug.
midi-removal-stall-test.ps1 — scripted, no hardware. Uses midi loopback create / remove to fire the same interface-removal path. Buckets samples into idle / create / removal windows, PASS/FAIL on the removal bucket against a 1.0 ms limit.
midi-removal-correctness-test.ps1 — snapshots the entire enumerable port table and requires it to return exactly to baseline after create/remove cycles.
midi-walk-cost.ps1 — replays the RefreshPortsForFlow SetupDi sequence to measure walk cost directly, with a monitor mode that runs across device changes.
All located in \build\tests\winmm-lock\
Caveat: the loopback test passes even with Items 2 and 3 present, because loopback removal involves no device teardown and therefore no slow cross-process send. It guards the PnP walk only. Do not treat a loopback PASS as proof the hardware case is fixed.
ETW decode recipe: tracerpt x.etl -o x.xml -of XML -y -lr (3 s for 414k events), then stream the XML line by line extracting SystemTime, ProcessID and <Data Name="Location">. Providers Microsoft.Windows.Midi2.WdmAud2 {e6443bc1-e9c5-5a3f-cfb6-abcd62e52e41} and Microsoft.Windows.Midi2.MidiSrv {f42d2441-aac3-5216-0150-3c0f50006b64}. The WdmAud2 provider can be captured unelevated since it is in the client process.
Unexplained, recorded for completeness
midiOutShortMsg also stalls 13–16 ms at a very consistent ~360–370 ms before each device arrival, in both pre-fix and post-fix runs. The ±120 ms trace windows contain no refreshes and only steady-state traffic, so this is not CMidiPorts::m_Lock. Most likely general USB/PnP installation load or midisrv activity. Not attributed, not addressed.
Notes
- All items are in shipping
wdmaud2.drv and need the usual KIR gate with // Start add / // End add markers.
- Items 2, 3 and 4 are independent and can be taken separately.
- None of these are regressions; all are pre-existing.
- Third-party impact is real and not hypothetical:
rtpMIDISvc calls midiOutGetDevCaps/midiInGetDevCaps 149 times per second continuously, each taking m_Lock.
Logging this bug against myself so this information doesn't get lost. This came out of the JUCE investigations.
CMidiPortslock contention stalls the WinMM MIDI send path during device changesComponent:
wdmaud2.drv—MidiSrvPorts.cpp,MidiSrvPort.cpp,MidiSrvPort.hFound: 2026-09-03, while analyzing JUCE's Windows MIDI back end (juce-framework/JUCE#1726). Not JUCE-specific — it affects every WinMM client.
Status: Partially fixed. One of three items addressed; two remain.
Summary
Every process that uses WinMM MIDI holds a single exclusive
CMidiPorts::m_Lock. That lock is taken by the message send path, by all enumeration entry points, and by the PnP notification callback. During a device arrival or removal the notification thread holds it long enough to stall unrelated callers by tens of milliseconds, includingmidiOutShortMsgon an audio thread.One cause — a full PnP walk per departing interface — has been fixed under
Feature_Servicing_MIDI2WinMMInterfaceRemovalPerf. Two causes remain, plus a related staleness question.Measurements
All on a machine with 33 MIDI OUT / 29 MIDI IN ports (39 / 35 device interfaces), unplugging a 4-out/3-in class-compliant controller.
midiOutShortMsgsent to a software loopback port from aHighest-priority thread;getNumDevspolled from a second thread.midiInGetNumDevsis the cleanest probe available: withFeature_Servicing_MIDI2NumDevsPerfenabled it takes the lock, reads a cached integer, and returns. No I/O, no RPC, no allocation. Any measurable duration is lock wait by construction.Baseline, no device activity:
midiOutShortMsgmedian 0.0155 ms, p99 0.049 ms, max 0.312 ms over 3,002 calls.Before any fix, six removals:
getNumDevspair, per removalgetNumDevspair, per arrivalmidiOutShortMsgmax30.8 ms is 5.8 buffer periods at 48 kHz / 256 samples.
PnP walk cost (measured directly by replaying the
RefreshPortsForFlowSetupDi sequence):The walk is not inflated by PnP activity (6% slower, and the storm maximum was lower than the quiet maximum). The stall was queueing, not an expensive walk.
ETW attribution.
CMidiPortsis per-process, so refreshes must be counted per PID. Four unrelated processes each performed their own for the same device events — the test app,rtpMIDISvc,TotalMixFX_x64,NIHostIntegrationAgent— 320 refreshes total in 60 seconds. Per process:That asymmetry is the mechanism: on removal the walks are back to back, and
wil::critical_sectionis not FIFO, so the notification thread releases and immediately re-acquires, starving waiters for 1.5–2.4× the total hold. On arrival the same work is spaced out and waiters slot in between.Client-observed port creation span: ~435 ms from first activity to all ports visible, measured twice within 2 ms (434.7 and 432.9 ms), for a 4-out/3-in device. In-ports and out-ports became visible ~63 ms apart.
Every method that takes
m_Lockm_Lockis awil::critical_section— exclusive only, no shared mode.MidiInterfaceChangeGetOpenedPortMidMessage/ModMessageGETNUMDEVSGetDevCapsGetDeviceInterface(DRV_QUERYDEVICEINTERFACE)Open/CloseShutdownGetMidiDeviceCountForwardModMessage(1335) callsGetOpenedPort(1349), which is why sending a MIDI message contends with PnP enumeration.Item 1 — FIXED: removal performed a full PnP walk per departing interface
MidiInterfaceChangecalledRefreshPortsForFlowonce per removal notification. Fixed underFeature_Servicing_MIDI2WinMMInterfaceRemovalPerfby net-newCMidiPorts::RemovePortForInterface, which erases the departed interface fromm_MidiPortInfo[flow]and recomputesm_MidiPortCount[flow]from the ordered map — the same state a refresh would produce for a single departure, with no enumeration. Gated at the call site; the original path is preserved in theelse.Result:
getNumDevsper removal, meanmidiOutShortMsgmaxETW confirms the removal path now performs zero refreshes and completes seven interface removals in under 1 ms.
Correctness verified independently: the full enumerable port table — every index, every name, both device counts — returns exactly to baseline after each of three create/remove cycles.
Item 2 — OPEN:
NotifyInterfaceRemovalis called while holding the ports lockMidiSrvPorts.cpp:102iteratesm_OpenPortscallingMidiSrvPort.h:33, which acquires the per-port lock — all whileCMidiPorts::m_Lockis still held.If any port's lock is held by an in-flight send (Item 3), the notification thread blocks on it while pinning the global ports lock, dragging in every other caller:
getNumDevs,GetDevCaps, other ports' sends.Fixable within wdmaud2. Snapshot the open-port
com_ptrs under the lock, release it, then notify. Needs its own KIR; it changes lock scope, so it is more invasive than Item 1.Helps enumeration callers and other ports. Does not help the thread whose own send is slow.
Item 3 — OPEN: the send path holds the port lock across a cross-process call
MidiSrvPort.cpp:849takes the per-portm_Lockand holds it acrossm_MidisrvTransport->SendMidiMessage(...). While midisrv tears down a removed device that call runs ~19 ms.This is the dominant remaining cost and is service-side. After Item 1, ETW shows the shape clearly at every hardware removal:
Out of scope for a wdmaud2-only change.
Item 4 — LOWER PRIORITY:
GETNUMDEVScan return a count that lags PnPWith
Feature_Servicing_MIDI2NumDevsPerfenabled,GETNUMDEVSreturns the cachedm_MidiPortCount[flow], updated only from the CM callback. That change was necessary — callers exist that putmidiInGetNumDevs()in a loop condition, which previously meant N+1 full enumerations per pass — but it removed an accidental self-healing property: a caller that raced the notification used to still get ground truth.Two windows, only one of which is ours:
SetupDiGetClassDevs(DIGCF_PRESENT), but this process'sCMEventCallbackhas not run, so the cache lags. Small.Window B dominates, which is why this is low priority.
If revisited: there is no PnP generation counter, so the shape is a cheap validity probe —
CM_Get_Device_Interface_ListWwithCM_GET_DEVICE_INTERFACE_LIST_PRESENT, hash the multi-sz, compare against the hash saved by the last full refresh, and only pay forRefreshPortsForFlowon mismatch. Add a 10–50 ms throttle so a looping caller cannot produce a PnP query per iteration.Rejected: querying midisrv through
m_MidisrvTransportfor the count — adds an RPC to a hot path, creates a new failure mode when the service is down, and lets the service view diverge from what is actually enumerable.Also considered and rejected
Converting
m_Lockto an SRW lock withlock_shared()on the read paths. This is the most general fix — readers would never wait on a refresh at all. It cannot be KIR-gated cleanly: switching lock primitive at runtime requires routing all ten acquisition sites through helpers, and one missed site is a silent data race. Revisit only if Items 2 and 3 prove insufficient.Repro and tests
midi-lock-probe.ps1— sends Active Sensing to a software loopback out port at ~500/s from aHighest-priority thread, timing every call; second thread pollsgetNumDevsand timestamps count changes. Requires physical plug/unplug.midi-removal-stall-test.ps1— scripted, no hardware. Usesmidi loopback create/removeto fire the same interface-removal path. Buckets samples into idle / create / removal windows, PASS/FAIL on the removal bucket against a 1.0 ms limit.midi-removal-correctness-test.ps1— snapshots the entire enumerable port table and requires it to return exactly to baseline after create/remove cycles.midi-walk-cost.ps1— replays theRefreshPortsForFlowSetupDi sequence to measure walk cost directly, with a monitor mode that runs across device changes.All located in
\build\tests\winmm-lock\Caveat: the loopback test passes even with Items 2 and 3 present, because loopback removal involves no device teardown and therefore no slow cross-process send. It guards the PnP walk only. Do not treat a loopback PASS as proof the hardware case is fixed.
ETW decode recipe:
tracerpt x.etl -o x.xml -of XML -y -lr(3 s for 414k events), then stream the XML line by line extractingSystemTime,ProcessIDand<Data Name="Location">. ProvidersMicrosoft.Windows.Midi2.WdmAud2{e6443bc1-e9c5-5a3f-cfb6-abcd62e52e41}andMicrosoft.Windows.Midi2.MidiSrv{f42d2441-aac3-5216-0150-3c0f50006b64}. The WdmAud2 provider can be captured unelevated since it is in the client process.Unexplained, recorded for completeness
midiOutShortMsgalso stalls 13–16 ms at a very consistent ~360–370 ms before each device arrival, in both pre-fix and post-fix runs. The ±120 ms trace windows contain no refreshes and only steady-state traffic, so this is notCMidiPorts::m_Lock. Most likely general USB/PnP installation load or midisrv activity. Not attributed, not addressed.Notes
wdmaud2.drvand need the usual KIR gate with// Start add/// End addmarkers.rtpMIDISvccallsmidiOutGetDevCaps/midiInGetDevCaps149 times per second continuously, each takingm_Lock.