Skip to content

Windows: pin Microsoft.ML.OnnxRuntime.DirectML back to 1.23.0 (1.24.x faults natively on session creation) - #2443

Merged
stakira merged 1 commit into
openutau:masterfrom
KakaruHayate:fix/win-directml-1230
Sep 24, 2026
Merged

stakira merged 1 commit into
openutau:masterfrom
KakaruHayate:fix/win-directml-1230

Conversation

@KakaruHayate

Copy link
Copy Markdown
Contributor

Fixes the native crash on Windows when the DirectML runner is used and a singer
is switched or loaded.

Symptom

The process disappears while an InferenceSession is being created. There is no
managed exception and nothing reaches the log: the file ends on an ordinary
informational line, and AppDomain.UnhandledException prints nothing. The only
trace left behind is a Windows ".NET Runtime" 1026 event with 0x80131506
(COR_E_EXECUTIONENGINE), i.e. a native fault that is fatalized at the P/Invoke
boundary and never becomes a catchable exception.

Switching the runner to CPU makes it go away completely, which points at the
DirectML execution provider rather than at the rendering code.

Why the package version is the only variable

Every other candidate was checked:

  • The CPU execution provider is unaffected. Same models, same projects, same
    session-creation path, no crash.
  • No voice model is to blame. InferenceSession profiling on a real
    DiffSinger dur.onnx shows 252 of 258 nodes still executing on
    DmlExecutionProvider under both 1.23.0 and 1.24.4, and the per-node
    assignment is identical
    . Nothing silently falls back to CPU in a way that
    would explain a difference between the two versions.
  • DirectML.dll is the same file. Both packages depend on
    Microsoft.AI.DirectML 1.15.4, so the native DirectML runtime is not the
    variable.
  • The DirectML EP sources are the same. Recursively comparing
    onnxruntime/core/providers/dml between v1.23.0 and v1.24.4 shows only
    two changed blobs
    , and both are additive: two declarations in
    DmlExecutionProvider/inc/DmlExecutionProvider.h and two exported APIs in
    dml_provider_factory.cc. Kernel implementations, allocators, the execution
    context and GraphDescBuilder are byte-identical, and graph partitioning
    behaves identically.
  • This repo's own registration code is not the variable. The
    AppendExecutionProvider(OrtEnv, OrtEpDevice[], opts) call introduced by
    feat(onnx): drop Vortice.DXGI & update dml to 1.23.0 #1774 is byte-identical across the branches involved.

So the NuGet package version is what is left. This is the change that made it
move: #1722 ("Add CUDA rendering support on Linux") pushed the Windows group from
1.23.0 to 1.24.4 while adding the Linux CUDA package. The fix here simply
restores the version that Windows had been on since #1774 and for the ~10 months
in between, so it is not new territory for this codebase.

Scope

Only the Windows group is touched.

  • Linux cannot move back: the package it uses
    (Microsoft.ML.OnnxRuntime.Gpu.Linux) has no 1.23.x release at all — it starts
    at 1.24.3 — and the CUDA detection added by Add CUDA rendering support on Linux #1722 expects the newer one.
  • macOS/CoreML is unrelated to this crash and keeps its current version.

Why this is a downgrade rather than a normal fix

To be explicit: this pins around an upstream regression, it does not fix it.
There is no upstream version to wait for — DirectML was moved to sustained
engineering
in 2025-09, Microsoft.ML.OnnxRuntime.DirectML stops at 1.24.4,
and the robustness fixes made after it (including
microsoft/onnxruntime#28007
"dml: add per-instance mutexes to fix concurrent session crashes", still open, as
well as #31719 and #32745) will never ship as a package upgrade. Given that, a
pin is the only mitigation that can actually reach users of the Windows DirectML
path. A proper fix means moving off DirectML — the CPU runner is known not to
crash — and that is a separate, larger change.

Testing

Windows 11 Pro 22H2 (build 22621), x64, NVIDIA RTX 2070, ONNX Runner set to
DirectML.

  • dotnet build OpenUtau -c Release on this branch: 0 errors.
  • Verified the produced artifact actually carries the older runtime:
    runtimes/win-x64/native/onnxruntime.dll is 1.23.20250926.7 and the managed
    Microsoft.ML.OnnxRuntime.dll is 1.23.0.0 (previously
    1.24.20260316.9 / 1.24.4).
  • With the 1.24.4 build the crash reproduces when switching singer; with the CPU
    runner and with the 1.23.0 build it does not reappear in the same usage.

I have only tested on Windows; the Linux and macOS groups are untouched and I did
not build or run them. Flagging this for the maintainers since it changes a
dependency version.

…on creation

The Windows DirectML build of ONNX Runtime 1.24.x can kill the process natively
while an InferenceSession is being created. There is no managed exception and
nothing is written to the log - the process simply disappears, leaving only a
Windows ".NET Runtime" 1026 event with 0x80131506. It reproduces reliably when
switching singer, i.e. when several sessions are created in quick succession.

What was ruled out while tracking this down:
 - the CPU execution provider is unaffected;
 - the 1.23.0 DirectML build is unaffected;
 - no model in the voicebank is to blame: runtime profiling shows 252/258 nodes
   still executing on DmlExecutionProvider under both 1.23.0 and 1.24.4, and the
   per-node assignment is identical;
 - the DirectML EP sources are byte-identical between v1.23.0 and v1.24.4 apart
   from two added exports, and graph partitioning behaves the same way.

1.24.4 is a dead end for fixes. DirectML was moved to sustained engineering in
2025-09, so this is the last release of Microsoft.ML.OnnxRuntime.DirectML, and
the robustness fixes made afterwards (upstream #28007 "per-instance mutexes to
fix concurrent session crashes", #31719, #32745, ...) will never ship as a
package upgrade.

Only the Windows group is touched. The Linux group cannot move back - the
package it uses (Microsoft.ML.OnnxRuntime.Gpu.Linux) has no 1.23.x release at
all - and the macOS/CoreML path is unrelated to this crash.
@KakaruHayate
KakaruHayate requested a review from a team September 24, 2026 06:08
@stakira
stakira merged commit 9a2bc12 into openutau:master Sep 24, 2026
3 checks passed
@KakaruHayate
KakaruHayate deleted the fix/win-directml-1230 branch September 24, 2026 06:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants