Skip to content

test(harness): hold the scheduler contract's fixtures by construction - #2411

Merged
DeusData merged 1 commit into
mainfrom
fix/harness-timeout-race-deterministic
Sep 30, 2026
Merged

DeusData merged 1 commit into
mainfrom
fix/harness-timeout-race-deterministic

Conversation

@DeusData

Copy link
Copy Markdown
Owner

The parallel-scheduler contract failed twice on test-windows CLANG64 while
running fake suites only:

Both verdicts were a race between a clock and runner latency. An audit of
the fixture found more of the same shape:

  • Every hanging fixture ended on its own after 30s. A starved scheduler
    could therefore see a "hung" leader exit 0, and a descendant the
    scheduler failed to kill could vanish before the leak check looked.
  • The 1s --timeout could kill a slow leader before it had printed its
    summary (asserted pass=1), spawned its descendant or written the
    descendant pid. That pid was then read back with int() from a file
    written create-then-write, so a reader could see no file or an empty one.
    The empty-file read recorded earlier was the scheduler's ready file,
    which 3b9052c already publishes atomically; the fixture's own writer
    had the same shape.
  • The descendant installed its SIGTERM-ignore during its own start-up. A
    SIGTERM that won that race skipped the SIGKILL escalation the POSIX leg
    exists to cover, silently and with the verdict unchanged.
  • On Windows, the stubborn-tree leak check was a single look right after
    taskkill, but TerminateProcess returns before its target has exited.

Fix it by construction, not by budget:

  • A gate. Every process that must outlive the scheduler's decision blocks
    reading stdin. That is a pipe whose only write end the contract holds,
    handed down through the scheduler, and the read returns only at EOF. The
    contract closes the pipe once the verdict is recorded. If the contract
    dies, EOF releases everything, so nothing is orphaned and no fixture
    ends on a clock.
  • The three hanging scenarios run under the scheduler's existing
    pre-terminate barrier. They are released only once the fixture has
    published <suite>.established, which happens after the summary is
    flushed, the descendant has confirmed it is armed, and its pid is known.
    The marker is written to a temp file and moved into place with
    os.replace, so a reader sees either no file or the whole file.
  • The Windows leak check waits on the descendant's process handle. That
    wait is exact, not a race: the descendant is held by the gate, so the
    scheduler's kill is the only way it can end.
  • run-test-wave.py runs taskkill through windows_taskkill_tree() with its
    own stable-state budget (WINDOWS_TASKKILL_SECONDS = 15), as 132e8fc did
    for the descendant probe. --kill-grace keeps bounding what its name
    says. Real runs already pass --kill-grace 15, so they are unchanged.
    The contract pins this structurally, the same way it pins the probe:
    exactly one taskkill site, no kill-grace there, and a helper bounded by
    the constant that fails closed on a timeout, a start failure or a
    non-zero exit.

Every assertion keeps its contract:

  • A leader that hangs after a green summary is recorded as rc=124 pass=1.
  • A zero-test child is recorded as rc=97.
  • The wave continues after a bounded failure.
  • A SIGTERM-resistant tree is killed in full.
  • A leader that exits between the timeout decision and termination is
    still cleaned up on POSIX. On Windows the scheduler refuses (rc=2,
    naming the cleanup failure) over a descendant that is really alive.
    A new self-check fails loudly if a held leader or descendant is not live
    while the barrier holds, that is, if the fixture no longer exercises what
    those assertions claim.

Proof, on macOS arm64 and a Linux arm64 container. The Windows branch was
reasoned through; the VM is down.

  • Torture hook (proof only) that delays each hanging leader 2s, past the
    1s timeout. The old fixture is RED 3/3 on every scenario (pass=0,
    FileNotFoundError on descendant.pid, "lost its bounded result"). The new
    one is green, including at 8s, and 10/10 under the hook plus CPU load.
  • The fix(pipeline): spill before the post-extraction phases when they would cross the budget (#2184) #2394 mechanism: after an injected 31s refusal latency the old
    descendant is gone and the new one is live. It exits 0.02s after the
    gate closes.
  • The taskkill pin is RED on the old scheduler and green on the new one.
  • Stress, 20 runs each under CPU load. The old full contract passed 20/20
    on macOS and 20/20 on Linux, so these hosts do not reproduce the CI
    timing. The rebuilt held scenarios passed 20/20 on macOS and 20/20 on
    Linux. The full contract after the change passes a single run on macOS
    and on Linux, and 20/20 on each leg under CPU load (4 load workers;
    the Linux container is capped at 4 CPUs).

Signed-off-by: Martin Vogel martin.vogel.tech@gmail.com

The parallel-scheduler contract failed twice on test-windows CLANG64 while
running fake suites only:

- PR #2394 (run 36364479687): "Windows timeout race refused without a
  surviving descendant to refuse over". The refusal came 38s after the
  forced leader exit (a cold powershell/CIM descendant probe, once in the
  wave loop and again in the cleanup pass), while the fixture's descendant
  was a `time.sleep(30)` that had already exited on its own.
- PR #2345 (run 36199749518): "suite 'hang_after_summary' taskkill could
  not prove process-tree cleanup". The scheduler bounded taskkill.exe with
  --kill-grace, which the contract passes as 1s, so a slow cold start of
  the tool was reported as a failed cleanup.

Both verdicts were a race between a clock and runner latency. An audit of
the fixture found more of the same shape:

- Every hanging fixture ended on its own after 30s. A starved scheduler
  could therefore see a "hung" leader exit 0, and a descendant the
  scheduler failed to kill could vanish before the leak check looked.
- The 1s --timeout could kill a slow leader before it had printed its
  summary (asserted pass=1), spawned its descendant or written the
  descendant pid. That pid was then read back with int() from a file
  written create-then-write, so a reader could see no file or an empty one.
  The empty-file read recorded earlier was the scheduler's ready file,
  which 3b9052c already publishes atomically; the fixture's own writer
  had the same shape.
- The descendant installed its SIGTERM-ignore during its own start-up. A
  SIGTERM that won that race skipped the SIGKILL escalation the POSIX leg
  exists to cover, silently and with the verdict unchanged.
- On Windows, the stubborn-tree leak check was a single look right after
  taskkill, but TerminateProcess returns before its target has exited.

Fix it by construction, not by budget:

- A gate. Every process that must outlive the scheduler's decision blocks
  reading stdin. That is a pipe whose only write end the contract holds,
  handed down through the scheduler, and the read returns only at EOF. The
  contract closes the pipe once the verdict is recorded. If the contract
  dies, EOF releases everything, so nothing is orphaned and no fixture
  ends on a clock.
- The three hanging scenarios run under the scheduler's existing
  pre-terminate barrier. They are released only once the fixture has
  published `<suite>.established`, which happens after the summary is
  flushed, the descendant has confirmed it is armed, and its pid is known.
  The marker is written to a temp file and moved into place with
  os.replace, so a reader sees either no file or the whole file.
- The Windows leak check waits on the descendant's process handle. That
  wait is exact, not a race: the descendant is held by the gate, so the
  scheduler's kill is the only way it can end.
- run-test-wave.py runs taskkill through windows_taskkill_tree() with its
  own stable-state budget (WINDOWS_TASKKILL_SECONDS = 15), as 132e8fc did
  for the descendant probe. --kill-grace keeps bounding what its name
  says. Real runs already pass --kill-grace 15, so they are unchanged.
  The contract pins this structurally, the same way it pins the probe:
  exactly one taskkill site, no kill-grace there, and a helper bounded by
  the constant that fails closed on a timeout, a start failure or a
  non-zero exit.

Every assertion keeps its contract:
- A leader that hangs after a green summary is recorded as rc=124 pass=1.
- A zero-test child is recorded as rc=97.
- The wave continues after a bounded failure.
- A SIGTERM-resistant tree is killed in full.
- A leader that exits between the timeout decision and termination is
  still cleaned up on POSIX. On Windows the scheduler refuses (rc=2,
  naming the cleanup failure) over a descendant that is really alive.
A new self-check fails loudly if a held leader or descendant is not live
while the barrier holds, that is, if the fixture no longer exercises what
those assertions claim.

Proof, on macOS arm64 and a Linux arm64 container. The Windows branch was
reasoned through; the VM is down.
- Torture hook (proof only) that delays each hanging leader 2s, past the
  1s timeout. The old fixture is RED 3/3 on every scenario (pass=0,
  FileNotFoundError on descendant.pid, "lost its bounded result"). The new
  one is green, including at 8s, and 10/10 under the hook plus CPU load.
- The #2394 mechanism: after an injected 31s refusal latency the old
  descendant is gone and the new one is live. It exits 0.02s after the
  gate closes.
- The taskkill pin is RED on the old scheduler and green on the new one.
- Stress, 20 runs each under CPU load. The old full contract passed 20/20
  on macOS and 20/20 on Linux, so these hosts do not reproduce the CI
  timing. The rebuilt held scenarios passed 20/20 on macOS and 20/20 on
  Linux. The full contract after the change passes a single run on macOS
  and on Linux, and 20/20 on each leg under CPU load (4 load workers;
  the Linux container is capped at 4 CPUs).

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
@DeusData
DeusData merged commit 878379a into main Sep 30, 2026
35 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant