Repository navigation
test(harness): hold the scheduler contract's fixtures by construction - #2411
Merged
Merged
Conversation
The parallel-scheduler contract failed twice on test-windows CLANG64 while running fake suites only: - PR #2394 (run 36364479687): "Windows timeout race refused without a surviving descendant to refuse over". The refusal came 38s after the forced leader exit (a cold powershell/CIM descendant probe, once in the wave loop and again in the cleanup pass), while the fixture's descendant was a `time.sleep(30)` that had already exited on its own. - PR #2345 (run 36199749518): "suite 'hang_after_summary' taskkill could not prove process-tree cleanup". The scheduler bounded taskkill.exe with --kill-grace, which the contract passes as 1s, so a slow cold start of the tool was reported as a failed cleanup. Both verdicts were a race between a clock and runner latency. An audit of the fixture found more of the same shape: - Every hanging fixture ended on its own after 30s. A starved scheduler could therefore see a "hung" leader exit 0, and a descendant the scheduler failed to kill could vanish before the leak check looked. - The 1s --timeout could kill a slow leader before it had printed its summary (asserted pass=1), spawned its descendant or written the descendant pid. That pid was then read back with int() from a file written create-then-write, so a reader could see no file or an empty one. The empty-file read recorded earlier was the scheduler's ready file, which 3b9052c already publishes atomically; the fixture's own writer had the same shape. - The descendant installed its SIGTERM-ignore during its own start-up. A SIGTERM that won that race skipped the SIGKILL escalation the POSIX leg exists to cover, silently and with the verdict unchanged. - On Windows, the stubborn-tree leak check was a single look right after taskkill, but TerminateProcess returns before its target has exited. Fix it by construction, not by budget: - A gate. Every process that must outlive the scheduler's decision blocks reading stdin. That is a pipe whose only write end the contract holds, handed down through the scheduler, and the read returns only at EOF. The contract closes the pipe once the verdict is recorded. If the contract dies, EOF releases everything, so nothing is orphaned and no fixture ends on a clock. - The three hanging scenarios run under the scheduler's existing pre-terminate barrier. They are released only once the fixture has published `<suite>.established`, which happens after the summary is flushed, the descendant has confirmed it is armed, and its pid is known. The marker is written to a temp file and moved into place with os.replace, so a reader sees either no file or the whole file. - The Windows leak check waits on the descendant's process handle. That wait is exact, not a race: the descendant is held by the gate, so the scheduler's kill is the only way it can end. - run-test-wave.py runs taskkill through windows_taskkill_tree() with its own stable-state budget (WINDOWS_TASKKILL_SECONDS = 15), as 132e8fc did for the descendant probe. --kill-grace keeps bounding what its name says. Real runs already pass --kill-grace 15, so they are unchanged. The contract pins this structurally, the same way it pins the probe: exactly one taskkill site, no kill-grace there, and a helper bounded by the constant that fails closed on a timeout, a start failure or a non-zero exit. Every assertion keeps its contract: - A leader that hangs after a green summary is recorded as rc=124 pass=1. - A zero-test child is recorded as rc=97. - The wave continues after a bounded failure. - A SIGTERM-resistant tree is killed in full. - A leader that exits between the timeout decision and termination is still cleaned up on POSIX. On Windows the scheduler refuses (rc=2, naming the cleanup failure) over a descendant that is really alive. A new self-check fails loudly if a held leader or descendant is not live while the barrier holds, that is, if the fixture no longer exercises what those assertions claim. Proof, on macOS arm64 and a Linux arm64 container. The Windows branch was reasoned through; the VM is down. - Torture hook (proof only) that delays each hanging leader 2s, past the 1s timeout. The old fixture is RED 3/3 on every scenario (pass=0, FileNotFoundError on descendant.pid, "lost its bounded result"). The new one is green, including at 8s, and 10/10 under the hook plus CPU load. - The #2394 mechanism: after an injected 31s refusal latency the old descendant is gone and the new one is live. It exits 0.02s after the gate closes. - The taskkill pin is RED on the old scheduler and green on the new one. - Stress, 20 runs each under CPU load. The old full contract passed 20/20 on macOS and 20/20 on Linux, so these hosts do not reproduce the CI timing. The rebuilt held scenarios passed 20/20 on macOS and 20/20 on Linux. The full contract after the change passes a single run on macOS and on Linux, and 20/20 on each leg under CPU load (4 load workers; the Linux container is capped at 4 CPUs). Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The parallel-scheduler contract failed twice on test-windows CLANG64 while
running fake suites only:
surviving descendant to refuse over". The refusal came 38s after the
forced leader exit (a cold powershell/CIM descendant probe, once in the
wave loop and again in the cleanup pass), while the fixture's descendant
was a
time.sleep(30)that had already exited on its own.not prove process-tree cleanup". The scheduler bounded taskkill.exe with
--kill-grace, which the contract passes as 1s, so a slow cold start of
the tool was reported as a failed cleanup.
Both verdicts were a race between a clock and runner latency. An audit of
the fixture found more of the same shape:
could therefore see a "hung" leader exit 0, and a descendant the
scheduler failed to kill could vanish before the leak check looked.
summary (asserted pass=1), spawned its descendant or written the
descendant pid. That pid was then read back with int() from a file
written create-then-write, so a reader could see no file or an empty one.
The empty-file read recorded earlier was the scheduler's ready file,
which 3b9052c already publishes atomically; the fixture's own writer
had the same shape.
SIGTERM that won that race skipped the SIGKILL escalation the POSIX leg
exists to cover, silently and with the verdict unchanged.
taskkill, but TerminateProcess returns before its target has exited.
Fix it by construction, not by budget:
reading stdin. That is a pipe whose only write end the contract holds,
handed down through the scheduler, and the read returns only at EOF. The
contract closes the pipe once the verdict is recorded. If the contract
dies, EOF releases everything, so nothing is orphaned and no fixture
ends on a clock.
pre-terminate barrier. They are released only once the fixture has
published
<suite>.established, which happens after the summary isflushed, the descendant has confirmed it is armed, and its pid is known.
The marker is written to a temp file and moved into place with
os.replace, so a reader sees either no file or the whole file.
wait is exact, not a race: the descendant is held by the gate, so the
scheduler's kill is the only way it can end.
own stable-state budget (WINDOWS_TASKKILL_SECONDS = 15), as 132e8fc did
for the descendant probe. --kill-grace keeps bounding what its name
says. Real runs already pass --kill-grace 15, so they are unchanged.
The contract pins this structurally, the same way it pins the probe:
exactly one taskkill site, no kill-grace there, and a helper bounded by
the constant that fails closed on a timeout, a start failure or a
non-zero exit.
Every assertion keeps its contract:
still cleaned up on POSIX. On Windows the scheduler refuses (rc=2,
naming the cleanup failure) over a descendant that is really alive.
A new self-check fails loudly if a held leader or descendant is not live
while the barrier holds, that is, if the fixture no longer exercises what
those assertions claim.
Proof, on macOS arm64 and a Linux arm64 container. The Windows branch was
reasoned through; the VM is down.
1s timeout. The old fixture is RED 3/3 on every scenario (pass=0,
FileNotFoundError on descendant.pid, "lost its bounded result"). The new
one is green, including at 8s, and 10/10 under the hook plus CPU load.
descendant is gone and the new one is live. It exits 0.02s after the
gate closes.
on macOS and 20/20 on Linux, so these hosts do not reproduce the CI
timing. The rebuilt held scenarios passed 20/20 on macOS and 20/20 on
Linux. The full contract after the change passes a single run on macOS
and on Linux, and 20/20 on each leg under CPU load (4 load workers;
the Linux container is capped at 4 CPUs).
Signed-off-by: Martin Vogel martin.vogel.tech@gmail.com