Summary
When an LLM call is cancelled by the caller's own signal / deadlineAt, LlmClient still falls back to the host LLM. The host fallback then ignores that signal and runs for timeoutMs + 5 s. A budget-bounded call (for example the turn-start retrieval filter) therefore spawns a host LLM request that can't succeed in time and can't be cancelled.
Code on main:
shouldFallback() (core/llm/client.ts) falls back on any LLM_TIMEOUT:
return (
err.code === ERROR_CODES.LLM_UNAVAILABLE ||
err.code === ERROR_CODES.LLM_RATE_LIMITED ||
err.code === ERROR_CODES.LLM_TIMEOUT
);
core/llm/fetcher.ts raises that same code for a caller-initiated abort ("<provider> request was cancelled", details.cancelled: true), so "the caller gave up" is treated as "the provider is down". By then opts.signal is already aborted, or deadlineAt has passed, and callHostFallback passes that same dead budget on.
- The host bridge drops the signal.
lazyHostLlmBridge.complete in bridge.mts forwards timeoutMs only, and stdio.serverRequest(method, params, { timeoutMs }) (bridge/stdio.ts) has no abort hook:
{ timeoutMs: (input.timeoutMs ?? 60_000) + 5_000 },
Observed
On a fork with query extraction bounded by the turn deadline, for the same reason: turn budget 5.75 s, a large (16k-char) cron-job prompt, and the primary cancelled by its budget signal after about 2.5 s. Then:
host.fallback_failed primary: { code: "llm_timeout", message: "openai_compatible request was cancelled", cancelled: true }
host: "serverRequest host.llm.complete timed out after 7462ms"
retrieval done totalMs: 10521 (budget 5750)
The host-side timeouts seen were 5,513 / 6,452 / 6,875 / 7,220 / 7,462 ms: each equals the remaining budget plus 5,000 ms. The cancelled primary also writes no llm.jsonl row, so from the logs it looks as if the primary was never called.
On main, llmFilterCandidates races its completion against the deadline (waitForFilterDeadline), so the filter itself returns on time. The fallback it triggered keeps running, though: it costs a full host LLM call, and in the Hermes adapter it holds the single serial host-handler worker (bridge_client.py, "one bounded, daemon worker"), which delays any host.llm.complete queued behind it. Any caller that awaits the client without such a race gets the full overrun as latency.
Suggested fix
- In
callWithFallback, don't fall back when the caller's budget is gone; rethrow the primary error instead:
const budgetGone = opts?.signal?.aborted || (opts?.deadlineAt !== undefined && Date.now() >= opts.deadlineAt);
if (!budgetGone && shouldFallback(err, config, provider.name)) { … }
(Checking details.cancelled alone would miss the deadline-passed case.)
- Give
serverRequest an optional signal that rejects and clears the pending entry. In lazyHostLlmBridge, pass input.signal and cap the timeout at deadlineAt - now when a deadline exists, with no +5 s padding.
Test idea
A fake primary that sleeps past deadlineAt, plus a registered host bridge. The call must reject by deadlineAt, and the host bridge must not be invoked. A provider error before the deadline must still fall back as it does today.
Summary
When an LLM call is cancelled by the caller's own
signal/deadlineAt,LlmClientstill falls back to the host LLM. The host fallback then ignores that signal and runs fortimeoutMs + 5 s. A budget-bounded call (for example the turn-start retrieval filter) therefore spawns a host LLM request that can't succeed in time and can't be cancelled.Code on
main:shouldFallback()(core/llm/client.ts) falls back on anyLLM_TIMEOUT:core/llm/fetcher.tsraises that same code for a caller-initiated abort ("<provider> request was cancelled",details.cancelled: true), so "the caller gave up" is treated as "the provider is down". By thenopts.signalis already aborted, ordeadlineAthas passed, andcallHostFallbackpasses that same dead budget on.lazyHostLlmBridge.completeinbridge.mtsforwardstimeoutMsonly, andstdio.serverRequest(method, params, { timeoutMs })(bridge/stdio.ts) has no abort hook:Observed
On a fork with query extraction bounded by the turn deadline, for the same reason: turn budget 5.75 s, a large (16k-char) cron-job prompt, and the primary cancelled by its budget signal after about 2.5 s. Then:
The host-side timeouts seen were 5,513 / 6,452 / 6,875 / 7,220 / 7,462 ms: each equals the remaining budget plus 5,000 ms. The cancelled primary also writes no
llm.jsonlrow, so from the logs it looks as if the primary was never called.On
main,llmFilterCandidatesraces its completion against the deadline (waitForFilterDeadline), so the filter itself returns on time. The fallback it triggered keeps running, though: it costs a full host LLM call, and in the Hermes adapter it holds the single serial host-handler worker (bridge_client.py, "one bounded, daemon worker"), which delays anyhost.llm.completequeued behind it. Any caller that awaits the client without such a race gets the full overrun as latency.Suggested fix
callWithFallback, don't fall back when the caller's budget is gone; rethrow the primary error instead:details.cancelledalone would miss the deadline-passed case.)serverRequestan optionalsignalthat rejects and clears the pending entry. InlazyHostLlmBridge, passinput.signaland cap the timeout atdeadlineAt - nowwhen a deadline exists, with no +5 s padding.Test idea
A fake primary that sleeps past
deadlineAt, plus a registered host bridge. The call must reject bydeadlineAt, and the host bridge must not be invoked. A provider error before the deadline must still fall back as it does today.