Skip to content

A failed tool call leaves no tool output, so the thread becomes unsendable after a runtime restart. #6803

Description

@Denis-VG

Description

A tool call that fails is persisted without any tool result: the session item keeps
status: "failed" and metadata.tool_result_for: null, while a successful call carries
its own id in that field. While the same runtime process stays alive this is invisible —
the failure is still in memory and the model receives it normally. After the runtime is
restarted, the next request is rebuilt from the persisted transcript and contains a
function_call with no matching function_call_output, so the provider rejects the
whole request:

Responses API request failed: Invalid request (400):
No tool output found for tool call call_00_iwWH8V1nyzeeIJbRbRlv5008.

The thread is then permanently wedged: every following turn fails in about one second,
and no Runtime API call repairs it.

Steps to reproduce

  1. Start the runtime:
    app-server --http --host 127.0.0.1 --port <port> --auth-token <token>
    (home passed through APPDATA/LOCALAPPDATA/USERPROFILE, as the vendor launchers do)
  2. POST /v1/threads with {"workspace": "<empty directory>"}
  3. POST /v1/threads/<id>/turns with
    {"prompt": "Read the file net-takogo-fajla-12345.txt in the working directory and summarise its contents."}
    The model calls read, the file does not exist, the tool fails.
    This turn completes normally.
  4. Restart app-server on the same home.
  5. POST /v1/threads/<id>/turns with {"prompt": "Answer with one word: alive."}
    → the turn fails with 400 No tool output found for tool call ….

A script that performs exactly these five steps is attached to this report; it prints
REPRODUCED on a machine where the bug is present (exit code 0). It talks to a real
provider — two turns, a fraction of a cent.

Expected behavior

A failed tool call must leave the transcript replayable. Either record a tool result for
it (is_error: true together with the failure text), or drop the call — or synthesise a
function_call_output — when the request is built. Restarting the runtime must not make
an existing thread unsendable.

Actual behavior

The persisted transcript contains a call with nothing answering it:

failed call:    {"kind":"tool_call","status":"failed",
                 "metadata":{"tool_use_id":"call_00_xSP…|a372d3f5-…",
                             "tool_result_for":null}}

completed call: {"kind":"tool_call","status":"completed",
                 "metadata":{"tool_use_id":"call_00_kwb…|b3dfeb42-…",
                             "tool_result_for":"call_00_kwb…|b3dfeb42-…"}}

Within one runtime process nothing breaks; after a restart every turn fails. There is no
repair route:

  • POST /v1/threads/<id>/turns/<turn>/tool-calls/<call_id>/result accepts only
    pending dynamic tool calls — for a failed call from an earlier turn it answers
    404 {"message":"No pending dynamic tool call '<id>'"} (and the thread reports
    pending_dynamic_tool_calls: []).
  • POST /v1/threads/<id>/fork copies the whole thread, orphan included; it ignores
    unknown body fields and offers no truncation.
  • Only POST /v1/threads/<id>/undo helps, and it removes exactly one exchange per call
    and returns a new thread.

Reproduced live on 2026-09-30, deepseek-flash:

turn 1 : completed
  tool_call failed     tool_result_for=None
  tool_call completed  tool_result_for='call_00_DVQd6AoxV6ia428T0Gwg8053|edaff0d0-…'
  tool_call completed  tool_result_for='call_00_cIVEAJs8WjwKBojU13Q30537|ab7657cf-…'
  tool_call completed  tool_result_for='call_01_4NofpBOZf2PERQj72T7p6773|13c6e540-…'
  tool_call completed  tool_result_for='call_00_WAq4k4cKMkEsg0y5SvWs4846|079c1917-…'
  failed calls with no tool_result_for: 1
--- restarting the runtime on the same home
turn   : turn_65d3d034 failed
  error failed  … (400): No tool output found for tool call call_00_iwWH8V1nyzeeIJbRbRlv5008
== REPRODUCED

Impact

Happens 100% of the time once a tool has failed and the app is restarted — after that the
thread cannot be used at all, and every new message is rejected in about a second.

Triggers seen so far, all of them ordinary: a code_execution call timing out after
120 s, a code_execution call failing input validation, and a read of a missing path.
The restart is not exotic either — it happens on every app restart, and our client
restarts the runtime whenever settings are saved.

The workaround costs the user history: N undo calls for a broken turn that sits N turns
from the end, each one leaving another thread behind (there is no delete).

Environment

  • OS: Windows 10 Pro, 10.0.19045 (AMD64)

  • codewhale version: codewhale 0.10.0 (1be1a703b975); runtime_api_version: 1.0

  • Install method: release binary (codew-windows-x64.exe), run as app-server --http

  • codewhale doctor summary (relevant lines only):

    Version Information:
      codewhale-tui: 0.10.0 (1be1a703b975)
    
    API Connectivity:
      · provider: deepseek
      · base_url: https://api.deepseek.com
      · model: deepseek-flash (resolved)
      · strict_tool_mode: disabled
    
    Tool Dependencies:
      ✓ Python: python → code_execution tool registered
      ✓ Node.js: present → js_execution tool registered
    
  • Model/provider: deepseek-flash via deepseek (https://api.deepseek.com)

  • Terminal app: none — the runtime is started headless by a browser client that talks to
    /v1/* and renders the SSE stream (the same thread had also been used from the TUI
    earlier that day)

  • Shell: PowerShell 5.1 on Windows

Logs, screenshots, or recordings

  • The reproduction script is attached to this report (steps 1–5 above, with the verdict).

  • Nothing comes from the provider side except the rejection itself, because the request is
    refused before it reaches the model:

    Responses API request failed: Invalid request (400):
    No tool output found for tool call call_00_iwWH8V1nyzeeIJbRbRlv5008. (request_id: …)
    
  • Notes:

    • max_history: 100 and auto_compact: false in this config are irrelevant — the
      reproduction runs on a 17-item thread.
    • The same missing result is recorded when the provider is a local OpenAI-compatible
      endpoint (that is where we first saw it); such an endpoint does not reject unpaired
      tool calls, so the hole stayed invisible there.
    • strict_tool_mode: disabled here — would enabling it catch a tool call that has no
      output before the request goes out?
    • Side observation, not investigated: after undo the runtime rewrote the session —
      turn ids changed while the thread id stayed the same.

repro_orphan_tool_call.py

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingreliabilityReliability, flaky behavior, retries, fallbacks, and robustnessruntime-apiRuntime API, app-server, SDK and exec automation surfacestoolsTool execution, tool schemas, tool UX, and built-in tool behavior

    Type

    No type

    Projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions