unsloth/studio/backend/tests/test_openai_tool_result_fallbacks.py
Daniel Han 41d24227cd
Studio: per-card web_search result + shell_call output fallback (OpenAI) (#5785)
* Studio: per-card web_search result + shell_call output fallback (OpenAI)

Two empty-output bugs in the OpenAI Responses tool-result rendering that
showed up clearly when a single prompt invoked 9 web_search + 4
code_execution + 1 image_generation in one turn. Reproduction shape in
the SQLite-stored chat history:

- 8 of 9 web_search tool-call records had result == "" (the cards
  rendered as empty cards in the thread)
- 4 of 4 code_execution (shell_call) records were missing the result
  key entirely (NoneType), so the cards that showed "Ran cat ..." style
  commands displayed the command line but no output panel at all
- image_generation worked, as did the very last web_search of the run

Root causes in studio/backend/core/inference/external_provider.py:

1. web_search_call's tool_end emitted result: "" by design, with the
   intent of overwriting only the LAST call at response.completed with
   the full citation list (the source-pill extractor on the frontend
   flatMaps across every web_search result, so a single non-empty
   result is enough for the trailing source pills). Side effect: every
   intermediate card renders empty in the thread. Fix: seed each call's
   own tool_end result with "Searching: <query>" so the per-card text
   is never empty, then keep the last-call overwrite path so the
   source-pill extractor still works. Falls back to empty when the
   model emits an action with no query, so the existing last-call path
   stays unchanged for that edge.

2. shell_call's tool_start was emitted from
   response.output_item.done for the call item, but tool_end lived in
   the separate response.output_item.done handler for shell_call_output.
   When OpenAI's Responses stream bundles the output array onto the
   shell_call item's own done event (no separate shell_call_output
   item), the previous handler emitted tool_start with no following
   tool_end. The card spun on "running" indefinitely and stored as
   NoneType in the thread DB. Fix: when the shell_call's done event
   carries an embedded output list, emit tool_end immediately from
   that. Track tool_end_emitted on the shell_calls map so a subsequent
   shell_call_output event (some streams ship both) is skipped instead
   of double-completing the card. A final flush at response.completed
   emits tool_end for any orphan shell_call that received neither
   bundled output nor a separate output event, so cards always finalise.

Tests (studio/backend/tests/test_openai_tool_result_fallbacks.py, 6
new):
- web_search: three calls, each card's result is its own Searching:
  query (no empties)
- web_search: last call still gets the aggregated citation block when
  url_citations arrive (pins the overwrite path)
- web_search: empty action.query falls back to result == "" (no junk
  Searching: placeholder)
- shell_call: bundled output on done emits a single tool_end with that
  output as the result text
- shell_call: bundled-then-separate output does not double-emit
  tool_end (subsequent shell_call_output is skipped)
- shell_call: orphan call with neither bundled nor separate output is
  flushed at response.completed so the card finalises

15/15 tests green when combined with the existing 9 in
test_openai_code_execution.py. Pre-commit + ruff format clean.

Scope: OpenAI Responses-API code path only. The Anthropic native
Messages-API path (_stream_anthropic) is untouched, as is the local
llama-server path. Local-model behaviour cannot regress because the
edited handlers only fire inside the OpenAI cloud branch.

* Studio: per-model external max_tokens cap + clamp on model switch

Two related external-provider issues that surfaced from the same
investigation as the per-card web_search / shell_call result bugs in
the previous commit:

A. Slider cap was a one-size-fits-all 32768 for every external model.

   provider-capabilities.ts kept a single EXTERNAL_MAX_OUTPUT_TOKENS
   constant (32k), well below what most providers actually accept. The
   docstring even called out the right per-provider numbers (Anthropic
   Opus 128k, GPT-5.x ~128k, Gemini 2.5 ~65k, DeepSeek 8k) but the
   code picked the lowest as a conservative floor. Effect: long
   generations from gpt-5.5 / claude-opus-4-7 silently truncated at
   32k even though the API would have served up to 128k.

   Fix: introduce getExternalMaxOutputTokens(providerType, modelId)
   returning the documented per-model cap. Patterns are checked
   longest-first so e.g. gpt-5.5-pro matches before gpt-5.5. Unknown
   provider/model combinations fall back to the existing 32k floor so
   no surprise increases for ids we don't know about.

   Per-model caps from the official docs:
   - OpenAI gpt-5.5 / gpt-5.5-pro: 128000
   - OpenAI gpt-5.4 / gpt-5.4-pro: 65536
   - OpenAI gpt-5.3: 16384
   - Anthropic claude-opus-4-7: 128000
   - Anthropic claude-opus-4-6 / sonnet-4-6 / opus-4-5 / sonnet-4-5 /
     haiku-4-5: 64000
   - Gemini 3.x family: 65535
   - DeepSeek: 8192
   - OpenRouter: strip provider/ prefix from the id and re-resolve

   The slider in chat-settings-sheet.tsx and the send-time clamp in
   chat-adapter.ts both call the new function so the slider's max=
   matches what the wire layer will accept.

B. Slider value lied after switching from a local model to external.

   When Studio auto-loads the helper Gemma-4-E2B-it on first chat,
   chat-adapter sets params.maxTokens to Gemma's context_length
   (262144 for Gemma 4). Switching the model picker to gpt-5.5 then
   flips the slider's max prop to the external cap, but the stored
   params.maxTokens is never reset. The numeric value next to the
   slider would render 262144 against a track that ended at the
   external cap. The send-time clamp brought the outbound max_tokens
   back down to the cap, so the API call was safe, but the displayed
   number had no relationship to what was actually being sent.

   Fix: chat-runtime-store.setCheckpoint now clamps params.maxTokens
   to getExternalMaxOutputTokens(...) on transitions into an external
   model. Looks up the provider via useExternalProvidersStore so we
   can derive providerType from the parsed external model id. No-op
   when the stored maxTokens is already at or below the new cap, so
   user-tuned values within range survive the switch.

Scope: pure frontend changes scoped to external-provider code paths.
Local model behaviour is untouched -- the ggufContextLength branch of
the slider's max= is unchanged, and setCheckpoint only mutates
maxTokens when isExternalModelId(modelId) is true. The send-time
clamp continues to be the safety net for any in-flight request that
crosses a model switch before the store-level clamp has applied.

Typecheck (tsc -b) clean; bun run build succeeds (2.13s).

Co-changes with the previous commit (7fe1adbf, per-card web_search +
shell_call output fallback) form a single PR: every empty-output and
silent-truncation issue surfaced from the same animal-popularity
prompt reproduction is now addressed in one branch.

* Studio: correct external max_tokens caps for Gemini and DeepSeek

Per-doc corrections to the per-model cap table added in 95da8d52:

- Gemini 3.x family: 65535 -> 65536, per
  https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro-preview
  (the published max_output_tokens is exactly 64K = 65536). The earlier
  65535 was an off-by-one rough cap.
- DeepSeek (deepseek-chat / deepseek-reasoner aliases): 8192 -> 384000,
  per https://api-docs.deepseek.com/quick_start/pricing. DeepSeek V4
  Flash / Pro both list MAX OUTPUT = 384K; the chat / reasoner ids are
  deprecated aliases for V4 Flash non-thinking / thinking modes. The
  8192 value was carried over from V3 and silently truncated V4 traffic
  at 2% of its actual ceiling.

Affects only the slider max and the send-time clamp for these provider
types. Other providers' caps unchanged. tsc -b clean.

* Studio: also flush orphan shell_calls on response.incomplete

Addresses gemini-code-assist[bot] high-priority inline review on PR
5785: the orphan-shell_call final flush added in 7fe1adbf landed only
in the response.completed branch. Truncated OpenAI Responses streams
emit response.incomplete instead (for example when the request hits
max_output_tokens), which left in-flight shell_call cards spinning
indefinitely in the UI.

Mirror the same flush block in the response.incomplete handler so the
truncated-stream path finalizes every pending tool card. The
tool_end_emitted guard keeps the path idempotent: if a shell_call
already completed via bundled output on its done event, the incomplete
flush is a no-op for it.

Two new tests in test_openai_tool_result_fallbacks.py:
- test_shell_call_flushed_on_response_incomplete_truncation pins the
  bug repro: an in-flight shell_call followed by response.incomplete
  must emit tool_end so the card finalizes.
- test_shell_call_incomplete_does_not_double_emit pins idempotency:
  a shell_call that completed via bundled output and is then followed
  by response.incomplete emits exactly one tool_end with the bundled
  result text.

17/17 tests green (8 fallback tests + 9 existing code-execution). Pre-
commit + ruff format clean.

* Studio: trim verbose comments across PR 5785 edits

Compress the in-code commentary added across this branch to one or two
lines per block; the verbose prose was easier as a PR description than
as inline noise. No behavioural changes: 17/17 tests still green, tsc -b
still clean.
2026-05-26 04:31:22 -07:00

372 lines
12 KiB
Python

# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""Regression tests for OpenAI Responses tool-result rendering.
Covers two bug classes: empty web_search cards (per-card result seeded
with "Searching: <query>") and orphan shell_call cards (bundled-output
fallback + final flush at response.completed / response.incomplete).
"""
import asyncio
import json
import httpx
from core.inference import external_provider as ep_mod
from core.inference.external_provider import ExternalProviderClient
def _drive(coro):
return asyncio.new_event_loop().run_until_complete(coro)
async def _collect(agen):
out = []
async for line in agen:
out.append(line)
return out
def _mock_http_client(monkeypatch, handler):
transport = httpx.MockTransport(handler)
monkeypatch.setattr(ep_mod, "_http_client", httpx.AsyncClient(transport = transport))
def _make_client(base_url: str = "https://api.openai.com/v1") -> ExternalProviderClient:
return ExternalProviderClient(
provider_type = "openai",
base_url = base_url,
api_key = "sk-test",
)
def _openai_sse(events: list[dict]) -> bytes:
chunks: list[str] = []
for event in events:
chunks.append(f"event: {event['type']}")
chunks.append(f"data: {json.dumps(event)}")
chunks.append("")
return ("\n".join(chunks) + "\n").encode("utf-8")
def _tool_events(lines: list[str]) -> list[dict]:
out: list[dict] = []
for line in lines:
if not line.startswith("data:"):
continue
raw = line[len("data:") :].strip()
if not raw or raw == "[DONE]":
continue
try:
parsed = json.loads(raw)
except json.JSONDecodeError:
continue
if isinstance(parsed, dict) and "_toolEvent" in parsed:
out.append(parsed["_toolEvent"])
return out
def _drive_stream(sse_events, enabled_tools, monkeypatch):
def handler(request):
return httpx.Response(
200,
content = _openai_sse(sse_events),
headers = {"content-type": "text/event-stream"},
)
_mock_http_client(monkeypatch, handler)
async def run():
client = _make_client()
return await _collect(
client._stream_openai_responses(
messages = [{"role": "user", "content": "x"}],
model = "gpt-5.5",
temperature = 0.7,
top_p = 0.95,
max_tokens = 4096,
enable_thinking = None,
reasoning_effort = None,
enabled_tools = enabled_tools,
)
)
return _drive(run())
# ── web_search per-card result ─────────────────────────────────────────
def test_web_search_each_call_carries_its_own_query_as_result(monkeypatch):
"""Each card carries its own `Searching: <query>` text; no empties."""
sse_events = [
{
"type": "response.output_item.done",
"item": {
"type": "web_search_call",
"id": "ws_1",
"action": {"query": "popular animals 2026"},
},
},
{
"type": "response.output_item.done",
"item": {
"type": "web_search_call",
"id": "ws_2",
"action": {"query": "most loved animals poll"},
},
},
{
"type": "response.output_item.done",
"item": {
"type": "web_search_call",
"id": "ws_3",
"action": {"query": "tiger ranking"},
},
},
{"type": "response.completed", "response": {}},
]
lines = _drive_stream(sse_events, ["web_search"], monkeypatch)
events = _tool_events(lines)
ends = [e for e in events if e["type"] == "tool_end"]
by_id = {e["tool_call_id"]: e for e in ends}
assert by_id["ws_1"]["result"] == "Searching: popular animals 2026"
assert by_id["ws_2"]["result"] == "Searching: most loved animals poll"
assert by_id["ws_3"]["result"] == "Searching: tiger ranking"
def test_web_search_last_call_overwritten_with_citations(monkeypatch):
"""Last call still gets the aggregated citation list; earlier calls
keep their per-call `Searching:` text."""
sse_events = [
{
"type": "response.output_item.done",
"item": {
"type": "web_search_call",
"id": "ws_1",
"action": {"query": "first query"},
},
},
{
"type": "response.output_item.done",
"item": {
"type": "web_search_call",
"id": "ws_2",
"action": {"query": "second query"},
},
},
{
"type": "response.output_text.annotation.added",
"annotation": {
"type": "url_citation",
"url": "https://example.com/a",
"title": "Example A",
},
},
{"type": "response.completed", "response": {}},
]
lines = _drive_stream(sse_events, ["web_search"], monkeypatch)
events = _tool_events(lines)
ends = [e for e in events if e["type"] == "tool_end"]
by_id: dict = {}
# Keep the LAST tool_end per id (the citation overwrite for ws_2).
for e in ends:
by_id[e["tool_call_id"]] = e
# First call keeps its own query.
assert by_id["ws_1"]["result"] == "Searching: first query"
# Last call gets overwritten with the citation block.
assert "Title: Example A" in by_id["ws_2"]["result"]
assert "URL: https://example.com/a" in by_id["ws_2"]["result"]
def test_web_search_empty_query_falls_back_to_empty_result(monkeypatch):
"""No query -> empty result (no `Searching:` placeholder)."""
sse_events = [
{
"type": "response.output_item.done",
"item": {
"type": "web_search_call",
"id": "ws_only",
"action": {},
},
},
{"type": "response.completed", "response": {}},
]
lines = _drive_stream(sse_events, ["web_search"], monkeypatch)
events = _tool_events(lines)
ends = [e for e in events if e["type"] == "tool_end"]
assert len(ends) == 1
assert ends[0]["result"] == ""
# ── shell_call output fallbacks ────────────────────────────────────────
def test_shell_call_emits_tool_end_when_output_bundled_on_done(monkeypatch):
"""Output bundled on the shell_call done event emits tool_end."""
sse_events = [
{
"type": "response.output_item.done",
"item": {
"type": "shell_call",
"id": "scall_bundled",
"action": {"commands": ["echo hi"]},
"output": [
{
"stdout": "hi\n",
"stderr": "",
"outcome": {"type": "exit", "exit_code": 0},
}
],
},
},
{"type": "response.completed", "response": {}},
]
lines = _drive_stream(sse_events, ["code_execution"], monkeypatch)
events = _tool_events(lines)
starts = [e for e in events if e["type"] == "tool_start"]
ends = [e for e in events if e["type"] == "tool_end"]
assert len(starts) == 1
assert starts[0]["tool_call_id"] == "scall_bundled"
assert len(ends) == 1
assert ends[0]["tool_call_id"] == "scall_bundled"
assert "hi" in ends[0]["result"]
def test_shell_call_bundled_then_separate_output_does_not_double_emit(monkeypatch):
"""Separate shell_call_output after bundled-output is a no-op."""
sse_events = [
{
"type": "response.output_item.done",
"item": {
"type": "shell_call",
"id": "scall_both",
"action": {"commands": ["echo bundle"]},
"output": [
{
"stdout": "bundle\n",
"stderr": "",
"outcome": {"type": "exit", "exit_code": 0},
}
],
},
},
{
"type": "response.output_item.done",
"item": {
"type": "shell_call_output",
"id": "scout_both",
"call_id": "scall_both",
"output": [
{
"stdout": "should not double-emit\n",
"stderr": "",
"outcome": {"type": "exit", "exit_code": 0},
}
],
},
},
{"type": "response.completed", "response": {}},
]
lines = _drive_stream(sse_events, ["code_execution"], monkeypatch)
events = _tool_events(lines)
ends = [e for e in events if e["type"] == "tool_end"]
assert len(ends) == 1
assert ends[0]["tool_call_id"] == "scall_both"
assert "bundle" in ends[0]["result"]
assert "should not double-emit" not in ends[0]["result"]
def test_shell_call_final_flush_on_completed_when_no_output_event(monkeypatch):
"""Orphan shell_call finalises via the response.completed flush."""
sse_events = [
{
"type": "response.output_item.added",
"item": {
"type": "shell_call",
"id": "scall_orphan",
"action": {"commands": ["true"]},
},
},
{
"type": "response.output_item.done",
"item": {
"type": "shell_call",
"id": "scall_orphan",
"action": {"commands": ["true"]},
"status": "completed",
},
},
{"type": "response.completed", "response": {}},
]
lines = _drive_stream(sse_events, ["code_execution"], monkeypatch)
events = _tool_events(lines)
ends = [e for e in events if e["type"] == "tool_end"]
assert any(e["tool_call_id"] == "scall_orphan" for e in ends)
def test_shell_call_flushed_on_response_incomplete_truncation(monkeypatch):
"""Truncated streams (response.incomplete) also flush orphan calls."""
sse_events = [
{
"type": "response.output_item.added",
"item": {
"type": "shell_call",
"id": "scall_truncated",
"action": {"commands": ["long_running"]},
},
},
{
"type": "response.output_item.done",
"item": {
"type": "shell_call",
"id": "scall_truncated",
"action": {"commands": ["long_running"]},
"status": "in_progress",
},
},
{
"type": "response.incomplete",
"response": {
"incomplete_details": {"reason": "max_output_tokens"},
},
},
]
lines = _drive_stream(sse_events, ["code_execution"], monkeypatch)
events = _tool_events(lines)
ends = [e for e in events if e["type"] == "tool_end"]
assert any(e["tool_call_id"] == "scall_truncated" for e in ends)
def test_shell_call_incomplete_does_not_double_emit(monkeypatch):
"""response.incomplete is idempotent against already-finalised calls."""
sse_events = [
{
"type": "response.output_item.done",
"item": {
"type": "shell_call",
"id": "scall_done",
"action": {"commands": ["echo done"]},
"output": [
{
"stdout": "done\n",
"stderr": "",
"outcome": {"type": "exit", "exit_code": 0},
}
],
},
},
{
"type": "response.incomplete",
"response": {
"incomplete_details": {"reason": "max_output_tokens"},
},
},
]
lines = _drive_stream(sse_events, ["code_execution"], monkeypatch)
events = _tool_events(lines)
ends = [e for e in events if e["type"] == "tool_end"]
assert len(ends) == 1
assert ends[0]["tool_call_id"] == "scall_done"
assert "done" in ends[0]["result"]