* fix(studio): forward OpenAI tools/tool_choice to llama-server (#4999)
Studio's /v1/chat/completions silently stripped standard OpenAI `tools`
and `tool_choice` fields, so clients using standard function calling
(opencode, Claude Code, Cursor, Continue, ...) never got structured
tool_calls back. Adds a client-side pass-through path mirroring the
existing Anthropic /v1/messages flow: when `tools` is present without
Studio's `enable_tools` shorthand, the request is forwarded to
llama-server verbatim so the client sees native id, finish_reason
("tool_calls"), delta.tool_calls, and accurate usage tokens.
Also wires Anthropic tool_choice forwarding: /v1/messages previously
accepted tool_choice on the request model but silently dropped it with
a warning. Translate the four Anthropic shapes to OpenAI format and
forward them so agentic clients can actually enforce tool use.
- ChatCompletionRequest: add tools, tool_choice, stop; extra="allow"
- ChatMessage: accept role="tool", optional tool_call_id / tool_calls /
name; content is now optional (assistant with only tool_calls)
- routes/inference.py: _openai_passthrough_stream /
_openai_passthrough_non_streaming helpers, routing branch in
openai_chat_completions, vision+tools via content-parts injection
- _build_passthrough_payload: tool_choice parameter (default "auto")
- anthropic_compat: anthropic_tool_choice_to_openai() translator
- tests/test_openai_tool_passthrough.py: Pydantic + translator unit tests
- tests/test_studio_api.py: 5 new E2E tests (non-stream, stream,
multi-turn, OpenAI SDK, Anthropic tool_choice=any regression)
* fix(studio): surface httpx transport errors from OpenAI passthrough
When the managed llama-server subprocess crashes mid-request, the
async pass-through helpers in routes/inference.py used to return a
bare 500 (non-streaming) or an "An internal error occurred" SSE chunk
(streaming) because _friendly_error only recognized the sync path's
"Lost connection to llama-server" substring -- httpx transport
failures (ConnectError / ReadError / RemoteProtocolError /
ReadTimeout) stringify differently and fell through to the generic
case.
- _friendly_error: map any httpx.RequestError subclass to the same
"Lost connection to the model server" message the sync chat path
emits. Placed before the substring heuristics so the streaming path
automatically picks it up via its existing except Exception catch.
- _openai_passthrough_non_streaming: wrap the httpx.AsyncClient.post
in a try/except httpx.RequestError and re-raise as HTTPException
502 with the friendly detail.
- tests/test_openai_tool_passthrough.py: new TestFriendlyErrorHttpx
class pinning the mapping for ConnectError, ReadError,
RemoteProtocolError, ReadTimeout, and confirming non-httpx paths
(context-size heuristic, generic fallback) are unchanged.
* fix(studio): close aiter_bytes/aiter_lines explicitly in passthroughs
The httpcore asyncgen cleanup fix in 5cedd9a5 is incomplete on Python
3.13 + httpcore 1.0.x: it switched to manual client/response lifecycle
but still used anonymous `async for raw_line in resp.aiter_lines():`
patterns in all three streaming paths. Python's async for does NOT
auto-close the iterator on break/return, so the aiter_lines /
aiter_bytes async generator remains alive, reachable only from the
surrounding coroutine frame. Once `_stream()` returns the frame is
GC'd and the orphaned asyncgen is finalized on a LATER GC pass in a
DIFFERENT asyncio task, where httpcore's
HTTP11ConnectionByteStream.aclose() enters anyio.CancelScope.__exit__
with a mismatched task and prints "Exception ignored in: <async
generator>" / "async generator ignored GeneratorExit" / "Attempted
to exit cancel scope in a different task" to the server log.
User observed this on /v1/messages after successful (status 200)
requests, with the traceback pointing at HTTP11ConnectionByteStream
.__aiter__ / .aclose inside httpcore.
Fix: save resp.aiter_lines() / resp.aiter_bytes() as a variable and
explicitly `await iter.aclose()` in the finally block BEFORE
resp.aclose() / client.aclose(). This closes the asyncgen inside the
current task's event loop, so the internal httpcore byte stream is
cleaned up before Python's asyncgen GC hook has anything orphaned to
finalize. Each aclose is wrapped in try/except Exception so nested
anyio cleanup noise can't bubble out.
Applied to all three streaming passthrough paths:
- _anthropic_passthrough_stream (/v1/messages client-side tool path)
- _openai_passthrough_stream (/v1/chat/completions client-side tool
path, new in this PR)
- openai_completions (/v1/completions bytes proxy from PR #4956)
* fix(studio): default ChatCompletionRequest.stream to false per OpenAI spec
OpenAI's /v1/chat/completions spec defaults `stream` to false, so
clients that omit the field (naive curl, minimal integrations) expect
a single JSON response back. Studio was defaulting to true, silently
switching those clients into SSE and breaking any parser that didn't
also handle streaming. ResponsesRequest and AnthropicMessagesRequest
already default to false correctly; only ChatCompletionRequest was
wrong.
Studio's own frontend always sets `stream` explicitly on every
chat-adapter / chat-api / runtime-provider call site, so the flip has
no UI impact. SDK users (OpenAI Python/JS SDK, opencode, Claude Code,
Cursor, Continue) also always pass `stream` explicitly, so they're
unaffected. The only clients feeling the change are raw-curl users
who were relying on the wrong default -- those get the correct OpenAI
behavior now.
Added a regression test pinning the default so it can't silently
flip back.
* fix(studio): reject images in OpenAI tool passthrough for text-only GGUFs
The new tool passthrough branch runs before _extract_content_parts,
skipping the existing not is_vision guard. Requests combining tools
with an image on a text-only tool-capable GGUF were forwarded to
llama-server, producing opaque upstream errors instead of the
pre-existing clear 400. Restore the guard inline at the dispatch
point, checking both legacy image_base64 and inline image_url parts.
* fix(studio): require tool_call_id on role=tool chat messages
Enforce the OpenAI spec rule that role="tool" messages must carry a
tool_call_id. Without it, upstream backends cannot associate a tool
result with the assistant's prior tool_calls entry and the request
fails in non-obvious ways through the passthrough path. Reject at the
request boundary with a 422 instead.
* fix(studio): harden OpenAI tool passthrough validation and error surfacing
Three related fixes called out by the PR review:
1. Preserve upstream status codes in the streaming passthrough. The
httpx request is now dispatched before the StreamingResponse is
constructed. Non-200 upstream responses and httpx RequestError
transport failures raise HTTPException with the real status
instead of being buried inside a 200 SSE error frame, so OpenAI
SDK clients see APIError/BadRequestError/... as expected.
2. Require non-empty content on user/system/tool messages. Per the
OpenAI spec, content may only be omitted on assistant messages
that carry tool_calls; enforce that at the request boundary so
malformed messages never reach the passthrough path.
3. Role-constrain tool-call metadata. tool_calls is only valid on
role=assistant, tool_call_id and name only on role=tool. Without
this, a user/system message with tool_calls would flip the
passthrough branch on and be forwarded to llama-server, surfacing
as an opaque upstream error.
* fix(studio): normalize image mode and passthrough JSON verbatim
Two Gemini-code-assist review findings on PR #5099:
1. Unconditionally convert decoded images to RGB before PNG encoding.
The prior code only handled RGBA, letting CMYK/I/F images crash
at img.save(format="PNG") and surface as opaque 400s. Applied to
both the passthrough helper and the non-passthrough GGUF path
that originally carried this pattern, keeping the two sites in
sync.
2. Return the upstream JSON body as raw bytes via Response rather
than parse-then-re-serialize with JSONResponse. Matches the
passthrough helper's "verbatim" contract and drops a redundant
round-trip.
---------
Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
521 lines
18 KiB
Python
521 lines
18 KiB
Python
# SPDX-License-Identifier: AGPL-3.0-only
|
|
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved.
|
|
|
|
"""
|
|
Anthropic Messages API ↔ OpenAI format translation utilities.
|
|
|
|
Pure functions and a stateful stream emitter — no FastAPI, no I/O.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import json
|
|
from typing import Any, Optional, Union
|
|
|
|
|
|
def anthropic_messages_to_openai(
|
|
messages: list[dict],
|
|
system: Optional[Union[str, list]] = None,
|
|
) -> list[dict]:
|
|
"""Convert Anthropic messages + system to OpenAI-format message dicts."""
|
|
result: list[dict] = []
|
|
|
|
# System prompt
|
|
if system:
|
|
if isinstance(system, str):
|
|
result.append({"role": "system", "content": system})
|
|
elif isinstance(system, list):
|
|
parts = []
|
|
for block in system:
|
|
if isinstance(block, dict) and block.get("type") == "text":
|
|
parts.append(block["text"])
|
|
elif isinstance(block, str):
|
|
parts.append(block)
|
|
if parts:
|
|
result.append({"role": "system", "content": "\n".join(parts)})
|
|
|
|
for msg in messages:
|
|
role = msg["role"] if isinstance(msg, dict) else msg.role
|
|
content = msg["content"] if isinstance(msg, dict) else msg.content
|
|
|
|
if isinstance(content, str):
|
|
result.append({"role": role, "content": content})
|
|
continue
|
|
|
|
# Content is a list of blocks
|
|
text_parts: list[str] = []
|
|
tool_calls: list[dict] = []
|
|
tool_results: list[dict] = []
|
|
|
|
for block in content:
|
|
b = block if isinstance(block, dict) else block.model_dump()
|
|
btype = b.get("type", "")
|
|
|
|
if btype == "text":
|
|
text_parts.append(b["text"])
|
|
elif btype == "tool_use":
|
|
tool_calls.append(
|
|
{
|
|
"id": b["id"],
|
|
"type": "function",
|
|
"function": {
|
|
"name": b["name"],
|
|
"arguments": json.dumps(b["input"]),
|
|
},
|
|
}
|
|
)
|
|
elif btype == "tool_result":
|
|
tc = b.get("content", "")
|
|
if isinstance(tc, list):
|
|
tc = " ".join(
|
|
p["text"]
|
|
for p in tc
|
|
if isinstance(p, dict) and p.get("type") == "text"
|
|
)
|
|
tool_results.append(
|
|
{
|
|
"role": "tool",
|
|
"tool_call_id": b["tool_use_id"],
|
|
"content": str(tc),
|
|
}
|
|
)
|
|
|
|
if role == "assistant":
|
|
msg_dict: dict[str, Any] = {"role": "assistant"}
|
|
if text_parts:
|
|
msg_dict["content"] = "\n".join(text_parts)
|
|
if tool_calls:
|
|
msg_dict["tool_calls"] = tool_calls
|
|
result.append(msg_dict)
|
|
elif role == "user":
|
|
if text_parts:
|
|
result.append({"role": "user", "content": "\n".join(text_parts)})
|
|
for tr in tool_results:
|
|
result.append(tr)
|
|
|
|
return result
|
|
|
|
|
|
def anthropic_tools_to_openai(tools: list) -> list[dict]:
|
|
"""Convert Anthropic tool definitions to OpenAI function-tool format."""
|
|
result = []
|
|
for t in tools:
|
|
td = t if isinstance(t, dict) else t.model_dump()
|
|
result.append(
|
|
{
|
|
"type": "function",
|
|
"function": {
|
|
"name": td["name"],
|
|
"description": td.get("description", ""),
|
|
"parameters": td.get("input_schema", {}),
|
|
},
|
|
}
|
|
)
|
|
return result
|
|
|
|
|
|
def anthropic_tool_choice_to_openai(tc: Any) -> Any:
|
|
"""Translate Anthropic `tool_choice` into OpenAI `tool_choice`.
|
|
|
|
Anthropic formats (all dict shapes with a ``type`` discriminator):
|
|
|
|
- ``{"type": "auto"}`` → ``"auto"``
|
|
- ``{"type": "any"}`` → ``"required"``
|
|
- ``{"type": "none"}`` → ``"none"``
|
|
- ``{"type": "tool", "name": "get_weather"}``
|
|
→ ``{"type": "function", "function": {"name": "get_weather"}}``
|
|
|
|
Returns ``None`` for ``None`` or any unrecognized shape (caller may
|
|
then fall back to its own default, typically ``"auto"``).
|
|
"""
|
|
if tc is None:
|
|
return None
|
|
if not isinstance(tc, dict):
|
|
return None
|
|
t = tc.get("type")
|
|
if t == "auto":
|
|
return "auto"
|
|
if t == "any":
|
|
return "required"
|
|
if t == "none":
|
|
return "none"
|
|
if t == "tool":
|
|
name = tc.get("name")
|
|
if not name:
|
|
return None
|
|
return {"type": "function", "function": {"name": name}}
|
|
return None
|
|
|
|
|
|
def build_anthropic_sse_event(event_type: str, data: dict) -> str:
|
|
"""Format a single Anthropic SSE event."""
|
|
return f"event: {event_type}\ndata: {json.dumps(data)}\n\n"
|
|
|
|
|
|
class AnthropicStreamEmitter:
|
|
"""Converts generator events from generate_chat_completion_with_tools()
|
|
into Anthropic Messages SSE strings."""
|
|
|
|
def __init__(self) -> None:
|
|
self.block_index: int = 0
|
|
self._text_block_open: bool = False
|
|
self._prev_text: str = ""
|
|
self._usage: dict = {}
|
|
|
|
def start(self, message_id: str, model: str) -> list[str]:
|
|
"""Emit message_start and open the first text content block."""
|
|
events = []
|
|
events.append(
|
|
build_anthropic_sse_event(
|
|
"message_start",
|
|
{
|
|
"type": "message_start",
|
|
"message": {
|
|
"id": message_id,
|
|
"type": "message",
|
|
"role": "assistant",
|
|
"content": [],
|
|
"model": model,
|
|
"stop_reason": None,
|
|
"stop_sequence": None,
|
|
"usage": {"input_tokens": 0, "output_tokens": 0},
|
|
},
|
|
},
|
|
)
|
|
)
|
|
events.extend(self._open_text_block())
|
|
return events
|
|
|
|
def feed(self, event: dict) -> list[str]:
|
|
"""Process one generator event, return SSE strings."""
|
|
etype = event.get("type", "")
|
|
if etype == "content":
|
|
return self._handle_content(event)
|
|
elif etype == "tool_start":
|
|
return self._handle_tool_start(event)
|
|
elif etype == "tool_end":
|
|
return self._handle_tool_end(event)
|
|
elif etype == "metadata":
|
|
self._usage = event.get("usage", {})
|
|
return []
|
|
# status events — no Anthropic equivalent
|
|
return []
|
|
|
|
def finish(self, stop_reason: str = "end_turn") -> list[str]:
|
|
"""Close any open block and emit message_delta + message_stop."""
|
|
events = []
|
|
if self._text_block_open:
|
|
events.append(self._close_block())
|
|
events.append(
|
|
build_anthropic_sse_event(
|
|
"message_delta",
|
|
{
|
|
"type": "message_delta",
|
|
"delta": {"stop_reason": stop_reason, "stop_sequence": None},
|
|
"usage": {
|
|
"output_tokens": self._usage.get("completion_tokens", 0),
|
|
},
|
|
},
|
|
)
|
|
)
|
|
events.append(
|
|
build_anthropic_sse_event(
|
|
"message_stop",
|
|
{
|
|
"type": "message_stop",
|
|
},
|
|
)
|
|
)
|
|
return events
|
|
|
|
def _handle_content(self, event: dict) -> list[str]:
|
|
cumulative = event.get("text", "")
|
|
new_text = cumulative[len(self._prev_text) :]
|
|
self._prev_text = cumulative
|
|
if not new_text:
|
|
return []
|
|
if not self._text_block_open:
|
|
events = self._open_text_block()
|
|
else:
|
|
events = []
|
|
events.append(
|
|
build_anthropic_sse_event(
|
|
"content_block_delta",
|
|
{
|
|
"type": "content_block_delta",
|
|
"index": self.block_index,
|
|
"delta": {"type": "text_delta", "text": new_text},
|
|
},
|
|
)
|
|
)
|
|
return events
|
|
|
|
def _handle_tool_start(self, event: dict) -> list[str]:
|
|
events = []
|
|
# Close current text block if open
|
|
if self._text_block_open:
|
|
events.append(self._close_block())
|
|
# Open a tool_use block
|
|
self.block_index += 1
|
|
events.append(
|
|
build_anthropic_sse_event(
|
|
"content_block_start",
|
|
{
|
|
"type": "content_block_start",
|
|
"index": self.block_index,
|
|
"content_block": {
|
|
"type": "tool_use",
|
|
"id": event.get("tool_call_id", ""),
|
|
"name": event.get("tool_name", ""),
|
|
"input": {},
|
|
},
|
|
},
|
|
)
|
|
)
|
|
# Emit the arguments as input_json_delta
|
|
args = event.get("arguments", {})
|
|
if args:
|
|
events.append(
|
|
build_anthropic_sse_event(
|
|
"content_block_delta",
|
|
{
|
|
"type": "content_block_delta",
|
|
"index": self.block_index,
|
|
"delta": {
|
|
"type": "input_json_delta",
|
|
"partial_json": json.dumps(args),
|
|
},
|
|
},
|
|
)
|
|
)
|
|
return events
|
|
|
|
def _handle_tool_end(self, event: dict) -> list[str]:
|
|
events = []
|
|
# Close the tool_use block
|
|
events.append(self._close_block())
|
|
# Emit custom tool_result event (non-standard, ignored by SDKs)
|
|
events.append(
|
|
build_anthropic_sse_event(
|
|
"tool_result",
|
|
{
|
|
"type": "tool_result",
|
|
"tool_use_id": event.get("tool_call_id", ""),
|
|
"content": event.get("result", ""),
|
|
},
|
|
)
|
|
)
|
|
# Open a new text block for the model's next response
|
|
self.block_index += 1
|
|
events.extend(self._open_text_block())
|
|
# Reset text tracking for the next synthesis turn
|
|
self._prev_text = ""
|
|
return events
|
|
|
|
def _open_text_block(self) -> list[str]:
|
|
self._text_block_open = True
|
|
return [
|
|
build_anthropic_sse_event(
|
|
"content_block_start",
|
|
{
|
|
"type": "content_block_start",
|
|
"index": self.block_index,
|
|
"content_block": {"type": "text", "text": ""},
|
|
},
|
|
)
|
|
]
|
|
|
|
def _close_block(self) -> str:
|
|
self._text_block_open = False
|
|
return build_anthropic_sse_event(
|
|
"content_block_stop",
|
|
{
|
|
"type": "content_block_stop",
|
|
"index": self.block_index,
|
|
},
|
|
)
|
|
|
|
|
|
class AnthropicPassthroughEmitter:
|
|
"""Converts llama-server's OpenAI-format streaming chunks into Anthropic SSE.
|
|
|
|
Used for the client-side tool-use pass-through path: the client (e.g. Claude
|
|
Code) sends its own tool definitions in the ``tools`` field and expects to
|
|
execute them itself. We forward them to llama-server and translate the
|
|
streaming response back to Anthropic format without executing anything.
|
|
"""
|
|
|
|
def __init__(self) -> None:
|
|
self.block_index: int = -1
|
|
self._current_block_type: Optional[str] = None # "text" | "tool_use" | None
|
|
self._tool_call_states: dict = {} # delta index -> {block_index, id, name}
|
|
self._usage: dict = {}
|
|
self._stop_reason: str = "end_turn"
|
|
|
|
def start(self, message_id: str, model: str) -> list[str]:
|
|
return [
|
|
build_anthropic_sse_event(
|
|
"message_start",
|
|
{
|
|
"type": "message_start",
|
|
"message": {
|
|
"id": message_id,
|
|
"type": "message",
|
|
"role": "assistant",
|
|
"content": [],
|
|
"model": model,
|
|
"stop_reason": None,
|
|
"stop_sequence": None,
|
|
"usage": {"input_tokens": 0, "output_tokens": 0},
|
|
},
|
|
},
|
|
)
|
|
]
|
|
|
|
def feed_chunk(self, chunk: dict) -> list[str]:
|
|
"""Process one OpenAI streaming chat.completion.chunk."""
|
|
events: list[str] = []
|
|
|
|
# usage-only chunks carry token totals
|
|
usage = chunk.get("usage")
|
|
if usage:
|
|
self._usage = usage
|
|
|
|
choices = chunk.get("choices") or []
|
|
if not choices:
|
|
return events
|
|
|
|
choice = choices[0]
|
|
delta = choice.get("delta") or {}
|
|
finish_reason = choice.get("finish_reason")
|
|
|
|
# ── Text content ──
|
|
content = delta.get("content")
|
|
if content:
|
|
if self._current_block_type != "text":
|
|
if self._current_block_type is not None:
|
|
events.append(self._close_current_block())
|
|
events.extend(self._open_text_block())
|
|
events.append(
|
|
build_anthropic_sse_event(
|
|
"content_block_delta",
|
|
{
|
|
"type": "content_block_delta",
|
|
"index": self.block_index,
|
|
"delta": {"type": "text_delta", "text": content},
|
|
},
|
|
)
|
|
)
|
|
|
|
# ── Tool calls (streaming deltas) ──
|
|
tool_calls = delta.get("tool_calls") or []
|
|
for tc in tool_calls:
|
|
tc_idx = tc.get("index", 0)
|
|
fn = tc.get("function") or {}
|
|
if tc_idx not in self._tool_call_states:
|
|
# New tool call — close prior block, open tool_use block
|
|
if self._current_block_type is not None:
|
|
events.append(self._close_current_block())
|
|
tc_id = tc.get("id", "")
|
|
tc_name = fn.get("name", "")
|
|
self.block_index += 1
|
|
self._current_block_type = "tool_use"
|
|
self._tool_call_states[tc_idx] = {
|
|
"block_index": self.block_index,
|
|
"id": tc_id,
|
|
"name": tc_name,
|
|
}
|
|
events.append(
|
|
build_anthropic_sse_event(
|
|
"content_block_start",
|
|
{
|
|
"type": "content_block_start",
|
|
"index": self.block_index,
|
|
"content_block": {
|
|
"type": "tool_use",
|
|
"id": tc_id,
|
|
"name": tc_name,
|
|
"input": {},
|
|
},
|
|
},
|
|
)
|
|
)
|
|
|
|
args_delta = fn.get("arguments", "")
|
|
if args_delta:
|
|
events.append(
|
|
build_anthropic_sse_event(
|
|
"content_block_delta",
|
|
{
|
|
"type": "content_block_delta",
|
|
"index": self._tool_call_states[tc_idx]["block_index"],
|
|
"delta": {
|
|
"type": "input_json_delta",
|
|
"partial_json": args_delta,
|
|
},
|
|
},
|
|
)
|
|
)
|
|
|
|
# ── Finish reason ──
|
|
if finish_reason:
|
|
if finish_reason == "tool_calls":
|
|
self._stop_reason = "tool_use"
|
|
elif finish_reason == "length":
|
|
self._stop_reason = "max_tokens"
|
|
else:
|
|
self._stop_reason = "end_turn"
|
|
|
|
return events
|
|
|
|
def finish(self) -> list[str]:
|
|
events: list[str] = []
|
|
if self._current_block_type is not None:
|
|
events.append(self._close_current_block())
|
|
events.append(
|
|
build_anthropic_sse_event(
|
|
"message_delta",
|
|
{
|
|
"type": "message_delta",
|
|
"delta": {
|
|
"stop_reason": self._stop_reason,
|
|
"stop_sequence": None,
|
|
},
|
|
"usage": {
|
|
"output_tokens": self._usage.get("completion_tokens", 0),
|
|
},
|
|
},
|
|
)
|
|
)
|
|
events.append(
|
|
build_anthropic_sse_event(
|
|
"message_stop",
|
|
{"type": "message_stop"},
|
|
)
|
|
)
|
|
return events
|
|
|
|
def _open_text_block(self) -> list[str]:
|
|
self.block_index += 1
|
|
self._current_block_type = "text"
|
|
return [
|
|
build_anthropic_sse_event(
|
|
"content_block_start",
|
|
{
|
|
"type": "content_block_start",
|
|
"index": self.block_index,
|
|
"content_block": {"type": "text", "text": ""},
|
|
},
|
|
)
|
|
]
|
|
|
|
def _close_current_block(self) -> str:
|
|
idx = self.block_index
|
|
self._current_block_type = None
|
|
return build_anthropic_sse_event(
|
|
"content_block_stop",
|
|
{
|
|
"type": "content_block_stop",
|
|
"index": idx,
|
|
},
|
|
)
|