unsloth/studio/backend/core/inference
Daniel Han 0be1e09421 Clean local stop list + cap Responses bridge parallel tool calls (PR #5711)
Round 17 reviewer consensus on two extensions of the round 16 cap:

1. Direct GGUF stop forwarding (3/10 + sibling findings) — every
   llama_cpp.py payload builder and routes/inference.py direct-GGUF
   call site pass `stop` through unfiltered, while the external
   provider helper and `_build_passthrough_payload` already strip
   empty / non-string entries. Add a shared `_clean_local_stop_list`
   helper in the route layer for the two callers, mirror the same
   inline filter in `llama_cpp.py`'s three payload builders. Stops
   `stop=["", "END"]` from a stale client 400'ing llama-server.

2. Responses bridge tool-call cap (3/10 streaming + 1/10 non-streaming)
   — `_responses_stream` iterated every streamed `delta.tool_calls`
   index and `_responses_non_streaming` translated every returned
   `message.tool_calls` entry, even when `parallel_tool_calls=false`.
   Latch the first tool-call index in the streaming bridge and drop
   subsequent siblings; cap to one in the non-streaming bridge.
   Matches the GGUF agentic-loop / Anthropic-passthrough caps.

Local-passthrough OpenAI paths (verbatim SSE / verbatim JSON) are
left alone because the contract is "raw upstream forwarding"; clients
calling /v1/chat/completions through Studio directly should still see
llama-server's native output.
2026-05-24 19:39:37 +00:00
..
__init__.py Final cleanup 2026-03-12 18:28:04 +00:00
_html_to_md.py fix: studio web search SSL failures and empty page content (#4754) 2026-04-01 06:12:02 -07:00
anthropic_compat.py Apply parallel_tool_calls cap to Anthropic passthrough + safetensors path (PR #5711) 2026-05-24 18:46:07 +00:00
audio_codecs.py Add native GGUF intake to Studio (#5246) 2026-05-04 11:46:18 +02:00
chat_template_helpers.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00
defaults.py studio: engage draft-mtp on vision MTP GGUFs (drop incorrect vision gate) (#5560) 2026-05-18 08:42:55 -07:00
external_provider.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-05-24 15:55:38 +00:00
inference.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00
key_exchange.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
llama_cpp.py Clean local stop list + cap Responses bridge parallel tool calls (PR #5711) 2026-05-24 19:39:37 +00:00
llama_server_args.py studio: add --spec-draft-n-max toggle for MTP speculative decoding (#5582) 2026-05-19 06:17:04 -07:00
mlx_inference.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00
orchestrator.py Apply parallel_tool_calls cap to Anthropic passthrough + safetensors path (PR #5711) 2026-05-24 18:46:07 +00:00
pricing.py Studio: per-session cost calculator + /api/providers/pricing endpoint (#5690) 2026-05-22 06:03:43 -07:00
providers.py Gemini stop cap is 4, matching the OpenAI compat layer (PR #5711) 2026-05-24 16:39:49 +00:00
safetensors_agentic.py Apply parallel_tool_calls cap to Anthropic passthrough + safetensors path (PR #5711) 2026-05-24 18:46:07 +00:00
tool_call_parser.py Revert "studio: tool calling for Llama-3, Mistral, Gemma 4 on safetensors + MLX (#5615)" (#5619) 2026-05-19 07:26:39 -07:00
tools.py studio: tighten sandbox blocklist precision (bash, hf upload, NOFILE) (#5487) 2026-05-18 00:01:17 -07:00
worker.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00