unsloth/studio/backend/core/inference
Daniel Han e2b7f5958b Studio: round 9 -- three P2 fixes from latest Codex bot review
1. Codex SSE wrapper terminates on exact `data: [DONE]` only.

   The old substring check `if "[DONE]" in line` would flip
   sent_done True when a normal model response carried the literal
   text "[DONE]" in delta.content (for example an explanation of
   the OpenAI stream sentinel). The real terminator was then
   suppressed, leaving OpenAI-compatible clients that finalise on
   the explicit sentinel hung on stream close. Now compares the
   stripped line to the exact `data: [DONE]` form.

2. Legacy `thread.run_streaming` path no longer returns an empty
   reply on completion-only streams.

   If the SDK exposes `thread.run_streaming` but the stream emits
   ONLY item.completed / agentMessage events with no message
   deltas, the loop previously exited with emitted_any False and
   never reached the agent-message fallback. The request returned
   200 with an empty assistant reply even though Codex produced a
   final answer. Mirror the canonical-path behavior: collect
   `_completed_agent_message_text` strings in a sidecar list and
   emit the last one when no deltas arrived. Match the canonical
   payload-extraction (`getattr(event, "payload", event)`) so the
   event-vs-payload SDK shape difference is handled the same way
   in both branches.

3. Parallel-calls fan-out propagates CodexUnavailableError so the
   route layer can return 503.

   When the SDK is not importable or the safety enums are missing
   without the dev opt-in, every worker raised the same
   CodexUnavailableError. The previous catch-all converted the
   error into a per-tab codex_tab_error event, the outer stream
   never raised, and clients saw a 200 with only tool events and
   an empty synthesis -- OpenAI-compatible consumers that ignore
   _toolEvent saw a successful empty reply. Now CodexUnavailableError
   re-raises out of the worker (no spurious per-tab error event),
   _await_workers re-raises it when EVERY worker hit the same
   setup failure, and the finally-block drain await propagates the
   exception out of the parallel function so the route's existing
   CodexUnavailableError handler can emit the right 503 SSE error
   frame. Per-tab runtime failures (model rejected, timeout, mid-
   stream SDK crash) still get swallowed into codex_tab_error
   events so a single bad model in the fan-out does not kill the
   others.

Test counts: 63/63 passing (60 round 6-8 plus 3 new round 9 regression
tests). Each new test was first run against a `git stash`-restored
pre-fix tree to confirm it catches the bug, then run against the
patched tree.
2026-05-27 06:58:06 +00:00
..
__init__.py Final cleanup 2026-03-12 18:28:04 +00:00
_html_to_md.py fix: studio web search SSL failures and empty page content (#4754) 2026-04-01 06:12:02 -07:00
anthropic_compat.py Studio: Claude Code Anthropic API tool compatibility (#5390) 2026-05-21 16:45:05 +04:00
audio_codecs.py Add native GGUF intake to Studio (#5246) 2026-05-04 11:46:18 +02:00
chat_template_helpers.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00
codex_availability.py Studio: round 7b -- tighten device-login log filter + harden timeout kill 2026-05-25 14:45:26 +00:00
codex_provider.py Studio: round 9 -- three P2 fixes from latest Codex bot review 2026-05-27 06:58:06 +00:00
defaults.py studio: engage draft-mtp on vision MTP GGUFs (drop incorrect vision gate) (#5560) 2026-05-18 08:42:55 -07:00
external_provider.py Studio: per-card web_search result + shell_call output fallback (OpenAI) (#5785) 2026-05-26 04:31:22 -07:00
inference.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00
key_exchange.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
llama_cpp.py Studio: stop the model from replying twice when it refuses (#5775) 2026-05-26 05:30:52 -07:00
llama_server_args.py Studio: expose --parallel / -np flag on unsloth studio run (#5737) 2026-05-26 23:13:45 -07:00
mlx_inference.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00
orchestrator.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00
pricing.py Studio: pricing follow-up to #5690 (longest-prefix match + chat-style usage keys) (#5722) 2026-05-25 23:39:58 -07:00
providers.py Studio: round 4 hardening for the Codex provider 2026-05-24 15:44:11 +00:00
safetensors_agentic.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00
tool_call_parser.py Revert "studio: tool calling for Llama-3, Mistral, Gemma 4 on safetensors + MLX (#5615)" (#5619) 2026-05-19 07:26:39 -07:00
tools.py studio: tighten sandbox blocklist precision (bash, hf upload, NOFILE) (#5487) 2026-05-18 00:01:17 -07:00
worker.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00