Commit graph

450 commits

Author SHA1 Message Date
pre-commit-ci[bot]
cd8284d5cd [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-27 13:31:21 +00:00
Daniel Han
f01011e4dd Studio: real Codex parallel-call tab UI + in-process SDK spoof
Two paired changes that finally make the Codex parallel-calls fan-out
visible as actual clickable tabs in the chat surface, plus a credit-
free spoof that lets the whole pipeline run in dev / CI without ever
touching the upstream API.

1. Real tab UI (frontend).

   The chat-adapter used to render the per-worker outputs as inline
   `[Codex tab 1/N] ...` text blocks in the assistant message body,
   which collapsed into one big run-on block once more than a handful
   of tokens had streamed. Now each `codex_*` SSE event is folded into
   `codexParallelState` and re-published as the `args.state` of a
   single tool-call part with `toolName === "codex_parallel"`. The
   assistant-ui surface dispatches that to the new
   `CodexParallelToolUI` wrapper, which mounts the existing
   `CodexParallelTabs` component -- one tab per worker, one Synthesis
   tab, click to switch. The stable `toolCallId` keeps assistant-ui
   updating the SAME card across stream yields rather than spawning
   new cards.

   `renderCodexTabsBlock` now returns the empty string so the message
   body no longer contains the labelled-text fallback (kept the
   function name so the rest of the adapter's `renderFullContent` /
   pin-signature paths are untouched).

2. Credit-free Codex SDK spoof (backend).

   New `studio/backend/core/inference/codex_spoof.py` exposes a drop-in
   subset of the upstream `openai_codex` surface (`AsyncCodex`,
   `AppServerConfig`, `ApprovalMode.deny_all`, `SandboxMode.read_only`,
   thread with `turn().stream()` + `run_streaming()` + `run()`) and
   emits deterministic per-tab streaming events tagged with the worker
   index, so flipping between tabs in the UI shows visibly distinct
   text. Activated by `UNSLOTH_CODEX_SPOOF=1`; `_import_codex` installs
   the spoof into `sys.modules` under both `openai_codex` and
   `codex_app_server` and the rest of the provider keeps running
   unchanged. OFF by default; production is unaffected.

   Six new tests cover the spoof itself (module install, env-flag
   gating, delta + completion event shape, per-tab tagging, provider
   import path, safety-kwargs resolution against the spoof). 69/69
   tests pass with and without the flag; TypeScript clean.
2026-05-27 13:30:53 +00:00
Daniel Han
ecf8bc767d Merge remote-tracking branch 'origin/main' into feat/codex-provider
# Conflicts:
#	studio/frontend/src/features/chat/api/chat-adapter.ts
2026-05-27 13:22:31 +00:00
Daniel Han
ab48465135
Studio: add Gemini provider with web_search, code_execution, prompt caching, and Nano Banana image generation (#5720)
* Studio: add Gemini provider with web_search, code_execution, prompt caching, and Nano Banana image generation

Wires Google's native Gemini API into Studio's external-provider stack
so users can pick gemini-2.5-pro / gemini-2.5-flash / gemini-2.5-flash-image
(Nano Banana) alongside the existing OpenAI / Anthropic / OpenRouter
providers. Gemini does not speak OpenAI Chat Completions on its primary
endpoint; the new `_stream_gemini` async generator translates between
the two shapes the same way `_stream_anthropic` handles the Messages API.

Backend:
- New `_stream_gemini` translator in external_provider.py. Converts
  OpenAI messages -> Gemini `contents` + `systemInstruction`; maps
  generationConfig (temperature / topP / topK / maxOutputTokens);
  forwards `tools: [{googleSearch: {}}]` for web_search and
  `{codeExecution: {}}` for code_execution; passes `cachedContent`
  through for prompt caching; sets `responseModalities=[TEXT, IMAGE]`
  for Nano Banana image generation.
- Translates streamed `GenerateContentResponse` SSE frames back into
  OpenAI chat.completion.chunk frames (text deltas, function_call ->
  tool_calls deltas, inlineData -> image_b64 tool_end envelope, usage
  chunk before [DONE]).
- Registry entry switched to native base URL
  `https://generativelanguage.googleapis.com/v1beta` with
  `openai_compatible: False` and the `x-goog-api-key` auth header.
  Model lineup curated to current 2.5 / 2.0 family + Nano Banana.

Frontend:
- Provider-capability matrix: Gemini supports temperature, top_p, top_k,
  presence_penalty (matches generationConfig); min_p / repetition_penalty
  hidden because the API does not accept them.
- `providerSupportsBuiltinWebSearch` / `providerSupportsBuiltinCodeExecution`
  / `providerSupportsBuiltinImageGeneration` extended for Gemini.
- Prompt caching toggle now also lit on Gemini.

Tests:
- 21 new tests in `test_gemini_provider.py` using httpx.MockTransport.
  Cover request body shape conversion, URL/header wiring, web_search
  forwarded as googleSearch, function-call translation both directions,
  prompt caching passthrough, image generation emitting image_b64,
  grounded-search citations -> tool_end, finish_reason mapping, and
  vision data URL -> inlineData translation.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: forward presence_penalty to Gemini and recover function name from tool_call_id

Two follow-up fixes for the Gemini provider:

  * Thread presence_penalty into _stream_gemini and set
    generationConfig.presencePenalty when non-zero. The OpenAI-side
    capability matrix already exposes the slider for Gemini, so the
    value was being collected and silently dropped on the way out.

  * When an OpenAI role=tool message omits 'name' and only carries
    'tool_call_id', recover the function name from the matching
    functionCall on the prior assistant turn. Gemini 400s on an empty
    functionResponse name.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: surface Gemini code execution parts as code_execution tool events

The Gemini stream parser only handled text/functionCall/inlineData
parts, so when the user toggled the Code pill on a Gemini model the
sandbox output (executableCode + codeExecutionResult parts) was
dropped on the floor while adjacent text reached the UI. Reviewers
flagged this as the headline feature being silently broken.

Translate both parts into the existing code_execution tool envelope
that CodeExecutionToolUI already consumes for OpenAI / Anthropic:

  * executableCode  -> tool_start with kind=code_execution and the
    source code under arguments.code. We mint a tool_call_id and
    stash it so the matching result block can pair to it.
  * codeExecutionResult -> tool_end on that id with the stdout under
    result. Non-OK outcomes (OUTCOME_FAILED / OUTCOME_DEADLINE_EXCEEDED)
    are prefixed onto the text so the failure is visible.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: native Gemini model catalog, function-call ids, and honest cache claim

Three follow-ups to the Gemini provider PR after the codex pass:

  * list_models() now translates Gemini's native /v1beta/models
    payload ({models[{name, baseModelId, displayName,
    supportedGenerationMethods}]}) into the OpenAI-compatible shape
    Studio expects. Without this the picker stayed empty for Gemini
    and fell back to hardcoded defaults. Embedding-only models are
    filtered out.

  * Forward the OpenAI tool_call id into Gemini's functionCall.id
    and mirror it onto functionResponse.id. Two parallel calls to
    the same function name can now be paired unambiguously on the
    follow-up turn.

  * Drop Gemini from the prompt-caching capability set. The wire
    flow requires a separate cachedContents POST first and the
    boolean Studio emits today is a no-op; the toggle should not
    advertise a feature it cannot apply. Leaves a pointer to the
    docs for the eventual two-step orchestration.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: distinct tool_calls index per emitted Gemini function call

Codex flagged that the Gemini stream parser hardcoded
tool_calls[0].index to 0 on every emitted functionCall. OpenAI
reassemblers key tool_calls by index when joining deltas, so two
parallel function calls in one assistant turn collapsed onto a
single slot and the second call's arguments overwrote the first.

Track the running count via len(emitted_function_call_ids) - 1
and emit it as the per-call index. The dedupe guard above (skip
when fc_id already in the set) means the index is monotonic and
stable for the lifetime of the stream. Regression test asserts
[0, 1] across two parallel calls in one candidate parts list.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: surface Gemini 3.5/3.1/3 + Nano Banana 2/Pro and plumb thinking budget

`gemini-2.0-flash` / `gemini-2.0-flash-exp` were retired by Google in 2026
(`/v1beta/models/gemini-2.0-flash:streamGenerateContent` returns HTTP 404
"no longer available to new users"), and the picker had nothing past the
2.x family. Verified against the live ListModels catalog: drop the retired
ids from `default_models` + allowlist and surface the chat-capable
3.5 / 3.1 / 3 families plus the Nano Banana image trio.

Also plumb `enable_thinking` / `reasoning_effort` into Gemini's
`generationConfig.thinkingConfig`. Without this, Gemini 3.5 Flash,
gemini-pro-latest, and the 3.x previews silently spend the caller's
`max_tokens` budget on hidden "thoughts" before emitting any visible
answer -- the chat shows a truncated stub like "The capital of" and
streams stop. Mapping:
  - enable_thinking=False / reasoning_effort=none -> thinkingBudget=0
    (Flash tier; Pro tier coerces to a small positive budget because
    the API 400s on 0 with "This model only works in thinking mode")
  - minimal/low/medium/high -> 512/2048/8192/24576 budget tokens
  - max/xhigh -> -1 (dynamic)
  - default (neither knob set) -> thinkingConfig omitted, model decides

Frontend `getExternalReasoningCapabilities` now surfaces a
`reasoning_effort` picker for every Gemini chat id (Pro tier hides the
"none" option; image-tier ids stay knob-less). Adds 6 unit tests
covering Flash/Pro effort mapping, the off-toggle coercion on Pro,
default omission, and the nano-banana-pro-preview alias routing
through the image modalities path. 28 -> 34 tests in
`test_gemini_provider.py`, all green; full backend suite still passes
(1459/1460; the unrelated test_help_output flake is pre-existing and
not in any file this PR touches).

Live verification against generativelanguage.googleapis.com on
2026-05-24 with `_stream_gemini` directly:
  text   gemini-3.5-flash           single PASS  multi PASS
  text   gemini-3.1-pro-preview     single PASS  multi PASS
  text   gemini-3.1-flash-lite      single PASS  multi PASS
  text   gemini-3-pro-preview       single PASS  multi PASS
  text   gemini-3-flash-preview     single PASS  multi PASS
  text   gemini-2.5-pro             single PASS  multi PASS
  text   gemini-2.5-flash           single PASS  multi PASS
  text   gemini-2.5-flash-lite      single PASS  multi PASS
  text   gemini-flash-latest        single PASS  multi PASS
  text   gemini-flash-lite-latest   single PASS  multi PASS
  text   gemini-pro-latest          single PASS  multi PASS
  image  gemini-2.5-flash-image     PASS (1082 KB png returned)
  image  gemini-3.1-flash-image-preview  PASS (Nano Banana 2)
  image  gemini-3-pro-image-preview      PASS (Nano Banana Pro)
  tool   web_search                 PASS
  tool   code_execution             PASS
  -> 16/16 e2e through the actual ExternalProviderClient code path.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: tighten Gemini provider after review (PR #5720)

Fixes a batch of bugs surfaced by a second-pass review on top of the
3.5/3.1/3 + Nano Banana 2/Pro additions in c6724dbd.

Backend (external_provider.py):
- Constructor normalises legacy /v1beta/openai base URLs to /v1beta so
  Gemini providers saved before the native switch keep working without
  a manual re-config.
- Skip thinkingConfig, googleSearch, and codeExecution on image-tier
  models (-image / nano-banana). The image responseModalities path is
  mutually exclusive with text-tool wiring and stale UI state would
  otherwise 400 the turn.
- _PRO_THINKING_PREFIXES now includes gemini-3.5-pro and uses anchored
  prefix matching (exact id or "<prefix>-...") so the image-tier
  gemini-3-pro-image-preview cannot accidentally match the pro guard.
- Gemini 3 functionCall thoughtSignature is round-tripped through the
  tool_calls envelope via extra_content.google.thought_signature on
  emit, and replayed as a sibling of functionCall on the next request.
- finishReason swaps STOP -> tool_calls when any functionCall was
  emitted on the same turn so OAI clients trigger tool execution
  (matches the OpenAI Chat Completions contract).
- usageMetadata.thoughtsTokenCount is rolled into output_tokens and
  surfaced on output_tokens_details.reasoning_tokens so total_tokens
  reflects the full billable spend instead of dropping the hidden
  reasoning slice.

Registry (providers.py):
- Drop gemini-3-pro-preview from default_models. Google shut it down
  on 2026-03-09 and auto-redirects to gemini-3.1-pro-preview; we
  surface the canonical id only.
- Add model_id_deny_exact = ("gemini-3-pro-preview",) so the live
  ListModels fetch does not re-surface the redirect alias.

Route schema (models/inference.py):
- enable_prompt_caching widened to Optional[Union[bool, str]] so the
  /v1/chat/completions caller can pass a Gemini cachedContent resource
  name (e.g. cachedContents/abc123). Without this widening _stream_gemini
  s string cachedContent passthrough was unreachable from the public
  route (bool_parsing 422). stream_chat_completion signature mirrors.

Frontend (provider-capabilities.ts, chat-page.tsx, chat-adapter.ts):
- providerSupportsBuiltinImageGeneration now also recognises
  nano-banana ids (nano-banana-pro-preview was hidden from the image
  pill before).
- providerSupportsBuiltinWebSearch takes the model id so Gemini image
  models hide the Search pill (mirrors the backend skip).
- providerSupportsBuiltinCodeExecution uses the same isGeminiImageModel
  guard for nano-banana ids.
- GEMINI_THINKING_PRO_PREFIXES gains gemini-3.5-pro; gemini-3-pro
  tightened to gemini-3-pro-preview to avoid the image-id overlap.
- Updated 3 callers of providerSupportsBuiltinWebSearch to thread the
  selected model id through.

Tests (test_gemini_provider.py): 34 -> 42, all green
- test_image_models_skip_thinking_config
- test_image_models_drop_text_only_tools
- test_gemini_35_pro_recognized_as_pro_thinking
- test_legacy_openai_base_url_normalized
- test_finish_reason_swaps_to_tool_calls_when_function_call_emitted
- test_thought_signature_round_trips_into_gemini_function_call
- test_thought_signature_emitted_in_tool_call_delta
- test_usage_chunk_includes_thoughts_tokens

Verification:
- Backend pytest 1518/1519 passing (one unrelated Qwen3.5 flash-attn
  test fails on main as well; nothing in this PR touches that path).
- Frontend npx tsc -b clean.
- Live e2e 16/16 against generativelanguage.googleapis.com through the
  patched _stream_gemini code path (all 11 chat models single + multi
  turn, all 3 image models returned image bytes, web_search and
  code_execution tools both emit the expected envelope).
- Live /api/providers/models against the patched backend surfaces 16
  ids (gemini-3-pro-preview correctly filtered via deny_exact).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: address second-pass review findings on Gemini (PR #5720)

Round-2 reviewer.py flagged a phantom web_search card on image
turns (12/12 reviewers), route-layer stripping of tool_calls /
tool_call_id / name, an over-narrow image-mode tool guard, and
silent safety blocks. This patch fixes all four.

Backend (external_provider.py):
- web_search_active is now derived from the outbound tools_array
  (whether googleSearch was actually forwarded), not the raw
  enabled_tools intent. Image-mode turns dropped the tool above so
  the inbound stream no longer emits a phantom "search complete"
  tool_start / tool_end on those turns.
- text_tools_allowed now uses is_image_model (covers both `-image`
  / `nano-banana` picker models AND text models that requested
  `image_generation` via enabled_tools). Verified against the live
  Gemini API which rejects both googleSearch and codeExecution
  alongside responseModalities=["TEXT","IMAGE"] with explicit 400s
  ("Search as tool is not enabled for this model", "Code execution
  is not enabled for this model").
- promptFeedback.blockReason is surfaced as a 400 content-filter
  error chunk instead of returning an empty successful assistant
  response. The streaming loop closes the response before exiting.

Route (routes/inference.py):
- _build_external_messages now propagates tool_calls (assistant),
  tool_call_id, and name (tool result) through every code path
  (string content, multimodal content, non-vision fallback). Without
  this Gemini 3 function-call round trips lost their thoughtSignature
  + tool_call_id at the route boundary, and functionResponse.name
  arrived empty on the second turn.
- Assistant messages with content=None and tool_calls populated are
  preserved as a synthetic empty-string content turn so the
  Gemini translator can rebuild the functionCall part.

Tests (test_gemini_provider.py): 42 -> 45, all green
- test_image_models_suppress_phantom_web_search_card
- test_image_generation_tool_drops_text_tools
- test_prompt_feedback_block_reason_surfaces_as_error

Verification:
- Backend pytest 1736 / 1736 (the two pre-existing unrelated fails
  on main, test_help_output and Qwen3.5 flash-attn pin, are skipped).
- Frontend npx tsc -b clean.
- Live e2e 16/16 against generativelanguage.googleapis.com:
  11 chat models single + multi turn, 3 image models returning
  image bytes, web_search and code_execution both PASS.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: fix third-pass Gemini findings (PR #5720)

Round 3 review follow-ups:

Backend (studio/backend/core/inference/external_provider.py):
- Close response AND aiter_lines iterator in a finally so normal,
  prompt-block, and cancellation exits all clean up (eliminates the
  RuntimeWarning about aclose never being awaited).
- Pair the synthetic web_search tool_start with a tool_end on the
  promptFeedback.blockReason path so the UI does not leave a stuck
  "searching..." spinner after the error toast.
- Preserve native id and thoughtSignature on executableCode and
  codeExecutionResult tool events under google.native_part, and pair
  the tool_end on the code-exec id so multi-turn code-execution
  replays do not lose Gemini-required history.
- Carry part-level thoughtSignature on text deltas via
  delta.extra_content.google.thought_signature and on inline image
  tool_end via google.thought_signature so Gemini 3 image editing
  and tool turns round-trip the signature on the next request.
- Guess remote image_url MIME from the URL path so PNG / WebP / GIF
  inputs are not silently relabeled as JPEG.
- Roll usageMetadata.toolUsePromptTokenCount into translated input
  tokens and surface thoughtsTokenCount as
  completion_tokens_details.reasoning_tokens in _build_usage_chunk.
- Only normalize the Google-hosted /v1beta/openai legacy base URL;
  custom proxies whose paths happen to end in /openai are left
  untouched.
- Forward ChatCompletionRequest.tools and tool_choice through
  stream_chat_completion into _stream_gemini, translating to
  tools[].functionDeclarations and toolConfig.functionCallingConfig.

Frontend:
- chat-adapter: when Gemini image-generation is enabled for the turn,
  also disable Search and Code so the request, builder, and active
  pills agree with what the backend actually sends (the backend
  already strips text tools when image_generation is in enabled_tools).
- chat-adapter: consume OpenAI-shape delta.tool_calls chunks so
  Gemini function-call deltas without text surface as tool-call parts.
- shared-composer: disable Search and Code pills while Gemini image
  mode is active so the UI matches the request.

Tests (studio/backend/tests/test_gemini_provider.py): adds coverage
for proxy base-url gating, remote image MIME inference,
toolUsePromptTokenCount, reasoning_tokens propagation, prompt-block
web_search tool_end pairing, native code-exec id/thoughtSignature
metadata, inline image thoughtSignature, text-chunk extra_content,
OpenAI tools/tool_choice translation, and image-model tool drop.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: Gemini 3 thinkingLevel + image-model Search grounding (PR #5720)

Gemini 3.x migrated to a string `thinkingConfig.thinkingLevel`
(MINIMAL/LOW/MEDIUM/HIGH) and rejects `thinkingBudget`+`thinkingLevel`
in the same request. Gemini 3 also cannot turn thinking fully off, so
the lowest position is "minimal" (Flash) or "low" (Pro rejects
"minimal").

- external_provider._stream_gemini: split thinking translation by
  family. Gemini 3.x (3 / 3.1 / 3.5 + gemini-pro-latest /
  gemini-flash-latest / gemini-flash-lite-latest) emits
  thinkingConfig.thinkingLevel; effort none/off coerces to "low" on
  Pro and "minimal" on Flash. Gemini 2.5 stays on thinkingBudget.
- external_provider._stream_gemini: allow `tools: [{googleSearch: {}}]`
  on the Gemini 3 image family (gemini-3-pro-image-preview,
  gemini-3.1-flash-image-preview, nano-banana-pro). Google's docs
  document Search grounding on these. codeExecution stays blocked
  on image mode (still mutually exclusive with responseModalities).
- provider-capabilities.ts: mirror the Gemini 3 effort ladders in
  resolveGeminiReasoningCapabilities (Pro: low/medium/high; Flash:
  minimal/low/medium/high; 2.5 Flash keeps the off-position).
- provider-capabilities.ts: providerSupportsBuiltinWebSearch now
  returns true on the documented Gemini 3 image models so the pill
  is reachable; older image ids (gemini-2.5-flash-image) still hide.

Tests: splits the existing thinkingBudget cases by family (Gemini 3
checks thinkingLevel; Gemini 2.5 keeps thinkingBudget), adds positive
googleSearch coverage for Gemini 3 image models and negative
googleSearch coverage for legacy image models.

References:
- https://ai.google.dev/gemini-api/docs/thinking
- https://ai.google.dev/gemini-api/docs/gemini-3
- https://ai.google.dev/gemini-api/docs/models/gemini-3-pro-image-preview

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: attach Gemini code_execution inline images to the code card (PR #5720)

When a text Gemini turn wires codeExecution and the sandbox produces a
matplotlib plot, the inline image part ships right after the
codeExecutionResult. Previously this surfaced as a separate empty
image_generation card. Track the most recent code_execution
tool_call_id + result text and, when an inline image follows with
code_execution active, emit a second tool_end on the same id that
appends the image as a data: URI under the `__IMAGES__:` marker the
chat-adapter already understands.

Image-picker turns (`-image` / `nano-banana`) keep the standalone
image_generation envelope so Nano Banana outputs render the same way.

Tests: covers the merged code-execution card emission with no
standalone image_generation event when code_execution is the active
tool.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: fix fourth-pass Gemini findings (PR #5720)

Round 4 review follow-ups:

Backend:
- `_is_openai_compatible` + `_auth_headers` detect Gemini connections
  pointed at a custom OpenAI-compatible proxy (non-Google host whose
  path ends in `/openai`) and route them through the OpenAI-compat
  surface with `Authorization: Bearer ...` instead of the native
  `_stream_gemini` translator + `x-goog-api-key`. Google-hosted Gemini
  keeps the native dispatch path it migrated to in this PR.
- `_stream_gemini` thinkingLevel handling for Gemini 3 Pro now coerces
  both "minimal" and "medium" effort to "low" / "high" respectively
  (Pro tier only accepts low/high per
  https://ai.google.dev/gemini-api/docs/thinking).
- `providers.py` `default_models` restores the advertised
  `gemini-3.5-pro` and the rolling `gemini-pro-latest` /
  `gemini-flash-latest` / `gemini-flash-lite-latest` aliases that the
  allowlist already admits.

Frontend:
- chat-adapter: lean on `providerSupportsBuiltinWebSearch` (which
  already encodes the Gemini 3 image-model Search allowance) instead
  of blanket-disabling Search whenever Gemini image mode is active.
  Code execution stays blocked because Gemini image mode rejects it.
- shared-composer: mirror the same gate -- only the Code pill is
  unconditionally disabled in Gemini image mode; the Search pill is
  driven by `supportsBuiltinWebSearch`.
- provider-capabilities: Gemini 3 Pro reasoning levels now expose only
  "low" and "high" (no Medium pill) to match the API.

Tests: covers the Gemini 3 Pro medium / minimal coercion, the custom
proxy OAI-compat dispatch + Authorization Bearer auth, and the
native-vs-proxy detection. Also closes the mocked httpx.AsyncClient
inside the test event loop so the Python 3.13 `aclose was never
awaited` warning no longer fires.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: fix fifth-pass Gemini findings (PR #5720)

Round 5 review follow-ups:

Backend:
- `_is_openai_compatible` + `_auth_headers` now treat ANY non-Google
  Gemini base URL as OpenAI-compat (LiteLLM / custom OAI gateways /
  OpenAI-compat vLLM routers), not just paths ending in `/openai`.
  Pre-existing saved Gemini proxies on `/v1` keep working.
- Gemini 3 thinkingLevel coercion narrowed to the documented
  inconsistencies: only "minimal" is coerced to "low" on Pro tier.
  "medium" passes through (Gemini 3.1 Pro accepts it per
  https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/gemini/3-1-pro).
- `_stream_gemini` only flips `responseModalities=[TEXT,IMAGE]` when
  the selected model is image-capable. A stale
  `enabled_tools=["image_generation"]` on a text model is silently
  dropped instead of producing an invalid Gemini request.
- `_stream_gemini` validates the model id against
  `[A-Za-z0-9._-]+` before URL interpolation so a model like
  `../cachedContents/x` cannot redirect the request to an unintended
  endpoint with the configured API key attached.
- Empty-text Gemini parts that still carry `thoughtSignature` emit a
  content-free delta with `extra_content.google.thought_signature` so
  Gemini 3 turns that end with a signature-only fragment do not lose
  the replay state.
- ConnectError / ReadTimeout / generic HTTPError paths in
  `_stream_gemini` now close the synthetic web_search tool_start
  with a matching tool_end before the error chunk so the UI does not
  leave a stuck "searching..." card on transport failure.
- `providers.py` default_models drop the non-existent
  `gemini-3.5-pro` (Google launched only `gemini-3.5-flash` at
  I/O 2026; Pro tier remains `gemini-3.1-pro-preview`).
- `routes/inference.py` only forwards `payload.top_k` when the caller
  explicitly set it on the request (Pydantic `model_fields_set`).
  Omitted top_k stays omitted, restoring the pre-PR behavior where
  Gemini uses its server default.
- `ChatCompletionRequest.enable_prompt_caching` adds a `mode="before"`
  validator that coerces the canonical string literals "true"/"false"
  back to bool so historical opt-out callers keep working after the
  field widened to `Union[bool, str]` for Gemini cache resource names.

Frontend:
- `providerSupportsBuiltinWebSearch` / Code / Image now accept the
  saved connection `baseUrl` and return false for custom OAI-compat
  Gemini proxies. Backend skips `_stream_gemini` for those bases, so
  native tool envelopes never reach them; hiding the pills keeps the
  request, builder, and UI consistent.
- `provider-capabilities.ts` Gemini 3 Pro effort ladder restores
  `["low", "medium", "high"]` to match Google's documented levels.
- Call sites in `chat-page.tsx` and `chat-adapter.ts` pass through
  `provider.baseUrl` so the proxy gate fires.

Tests: covers Gemini 3 Pro medium pass-through, custom proxy dispatch
on `/v1` and `/openai` bases, path-traversal model id rejection,
top_k omission when not explicit, text-model image_generation drop,
empty-text + thoughtSignature surfacing, and
enable_prompt_caching string coercion.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: fix sixth-pass Gemini findings (PR #5720)

Round 6 review follow-ups:

Frontend:
- chat-adapter `delta.tool_calls` accumulates fragments by `id` /
  `index` instead of pushing a new tool-call card per chunk. The
  standard OpenAI Chat Completions stream contract sends `id`/`name`
  on the first chunk and partial `function.arguments` on subsequent
  chunks; our previous handler parsed each fragment as a standalone
  tool call. Local llama.cpp and OAI-compat providers that stream
  fragments now reassemble into a single function-call part.
- chat-adapter also preserves `extra_content` on streamed tool-call
  deltas so Gemini 3 `thoughtSignature` survives to the next turn.
- provider-capabilities Gemini 3 Pro restores "medium" in the
  reasoning-effort ladder (Google's official Gemini API thinking
  doc lists low/medium/high for Gemini 3.1 Pro; my earlier round 4
  coercion was wrong).
- provider-capabilities orders `gemini-2.5-flash-lite` ahead of the
  broader `gemini-2.5-flash` prefix so Flash-Lite falls into the
  "no native thinking knob" branch as documented.

* Studio: round-trip Gemini tool_calls and tool results (PR #5720)

Recurring round 3-6 P1: the chat-adapter renders Gemini function-call
parts and code-execution events but `toOpenAIMessage` only serialized
text + image content, so the next turn lost the assistant
`tool_calls[]` (including Gemini 3's required
`extra_content.google.thought_signature`) and the matching
`role="tool"` result. Gemini 3 multi-turn function calling and code
execution failed validation on the second turn.

Frontend:
- types/api.ts widens OpenAIChatMessage to permit `role="tool"`,
  `tool_calls`, `tool_call_id`, `name`, and `content: null`. Adds
  OpenAIToolCallPart with `extra_content` for the Gemini round-trip.
- chat-adapter: new `toOpenAIMessages` expands an assistant turn with
  tool-call parts into [assistant w/ tool_calls + extra_content,
  role=tool result, ...]. tool result content is JSON-serialized so
  the backend translator can rebuild Gemini's `functionResponse`
  shape.
- chat-adapter outbound history now uses `flatMap(toOpenAIMessages)`
  so each assistant tool-call round-trips through the standard OAI
  shape the backend's `_stream_gemini` already understands.

* Studio: replay Gemini code_execution and image native parts on history (PR #5720)

Multi-turn Gemini history previously lost the native executableCode,
codeExecutionResult, and inlineData parts because the outbound
translator regenerated a generic functionCall for every assistant
tool_call. Stow the native dict on tool_end (frontend) and replay it
verbatim with thoughtSignature (backend) so follow-up turns preserve
the prior execution and image generation state. Skip role="tool"
fan-out for server-side builtin tools so Gemini does not 400 on a
functionResponse with no matching user-declared function.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: complete Gemini built-in tool replay round-trip (PR #5720)

Round 7 follow-up to the multi-turn native-part work. Three asymmetric
storage/consume gaps remained between the backend translator and the
chat adapter, so realistic Gemini follow-up turns degraded to generic
functionCalls instead of native history.

- Frontend collectAssistantToolCalls now drops web_search outright,
  drops code_execution / image_generation when the native part is
  missing, and promotes args.google to extra_content.google so the
  backend native_part replay branch actually fires.
- Backend image_generation tool_end now emits google.native_part
  with the inlineData (mimeType + base64) and thoughtSignature so the
  follow-up image-edit turn can replay the prior image as a native
  Gemini model part.
- Backend code-execution plot tool_end now stows google.native_part
  with the inlineData so the merged code-exec card can round-trip
  executableCode + codeExecutionResult + inlineData on the same id.
- Added regression tests for image-gen native-part replay and the
  code-exec plot native_part stow.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 8 Gemini follow-ups (PR #5720)

- Text-part thoughtSignature: stow on the assistant message during
  streaming and replay onto the last text part on the next turn so
  Gemini 3 strict function-calling does not reject history.
- Function declarations: recursively strip Gemini-unsupported OpenAPI
  keys (additionalProperties, $schema, $defs, strict, etc.) so OpenAI
  strict tools stop 400ing as INVALID_ARGUMENT on Gemini.
- OpenAI-compat fallback: forward tools/tool_choice so custom Gemini
  proxies (LiteLLM, gateways) keep function-calling.
- enable_prompt_caching: cover the Pydantic v1 legacy off/on/f/n/t/y
  string set so explicit opt-outs stay opt-out (Gemini was sending
  cachedContent: "off" otherwise).
- Frontend collectAssistantToolCalls / collectToolResultMessages: use
  google.native_part + result presence to disambiguate provider
  builtins from same-named user-declared functions.
- Added regression tests for text-signature replay and schema
  sanitization.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 9 Gemini follow-ups (PR #5720)

Two round-9 convergent finds across the 12 reviewers:

- Server-side web_search was leaking onto the next turn as a fake
  user functionCall/functionResponse. The previous heuristic (skip
  builtin only when no native_part AND no result) let it through
  because the synthetic tool card has a non-empty result string.
  Always skip web_search by name on both serializers, accept that a
  user-declared function literally named "web_search" must use a
  different name.
- Assistant `extra_content` was dropped by ChatMessage validation
  before _stream_gemini could replay text-part thought signatures.
  Add the field to ChatMessage and forward it through
  _build_external_messages so the multi-turn signature path actually
  carries data.

Includes a regression test for the ChatMessage round-trip.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 10 Gemini follow-ups (PR #5720)

Three convergent round-10 reviewer findings closed:

- Tag synthetic provider-side builtins with `args._server_tool=True`
  via a central helper that runs in every `_emit_tool_event` /
  `_emit_synthetic_tool_event` path. The frontend filter now skips
  on that marker instead of on the public tool name, so local
  llama.cpp `web_search` and OpenAI function tools literally named
  `web_search` / `code_execution` / `image_generation` round-trip
  cleanly while Gemini grounding / hosted code-exec / hosted image
  cards stay skipped.
- Gate Gemini image-mode (responseModalities=[TEXT,IMAGE]) on the
  Images pill (enabled_tools containing `image_generation`).
  Selecting an image-capable model with the pill off no longer forces
  image output the UI says is disabled.
- Frontend missing-key guard now exempts custom Gemini OAI-compat
  proxies (LiteLLM, gateways) the same way the backend already
  does, so a saved Gemini connection on `http://localhost:4000/v1`
  with no API key stops being blocked.

Existing tests updated to pass `enabled_tools=["image_generation"]`
on image-mode capture paths.

* Studio: round 11 Gemini follow-ups (PR #5720)

Four round-11 findings closed:

- Kimi _stream_kimi_web_search's local _synthetic_chunk helper now
  runs through _stamp_server_tool_marker so Kimi search history is
  not replayed as a fake user functionCall on the next turn (was an
  asymmetric miss after the round-10 tagging work).
- OpenAI Responses path (/v1/responses for gpt-5.x) forwards
  caller-supplied tools / tool_choice, translating the Chat
  Completions function-tool shape into the Responses native shape.
  Without this, standard OpenAI tools silently dropped on
  Responses-routed traffic.
- Decoupled the Gemini image-tier model-id guards (text-tool /
  thinking strip) from the Images pill flip
  (responseModalities=[TEXT,IMAGE]). gemini-2.5-flash-image with
  Search/Code on and the Images pill OFF no longer forwards
  googleSearch + thinkingConfig (Gemini 400s on those for legacy
  image ids).
- Gemini-only extra_content is now forwarded by
  _build_external_messages only when provider_type=="gemini" so
  Google's thought_signature does not leak into OpenAI / Mistral /
  Kimi / OpenRouter request bodies as an unknown field.

Added a regression test for the image-tier strict-guard split and
extended the extra_content test to cover the non-Gemini suppression.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 12 Gemini follow-ups (PR #5720)

Three round-12 convergent findings closed:

- extra_content leak to custom Gemini OAI-compat proxies (8/12
  reviewers). _build_external_messages now gates extra_content on
  the native generativelanguage.googleapis.com host, not just
  provider_type=="gemini", so LiteLLM / custom gateways routed
  through /chat/completions do not get an unknown top-level field.
- OpenAI Responses function-tool round-trip (5/12 reviewers). I
  added user `tools` forwarding in round 11 but did not parse the
  matching response.output_item.done items of type=function_call.
  The parser now translates them into Chat Completions
  delta.tool_calls and the terminal chunk reports
  finish_reason="tool_calls" when the model invoked a user
  function.
- Image-tier model with Images pill OFF (2/12). Google's image
  models default to text+image when responseModalities is omitted,
  so the previous fix silently still billed image output. Force
  responseModalities=["TEXT"] when the Images pill is off and the
  selected model is image-capable.

Updated the two pre-existing tests that pinned the synthetic-tool
arguments shape to include the new `_server_tool: True` marker, and
added a regression test for the Responses function-call output
translation.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 13 Gemini/Responses follow-ups (PR #5720)

Three round-13 convergent findings closed:

- OpenAI Responses function_call indices: my round-12 translator
  hardcoded every emitted tool_calls[*].index to 0, so parallel
  function calls collapsed for index-keyed clients. Track and
  increment function_call_index per emit (mirrors the Gemini
  branch's distinct-index pattern). 10/12 reviewers flagged.
- _SERVER_SIDE_BUILTIN_TOOL_NAMES now includes web_fetch so
  Anthropic-hosted web_fetch cards carry the _server_tool marker
  and the frontend history serializer doesn't replay them as fake
  user functions. 4 reviewers flagged.
- OpenAI Responses follow-up tool results now serialize as
  Responses-shape function_call / function_call_output items keyed
  by call_id, instead of Chat Completions role="tool" content.
  Skips assistant tool_calls tagged with _server_tool so hosted
  builtins don't round-trip as user functions. 2 reviewers flagged.

Updated the Anthropic code_execution and web_fetch test argument
pins to include the new _server_tool marker, and added two
regression tests (distinct indices on parallel function_call,
function_call_output round-trip).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 14 Gemini follow-ups (PR #5720)

Three round-14 findings closed:

- Remote `image_url` translation (5 reviewers convergent). Public
  HTTPS image URLs can't be sent as `fileData.fileUri` -- Gemini
  reserves that path for Files API URIs and YouTube. Fetch the
  bytes server-side and inline them as base64 `inlineData`,
  mirroring the pre-PR OpenAI-compat behaviour. YouTube URLs and
  generativelanguage.googleapis.com/v1beta/files/* stay as
  `fileData`.
- Nullable JSON Schema type arrays. OpenAI strict tools commonly
  use `"type": ["string", "null"]`; the Gemini sanitizer now
  flattens that to `"type": "string", "nullable": true` so strict
  function tools stop 400ing.
- Parallel functionResponses now ride on one user content block
  with multiple `functionResponse` parts, matching Google's
  parallel tool docs. Consecutive `role="tool"` messages merge
  into the previous user turn instead of splitting into separate
  Gemini user turns.

Three regression tests added (remote URL fetch + inline, Files
API / YouTube fileData preservation, schema nullable flattening,
parallel-tool grouping).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: SSRF harden Gemini remote image fetch (PR #5720)

Round 15 convergent finding (12/12 reviewers). My round-14 fix to
download user-controlled image URLs for inlineData inlining was an
SSRF / data-exfiltration path: no scheme check, no private-host
guard, no size cap, no Content-Type validation, redirects could
bounce to internal services, and the full URL was logged.

Replace the inline fetch with `_safe_fetch_image_for_gemini`:

- Require https:// (reject http, file, data, ftp, etc).
- Resolve the hostname via socket.getaddrinfo and reject if ANY
  resolved address is private / loopback / link-local / multicast /
  reserved / unspecified (covers 127.0.0.0/8, 10/8, 172.16/12,
  192.168/16, ::1, 169.254/16 metadata, RFC 6890).
- Block IP-literal URLs that resolve into those same ranges.
- Cap response body at 10 MB (Content-Length pre-check + streamed
  byte counter).
- Require Content-Type to start with `image/`.
- Disable redirect following so a 302 to a private host can't slip
  past the address check.
- Use a short 15s timeout and a tiny connection pool dedicated to
  these fetches.
- Log only the host name + error class -- no full URL, no signed
  querystring leak.

If the guard rejects, the image part is silently dropped (instead
of forwarding raw bytes or a fileData fallback). Files API URIs
and YouTube URLs still ride as `fileData.fileUri` unchanged.

Tests: replaced the live-fetch test with a `_safe_fetch_image_for_gemini`
monkeypatch, added four new SSRF-guard tests (non-https rejected,
loopback / private IP literals rejected, hostnames that resolve to
private IPs rejected).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 16 Gemini follow-ups (PR #5720)

- IP-pinned image fetch (`_safe_fetch_image_for_gemini`): reuse the
  validated-once-then-pin pattern from `tools._fetch_page_text` via
  `asyncio.to_thread`, so DNS rebinding between validation and the
  HTTP connect cannot redirect us at a private/metadata address.
  Catch malformed-bracketed IPv6 urlparse errors. Follow up to 4
  redirect hops with per-hop SSRF re-validation.
- Replace contains-substring detection of Gemini Files API + YouTube
  URLs with parsed scheme/host/path checks, so attacker URLs like
  `https://evil.example/path/youtube.com/x.png` no longer skip the
  safe-fetch path and serialize as `fileData.fileUri`.
- `_build_external_messages`: strip per-tool-call `extra_content`
  for non-native-Gemini providers; the Gemini-only
  `thought_signature` payload was leaking through `tool_calls[]`
  into /chat/completions on OpenAI, Anthropic, and custom Gemini
  OAI-compat gateways.
- `_server_tool` marker now gated on the function name being one of
  the canonical builtin names (`web_search`, `web_fetch`,
  `code_execution`, `image_generation`) AND the marker being set,
  so a user function whose schema happens to define an
  `_server_tool` field is no longer dropped. Frontend filter mirrors
  the same gate, plus a backward-compat fallback for pre-PR
  persisted server-tool cards (no marker) routed via name +
  native_part / web-tool heuristic.
- Gemini schema sanitizer collapses `anyOf: [{X}, {"type":"null"}]`
  to `{X, "nullable": true}` so Optional[X] tool args from
  OpenAI/Pydantic schemas no longer 400 the Gemini request.
- Frontend tool-result serializer emits `{"result":""}` for empty
  string outputs so the ChatMessage validator does not reject
  `role="tool"` with empty content.
- Coerce `medium` thinkingLevel to `high` for legacy
  `gemini-3-pro*` / `gemini-3-pro-preview*` (only low/high
  documented; shut down 2026-03-09); 3.1+ Pro still passes through.
- Hide Gemini native thinking ladder on custom OAI-compat Gemini
  gateways by routing `getExternalReasoningCapabilities` through
  `isGeminiCustomOpenAICompatBase(baseUrl)`; thread baseUrl through
  all four call sites.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 17 Gemini follow-ups (PR #5720)

- Frontend `collectAssistantToolCalls` and `collectToolResultMessages`
  no longer drop unmarked `web_search` / `web_fetch` cards by name
  alone: a user-defined function with one of those names must
  round-trip. Pre-PR persisted `code_execution` / `image_generation`
  cards still get filtered via a shape heuristic (kind/command/code/
  prompt fields) instead of bare name.
- `_build_external_messages._filter_tool_calls` now drops marked
  server-side builtin `tool_calls` entirely for non-native-Gemini
  providers, not just their `extra_content`. An assistant turn whose
  only payload was a marked builtin is dropped completely so the
  receiving provider does not see an orphan tool_call.
- `_stream_anthropic` translates OpenAI top-level `tool_calls` into
  Anthropic native `{type:"tool_use", id, name, input}` content
  blocks, and translates `role="tool"` follow-ups into `role:"user"`
  messages carrying a `tool_result` block. Anthropic's native
  Messages API rejects the OpenAI shapes.
- `_safe_fetch_image_for_gemini_sync` factors URL validation through
  `_safe_parse_https`, so malformed `port` access (e.g.
  `https://host:bad/x.png`) and malformed redirect targets (e.g. a
  302 to `https://[bad/x.png`) drop the image instead of raising mid-
  request.
- `tool_choice="none"` now disables hosted builtins (Gemini
  googleSearch / codeExecution and OpenAI Responses web_search /
  shell / image_generation), not just user function declarations.
- Schema sanitizer handles multi-type `anyOf` with null
  (`Union[str, int, None]`): keep the slim non-null anyOf and add
  `nullable: true` so Gemini does not reject `{"type":"null"}`.
- Image fetch falls back to the caller-provided MIME (guessed from
  URL extension) when the server omits Content-Type instead of
  dropping the image as `non-image content-type=<none>`.
- Per-request aggregate caps on remote image inlining (8 images,
  20MB total) so a single chat request cannot force unbounded
  backend downloads.
- Frontend exposes the reasoning ladder for `gemini-2.5-flash-lite`
  (`none/minimal/low/medium/high/max`) so the UI can drive the
  thinkingBudget the backend already supports.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 18 Gemini follow-ups (PR #5720)

- `tool_choice="none"` now opts out of hosted builtin tools on every
  provider path, not just Gemini and OpenAI Responses. Anthropic
  web_search / web_fetch / code_execution, Kimi `$web_search` early
  return, and OpenRouter `plugins:[{id:"web"}]` are all gated on
  `tool_choice_disabled`. Passing `enabled_tools=[...]` with
  `tool_choice="none"` no longer triggers provider-side search /
  code execution for any provider.
- `_stream_anthropic` accepts `tool_choice` and threads it through;
  the dispatcher in `stream_chat_completion` forwards it.
- Frontend `isServerSideBuiltinToolPart` simplified to drop only on
  (marker) OR (canonical name + native_part). The previous shape
  heuristic on `args.kind`/`args.command`/`args.code`/`args.prompt`
  dropped real user-declared `code_execution` / `image_generation`
  functions. Pre-PR persisted hosted cards lacking the marker now
  leak to non-native providers on switch -- preferred to silently
  deleting legitimate function-call history.
- Backend `_is_marked_server_builtin_tool_call` and the OpenAI
  Responses translator's matching filter accept BOTH `_server_tool`
  marker AND `args.google.native_part` as durable provider-side
  signals so Gemini code_execution / image_generation cards are
  still dropped on a provider switch.
- Per-request remote image count cap now counts ATTEMPTS, not just
  successful inlines, so 100 failing/slow URLs cannot each consume
  the 15s fetch timeout. Data: URL images now share the same count
  and byte caps as fetched remote URLs.
- OpenAI Responses translator tracks skipped server-builtin
  `function_call` ids and drops their matching `role="tool"`
  follow-ups, preventing orphan `function_call_output` items in the
  outbound body.
- Gemini schema sanitizer preserves multi-type unions with null:
  `{"type":["string","integer","null"]}` becomes
  `anyOf:[{string},{integer}] + nullable:true` instead of being
  flattened to the first non-null type.
- Gemini model id validation moved to the top of `_stream_gemini`
  so an invalid model id rejects the request before any remote
  image fetch / message translation side effect.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 19 Gemini follow-ups (PR #5720)

- `_build_external_messages` now skips an empty assistant turn when
  `_filter_tool_calls` drops every synthetic builtin tool_call (was
  guarded only on the `content is None` branch; the string-content
  and list-content branches still forwarded
  `{"role":"assistant","content":""}` which several providers
  reject). Also tracks the dropped server-builtin tool_call ids and
  skips the matching `role="tool"` follow-ups so the receiving
  provider does not see an orphan tool_result.
- OpenRouter `web_search_active` (the synthetic tool_start /
  tool_end emitter) is now also gated on `tool_choice_disabled` so
  a request with `tool_choice="none"` does not surface a fake
  web_search card in the chat UI even though the plugin was
  correctly stripped from the outbound body.
- `_stream_anthropic` translates an OpenAI role="tool" with list
  content (`content=[{"type":"text","text":"..."}]`) into a native
  `tool_result` block on a user message; previously only the
  string-content shape was translated, so list-content tool results
  were forwarded as invalid `role:"tool"` messages.
- Gemini `data:` URL image_url parts now require an `image/*` MIME
  type; a `data:text/html;base64,...` is dropped instead of being
  forwarded as `inlineData.mimeType="text/html"` (Gemini rejects
  the malformed image part). Symmetric with the fetched-remote
  image fetch path that already rejects non-image Content-Type.
- YouTube `fileData.fileUri` now declares `video/mp4` as the
  mimeType instead of `image/jpeg` guessed from the URL path. The
  YouTube/fileData input is the documented Gemini video path; the
  guessed image MIME made valid YouTube inputs malformed.
- OpenAI Responses translator preserves `response.output` ordering
  on assistant turns that emitted both text and a function_call:
  assistant text is now serialized BEFORE the function_call item
  so the subsequent function_call_output (the matching role=tool
  follow-up) lands in the right position. Previously the order
  was function_call -> assistant text -> function_call_output,
  which can confuse multi-turn function-calling flows.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 20 Gemini follow-ups (PR #5720)

Convergent reviewer findings from round 20:

- tool_choice="none" no longer flips responseModalities=[TEXT,IMAGE]
  on image-tier Gemini models. Forced-function tool_choice (e.g.
  {type:function, function:{name:lookup}}) also drops hosted Search /
  code execution from the Gemini body so the caller's pinned user
  function is not silently joined by hosted builtins.

- Gemini code-execution thoughtSignature replay now uses an ordered
  parts list (native_part.parts[]) so per-part signatures stay
  attached to the exact part Gemini emitted. The previous merged
  shape fanned one top-level thoughtSignature across executableCode
  + codeExecutionResult + inlineData and tripped Gemini 3 strict
  validators. Backward-compat fallback keeps pre-round-21 persisted
  history working: a legacy native_part with a single subpart still
  replays the signature on that subpart; merged legacy objects pin
  the signature to executableCode only.

- Remote-image fetch threads the remaining per-request byte budget
  into _safe_fetch_image_for_gemini, so over-budget URLs are
  refused via Content-Length pre-check / short read instead of
  fully downloaded then discarded after the aggregate cap check.

- Gemini role=tool with OpenAI list-form content
  ([{type:text,text:result}]) now flattens text parts before
  building functionResponse.response.result; previously the parts
  arrived as the result value instead of the actual tool output.

- Frontend chat-adapter merges native_part by concatenating parts
  lists (preserving per-part thoughtSignature). Wire types expose
  enable_prompt_caching as boolean|string (Gemini cached-content
  name) and OpenAIChatDelta now carries tool_calls and extra_content.

- Test test_openrouter_no_synthetic_web_search_event_on_tool_choice_none
  reads _toolEvent from the top-level SSE payload so a backend
  regression cannot mask the assertion.

Adds 7 regression tests covering image_generation gate, forced-function
gate, native_part list replay, legacy fallback, list-content
functionResponse flattening, fetch byte-budget threading, and wire
types.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Apply forced-function tool_choice gate to Anthropic, OpenRouter, Kimi

Previously only the Gemini path treated `tool_choice={"type":"function",
"function":{"name":...}}` as a hosted-tool opt-out. Anthropic,
OpenRouter, and Kimi still attached hosted web_search / web_fetch /
code_execution when the caller explicitly pinned a user function plus
`enabled_tools=[...]`. That contradicts the explicit function pin and
bills the caller for unwanted server-side calls.

Mirror the Gemini gate symmetrically:
  - Anthropic web_search / web_fetch / code_execution
  - OpenRouter `plugins:[{id:"web"}]` + the synthetic web_search SSE
    event the same path emits at stream close
  - Kimi `_stream_kimi_web_search` dispatch

Adds 4 regression tests:
  - test_anthropic_forced_function_tool_choice_drops_hosted_tools
  - test_openrouter_forced_function_tool_choice_drops_web_plugin
  - test_kimi_forced_function_tool_choice_skips_web_search_helper
  - test_openrouter_no_synthetic_web_search_event_on_forced_function_tool_choice

All 146 existing backend tests still pass.

* Strip Gemini-only synthetic tool history on local-GGUF dispatch

After a Gemini chat that ran code_execution / image_generation, switching
the same thread to a local GGUF model used to forward the synthetic
provider-side tool_calls (tagged with `args._server_tool` or carrying a
Gemini `args.google.native_part` payload) and the message-level
`extra_content` to llama-server. The receiving backend has no tool
declaration for those names and no use for Gemini thoughtSignature
metadata; in the worst case it can produce an orphan tool_call_id and a
confused continuation.

Add `_strip_provider_synthetic_tool_history()` and wire it through the
two local message builders:
  - `_openai_messages_for_passthrough`  (OAI-compat passthrough)
  - `_openai_messages_for_gguf_chat`    (standard GGUF chat path)

Real user-function `tool_calls` and their matching `role="tool"` replies
survive unchanged; only synthetic provider-side cards and Gemini-only
`extra_content` are stripped. If the synthetic call was the assistant
turn's only payload, the now-empty turn is dropped too so llama-server
does not reject the request.

Adds 2 regression tests:
  - test_strip_provider_synthetic_tool_history_drops_synthetic_only
  - test_strip_provider_synthetic_tool_history_drops_empty_assistant

142 existing backend tests still pass.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Disable Search/Code composer pills for Gemini image-tier models

For external Gemini image-tier models (gemini-2.5-flash-image,
gemini-3.x-image-preview, etc.), the backend unconditionally strips
code_execution and strips web_search on older image ids. Search is
still allowed on Gemini 3.x Pro/Flash image models, which
supportsBuiltinWebSearch already encodes per model.

Before this commit the composer pill gates were:
  searchDisabled = !modelLoaded || !(supportsTools || supportsBuiltinWebSearch)
  codeDisabled   = !modelLoaded || !(supportsTools || supportsBuiltinCodeExecution) || imageModeDisablesCode

`supportsTools` here is a local-runtime fallback that becomes true when
any tool-capable local model has been loaded in the session. With a
local tool-capable runtime active, switching the chat to an external
Gemini image-tier model used to leave Search/Code clickable, even
though the backend will silently drop the tool on the wire.

Detect "external provider is Gemini AND the model is image-tier" (via
supportsBuiltinImageGeneration) and gate the two pills strictly on the
provider's own builtin support in that case. Non-Gemini paths and
non-image Gemini models keep the supportsTools fallback unchanged.

* Apply forced-function tool_choice gate to OpenAI Responses path

Round 22 added the gate for Gemini / Anthropic / OpenRouter / Kimi but
missed the OpenAI Responses translator. When a caller pinned a user
function via `tool_choice={"type":"function","function":{"name":...}}`
plus `enabled_tools=["web_search","code_execution","image_generation"]`,
the Responses body still attached `{"type":"web_search"}`,
`{"type":"shell"}`, and `{"type":"image_generation"}` server tools. The
function pin should suppress those for the same privacy + billing reason
the other provider paths now do.

Compute `_responses_tool_choice_forced_function` next to
`_responses_tool_choice_none` and gate each hosted-tool append on
`_responses_hosted_builtins_allowed = not none and not forced_function`.
The fix has to be applied in TWO places: the initial body builder and
`_build_body()` (called by the container-expiry retry path). User
function declarations still flow through so the pin has something to
target, and the Responses-shape `{type:"function", name:"..."}`
`tool_choice` is forwarded unchanged.

Adds regression test `test_openai_responses_forced_function_tool_choice_drops_hosted_tools`.
All 166 existing backend tests across Gemini + Responses + image-gen +
code-exec suites still pass.

* Round 24 P1s: SSRF shared-address gap + extra_content text-only leak + custom-Gemini model list

Three convergent P1s from round 24 review:

1. SSRF: the shared SSRF validator in `tools._validate_and_resolve_host`
   used a denylist (is_private / loopback / link_local / multicast /
   reserved / unspecified). Python classifies shared address space
   (100.64.0.0/10 carrier-grade NAT, plus 240.0.0.0/4, benchmarking
   ranges, etc.) with `is_private=False` AND `is_global=False`. The new
   Gemini server-side image fetcher therefore accepts URLs whose
   hostname resolves to 100.64.0.1 in cloud/VPC deployments. Add
   `not ip.is_global` as the primary gate -- a single source of truth
   that covers every current and future non-global range.

2. _strip_provider_synthetic_tool_history previously only stripped
   message-level `extra_content` when the assistant turn had tool_calls.
   A plain text Gemini reply carrying
   `extra_content.google.thought_signature` flowed through to
   llama-server when the thread was switched to a local GGUF backend.
   Always strip message-level `extra_content` on assistant turns.

3. routes/providers.list_provider_models applied Gemini's native
   `model_id_allowlist` regex to every Gemini provider, including
   custom OAI-compatible bases (LiteLLM, deployment gateways). IDs like
   `google/gemini-2.5-flash` and team-prefixed deployment aliases got
   filtered out even though the chat-dispatch path now routes them via
   the OpenAI-compatible client. Skip registry-level model-id filters
   when the configured Gemini base_url host is not the canonical
   `generativelanguage.googleapis.com`, mirroring the chat-dispatch
   gate.

Three regression tests added:
  - test_validate_and_resolve_host_blocks_shared_address_space
  - test_strip_provider_synthetic_tool_history_drops_text_only_extra_content
  - test_gemini_custom_oai_compat_base_skips_native_allowlist

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Round 25 P1s: skip synthetic server-tool replay + inline $ref/$defs into Gemini schema

Two convergent reviewer findings on the native Gemini path:

1. _stream_gemini's tool_calls replay loop falls through to a generic
   functionCall emission whenever it sees an assistant tool_call. Marked
   server-side builtin cards (web_search / web_fetch tagged with
   _server_tool or args.google.native_part) hit that fallthrough with no
   replayable native_part, which produces an outbound functionCall whose
   name is not a declared user function. The Gemini turn 400s on the
   undeclared name. Guard the loop to drop those entries instead, while
   keeping the existing code_execution / image_generation native-part
   replay branch intact.

2. _sanitize_gemini_schema uses a strict allowlist that drops local
   $ref / $defs references. Pydantic-generated tool schemas hoist nested
   object shapes into $defs and reference them via {"$ref": "#/$defs/X"},
   so a property like address: {"$ref": "#/$defs/Address"} collapsed to
   {} on the wire and the model lost the nested fields, types, and
   required keys. Resolve local #/... pointers against the schema root
   and inline the referenced subtree, with local siblings overriding
   the reference (normal JSON Schema composition) and a seen-ref guard
   for self-referential schemas.

Added regression coverage:
- test_gemini_native_skips_synthetic_server_builtin_replay
- test_function_declarations_inline_local_refs_into_gemini_schema
- test_function_declarations_inline_local_refs_in_anyof_and_items
- test_function_declarations_self_referential_schema_terminates

All 145 Gemini provider tests pass; touched provider regression set
(OpenAI Responses, code execution, image generation, Anthropic code
execution, Anthropic web_fetch) also 43/43 green.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Round 26 P1s: drop orphan Gemini functionResponse + Anthropic /messages synthetic-history strip

Reviewer round 26 surfaced two convergent asymmetric-fix bugs.

1. _stream_gemini drops a synthetic server-tool tool_call (web_search /
   web_fetch tagged _server_tool) and also replays code_execution /
   image_generation tool_calls as Gemini-native executableCode /
   codeExecutionResult / inlineData parts. The matching role="tool"
   follow-up was still falling through to the generic functionResponse
   branch, producing either an orphan functionResponse (synthetic case)
   or a duplicate response pointing at a name with no
   functionDeclarations entry (native-part case). Both forms 400 the
   next Gemini turn. Track skipped + native-replayed tool_call_ids in
   _gemini_skip_tool_result_ids and short-circuit the role="tool"
   branch on a match.

2. The Anthropic-compatible local /v1/messages route only called
   _drop_empty_assistant_sentinels on the OpenAI-translated history,
   while the sibling /v1/chat/completions and GGUF passthrough builders
   chain that with _strip_provider_synthetic_tool_history. An Anthropic
   caller replaying a prior provider-side tool_use therefore forwarded
   fake builtin tool history straight into local llama-server. Apply
   the same strip on the Anthropic route after the
   anthropic_messages_to_openai conversion.

Regression coverage added:
- test_gemini_native_skips_orphan_function_response_for_dropped_builtin
- test_gemini_native_skips_orphan_function_response_for_native_part_replay

Gemini suite 147/147; touched provider regression set 43/43.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Round 27 P1s: native_part location fallback + Gemini image request budget for base64

Two convergent reviewer findings on the native Gemini path.

1. _stream_gemini's synthetic-builtin detector at lines 3519-3524
   recognizes args.google.native_part as a server-tool marker, but
   _native_part was only loaded from tc.extra_content.google.native_part.
   A direct OpenAI-compatible API caller or imported third-party thread
   round-trips the payload through function.arguments because
   tool_calls[].extra_content is not in the OpenAI spec. The round-25
   guard then saw a synthetic builtin with no _native_part and dropped
   the entire assistant turn, so the next native Gemini request lost
   the prior executableCode / inlineData / codeExecutionResult context.
   Fall back to args.google.native_part when extra_content path is
   missing, mirroring what the synthetic detector already accepts.

2. _GEMINI_REMOTE_IMAGE_MAX_TOTAL_BYTES capped DECODED bytes at 20MB.
   Gemini receives images base64-encoded inside JSON, and base64
   inflates payload size by ~4/3. With 20MB decoded the actual JSON
   body is ~26.7MB plus prompt overhead, well over Gemini's ~20MB
   request limit. Drop the decoded cap to 14MB so realistic multi-
   image turns stay safely under 20MB encoded.

Added regression test test_gemini_native_part_falls_back_to_args_google
covering an OpenAI-compat-shaped image_generation tool_call whose
native_part lives only in function.arguments.

Gemini suite 148/148.

* Fix TS build errors from main merge: restore imageParts + refusal return [] + cast image-edit ref

Three errors in chat-adapter.ts surfaced by the frontend tsc step after merging
main into feat/gemini-provider:

1. The Anthropic refusal early-return used main's  but
   toOpenAIMessages returns SerializedMessage[]; flip to .
2. Restore  -- the line
   was lost when removing main's conflict block from the function body.
3. selectedImageEditReference splice was inserting OpenAIChatMessage
   into a SerializedMessage[] array; the shapes differ on tool_calls.id
   nullability. Cast the reference message through unknown -- it carries
   no tool_calls, so the runtime payload is structurally compatible.

Reproduced locally with `tsc -b --pretty false` (now passes). Build
also failing in the in-repo `npm run build` step on PR CI; this commit
unblocks all 12 failing UI/API workflows.

* Tighten verbose comments in external_provider.py + chat-adapter.ts

Compress multi-line explanatory comments in the Gemini translator
and the chat adapter without changing any behaviour. All 148 Gemini
provider tests still pass; tsc --noEmit clean.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@users.noreply.github.com>
2026-05-27 06:01:24 -07:00
Dariton4000
dac2aeda1a
Studio: expose image size setting in training UI (#5743)
* Studio: add VLM image-size control for training

  Studio vision fine-tuning had no explicit way to cap image resolution, so
  users could not trade visual detail against context and memory use from the
  training UI, YAML config, or API payload. :) Add a nullable `vision_image_size`
  setting that keeps the current model default when unset and applies a
  max-side resize when provided.

  - Add `vision_image_size` to the training request model, route payload, backend
    training config, and frontend API/types plumbing.
  - Validate the value server-side as either null or an integer in the supported
    256-2048 range.
  - Surface an Image Size selector for vision LoRA training with Default plus
    common preset sizes.
  - Include the value in training start payloads only for image-dataset vision
    models, and serialize it into vision-aware YAML configs.
  - Map backend model defaults back into the training store and reset the value
    when reapplying model defaults.
  - Pass the resize through the Torch trainer via `UnslothVisionDataCollator`
    using max-dimension semantics.
  - Apply the same max-dimension resize in the MLX VLM path before mlx-vlm's
    internal collation, preserving aspect ratio and avoiding upscaling.
  - Add backend validation coverage and MLX resize-size tests for the new
    behavior.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: thread vision_image_size into DeepSeek OCR + writable MLX ndarray

- trainer.py: DeepSeek OCR collator now honors the new vision_image_size
  setting as image_size. Falls back to 640 when null. base_size stays at
  1024 and crop_mode stays True so the Gundam preset's dynamic cropping
  of large documents keeps working.
- worker.py: _resize_mlx_vlm_image returns np.array(image, copy=True)
  instead of np.asarray(image). The PIL view from np.asarray is not
  writable, which makes HF VLM processors emit "The given NumPy array
  is not writable, and PyTorch does not support non-writable tensors..."
  when they call torch.from_numpy. copy=True keeps the same shape and
  dtype but produces a writable buffer.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: align YAML export gate with API mapper + extend Image Size dropdown

- training-section.tsx: handleSaveConfig now passes
  isVisionModel && isDatasetImage === true to serializeConfigToYaml,
  matching buildTrainingStartPayload. Stops vision_image_size from
  leaking into exported YAML for text-only datasets where the API
  would have sent null.
- params-section.tsx: add 256 to visionImageSizePresets so the
  dropdown spans the validator's full [256, 2048] range. Also render
  a synthetic SelectItem for the current value when it was loaded
  from YAML or model defaults and is not in the preset list, so the
  controlled Select always shows the active size.

* Studio: validate vision_image_size in YAML/model-default loader

mapBackendModelConfigToTrainingPatch now mirrors the backend validator
at studio/backend/models/training.py:169 by dropping any value that is
not an integer in [256, 2048]. Pre-fix, an imported YAML like
vision_image_size: 4096 or 640.5 would land in the store and the UI
would happily display it, only to fail when Start Training posted to
the backend. With this guard the store never holds a value the backend
would reject.

* Studio: precise error messages for invalid vision_image_size inputs

Switch the field_validator to mode="before" so True/False surface as
bool (not Pydantic's coerced 1/0) and give a precise
"must be an integer or null" message instead of the misleading
"must be in [256, 2048] (got 1)". Also explicitly accepts numpy
Integral and integral Real scalars so YAML or programmatic callers
using numpy ints keep working.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: test that bool inputs yield the precise 'integer or null' error

Regression guard for the validator switch to mode="before". Pre-fix,
vision_image_size: True was rejected with "must be in [256, 2048]
(got 1)" because Pydantic coerced before our check ran. New test
asserts the message now reads "integer or null".

* Studio: tighten vision_image_size loader + YAML save + MLX rounding

Round 2 of follow-up review surfaced three usability issues:

- model-defaults.ts: switching to a model whose backend YAML omits
  vision_image_size now explicitly resets the store value to null.
  Pre-fix, a stale 2048 from a previous model would silently apply
  to the new run because every checked-in model-default file omits
  the key.
- training-section.tsx: handleSaveConfig now includes vision fields
  unless isDatasetImage is definitively false. isDatasetImage is null
  during dataset checks, after dataset edits, and on import; treating
  unknown as "drop" would silently lose the user's selection in those
  windows. Confirmed-text-only datasets still drop the value.
- worker.py: _mlx_vlm_max_resized_size now mirrors the Torch collator's
  integer formula (w * size + size_func // 2) // size_func instead of
  Python round(), which uses banker's rounding and disagreed by 1px on
  half-pixel inputs like 333x1000 with target 500 (was 166, now 167).
  Test_mlx_training_worker_config gains parity assertions.

* Studio: reset vision_image_size in the model-config error fallback path

mapBackendModelConfigToTrainingPatch resets stale image size on the
success path, but if the /api/models/config endpoint throws,
training-config-store.ts falls through to checkVisionModel and only
updates capability flags. Pre-fix that left a stale 2048 (or any
prior selection) in the store, so once dataset detection marked the
new dataset as image, the next training start would silently apply
the previous model's size. The error branch now also resets to the
DEFAULT_HYPERPARAMS.visionImageSize sentinel.

* Studio: revert DeepSeek OCR Image Size knob + move missing-key reset

Round 3 of the parallel-reviewer pass surfaced two issues that I had
introduced earlier in this PR's follow-ups.

- trainer.py: my prior change threaded vision_image_size into the
  DeepSeek OCR collator's image_size argument. The collator's
  (image_size, base_size, crop_mode) is a single preset
  (Tiny / Small / Base / Large / Gundam); changing image_size in
  isolation desynchronizes the per-crop pixel grid from num_queries
  downstream and produces wrong token grids on documents larger than
  the per-crop tile. The fix pins the collator back at the Gundam
  preset and logs a clear "ignored for DeepSeek OCR" notice when the
  user has selected a non-default Image Size.
- model-defaults.ts + training-config-store.ts: the round 4 fix that
  reset visionImageSize when a model YAML omitted the key also fired
  on same-model reloads (ensureModelDefaultsLoaded re-fires on page
  refresh), wiping a value the user had just selected. The reset is
  now in setSelectedModel, gated on selectedModel != previousModel,
  so true model switches still clear stale values while reloads keep
  the user's selection.

* Studio: extend DeepSeek OCR Image Size exclusion to MLX + frontend

Round 4 of the parallel-reviewer pass flagged that the Torch trainer
exclusion I added did not have a matching MLX guard, and that the UI
still offered the dropdown for DeepSeek OCR even though the backend
ignores it.

- worker.py: _run_mlx_training now mirrors the Torch exclusion. When
  the model name matches DeepSeek OCR, vision_image_size is forced
  back to None before _adapt_for_mlx_vlm sees it, so dataset images
  pass through unchanged just like the Torch path. Emits a clear
  status line when this happens.
- params-section.tsx: the Image Size Row is now gated on
  showVisionImageSize (showVisionLora && !isDeepseekOcr) instead of
  showVisionLora alone, so DeepSeek OCR users no longer see a control
  that silently has no effect.
- mappers.ts: buildTrainingStartPayload sends null for vision_image_size
  whenever the selected model is DeepSeek OCR, so the backend log line
  about ignoring the value never fires from a UI-driven start.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: tighten YAML import/save for vision_image_size

Two YAML-path asymmetries that could leak a stale image size into
training:

- parseYamlConfig now treats a missing training.vision_image_size as
  null. Without this, importing a YAML saved before this feature (or
  any config that omits the key) preserved whatever value the user had
  previously set on a different model. The model-defaults reload path
  still uses Object.hasOwn so same-model defaults reloads do not wipe
  a manual selection; only file import normalises the missing key.

- handleSaveConfig now passes a DeepSeek-OCR-specific guard to
  serializeConfigToYaml so saved YAML matches what the API mapper
  actually sends. Previously a state with visionImageSize set could
  emit the key even though Studio ignored it at training time for
  DeepSeek OCR, and a later import for a non-DeepSeek vision model
  would activate the stale value.

serializeConfigToYaml gains an optional third parameter
includeVisionImageSize defaulting to includeVisionFields, preserving
the existing 2-arg call signature for backwards compatibility.

* Studio: also reset vision_image_size when YAML lacks a training section

Round 9's parseYamlConfig normalization only fired when the YAML had a
training mapping that omitted vision_image_size. A lora-only or
logging-only YAML (or one with `training: null`) still left trainingObj
unset, the mapper saw no vision_image_size key, and the previously
selected store value persisted into the next training run.

Now an absent or null training section is synthesised as
{ vision_image_size: null } so model-defaults.ts always patches
visionImageSize back to Default on file import. Same-model defaults
reloads still preserve manual choices via the existing Object.hasOwn
gate in mapBackendModelConfigToTrainingPatch.

* Studio: unify parseYamlConfig non-object training handling

A fresh static review (Opus subagent) flagged P3-1: parseYamlConfig
only synthesised vision_image_size: null when raw.training was either
absent or a plain object missing the key. If raw.training is a scalar
or an array (malformed but still parseable), the value was passed
through unchanged, the mapper's Object.hasOwn returned false, and any
previously selected visionImageSize persisted - the same stale-state
leak the lora-only fallback was added to close.

Treat any non-plain-object raw.training (null, array, scalar) as a
malformed/missing section and reset to { vision_image_size: null }.

* Studio: tighten code comments for vision_image_size path

* Studio: tighten vision_image_size validator + restore lost comment context

Two issues surfaced by a fresh adversarial review of the validator:

1. v.strip().lstrip("+-").isdigit() let "++512" / "--256" / "+-+512"
   slip past the gate, then int("++512") raised an uncaught ValueError
   and Pydantic surfaced "invalid literal for int() with base 10: '++512'"
   instead of the contracted "vision_image_size must be an integer or null".

2. str.isdigit() returns True for Unicode digit families (full-width '512',
   Arabic-Indic '٥١٢', Devanagari '१०२४'), and int() coerces them, so the
   value reaching the backend wasn't the ASCII the user typed.

Replaced the lstrip+isdigit pair with re.fullmatch(r'[+-]?[0-9]+', stripped),
which rejects both shapes with the precise error and accepts the documented
ones ('256', '+512', ' 1024 '). Added 8 regression test cases covering
multi-sign strings, lone sign, and the three Unicode digit families.

Also restored comment context lost in f9c39331:
- model-defaults.ts: name studio/backend/models/training.py:_check_vision_image_size
  as the spec the [256, 2048] range mirrors, so a maintainer changing the
  cap in one file can find the other.
- training-section.tsx: enumerate the three windows in which isDatasetImage
  is null (before a check, after dataset edits, on import) so a future
  maintainer doesn't simplify the gate to `isCheckingDataset`.
- worker.py: qualify the writable-ndarray comment with "when a resize is
  requested" so it doesn't misadvertise the resize=None early-return.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2026-05-27 05:01:24 -07:00
pre-commit-ci[bot]
dd0aeec388 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-27 06:58:56 +00:00
Daniel Han
e2b7f5958b Studio: round 9 -- three P2 fixes from latest Codex bot review
1. Codex SSE wrapper terminates on exact `data: [DONE]` only.

   The old substring check `if "[DONE]" in line` would flip
   sent_done True when a normal model response carried the literal
   text "[DONE]" in delta.content (for example an explanation of
   the OpenAI stream sentinel). The real terminator was then
   suppressed, leaving OpenAI-compatible clients that finalise on
   the explicit sentinel hung on stream close. Now compares the
   stripped line to the exact `data: [DONE]` form.

2. Legacy `thread.run_streaming` path no longer returns an empty
   reply on completion-only streams.

   If the SDK exposes `thread.run_streaming` but the stream emits
   ONLY item.completed / agentMessage events with no message
   deltas, the loop previously exited with emitted_any False and
   never reached the agent-message fallback. The request returned
   200 with an empty assistant reply even though Codex produced a
   final answer. Mirror the canonical-path behavior: collect
   `_completed_agent_message_text` strings in a sidecar list and
   emit the last one when no deltas arrived. Match the canonical
   payload-extraction (`getattr(event, "payload", event)`) so the
   event-vs-payload SDK shape difference is handled the same way
   in both branches.

3. Parallel-calls fan-out propagates CodexUnavailableError so the
   route layer can return 503.

   When the SDK is not importable or the safety enums are missing
   without the dev opt-in, every worker raised the same
   CodexUnavailableError. The previous catch-all converted the
   error into a per-tab codex_tab_error event, the outer stream
   never raised, and clients saw a 200 with only tool events and
   an empty synthesis -- OpenAI-compatible consumers that ignore
   _toolEvent saw a successful empty reply. Now CodexUnavailableError
   re-raises out of the worker (no spurious per-tab error event),
   _await_workers re-raises it when EVERY worker hit the same
   setup failure, and the finally-block drain await propagates the
   exception out of the parallel function so the route's existing
   CodexUnavailableError handler can emit the right 503 SSE error
   frame. Per-tab runtime failures (model rejected, timeout, mid-
   stream SDK crash) still get swallowed into codex_tab_error
   events so a single bad model in the fan-out does not kill the
   others.

Test counts: 63/63 passing (60 round 6-8 plus 3 new round 9 regression
tests). Each new test was first run against a `git stash`-restored
pre-fix tree to confirm it catches the bug, then run against the
patched tree.
2026-05-27 06:58:06 +00:00
Daniel Han
f29faefb98 Merge branch 'main' into feat/codex-provider
Resolve three conflicts touched by main since the branch forked:

- studio/backend/core/inference/external_provider.py: take main's
  rewrite of _anthropic_citation_key (extended dedup keys covering
  end-char/page/block indices plus search_result_index) and the new
  _anthropic_supports_fast_mode helper. The branch had only the
  earlier shorter citation_key form so accepting main wholesale
  here loses nothing Codex-related.

- studio/backend/models/inference.py: keep BOTH the Codex
  parallel_calls field + _clamp_parallel_calls validator (from the
  branch) AND the new Anthropic fast_mode field (from main). They
  occupy different provider lanes.

- studio/frontend/src/features/chat/api/chat-adapter.ts: fold the
  Codex per-tab rendering pass (renderFullContent) into main's new
  orderAssistantContent positioning so tools land before text,
  generated images after, AND any Codex tab text accumulated in
  earlier _toolEvent frames is preserved through the synthesis
  content delta (round 8 render-order fix).

Backend codex tests: 60/60 passing. Anthropic citations / fast_mode
tests: 96/96 passing. Frontend builds cleanly.
2026-05-27 06:52:19 +00:00
Daniel Han
649b9f7808
Studio: expose --parallel / -np flag on unsloth studio run (#5737)
* Studio: expose --parallel / -np on `unsloth studio run`

The CLI was hardcoding `llama_parallel_slots=4` in `run_kwargs` at
`unsloth_cli/commands/studio.py`, leaving users unable to tune the
concurrent decode slot count even though the engine, KV-cache math,
and `studio.backend.run.run_server(llama_parallel_slots=...)`
plumbing all already accepted any N. This change adds a `--parallel`
/ `--n-parallel` / `-np` typer option (default 4 -- matches the
previous hardcoded value), forwards it into `run_kwargs`, and pins
the new surface with 4 unit tests.

Per-request state in `routes/inference.py` is already isolated
(`cancel_event` and `prev_text` are per-request locals in every
streaming handler; the `_lock` / `_serial_load_lock` only wrap
load/unload, not chat completions), so no concurrency refactor is
needed alongside this -- the engine layer already handles N
concurrent requests on one loaded model when llama-server is told
to.

Range guards: 1 <= N <= 64. With higher N each slot gets ctx/N KV
cache; users tuning this should be aware that per-call context
shrinks proportionally.

`unsloth studio` (the bare default command, no subcommand) still
defaults to llama_parallel_slots=1 via `run_server`'s own default;
this PR does not change that path -- it only exposes the knob on the
one-liner `studio run` command that already silently used 4.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Forward --parallel through venv re-exec and drop colliding short aliases

`unsloth studio run` re-execs into the Studio venv when invoked from
outside it (the common path). The arg-builder forwards every typer
option but the new --parallel, so the child re-execs at the default 4
and any user value is silently dropped. Worse: pre-PR users who
already pass `-np N` as a pass-through extra (where llama.cpp's
last-wins parsing made it stick) silently lose N after this PR lands.
Forward --parallel explicitly in the re-exec arg list.

While auditing the re-exec path, also drop the colliding 1-char
short aliases -m (--model) and -f (--frontend) plus the redundant
-hfr. Click's short-option clustering had been silently mis-parsing
~11 llama-server short flags via the pass-through path: -fa as
`-f a`, -mg 0 as `-m g` + stray 0, -fitt 1024 as `-f itt` + stray
1024, -hff path as `-f f` + stray `-h path`, -cmoe / -cram / -sm /
-ncmoe etc. The docstring promise ("any flag this command does not
recognize is forwarded verbatim") was silently violated.

-hf (2-char) is kept because Click treats multi-char shorts atomically
(no clustering of -hff / -hfv / -hffv / -hft) and -hf is documented
in basics/api/README.md. --model / --hf-repo / --frontend long forms
all unchanged. studio_default keeps -f because it has no pass-through.

Tests:
- test_studio_run_parallel_flag.py: 8 new re-exec coverage cases
  (all 3 aliases, 3 platforms via sys.platform mock, pre-PR `-np`
  regression, mixed with pass-through extras).
- test_studio_run_short_alias_clashes.py (new): surface checks that
  the removed shorts cannot reappear, plus 11 parametrized cases
  proving each previously-broken llama-server short flag now passes
  through verbatim, plus a happy-path test that documented -hf still
  works for `org/repo:variant` syntax.

All 27 tests pass. Negative test (revert either fix) shows the new
tests catch the regression.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix stale studio run docstring describing rejected llama-server flags

The pre-PR docstring listed --port, -c / --ctx-size, --api-key, -ngl,
--jinja, --flash-attn, --no-context-shift as "rejected with HTTP 400",
but only --port and --api-key (plus other networking / auth / model
identity / single-model UI flags) are actually in
studio/backend/core/inference/llama_server_args.py's denylist. -c /
-ngl / --jinja / --flash-attn / --no-context-shift are pass-through
and last-wins-override Studio's auto-set value.

Rewrite the docstring to match the real denylist groups and point at
the canonical source. Also add --parallel to one of the examples now
that it is a first-class flag.

* ci: broaden Linux + narrow Windows llama.cpp runtime patterns + trim #5741 comments (#5746)

* ci: broaden Linux llama.cpp runtime pattern to lib*.so*

#5741 patched the explicit Linux pattern list to add
``libllama-*-impl.so*`` after ggml-org/llama.cpp#23462 (between
b9279 and b9283) split each binary's entry code into a paired
``lib<binary>-impl.so`` shared library. Same class of upstream
repackaging will hit us again whenever a new shared lib is added.

Mirror what macOS already does and replace the per-lib list with a
single ``lib*.so*`` glob. ``copy_globs`` (line 3614) unions
patterns, so the per-variant ``libggml-cuda.so*`` / ``libggml-hip.so*``
entries were never filtering anything; the spec lives in
``runtime_payload_health_groups`` (line 5209) which keeps the
explicit minimum-required list per variant.

Dry-run against b9296-bin-ubuntu-x64.tar.gz: 40 files copied (all
ggml, llama, mtmd, impl variants + the two binaries we ship), 22
skipped (other CLIs, rpc-server, LICENSE). Functionally equal to
the post-#5741 set.

* cleanup: trim #5741 comments on the pydantic split

Comments added in #5741 explained the original bug in full each
time. They are mostly redundant with the commit message and the PR.
Trim them to one short paragraph per site.

No behavior change.

* ci: narrow Windows runtime pattern to llama-server.exe + llama-quantize.exe

Studio only invokes llama-server and llama-quantize. Mac and Linux
already filter to those two binaries; Windows was the odd one out
with ``*.exe`` copying every CLI upstream ships (llama-cli,
llama-bench, llama-mtmd-cli, ...).

Dry-run on b9296 (win cpu-x64, cpu-arm64, cuda-13.1, hip-radeon):
20 unused EXEs skipped per variant, all DLLs (incl. the new
llama-*-impl.dll family) still copied via ``*.dll``.

``existing_install_matches_choice`` already checks llama-server.exe
exists explicitly (line 5297), so the health gate is unchanged.

* Lower default weight_decay in RL config from 0.01 to 0.001 (#5747)

In full FT, AdamW weight decay shrinks the parameter directly so the
implicit prior is W -> 0. In LoRA the trained parameters are A and B
while the effective weight is W = W_init + (alpha/r) * B @ A; decaying
A and B separately drives BA -> 0, hence W -> W_init rather than 0.
The previous default of 0.01 inherited from full-FT recipes adds a
measurable pull on the merged adapter back toward the base model over
a few thousand steps. 0.001 keeps a small Frobenius-norm prior on
||A||^2 + ||B||^2 for numerical stability without meaningfully biasing
the merged weight toward init, and aligns with the value used across
the unsloth notebook templates.

* Studio: strip orphan tool_call XML leaking into visible content (#5735)

* Studio: strip orphan tool_call XML from streamed visible content

The speculative-buffer state machine in
`studio/backend/core/inference/llama_cpp.py` can slice a tool_call XML
block between the silent DRAINING path and the user-visible
content_accum, depending on when in the model's emission the BUFFERING
-> STREAMING -> DRAINING transitions fire. Three leak shapes were
observed in a 2026-05-22 sweep of 900 Qwen3.5 / Qwen3.6 GGUF runs:

  Pre-fix XML leak rate: 20/900 (2.22%), concentrated 6.7% on the
  larger Q8 / MTP configs:

    Qwen3.6-35B-A3B Q8_0         4/60  (6.7%)
    Qwen3.6-35B-A3B-MTP Q4       4/60  (6.7%)
    Qwen3.5-35B-A3B Q8_0         3/60  (5.0%)
    Qwen3.6-27B Q8_0             3/60  (5.0%)

The existing `_TOOL_XML_RE` only matched well-formed
`<tool_call>...</tool_call>` and `<function=...></function>` pairs, so
unterminated openings (close was DRAINED) and orphan closes (opening
was DRAINED) survived the strip and reached the user.

Fix relaxes the regex to also strip:
  1. Orphan opening up to end-of-string: `(?:</tool_call>|\Z)`
  2. Orphan closing tag: bare `</tool_call>` / `</function>`

Verified on the full sweep: 20/900 -> 0/900 (100% of detected leaks
eliminated). 16 unit tests in `test_tool_xml_strip.py` pin all three
leak shapes plus the well-formed cases, plus parametrised checks on
the 5 actual real-world leak samples from the sweep data.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: strip tail-only </parameter> orphan + tighten regex

The 2026-05-22 gdpval sweep surfaced a 4th XML-leak shape not caught
by the earlier regex: a bare `</parameter>\n\n` at end-of-buffer (7
of 192 trials, all Qwen3.5-27B + a few Qwen3.6-27B). The model emits
the full `<tool_call><function=...><parameter=...>...content...
</parameter></function></tool_call>` envelope, the speculative buffer
DRAINS the opening tags as intended, but EOS (max_tokens cutoff)
truncates the outer `</function></tool_call>` close, leaving just
`</parameter>` as the visible tail.

We strip this ONLY when end-anchored (`\s*\Z`) so legitimate
mid-text uses (user code samples, documentation discussing the
Qwen tool-call XML shape) survive. Verified on the 192-trial
gdpval corpus: before=7, after=0.

While at it, fold the five top-level alternations into three by
sharing tag-name and prefix subgroups:

  <tool_call>...    + <function=\w+>...    +    -->  <(?:tool_call|function=\w+)>...
  </tool_call>      | </function>                  -->  </(?:tool_call|function)>

Semantically identical (verified by replay over the 192-trial
corpus + adversarial inputs, 0 diffs) and 1.34x faster on real
workloads. Backtracking-safety pinned by two new perf guards
(256KB '<' spam, 1000x orphan opens).

Tests: 16 -> 28 (6 new functional + 4 well-formed-vs-orphan +
2 perf guards).

* Tighten comments in XML-strip regex and tests

Code says what it does; comments were repeating it. Strip the verbose
explanations down to the WHY-only bits (engine quirk, tail-anchor
rationale, real-world source of each test sample). No code changes.

inference.py:  21 -> 12 lines around _TOOL_XML_RE
test_tool_xml_strip.py: 343 -> 259 lines (-84)
Tests: 28/28 still pass.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>

* Address review: deny pass-through --parallel, preserve legacy short aliases, fix test harness

Round 1 review fixes for #5737:

1. Deny --parallel / --n-parallel / -np in the pass-through validator.
   Without this, `unsloth studio run --model X --parallel 8 -- --parallel
   999` would last-win-override the running llama-server slot count while
   Studio's app.state.llama_parallel_slots and KV-cache fitting stay at
   the typer value (8), so the resource plan and the running process
   disagree. Also bypasses the typer 1..64 range guard. Reject so the
   only path is the first-class typer flag.

2. Backwards-compat shim for -m / -hfr / -f. Dropping the short aliases
   from typer broke any script using `unsloth studio run -m X` or
   `-hfr Y` or `-f dist`. Add _consume_legacy_short_aliases which pops
   EXACT whole-token matches (or `-x=value` inline form) from ctx.args
   into the corresponding typer parameter. Clustered tokens (`-fa`,
   `-mg`, `-fitt`, ...) are left in the pass-through tail unchanged.
   --model becomes Optional with an explicit missing-required check
   after the preprocessor so legacy `-m X` still satisfies the
   "must specify a model" requirement.

3. Drop mix_stderr from CliRunner. Typer 0.25.1 / Click 8.4.1 removed
   the kwarg; the test harness raised TypeError before exercising the
   PR behaviour. Tests run cleanly on current and older Typer/Click.

4. Correct the -np regression test docstring. Pre-PR `-np 8` was
   clustered by Click as `-p 8` (port=8) + stray `-n`, silently
   breaking the port binding -- not "passed through as 8 slots". The
   post-PR assertion (child gets --parallel 8) is unchanged.

5. Update studio run docstring listing rejected flags so it now
   correctly includes --parallel / -np / --n-parallel.

New tests:
- test_llama_server_args.py: parametrized denylist coverage for
  --parallel / --n-parallel / -np including equals-form, including
  out-of-range bypass attempts (999, 0). is_managed_flag flips True.
- test_studio_run_short_alias_clashes.py: legacy -m / -hfr / -f
  promote to typer params; --model X + -m Y conflict errors; clustered
  -mg / -fa / -fitt still pass through (the original bug fix holds).

132 tests pass (98 backend + 34 cli).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Extend legacy-alias shim tests for repo:variant, inline value form, and missing model

Three additional edge cases for the -m / -hfr / -f preprocessor:
- `-m unsloth/foo:UD-Q4_K_XL` round-trips through both the preprocessor
  and _split_repo_variant so the child sees --model + --gguf-variant.
- `-m=foo` inline value form is promoted just like `-m foo`.
- Missing --model after the preprocessor raises typer.Exit(2) cleanly
  (replacing typer's pre-PR required-flag enforcement now that --model
  is Optional to allow the legacy promotion path).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Scrub .github/workflows for staging push (matches staging base)

* Fix studio CLI argv handling and pass-through docstring drift

- studio/backend/core/inference/llama_server_args.py: drop the stale
  ``-np``/``--parallel`` entry from the docstring's pass-through tunable
  list. These flags moved into _DENYLIST_GROUPS so the docstring now
  contradicts the validator and would mislead future maintainers
  debugging the ValueError from validate_extra_args(["--parallel","8"]).
  The deleted wording was introduced by dbea77e34 ("Studio: forward
  llama-server args from `unsloth studio run`, activate `unsloth run`,
  and allow passing model:quant to load models") when --parallel was
  still a documented pass-through; the same commit's "quant" reference
  is about the model:quant syntax, unrelated to the parallel slot
  wording being deleted here.

- unsloth_cli/commands/studio.py: add _expand_attached_np_short next to
  _consume_legacy_short_aliases. Both work around Click's short-option
  clustering for this command -- the legacy preprocessor for `-m` / `-f`
  / `-hfr` and this one for the attached `-np<N>` form. Click clusters
  `-np8` as `-n -p 8` because `-p` is the typer short for `--port`,
  silently setting port=8 and dropping the parallel value; rewriting the
  attached form into separated `-np <N>` in sys.argv before Click
  parses preserves the user's value. Space/equals forms (`-np 8`,
  `-np=8`) already work and are left alone.

- unsloth_cli/__init__.py: import _expand_attached_np_short from the
  studio command and run it only when argv[0] looks like the unsloth
  console-script or workspace cli.py, so importing this module from a
  notebook or pytest run does not mutate the caller's argv.

* Tighten the -np canonicaliser comments

Drop the helper's co-location sentence (location is self-evident from
grep) and shorten the entry-gate rationale to one short sentence
covering the why.

* Sync .github/workflows with upstream author branch

* Sync .github/workflows with upstream author branch

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Bump install.sh / install.ps1 pin to unsloth>=2026.5.7 (#5753)

PyPI release unsloth 2026.5.7 is now live. Bumps the pinned floor in
install.sh and install.ps1 from unsloth>=2026.5.6 to unsloth>=2026.5.7
so fresh installs resolve to the new wheel.

Tagged on main as v0.1.416-beta.

* Catch attached `-np<N>` form in backend pass-through validator

The CLI-side `_expand_attached_np_short` rewrites `-np8` to `-np 8`
before Click parses, but HTTP /load `llama_extra_args=["-np8"]` goes
straight to `validate_extra_args` which only matched the exact token.
Reproducer: `validate_extra_args(["-np8"])` previously returned
`["-np8"]` instead of raising; once forwarded to llama-server it
last-win-overrode Studio's slot count while
`app.state.llama_parallel_slots` stayed at the typer value.

Normalise `-np<digits>` to `-np` in `_flag_name` so the denylist
catches the attached form alongside `-np`, `-np=8`, `--parallel`,
`--parallel=8`, and `--n-parallel`. Tests parametrize the new form
including out-of-range values.

* Restore _consume_legacy_short_aliases unit tests + _expand_attached_np_short tests

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Restore .github/workflows from origin/main

Earlier merge from claude_review's staging-scrub commits accidentally
deleted production CI workflows. Restore them to main's state.

* Scrub .github/workflows for staging push (matches staging base)

* Sync .github/workflows with upstream author branch

* Round 5+6: broaden -np gate to exact basenames + runtime parallel test

Reviewer-flagged improvements squashed into one commit so the auto-push
review bot doesn't keep stomping the branch:

- unsloth_cli/__init__.py: exact-basename match instead of
  endswith('cli.py'). Covers unsloth, unsloth.exe, unsloth-cli,
  unsloth-cli.exe, cli.py, unsloth-cli.py. A third-party mycli.py that
  happens to import unsloth_cli no longer has its argv mutated.

- unsloth_cli/tests/test_studio_run_parallel_flag.py: parametrised
  runtime test (N in {1, 4, 8, 64}) that fakes the in-venv path and
  asserts run_server is invoked with llama_parallel_slots=N.
  Complements the existing source-text check so refactors that preserve
  runtime semantics don't trip a false failure.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Round 7: respect '--' end-of-options and reject flag-as-value

Round 7 reviewer flagged three legitimate edge cases:

- _expand_attached_np_short rewrote post-'--' tokens. Convention: '--'
  ends option processing; payload after it is raw. Stop the loop there.

- _consume_legacy_short_aliases promoted post-'--' legacy aliases for
  the same reason. Treat post-'--' tail as raw.

- Legacy '-m -fa' silently consumed '-fa' as the model name, hiding
  the real CLI shape error. Reject any next-token that starts with '-'
  (except the lone '-' stdin/path sentinel) with a clear BadParameter.

Also expanded the missing-model error string to mention the still-
supported legacy '-m' / '-hfr' aliases so users hitting that diagnostic
on legacy scripts get the right migration hint.

Added four regression tests covering each new behaviour.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Round 8: soften flag-as-value to long-form only + normalise is_managed_flag

Round 8 reviewer flagged two cleanups:

- _consume_legacy_short_aliases rejected any next token starting with
  '-' as a flag, which would break legitimate values like '-foo'
  (path or model name with leading dash). Narrow the rejection to
  '--long' tokens only; '-x' short forms still pass through.

- is_managed_flag did raw _DENYLIST membership while validate_extra_args
  goes through _flag_name first, so '-np8' / '--parallel=8' /
  '--port=9000' classified as not-managed by the helper but rejected
  by the validator. Route is_managed_flag through _flag_name so the
  two helpers agree on every form callers might use.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Round 9: also catch -np-1 / -np+1 signed attached forms in denylist

Round 9 reviewer noticed _flag_name normalised -np<digits> but missed
signed variants -np-1 and -np+1, so validate_extra_args waved them
through while rejecting --parallel -1. llama.cpp would error out on
negative slot counts anyway, but the validator should classify every
form of the managed flag identically so the boundary is consistent.

* Round 10: signed -np in CLI canonicaliser + reject empty inline aliases

Round 10 reviewer flagged two real issues:

- _expand_attached_np_short rewrote only -np<digits>; signed forms
  -np-1 / -np+1 fell through. Backend _flag_name already classifies
  them as managed, so the CLI rewriter must too -- otherwise Click
  clusters -np-1 into -n -p -1 (port=-1) and never reaches the
  backend validator at all.

- -m= / -hfr= / -f= empty inline forms were accepted and produced
  --model '' / --frontend '' (then Path('') silently became '.') on
  re-exec. Reject empty inline values at the preprocessor with a
  clear BadParameter so the malformed input fails fast.

Both behaviours pinned with parametrised regression tests.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Expose --parallel on plain `unsloth studio` for API-path parity

The PR added --parallel to `unsloth studio run` but the plain
`unsloth studio` callback (used for API-only / bare-server launches)
still hardcoded llama_parallel_slots to its run_server default. With
--parallel now denied as a llama_extra_args pass-through, that flow
had no first-class way to raise concurrency.

- unsloth_cli/commands/studio.py: add --parallel / --n-parallel typer
  Option (default 4, range 1..64) to studio_default, forward through
  the venv re-exec, and pass llama_parallel_slots= to run_server in
  the in-venv path.
- studio/backend/run.py: argparse --parallel / --n-parallel with the
  same range guard so the spawned child accepts the forwarded flag.
- unsloth_cli/tests/test_studio_run_parallel_flag.py: test pins the
  new option presence, aliases, default and range guards.

* Round 12: narrow entry-point gate, preserve pre-PR plain-studio default, drop brittle source-text test

Three Opus subagent reviewers (security / backcompat / code-quality)
flagged the same handful of real issues. Consensus fixes:

- unsloth_cli/__init__.py: narrow the -np canonicaliser gate to just
  {unsloth, unsloth.exe} (the only pyproject-declared console_script).
  The previous cli.py / unsloth-cli.py entries would silently rewrite
  sys.argv for any third-party myproj/cli.py that happens to import
  unsloth_cli. Dev users running python cli.py ... -np N still work
  via the space form, which parses without the rewrite.

- unsloth_cli/commands/studio.py + studio/backend/run.py: restore the
  pre-PR llama_parallel_slots default of 1 on plain unsloth studio and
  python studio/backend/run.py. unsloth studio run keeps its
  hardcoded-pre-PR default of 4. Without this, my earlier API-path
  parity commit silently dropped per-call context to ctx/4 for the
  plain-studio flow.

- unsloth_cli/tests/test_studio_run_parallel_flag.py: drop the brittle
  source-text grep test (test_run_kwargs_use_parallel_value). The
  parametrised runtime test test_in_venv_path_passes_parallel_to_run_server
  already pins the same intent against actual behaviour.

- unsloth_cli/tests/test_studio_run_short_alias_clashes.py: pin the
  narrow entry-point gate with a parametrised negative test covering
  seven third-party argv[0] basenames (cli.py, /path/myproj/cli.py,
  pytest, unsloth-cli, etc.). Re-broadening the gate now trips a
  test instead of silently mutating an unrelated CLI's argv.

* Round 13: shared parallel constants, denylist invariant test, defence-in-depth

Three Opus subagent reviewers (adversarial-user / maintenance /
cross-file consistency) flagged a consistent set of cleanups; folded
into one commit to avoid the pre-commit.ci force-push race.

unsloth_cli/commands/studio.py:
- Extract _PARALLEL_MIN / _PARALLEL_MAX / _PARALLEL_DEFAULT_RUN /
  _PARALLEL_DEFAULT_PLAIN module-level constants and use them in both
  typer Options (plain studio_default = 1, studio run = 4).
- _expand_attached_np_short now rewrites -np<junk> when the suffix
  starts with a digit (or signed digit) so '-np8x' surfaces as a
  clean '-np takes an int' typer error instead of a baffling
  '--port invalid' complaint after Click clusters '-n -p 8x'.
- Re-exec forwarding emits --load-in-4bit / --no-load-in-4bit
  explicitly in both directions; previously the True default relied
  on both layers sharing the same default forever.
- run() docstring now explicitly says --parallel / -np pass-through
  via llama_extra_args is denied (use the typer flag above).

studio/backend/run.py:
- Mirror the parallel constants and route the argparse default,
  range check, and error message through them. Help text mentions
  the asymmetry with 'unsloth studio run' so direct-launch dev users
  aren't confused by Default 1 in isolation.

studio/backend/core/inference/llama_server_args.py:
- _flag_name strips surrounding whitespace before denylist lookup so
  a caller can't slip a managed flag past the boundary with a
  trailing space (the trimmed form is what downstream parsers see).

Tests:
- New typer-aliases-subset-of-denylist invariant: every alias the
  typer Option claims as --parallel on run() MUST be in the backend
  parallel denylist group. Catches the failure mode where someone
  adds a new alias and forgets the boundary.
- Extended denylist parametrize to cover ~14 previously untested
  aliases (-mu, -dr, -hfv/-hfrv/-hffv family, -mmu, full --ui group,
  --models-preset / --models-autoload / --no-models-autoload).
- Whitespace-padded denylist rejection (' --parallel', '-np ', etc).
- --load-in-4bit re-exec test pinning both polarities + default.
- -np<junk> argv rewriter regression tests.
- Cross-reference headers between the two test files.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: repair mlx studio base export save_method (#5727)

* Round 14: align backend -np recogniser with CLI rewriter + reject parent --parallel

Round 14 (reviewer.py --parallel 20 with gpt-5.3-codex-spark) flagged
two real P1s and a stale-rebase warning. All three addressed.

- studio/backend/core/inference/llama_server_args.py: widen
  _flag_name so -np<digit-prefix> with trailing junk (-np8x,
  -np-1foo, -np+1bar, -np9zzz) classifies as managed flag -np,
  matching the CLI _expand_attached_np_short rewriter. Without this,
  POST /api/inference/load with llama_extra_args=['-np8x'] slipped
  past the boundary while the CLI canonicalised the same form. The
  two sides now agree on every digit-prefix form.

- unsloth_cli/commands/studio.py: reject --parallel on the
  studio group when a subcommand is invoked. Pre-PR the studio
  callback had no --parallel; my Round 12 addition made
  'unsloth studio --parallel 8 run ...' silently drop the 8
  because typer doesn't propagate parent options into subcommand
  kwargs. Now errors with exit 2 and a message pointing the
  operator at the correct invocation
  ('unsloth studio run --parallel 8 ...').

- Picked up origin/main via merge (parent commit 0caf0526): the
  pre-flight stale-rebase detector found 2 lines on main in
  studio/backend/core/export/export.py missing from PR HEAD.
  Merged cleanly with no conflicts.

Tests:
- Parametrised denylist coverage for -np<digit-prefix>+junk forms.
- New runtime test confirms exit 2 + helpful error when the group
  --parallel is supplied alongside an invoked subcommand.
- Test that the default group --parallel value still lets a
  subcommand resolve (no false-positive rejection).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: tighten code comments across --parallel PR

Comment-only pass over the seven PR-touched files; trim verbose
docstrings, collapse multi-line section dividers, and drop
redundant prose that the code already conveys. No behaviour change.

* Studio: trim remaining verbose docstrings missed in last pass

Shorten the test_studio_run_parallel_flag.py module docstring and
the `Re-exec arg-builder coverage` block. No behaviour change.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: second comment-tightening pass across PR-touched code

Trim docstrings and inline comments in studio.py, run.py,
llama_server_args.py, and unsloth_cli/__init__.py. No behaviour change;
all 215 tests still pass.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: deny --embedding / --rerank / --tools pass-through

`--embedding` and `--rerank` flip llama-server into single-endpoint
mode, which breaks Studio's /v1/chat/completions hop. llama-server's
own `--tools` flag silently stacks on top of Studio's tool policy
resolved by `--enable-tools` / `--disable-tools`.

Add all three (plus the `--embeddings` / `--reranking` plural aliases)
to the boundary denylist so HTTP /load and pass-through extras both
reject them cleanly instead of silently desyncing the server surface.

Test added to the existing `test_denylist_rejects_all_aliases`
parametrize. 220 tests pass.

* Studio: make PR-touched tests robust to minimal envs + Windows

Two cross-OS CI findings:

1. `test_typer_parallel_aliases_are_subset_of_backend_denylist` was
   doing `from core.inference.llama_server_args import _DENYLIST_GROUPS`
   which triggers `core/inference/__init__.py` and pulls in the full
   backend chain (fastapi / structlog / loggers / utils.hardware).
   The invariant only needs the constants tuple, so load the module
   directly via `importlib.util.spec_from_file_location` -- the test
   now runs with just typer + pytest installed.

2. `test_legacy_frontend_alias_still_promotes_to_frontend` asserted
   the literal string `"/tmp/dist"` after the value round-trips through
   `Path()`. On Windows `str(Path("/tmp/dist"))` is `"\tmp\dist"`, so
   the assertion tripped on the same logical path. Compare via
   `Path(x) == Path("/tmp/dist")` so the test passes on every OS.

Both surfaced by the staging-4 cross-OS CI; no production-code change.
220 tests still pass locally.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: load llama_server_args.py directly in its unit tests

Same fix as the previous CLI-test commit: import the module via
`importlib.util.spec_from_file_location` instead of
`from core.inference.llama_server_args import ...`, so the test no
longer needs the full backend chain (fastapi / structlog / loggers /
utils.hardware) installed via `core/inference/__init__.py`.

The boundary validator is intentionally dependency-free; its unit
tests should reflect that.

* Fix test_main_composer_has_dir_auto anchor after PR #5784

PR #5784 ("Improve image generation UI") rewrote the message-input
textarea's static `aria-label="Message input"` into a JSX conditional
`aria-label={overlay ? "Image edit instructions" : "Message input"}`
but did not update the RTL bidi-attribute regression test, leaving
the literal-string `find('aria-label="Message input"')` anchor with
no match. The `Repo tests (CPU)` job has been red on main since.

Anchor on the inner `"Message input"` string literal instead -- it
survives both spellings and still pins the same textarea element so
the `dir="auto"` assertion has the right block to inspect.

Verified by re-running the exact CI command:
  954 passed, 3 skipped, 23 deselected (was 948 passed, 1 failed).

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Long Yixing <longyixing331@gmail.com>
2026-05-26 23:13:45 -07:00
oobabooga
ef6744253f
Studio: stop the model from replying twice when it refuses (#5775) 2026-05-26 05:30:52 -07:00
Daniel Han
41d24227cd
Studio: per-card web_search result + shell_call output fallback (OpenAI) (#5785)
* Studio: per-card web_search result + shell_call output fallback (OpenAI)

Two empty-output bugs in the OpenAI Responses tool-result rendering that
showed up clearly when a single prompt invoked 9 web_search + 4
code_execution + 1 image_generation in one turn. Reproduction shape in
the SQLite-stored chat history:

- 8 of 9 web_search tool-call records had result == "" (the cards
  rendered as empty cards in the thread)
- 4 of 4 code_execution (shell_call) records were missing the result
  key entirely (NoneType), so the cards that showed "Ran cat ..." style
  commands displayed the command line but no output panel at all
- image_generation worked, as did the very last web_search of the run

Root causes in studio/backend/core/inference/external_provider.py:

1. web_search_call's tool_end emitted result: "" by design, with the
   intent of overwriting only the LAST call at response.completed with
   the full citation list (the source-pill extractor on the frontend
   flatMaps across every web_search result, so a single non-empty
   result is enough for the trailing source pills). Side effect: every
   intermediate card renders empty in the thread. Fix: seed each call's
   own tool_end result with "Searching: <query>" so the per-card text
   is never empty, then keep the last-call overwrite path so the
   source-pill extractor still works. Falls back to empty when the
   model emits an action with no query, so the existing last-call path
   stays unchanged for that edge.

2. shell_call's tool_start was emitted from
   response.output_item.done for the call item, but tool_end lived in
   the separate response.output_item.done handler for shell_call_output.
   When OpenAI's Responses stream bundles the output array onto the
   shell_call item's own done event (no separate shell_call_output
   item), the previous handler emitted tool_start with no following
   tool_end. The card spun on "running" indefinitely and stored as
   NoneType in the thread DB. Fix: when the shell_call's done event
   carries an embedded output list, emit tool_end immediately from
   that. Track tool_end_emitted on the shell_calls map so a subsequent
   shell_call_output event (some streams ship both) is skipped instead
   of double-completing the card. A final flush at response.completed
   emits tool_end for any orphan shell_call that received neither
   bundled output nor a separate output event, so cards always finalise.

Tests (studio/backend/tests/test_openai_tool_result_fallbacks.py, 6
new):
- web_search: three calls, each card's result is its own Searching:
  query (no empties)
- web_search: last call still gets the aggregated citation block when
  url_citations arrive (pins the overwrite path)
- web_search: empty action.query falls back to result == "" (no junk
  Searching: placeholder)
- shell_call: bundled output on done emits a single tool_end with that
  output as the result text
- shell_call: bundled-then-separate output does not double-emit
  tool_end (subsequent shell_call_output is skipped)
- shell_call: orphan call with neither bundled nor separate output is
  flushed at response.completed so the card finalises

15/15 tests green when combined with the existing 9 in
test_openai_code_execution.py. Pre-commit + ruff format clean.

Scope: OpenAI Responses-API code path only. The Anthropic native
Messages-API path (_stream_anthropic) is untouched, as is the local
llama-server path. Local-model behaviour cannot regress because the
edited handlers only fire inside the OpenAI cloud branch.

* Studio: per-model external max_tokens cap + clamp on model switch

Two related external-provider issues that surfaced from the same
investigation as the per-card web_search / shell_call result bugs in
the previous commit:

A. Slider cap was a one-size-fits-all 32768 for every external model.

   provider-capabilities.ts kept a single EXTERNAL_MAX_OUTPUT_TOKENS
   constant (32k), well below what most providers actually accept. The
   docstring even called out the right per-provider numbers (Anthropic
   Opus 128k, GPT-5.x ~128k, Gemini 2.5 ~65k, DeepSeek 8k) but the
   code picked the lowest as a conservative floor. Effect: long
   generations from gpt-5.5 / claude-opus-4-7 silently truncated at
   32k even though the API would have served up to 128k.

   Fix: introduce getExternalMaxOutputTokens(providerType, modelId)
   returning the documented per-model cap. Patterns are checked
   longest-first so e.g. gpt-5.5-pro matches before gpt-5.5. Unknown
   provider/model combinations fall back to the existing 32k floor so
   no surprise increases for ids we don't know about.

   Per-model caps from the official docs:
   - OpenAI gpt-5.5 / gpt-5.5-pro: 128000
   - OpenAI gpt-5.4 / gpt-5.4-pro: 65536
   - OpenAI gpt-5.3: 16384
   - Anthropic claude-opus-4-7: 128000
   - Anthropic claude-opus-4-6 / sonnet-4-6 / opus-4-5 / sonnet-4-5 /
     haiku-4-5: 64000
   - Gemini 3.x family: 65535
   - DeepSeek: 8192
   - OpenRouter: strip provider/ prefix from the id and re-resolve

   The slider in chat-settings-sheet.tsx and the send-time clamp in
   chat-adapter.ts both call the new function so the slider's max=
   matches what the wire layer will accept.

B. Slider value lied after switching from a local model to external.

   When Studio auto-loads the helper Gemma-4-E2B-it on first chat,
   chat-adapter sets params.maxTokens to Gemma's context_length
   (262144 for Gemma 4). Switching the model picker to gpt-5.5 then
   flips the slider's max prop to the external cap, but the stored
   params.maxTokens is never reset. The numeric value next to the
   slider would render 262144 against a track that ended at the
   external cap. The send-time clamp brought the outbound max_tokens
   back down to the cap, so the API call was safe, but the displayed
   number had no relationship to what was actually being sent.

   Fix: chat-runtime-store.setCheckpoint now clamps params.maxTokens
   to getExternalMaxOutputTokens(...) on transitions into an external
   model. Looks up the provider via useExternalProvidersStore so we
   can derive providerType from the parsed external model id. No-op
   when the stored maxTokens is already at or below the new cap, so
   user-tuned values within range survive the switch.

Scope: pure frontend changes scoped to external-provider code paths.
Local model behaviour is untouched -- the ggufContextLength branch of
the slider's max= is unchanged, and setCheckpoint only mutates
maxTokens when isExternalModelId(modelId) is true. The send-time
clamp continues to be the safety net for any in-flight request that
crosses a model switch before the store-level clamp has applied.

Typecheck (tsc -b) clean; bun run build succeeds (2.13s).

Co-changes with the previous commit (7fe1adbf, per-card web_search +
shell_call output fallback) form a single PR: every empty-output and
silent-truncation issue surfaced from the same animal-popularity
prompt reproduction is now addressed in one branch.

* Studio: correct external max_tokens caps for Gemini and DeepSeek

Per-doc corrections to the per-model cap table added in 95da8d52:

- Gemini 3.x family: 65535 -> 65536, per
  https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro-preview
  (the published max_output_tokens is exactly 64K = 65536). The earlier
  65535 was an off-by-one rough cap.
- DeepSeek (deepseek-chat / deepseek-reasoner aliases): 8192 -> 384000,
  per https://api-docs.deepseek.com/quick_start/pricing. DeepSeek V4
  Flash / Pro both list MAX OUTPUT = 384K; the chat / reasoner ids are
  deprecated aliases for V4 Flash non-thinking / thinking modes. The
  8192 value was carried over from V3 and silently truncated V4 traffic
  at 2% of its actual ceiling.

Affects only the slider max and the send-time clamp for these provider
types. Other providers' caps unchanged. tsc -b clean.

* Studio: also flush orphan shell_calls on response.incomplete

Addresses gemini-code-assist[bot] high-priority inline review on PR
5785: the orphan-shell_call final flush added in 7fe1adbf landed only
in the response.completed branch. Truncated OpenAI Responses streams
emit response.incomplete instead (for example when the request hits
max_output_tokens), which left in-flight shell_call cards spinning
indefinitely in the UI.

Mirror the same flush block in the response.incomplete handler so the
truncated-stream path finalizes every pending tool card. The
tool_end_emitted guard keeps the path idempotent: if a shell_call
already completed via bundled output on its done event, the incomplete
flush is a no-op for it.

Two new tests in test_openai_tool_result_fallbacks.py:
- test_shell_call_flushed_on_response_incomplete_truncation pins the
  bug repro: an in-flight shell_call followed by response.incomplete
  must emit tool_end so the card finalizes.
- test_shell_call_incomplete_does_not_double_emit pins idempotency:
  a shell_call that completed via bundled output and is then followed
  by response.incomplete emits exactly one tool_end with the bundled
  result text.

17/17 tests green (8 fallback tests + 9 existing code-execution). Pre-
commit + ruff format clean.

* Studio: trim verbose comments across PR 5785 edits

Compress the in-code commentary added across this branch to one or two
lines per block; the verbose prose was easier as a PR description than
as inline noise. No behavioural changes: 17/17 tests still green, tsc -b
still clean.
2026-05-26 04:31:22 -07:00
Wasim Yousef Said
31ac558a73
Recipe Studio local model selector (#5769)
* feat(recipes): round-trip local model variants

* feat(recipes): add local model selector

* feat(recipes): wire selector into model editors

* fix(recipes): clear stale model state on relink

* feat(recipes): load selected local models for jobs

* chore(frontend): simplify biome scripts

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(recipes): handle local selector edge cases

* fix(recipes): polish local model selector behavior

* fix(recipes): delay local model restore until terminal runs

* fix(recipes): accept resolved default gguf variants

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-26 02:37:24 -07:00
Wasim Yousef Said
f364b08b6c
Support follow-up edits for generated images (#5712)
* Support follow-up edits for generated images

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix generated image edit references

* Fix image generation tool guard

* Replay OpenAI image reasoning refs

* Capture streamed OpenAI reasoning refs

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Trigger Gather Town test

* Use OpenAI response context for image edits

* Studio: improve generated image card UI

* Studio: refine generated image overlay and edit context

* Studio: animate generated image loading state

* Studio: bind generated-image edits to selected image

* Studio: drop empty external assistant payloads

* Studio: guard generated image clipboard MIME

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Simplify image generation edit wiring

* Remove temporary image generation test changes

* Preserve explicit image edit references

* Scope image edit references to threads

* Harden OpenAI image edit replay handling

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Update image generation tool event test

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-26 02:34:06 -07:00
Daniel Han
542d74370f
Studio: pricing follow-up to #5690 (longest-prefix match + chat-style usage keys) (#5722)
* Studio: longest-prefix pricing match + accept chat-style usage keys

Two P1 / High follow-ups from PR 5690 review feedback:

1. Pricing prefix lookup returned the first key it iterated, so
   dated snapshots like ``gpt-5.4-mini-2026-04-23`` collided with
   the shorter ``gpt-5.4`` entry and overbilled by 3x+. Sort the
   table keys longest-first so the most specific entry wins.

2. ``calculate_cost`` only read ``input_tokens`` / ``output_tokens``,
   but Studio's OpenAI-Chat-style usage envelope re-emits
   ``prompt_tokens`` / ``completion_tokens`` (the OpenAI Chat
   Completions vocabulary). Callers handing in the chat-style
   shape silently got a zeroed bill. Accept either pair so the
   calculator works against both raw upstream usage and the
   Studio-translated envelope.

Tests (4 new in test_pricing.py): dated mini/pro snapshots inherit
the right rate; chat-style usage keys price correctly; raw key wins
when both shapes are present.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: dedupe cache buckets when costing chat-style Anthropic usage

When the caller hands in Studio's chat-style envelope (``prompt_tokens``
emitted by ``_build_usage_chunk``) for Anthropic, that value already
folds ``cache_creation_input_tokens`` + ``cache_read_input_tokens`` into
the total. The previous follow-up accepted the chat-style key but then
re-added both cache buckets in ``billable_input_tokens`` and ``input_usd``,
double-counting cache tokens on every Anthropic chat-style call.

Detect which envelope landed (``input_tokens`` present = raw upstream;
absent + ``prompt_tokens`` present = Studio chat-style) and peel the
cache buckets off for Anthropic before the downstream math so both
envelopes produce identical costs.

OpenAI: ``input_tokens`` and Studio's ``prompt_tokens`` both already
include ``cache_read`` and exclude any notional ``cache_creation``, so
the OpenAI path stays a straight passthrough.

Tests (2 new): both envelopes match for Anthropic on a triple
(uncached + cache_creation + cache_read); OpenAI envelopes match on a
cached-tokens fixture.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: prefer raw output_tokens over chat-style completion_tokens

Codex flagged that the previous fallback chain
'usage.get("output_tokens") or usage.get("completion_tokens")'
treats an explicit 0 as missing -- a mixed-envelope payload where
'output_tokens' is 0 but 'completion_tokens' is non-zero (or
stale) bills the wrong amount. Mirror the has_input_tokens
precedence pattern: when the raw key is present we use it even at
0; otherwise fall back to completion_tokens.

* Studio: read OpenAI cached tokens from prompt_tokens_details too

Codex flagged that the chat-style OpenAI envelope Studio re-emits
via _build_usage_chunk surfaces cached prompt tokens under
prompt_tokens_details.cached_tokens, not input_tokens_details. The
OpenAI branch only checked input_tokens_details, so a cache-heavy
chat-style turn billed every cached token at the full input rate
instead of the 0.1x cache_read discount.

Walk both keys when discovering the cached count. New regression
test pins that the two envelopes price identically for a turn with
80k of 100k tokens cached.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: tighten pricing prefix match + clamp corrupt usage

Three follow-ups on the longest-prefix pricing match landed in this PR:

- Prefix match now requires a dash boundary or end-of-string. The
  longest-key sort alone still falsely landed "claude-opus-4-15" on
  the "claude-opus-4-1" row, and "gpt-5.5-prod" on the "gpt-5.5-pro"
  row (a 6x overcharge). Demanding the next character be "-" rules
  out the lookalikes while keeping dated snapshots
  ("gpt-5.4-mini-2026-04-23", "claude-opus-4-7-20260414") landing on
  their canonical row.
- Clamp every token count to >= 0. A corrupted upstream payload
  (negative cached count, off-by-one in a fixture) could previously
  produce a negative bill that masked real spend in the session
  total tooltip.
- Tolerate a non-dict "cache_creation" (e.g. an upstream proxy
  folded the field down to a single int). The current code raised
  AttributeError mid-turn; now it falls back to the 5m-default
  bucket so the rest of the cost calculation still runs.

Adds tests/test_pricing_edge.py with 20 adversarial cases covering
the boundary check, negative / None / zero token values across both
envelopes, cache_read > prompt corruption, the OpenAI long-context
threshold crossover on cache-inflated billable input, malformed
sub-objects, and unknown-provider degradation. Combined suite is
51 tests, all green.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Surface Anthropic cache-read fallback and forward 1h breakdown

Two correctness gaps surfaced on the chat-style usage envelope:

1) Anthropic cache_read fell through to "uncached input" pricing when
   the envelope arrived without the native ``cache_read_input_tokens``
   key (e.g. via a proxy that only emits the mirrored
   ``prompt_tokens_details.cached_tokens`` block). Studio's canonical
   ``_build_usage_chunk`` always sets both so production traffic was
   never affected, but the calculator should accept either as a
   defense-in-depth measure. Add a fallback to read the mirrored
   field when the native one is missing or zero; the native key still
   wins when both are present so the math stays deterministic.

2) ``_build_usage_chunk`` dropped the ``cache_creation`` 5m / 1h
   breakdown. Downstream ``calculate_cost`` then could not apply the
   2x 1h premium and silently fell back to the 5m default,
   underbilling 1h cache writes by 2x on chat-style traffic. Forward
   the breakdown verbatim when the upstream usage carries it.

Tests grow by 4 (20 -> 24): two for the prompt_tokens_details
fallback (with native-precedence pin), one for the chunk shape, one
for the end-to-end pricing parity check at 1h.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Add Anthropic fast_mode pricing multiplier

PR 5715 wires the fast-mode-2026-02-01 beta header + speed:"fast"
field through to Anthropic, but the cost calculator never learnt
about the matching 6x premium documented at
https://platform.claude.com/docs/en/build-with-claude/fast-mode
(Opus 4.7 standard $5/$25 per MTok, fast $30/$150).

This adds:
- ANTHROPIC_FAST_MODE_MULT = 6.0 constant.
- calculate_cost(..., fast_mode=True) applies the 6x to base input
  AND output rates before any cache multipliers (cache mults stack
  on top of fast per Anthropic docs).
- Provider+model gate: silently no-op on every model that is not
  claude-opus-4-6 / claude-opus-4-7 so a stray fast_mode=True on
  Sonnet/Haiku can never over-charge.
- model_priced label tagged "(fast)" so the cost tooltip can
  surface which rate fired.
- pricing_snapshot now exposes fast_mode_mult so the frontend cost
  panel doesn't have to hard-code 6.

7 new edge tests pin the math; existing 55 still pass.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Honor explicit zero cache_read_input_tokens on Anthropic envelopes

The previous follow-up fell back to ``prompt_tokens_details.cached_tokens``
whenever the native ``cache_read_input_tokens`` was missing OR equal to 0,
even though the commit message stated the native key always wins when
present. A proxy that forwards a stale ``prompt_tokens_details`` block
alongside an authoritative ``cache_read_input_tokens: 0`` would then
inflate cache_read past the real native count, posting a false cache_read
line and bumping billable_input_tokens. Switch the gate to native-key
presence so an explicit zero stays authoritative; the mirror only kicks
in when the native key is absent. Add a regression test pinning the
explicit-zero precedence.

* Move fast_mode pricing back to #5715

The fast_mode 6x multiplier landed in two places at once -- here
(f66df7ba) and on #5715 (4f1afdb5) -- since both audits ran in
parallel. Drop the duplicate from this branch so the change lives
in its natural home (#5715, which introduces fast_mode itself);
this PR stays focused on the cache-read fallback + 1h breakdown.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Shorten pricing comments for PR #5722

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-25 23:39:58 -07:00
Daniel Han
4854d4579f
Studio: surface Anthropic document citations inline + in Sources panel (#5718)
* Studio: surface Anthropic document citations inline + in Sources panel

Anthropic's Messages API streams ``citations_delta`` events on
``content_block_delta`` when the request enables
``citations: {enabled: true}`` on document blocks. Each event carries
one citation pointing at the source document; previously they were
silently dropped, so reader-visible references never reached the chat
UI even when the model was citing properly.

The proxy now:
- dedupes by the type-specific anchor (char_location / page_location /
  content_block_location / search_result_location) so re-cites of the
  same span collapse onto a single footnote;
- injects ``[N]`` inline right after the matching text run;
- forwards the full list as a synthetic ``document_citations``
  tool_event at ``message_stop`` so the Sources panel can render
  per-document footnotes next to web_search / web_fetch citations.

Streams that never emit ``citations_delta`` stay byte-identical.

References:
- https://platform.claude.com/docs/en/build-with-claude/citations
- https://platform.claude.com/docs/en/build-with-claude/search-results

Tests (5 in test_anthropic_citations.py): passthrough, single
char_location, dedup of repeat citations, distinct sources get
distinct numbers, search_result_location supported.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: surface Anthropic document_citations in the Sources panel

The PR added a backend _toolEvent.type='document_citations' on
message_stop and an inline [N] marker in the assistant text, but the
chat-adapter only handles container_*/tool_*/sources from
web_search and web_fetch tool calls. Reviewers flagged that the
inline [N] markers had no matching footnote entries in the Sources
panel.

Capture the new event into a documentCitationParts buffer, convert
each citation dict into a Sources-panel source entry (using
document_title or search-result source URL plus cited_text as the
snippet), dedupe by id, and append to the final yield alongside
the existing web_search/web_fetch sourceParts.

* Studio: dedupe search_result_location citations by search_result_index

Anthropic's documented search_result_location citation shape carries
search_result_index, source, title, and start/end_block_index --
NOT document_index/document_title. The previous key keyed on
document_index + document_title + source + start_block_index, so
two distinct search results from the same source collapsed onto the
same footnote and the second [N] marker was lost.

Switch the search_result_location branch to key on the documented
fields, and pin the behaviour with a regression test asserting that
two citations sharing source/title but with different
search_result_index get distinct [1] [2] markers.

* Studio: keep each citation distinct across the end-anchor

Codex follow-ups on the citations PR:

  * Backend _anthropic_citation_key now includes the end anchor for
    every variant (end_char_index, end_page_number,
    end_block_index). Anthropic ranges are start-AND-end pairs, so
    a same-start / different-end pair is two distinct citations
    that previously collapsed onto one footnote.

  * Frontend documentCitationToSource ids include the position
    fields (search_result_index, start/end char/page/block) instead
    of being keyed on URL alone. Two citations from the same
    document or two search_result_locations with the same source
    now produce distinct Sources-panel entries, matching the
    inline [N] numbering.

* Studio: key Sources list by per-citation id instead of url

Codex flagged that the Sources renderer keys badges on source.url,
so two Anthropic document citations sharing the same source URL
collide as React keys and one badge gets dropped (or duplicated).

The chat-adapter already mints a per-citation id that folds the
position fields (search_result_index, start/end char/page/block)
into the URL, so the two citations have distinct ids even when
their URL matches. Plumb that id through SourceData and use it as
the React key for both the measurement badges and the visible
SourceBadge list. Falls back to the URL when no id is supplied
(web_search and web_fetch source parts).

* Studio: enable Anthropic doc citations on input_document blocks

Plumb citations: {enabled: true} onto the translated Anthropic document
block (both base64 and URL source branches) so the upstream actually
emits citations_delta events. Without this opt-in the inline [N] +
Sources panel plumbing added in this PR is a no-op for real user
PDF / doc uploads.

Refs https://platform.claude.com/docs/en/build-with-claude/citations

Also add edge-case coverage for the citations_delta path:
malformed citations, mixed types per document, reversed indices,
missing document_index, non-int block indices, unknown citation
type, internal _key never leaking, footnote numbering across
content blocks, and the input_document wire-through itself.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Reject unsafe citation sources, bound cited_text payload

Three follow-ups on top of #5718 surfaced by a deeper review pass:

1) javascript: / data: / vbscript: in citation source is XSS-able.
   ``documentCitationToSource`` was assigning ``cit.source`` straight
   into ``Source.url`` and rendering it as an <a href>. A hostile
   model emitting ``cit.source = "javascript:alert(document.domain)"``
   would execute on click (openLink only intercepts URLs that contain
   "://" or start with "mailto:", which both miss the javascript:
   scheme). Restrict the navigable path to http(s):// only; anything
   else falls back to the existing #anthropic-doc anchor and the
   source title still renders the raw identifier for context. Also
   reject CR/LF inside the URL string.

2) Frontend sources collapse distinct backend footnotes when the
   citation type differs but positions match. char_location(0,5) and
   page_location(0,5) over the same source previously deduped into
   one entry because the id only carried position. Fold citation
   type into the id anchor so the 1:1 mapping with inline [N]
   markers is preserved across every citation shape.

3) ``cited_text`` was forwarded unbounded inside the synthetic
   document_citations tool_event. The Sources panel trims to 240
   chars for display anyway; for large RAG / search_result spans
   (~10kB cited_text is plausible) this inflates SSE bytes 40x
   for no UI benefit. Truncate server-side at 512 chars with an
   ellipsis so the description-trim downstream still has room to
   work and the wire stays bounded.

Tests grow from 21 to 22; existing 7 + edge 15 still green. Frontend
typecheck clean.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: apply http(s) URL guard to all Sources-panel link sources

The previous round only filtered ``cit.source`` inside
``documentCitationToSource``. Two parallel code paths still copied
provider/tool-controlled ``URL:`` text directly into clickable
``<a href>`` Sources-panel links:

  * ``parseSourcesFromResult`` in chat-adapter.ts (legacy web_search /
    web_fetch tool result parser)
  * ``parseSearchResults`` in tool-ui-web-search.tsx (inline tool card)

A hostile tool response like ``URL: javascript:alert(1)`` or
``URL: data:text/html,...`` was therefore still rendered as a
navigable badge in the Sources panel.

Centralise the safe-URL test (``isSafeNavigableSourceUrl``,
``isSafeHttpUrl``) using ``new URL()`` + protocol allowlist + CR/LF
rejection, and apply it to both parsers. Unsafe blocks are dropped
rather than rewritten to a hash anchor because the web_search /
web_fetch parsers have no document-index fallback.

Citation conversion now uses the same helper so the in-place
http(s) regex and CR/LF check stay in one place.

* Shorten citation comments for PR #5718

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-25 23:39:02 -07:00
Daniel Han
7d3c472461
Studio: standalone Fetch pill for Anthropic web_fetch (#5742)
* Studio: surface Anthropic web_fetch as a standalone Fetch pill

web_fetch used to be silently bundled with the Search pill on the
assumption that "search returns URLs, fetch reads them" is the
typical workflow. Two problems with that:

- Anthropic bills each web_fetch invocation separately from
  web_search hits, so combining them made the per-message cost
  surface ambiguous.
- It blocked "just fetch this one URL" workflows where the user
  already knows the page they want read and does not want a search
  round-trip.

Adds:

- `webFetchToolsEnabled` to the chat-runtime-store, persisted to
  localStorage under `unsloth_chat_web_fetch_tools_enabled`, with a
  matching `supportsBuiltinWebFetch` capability flag and a
  `setWebFetchToolsEnabled` setter.
- A new Fetch pill in the chat composer, rendered next to Images and
  only when the active provider returns true from
  `providerSupportsBuiltinWebFetch` (Anthropic today). The pill
  defaults off so per-fetch billing is always a deliberate opt-in.
- chat-page bootstraps `webFetchToolsEnabled` from the same stored-
  preference fallback the other pills use.
- chat-adapter reads `webFetchToolsEnabled` directly when deciding
  whether to append "web_fetch" to `enabled_tools`, decoupling it
  from `toolsEnabled` (Search).

Backend translation is unchanged: when `enabled_tools` already
contains "web_fetch", `_stream_anthropic` appends the
`web_fetch_20250910` / `web_fetch_20260209` tool exactly as before
(test_anthropic_web_fetch.py pins the standalone-only path at
`test_web_fetch_tool_appended_to_request_body` and the combined
path at `test_web_fetch_combined_with_web_search_and_code_execution`).
Frontend tsc passes.

* ci: re-trigger after transient GitHub API HTTP flake (checkout + ggml-org release fetch)

* Studio: include web_fetch in the disabled-tool guard axis

Reviewer P1 / High on PR #5742 (codex + gemini): after introducing
the standalone Fetch pill, `disabledToolGuard` still only branched on
`webSearchEnabledForThisTurn`. With Fetch ON and Search OFF the
system prompt would tell Claude "you do not have web search or web
fetch tools in this conversation", which contradicts the actual tool
schema being sent and suppresses `web_fetch` tool calls, defeating
the standalone-fetch workflow this PR adds.

Treat search and fetch as a single "any web tool enabled" axis. The
guard only needs to warn the model when no web tool is wired in for
this turn; once either pill is on the model can pick the right one
from the tool schema. The existing `webLabel` already covers both
names, so the user-visible guard text stays accurate in every
combination.

tsc clean.

* ci: re-trigger after transient infra flake on Windows prebuilt / actions/checkout

* Studio: route web_fetch through per-model version dispatch

The web_fetch tool body in `_stream_anthropic` hardcoded
`web_fetch_20250910` instead of calling `_anthropic_web_fetch_version`,
so Opus 4.6 / 4.7 and Sonnet 4.6 missed the `web_fetch_20260209`
dynamic-filtering variant. The picker, the unit tests for it, and a
deliberate "follow-up" note in `test_anthropic_web_fetch.py` already
existed; this just threads it through the emission site.

Mirrors how web_search and code_execution are dispatched per model.
Old models still resolve to `web_fetch_20250910` and continue to work.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Shorten web_fetch comments for PR #5742

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-25 23:37:26 -07:00
Daniel Han
063e1e497b
Studio: rewrite OpenAI Responses citation markers to markdown links (#5713)
* Studio: rewrite OpenAI Responses citation markers to markdown links

OpenAI's /v1/responses stream interleaves text deltas with inline
citation markers built from private-use codepoints (U+E200 / U+E201 /
U+E202) shaped like `citeSOURCE_ID`. The codepoints render
as garbled "E202" glyphs or empty boxes in most fonts, and the
markdown layer further strips them, leaving run-on text like
"citeturn1view0turn1view1turn3view0...". The url list still arrived in
the Sources panel via url_citation annotations, but the inline cite
hand-off into the prose was unreadable.

Rewrite each marker into `[N](URL)` when the matching url_citation
has already been recorded on this stream, and drop the marker
silently otherwise. The lookup uses a new `source_id` field captured
on `_record_url_citation` (accepts source_id / id / locator across
Responses API revisions). Annotations are now applied BEFORE the
delta text is rewritten so that markers and their resolving
annotation arriving in the same SSE event still resolve.

Reference: https://developers.openai.com/api/docs/guides/citation-formatting

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Preserve every source_id alias for a deduplicated url_citation

OpenAI's Responses stream cites the same URL under multiple
source_id markers when the model references different spans of the
same page. The previous dedup-by-URL kept only the first alias and
dropped the rest, so subsequent markers for the same URL never
resolved and got stripped from the prose. Switch the citation
record to a ``source_ids`` list and append new aliases on every
duplicate. The rewriter resolves any alias back to the same
citation number so the inline markers all collapse onto one footnote
rather than fanning out into bogus repeats.

Also collapse the two passes over ``all_url_citations`` in
``_record_url_citation`` into a single loop for clarity. Adds two
regression tests covering the alias-collision and mixed-shape cases.

* ci: re-trigger after flake in Studio GGUF Tool calling (rebased on main #5741 already)

* ci: re-run after transient CodeQL Python checkout auth flake

* Fix split-marker buffer + multi-source ids for PR #5713

The original rewriter only handles markers that arrive whole inside a
single response.output_text.delta event. OpenAI's stream chunks text
on byte-buffer boundaries with no awareness of the marker grammar,
so a marker can straddle two deltas (delta-1 ends with
"citetu", delta-2 starts with "rn0view0"). Each delta
was rewritten in isolation, so the half-marker leaked as garbled
"E200/E202" glyphs in the rendered prose.

Buffer the unterminated tail across deltas and concatenate it onto
the front of the next one so the rewriter sees a complete marker.
Flush the held-over tail on response.completed / response.incomplete /
[DONE], stripping any leftover private-use bytes so a never-closed
marker (truncated stream, missing annotation) never leaks.

Also handle the multi-source marker shape from the OpenAI docs --
citeid1id2 should expand to one bracket
link per resolvable id. The previous regex captured only the first
source id and silently dropped id2/id3.

Reference: https://developers.openai.com/api/docs/guides/citation-formatting

Tests: 21 new cases covering multi-source, locator suffix, marker
split across two and three deltas, unterminated marker on truncation,
late annotation resolving a buffered marker, idempotency, and the
head/tail split helper directly.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Defer citation segments until url_citation annotation arrives

The split-marker buffer already concatenates a marker that straddles
two response.output_text.delta events. But when the annotation event
for a url_citation arrives AFTER the delta that contains its inline
marker (the typical OpenAI Responses ordering), the rewriter still
saw an empty lookup table at delta time and silently stripped the
marker. The URL kept showing up in the sources panel but the inline
link reference was permanently gone.

Add _rewrite_citation_markers_partial which leaves an unresolved
marker verbatim and reports has_unresolved=True. The streaming loop
buffers any closed segment that contains an unresolved marker into a
pending_citation_segments FIFO and drains the queue on every later
annotation event, on response.completed, on response.incomplete, and
on the [DONE] sentinel. Drain order is preserved so later clean text
does not leapfrog an earlier deferred segment. End-of-stream forces a
strip so no codepoint leaks if the annotation never arrived.

Add six regression tests covering single-pass resolution, the late-
annotation two-pass case, multi-source markers with partial
resolution, mixed known and pending markers in one segment, and
idempotency on marker-free input.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Drop unterminated citation tail to prevent cite-prefix plain-text leak

`_flush_pending_marker_tail` stripped the three private-use citation
codepoints from the held-over buffer, but left the literal ``cite``
keyword plus the source id behind as plain text. A stream ending
mid-marker therefore emitted user-visible garbage like
``Some text citeturn0view0`` instead of the intended clean prose.

``pending_marker_tail`` is by construction the suffix that starts at
an unclosed ``\\ue200`` opener -- the split helper guarantees there is
no closing ``\\ue201`` byte. Without that close the marker is
meaningless: the source id cannot be resolved to a URL and the user
prose before the opener was already emitted as ``head`` on the
originating delta. Bail out before the strip step and return the
empty string. As a belt-and-braces measure also drop any orphan
``cite<sid>`` literal at the head of the buffer in case a future
caller passes a partially-terminated tail.

Update the matching ``_simulate_delta_stream`` harness in the edge
tests so it mirrors the new flush logic, and add four regression
tests covering unterminated marker with surrounding prose, marker-
only inputs, prefix-only outputs, and the split-then-close path that
still must resolve to a link.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Defer multi-source markers until all ids resolve for PR #5713

`_rewrite_citation_markers_partial` previously treated a marker as
resolved when even one token in a multi-source marker resolved,
dropping any still-pending source ids. In streamed Responses events
the annotations for a multi-source marker can arrive across separate
`annotation.added` chunks, so the caller no longer buffered that
segment for retry and the late source id was lost from the inline
citation entirely.

Flag the marker unresolved whenever any token misses the lookup so
the streamer keeps the segment pending. End-of-stream force flush
still drops unresolved tokens through `_replace_openai_citation_markers`
so locator-style suffixes (which look like unresolved ids at the token
level but only appear at end-of-stream) render cleanly.

Updated the multi-source test to assert the new pending-then-flush
behavior; locator output now lands at force-flush rather than mid
stream.

* Shorten citation marker comments for PR #5713

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-25 23:37:16 -07:00
Daniel Han
7d1b68079e
Studio: Anthropic fast_mode toggle and streaming refusal handling (#5715)
* Studio: add Anthropic fast_mode toggle + surface streaming refusals

Fast mode (beta `fast-mode-2026-02-01`) lets Claude Opus 4.6 and 4.7
generate output tokens up to 2.5x faster at 6x standard Opus
pricing. The toggle lives in Configuration → Provider when the
selected Anthropic model is Opus 4.6 or 4.7 and is otherwise
hidden. Backend gates the same prefixes a second time so a stale
frontend cannot make Anthropic 400 the request, and the
`fast-mode-2026-02-01` beta header is merged onto whatever other
betas the request already needed (code-execution, compaction).

Streaming refusals (`message_delta.delta.stop_reason="refusal"` on
Claude 4 models) now surface a short user-facing notice in the
assistant message before the translated OpenAI chunk emits the
existing `finish_reason="content_filter"`. Previously the chat
bubble truncated silently because the SSE stopped mid-stream with
no visible explanation. Per the upstream docs the conversation
must be reset before continuing, so the notice tells the user
exactly that.

Reference:
- https://platform.claude.com/docs/en/build-with-claude/fast-mode
- https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/handle-streaming-refusals

Tests:
- studio/backend/tests/test_anthropic_fast_mode_and_refusal.py (8 cases
  pinning fast_mode pass-through on 4.6/4.7, silent drop on Sonnet /
  Haiku / older Opus / None / False, and the refusal notice + finish
  reason on a synthetic refusal stream).

* Studio: drop refused Anthropic turns from the next request

Anthropic's streaming-refusal guidance says the refused assistant
turn must be removed or updated before the next call -- otherwise
the safety classifier keeps refusing. The PR only added a
user-visible notice; the partial assistant output (plus the notice
itself) still rode the next request via toOpenAIMessage.

Tag the refusal turn with an HTML-comment sentinel emitted alongside
the notice. The chat-adapter checks for that sentinel in
toOpenAIMessage and returns null, so the refused turn is excluded
from outboundMessages. The notice still renders in the transcript
(HTML comments don't display), so users keep the explanation.

* Studio: filter None finish_reason entries in test helper

test_refusal_maps_to_content_filter expects only ['content_filter']
in the finish_reasons list, but the post-PR refusal path emits a
user-visible content notice chunk first. Every _content_chunk
carries 'finish_reason: None' by construction; the helper was
appending those, so the assertion saw [None, 'content_filter']
instead of ['content_filter'].

None is not a finish reason -- it's just mid-stream delta noise.
Skip None values in _finish_reasons so the helper reflects what
the test names actually claim to check. Same fix applies cleanly
to the other helper usages (pause_turn test expects [] and the
sibling stop test expects ['stop'], both unaffected).

* Studio: cover Anthropic fast-mode edge cases

Adds 19 cases on top of the 9 in test_anthropic_fast_mode_and_refusal.
The base file pins the happy path; this file fills in the cliffs:

* Dated-snapshot prefix matching: claude-opus-4-7-2026-02-01 and
  claude-opus-4-6-2026-02-01 still gate fast_mode through, while
  claude-opus-4-5-2025-08-01 and claude-sonnet-4-6-2026-02-01 do not.
* Strict opt-in: a future claude-opus-4-8 or claude-opus-5 does NOT
  auto-enable fast_mode -- the prefix tuple must be bumped explicitly
  when a new family is whitelisted upstream.
* Beta-header merge: fast_mode coexists with code-execution-2025-08-25
  and compact-2026-01-12 in one comma-separated anthropic-beta header
  with no duplicates and no truncation. Pins the value to the exact
  fast-mode-2026-02-01 docs token so a typo would fail CI.
* Non-destruction: fast_mode=None produces byte-identical outbound
  body and headers to the version that omits the argument entirely.
  Same for fast_mode=False. Guarantees the upgrade path is
  non-breaking on existing Anthropic streams.
* Refusal stream ordering: the user-visible notice precedes the
  finish_reason chunk so a streaming UI paints text before flipping
  to content_filter. Refusal sentinel emitted exactly once. Notice
  rides a normal content delta chunk with finish_reason still null.
  Partial assistant deltas survive before the notice.
* Provider-side refusal coverage: a refusal on Sonnet (not just Opus)
  still emits the notice + sentinel + content_filter mapping, since
  refusal handling is not gated on fast-mode capability.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Persist fastMode, drop refused user message on retry

Two follow-ups on #5715:

1) sanitizeInferenceParams stripped fastMode. fastMode is in
   PERSISTED_INFERENCE_PARAM_KEYS but the storage sanitizer only kept
   numeric fields plus systemPrompt and trustRemoteCode, so the new
   toggle was silently dropped on reload and on the
   /api/chat/settings round-trip. Save it the same way trustRemoteCode
   is saved.

2) Refusal recovery now also drops the triggering user turn.
   Returning null from toOpenAIMessage on the assistant side left the
   user prompt that caused the refusal in the outbound history, so
   the very next request would re-trigger the same classifier.
   Anthropic's refusal-handling guidance is explicit on this: remove
   the refused turn AND the user message that triggered it before
   the next call. Implemented via a pre-pass that pops the trailing
   user message when an assistant carries the refusal sentinel.

Typecheck clean.

* Studio: out-of-band refusal signal + fast-mode prefix/usage/pricing fixes

The text sentinel for the Anthropic refusal drop signal was spoofable:
any assistant message containing the literal
<!--studio:anthropic-refusal--> would prune the prior user + assistant
pair on the next request. Move the signal onto a separate _toolEvent
chunk that the chat adapter latches into
assistant.metadata.custom.anthropicRefusal; assistant text can no
longer control the pruner.

Tighten the fast-mode model gate (backend + frontend) to require a "-"
family boundary so claude-opus-4-70 / claude-opus-4-7b style IDs do
not get speed: "fast" on a naive startswith match.

Use survivingMessages for the image / audio attachment scan so a
refused user turn does not gate or mis-attribute the next non-refused
turn.

Propagate Anthropic usage.speed onto the OpenAI-style usage chunk and
apply the documented 6x fast-mode multiplier in the cost calculator
(stacks with prompt-cache multipliers per the docs); expose the new
multiplier on the pricing snapshot for the UI tooltip.

Tests cover the tool-event chunk shape, the prefix-collision rejects,
usage.speed propagation, the 6x pricing math, and that the visible
refusal text carries no embedded sentinel.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Shorten fast-mode and refusal comments for PR #5715

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-25 23:37:12 -07:00
Daniel Han
b17765b5ed Studio: round 8 -- replay guard on non-visible events + Add-flow Codex preselect
Two P1 fixes from the round 8 reviewer pass:

1. _stream_thread_run no longer replays a Codex turn that fired
   non-visible events before crashing.

   The replay guard only tracked `emitted_any` (visible text). A
   Codex turn that emitted, say, a command.delta or file.delta event
   first -- both filtered to "" by _coerce_text -- and THEN crashed
   would leave emitted_any=False and fall through to the buffered
   `thread.run(prompt)` fallback, re-executing the same turn and
   duplicating its side effects (shell commands, file writes,
   tool calls). This is exactly the case the guard was added to
   prevent in earlier rounds; the missing bit was tracking
   "the turn ran at all", not just "the turn yielded text".

   Fix: add a separate turn_started flag that flips True the moment
   we ask the SDK for a turn handle or observe any event from a
   streaming helper. When the buffered fallback is gated on
   turn_started instead of emitted_any, a partial-turn crash
   correctly stops without replaying. Regression test reproduces
   the bug against the pre-fix code (assertion catches the extra
   thread.run call) and locks the fix in.

2. openAddProvider now mirrors the providerType-change effect's
   Codex pre-check.

   The first-run UX fix from `26799d9a` pre-checked every Codex
   default model in the providerType-change effect, but
   openAddProvider() calls resetForm() (which clears
   selectedModelIds) and then only restores availableModels, not
   selectedModelIds. If the user closes the Add connection form
   and re-opens it while Codex is still the current providerType,
   the effect does not re-run, so the form opens with Codex
   defaults available but none selected -- the "Add at least one
   model ID" save guard then blocks the Save click.

   Fix: openAddProvider now seeds selectedModelIds with the full
   default-models list when the provider is Codex, matching the
   providerType-change effect so the two entry paths produce the
   same first-run state.
2026-05-25 15:24:49 +00:00
Daniel Han
dee1b68b6d Studio: round 7b -- tighten device-login log filter + harden timeout kill
Two more P1 follow-ups from the round 7 reviewer pass:

1. Device-login log filter no longer leaks sensitive lines.

   The old `_safe_to_forward` used unanchored substring matches like
   `"logged in"`, so a line such as

     Not logged in: refresh_token=rt_LEAK auth.json=/home/u/.codex/auth.json

   slipped through the safe-vocabulary filter and was streamed to the
   browser. A malicious codex shim earlier on PATH can print that
   line trivially, defeating the "opaque output stays in backend
   logs" safety guarantee the route claimed.

   Round 7b fix: anchored regex set (must start with one of the
   known upstream phrases) plus an explicit blocklist for
   refresh_token / access_token / api_key / secret / auth.json / the
   codex config dir / "not logged in" / "not authenticated". A line
   that matches the blocklist is dropped regardless of which safe
   pattern would otherwise have accepted it. New tests reconstruct
   the regex set inline and assert both the leak cases drop and the
   clean upstream phrases pass.

2. `_run_cli` timeout cleanup no longer 500s on a kill race.

   `_run_cli` would call `proc.kill()` after `os.killpg(pid, SIGTERM)`
   reaped the process group. If the SIGTERM landed first, the
   subsequent `proc.kill()` raised `ProcessLookupError` and bubbled
   out of `_run_cli`, turning `/api/codex/status` into a 500 during
   a timeout race. The device-login cleanup at codex_provider.py
   already wraps the same destructive call in a try / except. Mirror
   that exception guard here so the two timeout paths behave the
   same.
2026-05-25 14:45:26 +00:00
pre-commit-ci[bot]
8cef479423 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-25 14:43:25 +00:00
Daniel Han
f05517ac79 Studio: round 7 Codex hardening (cross-wrapper env scrub + device URL allowlist)
Two P1 fixes from the round 7 reviewer pass:

1. _ScrubbedEnvAsyncCodex no longer leaks secrets across overlapping
   sessions.

   The fallback env-scrub wrapper (used when the installed SDK build
   does not accept AppServerConfig(env=...)) refcounts deleted env
   vars under a shared lock so concurrent fan-out workers do not
   restore Studio's secrets while a peer is still inside SDK startup.
   The round 6 implementation enumerated keys to scrub via
   _codex_sdk_env_override() which only returns keys currently in
   os.environ. If wrapper A had already deleted HF_TOKEN before
   wrapper B entered, B's overrides dict no longer contained HF_TOKEN,
   B never bumped the refcount for it, and A's exit restored
   HF_TOKEN into the process env while B was still running -- so
   B's spawned codex app-server inherited the secret.

   Round 7 fix: under the lock, the union of (a) the current
   overrides dict and (b) every key still refcounted by an earlier
   wrapper is the set of keys this session must scrub. Originals are
   tracked module-level rather than per-instance so the last wrapper
   to release a key always restores the right pre-scrub value
   regardless of who first saw it. New regression test reproduces
   the leak against the round 6 code (asserts refcount == 2 after
   B enters; old code records 1) and locks the fix in.

2. Device-auth verification URL is now host-allowlisted.

   The codex login --device-auth output parser pulled any https URL
   matching /device, /activate, or /verify out of the CLI's stdout
   and emitted it as a device_url event. The frontend rendered that
   URL as an "Open verification page" CTA the user can click. A
   compromised codex shim earlier on PATH could print
   https://evil.example/activate?code=ABCD and Studio would surface
   the phishing link verbatim, even though every other login output
   line goes through a strict safe-vocabulary filter.

   Round 7 fix: device_url events only fire for URLs whose host is
   on a small allowlist (auth.openai.com / chatgpt.com over https).
   Anything else is logged at warn and dropped. Tests cover the
   known-good upstream URLs, several attacker patterns (lookalike
   subdomains, http downgrade, javascript:), and garbage input.
2026-05-25 14:43:10 +00:00
Daniel Han
2eaf1bbd31 Studio: pass codex_bin to AppServerConfig so PATH-only codex installs work
Reproduces with pip install openai-codex --no-deps (lightweight install
that skips the pinned openai-codex-cli-bin runtime) or any host where
the codex CLI is installed via npm i -g @openai/codex / Homebrew /
manual download. Studio constructs AsyncCodex(config=AppServerConfig(
env=...)) without codex_bin, so the SDK runs _installed_codex_path
which 'from codex_cli_bin import bundled_codex_path' and raises
FileNotFoundError: Unable to locate the pinned Codex runtime. Install
the published SDK build with its openai-codex-cli-bin dependency, or
set AppServerConfig.codex_bin explicitly. -- even though a perfectly
good codex is on PATH and was the binary the availability probe
already verified.

Fix: resolve shutil.which("codex") and pass it as
AppServerConfig(codex_bin=...). The PR's availability probe already
returns that exact path in /api/codex/status.cli_path, so we are
giving the SDK back the binary the user can see in the Connections
form. Falls back to AppServerConfig(env=...) (no codex_bin) when the
SDK build does not accept the kwarg yet, and falls back to PATH lookup
returning None on hosts without codex on PATH (in which case the
availability probe would have reported installed=false and Studio
never gets here).

Test fixture: also inject the fake module under openai_codex (the
canonical name the production importer prefers) so the test does not
silently exercise the real SDK on developer venvs that have
pip install openai-codex already done.

End-to-end verified live: pip install -e openai/codex sdk/python plus
Studio with codex CLI on PATH yielded ROUND_TRIP_OK streaming for
gpt-5.4-mini through the OpenAI-compat completions route.
2026-05-25 13:47:16 +00:00
Long Yixing
748aa1c482
fix: repair mlx studio base export save_method (#5727) 2026-05-25 04:04:07 -07:00
Daniel Han
2874abbfec Studio: fail closed when Codex SDK cannot enforce safety pins
Round 6 reviewer noted that the warn-and-proceed path in
`_start_thread_with_system` is "failing open" on a server-side
chat surface: an SDK rev that does not expose ApprovalMode or
SandboxMode would log a warning then call `thread_start(model=...)`
with NO safety kwargs, letting the model run under the SDK's
`auto_review` default. For a route that takes a user-controlled
prompt and can spawn shell commands or file writes, that is the
wrong tradeoff.

Now fails closed: when `_safe_thread_safety_kwargs()` returns the
empty dict the helper raises `CodexUnavailableError`, which the
route layer translates to a 503 with a clear error message telling
the operator to upgrade `openai_codex` (or set the explicit
override env var). The error message names the override so users
who hit this on a pre-release alpha can opt in with eyes open
rather than discovering the unsafe default after the fact.

`UNSLOTH_CODEX_ALLOW_UNSAFE_DEFAULTS=1` is the deliberately-verbose
escape hatch. Variable name long and explicit so it does not creep
into production environments by accident, kept on the codex
subprocess safe-list so the round 6 SDK env-scrub wrapper does not
delete it before the gate sees it.

Tests: 50 cases total (was 49). The previous old-SDK test was
renamed and replaced by two new ones:
- `test_thread_start_fails_closed_when_safety_unavailable` asserts
  the raise fires and `thread_start` is never called.
- `test_thread_start_allows_unsafe_defaults_with_explicit_opt_in`
  asserts the override env var lets the request through and
  `thread_start` runs without the safety kwargs (with a logged
  warning).
The `_install_fake_codex_sdk` helper now injects fake ApprovalMode
and SandboxMode by default so the general translation tests do
not need to opt into the override; the two round-6b tests above
pass `with_safety_enums=False` to exercise the fail-closed branch.
2026-05-24 17:45:20 +00:00
Daniel Han
8ee60019a4 Studio: round 6 Codex hardening (5 follow-ups)
Reviewer round 6 surfaced five real follow-ups on top of rounds 5
through 5g. Each one is a fix for an asymmetric guard or a wrong-
shape lookup in the new Codex provider code:

1. Re-gate `installed=True` on having BOTH the SDK AND a `codex`
   binary on PATH. Round 5 widened the gate to SDK-only, but the
   login route still shells out to the binary, so an SDK-only host
   would surface a Codex row whose Sign-in button immediately
   failed with "codex CLI not found on PATH". The canonical
   `openai-codex` package depends on `openai-codex-cli-bin` which
   places the shim on PATH for free, so the common install still
   lights up; the gate just refuses to advertise a provider Studio
   cannot actually drive end-to-end.

2. `_safe_thread_safety_kwargs` now also probes `.api` and
   `.generated.v2_all` for `SandboxMode`. Upstream `openai_codex`
   exports `ApprovalMode` at the top level but `SandboxMode` lives
   under `openai_codex.generated.v2_all`. The previous lookup
   returned `{}` on the canonical SDK install, so every thread_start
   ran with the unsafe `auto_review` default. Submodule probe
   resolves the canonical layout and keeps backwards-compat with
   builds that DID re-export at the top level.

3. `_coerce_text` now applies the answer-event-type filter on the
   object path too. The upstream SDK emits typed payload classes
   like `CommandExecutionOutputDelta`, `FileChangeDelta`,
   `ToolCallDelta`, `PatchApplyDelta`, etc., all of which carry a
   `.delta` string of local stdout / file paths / tool args. The
   dict path already filtered these out; the object path used to
   return `.delta` unconditionally, so a real SDK install could
   leak tool output into the visible chat reply.

4. `_ScrubbedEnvAsyncCodex` is now process-wide concurrency-safe
   AND fails-closed if the SDK constructor raises:
   - Refcount each scrubbed key under an `asyncio.Lock` so a fan-out
     wrapper that exits early cannot restore a secret while another
     wrapper is still inside SDK startup (round 6 reproduced this:
     wrapper A exited, wrapper B's SDK saw the restored HF_TOKEN).
   - Move `_async_codex_cls()` and its `__aenter__` INSIDE a
     try/except in `__aenter__`; on failure, run the release path
     so the scrubbed env vars are restored even though `__aexit__`
     never fires for the failed construction.

5. `_run_cli` now detaches into its own process group via
   `start_new_session=True` (Unix) / `CREATE_NEW_PROCESS_GROUP`
   (Windows) and kills the whole group on timeout, matching the
   protected path in `stream_codex_device_login`. A shimmed
   `codex login status` that forks a helper and blocks no longer
   leaves the child running after we killed the parent.

Frontend follow-up: chat-adapter now routes every rendered yield
through a `renderFullContent()` helper so the Codex per-tab text
accumulated in earlier `_toolEvent` frames is preserved when the
synthesis content delta arrives. Previously the next regular
content yield rebuilt `parts` from `cumulativeText` alone and the
tab section vanished from the final assistant message.

Tests: 49 cases total (was 47). New regressions:
- `installed_requires_both_cli_and_sdk` (round 6 revert).
- `safety_kwargs_finds_sandbox_mode_in_submodule` (canonical SDK
  layout where `SandboxMode` is in `.generated.v2_all`).
- `scrubbed_env_construction_failure_restores_env` (no permanent
  env leak when the SDK constructor raises).
- Expanded `coerce_text_drops_non_answer_event_types` to also
  exercise the object-shape code path with `CommandExecutionOutputDelta`,
  `FileChangeDelta`, `ToolCallDelta`, `PatchApplyDelta`,
  `PlanUpdateDelta`, `AgentReasoningDelta`, plus the
  positive `AgentMessageDelta` allow-through.
2026-05-24 17:30:01 +00:00
Daniel Han
aa258b983d Studio: account for every Codex turn in fan-out usage chunk
The fan-out path used to report usage as if a single Codex call had
run: `prompt_tokens = max(1, len(prompt)//4)` and
`completion_tokens = len(synthesis)//4`. In reality it had spawned
N parallel worker turns (each carrying the same prompt) plus a
synthesis turn that re-sent the prompt and every tab's output. For
`parallel_calls=20` that meant the cost / context widget under-
reported the request by roughly 20x.

Now sums:

- `prompt_tokens ≈ (N * prompt + synthesis_prompt) / 4` where
  `synthesis_prompt = sum(tab_outputs) + prompt`.
- `completion_tokens ≈ (sum(tab_output_chars) + synthesis_chars) / 4`.

Tests: new regression
`test_parallel_usage_accounts_for_all_calls` runs a 4-way fan-out
against a fake SDK with deterministic chunk lengths and asserts
the reported usage scales with N, not the single-call shape.
2026-05-24 17:11:42 +00:00
Daniel Han
f97a800d5b Studio: round 5e Codex hardening (4 follow-ups)
1. parallel_calls validator now clamps instead of 422-rejecting.
   The Pydantic schema was `int Field(ge=1, le=20)`, which was a
   regression from the pre-PR OpenAI-extra behaviour: a non-Codex
   client that sent the field with a legacy value like 0 (or a
   stray string from a misconfigured wrapper) now got a 422 even
   though the route silently ignores the field on every non-Codex
   provider. Replaced with a `field_validator(mode="before")` that
   coerces any input to the [1, 20] range, keeping the schema docs
   self-documenting while accepting legacy inputs.

2. Buffered Codex result with `final_response=None` no longer
   leaks `TurnResult(...)` Python object repr into the chat. The
   upstream SDK documents `TurnResult.final_response` as nullable
   for turns that perform tool work without producing a final
   assistant message; the previous `... or str(result)` fallback
   would render the repr as visible assistant text. New
   `_buffered_result_text` helper returns the empty string in that
   case so the stream finishes cleanly with no extra content
   chunk. Same fix applied to `_run_codex_synthesis`.

3. Device-login SSE no longer forwards arbitrary subprocess output
   to the browser. The previous code yielded every CLI line under
   `{type:"log"}`, which on a shimmed binary could leak refresh
   tokens, auth JSON, or local config paths into the authenticated
   stream. Filtered to a known-safe vocabulary ("Welcome to
   Codex", "Initializing", "Successfully logged in", etc.).
   `device_url` and `device_code` events still fire as before.

4. CodexLoginButton no longer calls `window.open` from inside an
   awaited SSE handler. Browser popup blockers (Firefox, Safari,
   Chrome strict) silently block popups triggered outside a fresh
   user gesture, so the auto-open was unreliable. Replaced with a
   prominent "Open verification page" button styled as an anchor;
   the click handler is a real user gesture and is never blocked.
   The URL string is still shown below the button for copy/paste.

Tests: 46 cases total (was 43). New regressions cover the
parallel_calls clamp path on three garbage inputs, the buffered
TurnResult-with-None-final repr leak guard, and the device-login
log filter (asserts refresh tokens / auth.json paths are dropped
while known-safe progress lines pass through).
2026-05-24 16:59:46 +00:00
Daniel Han
8c1c63a64d Studio: surface Codex final agent message when no deltas stream
The canonical openai_codex SDK can complete a turn successfully
without emitting any `message.delta` events: the final assistant
text arrives only as an `ItemCompletedNotification` whose item is
an `agentMessage`. Before this change `_stream_thread_run` would
loop through the stream, see no delta text, return, and Studio
would emit only the empty usage + stop + `[DONE]` frames -- the
user sees a blank reply for what was actually a complete answer.

Track agent-message texts collected during the stream loop and, if
no streamed deltas came through, yield the last one before
returning. The buffered `thread.run()` fallback is still gated by
the existing `emitted_any` flag so it never replays a turn that
already executed side effects (file writes, shell commands).

`_completed_agent_message_text` accepts both the upstream object
shape (`ItemCompletedNotification(item.root.text=...)`) and the
dict shape pre-release builds and tests use, so it works across SDK
revs without an explicit version gate.

Tests: new regression
`test_empty_stream_falls_back_to_completed_agent_message` exercises
a fake SDK whose `turn().stream()` yields only an `item.completed`
event with an `agentMessage`; the test asserts the final text
reaches the visible chat output and that the buffered `run()` path
is NOT re-executed.
2026-05-24 16:44:22 +00:00
Daniel Han
c5289f0249 Studio: pin Codex thread approvals + sandbox to safe defaults
The upstream openai_codex SDK defaults `approval_mode` to
`ApprovalMode.auto_review` (described in the SDK docs as "automatically
execute tools when permission escalations occur, without user
intervention") and leaves `sandbox` unset. Studio drives Codex from
a server-side chat request with no per-action approval UI, so leaving
those at the SDK defaults would let a model decide on its own to run
shell commands, write files, or hit the network on the operator's
machine.

This wires every `thread_start` call (single-turn, parallel-worker,
synthesis) through a helper that pins:

- `approval_mode = ApprovalMode.deny_all` -- reject any tool /
  command escalation rather than auto-approving it.
- `sandbox = SandboxMode.read_only` -- the policy that bans file
  writes and disables network.

The kwargs are looked up dynamically: when the installed SDK is too
old to expose either enum we log a structured warning and proceed
without them rather than refusing to run, so users on pre-release
alpha builds are not bricked. Once the canonical openai-codex SDK
is what every install pulls, the warning will be silent and the
safety pins will always apply.

Tests: three new regressions in TestCodexHardenedRegressions cover
the safe-pin path on a fake SDK that exposes the enums, the
warn-and-proceed path on a fake SDK that does not, and the same
pins on the synthesis turn so a fan-out tab cannot sneak an unsafe
default into the unification step.
2026-05-24 16:40:48 +00:00
Daniel Han
1a107651c4 Studio: round 5 hardening for the Codex provider
Five tightening fixes driven by the reviewer pass on top of round 4:

1. Drop OPENAI_API_KEY from the codex subprocess safe-list.
   The OpenAI provider key belongs to the OpenAI provider; a shimmed
   `codex` binary on PATH must not receive it. Users who want to wire
   the same key into Codex now set CODEX_OPENAI_API_KEY, which is the
   one OpenAI-shaped key we still forward.

2. Fail-closed env scrub on the SDK path.
   When AppServerConfig is missing from the installed openai_codex
   build, the bare AsyncCodex() constructor used to inherit the full
   os.environ via the SDK's internal os.environ.copy(). Replaced the
   fallback with a _ScrubbedEnvAsyncCodex wrapper that swaps
   os.environ for the lifetime of the session so HF_TOKEN, GH_TOKEN,
   WANDB_API_KEY etc never reach the spawned app-server.

3. Prefer the upstream-canonical base_instructions kwarg.
   The real openai_codex SDK takes the system prompt as
   `base_instructions`; our previous helper only knew `system`. Now
   tries base_instructions first, falls back to system, then inlines
   the system text in the user prompt as a last resort.

4. Filter visible text to answer-bearing event types only.
   _coerce_text used to render any payload that exposed a `delta` /
   `text` / `content` field, which let command output, file paths and
   tool call arguments leak into the Chat Completions reply. Gated on
   a _ANSWER_EVENT_TYPES allow-list (message.delta, completed,
   text_delta, etc.); untyped legacy dicts still pass through.

5. Treat the SDK as the install gate.
   openai-codex-cli-bin ships the codex runtime that backs
   AsyncCodex(...), so SDK alone is sufficient to drive the provider.
   `installed` no longer also requires a standalone codex binary on
   PATH; cli_path remains reported separately so the UI can still
   show whether a CLI is also installed.

Also added the matching positive-match regex line ("Authenticated:
Yes") for one more login-status wording the CLI ships in some
locales.

Tests: grown to 39 cases. New regressions cover the SDK-only install
gate, base_instructions kwarg priority + system fallback, the
fail-closed env scrub wrapper, the answer-only delta filter, the
Authenticated: Yes wording, and the CODEX_OPENAI_API_KEY-vs-
OPENAI_API_KEY split.
2026-05-24 16:20:27 +00:00
pre-commit-ci[bot]
593fc9edba [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 15:44:59 +00:00
Daniel Han
67b837e9f1 Merge branch 'feat/codex-provider' of https://github.com/unslothai/unsloth into feat/codex-provider 2026-05-24 15:44:11 +00:00
Daniel Han
4be807bbd8 Studio: round 4 hardening for the Codex provider
Fourth reviewer.py pass surfaced one more security finding and a
handful of correctness gaps. Each is small but the env-scrub for the
SDK path closes the asymmetric-fix loop opened in the previous round.

* Codex SDK construction now passes an `AppServerConfig(env=...)`
  that overrides every non-safe-listed env key to an empty string.
  Upstream openai/codex/sdk/python/client.py builds the spawn env as
  `os.environ.copy()` then `env.update(self.config.env)`, so this
  scrubs HF_TOKEN / GH_TOKEN / WANDB_API_KEY / ANTHROPIC_API_KEY etc.
  out of the codex app-server subprocess env on the chat / parallel
  / synthesis paths, matching the CLI/login paths from the previous
  round. The helper falls back to bare `AsyncCodex()` when the SDK
  version does not expose AppServerConfig, with logged warning.

* Install hint now names the actual upstream PyPI project,
  `openai-codex` (canonical), with `codex_app_server` documented as
  the legacy alias. The probe still accepts both import names so
  forward compat is preserved.

* Device-auth URL regex broadened to accept upstream's current
  `chatgpt.com/activate` shape and any `/device|/activate|/verify`
  variant, not just `/codex/device`. The frontend can now open the
  verification page on CLI builds that print the documented
  ChatGPT-style URL.

* `_run_codex_synthesis` now takes a `system` arg and forwards it
  to `thread_start(system=...)`, falling back to a prompt-prefix on
  older SDK revs that reject the kwarg. Previously a fan-out with
  "Always answer in Spanish" produced Spanish per-tab attempts but
  an English synthesis.

* `_detect_logged_in` negative regex now also matches "Not signed
  in", "Please sign in" (alternative localisations / future CLI
  releases). Same word-boundary anchoring as before.

* Frontend `CodexLoginEvent` union gains `device_code` and a `code`
  field. `CodexLoginButton` now renders the one-time code under the
  verification URL so users on a headless / remote install can copy
  the code without scraping the log pane. Also fixes a closure-stale
  bug where setError(message) was followed by a stale `error` read,
  losing specific backend errors; the new path keeps `lastStreamError`
  inside the closure.

* Replaced four hardcoded `/mnt/disks/...` paths in the new
  regression tests with `_backend_file()` resolved from `__file__`,
  so the suite runs in any checkout (CI, local dev, the review
  worker tree). Found by the round-4 reviewer.

Four new pytest cases pin the behaviour:
`test_not_signed_in_wording_also_handled`,
`test_device_url_accepts_generic_verification_url`,
`test_synthesis_call_forwards_system_prompt`, and
`test_sdk_env_scrubbed_via_appserverconfig`. 32/32 codex_provider
tests pass; `tsc --noEmit` clean.
2026-05-24 15:44:11 +00:00
pre-commit-ci[bot]
4b4c8553f8 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 15:16:52 +00:00
Daniel Han
fd8f25f507 Studio: scrub Codex subprocess env, guard partial-stream replay, abort login on unmount
Third reviewer.py pass found three remaining sharp edges. Each fix
is small and paired with a regression test where applicable.

* Codex subprocess env is now scrubbed to a safe-list before spawn.
  Both `_run_cli` in codex_availability and the device-auth spawn
  in stream_codex_device_login switch from `env=os.environ.copy()`
  to `env=_codex_subprocess_env()`, which forwards only PATH /
  HOME / USER / Windows-equivalents / CODEX_HOME / OPENAI_API_KEY /
  OPENAI_BASE_URL. Other-provider secrets like HF_TOKEN, GH_TOKEN,
  WANDB_API_KEY, ANTHROPIC_API_KEY no longer reach the local codex
  binary, so a shimmed `codex` earlier on PATH cannot harvest them.

* `_stream_thread_run` now tracks `emitted_any` and refuses to fall
  through to the buffered `await thread.run(prompt)` after either
  streaming helper has already yielded text. Previously a network
  glitch mid-stream re-executed the same Codex turn, which can
  duplicate file writes, shell commands, and other Codex side
  effects. The buffered path is now reserved for the zero-output
  case (no streaming helper resolved, or streaming returned empty).

* `CodexLoginButton` now aborts the SSE reader on unmount via a
  useEffect cleanup that calls `abortRef.current?.abort()`. The
  underlying `codex login --device-auth` subprocess no longer
  keeps streaming (and holding a device-auth session) after the
  dialog closes.

Two new pytest cases pin the behaviour: `test_codex_subprocess_env_scrubbed`
sets HF/GH/WANDB/ANTHROPIC keys and asserts none reach the codex
env while OPENAI_API_KEY / CODEX_HOME survive; and
`test_partial_stream_failure_does_not_replay_turn` injects a fake
`turn().stream()` that yields "partial output " then raises, and
asserts `thread.run()` is never called. 28/28 codex_provider tests
pass; `tsc --noEmit` clean.
2026-05-24 15:16:38 +00:00
Daniel Han
d6c47f6664 Studio: wire Codex provider through the UI end-to-end
Post-review pass driven by reviewer.py. The original PR shipped the
backend codex provider, the registry entry (with `hidden:true`), the
status API, and the `CodexParallelTabs` component, but the chat UI
never surfaced the row, required an API key for the connection, and
never sent `parallel_calls` over the wire. Also fixes a CodeQL leak
in the parallel fan-out error path and adds the canonical streaming
hook upstream actually exposes.

Frontend
* chat-providers-dialog.tsx now calls `/api/codex/status` alongside
  `/api/providers/registry`. When the host has Codex installed the
  Add connection dialog gains a synthetic Codex row (curated model
  list comes from `supported_models`) so the picker is reachable.
* The Add / Edit connection guards now skip the API-key requirement
  for Codex the same way they do for the custom OpenAI-compat
  presets; the field itself is also hidden so the user is not asked
  for a key Studio will not use.
* chat-adapter.ts now also exempts Codex from the "Missing API key"
  pre-flight, and emits `parallel_calls` on the outgoing request
  when the selected connection is Codex (clamped to [1, 20] by the
  shared helper, defaults to 1).
* external-providers.ts adds `codexParallelCalls` to
  ExternalProviderConfig so future composer UI can persist the
  user's pick per connection.

Backend
* `_stream_thread_run` now tries `thread.turn(prompt).stream()`
  first, mirroring the canonical openai_codex API
  (`openai/codex/sdk/python/src/openai_codex/api.py`). The legacy
  `thread.run_streaming(prompt)` path is kept as a fallback and the
  buffered `await thread.run(prompt)` stays as the last resort.
* `_stream_codex_parallel` no longer echoes `str(exc)` in the
  `codex_tab_error` SSE event. Per-tab failures now surface a
  generic "Codex tab failed" message plus an `exception_type`
  discriminator; `CodexUnavailableError` is the only exception
  whose text is forwarded verbatim because it is a user-actionable
  install hint with no sensitive content (CodeQL
  `py/information-exposure-through-exception`).

Tests
* New `TestCodexHardenedRegressions::test_parallel_tab_error_sanitised`
  injects a fake SDK that raises with a path-like message and
  asserts the SSE frames do not echo it.
* New `TestCodexHardenedRegressions::test_thread_turn_stream_path_taken`
  verifies the canonical `thread.turn(prompt).stream()` hook is
  preferred over the legacy helper.

All 26 codex_provider tests pass. Frontend `tsc --noEmit` clean.
2026-05-24 14:46:03 +00:00
Daniel Han
b6577a6287 Studio: name the actual PyPI package in the Codex install hint
The OpenAI Codex Python SDK ships on PyPI as
`openai-codex-app-server-sdk`, not `openai-codex` (which is the
GitHub repo project name in pyproject.toml). The runtime binary
ships separately as `openai-codex-cli-bin`. Both packages expose
the import name `openai_codex`; the older docs reference
`codex_app_server` so we keep probing both.

Update the `CodexUnavailableError` message and the provider
registry notes so a user hitting the unavailable path gets a
copy-pasteable `pip install` command. No behaviour change.
2026-05-24 14:25:04 +00:00
pre-commit-ci[bot]
861da31fcf [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 14:13:24 +00:00
Daniel Han
b8cd677397 Studio: harden Codex provider against upstream CLI/SDK shape
Followups on the post-merge review pass for the Codex SDK chat
provider. Verified against codex-cli 0.133.0 + the upstream
`openai/codex` Rust + Python sources, then pinned each fix with
a regression test in `test_codex_provider.py` (24/24 passing).

* Probe both `openai_codex` (canonical upstream Python package at
  `openai/codex/sdk/python`) and the legacy `codex_app_server`
  alias. Without this the availability probe always reported
  `sdk_importable: false` even when the SDK was installed, so the
  provider was permanently hidden.
* Switch the device-auth and login-status invocations from
  `codex auth login --device-auth` / `codex auth status` to the
  real upstream subcommands `codex login --device-auth` and
  `codex login status`. The former path returns
  `unrecognized subcommand 'auth'` on a real CLI.
* Strip ANSI control sequences before extracting the device URL
  (upstream wraps the URL in `\x1b[34m...\x1b[0m`) and tighten the
  pattern to the canonical `.../codex/device` shape. Also surface
  the one-time code as a `device_code` SSE event so the UI can
  show it alongside the URL.
* Fix `_detect_logged_in` substring footgun: `"logged in" in
  combined` matched inside `"not logged in"`, flipping logged-out
  users to logged-in. Anchor on word boundaries with negative
  prefixes winning regardless of return code.
* Cancel in-flight fan-out workers on SSE disconnect. Previously
  every parallel Codex turn ran to completion against a
  disconnected client and burned quota; now `_stream_codex_parallel`
  cancels its worker + drain tasks in a try/finally on
  `CancelledError`/`GeneratorExit`.
* Tear down the device-login subprocess on disconnect via
  `start_new_session=True` + `os.killpg(SIGTERM)` (Unix) or
  `CREATE_NEW_PROCESS_GROUP` + `CTRL_BREAK_EVENT` (Windows), with
  a bounded `proc.wait()` and `proc.kill()` fallback. Previously
  `finally: await proc.wait()` blocked the SSE close path because
  `codex login --device-auth` only exits on user action.
* Render the full conversation transcript in `_last_user_prompt`
  instead of returning only the most recent user message. The PR
  opens a fresh thread per request so prior assistant turns were
  dropped, degrading multi-turn chats to single-shot prompts.
  Single-turn input is unchanged.
* Make `ChatCompletionRequest.parallel_calls` default to 1 (`int`
  with `ge=1, le=20`) instead of `Optional[int] = None`. The
  runtime already coerced `None` -> 1, but the schema now matches
  the documented `[1, 20]` range.
* Replace the registry's hardcoded `default_models` (which
  contained `o3`, not in the upstream catalog) with the current
  `gpt-5.5 / 5.4 / 5.4-mini / 5.3-codex / 5.2` set from
  `codex-rs/models-manager/models.json`.
* Stop echoing `str(exc)` in SSE error frames in both
  `routes/inference.py` and `routes/codex.py`. The Codex SDK can
  raise with local paths, env-var content, or traceback fragments
  (CodeQL `py/information-exposure-through-exception`). Surface a
  generic message + `exception_type` discriminator; log the full
  reason server-side via `logger.error(..., exc_type=..., error=...)`.

Doc / comment updates throughout to refer to `codex login` /
`openai_codex` rather than the older incorrect strings.

Tested: pytest 24 cases in `test_codex_provider.py` (the original
14 + 10 new `TestCodexHardenedRegressions`) plus the rest of the
Studio-backend test suite the PR touches (209 passing). Also
verified live against Studio launched from this branch on a
Blackwell B200 via `UNSLOTH_STUDIO_HOME=$WORKSPACE/temp/...
./install.sh --local` then a Playwright probe.
2026-05-24 14:13:03 +00:00
Daniel Han
4188f916f4 wip: anthropic citation helper 2026-05-23 14:00:31 +00:00
pre-commit-ci[bot]
ea7ae85562 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-23 14:00:31 +00:00
Daniel Han
cbc3c43655 Studio: add Codex SDK as a chat provider with parallel-calls fan-out
Wires the OpenAI Codex CLI / Python SDK (codex_app_server) into Studio
as a new chat provider type. Hosts that don't have the CLI or the SDK
installed never see the entry; on logged-out hosts the provider config
dialog renders a device-auth Sign-in button that surfaces the
verification URL and streams CLI progress back over SSE.

Backend
- new core/inference/codex_availability.py probes the CLI + SDK and
  reports {installed, logged_in, version, supported_models}; it never
  imports codex_app_server at module top level so the rest of the
  backend keeps starting cleanly on hosts that don't have the SDK.
- new core/inference/codex_provider.py wraps AsyncCodex and translates
  Codex events into OpenAI chat-completion chunks. Supports the
  thread.run_streaming path with a non-streaming fallback for older
  SDK revs.
- parallel_calls > 1 fans the turn out across N tasks (capped at 20)
  via asyncio.gather and emits codex_tab_open / codex_tab_chunk /
  codex_tab_close tool-events per attempt plus a final codex_gather
  synthesis event. A separate standalone Codex call produces the
  unified answer.
- new routes/codex.py exposes GET /api/codex/status and POST
  /api/codex/login. The login route shells out to
  codex auth login --device-auth and streams events; the first event
  carries the verification URL so the frontend can window.open it.
- ChatCompletionRequest gains a parallel_calls field bounded [1, 20]
  by pydantic. The codex registry entry stays hidden by default; the
  /api/codex/status probe is the authoritative gate.
- routes/inference.py dispatches provider_type=codex through the
  local CLI/SDK pipeline instead of the standard HTTP client, with
  graceful error surfacing for CodexUnavailableError.

Frontend
- new api/codex-api.ts exposes fetchCodexStatus() and an async
  generator streamCodexDeviceLogin() that drives the SSE stream and
  yields parsed events.
- new components/codex-parallel-tabs.tsx renders the tabbed parallel-
  calls UI with a Synthesis tab highlighted once the codex_gather
  event arrives. Pure reducer keeps the state transitions unit-
  testable.
- new components/codex-login-button.tsx posts to /api/codex/login,
  opens the verification URL in a new tab via window.open, and shows
  the streamed CLI log as it lands.
- external-providers.ts exports CODEX_PROVIDER_TYPE,
  CODEX_MAX_PARALLEL_CALLS, isCodexProviderType, and
  clampCodexParallelCalls. Codex is marked text-only so the composer
  hides image-attach affordances when selected.

Tests
- tests/test_codex_provider.py (14 cases) covers the availability
  probe across the four install / login states, the streaming +
  parallel-calls translation against a fake codex_app_server module
  injected into sys.modules, the [1, 20] pydantic clamp, the
  CodexUnavailableError surfacing path, and the parallel_calls=1
  single-call shape (no tab tool-events).
2026-05-23 14:00:31 +00:00
Daniel Han
ebe504b558
Studio: PDF / document attachments for Anthropic + OpenAI (#5689)
* Studio: PDF / document attachments for Anthropic + OpenAI

Studio's local-GGUF chat already supports image attachments via the
`image_url` content part shape. PDFs and other documents had no
plumbing for the external-provider path: there was no normalised
content type the frontend could send that translated to Anthropic's
native `document` block or OpenAI's `input_file`.

Add a Studio-side `input_document` content part on assistant /
user messages with three shapes:

  {type: "input_document",
   file_data: "data:application/pdf;base64,<DATA>",
   filename?: "name.pdf",
   media_type?: "application/pdf"}

  {type: "input_document",
   file_url: "https://example.com/doc.pdf",
   filename?: "doc.pdf"}

Translation:

- Anthropic Messages API: emits a `document` block with
  `{source: {type:"base64", media_type, data}}` or
  `{source: {type:"url", url}}`, plus an optional `title` from
  `filename`. PDFs are extracted server-side by Anthropic per their
  vision/document docs and counted toward input tokens.
- OpenAI Responses API: emits `{type:"input_file", file_data |
  file_url, filename?}`. PDFs are extracted server-side.

Empty / unparseable `input_document` parts are silently dropped so
a malformed frontend payload can't blow up the request.

Tests:

- New `test_multimodal_document.py` with 6 cases pinning the
  outbound body shape for base64 + URL inputs on both providers,
  and the empty-part drop behavior on both.
- The Anthropic assertions strip the prompt-cache wrapper
  (`cache_control:{type:ephemeral}` that the tail-message caching
  layer adds) before comparing the document core fields, so this
  test stays focused on the translation, not the caching layer.

Live verified end-to-end against both providers: a 363-byte
single-page "HELLO" PDF, base64-encoded, attached as a `document`
block to Opus 4.7 and as an `input_file` to gpt-5.5. Both models
correctly extracted the word "HELLO" from the PDF.

Follow-up (out of scope):

- Pydantic schema entry on ChatMessage.content for `input_document`
  (today it rides through because ChatCompletionRequest uses
  extra=allow). Will tighten when the frontend attach button lands.
- Frontend file-picker UX for non-image attachments on the external
  provider path.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: gate empty-content msg + skip empty data-URI payload

Gemini High + Codex P2 on PR #5689:

1. Anthropic translation appended an empty `anthropic_parts` array
   when every part was dropped (e.g. user sent only an unparseable
   input_document). Anthropic 400s on "messages.N.content: at least
   one block is required". Skip the whole-message append when no
   parts survived. The OpenAI Responses path already had the
   equivalent guard, so this brings the two providers into parity.

2. `data:application/pdf;base64,` with no payload (or whitespace-only)
   parses to an empty `source.data` string. Anthropic rejects that
   with 400 as well. Skip the document block before constructing it.

Plus 2 new test cases pinning both behaviors:

- `test_anthropic_empty_only_document_drops_whole_message`: confirms
  a turn whose only content is an unparseable input_document does
  NOT make it onto the outbound `messages` array.
- `test_anthropic_empty_data_uri_payload_is_dropped`: confirms an
  empty-payload data-URI is filtered out at translation time.

(Note re: gemini's other High note about adding `input_document` to
the Pydantic ContentPart union -- ChatCompletionRequest is configured
with `extra=allow` so the part rides through today. Tightening the
union belongs with the frontend attach-button PR that surfaces the
field; called out as follow-up in the PR description.)

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: register input_document in ContentPart + builder

Reviewer caught that the translation code on the external_provider
side was unreachable from a real ChatCompletionRequest:

- ContentPart is a discriminated Union of (text, image_url) only, so
  any `{"type": "input_document", ...}` part was rejected by Pydantic
  at request parsing with a discriminator error before the helper
  could see it.
- _build_external_messages in routes/inference.py only walked text
  and image_url parts, so even with a permissive schema the document
  parts would have been silently dropped instead of forwarded to
  the per-provider translator.

Fixes:

- Add InputDocumentContentPart with optional file_data / file_url /
  filename / media_type and Tag("input_document") on the Union.
- Extend _build_external_messages to pass input_document through as
  a plain dict for vision-capable providers (so external_provider's
  existing Anthropic `document` and OpenAI Responses `input_file`
  mappers actually run) and strip them on non-vision providers.

Tests added: schema accepts input_document, builder passes it to
vision providers, builder strips it on non-vision providers.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: validate file_data before preferring over file_url

Codex P2 caught that the OpenAI input_document translator treats any
truthy file_data as valid and never falls back to file_url. That
means a malformed `data:application/pdf;base64,` (empty payload) or
a whitespace-only data URI gets forwarded as `file_data=""` and
400s the whole turn, AND silently discards a perfectly recoverable
file_url on the same part.

Mirror the Anthropic-side guard onto the OpenAI Responses path:
treat any "data:" URI with no actual base64 payload as missing and
fall through to file_url. Standalone-empty data URIs (no fallback)
are dropped entirely instead of being sent to the wire.

Tests added: empty data URI + valid file_url -> file_url wins,
whitespace-only data URI + valid file_url -> file_url wins,
empty data URI without fallback -> part is dropped.

* Address review: Anthropic side also falls back to file_url on empty data URI

Codex P2 follow-up to my earlier fix: I added the empty-data-URI ->
file_url fallback to the OpenAI Responses translator but missed
the Anthropic translator, which still `continue`d on empty payloads
and discarded an otherwise valid file_url on the same part. Result:
when the frontend supplied both file_data (placeholder / broken)
AND a working file_url, Anthropic silently lost the attachment;
when the message contained only that part, the whole message could
be dropped before reaching the wire.

Mirrored the OpenAI guard: any "data:" URI with no actual base64
payload (`data:application/pdf;base64,` or whitespace-only) is
treated as missing, and the file_url branch takes over. The
all-parts-dropped guard further down already handles the
no-fallback case.

Tests added: empty data URI + valid file_url -> URL source on the
wire with the filename preserved; whitespace-only data URI + valid
file_url -> URL source on the wire.

* Address review: gate input_document passthrough to anthropic + openai

Codex P1: only `_stream_anthropic` and `_stream_openai_responses`
have explicit translation logic for input_document parts (the former
maps to {type:"document", source:...}, the latter to
{type:"input_file", file_data|file_url}). Every other provider
(gemini / mistral / kimi / openrouter / deepseek / qwen / custom)
goes through the generic /chat/completions passthrough that forwards
`messages` verbatim, so any input_document part on a non-vision
route on those providers would 400 with an unknown content_part
type.

Added `_INPUT_DOCUMENT_PROVIDERS = frozenset({"anthropic", "openai"})`
constant and gated the pass-through branch on `provider_type in
_INPUT_DOCUMENT_PROVIDERS`. Every other provider strips the part
(text content survives). Threaded provider_type through from
_proxy_to_external_provider's call site.

Tests updated: vision + provider in {anthropic, openai} still
forwards; six unmapped providers (gemini/mistral/kimi/openrouter/
deepseek/qwen) strip the part; missing provider_type strips
defensively. The existing non-vision drop test still passes.

* Fix stale web_fetch tool-version assertion after merging main

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:22:57 -07:00
Daniel Han
e86f3c5dc7
Studio: wire OpenAI Responses server-side context compaction (#5687)
* Studio: wire OpenAI Responses server-side context compaction

The OpenAI Responses API accepts a `context_management` field that
enables server-side compaction. When the rendered prompt crosses the
configured threshold, the API runs a server-side compaction step and
the request continues against the compacted prefix. No beta header
and no dated version pin are required, per the docs.

Changes:

- Add `compaction_threshold: Optional[int]` (ge=1_000, le=2_000_000)
  to ChatCompletionRequest. Thread through `routes/inference.py` ->
  `stream_chat_completion` -> `_stream_openai_responses`.
- In `_stream_openai_responses`, when threshold is set AND the base
  URL points at cloud OpenAI (api.openai.com), attach
  `context_management: [{type:"compaction", compact_threshold:N}]`
  to the outbound body. Non-cloud bases (ollama, llama.cpp, "custom"
  presets) silently drop the field so we don't 400 those servers.
- Add `test_openai_compaction.py` with 4 cases: cloud OpenAI sets
  the field verbatim, low-threshold probe passes through (we don't
  clamp on the OpenAI side because the API accepts whatever),
  non-cloud base drops the field, omitted threshold leaves body
  untouched.

Live verified against the real OpenAI API on gpt-5.5:
`context_management:[{type:"compaction", compact_threshold:200000}]`
returns 200 with no error.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: accept Azure OpenAI base URLs + raise compaction floor

Two reviewer follow-ups on the OpenAI compaction PR:

1. The `is_openai_cloud = "api.openai.com" in self.base_url` check
   excluded Azure OpenAI Foundry, even though Azure exposes the
   same /v1/responses extensions (context_management,
   prompt_cache_retention, container shell). Users on Azure saw
   their compaction toggle silently no-op. Broadened the check to
   also match `*.openai.azure.com` and made it case-insensitive so
   URLs copy-pasted from the Azure portal still resolve. Non-cloud
   OpenAI-compatible servers (ollama / llama.cpp / vLLM / "custom"
   preset) still fall outside the gate.

2. The schema floor on compaction_threshold was ge=1_000, which is
   well below the upstream Responses API's effective minimum
   (vercel/ai#12486, langchain-ai/langchain#35464 report
   `compact_threshold is not enabled` 400s on Azure at 100k; cloud
   uses 200k as the canonical example). Raised the floor to 10k
   so obvious typos surface as a clean 422 from FastAPI rather than
   an opaque upstream 400 the user has to debug from the SSE
   stream.

Tests added: Azure base URL carries both context_management and
prompt_cache_retention; mixed-case Azure URLs match; schema rejects
9_999 and accepts 10_000.

* Address review: drop schema-level compaction floor (cross-provider regression)

Codex P2 follow-up on the previous floor bump: ge=10_000 was
enforced globally at the ChatCompletionRequest layer, but the field
is documented as a no-op on every non-cloud OpenAI base and every
non-OpenAI provider. With the global floor, an Anthropic / ollama
/ llama.cpp / custom request that happens to carry compaction_threshold
below 10k was rejected with 422 at request validation time instead
of being silently ignored as the description promised.

Reverted the schema floor to ge=1 (any positive int) and rewrote
the description to call out per-provider routing: OpenAI cloud's
effective floor is around 200k and surfaces upstream 400s below
that; _stream_anthropic clamps sub-50k values up. Per-provider
helpers stay the single source of truth on the floor.

Test updated to pin: zero is still rejected, but every positive
value (1, 5_000, 9_999, 10_000, 200_000) passes schema validation.

* Address CodeQL: hostname-anchored OpenAI cloud detection

CodeQL py/incomplete-url-substring-sanitization fired on
`".openai.azure.com" in _base`. An attacker who controls the
configured base_url could slip cloud-only request body fields
(prompt_cache_retention, context_management compaction, container
shell) to an arbitrary server with:

  https://evil.com/api.openai.com/v1
  https://api.openai.com.attacker.com/v1
  https://attacker.com/.openai.azure.com/v1
  https://my-resource.openai.azure.com.attacker.com/openai/v1

Replaced the substring check with a `_is_openai_family_cloud`
helper that runs urllib.parse.urlparse on the URL and matches the
lowercased hostname exactly (`api.openai.com`) or via `endswith`
on the leading-dot suffix (`.openai.azure.com`). Both halves are
host-anchored so path / fake-subdomain bypasses fail.

Test added: every attacker-controlled bypass shape above must NOT
carry context_management OR prompt_cache_retention on the wire.
Existing Azure and openai.com tests still pass.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: scope compaction_threshold description to OpenAI on this branch

Codex P2: the field description on this PR mentioned Anthropic
compaction behavior, but the Anthropic wiring lives on PR 5686
(separate branch). On feat/openai-compaction alone, _stream_anthropic
has no compaction_threshold parameter, so the field is silently
ignored for Anthropic requests and the doc claim was misleading.

Trimmed the description to OpenAI cloud + Azure Foundry only on
this branch. PR 5686 already re-adds the Anthropic clause via its
own change, so the rebase / merge order on main will land the
combined description naturally once both PRs ship.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:20:45 -07:00
Daniel Han
9a737facaf
Studio: wire Anthropic server-side context compaction (#5686)
* Studio: wire Anthropic server-side context compaction

Anthropic ships server-side context compaction as a beta
(`compact-2026-01-12`). When the rendered prompt crosses the
configured input-token threshold, Anthropic runs an extra LLM pass
that summarises older turns and the request continues against the
compacted prefix. The response carries the original top-level fields
plus a new `context_management` block (with `applied_edits`) and
`usage.iterations[]` accounting per pass.

Per the docs the feature is currently supported on Opus 4.6, Opus 4.7,
Sonnet 4.6, and Mythos preview. The minimum threshold is 50k tokens;
under-50k requests 400.

Changes:

- Add prefix gate + helper `_anthropic_supports_compaction` plus
  constants `_ANTHROPIC_COMPACTION_PREFIXES`, `_ANTHROPIC_COMPACTION_BETA`,
  `_ANTHROPIC_COMPACTION_TYPE`, `_ANTHROPIC_COMPACTION_MIN`.
- Add `compaction_threshold: Optional[int]` to ChatCompletionRequest
  (50k ge bound, 2M le bound). Thread through `routes/inference.py`
  -> `stream_chat_completion` -> `_stream_anthropic`.
- In `_stream_anthropic`, when threshold is set AND the model
  accepts compaction, attach `context_management.edits[{type:
  "compact_20260112", trigger:{type:"input_tokens", value:N}}]` to
  the outbound body. Sub-50k values are clamped up to 50k to keep
  the request well-formed.
- Refactor the anthropic-beta header builder to merge any combination
  of `code-execution-2025-08-25` + `compact-2026-01-12` flags into
  one header value. Unrelated betas added at the registry level still
  pass through.
- Add `test_anthropic_compaction.py` with 16 cases: gate matrix
  (every doc-listed model), correct body shape, threshold clamping,
  beta header merge with code execution, silent no-op on unsupported
  models, omitted-threshold pass-through.

Live verified end-to-end against the real Anthropic API:
`compact_20260112` accepted on Opus 4.7, response carries
`context_management.applied_edits` + `usage.iterations[]` as
documented. (The first WebFetch-summarised version of these docs
suggested `compact_20260120`; the actual API only accepts
`compact_20260112`, matching the beta-header date. Worth pinning
behind a test so a future doc update can't drift back.)

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: drop ge=50_000 clamp + parse usage.iterations[]

Two reviewer follow-ups on the compaction PR:

1. Pydantic ge=50_000 on compaction_threshold was dead code.
   FastAPI rejected sub-50k threshold values with a 422 before the
   `max(int(...), _ANTHROPIC_COMPACTION_MIN)` clamp in
   _stream_anthropic could ever fire. Relaxed the floor to ge=1 so
   the in-helper clamp actually does its job; the schema comment
   now explains why this is intentional. Added a regression test
   that posts a value of 1 and 49_999 through the real request
   schema.

2. Anthropic publishes per-iteration token counts in
   `usage.iterations[]` whenever a fresh compaction has run, and
   the top-level input_tokens / output_tokens cover only the
   `message` iteration -- billing must add the compaction
   iterations on top. Aggregate compaction iteration tokens into
   `last_usage["compaction_input_tokens" / "compaction_output_tokens"]`
   so the cost surface (PR 5690) can read them without re-walking
   the array, and surface both figures in the closing stream
   summary log. Added two tests: one that pins the aggregation on a
   compacted turn and one that pins `None` when no fresh
   iterations land (so re-applied compaction blocks don't double-bill).

Sourcing: https://platform.claude.com/docs/en/build-with-claude/compaction

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: round-trip Anthropic compaction blocks across turns

Codex P1: once context_management is enabled and Anthropic runs
server-side compaction mid-stream, the response carries a
`{type:"compaction", content:"<summary>"}` content block on the
assistant message. The translator only handled text_delta and
input_json_delta on content_block_delta, so the compaction block
was silently dropped. Worse, the request schema's ContentPart
discriminated Union didn't accept `type:"compaction"`, and
_build_external_messages didn't pass it through, so even a
hand-crafted assistant message carrying the block would 422 at
parse time. Net result: Anthropic re-compacted from scratch on
every subsequent turn, wasting input tokens and reasoning budget.

End-to-end backend wiring of the round-trip:

1. SSE translator. _stream_anthropic now tracks a `current_compaction`
   state slot. content_block_start with type=="compaction" seeds it
   (Anthropic may include the summary on the start event AND/OR
   stream it via text_delta events on the same block index --
   handle both). text_delta inside a compaction block routes into
   the compaction buffer instead of the user-visible content
   stream, since the summary is opaque internal state, not
   assistant prose. content_block_stop emits a `compaction_block`
   tool_event carrying the full summary so the chat-adapter can
   persist it. compaction_blocks_seen is surfaced in the closing
   summary log.

2. Pydantic schema. Added CompactionContentPart with Tag("compaction")
   on the ContentPart Union so requests carrying the block parse
   cleanly. Required `content` field with a docstring pointing at
   the Anthropic docs.

3. Message builder. _build_external_messages forwards compaction
   parts on both vision and non-vision paths; the per-provider
   stream helper decides whether to forward to the wire (Anthropic
   does; other providers ignore the part). When a non-vision route
   ends up with a single text part, collapse back to a string
   so providers that don't accept content arrays still get the
   expected shape.

4. _stream_anthropic outbound translator. {type:"compaction"} parts
   on an assistant message land on the wire verbatim. Empty/missing
   `content` is skipped so a malformed stored block can't 400
   Anthropic.

Tests added (5): stream emits compaction_block tool event with the
summary intact; user-visible content stream does NOT carry the
summary text; outbound body forwards compaction parts verbatim on
the next turn; Pydantic schema accepts the part; builder passes
it through on both vision and non-vision provider routes.

Frontend follow-up: the chat-adapter needs to persist the
compaction_block tool_event onto the stored assistant message so
turn N+1 includes it in payload.messages. Pinned in the PR
description.

Sourcing: https://platform.claude.com/docs/en/build-with-claude/compaction

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: gate compaction-part passthrough to Anthropic only

Codex P1: my previous round-trip change preserved {type:"compaction"}
parts on every provider route in _build_external_messages. That
meant a chat history with prior compaction state silently leaked
the Anthropic-specific block to OpenAI/DeepSeek/Mistral/Gemini/
Kimi/OpenRouter on a provider switch, where generic
/chat/completions passthrough hands the unknown content type to
the upstream API and 400s the whole turn.

Added a `provider_type` kwarg to _build_external_messages and
gated the compaction forwarder on `provider_type == "anthropic"`.
Every other value (including the legacy None for callers that
don't pass it yet) strips the part. The Anthropic stream helper
still maps it to a native `compaction` block on the wire.

Threaded provider_type through from _proxy_to_external_provider's
call site.

Tests updated: vision + provider="anthropic" still forwards; six
non-anthropic providers strip the part; missing provider_type
strips defensively; non-vision + anthropic still forwards; non-vision
+ non-anthropic collapses back to a text string.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:19:09 -07:00
Daniel Han
a2d2b7866f
Studio: wire Anthropic web_fetch server-side tool (#5671)
* Studio: wire Anthropic web_fetch server-side tool

Studio's Anthropic passthrough only forwarded web_search and
code_execution when enabled_tools was set. Asking Claude through Studio
to fetch a URL produced no fetch (the tool was not in the outbound
tools array), so users had to fall back to web_search even when they
already had the exact URL they wanted.

This change opts in web_fetch_20250910 when enabled_tools contains
"web_fetch". The new tool entry is appended alongside any existing
web_search / code_execution entries:

  {"type": "web_fetch_20250910", "name": "web_fetch", "max_uses": 5}

No anthropic-beta header is required (web_fetch is GA); the existing
code-execution-2025-08-25 flag continues to merge cleanly when both
tools are enabled in the same turn.

SSE translation mirrors the web_search path. A `server_tool_use` block
with name="web_fetch" emits a `tool_start` _toolEvent carrying the
URL the model asked to fetch; the matching `web_fetch_tool_result`
block emits a `tool_end` _toolEvent whose result string follows the
Title / URL / Snippet shape parseSourcesFromResult on the frontend
already expects, so the source pill renders identically. Error blocks
(`web_fetch_tool_error`) are surfaced as "Error: <error_code>" matching
the code_execution error path.

The final "Anthropic stream complete" log line picks up web_fetch_
requested / web_fetch_invocations / web_fetch_urls so support reports
of "the model did not fetch anything" can be triaged from the log.

Verified end to end against claude-haiku-4-5 with
`enabled_tools=["web_fetch"]`: the model emitted tool_start with
url=https://example.com and tool_end with the page Title + URL +
Snippet, plus the assistant message correctly read back "Example
Domain" as the title.

Tests:
- 5 new unit tests in test_anthropic_web_fetch.py covering tool
  registration, the combined web_search + web_fetch + code_execution
  request body, the pill-off case, and SSE translation for both
  success and error paths.
- All 242 existing Anthropic + OpenAI provider tests still pass.

The enabled_tools field description in models/inference.py is updated
so OpenAPI consumers see the new option.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* web_fetch: title fallback to URL, log parse failures, drop dead checks

Three review nits on the previous commit:

1. `_format_web_fetch_result` left `title` empty when Anthropic omitted
   `document.title`. The frontend `parseSourcesFromResult` only emits
   a source pill when both `Title:` and `URL:` lines are present, so
   fetches against pages without an HTML title tag silently lost
   their citation in the UI. Fall back to `title = title or url`,
   matching the web_search formatter.

2. The broad `except Exception` around `json.loads(buffer)` for the
   web_fetch input swallowed the failure with no trace. Log at debug
   so a malformed partial_json buffer can be triaged from the server
   log without changing behavior.

3. `inner` was already sanitised to a dict at the matching
   content_block_start and `_format_web_fetch_result` always returns
   a non-empty string (defaulting to "(fetch complete)"), so the
   `isinstance(inner, dict) else {}` guard and the
   `result_text or "(fetch complete)"` fallback at the emit site
   were dead code. Removed.

Added a test exercising the titleless path so the fallback stays
covered.

* chat-adapter: emit source pills for web_fetch tool calls

`parseSourcesFromResult` was only wired up for tool calls where
`toolName === "web_search"`, so the Title / URL / Snippet block the
backend formatter emits for `web_fetch_tool_result` never reached the
source-pill renderer. Users saw the raw tool result in the tool card
but the dedicated source-pill row at the message tail stayed empty.

Both web_search and web_fetch ship the same text shape today, so the
fix is to broaden the gate.

* Address review: wire web_fetch from Search pill + fix pause_turn truncation

Two reviewer follow-ups on the Anthropic web_fetch PR:

1. The backend tool wiring landed but the frontend chat-adapter
   never put `web_fetch` in `enabled_tools`, so toggling the Search
   pill only ever attached `web_search` -- web_fetch was unreachable
   from the UI. Added providerSupportsBuiltinWebFetch() (Anthropic
   today) and paired the entry with the existing Search pill, since
   the canonical workflow is "search returns URLs, fetch reads
   them" and there is no separate UI toggle yet.

2. `pause_turn` from Anthropic's stop_reason vocabulary fell through
   the finish_reason map's "stop" default, which the OpenAI-format
   client renders as end-of-message and truncates the answer. Per
   the docs pause_turn means "Claude paused a long server-tool
   turn (web_search / web_fetch) and will resume". Mapped to None
   and skipped the chunk emission so the SSE stream still ends with
   [DONE] on message_stop but no terminal finish_reason lands on
   the client. While there: added explicit mappings for `tool_use`
   (-> tool_calls) and `refusal` (-> content_filter) which were
   also falling through to "stop".

Tests added: pause_turn emits no finish_reason, end_turn still
emits "stop", refusal maps to "content_filter".

Sourcing: https://platform.claude.com/docs/en/api/messages#response-stop-reason

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:03:48 -07:00
Daniel Han
2201fd687b
Studio: per-session cost calculator + /api/providers/pricing endpoint (#5690)
* Studio: per-session cost calculator + /api/providers/pricing endpoint

Neither the Anthropic Messages API nor the OpenAI Responses API
reports a `cost` field on the response. Both expose detailed token
counts (input, output, cache hits, server-tool invocations); pricing
multipliers live in the provider docs. The frontend's "cost so far"
display was impossible without scraping the server log.

Land the math + a snapshot endpoint so the cost calculator can run
client-side from the existing usage chunk plumbing. The actual UI
hookup belongs in a frontend follow-up (and is gated on PR #5670's
usage-chunk emission landing so the frontend sees the usage block
in the first place).

Changes:

- New `core/inference/pricing.py` with:
  - Per-MTok base pricing tables for every active Anthropic and
    gpt-5.x family member. Dated snapshots inherit the canonical-id
    price via prefix match so future snapshots cost the same as the
    canonical id until pricing changes.
  - Shared multipliers for Anthropic cache writes (5m: 1.25x, 1h: 2x)
    and reads (0.1x); OpenAI cache reads (0.1x); Anthropic server
    tool surcharges ($10 / 1k web_search, $0.05 / hour code_exec
    beyond the 50-hour daily free tier).
  - `calculate_cost(provider, model, usage)` returns a per-turn USD
    breakdown plus billable token counts, with priced=False for
    unknown models so the UI can still render token counts.
  - `pricing_snapshot()` returns the whole table for the frontend
    so it doesn't re-implement the multipliers.
- New `GET /api/providers/pricing` returning the snapshot, scoped
  behind the existing auth dependency.
- New `backend/tests/test_pricing.py` with 12 cases pinning the
  math against documented values: base input/output multiplication,
  5m / 1h / read multipliers, default-to-5m fallback when the
  breakdown is absent, web_search per-1k pricing, code_execution
  per-hour pricing, dated-snapshot fallback, OpenAI cache-read
  discount accounting (cached tokens subtracted from full-price
  bucket and re-billed at 0.1x), unknown model graceful-degrade,
  and the snapshot endpoint shape.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: verified OpenAI pricing + fix billable input double-count

Address the cost-calculator review:

- OpenAI prices were 2-6x under the actual published rates.
  Cross-checked the live developers.openai.com/api/docs/pricing page
  and replaced every entry. gpt-5.5 is 5/30, gpt-5.5-pro is 30/180,
  gpt-5.4 is 2.5/15, gpt-5.4-mini 0.75/4.5, gpt-5.4-nano 0.20/1.25,
  gpt-5.3-codex 1.75/14. Added chat-latest alias to the canonical
  chat-snapshot rate. Dropped o3 / o4 / gpt-4.5 rows that are no
  longer listed on the page; calculator returns priced=False instead
  of silently billing at zero.

- billable_input_tokens was double-counting cached tokens for
  OpenAI. Anthropic excludes cache_* buckets from input_tokens so
  we add them; OpenAI folds cache_read_input_tokens into
  input_tokens already, so the tooltip read 1.8M for a 1.0M bill.
  Branched the math by provider and added a regression test.

Sourcing notes in the module docstring updated.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: canonical 4.5 ids, long-context tier, OpenAI tool fees

Three Codex P1 follow-ups on the cost calculator:

1. Canonical Anthropic 4.5 ids missing from ANTHROPIC_PRICING.
   claude-opus-4-5 / claude-sonnet-4-5 / claude-haiku-4-5 (no date
   suffix) are the ids used by backend defaults
   (PROVIDER_REGISTRY['anthropic'].default_models), but the table
   only had the dated forms. _lookup's prefix fallback doesn't help
   because the canonical id is SHORTER than the dated key, so
   str.startswith goes the wrong way and the calculator returned
   priced=False + zero cost. Added the canonical aliases for
   opus-4-5, sonnet-4-5, haiku-4-5, and opus-4-1.

2. OpenAI long-context tier. gpt-5.5 and gpt-5.4 cross over at
   272k input tokens to a 2x input / 1.5x output rate (gpt-5.5:
   $5/$30 -> $10/$45; gpt-5.4: $2.50/$15 -> $5/$22.50). Turns past
   the threshold were systematically undercounted at headline
   rates. Added long_context_threshold / long_context_input_per_mtok /
   long_context_output_per_mtok columns and a tier-selection step
   in calculate_cost; model_priced gains a "(long-context >272000)"
   suffix when the higher tier applies so the tooltip can show
   which rate was used. gpt-5.5-pro / gpt-5.4-pro / mini / nano /
   codex have no published long-context tier today, so they keep a
   single rate.

3. OpenAI server-tool surcharges. web_search is $10/1000 calls and
   the hosted shell container is $0.03 per 20-minute session on the
   default 1g tier (~$0.09/hr). server_tools_usd was previously
   stuck at 0.0 for OpenAI even when web_search and shell tools
   fired, so sessions with tool use understated cost. Added
   OPENAI_WEB_SEARCH_USD_PER_1K and OPENAI_CONTAINER_USD_PER_HOUR
   constants plus a parallel of the Anthropic surcharge block that
   reads counts from usage["openai_tool_use"]. The SSE translator
   wires the counts in a follow-up commit; the calculator is now
   ready for them. pricing_snapshot also exposes both constants so
   the frontend tooltip can render the per-call rate.

Existing tests updated to stay in the short-context tier where they
were testing base rates; new tests pin canonical 4.5 lookups,
long-context crossover on gpt-5.5/gpt-5.4, the absence of crossover
on mini/nano/codex, and OpenAI tool surcharges (web_search,
container hours, combined total).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:03:43 -07:00
Daniel Han
5b41872e8b
Studio: wire OpenAI image_generation tool (#5688)
* Studio: wire OpenAI image_generation tool

OpenAI's Responses API exposes server-side image generation as a
tool entry (`{type: "image_generation"}`); the result comes back as
an `image_generation_call` output item with the base64 image on
`result`, the actual prompt used on `revised_prompt`, plus `size`,
`quality`, `output_format`, `background`. The model decides when to
call the tool based on the user's request; rendering uses one of
the gpt-image-* backbones server-side.

Available on every gpt-5.x family member plus gpt-4.1, gpt-4o, o3,
o4-mini per the docs.

Changes:

- Append `{type:"image_generation"}` to the Responses request tools
  array when `enabled_tools` carries `image_generation` AND the base
  URL points at cloud OpenAI. Non-cloud bases (ollama, llama.cpp,
  "custom" presets that collapse to provider="openai") silently drop
  the tool to avoid 400s.
- Mirror the same logic in `_build_body` (the post-expiry retry
  builder) so retries carry the same tool set as the original
  attempt.
- Handle `image_generation_call` items in
  `response.output_item.done`: emit `tool_start` with
  `arguments:{kind:"image", prompt:<revised_prompt>}` and `tool_end`
  with `image_b64`, `image_mime`, `size`, `quality`, `background`
  so the chat adapter can render an inline preview. Image bytes go
  on the tool_end chunk; no extra fields on the chat-completions
  envelope so the OpenAI SDK shape stays clean.
- Add `import time` (used for synthesised tool_call_id fallback).
- Add `test_openai_image_generation.py` with 5 cases: tool entry on
  cloud OpenAI, combined with web_search + code_execution
  (verifies all three coexist), non-cloud drop, omitted pill leaves
  body untouched, output item translation produces the expected
  tool_start + tool_end chunks.

Live verified end-to-end: `gpt-5.4-mini` with `image_generation`
tool returned an `image_generation_call` carrying ~1MB of base64
PNG plus the gpt-image backbone's revised prompt.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Use time.time_ns() for synthesised image_generation tool_call_id

Gemini medium on PR #5688: `int(time.time() * 1000)` has 1ms
resolution; two image generations resolving in the same millisecond
would collide on the synthesised id. Bump to nanoseconds.

(In practice the upstream `image_generation_call` item always carries
its own `id`; the synthesised fallback only fires when OpenAI omits
it -- rare, but cheap to harden.)

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:03:38 -07:00
Daniel Han
b8dde0a835
Studio: support Anthropic 1h cache TTL via prompt_cache_ttl (#5685)
* Studio: support Anthropic 1h cache TTL via prompt_cache_ttl field

Anthropic exposes two ephemeral cache pools per request: the default
5-minute pool, and a 1-hour pool selected by attaching `ttl:"1h"` to
the `cache_control` marker. 1h writes are billed at 2x base input vs
1.25x for 5m, but reads stay at 0.1x for both, so a single extra read
landing more than 5 minutes after the write pays off the premium.

Studio hardcoded the 5m pool via `cache_control: {type:"ephemeral"}`
on both breakpoints. For chats with multi-minute idle gaps (people
juggling tabs, long-running tool calls between turns), the cache
expires before the next turn and every read becomes a cache_creation,
not a cache_read -- exactly the case where the 1h pool wins.

Changes:

- Add `prompt_cache_ttl: Optional[Literal["5m", "1h"]]` to
  ChatCompletionRequest. Default (None) preserves today's 5m behavior.
- Thread through `routes/inference.py` ->
  `stream_chat_completion` -> `_stream_anthropic`.
- Build a shared `cache_marker` dict in `_stream_anthropic`; attach
  `ttl` only when the request asks for one of the two valid values.
  Unknown TTL strings are silently dropped to avoid sending malformed
  markers (the upstream API would 400).
- Apply the same marker to both existing breakpoints (system block at
  line 1175 and the latest-message tail at line 1198 / 1213) so the
  pool selection is consistent across the whole prefix.
- Add `test_anthropic_cache_ttl.py` with 11 parametrized cases
  pinning the outbound body shape: omitted -> default marker;
  explicit `5m`/`1h` -> ttl field set; unknown values dropped;
  caching off -> no markers at all.

Verified upstream that `cache_control: {type:"ephemeral", ttl:"1h"}`
is accepted by the Anthropic API today; no beta header required.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Relax prompt_cache_ttl to Optional[str] (Codex P1)

Declaring `prompt_cache_ttl` as `Optional[Literal["5m", "1h"]]` made
FastAPI/Pydantic 422 the request before _stream_anthropic could even
see the field. The whole point of the downstream drop-unknown-values
behaviour was to keep a stale frontend from crashing the request;
the strict Literal at the request layer defeated that.

Loosen the schema to Optional[str]; the existing in-helper guard
already restricts forwarded values to {"5m", "1h"} (everything else
is silently dropped). Test suite stays unchanged -- the bogus-value
cases in test_anthropic_cache_ttl.py already pass arbitrary strings
through and assert they are dropped before the wire.

* Address review: confirm extended-cache-ttl beta header is GA

Reviewer asked whether the 1h cache TTL still requires the
`extended-cache-ttl-2025-04-11` anthropic-beta header. Investigated:

- Live-tested api.anthropic.com on claude-opus-4-7 (2026-05-22)
  with cache_control={type:"ephemeral", ttl:"1h"} and NO beta
  header. Got status 200 and ephemeral_1h_input_tokens populated
  on the create turn, plus cache_read_input_tokens populated on
  the reuse turn.
- Cross-checked the current prompt-caching docs: no mention of
  any beta header on the 1h TTL path.

Conclusion: the gate has been promoted to GA. The code already
does not send the beta header (the cache_marker dict only carries
`type`/`ttl`), so no wire change is needed. Pinned the contract
with two regression tests that assert the header is NOT on the
outbound request, and added a docstring note explaining the
investigation outcome so a future reader does not re-add it.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:03:32 -07:00