Commit graph

8 commits

Author SHA1 Message Date
Daniel Han
187144d4e7
Reduce and tighten code comments and docstrings repo-wide (#6095)
Trim and tighten code comments and docstrings across the repository. Comment-only: every changed file verified code-identical to main via AST/token comparison.
2026-06-08 23:09:51 -07:00
Daniel Han
8292e699e4
Studio: make code comments and docstrings more succinct (#6029)
Trim and tighten code comments and docstrings across studio/ Python. Comment-only: every changed file verified code-identical to main via AST/token comparison.
2026-06-08 23:07:28 -07:00
Daniel Han
3ce187da02
Formatting: ruff line-length 100, kwarg-spacing passes, drop blank after short local imports (#6079)
Raise ruff line-length to 100 and extend the local pre-commit format pipeline (def-signature magic-comma normalization, short multi-line assert collapse, kwarg '=' spacing, blank-line-after-short-import removal, adjacent string-literal / f-string+plain merge, redundant-pass pruning). Every transform re-checks the file AST and is dropped if it would differ; the whole-repo reformat is verified AST-identical per file and idempotent.
2026-06-08 04:24:13 -07:00
Daniel Han
ab48465135
Studio: add Gemini provider with web_search, code_execution, prompt caching, and Nano Banana image generation (#5720)
* Studio: add Gemini provider with web_search, code_execution, prompt caching, and Nano Banana image generation

Wires Google's native Gemini API into Studio's external-provider stack
so users can pick gemini-2.5-pro / gemini-2.5-flash / gemini-2.5-flash-image
(Nano Banana) alongside the existing OpenAI / Anthropic / OpenRouter
providers. Gemini does not speak OpenAI Chat Completions on its primary
endpoint; the new `_stream_gemini` async generator translates between
the two shapes the same way `_stream_anthropic` handles the Messages API.

Backend:
- New `_stream_gemini` translator in external_provider.py. Converts
  OpenAI messages -> Gemini `contents` + `systemInstruction`; maps
  generationConfig (temperature / topP / topK / maxOutputTokens);
  forwards `tools: [{googleSearch: {}}]` for web_search and
  `{codeExecution: {}}` for code_execution; passes `cachedContent`
  through for prompt caching; sets `responseModalities=[TEXT, IMAGE]`
  for Nano Banana image generation.
- Translates streamed `GenerateContentResponse` SSE frames back into
  OpenAI chat.completion.chunk frames (text deltas, function_call ->
  tool_calls deltas, inlineData -> image_b64 tool_end envelope, usage
  chunk before [DONE]).
- Registry entry switched to native base URL
  `https://generativelanguage.googleapis.com/v1beta` with
  `openai_compatible: False` and the `x-goog-api-key` auth header.
  Model lineup curated to current 2.5 / 2.0 family + Nano Banana.

Frontend:
- Provider-capability matrix: Gemini supports temperature, top_p, top_k,
  presence_penalty (matches generationConfig); min_p / repetition_penalty
  hidden because the API does not accept them.
- `providerSupportsBuiltinWebSearch` / `providerSupportsBuiltinCodeExecution`
  / `providerSupportsBuiltinImageGeneration` extended for Gemini.
- Prompt caching toggle now also lit on Gemini.

Tests:
- 21 new tests in `test_gemini_provider.py` using httpx.MockTransport.
  Cover request body shape conversion, URL/header wiring, web_search
  forwarded as googleSearch, function-call translation both directions,
  prompt caching passthrough, image generation emitting image_b64,
  grounded-search citations -> tool_end, finish_reason mapping, and
  vision data URL -> inlineData translation.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: forward presence_penalty to Gemini and recover function name from tool_call_id

Two follow-up fixes for the Gemini provider:

  * Thread presence_penalty into _stream_gemini and set
    generationConfig.presencePenalty when non-zero. The OpenAI-side
    capability matrix already exposes the slider for Gemini, so the
    value was being collected and silently dropped on the way out.

  * When an OpenAI role=tool message omits 'name' and only carries
    'tool_call_id', recover the function name from the matching
    functionCall on the prior assistant turn. Gemini 400s on an empty
    functionResponse name.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: surface Gemini code execution parts as code_execution tool events

The Gemini stream parser only handled text/functionCall/inlineData
parts, so when the user toggled the Code pill on a Gemini model the
sandbox output (executableCode + codeExecutionResult parts) was
dropped on the floor while adjacent text reached the UI. Reviewers
flagged this as the headline feature being silently broken.

Translate both parts into the existing code_execution tool envelope
that CodeExecutionToolUI already consumes for OpenAI / Anthropic:

  * executableCode  -> tool_start with kind=code_execution and the
    source code under arguments.code. We mint a tool_call_id and
    stash it so the matching result block can pair to it.
  * codeExecutionResult -> tool_end on that id with the stdout under
    result. Non-OK outcomes (OUTCOME_FAILED / OUTCOME_DEADLINE_EXCEEDED)
    are prefixed onto the text so the failure is visible.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: native Gemini model catalog, function-call ids, and honest cache claim

Three follow-ups to the Gemini provider PR after the codex pass:

  * list_models() now translates Gemini's native /v1beta/models
    payload ({models[{name, baseModelId, displayName,
    supportedGenerationMethods}]}) into the OpenAI-compatible shape
    Studio expects. Without this the picker stayed empty for Gemini
    and fell back to hardcoded defaults. Embedding-only models are
    filtered out.

  * Forward the OpenAI tool_call id into Gemini's functionCall.id
    and mirror it onto functionResponse.id. Two parallel calls to
    the same function name can now be paired unambiguously on the
    follow-up turn.

  * Drop Gemini from the prompt-caching capability set. The wire
    flow requires a separate cachedContents POST first and the
    boolean Studio emits today is a no-op; the toggle should not
    advertise a feature it cannot apply. Leaves a pointer to the
    docs for the eventual two-step orchestration.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: distinct tool_calls index per emitted Gemini function call

Codex flagged that the Gemini stream parser hardcoded
tool_calls[0].index to 0 on every emitted functionCall. OpenAI
reassemblers key tool_calls by index when joining deltas, so two
parallel function calls in one assistant turn collapsed onto a
single slot and the second call's arguments overwrote the first.

Track the running count via len(emitted_function_call_ids) - 1
and emit it as the per-call index. The dedupe guard above (skip
when fc_id already in the set) means the index is monotonic and
stable for the lifetime of the stream. Regression test asserts
[0, 1] across two parallel calls in one candidate parts list.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: surface Gemini 3.5/3.1/3 + Nano Banana 2/Pro and plumb thinking budget

`gemini-2.0-flash` / `gemini-2.0-flash-exp` were retired by Google in 2026
(`/v1beta/models/gemini-2.0-flash:streamGenerateContent` returns HTTP 404
"no longer available to new users"), and the picker had nothing past the
2.x family. Verified against the live ListModels catalog: drop the retired
ids from `default_models` + allowlist and surface the chat-capable
3.5 / 3.1 / 3 families plus the Nano Banana image trio.

Also plumb `enable_thinking` / `reasoning_effort` into Gemini's
`generationConfig.thinkingConfig`. Without this, Gemini 3.5 Flash,
gemini-pro-latest, and the 3.x previews silently spend the caller's
`max_tokens` budget on hidden "thoughts" before emitting any visible
answer -- the chat shows a truncated stub like "The capital of" and
streams stop. Mapping:
  - enable_thinking=False / reasoning_effort=none -> thinkingBudget=0
    (Flash tier; Pro tier coerces to a small positive budget because
    the API 400s on 0 with "This model only works in thinking mode")
  - minimal/low/medium/high -> 512/2048/8192/24576 budget tokens
  - max/xhigh -> -1 (dynamic)
  - default (neither knob set) -> thinkingConfig omitted, model decides

Frontend `getExternalReasoningCapabilities` now surfaces a
`reasoning_effort` picker for every Gemini chat id (Pro tier hides the
"none" option; image-tier ids stay knob-less). Adds 6 unit tests
covering Flash/Pro effort mapping, the off-toggle coercion on Pro,
default omission, and the nano-banana-pro-preview alias routing
through the image modalities path. 28 -> 34 tests in
`test_gemini_provider.py`, all green; full backend suite still passes
(1459/1460; the unrelated test_help_output flake is pre-existing and
not in any file this PR touches).

Live verification against generativelanguage.googleapis.com on
2026-05-24 with `_stream_gemini` directly:
  text   gemini-3.5-flash           single PASS  multi PASS
  text   gemini-3.1-pro-preview     single PASS  multi PASS
  text   gemini-3.1-flash-lite      single PASS  multi PASS
  text   gemini-3-pro-preview       single PASS  multi PASS
  text   gemini-3-flash-preview     single PASS  multi PASS
  text   gemini-2.5-pro             single PASS  multi PASS
  text   gemini-2.5-flash           single PASS  multi PASS
  text   gemini-2.5-flash-lite      single PASS  multi PASS
  text   gemini-flash-latest        single PASS  multi PASS
  text   gemini-flash-lite-latest   single PASS  multi PASS
  text   gemini-pro-latest          single PASS  multi PASS
  image  gemini-2.5-flash-image     PASS (1082 KB png returned)
  image  gemini-3.1-flash-image-preview  PASS (Nano Banana 2)
  image  gemini-3-pro-image-preview      PASS (Nano Banana Pro)
  tool   web_search                 PASS
  tool   code_execution             PASS
  -> 16/16 e2e through the actual ExternalProviderClient code path.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: tighten Gemini provider after review (PR #5720)

Fixes a batch of bugs surfaced by a second-pass review on top of the
3.5/3.1/3 + Nano Banana 2/Pro additions in c6724dbd.

Backend (external_provider.py):
- Constructor normalises legacy /v1beta/openai base URLs to /v1beta so
  Gemini providers saved before the native switch keep working without
  a manual re-config.
- Skip thinkingConfig, googleSearch, and codeExecution on image-tier
  models (-image / nano-banana). The image responseModalities path is
  mutually exclusive with text-tool wiring and stale UI state would
  otherwise 400 the turn.
- _PRO_THINKING_PREFIXES now includes gemini-3.5-pro and uses anchored
  prefix matching (exact id or "<prefix>-...") so the image-tier
  gemini-3-pro-image-preview cannot accidentally match the pro guard.
- Gemini 3 functionCall thoughtSignature is round-tripped through the
  tool_calls envelope via extra_content.google.thought_signature on
  emit, and replayed as a sibling of functionCall on the next request.
- finishReason swaps STOP -> tool_calls when any functionCall was
  emitted on the same turn so OAI clients trigger tool execution
  (matches the OpenAI Chat Completions contract).
- usageMetadata.thoughtsTokenCount is rolled into output_tokens and
  surfaced on output_tokens_details.reasoning_tokens so total_tokens
  reflects the full billable spend instead of dropping the hidden
  reasoning slice.

Registry (providers.py):
- Drop gemini-3-pro-preview from default_models. Google shut it down
  on 2026-03-09 and auto-redirects to gemini-3.1-pro-preview; we
  surface the canonical id only.
- Add model_id_deny_exact = ("gemini-3-pro-preview",) so the live
  ListModels fetch does not re-surface the redirect alias.

Route schema (models/inference.py):
- enable_prompt_caching widened to Optional[Union[bool, str]] so the
  /v1/chat/completions caller can pass a Gemini cachedContent resource
  name (e.g. cachedContents/abc123). Without this widening _stream_gemini
  s string cachedContent passthrough was unreachable from the public
  route (bool_parsing 422). stream_chat_completion signature mirrors.

Frontend (provider-capabilities.ts, chat-page.tsx, chat-adapter.ts):
- providerSupportsBuiltinImageGeneration now also recognises
  nano-banana ids (nano-banana-pro-preview was hidden from the image
  pill before).
- providerSupportsBuiltinWebSearch takes the model id so Gemini image
  models hide the Search pill (mirrors the backend skip).
- providerSupportsBuiltinCodeExecution uses the same isGeminiImageModel
  guard for nano-banana ids.
- GEMINI_THINKING_PRO_PREFIXES gains gemini-3.5-pro; gemini-3-pro
  tightened to gemini-3-pro-preview to avoid the image-id overlap.
- Updated 3 callers of providerSupportsBuiltinWebSearch to thread the
  selected model id through.

Tests (test_gemini_provider.py): 34 -> 42, all green
- test_image_models_skip_thinking_config
- test_image_models_drop_text_only_tools
- test_gemini_35_pro_recognized_as_pro_thinking
- test_legacy_openai_base_url_normalized
- test_finish_reason_swaps_to_tool_calls_when_function_call_emitted
- test_thought_signature_round_trips_into_gemini_function_call
- test_thought_signature_emitted_in_tool_call_delta
- test_usage_chunk_includes_thoughts_tokens

Verification:
- Backend pytest 1518/1519 passing (one unrelated Qwen3.5 flash-attn
  test fails on main as well; nothing in this PR touches that path).
- Frontend npx tsc -b clean.
- Live e2e 16/16 against generativelanguage.googleapis.com through the
  patched _stream_gemini code path (all 11 chat models single + multi
  turn, all 3 image models returned image bytes, web_search and
  code_execution tools both emit the expected envelope).
- Live /api/providers/models against the patched backend surfaces 16
  ids (gemini-3-pro-preview correctly filtered via deny_exact).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: address second-pass review findings on Gemini (PR #5720)

Round-2 reviewer.py flagged a phantom web_search card on image
turns (12/12 reviewers), route-layer stripping of tool_calls /
tool_call_id / name, an over-narrow image-mode tool guard, and
silent safety blocks. This patch fixes all four.

Backend (external_provider.py):
- web_search_active is now derived from the outbound tools_array
  (whether googleSearch was actually forwarded), not the raw
  enabled_tools intent. Image-mode turns dropped the tool above so
  the inbound stream no longer emits a phantom "search complete"
  tool_start / tool_end on those turns.
- text_tools_allowed now uses is_image_model (covers both `-image`
  / `nano-banana` picker models AND text models that requested
  `image_generation` via enabled_tools). Verified against the live
  Gemini API which rejects both googleSearch and codeExecution
  alongside responseModalities=["TEXT","IMAGE"] with explicit 400s
  ("Search as tool is not enabled for this model", "Code execution
  is not enabled for this model").
- promptFeedback.blockReason is surfaced as a 400 content-filter
  error chunk instead of returning an empty successful assistant
  response. The streaming loop closes the response before exiting.

Route (routes/inference.py):
- _build_external_messages now propagates tool_calls (assistant),
  tool_call_id, and name (tool result) through every code path
  (string content, multimodal content, non-vision fallback). Without
  this Gemini 3 function-call round trips lost their thoughtSignature
  + tool_call_id at the route boundary, and functionResponse.name
  arrived empty on the second turn.
- Assistant messages with content=None and tool_calls populated are
  preserved as a synthetic empty-string content turn so the
  Gemini translator can rebuild the functionCall part.

Tests (test_gemini_provider.py): 42 -> 45, all green
- test_image_models_suppress_phantom_web_search_card
- test_image_generation_tool_drops_text_tools
- test_prompt_feedback_block_reason_surfaces_as_error

Verification:
- Backend pytest 1736 / 1736 (the two pre-existing unrelated fails
  on main, test_help_output and Qwen3.5 flash-attn pin, are skipped).
- Frontend npx tsc -b clean.
- Live e2e 16/16 against generativelanguage.googleapis.com:
  11 chat models single + multi turn, 3 image models returning
  image bytes, web_search and code_execution both PASS.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: fix third-pass Gemini findings (PR #5720)

Round 3 review follow-ups:

Backend (studio/backend/core/inference/external_provider.py):
- Close response AND aiter_lines iterator in a finally so normal,
  prompt-block, and cancellation exits all clean up (eliminates the
  RuntimeWarning about aclose never being awaited).
- Pair the synthetic web_search tool_start with a tool_end on the
  promptFeedback.blockReason path so the UI does not leave a stuck
  "searching..." spinner after the error toast.
- Preserve native id and thoughtSignature on executableCode and
  codeExecutionResult tool events under google.native_part, and pair
  the tool_end on the code-exec id so multi-turn code-execution
  replays do not lose Gemini-required history.
- Carry part-level thoughtSignature on text deltas via
  delta.extra_content.google.thought_signature and on inline image
  tool_end via google.thought_signature so Gemini 3 image editing
  and tool turns round-trip the signature on the next request.
- Guess remote image_url MIME from the URL path so PNG / WebP / GIF
  inputs are not silently relabeled as JPEG.
- Roll usageMetadata.toolUsePromptTokenCount into translated input
  tokens and surface thoughtsTokenCount as
  completion_tokens_details.reasoning_tokens in _build_usage_chunk.
- Only normalize the Google-hosted /v1beta/openai legacy base URL;
  custom proxies whose paths happen to end in /openai are left
  untouched.
- Forward ChatCompletionRequest.tools and tool_choice through
  stream_chat_completion into _stream_gemini, translating to
  tools[].functionDeclarations and toolConfig.functionCallingConfig.

Frontend:
- chat-adapter: when Gemini image-generation is enabled for the turn,
  also disable Search and Code so the request, builder, and active
  pills agree with what the backend actually sends (the backend
  already strips text tools when image_generation is in enabled_tools).
- chat-adapter: consume OpenAI-shape delta.tool_calls chunks so
  Gemini function-call deltas without text surface as tool-call parts.
- shared-composer: disable Search and Code pills while Gemini image
  mode is active so the UI matches the request.

Tests (studio/backend/tests/test_gemini_provider.py): adds coverage
for proxy base-url gating, remote image MIME inference,
toolUsePromptTokenCount, reasoning_tokens propagation, prompt-block
web_search tool_end pairing, native code-exec id/thoughtSignature
metadata, inline image thoughtSignature, text-chunk extra_content,
OpenAI tools/tool_choice translation, and image-model tool drop.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: Gemini 3 thinkingLevel + image-model Search grounding (PR #5720)

Gemini 3.x migrated to a string `thinkingConfig.thinkingLevel`
(MINIMAL/LOW/MEDIUM/HIGH) and rejects `thinkingBudget`+`thinkingLevel`
in the same request. Gemini 3 also cannot turn thinking fully off, so
the lowest position is "minimal" (Flash) or "low" (Pro rejects
"minimal").

- external_provider._stream_gemini: split thinking translation by
  family. Gemini 3.x (3 / 3.1 / 3.5 + gemini-pro-latest /
  gemini-flash-latest / gemini-flash-lite-latest) emits
  thinkingConfig.thinkingLevel; effort none/off coerces to "low" on
  Pro and "minimal" on Flash. Gemini 2.5 stays on thinkingBudget.
- external_provider._stream_gemini: allow `tools: [{googleSearch: {}}]`
  on the Gemini 3 image family (gemini-3-pro-image-preview,
  gemini-3.1-flash-image-preview, nano-banana-pro). Google's docs
  document Search grounding on these. codeExecution stays blocked
  on image mode (still mutually exclusive with responseModalities).
- provider-capabilities.ts: mirror the Gemini 3 effort ladders in
  resolveGeminiReasoningCapabilities (Pro: low/medium/high; Flash:
  minimal/low/medium/high; 2.5 Flash keeps the off-position).
- provider-capabilities.ts: providerSupportsBuiltinWebSearch now
  returns true on the documented Gemini 3 image models so the pill
  is reachable; older image ids (gemini-2.5-flash-image) still hide.

Tests: splits the existing thinkingBudget cases by family (Gemini 3
checks thinkingLevel; Gemini 2.5 keeps thinkingBudget), adds positive
googleSearch coverage for Gemini 3 image models and negative
googleSearch coverage for legacy image models.

References:
- https://ai.google.dev/gemini-api/docs/thinking
- https://ai.google.dev/gemini-api/docs/gemini-3
- https://ai.google.dev/gemini-api/docs/models/gemini-3-pro-image-preview

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: attach Gemini code_execution inline images to the code card (PR #5720)

When a text Gemini turn wires codeExecution and the sandbox produces a
matplotlib plot, the inline image part ships right after the
codeExecutionResult. Previously this surfaced as a separate empty
image_generation card. Track the most recent code_execution
tool_call_id + result text and, when an inline image follows with
code_execution active, emit a second tool_end on the same id that
appends the image as a data: URI under the `__IMAGES__:` marker the
chat-adapter already understands.

Image-picker turns (`-image` / `nano-banana`) keep the standalone
image_generation envelope so Nano Banana outputs render the same way.

Tests: covers the merged code-execution card emission with no
standalone image_generation event when code_execution is the active
tool.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: fix fourth-pass Gemini findings (PR #5720)

Round 4 review follow-ups:

Backend:
- `_is_openai_compatible` + `_auth_headers` detect Gemini connections
  pointed at a custom OpenAI-compatible proxy (non-Google host whose
  path ends in `/openai`) and route them through the OpenAI-compat
  surface with `Authorization: Bearer ...` instead of the native
  `_stream_gemini` translator + `x-goog-api-key`. Google-hosted Gemini
  keeps the native dispatch path it migrated to in this PR.
- `_stream_gemini` thinkingLevel handling for Gemini 3 Pro now coerces
  both "minimal" and "medium" effort to "low" / "high" respectively
  (Pro tier only accepts low/high per
  https://ai.google.dev/gemini-api/docs/thinking).
- `providers.py` `default_models` restores the advertised
  `gemini-3.5-pro` and the rolling `gemini-pro-latest` /
  `gemini-flash-latest` / `gemini-flash-lite-latest` aliases that the
  allowlist already admits.

Frontend:
- chat-adapter: lean on `providerSupportsBuiltinWebSearch` (which
  already encodes the Gemini 3 image-model Search allowance) instead
  of blanket-disabling Search whenever Gemini image mode is active.
  Code execution stays blocked because Gemini image mode rejects it.
- shared-composer: mirror the same gate -- only the Code pill is
  unconditionally disabled in Gemini image mode; the Search pill is
  driven by `supportsBuiltinWebSearch`.
- provider-capabilities: Gemini 3 Pro reasoning levels now expose only
  "low" and "high" (no Medium pill) to match the API.

Tests: covers the Gemini 3 Pro medium / minimal coercion, the custom
proxy OAI-compat dispatch + Authorization Bearer auth, and the
native-vs-proxy detection. Also closes the mocked httpx.AsyncClient
inside the test event loop so the Python 3.13 `aclose was never
awaited` warning no longer fires.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: fix fifth-pass Gemini findings (PR #5720)

Round 5 review follow-ups:

Backend:
- `_is_openai_compatible` + `_auth_headers` now treat ANY non-Google
  Gemini base URL as OpenAI-compat (LiteLLM / custom OAI gateways /
  OpenAI-compat vLLM routers), not just paths ending in `/openai`.
  Pre-existing saved Gemini proxies on `/v1` keep working.
- Gemini 3 thinkingLevel coercion narrowed to the documented
  inconsistencies: only "minimal" is coerced to "low" on Pro tier.
  "medium" passes through (Gemini 3.1 Pro accepts it per
  https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/gemini/3-1-pro).
- `_stream_gemini` only flips `responseModalities=[TEXT,IMAGE]` when
  the selected model is image-capable. A stale
  `enabled_tools=["image_generation"]` on a text model is silently
  dropped instead of producing an invalid Gemini request.
- `_stream_gemini` validates the model id against
  `[A-Za-z0-9._-]+` before URL interpolation so a model like
  `../cachedContents/x` cannot redirect the request to an unintended
  endpoint with the configured API key attached.
- Empty-text Gemini parts that still carry `thoughtSignature` emit a
  content-free delta with `extra_content.google.thought_signature` so
  Gemini 3 turns that end with a signature-only fragment do not lose
  the replay state.
- ConnectError / ReadTimeout / generic HTTPError paths in
  `_stream_gemini` now close the synthetic web_search tool_start
  with a matching tool_end before the error chunk so the UI does not
  leave a stuck "searching..." card on transport failure.
- `providers.py` default_models drop the non-existent
  `gemini-3.5-pro` (Google launched only `gemini-3.5-flash` at
  I/O 2026; Pro tier remains `gemini-3.1-pro-preview`).
- `routes/inference.py` only forwards `payload.top_k` when the caller
  explicitly set it on the request (Pydantic `model_fields_set`).
  Omitted top_k stays omitted, restoring the pre-PR behavior where
  Gemini uses its server default.
- `ChatCompletionRequest.enable_prompt_caching` adds a `mode="before"`
  validator that coerces the canonical string literals "true"/"false"
  back to bool so historical opt-out callers keep working after the
  field widened to `Union[bool, str]` for Gemini cache resource names.

Frontend:
- `providerSupportsBuiltinWebSearch` / Code / Image now accept the
  saved connection `baseUrl` and return false for custom OAI-compat
  Gemini proxies. Backend skips `_stream_gemini` for those bases, so
  native tool envelopes never reach them; hiding the pills keeps the
  request, builder, and UI consistent.
- `provider-capabilities.ts` Gemini 3 Pro effort ladder restores
  `["low", "medium", "high"]` to match Google's documented levels.
- Call sites in `chat-page.tsx` and `chat-adapter.ts` pass through
  `provider.baseUrl` so the proxy gate fires.

Tests: covers Gemini 3 Pro medium pass-through, custom proxy dispatch
on `/v1` and `/openai` bases, path-traversal model id rejection,
top_k omission when not explicit, text-model image_generation drop,
empty-text + thoughtSignature surfacing, and
enable_prompt_caching string coercion.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: fix sixth-pass Gemini findings (PR #5720)

Round 6 review follow-ups:

Frontend:
- chat-adapter `delta.tool_calls` accumulates fragments by `id` /
  `index` instead of pushing a new tool-call card per chunk. The
  standard OpenAI Chat Completions stream contract sends `id`/`name`
  on the first chunk and partial `function.arguments` on subsequent
  chunks; our previous handler parsed each fragment as a standalone
  tool call. Local llama.cpp and OAI-compat providers that stream
  fragments now reassemble into a single function-call part.
- chat-adapter also preserves `extra_content` on streamed tool-call
  deltas so Gemini 3 `thoughtSignature` survives to the next turn.
- provider-capabilities Gemini 3 Pro restores "medium" in the
  reasoning-effort ladder (Google's official Gemini API thinking
  doc lists low/medium/high for Gemini 3.1 Pro; my earlier round 4
  coercion was wrong).
- provider-capabilities orders `gemini-2.5-flash-lite` ahead of the
  broader `gemini-2.5-flash` prefix so Flash-Lite falls into the
  "no native thinking knob" branch as documented.

* Studio: round-trip Gemini tool_calls and tool results (PR #5720)

Recurring round 3-6 P1: the chat-adapter renders Gemini function-call
parts and code-execution events but `toOpenAIMessage` only serialized
text + image content, so the next turn lost the assistant
`tool_calls[]` (including Gemini 3's required
`extra_content.google.thought_signature`) and the matching
`role="tool"` result. Gemini 3 multi-turn function calling and code
execution failed validation on the second turn.

Frontend:
- types/api.ts widens OpenAIChatMessage to permit `role="tool"`,
  `tool_calls`, `tool_call_id`, `name`, and `content: null`. Adds
  OpenAIToolCallPart with `extra_content` for the Gemini round-trip.
- chat-adapter: new `toOpenAIMessages` expands an assistant turn with
  tool-call parts into [assistant w/ tool_calls + extra_content,
  role=tool result, ...]. tool result content is JSON-serialized so
  the backend translator can rebuild Gemini's `functionResponse`
  shape.
- chat-adapter outbound history now uses `flatMap(toOpenAIMessages)`
  so each assistant tool-call round-trips through the standard OAI
  shape the backend's `_stream_gemini` already understands.

* Studio: replay Gemini code_execution and image native parts on history (PR #5720)

Multi-turn Gemini history previously lost the native executableCode,
codeExecutionResult, and inlineData parts because the outbound
translator regenerated a generic functionCall for every assistant
tool_call. Stow the native dict on tool_end (frontend) and replay it
verbatim with thoughtSignature (backend) so follow-up turns preserve
the prior execution and image generation state. Skip role="tool"
fan-out for server-side builtin tools so Gemini does not 400 on a
functionResponse with no matching user-declared function.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: complete Gemini built-in tool replay round-trip (PR #5720)

Round 7 follow-up to the multi-turn native-part work. Three asymmetric
storage/consume gaps remained between the backend translator and the
chat adapter, so realistic Gemini follow-up turns degraded to generic
functionCalls instead of native history.

- Frontend collectAssistantToolCalls now drops web_search outright,
  drops code_execution / image_generation when the native part is
  missing, and promotes args.google to extra_content.google so the
  backend native_part replay branch actually fires.
- Backend image_generation tool_end now emits google.native_part
  with the inlineData (mimeType + base64) and thoughtSignature so the
  follow-up image-edit turn can replay the prior image as a native
  Gemini model part.
- Backend code-execution plot tool_end now stows google.native_part
  with the inlineData so the merged code-exec card can round-trip
  executableCode + codeExecutionResult + inlineData on the same id.
- Added regression tests for image-gen native-part replay and the
  code-exec plot native_part stow.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 8 Gemini follow-ups (PR #5720)

- Text-part thoughtSignature: stow on the assistant message during
  streaming and replay onto the last text part on the next turn so
  Gemini 3 strict function-calling does not reject history.
- Function declarations: recursively strip Gemini-unsupported OpenAPI
  keys (additionalProperties, $schema, $defs, strict, etc.) so OpenAI
  strict tools stop 400ing as INVALID_ARGUMENT on Gemini.
- OpenAI-compat fallback: forward tools/tool_choice so custom Gemini
  proxies (LiteLLM, gateways) keep function-calling.
- enable_prompt_caching: cover the Pydantic v1 legacy off/on/f/n/t/y
  string set so explicit opt-outs stay opt-out (Gemini was sending
  cachedContent: "off" otherwise).
- Frontend collectAssistantToolCalls / collectToolResultMessages: use
  google.native_part + result presence to disambiguate provider
  builtins from same-named user-declared functions.
- Added regression tests for text-signature replay and schema
  sanitization.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 9 Gemini follow-ups (PR #5720)

Two round-9 convergent finds across the 12 reviewers:

- Server-side web_search was leaking onto the next turn as a fake
  user functionCall/functionResponse. The previous heuristic (skip
  builtin only when no native_part AND no result) let it through
  because the synthetic tool card has a non-empty result string.
  Always skip web_search by name on both serializers, accept that a
  user-declared function literally named "web_search" must use a
  different name.
- Assistant `extra_content` was dropped by ChatMessage validation
  before _stream_gemini could replay text-part thought signatures.
  Add the field to ChatMessage and forward it through
  _build_external_messages so the multi-turn signature path actually
  carries data.

Includes a regression test for the ChatMessage round-trip.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 10 Gemini follow-ups (PR #5720)

Three convergent round-10 reviewer findings closed:

- Tag synthetic provider-side builtins with `args._server_tool=True`
  via a central helper that runs in every `_emit_tool_event` /
  `_emit_synthetic_tool_event` path. The frontend filter now skips
  on that marker instead of on the public tool name, so local
  llama.cpp `web_search` and OpenAI function tools literally named
  `web_search` / `code_execution` / `image_generation` round-trip
  cleanly while Gemini grounding / hosted code-exec / hosted image
  cards stay skipped.
- Gate Gemini image-mode (responseModalities=[TEXT,IMAGE]) on the
  Images pill (enabled_tools containing `image_generation`).
  Selecting an image-capable model with the pill off no longer forces
  image output the UI says is disabled.
- Frontend missing-key guard now exempts custom Gemini OAI-compat
  proxies (LiteLLM, gateways) the same way the backend already
  does, so a saved Gemini connection on `http://localhost:4000/v1`
  with no API key stops being blocked.

Existing tests updated to pass `enabled_tools=["image_generation"]`
on image-mode capture paths.

* Studio: round 11 Gemini follow-ups (PR #5720)

Four round-11 findings closed:

- Kimi _stream_kimi_web_search's local _synthetic_chunk helper now
  runs through _stamp_server_tool_marker so Kimi search history is
  not replayed as a fake user functionCall on the next turn (was an
  asymmetric miss after the round-10 tagging work).
- OpenAI Responses path (/v1/responses for gpt-5.x) forwards
  caller-supplied tools / tool_choice, translating the Chat
  Completions function-tool shape into the Responses native shape.
  Without this, standard OpenAI tools silently dropped on
  Responses-routed traffic.
- Decoupled the Gemini image-tier model-id guards (text-tool /
  thinking strip) from the Images pill flip
  (responseModalities=[TEXT,IMAGE]). gemini-2.5-flash-image with
  Search/Code on and the Images pill OFF no longer forwards
  googleSearch + thinkingConfig (Gemini 400s on those for legacy
  image ids).
- Gemini-only extra_content is now forwarded by
  _build_external_messages only when provider_type=="gemini" so
  Google's thought_signature does not leak into OpenAI / Mistral /
  Kimi / OpenRouter request bodies as an unknown field.

Added a regression test for the image-tier strict-guard split and
extended the extra_content test to cover the non-Gemini suppression.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 12 Gemini follow-ups (PR #5720)

Three round-12 convergent findings closed:

- extra_content leak to custom Gemini OAI-compat proxies (8/12
  reviewers). _build_external_messages now gates extra_content on
  the native generativelanguage.googleapis.com host, not just
  provider_type=="gemini", so LiteLLM / custom gateways routed
  through /chat/completions do not get an unknown top-level field.
- OpenAI Responses function-tool round-trip (5/12 reviewers). I
  added user `tools` forwarding in round 11 but did not parse the
  matching response.output_item.done items of type=function_call.
  The parser now translates them into Chat Completions
  delta.tool_calls and the terminal chunk reports
  finish_reason="tool_calls" when the model invoked a user
  function.
- Image-tier model with Images pill OFF (2/12). Google's image
  models default to text+image when responseModalities is omitted,
  so the previous fix silently still billed image output. Force
  responseModalities=["TEXT"] when the Images pill is off and the
  selected model is image-capable.

Updated the two pre-existing tests that pinned the synthetic-tool
arguments shape to include the new `_server_tool: True` marker, and
added a regression test for the Responses function-call output
translation.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 13 Gemini/Responses follow-ups (PR #5720)

Three round-13 convergent findings closed:

- OpenAI Responses function_call indices: my round-12 translator
  hardcoded every emitted tool_calls[*].index to 0, so parallel
  function calls collapsed for index-keyed clients. Track and
  increment function_call_index per emit (mirrors the Gemini
  branch's distinct-index pattern). 10/12 reviewers flagged.
- _SERVER_SIDE_BUILTIN_TOOL_NAMES now includes web_fetch so
  Anthropic-hosted web_fetch cards carry the _server_tool marker
  and the frontend history serializer doesn't replay them as fake
  user functions. 4 reviewers flagged.
- OpenAI Responses follow-up tool results now serialize as
  Responses-shape function_call / function_call_output items keyed
  by call_id, instead of Chat Completions role="tool" content.
  Skips assistant tool_calls tagged with _server_tool so hosted
  builtins don't round-trip as user functions. 2 reviewers flagged.

Updated the Anthropic code_execution and web_fetch test argument
pins to include the new _server_tool marker, and added two
regression tests (distinct indices on parallel function_call,
function_call_output round-trip).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 14 Gemini follow-ups (PR #5720)

Three round-14 findings closed:

- Remote `image_url` translation (5 reviewers convergent). Public
  HTTPS image URLs can't be sent as `fileData.fileUri` -- Gemini
  reserves that path for Files API URIs and YouTube. Fetch the
  bytes server-side and inline them as base64 `inlineData`,
  mirroring the pre-PR OpenAI-compat behaviour. YouTube URLs and
  generativelanguage.googleapis.com/v1beta/files/* stay as
  `fileData`.
- Nullable JSON Schema type arrays. OpenAI strict tools commonly
  use `"type": ["string", "null"]`; the Gemini sanitizer now
  flattens that to `"type": "string", "nullable": true` so strict
  function tools stop 400ing.
- Parallel functionResponses now ride on one user content block
  with multiple `functionResponse` parts, matching Google's
  parallel tool docs. Consecutive `role="tool"` messages merge
  into the previous user turn instead of splitting into separate
  Gemini user turns.

Three regression tests added (remote URL fetch + inline, Files
API / YouTube fileData preservation, schema nullable flattening,
parallel-tool grouping).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: SSRF harden Gemini remote image fetch (PR #5720)

Round 15 convergent finding (12/12 reviewers). My round-14 fix to
download user-controlled image URLs for inlineData inlining was an
SSRF / data-exfiltration path: no scheme check, no private-host
guard, no size cap, no Content-Type validation, redirects could
bounce to internal services, and the full URL was logged.

Replace the inline fetch with `_safe_fetch_image_for_gemini`:

- Require https:// (reject http, file, data, ftp, etc).
- Resolve the hostname via socket.getaddrinfo and reject if ANY
  resolved address is private / loopback / link-local / multicast /
  reserved / unspecified (covers 127.0.0.0/8, 10/8, 172.16/12,
  192.168/16, ::1, 169.254/16 metadata, RFC 6890).
- Block IP-literal URLs that resolve into those same ranges.
- Cap response body at 10 MB (Content-Length pre-check + streamed
  byte counter).
- Require Content-Type to start with `image/`.
- Disable redirect following so a 302 to a private host can't slip
  past the address check.
- Use a short 15s timeout and a tiny connection pool dedicated to
  these fetches.
- Log only the host name + error class -- no full URL, no signed
  querystring leak.

If the guard rejects, the image part is silently dropped (instead
of forwarding raw bytes or a fileData fallback). Files API URIs
and YouTube URLs still ride as `fileData.fileUri` unchanged.

Tests: replaced the live-fetch test with a `_safe_fetch_image_for_gemini`
monkeypatch, added four new SSRF-guard tests (non-https rejected,
loopback / private IP literals rejected, hostnames that resolve to
private IPs rejected).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 16 Gemini follow-ups (PR #5720)

- IP-pinned image fetch (`_safe_fetch_image_for_gemini`): reuse the
  validated-once-then-pin pattern from `tools._fetch_page_text` via
  `asyncio.to_thread`, so DNS rebinding between validation and the
  HTTP connect cannot redirect us at a private/metadata address.
  Catch malformed-bracketed IPv6 urlparse errors. Follow up to 4
  redirect hops with per-hop SSRF re-validation.
- Replace contains-substring detection of Gemini Files API + YouTube
  URLs with parsed scheme/host/path checks, so attacker URLs like
  `https://evil.example/path/youtube.com/x.png` no longer skip the
  safe-fetch path and serialize as `fileData.fileUri`.
- `_build_external_messages`: strip per-tool-call `extra_content`
  for non-native-Gemini providers; the Gemini-only
  `thought_signature` payload was leaking through `tool_calls[]`
  into /chat/completions on OpenAI, Anthropic, and custom Gemini
  OAI-compat gateways.
- `_server_tool` marker now gated on the function name being one of
  the canonical builtin names (`web_search`, `web_fetch`,
  `code_execution`, `image_generation`) AND the marker being set,
  so a user function whose schema happens to define an
  `_server_tool` field is no longer dropped. Frontend filter mirrors
  the same gate, plus a backward-compat fallback for pre-PR
  persisted server-tool cards (no marker) routed via name +
  native_part / web-tool heuristic.
- Gemini schema sanitizer collapses `anyOf: [{X}, {"type":"null"}]`
  to `{X, "nullable": true}` so Optional[X] tool args from
  OpenAI/Pydantic schemas no longer 400 the Gemini request.
- Frontend tool-result serializer emits `{"result":""}` for empty
  string outputs so the ChatMessage validator does not reject
  `role="tool"` with empty content.
- Coerce `medium` thinkingLevel to `high` for legacy
  `gemini-3-pro*` / `gemini-3-pro-preview*` (only low/high
  documented; shut down 2026-03-09); 3.1+ Pro still passes through.
- Hide Gemini native thinking ladder on custom OAI-compat Gemini
  gateways by routing `getExternalReasoningCapabilities` through
  `isGeminiCustomOpenAICompatBase(baseUrl)`; thread baseUrl through
  all four call sites.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 17 Gemini follow-ups (PR #5720)

- Frontend `collectAssistantToolCalls` and `collectToolResultMessages`
  no longer drop unmarked `web_search` / `web_fetch` cards by name
  alone: a user-defined function with one of those names must
  round-trip. Pre-PR persisted `code_execution` / `image_generation`
  cards still get filtered via a shape heuristic (kind/command/code/
  prompt fields) instead of bare name.
- `_build_external_messages._filter_tool_calls` now drops marked
  server-side builtin `tool_calls` entirely for non-native-Gemini
  providers, not just their `extra_content`. An assistant turn whose
  only payload was a marked builtin is dropped completely so the
  receiving provider does not see an orphan tool_call.
- `_stream_anthropic` translates OpenAI top-level `tool_calls` into
  Anthropic native `{type:"tool_use", id, name, input}` content
  blocks, and translates `role="tool"` follow-ups into `role:"user"`
  messages carrying a `tool_result` block. Anthropic's native
  Messages API rejects the OpenAI shapes.
- `_safe_fetch_image_for_gemini_sync` factors URL validation through
  `_safe_parse_https`, so malformed `port` access (e.g.
  `https://host:bad/x.png`) and malformed redirect targets (e.g. a
  302 to `https://[bad/x.png`) drop the image instead of raising mid-
  request.
- `tool_choice="none"` now disables hosted builtins (Gemini
  googleSearch / codeExecution and OpenAI Responses web_search /
  shell / image_generation), not just user function declarations.
- Schema sanitizer handles multi-type `anyOf` with null
  (`Union[str, int, None]`): keep the slim non-null anyOf and add
  `nullable: true` so Gemini does not reject `{"type":"null"}`.
- Image fetch falls back to the caller-provided MIME (guessed from
  URL extension) when the server omits Content-Type instead of
  dropping the image as `non-image content-type=<none>`.
- Per-request aggregate caps on remote image inlining (8 images,
  20MB total) so a single chat request cannot force unbounded
  backend downloads.
- Frontend exposes the reasoning ladder for `gemini-2.5-flash-lite`
  (`none/minimal/low/medium/high/max`) so the UI can drive the
  thinkingBudget the backend already supports.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 18 Gemini follow-ups (PR #5720)

- `tool_choice="none"` now opts out of hosted builtin tools on every
  provider path, not just Gemini and OpenAI Responses. Anthropic
  web_search / web_fetch / code_execution, Kimi `$web_search` early
  return, and OpenRouter `plugins:[{id:"web"}]` are all gated on
  `tool_choice_disabled`. Passing `enabled_tools=[...]` with
  `tool_choice="none"` no longer triggers provider-side search /
  code execution for any provider.
- `_stream_anthropic` accepts `tool_choice` and threads it through;
  the dispatcher in `stream_chat_completion` forwards it.
- Frontend `isServerSideBuiltinToolPart` simplified to drop only on
  (marker) OR (canonical name + native_part). The previous shape
  heuristic on `args.kind`/`args.command`/`args.code`/`args.prompt`
  dropped real user-declared `code_execution` / `image_generation`
  functions. Pre-PR persisted hosted cards lacking the marker now
  leak to non-native providers on switch -- preferred to silently
  deleting legitimate function-call history.
- Backend `_is_marked_server_builtin_tool_call` and the OpenAI
  Responses translator's matching filter accept BOTH `_server_tool`
  marker AND `args.google.native_part` as durable provider-side
  signals so Gemini code_execution / image_generation cards are
  still dropped on a provider switch.
- Per-request remote image count cap now counts ATTEMPTS, not just
  successful inlines, so 100 failing/slow URLs cannot each consume
  the 15s fetch timeout. Data: URL images now share the same count
  and byte caps as fetched remote URLs.
- OpenAI Responses translator tracks skipped server-builtin
  `function_call` ids and drops their matching `role="tool"`
  follow-ups, preventing orphan `function_call_output` items in the
  outbound body.
- Gemini schema sanitizer preserves multi-type unions with null:
  `{"type":["string","integer","null"]}` becomes
  `anyOf:[{string},{integer}] + nullable:true` instead of being
  flattened to the first non-null type.
- Gemini model id validation moved to the top of `_stream_gemini`
  so an invalid model id rejects the request before any remote
  image fetch / message translation side effect.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 19 Gemini follow-ups (PR #5720)

- `_build_external_messages` now skips an empty assistant turn when
  `_filter_tool_calls` drops every synthetic builtin tool_call (was
  guarded only on the `content is None` branch; the string-content
  and list-content branches still forwarded
  `{"role":"assistant","content":""}` which several providers
  reject). Also tracks the dropped server-builtin tool_call ids and
  skips the matching `role="tool"` follow-ups so the receiving
  provider does not see an orphan tool_result.
- OpenRouter `web_search_active` (the synthetic tool_start /
  tool_end emitter) is now also gated on `tool_choice_disabled` so
  a request with `tool_choice="none"` does not surface a fake
  web_search card in the chat UI even though the plugin was
  correctly stripped from the outbound body.
- `_stream_anthropic` translates an OpenAI role="tool" with list
  content (`content=[{"type":"text","text":"..."}]`) into a native
  `tool_result` block on a user message; previously only the
  string-content shape was translated, so list-content tool results
  were forwarded as invalid `role:"tool"` messages.
- Gemini `data:` URL image_url parts now require an `image/*` MIME
  type; a `data:text/html;base64,...` is dropped instead of being
  forwarded as `inlineData.mimeType="text/html"` (Gemini rejects
  the malformed image part). Symmetric with the fetched-remote
  image fetch path that already rejects non-image Content-Type.
- YouTube `fileData.fileUri` now declares `video/mp4` as the
  mimeType instead of `image/jpeg` guessed from the URL path. The
  YouTube/fileData input is the documented Gemini video path; the
  guessed image MIME made valid YouTube inputs malformed.
- OpenAI Responses translator preserves `response.output` ordering
  on assistant turns that emitted both text and a function_call:
  assistant text is now serialized BEFORE the function_call item
  so the subsequent function_call_output (the matching role=tool
  follow-up) lands in the right position. Previously the order
  was function_call -> assistant text -> function_call_output,
  which can confuse multi-turn function-calling flows.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: round 20 Gemini follow-ups (PR #5720)

Convergent reviewer findings from round 20:

- tool_choice="none" no longer flips responseModalities=[TEXT,IMAGE]
  on image-tier Gemini models. Forced-function tool_choice (e.g.
  {type:function, function:{name:lookup}}) also drops hosted Search /
  code execution from the Gemini body so the caller's pinned user
  function is not silently joined by hosted builtins.

- Gemini code-execution thoughtSignature replay now uses an ordered
  parts list (native_part.parts[]) so per-part signatures stay
  attached to the exact part Gemini emitted. The previous merged
  shape fanned one top-level thoughtSignature across executableCode
  + codeExecutionResult + inlineData and tripped Gemini 3 strict
  validators. Backward-compat fallback keeps pre-round-21 persisted
  history working: a legacy native_part with a single subpart still
  replays the signature on that subpart; merged legacy objects pin
  the signature to executableCode only.

- Remote-image fetch threads the remaining per-request byte budget
  into _safe_fetch_image_for_gemini, so over-budget URLs are
  refused via Content-Length pre-check / short read instead of
  fully downloaded then discarded after the aggregate cap check.

- Gemini role=tool with OpenAI list-form content
  ([{type:text,text:result}]) now flattens text parts before
  building functionResponse.response.result; previously the parts
  arrived as the result value instead of the actual tool output.

- Frontend chat-adapter merges native_part by concatenating parts
  lists (preserving per-part thoughtSignature). Wire types expose
  enable_prompt_caching as boolean|string (Gemini cached-content
  name) and OpenAIChatDelta now carries tool_calls and extra_content.

- Test test_openrouter_no_synthetic_web_search_event_on_tool_choice_none
  reads _toolEvent from the top-level SSE payload so a backend
  regression cannot mask the assertion.

Adds 7 regression tests covering image_generation gate, forced-function
gate, native_part list replay, legacy fallback, list-content
functionResponse flattening, fetch byte-budget threading, and wire
types.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Apply forced-function tool_choice gate to Anthropic, OpenRouter, Kimi

Previously only the Gemini path treated `tool_choice={"type":"function",
"function":{"name":...}}` as a hosted-tool opt-out. Anthropic,
OpenRouter, and Kimi still attached hosted web_search / web_fetch /
code_execution when the caller explicitly pinned a user function plus
`enabled_tools=[...]`. That contradicts the explicit function pin and
bills the caller for unwanted server-side calls.

Mirror the Gemini gate symmetrically:
  - Anthropic web_search / web_fetch / code_execution
  - OpenRouter `plugins:[{id:"web"}]` + the synthetic web_search SSE
    event the same path emits at stream close
  - Kimi `_stream_kimi_web_search` dispatch

Adds 4 regression tests:
  - test_anthropic_forced_function_tool_choice_drops_hosted_tools
  - test_openrouter_forced_function_tool_choice_drops_web_plugin
  - test_kimi_forced_function_tool_choice_skips_web_search_helper
  - test_openrouter_no_synthetic_web_search_event_on_forced_function_tool_choice

All 146 existing backend tests still pass.

* Strip Gemini-only synthetic tool history on local-GGUF dispatch

After a Gemini chat that ran code_execution / image_generation, switching
the same thread to a local GGUF model used to forward the synthetic
provider-side tool_calls (tagged with `args._server_tool` or carrying a
Gemini `args.google.native_part` payload) and the message-level
`extra_content` to llama-server. The receiving backend has no tool
declaration for those names and no use for Gemini thoughtSignature
metadata; in the worst case it can produce an orphan tool_call_id and a
confused continuation.

Add `_strip_provider_synthetic_tool_history()` and wire it through the
two local message builders:
  - `_openai_messages_for_passthrough`  (OAI-compat passthrough)
  - `_openai_messages_for_gguf_chat`    (standard GGUF chat path)

Real user-function `tool_calls` and their matching `role="tool"` replies
survive unchanged; only synthetic provider-side cards and Gemini-only
`extra_content` are stripped. If the synthetic call was the assistant
turn's only payload, the now-empty turn is dropped too so llama-server
does not reject the request.

Adds 2 regression tests:
  - test_strip_provider_synthetic_tool_history_drops_synthetic_only
  - test_strip_provider_synthetic_tool_history_drops_empty_assistant

142 existing backend tests still pass.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Disable Search/Code composer pills for Gemini image-tier models

For external Gemini image-tier models (gemini-2.5-flash-image,
gemini-3.x-image-preview, etc.), the backend unconditionally strips
code_execution and strips web_search on older image ids. Search is
still allowed on Gemini 3.x Pro/Flash image models, which
supportsBuiltinWebSearch already encodes per model.

Before this commit the composer pill gates were:
  searchDisabled = !modelLoaded || !(supportsTools || supportsBuiltinWebSearch)
  codeDisabled   = !modelLoaded || !(supportsTools || supportsBuiltinCodeExecution) || imageModeDisablesCode

`supportsTools` here is a local-runtime fallback that becomes true when
any tool-capable local model has been loaded in the session. With a
local tool-capable runtime active, switching the chat to an external
Gemini image-tier model used to leave Search/Code clickable, even
though the backend will silently drop the tool on the wire.

Detect "external provider is Gemini AND the model is image-tier" (via
supportsBuiltinImageGeneration) and gate the two pills strictly on the
provider's own builtin support in that case. Non-Gemini paths and
non-image Gemini models keep the supportsTools fallback unchanged.

* Apply forced-function tool_choice gate to OpenAI Responses path

Round 22 added the gate for Gemini / Anthropic / OpenRouter / Kimi but
missed the OpenAI Responses translator. When a caller pinned a user
function via `tool_choice={"type":"function","function":{"name":...}}`
plus `enabled_tools=["web_search","code_execution","image_generation"]`,
the Responses body still attached `{"type":"web_search"}`,
`{"type":"shell"}`, and `{"type":"image_generation"}` server tools. The
function pin should suppress those for the same privacy + billing reason
the other provider paths now do.

Compute `_responses_tool_choice_forced_function` next to
`_responses_tool_choice_none` and gate each hosted-tool append on
`_responses_hosted_builtins_allowed = not none and not forced_function`.
The fix has to be applied in TWO places: the initial body builder and
`_build_body()` (called by the container-expiry retry path). User
function declarations still flow through so the pin has something to
target, and the Responses-shape `{type:"function", name:"..."}`
`tool_choice` is forwarded unchanged.

Adds regression test `test_openai_responses_forced_function_tool_choice_drops_hosted_tools`.
All 166 existing backend tests across Gemini + Responses + image-gen +
code-exec suites still pass.

* Round 24 P1s: SSRF shared-address gap + extra_content text-only leak + custom-Gemini model list

Three convergent P1s from round 24 review:

1. SSRF: the shared SSRF validator in `tools._validate_and_resolve_host`
   used a denylist (is_private / loopback / link_local / multicast /
   reserved / unspecified). Python classifies shared address space
   (100.64.0.0/10 carrier-grade NAT, plus 240.0.0.0/4, benchmarking
   ranges, etc.) with `is_private=False` AND `is_global=False`. The new
   Gemini server-side image fetcher therefore accepts URLs whose
   hostname resolves to 100.64.0.1 in cloud/VPC deployments. Add
   `not ip.is_global` as the primary gate -- a single source of truth
   that covers every current and future non-global range.

2. _strip_provider_synthetic_tool_history previously only stripped
   message-level `extra_content` when the assistant turn had tool_calls.
   A plain text Gemini reply carrying
   `extra_content.google.thought_signature` flowed through to
   llama-server when the thread was switched to a local GGUF backend.
   Always strip message-level `extra_content` on assistant turns.

3. routes/providers.list_provider_models applied Gemini's native
   `model_id_allowlist` regex to every Gemini provider, including
   custom OAI-compatible bases (LiteLLM, deployment gateways). IDs like
   `google/gemini-2.5-flash` and team-prefixed deployment aliases got
   filtered out even though the chat-dispatch path now routes them via
   the OpenAI-compatible client. Skip registry-level model-id filters
   when the configured Gemini base_url host is not the canonical
   `generativelanguage.googleapis.com`, mirroring the chat-dispatch
   gate.

Three regression tests added:
  - test_validate_and_resolve_host_blocks_shared_address_space
  - test_strip_provider_synthetic_tool_history_drops_text_only_extra_content
  - test_gemini_custom_oai_compat_base_skips_native_allowlist

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Round 25 P1s: skip synthetic server-tool replay + inline $ref/$defs into Gemini schema

Two convergent reviewer findings on the native Gemini path:

1. _stream_gemini's tool_calls replay loop falls through to a generic
   functionCall emission whenever it sees an assistant tool_call. Marked
   server-side builtin cards (web_search / web_fetch tagged with
   _server_tool or args.google.native_part) hit that fallthrough with no
   replayable native_part, which produces an outbound functionCall whose
   name is not a declared user function. The Gemini turn 400s on the
   undeclared name. Guard the loop to drop those entries instead, while
   keeping the existing code_execution / image_generation native-part
   replay branch intact.

2. _sanitize_gemini_schema uses a strict allowlist that drops local
   $ref / $defs references. Pydantic-generated tool schemas hoist nested
   object shapes into $defs and reference them via {"$ref": "#/$defs/X"},
   so a property like address: {"$ref": "#/$defs/Address"} collapsed to
   {} on the wire and the model lost the nested fields, types, and
   required keys. Resolve local #/... pointers against the schema root
   and inline the referenced subtree, with local siblings overriding
   the reference (normal JSON Schema composition) and a seen-ref guard
   for self-referential schemas.

Added regression coverage:
- test_gemini_native_skips_synthetic_server_builtin_replay
- test_function_declarations_inline_local_refs_into_gemini_schema
- test_function_declarations_inline_local_refs_in_anyof_and_items
- test_function_declarations_self_referential_schema_terminates

All 145 Gemini provider tests pass; touched provider regression set
(OpenAI Responses, code execution, image generation, Anthropic code
execution, Anthropic web_fetch) also 43/43 green.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Round 26 P1s: drop orphan Gemini functionResponse + Anthropic /messages synthetic-history strip

Reviewer round 26 surfaced two convergent asymmetric-fix bugs.

1. _stream_gemini drops a synthetic server-tool tool_call (web_search /
   web_fetch tagged _server_tool) and also replays code_execution /
   image_generation tool_calls as Gemini-native executableCode /
   codeExecutionResult / inlineData parts. The matching role="tool"
   follow-up was still falling through to the generic functionResponse
   branch, producing either an orphan functionResponse (synthetic case)
   or a duplicate response pointing at a name with no
   functionDeclarations entry (native-part case). Both forms 400 the
   next Gemini turn. Track skipped + native-replayed tool_call_ids in
   _gemini_skip_tool_result_ids and short-circuit the role="tool"
   branch on a match.

2. The Anthropic-compatible local /v1/messages route only called
   _drop_empty_assistant_sentinels on the OpenAI-translated history,
   while the sibling /v1/chat/completions and GGUF passthrough builders
   chain that with _strip_provider_synthetic_tool_history. An Anthropic
   caller replaying a prior provider-side tool_use therefore forwarded
   fake builtin tool history straight into local llama-server. Apply
   the same strip on the Anthropic route after the
   anthropic_messages_to_openai conversion.

Regression coverage added:
- test_gemini_native_skips_orphan_function_response_for_dropped_builtin
- test_gemini_native_skips_orphan_function_response_for_native_part_replay

Gemini suite 147/147; touched provider regression set 43/43.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Round 27 P1s: native_part location fallback + Gemini image request budget for base64

Two convergent reviewer findings on the native Gemini path.

1. _stream_gemini's synthetic-builtin detector at lines 3519-3524
   recognizes args.google.native_part as a server-tool marker, but
   _native_part was only loaded from tc.extra_content.google.native_part.
   A direct OpenAI-compatible API caller or imported third-party thread
   round-trips the payload through function.arguments because
   tool_calls[].extra_content is not in the OpenAI spec. The round-25
   guard then saw a synthetic builtin with no _native_part and dropped
   the entire assistant turn, so the next native Gemini request lost
   the prior executableCode / inlineData / codeExecutionResult context.
   Fall back to args.google.native_part when extra_content path is
   missing, mirroring what the synthetic detector already accepts.

2. _GEMINI_REMOTE_IMAGE_MAX_TOTAL_BYTES capped DECODED bytes at 20MB.
   Gemini receives images base64-encoded inside JSON, and base64
   inflates payload size by ~4/3. With 20MB decoded the actual JSON
   body is ~26.7MB plus prompt overhead, well over Gemini's ~20MB
   request limit. Drop the decoded cap to 14MB so realistic multi-
   image turns stay safely under 20MB encoded.

Added regression test test_gemini_native_part_falls_back_to_args_google
covering an OpenAI-compat-shaped image_generation tool_call whose
native_part lives only in function.arguments.

Gemini suite 148/148.

* Fix TS build errors from main merge: restore imageParts + refusal return [] + cast image-edit ref

Three errors in chat-adapter.ts surfaced by the frontend tsc step after merging
main into feat/gemini-provider:

1. The Anthropic refusal early-return used main's  but
   toOpenAIMessages returns SerializedMessage[]; flip to .
2. Restore  -- the line
   was lost when removing main's conflict block from the function body.
3. selectedImageEditReference splice was inserting OpenAIChatMessage
   into a SerializedMessage[] array; the shapes differ on tool_calls.id
   nullability. Cast the reference message through unknown -- it carries
   no tool_calls, so the runtime payload is structurally compatible.

Reproduced locally with `tsc -b --pretty false` (now passes). Build
also failing in the in-repo `npm run build` step on PR CI; this commit
unblocks all 12 failing UI/API workflows.

* Tighten verbose comments in external_provider.py + chat-adapter.ts

Compress multi-line explanatory comments in the Gemini translator
and the chat adapter without changing any behaviour. All 148 Gemini
provider tests still pass; tsc --noEmit clean.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@users.noreply.github.com>
2026-05-27 06:01:24 -07:00
Daniel Han
7d3c472461
Studio: standalone Fetch pill for Anthropic web_fetch (#5742)
* Studio: surface Anthropic web_fetch as a standalone Fetch pill

web_fetch used to be silently bundled with the Search pill on the
assumption that "search returns URLs, fetch reads them" is the
typical workflow. Two problems with that:

- Anthropic bills each web_fetch invocation separately from
  web_search hits, so combining them made the per-message cost
  surface ambiguous.
- It blocked "just fetch this one URL" workflows where the user
  already knows the page they want read and does not want a search
  round-trip.

Adds:

- `webFetchToolsEnabled` to the chat-runtime-store, persisted to
  localStorage under `unsloth_chat_web_fetch_tools_enabled`, with a
  matching `supportsBuiltinWebFetch` capability flag and a
  `setWebFetchToolsEnabled` setter.
- A new Fetch pill in the chat composer, rendered next to Images and
  only when the active provider returns true from
  `providerSupportsBuiltinWebFetch` (Anthropic today). The pill
  defaults off so per-fetch billing is always a deliberate opt-in.
- chat-page bootstraps `webFetchToolsEnabled` from the same stored-
  preference fallback the other pills use.
- chat-adapter reads `webFetchToolsEnabled` directly when deciding
  whether to append "web_fetch" to `enabled_tools`, decoupling it
  from `toolsEnabled` (Search).

Backend translation is unchanged: when `enabled_tools` already
contains "web_fetch", `_stream_anthropic` appends the
`web_fetch_20250910` / `web_fetch_20260209` tool exactly as before
(test_anthropic_web_fetch.py pins the standalone-only path at
`test_web_fetch_tool_appended_to_request_body` and the combined
path at `test_web_fetch_combined_with_web_search_and_code_execution`).
Frontend tsc passes.

* ci: re-trigger after transient GitHub API HTTP flake (checkout + ggml-org release fetch)

* Studio: include web_fetch in the disabled-tool guard axis

Reviewer P1 / High on PR #5742 (codex + gemini): after introducing
the standalone Fetch pill, `disabledToolGuard` still only branched on
`webSearchEnabledForThisTurn`. With Fetch ON and Search OFF the
system prompt would tell Claude "you do not have web search or web
fetch tools in this conversation", which contradicts the actual tool
schema being sent and suppresses `web_fetch` tool calls, defeating
the standalone-fetch workflow this PR adds.

Treat search and fetch as a single "any web tool enabled" axis. The
guard only needs to warn the model when no web tool is wired in for
this turn; once either pill is on the model can pick the right one
from the tool schema. The existing `webLabel` already covers both
names, so the user-visible guard text stays accurate in every
combination.

tsc clean.

* ci: re-trigger after transient infra flake on Windows prebuilt / actions/checkout

* Studio: route web_fetch through per-model version dispatch

The web_fetch tool body in `_stream_anthropic` hardcoded
`web_fetch_20250910` instead of calling `_anthropic_web_fetch_version`,
so Opus 4.6 / 4.7 and Sonnet 4.6 missed the `web_fetch_20260209`
dynamic-filtering variant. The picker, the unit tests for it, and a
deliberate "follow-up" note in `test_anthropic_web_fetch.py` already
existed; this just threads it through the emission site.

Mirrors how web_search and code_execution are dispatched per model.
Old models still resolve to `web_fetch_20250910` and continue to work.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Shorten web_fetch comments for PR #5742

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-25 23:37:26 -07:00
Daniel Han
7d1b68079e
Studio: Anthropic fast_mode toggle and streaming refusal handling (#5715)
* Studio: add Anthropic fast_mode toggle + surface streaming refusals

Fast mode (beta `fast-mode-2026-02-01`) lets Claude Opus 4.6 and 4.7
generate output tokens up to 2.5x faster at 6x standard Opus
pricing. The toggle lives in Configuration → Provider when the
selected Anthropic model is Opus 4.6 or 4.7 and is otherwise
hidden. Backend gates the same prefixes a second time so a stale
frontend cannot make Anthropic 400 the request, and the
`fast-mode-2026-02-01` beta header is merged onto whatever other
betas the request already needed (code-execution, compaction).

Streaming refusals (`message_delta.delta.stop_reason="refusal"` on
Claude 4 models) now surface a short user-facing notice in the
assistant message before the translated OpenAI chunk emits the
existing `finish_reason="content_filter"`. Previously the chat
bubble truncated silently because the SSE stopped mid-stream with
no visible explanation. Per the upstream docs the conversation
must be reset before continuing, so the notice tells the user
exactly that.

Reference:
- https://platform.claude.com/docs/en/build-with-claude/fast-mode
- https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/handle-streaming-refusals

Tests:
- studio/backend/tests/test_anthropic_fast_mode_and_refusal.py (8 cases
  pinning fast_mode pass-through on 4.6/4.7, silent drop on Sonnet /
  Haiku / older Opus / None / False, and the refusal notice + finish
  reason on a synthetic refusal stream).

* Studio: drop refused Anthropic turns from the next request

Anthropic's streaming-refusal guidance says the refused assistant
turn must be removed or updated before the next call -- otherwise
the safety classifier keeps refusing. The PR only added a
user-visible notice; the partial assistant output (plus the notice
itself) still rode the next request via toOpenAIMessage.

Tag the refusal turn with an HTML-comment sentinel emitted alongside
the notice. The chat-adapter checks for that sentinel in
toOpenAIMessage and returns null, so the refused turn is excluded
from outboundMessages. The notice still renders in the transcript
(HTML comments don't display), so users keep the explanation.

* Studio: filter None finish_reason entries in test helper

test_refusal_maps_to_content_filter expects only ['content_filter']
in the finish_reasons list, but the post-PR refusal path emits a
user-visible content notice chunk first. Every _content_chunk
carries 'finish_reason: None' by construction; the helper was
appending those, so the assertion saw [None, 'content_filter']
instead of ['content_filter'].

None is not a finish reason -- it's just mid-stream delta noise.
Skip None values in _finish_reasons so the helper reflects what
the test names actually claim to check. Same fix applies cleanly
to the other helper usages (pause_turn test expects [] and the
sibling stop test expects ['stop'], both unaffected).

* Studio: cover Anthropic fast-mode edge cases

Adds 19 cases on top of the 9 in test_anthropic_fast_mode_and_refusal.
The base file pins the happy path; this file fills in the cliffs:

* Dated-snapshot prefix matching: claude-opus-4-7-2026-02-01 and
  claude-opus-4-6-2026-02-01 still gate fast_mode through, while
  claude-opus-4-5-2025-08-01 and claude-sonnet-4-6-2026-02-01 do not.
* Strict opt-in: a future claude-opus-4-8 or claude-opus-5 does NOT
  auto-enable fast_mode -- the prefix tuple must be bumped explicitly
  when a new family is whitelisted upstream.
* Beta-header merge: fast_mode coexists with code-execution-2025-08-25
  and compact-2026-01-12 in one comma-separated anthropic-beta header
  with no duplicates and no truncation. Pins the value to the exact
  fast-mode-2026-02-01 docs token so a typo would fail CI.
* Non-destruction: fast_mode=None produces byte-identical outbound
  body and headers to the version that omits the argument entirely.
  Same for fast_mode=False. Guarantees the upgrade path is
  non-breaking on existing Anthropic streams.
* Refusal stream ordering: the user-visible notice precedes the
  finish_reason chunk so a streaming UI paints text before flipping
  to content_filter. Refusal sentinel emitted exactly once. Notice
  rides a normal content delta chunk with finish_reason still null.
  Partial assistant deltas survive before the notice.
* Provider-side refusal coverage: a refusal on Sonnet (not just Opus)
  still emits the notice + sentinel + content_filter mapping, since
  refusal handling is not gated on fast-mode capability.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Persist fastMode, drop refused user message on retry

Two follow-ups on #5715:

1) sanitizeInferenceParams stripped fastMode. fastMode is in
   PERSISTED_INFERENCE_PARAM_KEYS but the storage sanitizer only kept
   numeric fields plus systemPrompt and trustRemoteCode, so the new
   toggle was silently dropped on reload and on the
   /api/chat/settings round-trip. Save it the same way trustRemoteCode
   is saved.

2) Refusal recovery now also drops the triggering user turn.
   Returning null from toOpenAIMessage on the assistant side left the
   user prompt that caused the refusal in the outbound history, so
   the very next request would re-trigger the same classifier.
   Anthropic's refusal-handling guidance is explicit on this: remove
   the refused turn AND the user message that triggered it before
   the next call. Implemented via a pre-pass that pops the trailing
   user message when an assistant carries the refusal sentinel.

Typecheck clean.

* Studio: out-of-band refusal signal + fast-mode prefix/usage/pricing fixes

The text sentinel for the Anthropic refusal drop signal was spoofable:
any assistant message containing the literal
<!--studio:anthropic-refusal--> would prune the prior user + assistant
pair on the next request. Move the signal onto a separate _toolEvent
chunk that the chat adapter latches into
assistant.metadata.custom.anthropicRefusal; assistant text can no
longer control the pruner.

Tighten the fast-mode model gate (backend + frontend) to require a "-"
family boundary so claude-opus-4-70 / claude-opus-4-7b style IDs do
not get speed: "fast" on a naive startswith match.

Use survivingMessages for the image / audio attachment scan so a
refused user turn does not gate or mis-attribute the next non-refused
turn.

Propagate Anthropic usage.speed onto the OpenAI-style usage chunk and
apply the documented 6x fast-mode multiplier in the cost calculator
(stacks with prompt-cache multipliers per the docs); expose the new
multiplier on the pricing snapshot for the UI tooltip.

Tests cover the tool-event chunk shape, the prefix-collision rejects,
usage.speed propagation, the 6x pricing math, and that the visible
refusal text carries no embedded sentinel.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Shorten fast-mode and refusal comments for PR #5715

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-25 23:37:12 -07:00
Daniel Han
9a737facaf
Studio: wire Anthropic server-side context compaction (#5686)
* Studio: wire Anthropic server-side context compaction

Anthropic ships server-side context compaction as a beta
(`compact-2026-01-12`). When the rendered prompt crosses the
configured input-token threshold, Anthropic runs an extra LLM pass
that summarises older turns and the request continues against the
compacted prefix. The response carries the original top-level fields
plus a new `context_management` block (with `applied_edits`) and
`usage.iterations[]` accounting per pass.

Per the docs the feature is currently supported on Opus 4.6, Opus 4.7,
Sonnet 4.6, and Mythos preview. The minimum threshold is 50k tokens;
under-50k requests 400.

Changes:

- Add prefix gate + helper `_anthropic_supports_compaction` plus
  constants `_ANTHROPIC_COMPACTION_PREFIXES`, `_ANTHROPIC_COMPACTION_BETA`,
  `_ANTHROPIC_COMPACTION_TYPE`, `_ANTHROPIC_COMPACTION_MIN`.
- Add `compaction_threshold: Optional[int]` to ChatCompletionRequest
  (50k ge bound, 2M le bound). Thread through `routes/inference.py`
  -> `stream_chat_completion` -> `_stream_anthropic`.
- In `_stream_anthropic`, when threshold is set AND the model
  accepts compaction, attach `context_management.edits[{type:
  "compact_20260112", trigger:{type:"input_tokens", value:N}}]` to
  the outbound body. Sub-50k values are clamped up to 50k to keep
  the request well-formed.
- Refactor the anthropic-beta header builder to merge any combination
  of `code-execution-2025-08-25` + `compact-2026-01-12` flags into
  one header value. Unrelated betas added at the registry level still
  pass through.
- Add `test_anthropic_compaction.py` with 16 cases: gate matrix
  (every doc-listed model), correct body shape, threshold clamping,
  beta header merge with code execution, silent no-op on unsupported
  models, omitted-threshold pass-through.

Live verified end-to-end against the real Anthropic API:
`compact_20260112` accepted on Opus 4.7, response carries
`context_management.applied_edits` + `usage.iterations[]` as
documented. (The first WebFetch-summarised version of these docs
suggested `compact_20260120`; the actual API only accepts
`compact_20260112`, matching the beta-header date. Worth pinning
behind a test so a future doc update can't drift back.)

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: drop ge=50_000 clamp + parse usage.iterations[]

Two reviewer follow-ups on the compaction PR:

1. Pydantic ge=50_000 on compaction_threshold was dead code.
   FastAPI rejected sub-50k threshold values with a 422 before the
   `max(int(...), _ANTHROPIC_COMPACTION_MIN)` clamp in
   _stream_anthropic could ever fire. Relaxed the floor to ge=1 so
   the in-helper clamp actually does its job; the schema comment
   now explains why this is intentional. Added a regression test
   that posts a value of 1 and 49_999 through the real request
   schema.

2. Anthropic publishes per-iteration token counts in
   `usage.iterations[]` whenever a fresh compaction has run, and
   the top-level input_tokens / output_tokens cover only the
   `message` iteration -- billing must add the compaction
   iterations on top. Aggregate compaction iteration tokens into
   `last_usage["compaction_input_tokens" / "compaction_output_tokens"]`
   so the cost surface (PR 5690) can read them without re-walking
   the array, and surface both figures in the closing stream
   summary log. Added two tests: one that pins the aggregation on a
   compacted turn and one that pins `None` when no fresh
   iterations land (so re-applied compaction blocks don't double-bill).

Sourcing: https://platform.claude.com/docs/en/build-with-claude/compaction

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: round-trip Anthropic compaction blocks across turns

Codex P1: once context_management is enabled and Anthropic runs
server-side compaction mid-stream, the response carries a
`{type:"compaction", content:"<summary>"}` content block on the
assistant message. The translator only handled text_delta and
input_json_delta on content_block_delta, so the compaction block
was silently dropped. Worse, the request schema's ContentPart
discriminated Union didn't accept `type:"compaction"`, and
_build_external_messages didn't pass it through, so even a
hand-crafted assistant message carrying the block would 422 at
parse time. Net result: Anthropic re-compacted from scratch on
every subsequent turn, wasting input tokens and reasoning budget.

End-to-end backend wiring of the round-trip:

1. SSE translator. _stream_anthropic now tracks a `current_compaction`
   state slot. content_block_start with type=="compaction" seeds it
   (Anthropic may include the summary on the start event AND/OR
   stream it via text_delta events on the same block index --
   handle both). text_delta inside a compaction block routes into
   the compaction buffer instead of the user-visible content
   stream, since the summary is opaque internal state, not
   assistant prose. content_block_stop emits a `compaction_block`
   tool_event carrying the full summary so the chat-adapter can
   persist it. compaction_blocks_seen is surfaced in the closing
   summary log.

2. Pydantic schema. Added CompactionContentPart with Tag("compaction")
   on the ContentPart Union so requests carrying the block parse
   cleanly. Required `content` field with a docstring pointing at
   the Anthropic docs.

3. Message builder. _build_external_messages forwards compaction
   parts on both vision and non-vision paths; the per-provider
   stream helper decides whether to forward to the wire (Anthropic
   does; other providers ignore the part). When a non-vision route
   ends up with a single text part, collapse back to a string
   so providers that don't accept content arrays still get the
   expected shape.

4. _stream_anthropic outbound translator. {type:"compaction"} parts
   on an assistant message land on the wire verbatim. Empty/missing
   `content` is skipped so a malformed stored block can't 400
   Anthropic.

Tests added (5): stream emits compaction_block tool event with the
summary intact; user-visible content stream does NOT carry the
summary text; outbound body forwards compaction parts verbatim on
the next turn; Pydantic schema accepts the part; builder passes
it through on both vision and non-vision provider routes.

Frontend follow-up: the chat-adapter needs to persist the
compaction_block tool_event onto the stored assistant message so
turn N+1 includes it in payload.messages. Pinned in the PR
description.

Sourcing: https://platform.claude.com/docs/en/build-with-claude/compaction

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: gate compaction-part passthrough to Anthropic only

Codex P1: my previous round-trip change preserved {type:"compaction"}
parts on every provider route in _build_external_messages. That
meant a chat history with prior compaction state silently leaked
the Anthropic-specific block to OpenAI/DeepSeek/Mistral/Gemini/
Kimi/OpenRouter on a provider switch, where generic
/chat/completions passthrough hands the unknown content type to
the upstream API and 400s the whole turn.

Added a `provider_type` kwarg to _build_external_messages and
gated the compaction forwarder on `provider_type == "anthropic"`.
Every other value (including the legacy None for callers that
don't pass it yet) strips the part. The Anthropic stream helper
still maps it to a native `compaction` block on the wire.

Threaded provider_type through from _proxy_to_external_provider's
call site.

Tests updated: vision + provider="anthropic" still forwards; six
non-anthropic providers strip the part; missing provider_type
strips defensively; non-vision + anthropic still forwards; non-vision
+ non-anthropic collapses back to a text string.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:19:09 -07:00
Daniel Han
a2d2b7866f
Studio: wire Anthropic web_fetch server-side tool (#5671)
* Studio: wire Anthropic web_fetch server-side tool

Studio's Anthropic passthrough only forwarded web_search and
code_execution when enabled_tools was set. Asking Claude through Studio
to fetch a URL produced no fetch (the tool was not in the outbound
tools array), so users had to fall back to web_search even when they
already had the exact URL they wanted.

This change opts in web_fetch_20250910 when enabled_tools contains
"web_fetch". The new tool entry is appended alongside any existing
web_search / code_execution entries:

  {"type": "web_fetch_20250910", "name": "web_fetch", "max_uses": 5}

No anthropic-beta header is required (web_fetch is GA); the existing
code-execution-2025-08-25 flag continues to merge cleanly when both
tools are enabled in the same turn.

SSE translation mirrors the web_search path. A `server_tool_use` block
with name="web_fetch" emits a `tool_start` _toolEvent carrying the
URL the model asked to fetch; the matching `web_fetch_tool_result`
block emits a `tool_end` _toolEvent whose result string follows the
Title / URL / Snippet shape parseSourcesFromResult on the frontend
already expects, so the source pill renders identically. Error blocks
(`web_fetch_tool_error`) are surfaced as "Error: <error_code>" matching
the code_execution error path.

The final "Anthropic stream complete" log line picks up web_fetch_
requested / web_fetch_invocations / web_fetch_urls so support reports
of "the model did not fetch anything" can be triaged from the log.

Verified end to end against claude-haiku-4-5 with
`enabled_tools=["web_fetch"]`: the model emitted tool_start with
url=https://example.com and tool_end with the page Title + URL +
Snippet, plus the assistant message correctly read back "Example
Domain" as the title.

Tests:
- 5 new unit tests in test_anthropic_web_fetch.py covering tool
  registration, the combined web_search + web_fetch + code_execution
  request body, the pill-off case, and SSE translation for both
  success and error paths.
- All 242 existing Anthropic + OpenAI provider tests still pass.

The enabled_tools field description in models/inference.py is updated
so OpenAPI consumers see the new option.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* web_fetch: title fallback to URL, log parse failures, drop dead checks

Three review nits on the previous commit:

1. `_format_web_fetch_result` left `title` empty when Anthropic omitted
   `document.title`. The frontend `parseSourcesFromResult` only emits
   a source pill when both `Title:` and `URL:` lines are present, so
   fetches against pages without an HTML title tag silently lost
   their citation in the UI. Fall back to `title = title or url`,
   matching the web_search formatter.

2. The broad `except Exception` around `json.loads(buffer)` for the
   web_fetch input swallowed the failure with no trace. Log at debug
   so a malformed partial_json buffer can be triaged from the server
   log without changing behavior.

3. `inner` was already sanitised to a dict at the matching
   content_block_start and `_format_web_fetch_result` always returns
   a non-empty string (defaulting to "(fetch complete)"), so the
   `isinstance(inner, dict) else {}` guard and the
   `result_text or "(fetch complete)"` fallback at the emit site
   were dead code. Removed.

Added a test exercising the titleless path so the fallback stays
covered.

* chat-adapter: emit source pills for web_fetch tool calls

`parseSourcesFromResult` was only wired up for tool calls where
`toolName === "web_search"`, so the Title / URL / Snippet block the
backend formatter emits for `web_fetch_tool_result` never reached the
source-pill renderer. Users saw the raw tool result in the tool card
but the dedicated source-pill row at the message tail stayed empty.

Both web_search and web_fetch ship the same text shape today, so the
fix is to broaden the gate.

* Address review: wire web_fetch from Search pill + fix pause_turn truncation

Two reviewer follow-ups on the Anthropic web_fetch PR:

1. The backend tool wiring landed but the frontend chat-adapter
   never put `web_fetch` in `enabled_tools`, so toggling the Search
   pill only ever attached `web_search` -- web_fetch was unreachable
   from the UI. Added providerSupportsBuiltinWebFetch() (Anthropic
   today) and paired the entry with the existing Search pill, since
   the canonical workflow is "search returns URLs, fetch reads
   them" and there is no separate UI toggle yet.

2. `pause_turn` from Anthropic's stop_reason vocabulary fell through
   the finish_reason map's "stop" default, which the OpenAI-format
   client renders as end-of-message and truncates the answer. Per
   the docs pause_turn means "Claude paused a long server-tool
   turn (web_search / web_fetch) and will resume". Mapped to None
   and skipped the chunk emission so the SSE stream still ends with
   [DONE] on message_stop but no terminal finish_reason lands on
   the client. While there: added explicit mappings for `tool_use`
   (-> tool_calls) and `refusal` (-> content_filter) which were
   also falling through to "stop".

Tests added: pause_turn emits no finish_reason, end_turn still
emits "stop", refusal maps to "content_filter".

Sourcing: https://platform.claude.com/docs/en/api/messages#response-stop-reason

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:03:48 -07:00