Commit graph

5,470 commits

Author SHA1 Message Date
pre-commit-ci[bot]
eaaf7142f6 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-27 07:41:40 +00:00
Daniel Han
64328962d0 Expand local-backend coverage further: 10 more knobs from vLLM + llama.cpp live docs (PR #5711)
Round 4 expansion driven by direct fetches of the canonical
SamplingParams + server README pages cited in the user's request.
Adds ten more knobs the docs explicitly support but the panel doesn't
surface yet:

Knob (wire name)              llama.cpp  vLLM   Ollama   Source
----------------------------- ---------- ------ -------- ----------------
skip_special_tokens           no         yes    no       vLLM SamplingParams
spaces_between_special_tokens no         yes    no       vLLM SamplingParams
include_stop_str_in_output    no         yes    no       vLLM SamplingParams
truncate_prompt_tokens        no         yes    no       vLLM SamplingParams
n_keep                        yes        no     no       llama.cpp README
n_probs                       yes        no     no       llama.cpp README
cache_prompt                  yes        no     no       llama.cpp README
return_tokens                 yes        no     no       llama.cpp README
timings_per_token             yes        no     no       llama.cpp README
post_sampling_probs           yes        no     no       llama.cpp README

Backend rationale:
  - vLLM's documented SamplingParams class at
    https://docs.vllm.ai/en/latest/api/vllm/sampling_params/ lists
    skip_special_tokens (default True), spaces_between_special_tokens
    (True), include_stop_str_in_output (False), truncate_prompt_tokens
    (None). All four are vLLM-only; llama-server's README does not
    document them and Ollama's openai/openai.go translator does not
    forward them.
  - llama-server's README at
    https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
    lists n_keep, n_probs, cache_prompt, return_tokens, timings_per_token
    and post_sampling_probs as documented per-request fields. vLLM's
    SamplingParams has no analog, and Ollama's OAI translator drops them.

Capability matrix:
  LLAMA_CPP_CAPABILITIES: 6 llama-only true + 4 vLLM-only false.
  VLLM_CAPABILITIES:       4 vLLM-only true + 6 llama-only false.
  OLLAMA_CAPABILITIES:     all 10 off (OAI translator drops all of them).
  Every other bucket:      all 10 off.

Skip-when-default rules (mirror upstream defaults):
  skip_special_tokens / spaces_between_special_tokens / cache_prompt:
    default true upstream — forward only when explicitly false.
  include_stop_str_in_output / return_tokens / timings_per_token /
    post_sampling_probs: default false — forward only when true.
  truncate_prompt_tokens / n_probs: 0 / null = unset — forward when > 0.
  n_keep: accepts -1 for "keep all", so the gate is value != 0.

Frontend:
  - ProviderCapabilities interface +10 flags.
  - InferenceParams +10 nullable fields (3 numeric + 7 boolean), all
    null in DEFAULT_INFERENCE_PARAMS.
  - OpenAIChatCompletionsRequest wire shape +10 optional fields.
  - chat-adapter forwards each in both the external (capability-aware)
    and local (capability-bypass) branches.
  - chat-settings-storage adds the 3 numeric keys to the existing
    nullable-number loop and 7 boolean keys to a new nullable-boolean
    loop (alongside ignoreEos).

Backend:
  - ChatCompletionRequest +10 Optional Fields with pydantic bounds
    (truncate_prompt_tokens ge=1, n_probs ge=0; booleans unbounded;
    n_keep accepts -1 so no lower bound).
  - llama_cpp.py three payload builders (generate_chat_stream + the
    tool-loop payload block + the final-pass stream_payload) each
    accept and forward the 10 new kwargs.
  - routes/inference.py _build_passthrough_payload accepts and forwards
    the 10; both per-request call sites (lines ~2591, ~2790) thread
    them from the request payload into the llama_cpp methods.

Test: test_local_passthrough_forwards_vllm_output_and_llama_cpp_
  instrumentation round-trips all 10 fields with explicit values
  matching each backend's upstream default and confirms each is absent
  from the body when unset.

65/65 sampling_params_routing tests pass; frontend tsc clean.

Total local-backend knob coverage now (this PR):
  Standard:    temperature, top_p, top_k, min_p, repetition_penalty,
               presence_penalty, frequency_penalty, seed, stop,
               parallel_tool_calls (10)
  llama.cpp:   typical_p, top_n_sigma, repeat_last_n, dynatemp_range,
               dynatemp_exponent, mirostat, mirostat_tau, mirostat_eta,
               dry_multiplier, dry_base, dry_allowed_length,
               dry_penalty_last_n, xtc_probability, xtc_threshold,
               min_keep, ignore_eos, min_tokens, n_keep, n_probs,
               cache_prompt, return_tokens, timings_per_token,
               post_sampling_probs (23)
  vLLM-extra:  ignore_eos, min_tokens, skip_special_tokens,
               spaces_between_special_tokens, include_stop_str_in_output,
               truncate_prompt_tokens (6)
  OpenRouter:  top_a (1)

Deferred for future PRs (require array / object field shape):
  - llama.cpp DRY sequence_breakers (string array)
  - llama.cpp samplers ordering (string array)
  - llama.cpp / vLLM logit_bias (dict)
  - llama.cpp grammar (string) + json_schema (object)
  - vLLM guided_json / guided_regex / guided_choice / guided_grammar
  - vLLM allowed_token_ids / bad_words / stop_token_ids (int / str arrays)
  - OpenAI / Ollama logprobs + top_logprobs (bool + int pairing)
  - n / best_of (need SSE multi-choice handling first)
2026-05-27 07:41:25 +00:00
Daniel Han
3674e11f07 Expand local-backend sampler coverage: DRY + XTC + min_keep + ignore_eos + min_tokens (PR #5711)
Round 3 expansion driven by direct fetches of the llama.cpp server README,
vLLM's SamplingParams source, and Ollama's openai.go OAI translator.
Adds nine new sampling/control knobs with per-backend capability gating:

Knob (wire name)       llama.cpp  vLLM   Ollama   Source
---------------------- ---------- ------ -------- -----------------------
dry_multiplier         yes        no     no       llama.cpp README
dry_base               yes        no     no       llama.cpp README
dry_allowed_length     yes        no     no       llama.cpp README
dry_penalty_last_n     yes        no     no       llama.cpp README
xtc_probability        yes        no     no       llama.cpp README
xtc_threshold          yes        no     no       llama.cpp README
min_keep               yes        no     no       llama.cpp README
ignore_eos             yes        yes    no       llama.cpp + vLLM SamplingParams
min_tokens             yes        yes    no       llama.cpp + vLLM SamplingParams

Backend-side rationale:
  - llama.cpp: full chain documented at
    https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
  - vLLM: SamplingParams source confirms ignore_eos + min_tokens; the
    other seven have no field in
    https://github.com/vllm-project/vllm/blob/main/vllm/sampling_params.py
  - Ollama: openai/openai.go FromChatRequest copies only the OpenAI
    subset (temp/top_p/seed/freq/pres/max_tokens/logprobs/topLogprobs/
    response_format/reasoning_effort) on the /v1/chat/completions path
    Studio uses. All nine new knobs are silently dropped, so the
    OLLAMA_CAPABILITIES bucket keeps them off.

Frontend:
  - ProviderCapabilities interface gains 9 boolean flags.
  - InferenceParams gains 9 nullable fields (8 numeric + ignoreEos
    boolean), all defaulting to null in DEFAULT_INFERENCE_PARAMS.
  - OpenAIChatCompletionsRequest wire shape gains 9 optional fields
    with doc comments.
  - LLAMA_CPP_CAPABILITIES: all 9 on. VLLM_CAPABILITIES: 2 on
    (ignoreEos + minTokens) via inheritance, 7 off via explicit
    override. OLLAMA_CAPABILITIES: all 9 off (inherits + overrides
    ignoreEos/minTokens). Every other bucket (openai cloud / chat,
    anthropic, gemini, mistral, kimi, deepseek, openrouter) gets all
    9 off explicitly.
  - chat-adapter.ts gates each knob in both the external (capability-
    aware) and local (unconditional-when-meaningful) branches.
    Skip-when-default rules:
      dry_multiplier > 0 unlocks the 4-field DRY chain
      xtc_probability > 0 unlocks the 2-field XTC chain
      min_keep > 0, min_tokens > 0 forward only when set higher than 0
      ignore_eos forwards only when explicitly true
  - chat-settings-storage.ts persists all 9 keys (8 numeric in the
    existing nullable-number loop, ignoreEos with its own boolean
    handler).

Backend:
  - ChatCompletionRequest gains 9 Optional Field declarations with
    pydantic ge/le bounds (dry_multiplier ge=0; dry_base ge=1; xtc_*
    ge=0 le=1; min_keep / min_tokens / dry_allowed_length ge=0).
  - llama_cpp.py: three payload builders (generate_chat_stream + the
    two payload-construction blocks inside the tool-loop stream) each
    accept the 9 new kwargs and forward via `if x is not None`.
  - routes/inference.py: _build_passthrough_payload accepts the 9 new
    kwargs and forwards into the body. Two call sites that thread
    sampler params from the request payload (lines 2581, 2771) are
    extended to forward the 9 new fields.

Test:
  - test_local_passthrough_forwards_dry_xtc_min_keep_eos_min_tokens
    round-trips all 9 fields through _build_passthrough_payload and
    confirms each is absent when unset (so llama-server / vLLM apply
    their own defaults).

64/64 sampling_params_routing tests pass; frontend tsc clean.

Deferred for future PRs (require array / object field shape):
  - llama.cpp DRY sequence_breakers (string array)
  - llama.cpp samplers ordering (string array)
  - llama.cpp / vLLM logit_bias (dict)
  - llama.cpp n_probs + OpenAI logprobs/top_logprobs
  - llama.cpp grammar (string) + json_schema (object)
  - vLLM guided_json / guided_regex / guided_choice / guided_grammar
  - vLLM allowed_token_ids / bad_words / stop_token_ids
2026-05-27 07:04:29 +00:00
Daniel Han
c22f6e48ff Apply round-2 audit fixes: per-model OpenAI caps + o-series effort + Ollama bucket (PR #5711)
Second 5-Opus reviewer round. Applying high-confidence fixes; speculative
items (gpt-5.5-pro effort restriction, o3 image_generation gating,
o-series parallel_tool_calls per-model, gpt-5.x new model prefixes,
Anthropic fast-mode + Priority exclusion UI gate, Gemini service_tier,
Kimi k2.5 toggleable thinking) deferred to follow-up because they need
type-system changes, more verification, or backend wire work.

OpenAI max-output caps — replace the 3-line table with one driven by
direct dev.openai.com per-model fetches (cross-checked against the Azure
Foundry reasoning table):

  - gpt-5.4 / gpt-5.4-pro / gpt-5.4-mini / gpt-5.4-nano: 65536 -> 128000
    (https://developers.openai.com/api/docs/models/gpt-5.4 "128,000 max
    output tokens"; Azure table same).
  - gpt-5.3-codex: 16384 -> 128000
    (https://developers.openai.com/api/docs/models/gpt-5.3-codex).
  - gpt-5 / gpt-5.1 / gpt-5.2: 32k default -> 128000
    (https://developers.openai.com/api/docs/models/gpt-5.2 confirms
    128k; Azure table extends to gpt-5/5.1).
  - gpt-5.3-chat-latest and gpt-5.1-chat keep 16384 (chat-class
    variants per Azure context table row).
  - o1 / o3 / o3-mini / o3-pro / o4-mini / codex-mini: 32k default ->
    100000 (https://developers.openai.com/api/docs/models/o3 "100,000
    max output tokens"; Azure o-series table same).

Implementation: list the two 16k chat-latest ids first so the broader
`gpt-5` 128k entry doesn't shadow them.

OpenAI reasoning_effort levels:

  - gpt-5.3-codex: drop "none" from levels + flip supportsOff to false.
    Dev page lists the enum as low/medium/high/xhigh only — `none` is
    not in the codex variant.
  - o-series bucket: change prefix from ["o3"] to
    ["o1","o3","o4","codex-mini"]. Previously o1 / o4-mini / codex-mini
    fell into NO_REASONING_CAPS so the panel HID the effort slider for
    them — real UX regression for users on those ids. Azure o-series
    table confirms all four accept low/medium/high reasoning_effort.

DeepSeek default_models:

  - Add deepseek-v4-pro + deepseek-v4-flash alongside the legacy
    deepseek-chat / deepseek-reasoner aliases. The latter retire on
    2026-07-24 per https://api-docs.deepseek.com/updates; surfacing
    both lets the picker keep working on cutover.

Local backend bucket split (Ollama-stricter):

  - Splits the round-1 VLLM_OLLAMA_CAPABILITIES into a vLLM-specific
    bucket (keeps top_k / min_p / repetition_penalty / seed on; vLLM's
    SamplingParams supports all four) and an Ollama-specific bucket
    that ALSO hides top_k / min_p / repetition_penalty. Ollama's OAI
    translator (ollama/openai/openai.go FromChatRequest) only copies
    the documented OpenAI subset on the /v1/chat/completions path that
    Studio uses; the three knobs are silently dropped even though
    native /api/chat would forward them via `options`. Hiding them is
    the smaller fix vs adding a backend /api/chat rewrite path.

Reviewer claims verified wrong, skipped:

  - _ANTHROPIC_NEW_CODE_EXEC_PREFIXES already lists opus-4-7, opus-4-6,
    sonnet-4-6 (external_provider.py:337-339). No-op.
  - Mistral `seed` already renamed to `random_seed` by backend at
    external_provider.py:772. No-op.
  - OpenRouter `isOpenRouterMandatoryReasoningModel` uses `Set.has()`
    exact match, not prefix match, so deepseek/deepseek-r1-distill-*
    cannot accidentally hit the always-on guard. No-op.

Tests: 63/63 sampling_params_routing tests pass; frontend tsc clean.
2026-05-27 06:49:15 +00:00
Daniel Han
0234bef047 Apply 5-reviewer audit fixes to per-provider capability buckets (PR #5711)
Five independent reviewers cross-checked every provider's per-model
sampling-knob exposure against live docs (OpenAI, Anthropic, Gemini,
DeepSeek, Kimi, Mistral, OpenRouter, llama.cpp, vLLM, Ollama).
Applying the high-confidence drift fixes here; speculative items (pro
model effort restrictions, gpt-5.3 cap, OpenAI verbosity / o-series
output cap, Gemini topK / service_tier) are deferred to a follow-up
because they need backend wire changes or unverified doc claims.

Anthropic:
  - Move claude-opus-4-6 from the 64k group into the 128k group (live
    legacy table shows Opus 4.6 Max output = 128k tokens).
    https://platform.claude.com/docs/en/about-claude/models/overview
  - Add claude-sonnet-4 to the 64k group (was falling through to 32k
    default; live legacy table shows Sonnet 4 Max output = 64k tokens).
  - Extend ANTHROPIC_REASONING_MODELS with legacy claude-opus-4-1 /
    claude-opus-4 / claude-sonnet-4 at none/low/medium/high (live
    legacy table marks Extended thinking = Yes for all three).

OpenAI:
  - Split the gpt-5/gpt-5.1/gpt-5.2 reasoning bucket. Per Azure docs
    footnote ^7^, "minimal is only supported with the original GPT-5
    reasoning models. minimal is not supported with gpt-5.1 or greater".
    gpt-5.1 / gpt-5.2 now get none/low/medium/high/xhigh with
    supportsOff=true; bare gpt-5 keeps minimal/low/medium/high
    supportsOff=false. Ordering puts gpt-5.1 / gpt-5.2 before gpt-5 in
    the find() loop so the longer prefix matches first.
    https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/reasoning

DeepSeek:
  - Hide `seed` and `parallel_tool_calls` in the deepseek capability
    bucket. Neither field is in the current /chat/completions schema
    (body fields: messages, model, thinking, max_tokens, response_format,
    stop, stream, stream_options, temperature, top_p, tools, tool_choice,
    logprobs, top_logprobs, user_id). Surfacing them in the UI would be
    the silent-drop UX the file header warns against.
    https://api-docs.deepseek.com/api/create-chat-completion

Mistral:
  - magistral-medium-latest / magistral-small-latest are NATIVE
    always-on reasoning models; injecting reasoning_effort returns 422
    upstream. Switch both to withEnableThinkingStyle({reasoningAlwaysOn:
    true}) instead of the old none/medium/high effort ladder.
  - mistral-small-latest / mistral-medium-latest / mistral-vibe-cli-latest
    expose the documented three-tier adjustable ladder
    (none/low/medium/high), not the truncated none/high pair that was
    here before. mistral-medium-latest was not handled at all and now
    sits in the same bucket as small.
    https://docs.mistral.ai/studio-api/conversations/reasoning
    https://mistral.ai/news/magistral

OpenRouter:
  - Drop google/gemini-pro-latest from OPENROUTER_MANDATORY_REASONING_
    MODELS; the gateway 404s the id today
    (https://openrouter.ai/google/gemini-pro-latest). Removing rather
    than re-pinning to a versioned id that may rotate again.

Local backends:
  - Split LOCAL_LLAMA_CAPABILITIES into LLAMA_CPP_CAPABILITIES (full
    chain — for llama_cpp + custom) and VLLM_OLLAMA_CAPABILITIES (OpenAI
    subset + top_k/min_p/repetition_penalty/seed, no extended samplers).
    vLLM's SamplingParams has no typical_p / top_n_sigma / repeat_last_n
    / dynatemp_* / mirostat* fields, and Ollama's OpenAI translator
    (ollama/openai/openai.go FromChatRequest) only copies the OpenAI
    subset. Surfacing the eight extra sliders for vllm / ollama was
    silent-drop UX.

Tests:
  - test_deepseek_payload_omits_seed_and_parallel_tool_calls: read the
    TS file as text and assert the bucket has seed:false and
    parallelToolCalls:false. Backend has no JS engine; this is the
    cheapest way to lock the wire-drop invariant.
  - 63/63 sampling_params_routing tests pass; frontend tsc clean.
2026-05-27 06:31:40 +00:00
Daniel Han
22111744a4 Narrow Anthropic 4.7 sampling-removed gate to Opus only (PR #5711)
The 4.7 generation only shipped Claude Opus 4.7; Sonnet stops at 4.6
and Haiku at 4.5 per
https://platform.claude.com/docs/en/about-claude/models/overview.
The earlier `^claude-(?:opus|sonnet|haiku)-4-7` regex on both the
backend strip (_ANTHROPIC_4_7_SAMPLING_REMOVED in external_provider.py)
and the frontend mirror (ANTHROPIC_4_7_SAMPLING_REMOVED_REGEX in
provider-capabilities.ts) would have pre-emptively hidden temperature
/ top_p / top_k for any future claude-sonnet-4-7 or claude-haiku-4-7
id, even though Anthropic has explicitly not extended the sampling
removal beyond Opus. Tighten both regexes to `^claude-opus-4-7(?:[-.]|$)`
and update the routing-test pin so claude-sonnet-4-7 and claude-haiku-4-7
are in `should_not_match`. If those ids ever ship and adopt the same
removal, widening the regex is one-line.
2026-05-27 05:43:38 +00:00
pre-commit-ci[bot]
6b300699bf [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-26 14:36:42 +00:00
Daniel Han
60085606a7 Merge branch 'main' into feat/expose-sampling-params-core
Single conflict on studio/backend/core/inference/external_provider.py:
main added `previous_response_id` plumbing for OpenAI Responses chaining
in the same body-builder block PR 5711 uses for service_tier /
parallel_tool_calls. Kept both sets — service_tier+parallel_tool_calls
write first, then previous_response_id appends. No behavioural change
to either feature.

163 backend routing tests still pass; frontend tsc clean.
2026-05-26 14:36:15 +00:00
Daniel Han
facdff9ad7 Add extended llama.cpp samplers + OpenRouter top_a (PR #5711)
Cross-checked every supported sampling field against each provider's
live docs + LiteLLM's drop_params surface + the llama.cpp server
README. Pulled in the most-asked-for samplers that the PR was missing.

New ProviderCapabilities flags (default false on every SaaS provider
since none accept these):
  - typicalP            (already shipped one commit prior)
  - topNSigma           llama.cpp `top_n_sigma`
  - repeatLastN         llama.cpp `repeat_last_n` (paired w/ repeat_penalty)
  - dynatempRange       llama.cpp `dynatemp_range`
  - dynatempExponent    llama.cpp `dynatemp_exponent`
  - mirostat            llama.cpp `mirostat` mode (0/1/2)
  - mirostatTau         llama.cpp `mirostat_tau`
  - mirostatEta         llama.cpp `mirostat_eta`
  - topA                OpenRouter `top_a` (alternate dynamic-top-P)

Capability bucketing split: ALL_SUPPORTED retired in favor of
  - LOCAL_LLAMA_CAPABILITIES  -> custom / vllm / ollama / llama_cpp
    (full llama.cpp sampler chain, top_a off — not a llama.cpp field)
  - OPENROUTER_CAPABILITIES   -> openrouter
    (gateway's documented set incl. top_a, llama.cpp-only knobs off
     because OpenRouter docs don't list them and they'd be silently
     dropped on most underlying routes)

InferenceParams gains 8 nullable-number fields (mirroring `seed`'s
"null = unset, finite-number = forwarded" shape). DEFAULT_INFERENCE_PARAMS
defaults each to null. Persistence handler in chat-settings-storage
mirrors typicalP's nullable-float handling for all 8.

Backend:
  - 8 new ChatCompletionRequest fields with appropriate `ge`/`le`
    validators (mirostat 0..2, ranges 0.0..1.0 where applicable).
  - llama_cpp.py: signatures + payload forwarding extended on all
    three builders (chat-completion, agentic tool-loop, final-pass)
    so the new fields survive the local tool-loop too. `is not None`
    gating so defaults (e.g. mirostat=0) reach the wire only when the
    caller explicitly opted in.
  - routes/inference.py: _build_passthrough_payload extends to the
    extended sampler chain; 3 call sites (generate_chat_completion,
    generate_chat_completion_with_tools, _build_passthrough_payload)
    forward each field from `payload.*`.

Frontend chat-adapter: external branch forwards only when capability
allows (so OpenRouter gets top_a but not mirostat, local gets mirostat
but not top_a); local branch forwards unconditionally when the value
is meaningful (e.g. mirostat != 0, dynatemp_range > 0).

Test pinning the new field round-trip through _build_passthrough_payload
added; full PR-touched suite now 163 passing (was 161).

References:
  - llama.cpp server params: https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
  - OpenRouter params:       https://openrouter.ai/docs/api/reference/parameters
  - LiteLLM provider params: https://docs.litellm.ai/docs/completion/input
2026-05-26 14:34:32 +00:00
Daniel Han
e57a1a73b8 Update _utils.py v0.1.42-beta 2026-05-26 07:25:14 -07:00
Daniel Han
849da89605
Fix unsloth studio update silently downgrading on macOS arm64 (#5767)
* Fix unsloth studio update silently downgrading on macOS arm64

Root cause: studio/install_python_stack.py's "Updating base packages"
step passes `--upgrade-package unsloth -r base.txt -c constraints.txt`
with base.txt's `unsloth` and `unsloth-zoo` entries unpinned. On macOS
arm64 the resolver silently backtracks to an older unsloth (2026.5.2 or
even 2025.7.2) whenever a transitive constraint (the most common one is
bitsandbytes wheel availability: 0.49.0+ ships macosx_14_0_arm64 wheels,
older versions do not) makes the unpinned requirement satisfiable by an
older release. install.sh already maintains an explicit `unsloth>=N.N.N`
floor for the same reason, but the floor was missing from the in-venv
update path.

Reproduced on macos-14 across 2026.3.18 / 2026.4.8 / 2026.5.2 / 2026.5.6
starting states. All four ended on unsloth==2026.5.2 after a clean
`unsloth studio update` invocation (2026.5.6 was a true downgrade,
others were stale or partial advances).

Fix mirrors install.sh: query PyPI at runtime for the current latest
version of unsloth and unsloth-zoo, then pass `unsloth>=<latest>` and
`unsloth-zoo>=<latest>` as extra positional pins alongside the existing
`--upgrade-package` flags. Network failures fall back to the historical
unpinned behaviour so offline installs continue to work. Applied to all
three upgrade branches (standard update, local-repo overlay, no-torch).

Also fix the cosmetic `Hardware detected: MLX -- Apple Silicon (i386)`
banner. platform.processor() reads `uname -p` which returns "i386" on
many universal2-shaped Python builds even on a native arm64 interpreter;
platform.machine() is the reliable source ("arm64" once is_apple_silicon
has gated us).

* Dedup floor-pin call sites + LRU cache PyPI lookup

Three upgrade branches each rebuilt the same conditional `unsloth>=` /
`unsloth-zoo>=` arg list with two PyPI round-trips per branch -- six
round-trips per `unsloth studio update` invocation. Extract a
`_pin_floor_args(*, include_unsloth=True)` helper and wrap
`_resolve_latest_pypi_version` in `functools.lru_cache` so the three
branches share a single PyPI request per package.

Functionally equivalent; pure cleanup on top of the previous commit.

* Warn when PyPI is unreachable so the silent fallback is visible

If `_resolve_latest_pypi_version` returns None for either lookup the
floor args are silently dropped, which restores the pre-fix resolver
behaviour. Print a single cyan `warning` line in `_pin_floor_args` when
that happens so users behind a proxy / captive portal / firewalled
PyPI mirror know the upgrade has degraded -- and can supply network
egress or a `--index-url` mirror and retry.

* Soft floor with unpinned-fallback for hosts where floor is unsatisfiable

Reviewer found that the unconditional unsloth-zoo>=LATEST floor turns
a previously-resolvable macOS 13 arm64 update into a hard resolver
failure: unsloth-zoo 2026.5.4 requires mlx-vlm>=0.4.4 -> mlx>=0.30.0,
and mlx 0.30+ only publishes macosx_14_0_arm64 wheels. The pre-fix
behaviour backtracked to an older unsloth instead of erroring. We
should not turn "stale" into "fail".

Add pip_install_with_floor_fallback: first try the install with the
floor appended; if the resolver cannot satisfy it (subprocess exit
code != 0), retry the install without the floor and print a clear
warning. The fall-through preserves the legacy "succeed-but-stale"
contract on hosts where wheel availability is the bottleneck.

Also extend pip_install_try with a req= kwarg so the floor attempt
can pass `-r base.txt` like pip_install does, and add an
UNSLOTH_NO_PYPI_FLOOR=1 opt-out for air-gapped CI / corporate PyPI
mirrors that intentionally do not expose pypi.org directly.

All three upgrade branches (standard, local-repo, no-torch) now go
through the helper so the fallback behaviour is consistent.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Add second fallback level: floor without constraints

macOS arm64 floored attempt with -c constraints.txt fails because the
single-env constraint `transformers==4.57.6` conflicts with the new
unsloth-zoo 2026.5.4 -> mlx-vlm 0.4.4+ -> transformers>=5.1.0 chain.
First fallback level retries the floored install without constraints
(transformers freely resolves to a mlx-vlm-compatible version);
downstream pip_install calls still apply constraints.txt to anything
that doesn't transitively conflict.

If THAT still fails (wheel availability rather than constraint
conflict), drop the floor and fall back unpinned as before.

Verified locally with uv pip compile against aarch64-apple-darwin
python-3.13: strict-constrained floor errors, no-constraint floor
resolves cleanly to unsloth==2026.5.7 + unsloth-zoo==2026.5.4 +
transformers==5.5.0 + mlx-vlm==0.5.0.

* setup.sh/.ps1: also gate fast-path on unsloth-zoo being up to date

The version-check fast-path in setup.sh / setup.ps1 only looked at
unsloth itself. If unsloth was at the PyPI latest but unsloth-zoo was
stale, the gate set _SKIP_PYTHON_DEPS=true and install_python_stack.py
never ran -- so the new floor pin from PR #5767 had no effect for the
exact "unsloth at latest, zoo behind" state several reviewers flagged.

Probe both packages' installed-vs-latest versions and only skip the
deps step when BOTH match. When either is behind, fall through to
install_python_stack.py so the new resolver fix gets a chance to run.

Verified setup.sh with `bash -n`; the setup.ps1 change uses PowerShell
if-expressions for the null-default pattern rather than bash-style
${var:-default} which is not valid PowerShell.

* Skip unsloth-zoo floor too for custom no-torch test packages

Reviewer found the asymmetric guard: the no-torch branch was already
gating the unsloth floor on package_name == "unsloth" (test side
packages may not publish to PyPI), but the unsloth-zoo floor was
still added unconditionally. A custom no-torch update that ships its
own forked zoo metadata could now hit a public PyPI floor that does
not match the fork's published version.

Add a symmetric `include_zoo` parameter to `_pin_floor_args` and
gate both pins on the same `package_name == "unsloth"` check.

* Address review feedback: simpler except clause + private-index note

Gemini flagged TimeoutError in the PyPI fetch exception list. OSError already
covers socket timeouts and the 3.11+ TimeoutError subclass on every supported
Python, so drop the redundant entry and explain what each remaining exception
catches.

Codex flagged that floor lookups against pypi.org could break installs behind
a lagging private mirror. Step 3 of pip_install_with_floor_fallback already
recovers transparently in that case; expand the docstring so the behavior is
discoverable without reading the body.

* extras-no-deps: skip transformers==4.57.6 on macOS arm64

Reviewer flagged that the resolver-selected transformers from the
no-constraints base step on macOS arm64 (transformers 5.x for mlx-vlm
0.4.4+) gets silently downgraded back to 4.57.6 by extras-no-deps.txt
during the very next step, breaking mlx-vlm imports at runtime even
though unsloth itself reports as latest.

Add a PEP 508 platform marker so the pin only applies off macOS arm64.
constraints.txt still enforces 4.57.6 everywhere else; mlx-vlm only
publishes wheels for darwin arm64, so other platforms are unaffected.

* setup.sh/.ps1: gate fast-path zoo probe on _PKG_NAME == unsloth

Reviewer found the asymmetric custom-package regression: the new
zoo-aware fast-path probes public unsloth-zoo unconditionally, but a
custom STUDIO_PACKAGE_NAME side build may ship its own zoo fork via
dependency metadata and not install public unsloth-zoo at all. The
previous behaviour (skip Python deps if the custom package itself is at
its declared latest) is preserved by only running the zoo probe when
the managed package literally IS unsloth.

Matches the include_zoo gate already in _pin_floor_args() at
install_python_stack.py.

* install_python_stack: all-or-nothing floor + uv-to-pip retry

Two reviewer findings on the floor-pin helpers:

1. _pin_floor_args() previously kept a half-floor if one PyPI lookup
   succeeded and the other failed. With unsloth at latest but the zoo
   lookup down, the resolver could still backtrack zoo while we
   required unsloth at latest, defeating the pin. Return [] on any
   lookup failure so the unpinned legacy path runs cleanly.

2. pip_install_try() ran ONLY uv when USE_UV was true; a uv-specific
   failure short-circuited to False even when pip itself could have
   applied the floor. Mirror pip_install()'s uv-to-pip fallback: try
   uv, fall through to pip on non-zero exit, and only then give up.

* extras-no-deps: rewrite marker without `not` for PEP 508 parsers

pip's vendored packaging rejects `not (...)` in PEP 508 markers; the
grammar only specifies `and` / `or` between boolean atoms. The staging
macos-14 matrix failed every job at "Installing extras (no-deps)" with
`Expected a marker variable or quoted string`. Apply De Morgan's law
so the marker uses `or` between two `!=` checks, which both pip and
uv parse cleanly. Behaviour identical: skip the 4.57.6 pin only on
darwin arm64; pin everywhere else.

* constraints: skip transformers==4.57.6 pin on macOS arm64 too

Marker-gating the extras-no-deps.txt pin was not sufficient. Every
subsequent pip_install in the update pipeline passes
-c single-env/constraints.txt, and constraints.txt itself pinned
transformers==4.57.6 unconditionally. The latest staging-2 run shows
the base step's no-constraints fallback installed transformers 5.5.0
correctly, but a later constrained step (extras / studio / data-designer
deps) silently downgraded it back to 4.57.6, leaving mlx-vlm 0.5.0
in the venv with an unsatisfied transformers>=5.5.0 requirement.

Apply the same `sys_platform != "darwin" or platform_machine != "arm64"`
marker to the constraints.txt entry so it is inert on darwin arm64.
Other platforms still pin 4.57.6 because mlx-vlm only publishes wheels
for darwin arm64; no other platform is affected.

* constraints: carve out darwin arm64 from every == pin

Marker-gating only transformers was not enough; staging-2 still failed
with the same `transformers==4.57.6 in venv after the update` outcome
because the resolver hit a `huggingface-hub==0.36.2` (and adjacent)
conflict with mlx-vlm's `huggingface-hub>=1.5.0` requirement, then
fell back to a stale stack even after my no-constraints level fired
on the base step.

Apply the same `sys_platform != "darwin" or platform_machine != "arm64"`
marker to every == pin in constraints.txt. Range pins (mcp, fastmcp,
websockets) stay active everywhere because they do not conflict with
the mlx-vlm chain. mlx-vlm only publishes wheels for darwin arm64, so
no other platform is affected.

* install_python_stack: also --upgrade-package transformers and mlx-vlm

Staging-2 showed that even after the constraints.txt carve-out for
darwin arm64, the venv still ended up with the OLD `transformers==4.57.6`
paired with a NEW `mlx-vlm==0.5.0` from unsloth-zoo's transitive
upgrade. The resolver's --upgrade-package flag only freshens the named
packages and their newly-pulled transitive deps; transformers was
already installed at a version that satisfied unsloth-zoo's range
(`>=4.51.3,<=5.5.0` with exclusions), so the resolver did not upgrade
it -- even though mlx-vlm 0.5.0 requires `transformers>=5.5.0`.

Add `--upgrade-package transformers` and `--upgrade-package mlx-vlm`
to all three base-step branches. Both are no-ops when the package is
absent (mlx-vlm only ships wheels on darwin arm64); on darwin arm64
this is what nudges the resolver to upgrade both together so the
final venv is internally consistent. On Linux/Windows, transformers
stays at 4.57.6 because constraints.txt still pins it there and
mlx-vlm never enters the resolution.

* install_python_stack: explicit mlx-vlm + transformers realign on macOS arm64

Even with --upgrade-package hints, uv leaves the venv with the
already-installed transformers (4.57.6 inherited from the OLD venv's
constrained install) when that version still happens to satisfy
unsloth's own metadata range -- but it does not also re-resolve
mlx-vlm's stricter `transformers>=5.5.0` requirement, so the venv
ends up with mlx-vlm 0.5.0 paired with transformers 4.57.6 and
mlx-vlm imports break at runtime.

After the base step, on darwin arm64 only, run an explicit
`pip install --upgrade mlx-vlm transformers` with constrain=False.
This forces both packages through the resolver again as direct
top-level requirements, so transformers is pulled up to whatever
mlx-vlm's metadata requires (5.5.0 today). No effect on any other
platform because mlx-vlm has no wheels off darwin arm64 and the
branch is gated on IS_MAC_ARM.

* requirements: marker-gate every == pin that conflicts with mlx-vlm chain

Staging-2 kept ending up with transformers==4.57.6 even after the
realign step, because studio.txt unconditionally pins
huggingface-hub==0.36.2 (and datasets==4.3.0). Installing studio.txt
with constraints active pulls the resolver back to a huggingface-hub
that only recent transformers (4.x) supports, which silently downgrades
the realigned 5.5.0 to 4.57.6 -- exactly the inconsistency we tried to
prevent.

Also extras-no-deps.txt still pinned trl==0.23.1 unconditionally; the
0.23.1 wheel transitively requires huggingface-hub<1, same coupling.

Marker-gate all three. The carve-out is identical to constraints.txt's:
inactive on darwin arm64 (where the mlx-vlm chain dictates newer
versions), active everywhere else (where Linux/Windows users rely on
the single-env pins). mlx-vlm only publishes wheels for darwin arm64
so no other platform is affected.

* realign: --force-reinstall mlx-vlm + transformers + huggingface_hub

Plain --upgrade does not force uv to re-resolve mlx-vlm's transformers
requirement when the already-installed transformers happens to satisfy
unsloth's own range. Switch to --force-reinstall on the three packages
so the resolver tears them down and brings them back together with
consistent versions. Include huggingface_hub because transformers 5.x
requires hf-hub>=1.5.0 and the resolver would not touch it otherwise.

* realign: pin transformers via mlx-vlm's own metadata spec

`pip install --force-reinstall mlx-vlm transformers` still resolved to
an already-installed transformers 4.57.6 because uv treats it as
satisfying unsloth's transformers range without re-checking mlx-vlm's
stricter requirement. Pull mlx-vlm's actual transformers specifier
from its installed metadata at runtime and pass it as an explicit
version requirement (e.g. `transformers>=5.5.0` for mlx-vlm 0.5.0).
That removes the resolver's wiggle room: it MUST pick a transformers
satisfying mlx-vlm AND unsloth, which on darwin arm64 with the latest
unsloth-zoo means transformers==5.5.0. Falls back to unpinned
`transformers` if metadata read fails, so this never errors.

* realign: uninstall-then-install to bypass uv's incumbent bias

Every flag-based approach failed: --upgrade, --upgrade-package,
--force-reinstall, and even an explicit `transformers>=5.5.0`
requirement all left the venv with transformers==4.57.6 because uv
treats the already-installed version as satisfying unsloth-zoo's
range and refuses to disturb it, even when it does not satisfy
mlx-vlm's stricter requirement.

Replace the realign step with an explicit uninstall of the conflicting
trio (transformers / mlx-vlm / huggingface_hub) followed by a fresh
install. With no transformers in the venv, the resolver MUST pick a
version satisfying every installed package's metadata, which on
darwin arm64 with the latest unsloth-zoo is uniquely 5.5.0.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Trim verbose comments across PR #5767 changes

* Simplify mac-arm64 fix: install MLX stack with --no-deps

The previous approach (PyPI floor pin + 3-level fallback + macOS arm64
realign step + marker carve-outs on every == pin) was fighting symptoms.
The root cause is that unsloth-zoo declares mlx-vlm>=0.4.4 as a darwin
arm64 dep, and mlx-vlm 0.5.0's metadata pulls in transformers>=5.5.0,
which conflicts with the main venv's transformers==4.57.6 pin and forces
the resolver to backtrack unsloth.

Severing that chain at its source: install mlx + mlx-metal + mlx-lm +
mlx-vlm with --no-deps BEFORE unsloth-zoo. The resolver sees mlx-vlm
already installed (>=0.4.4) and never inspects its transformers metadata.
Per-model transformers version routing is already handled at runtime by
the side-car venvs in utils/transformers_version.py (.venv_t5_530 for
Ministral/GLM/Qwen3 MoE, .venv_t5_550 for Gemma 4).

Net change: -224 / +71 lines across install.sh, install_python_stack.py
and the three requirements files.

Reverted:
- _resolve_latest_pypi_version + _pin_floor_args + pip_install_with_floor_fallback
- macOS arm64 realign step (pip uninstall + reinstall)
- --upgrade-package transformers --upgrade-package mlx-vlm in base steps
- All ; sys_platform != "darwin" or platform_machine != "arm64" markers
  in constraints.txt, studio.txt, extras-no-deps.txt
- pip_install_try restored to its pre-PR signature

Added:
- install.sh: Apple Silicon MLX --no-deps install before unsloth (both
  fresh and migrated branches)
- install_python_stack.py: same step gated on IS_MAC_ARM and not skip_base

Kept (independent bugs):
- setup.sh / setup.ps1 dual-package zoo version check
- platform.processor() -> platform.machine() hardware-detect fix

* Minimise PR to mac-arm64-specific changes only

Revert setup.sh and setup.ps1 to main -- the dual-package zoo check was
defensive and not strictly needed once mlx-vlm is installed --no-deps
(the resolver-backtrack scenario that produced stale zoo no longer happens).

Tighten remaining comments in install.sh and install_python_stack.py.

Final PR-attributable changes:
  install.sh                                  +24/-5  (MLX --no-deps in 2 places)
  studio/install_python_stack.py              +19    (MLX --no-deps + IS_MAC_ARM)
  studio/backend/utils/hardware/hardware.py    +6/-6 (processor() -> machine())
  studio/backend/requirements/*.txt            unchanged

* Revert "Minimise PR to mac-arm64-specific changes only"

This reverts commit 9470daa855.

* Revert "Simplify mac-arm64 fix: install MLX stack with --no-deps"

This reverts commit f8a43b87e8.

* Revert "Trim verbose comments across PR #5767 changes"

This reverts commit c3f293a10f.

* Simplify mac-arm64 fix: --no-deps MLX + METADATA patch

Root cause: unsloth-zoo declares mlx-vlm>=0.4.4 as a darwin-arm64 dep, and
mlx-vlm 0.5.0's published metadata declares transformers>=5.5.0. Every
subsequent resolver run with constraints.txt's transformers==4.57.6 sees
the conflict and backtracks unsloth to escape it (user-reported downgrade).

The aggressive pin doesn't reflect what mlx-vlm actually requires at
top-level import time -- the symbols it loads (AutoProcessor, AutoTokenizer,
ProcessorMixin, BatchFeature) are stable across transformers 4.51+. Model-
specific submodules that genuinely need 5.x APIs are only loaded once the
3-tier transformers dispatcher (utils/transformers_version.py) has activated
the matching .venv_t5_530 / .venv_t5_550 side-car at runtime.

Fix: on Apple Silicon, install the MLX stack with --no-deps then rewrite
mlx-vlm/mlx-lm's installed METADATA to declare transformers>=4.51.3. Now
the resolver sees mlx-vlm 0.5.0 as compatible with the main venv's
transformers==4.57.6 and there's nothing to backtrack.

Reverts the previous heavy machinery:
- _resolve_latest_pypi_version, _pin_floor_args, pip_install_with_floor_fallback
- macOS arm64 realign step (pip uninstall + reinstall)
- --upgrade-package transformers --upgrade-package mlx-vlm in base steps
- All ; sys_platform != "darwin" or platform_machine != "arm64" markers
  in constraints.txt / studio.txt / extras-no-deps.txt
- setup.sh / setup.ps1 dual-package zoo check (Windows never had the bug;
  with this fix in place stale zoo no longer happens on macOS either)
- pip_install_try restored to pre-PR signature

Kept:
- install.sh: MLX --no-deps install in fresh + migrated branches
- install_python_stack.py: same step gated on IS_MAC_ARM and not skip_base
- _relax_mlx_metadata() helper, called immediately after each MLX install
- studio/backend/utils/hardware/hardware.py: platform.processor() ->
  platform.machine() cosmetic fix

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Use UV_OVERRIDE to relax mlx-vlm transformers pin

uv supports --overrides / UV_OVERRIDE which globally overrides any package's
stated dependency requirement. mlx-vlm 0.5.0 declares transformers>=5.5.0
and mlx-lm 0.31.3 declares transformers>=5.0.0; neither is true at top-level
import time (their imports use AutoProcessor / AutoTokenizer / ProcessorMixin /
BatchFeature which are stable across transformers 4.51+). Per-model 5.x
routing is handled at runtime via the .venv_t5_530 / .venv_t5_550 side-cars.

Override file (overrides-darwin-arm64.txt) declares transformers>=4.51.3 ;
exported via UV_OVERRIDE env var on Apple Silicon by both install.sh and
install_python_stack.py. uv then resolves mlx-vlm as compatible with the main
venv's transformers==4.57.6 (constraints.txt) and unsloth advances cleanly to
LATEST.

Drops, vs. the previous attempts:
- _resolve_latest_pypi_version + _pin_floor_args + pip_install_with_floor_fallback
  (floor-pin machinery -- replaced by single UV_OVERRIDE line)
- macOS arm64 realign step (pip uninstall + reinstall)
- --upgrade-package transformers --upgrade-package mlx-vlm in base steps
- All ; sys_platform != "darwin" or platform_machine != "arm64" markers
- _relax_mlx_metadata() helper + sed METADATA patch (uv reads from index, not
  dist-info, so dist-info patches were ineffective)

Kept:
- install.sh / install_python_stack.py: MLX latest install on Apple Silicon
  (now without --no-deps, the override lets the resolver pick a consistent set)
- studio/backend/utils/hardware/hardware.py: platform.machine() cosmetic fix

* Trim UV_OVERRIDE comments; bump override floor to 4.57.6

Match the main venv's constraints.txt pin exactly so the override file
reads as the actual installed version rather than mlx-vlm's API floor.
Comments collapsed to one-liners where possible.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-26 07:23:13 -07:00
Daniel Han
a02320aa7b Add typical_p sampler (local) + DeepSeek-reasoner per-model gating (PR #5711)
Two follow-ups from a closer reading of each provider's published
sampling surface against llama.cpp's own server README.

typical_p (locally typical sampling, `typ_p` in the llama.cpp sampler
chain):
  - New ProviderCapabilities.typicalP flag; defaults false on every
    SaaS provider (none accept the field) and true only on the
    permissive local buckets (custom, vllm, ollama, llama_cpp,
    openrouter via ALL_SUPPORTED). InferenceParams.typicalP is
    nullable number (null = unset; 1.0 = llama-server default, also
    treated as no-op when forwarding).
  - Backend: new ChatCompletionRequest.typical_p Field (0.0..1.0).
    Threaded through all three llama_cpp.py payload builders
    (chat-completion, agentic tool-loop, final-pass) so the field
    survives the local tool-loop too. _build_passthrough_payload in
    routes/inference.py picks it up and only writes the body when
    the caller set a value; left absent it falls back to llama-server
    default. Three route call sites (generate_chat_completion,
    generate_chat_completion_with_tools, _build_passthrough_payload)
    forward payload.typical_p.
  - Frontend: chat-adapter forwards on both branches (external opt-in
    only when capability allows + value != 1; local forwards
    unconditionally when set and != 1). OpenAIChatCompletionsRequest
    grows a `typical_p?` field. Persisted via chat-settings-storage
    mirroring the seed nullable-float handler.
  - Test: pin _build_passthrough_payload's forward + absent behavior.

DeepSeek per-model gating:
  - deepseek-reasoner / deepseek-r1 silently ignore temperature, top_p,
    presence_penalty, frequency_penalty per
    https://api-docs.deepseek.com/guides/reasoning_model — mirror the
    OpenAI / Claude 4.7 per-model approach: getProviderCapabilities
    downshifts these ids to a stripped capability set so the panel
    does not offer knobs the upstream silently drops.

161+1 sampling-routing tests pass; frontend tsc clean.

Refs:
  - llama.cpp server params: https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
  - DeepSeek reasoner restrictions: https://api-docs.deepseek.com/guides/reasoning_model
2026-05-26 14:12:40 +00:00
oobabooga
953c8bfa45
studio/frontend: add space above the scroll-to-bottom arrow in chat (#5776) 2026-05-26 06:02:03 -07:00
oobabooga
ef6744253f
Studio: stop the model from replying twice when it refuses (#5775) 2026-05-26 05:30:52 -07:00
oobabooga
eacff8b827
studio/frontend: show "1 second" instead of "1 seconds" in thinking blocks (#5777)
* studio/frontend: show "1 second" instead of "1 seconds" in thinking blocks

* Update studio/frontend/src/components/assistant-ui/reasoning.tsx

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-26 05:30:22 -07:00
Daniel Han
c49cc6daf5
Studio: auto-recover when shadowed 'unsloth' on PATH hides the frontend dist (#5782)
* Studio: auto-recover when shadowed 'unsloth' on PATH hides the frontend dist

The CLI launcher derives `_PACKAGE_ROOT` from where `unsloth_cli` imports
from, and `studio/backend/run.py` derives its default `frontend_path` from
`Path(__file__).resolve().parent.parent / "frontend" / "dist"`. When
another `unsloth` (a separate venv with `pip install unsloth`, a system
install, an older venv earlier on PATH) wins `which unsloth`, both
resolve into a site-packages tree that ships frontend source files but no
vite-built `dist/`. The backend warned `[WARNING] Frontend not found at
...` and then happily served 200 on every `/api/*` route while returning
`{"detail":"Not Found"}` on `/`. The 404 was silent to users -- the
process was healthy, the log line scrolled by, and the only symptom was a
blank browser tab.

This is a real situation: many devboxes carry a workspace venv with
`unsloth` installed years before the user runs `curl|sh` to install
Studio. The installer-managed binary at `~/.local/bin/unsloth` exists
but loses to the older venv on PATH order.

Three layers of fix, additive:

Layer C -- runtime auto-discovery (unsloth_cli + run.py)
The CLI now resolves `--frontend` explicitly before spawning `run.py`,
probing in order: package-local default, installer venv site-packages
(`$STUDIO_HOME/unsloth_studio/lib/python*/site-packages/...` and the
Windows `Lib/site-packages/...` equivalent), and editable-install source
roots read from `__editable___*_finder.py` MAPPING dicts in the installer
venv. `run.py` does the same probe as a backstop for direct `python
run.py` invocations.

Layer E -- loud structured error
The silent `[WARNING]` is replaced with a `SystemExit` that names every
candidate path tried and lists the four one-line fixes (run the absolute
path, pass `--frontend`, pass `--api-only`, reinstall). Suppressed only
in `--api-only` mode where no UI is served by design.

Layer F -- installer self-check (install.sh + install.ps1)
At the tail of install, both installers compare `command -v unsloth`
(POSIX) / `Get-Command unsloth` (PowerShell) against the just-installed
binary. If a different path wins, a yellow `warning` block names the
shadowing binary and prints the alias / absolute-path / PATH-reorder
fixes. install.sh uses the venv Python for path canonicalization so it
also works on macOS (BSD `readlink` has no `-f`).

Cross-platform notes:
- Glob patterns probe both `lib/python*/site-packages` (POSIX) and
  `Lib/site-packages` (Windows).
- Canonical-binary path branches on `sys.platform == "win32"` to pick
  `unsloth.exe` over `unsloth`.
- install.sh fixed for macOS; install.ps1 is the Windows analog.

Tests: `studio/backend/tests/test_frontend_resolution.py` covers five
cases via AST-load of the helpers (no uvicorn / FastAPI import needed,
matching `test_host_defaults.py`'s style):
1. Resolver returns None when nothing exists anywhere.
2. Resolver picks the first existing candidate when the default works.
3. Fallback to `$UNSLOTH_STUDIO_HOME` site-packages dist when the default
   is missing.
4. Fallback to an editable-install source root via MAPPING parsing.
5. Resolver tolerates a non-existent `$UNSLOTH_STUDIO_HOME`.

All 5 new + 2 existing host-default tests pass.

* Studio: address review feedback on PR 5782 (Windows hardlink, Win path hint, broader tests)

Four parallel platform reviews (Windows, Linux, macOS, general) on the
initial commit surfaced a small batch of correctness items, all addressed
here:

Windows install.ps1 (medium severity, false positive on every install):
The user-facing shim at $StudioHome\bin\unsloth.exe is a hardlink to
$VenvDir\Scripts\unsloth.exe (created at line 1582). Resolve-Path does not
de-duplicate hardlinks, so the previous string compare always saw the two
paths as different and the new "another 'unsloth' wins on PATH" warning
would fire on every fresh Windows install. Switched to content-hash
equality via Get-FileHash, which collapses hardlinks, symlinks, and
identical copies to a single identity. Also restricted the probe to
Get-Command -CommandType Application so PowerShell aliases / functions /
scripts named "unsloth" don't false-trigger.

Windows run.py SystemExit hint (medium severity, defeats the recovery UX):
The structured error printed Path(STUDIO_HOME)/"unsloth_studio"/"bin"/
"unsloth.exe" on every platform, but on Windows the installer places the
shim at $STUDIO_HOME/bin/unsloth.exe (no unsloth_studio segment) and the
venv binary at $STUDIO_HOME/unsloth_studio/Scripts/unsloth.exe. The hint
pointed at a non-existent path on Windows. Branch on sys.platform ==
"win32" to emit the real shim location; Linux / macOS keep the unsloth_
studio/bin/unsloth layout.

MAPPING regex robustness (low):
[^\n]* silently failed if a future setuptools / black reformat wrapped
the MAPPING dict across multiple lines. Tightened to [^}]* + re.DOTALL,
which still rejects nested dicts (setuptools never emits those for
editable installs) but tolerates either single- or multi-line literals.

install.sh broken-venv edge case (low, macOS reviewer):
Previously _canon fell back to echoing the raw input when the venv python
failed, which would make two symlinked-but-identical paths look different
and false-trigger the warning. Now _canon returns empty on failure and
the caller skips the whole comparison if either side is unresolvable.

argparse default + log readability (nits):
run.py's argparse --frontend default now reuses the module-level
_DEFAULT_FRONTEND_PATH constant so it stays in lockstep with run_server's
default. The [OK] log message resolves the chosen path so support output
is always absolute.

Tests grow from 5 to 8 in studio/backend/tests/test_frontend_resolution.
py (10/10 with the existing host-default tests):

- Windows-layout fallback: Lib/site-packages with capital L.
- Multi-line MAPPING dict: locks in the [^}]* + re.DOTALL behaviour.
- SystemExit message contract: every actionable fix string and the
  attempted-paths list must appear; pins the user-facing recovery
  message so a future refactor doesn't drop a bullet.

End-to-end re-verified on this box: shadowing workspace_22/bin/unsloth
still serves 200 on / through the editable-finder fallback, with the
follow-up resolve-then-log change yielding [OK] Frontend loaded from
/mnt/disks/unslothai/ubuntu/unsloth/studio/frontend/dist.

Out of scope (called out by reviewers but deferred):

- _resolve_frontend_path candidate ordering still tries _PACKAGE_ROOT
  first. For the rare case where a shadowing install carries an older
  built dist, this serves the stale UI instead of the fresh one. Fix is
  non-trivial (the --local workflow intentionally wants _PACKAGE_ROOT to
  win when the cloned repo is the source of truth), so leaving it for a
  follow-up.
- studio/backend/colab.py still bails out on missing frontend instead of
  routing through the new resolver. Pre-existing behaviour, separate PR.
- _resolve_frontend_path is duplicated across run.py and unsloth_cli/
  commands/studio.py. Minor maintenance concern; consolidation is
  natural in a later refactor.

* Studio: guard ast.literal_eval result with isinstance(dict)

Addresses gemini-code-assist[bot] high-priority inline review on PR 5782
flagging that `mapping.get('studio')` could raise AttributeError if the
MAPPING regex matched a brace-delimited literal that ast.literal_eval
parsed as a non-dict (set, list, None). The regex `\{[^}]*\}` happily
matches `{1, 2, 3}` and literal_eval returns a set; the previous code
then crashed on .get().

Setuptools's editable-install template only emits dict literals so this
is defensive rather than a live bug, but the guard is one line per call
site and prevents a future template change from taking out backend
startup or CLI invocation.

Both call sites (studio/backend/run.py:558 and
unsloth_cli/commands/studio.py:234) now bail out on the finder file when
isinstance(mapping, dict) is False; the resolver keeps probing the
remaining finders, so a malformed entry in one finder cannot poison the
discovery of a good one elsewhere.

Adds test_resolver_does_not_crash_on_non_dict_mapping_literal to
test_frontend_resolution.py, which writes one bad finder (MAPPING is a
set literal) alongside one good finder (MAPPING is a real dict) and
asserts the resolver returns the good finder's dist path. Without the
guard this test crashes with AttributeError; with the guard it passes.

11/11 tests green.
2026-05-26 05:29:42 -07:00
Wasim Yousef Said
fb65fed3b0
Keep generated image loading dots animated (#5786) 2026-05-26 04:37:18 -07:00
Daniel Han
41d24227cd
Studio: per-card web_search result + shell_call output fallback (OpenAI) (#5785)
* Studio: per-card web_search result + shell_call output fallback (OpenAI)

Two empty-output bugs in the OpenAI Responses tool-result rendering that
showed up clearly when a single prompt invoked 9 web_search + 4
code_execution + 1 image_generation in one turn. Reproduction shape in
the SQLite-stored chat history:

- 8 of 9 web_search tool-call records had result == "" (the cards
  rendered as empty cards in the thread)
- 4 of 4 code_execution (shell_call) records were missing the result
  key entirely (NoneType), so the cards that showed "Ran cat ..." style
  commands displayed the command line but no output panel at all
- image_generation worked, as did the very last web_search of the run

Root causes in studio/backend/core/inference/external_provider.py:

1. web_search_call's tool_end emitted result: "" by design, with the
   intent of overwriting only the LAST call at response.completed with
   the full citation list (the source-pill extractor on the frontend
   flatMaps across every web_search result, so a single non-empty
   result is enough for the trailing source pills). Side effect: every
   intermediate card renders empty in the thread. Fix: seed each call's
   own tool_end result with "Searching: <query>" so the per-card text
   is never empty, then keep the last-call overwrite path so the
   source-pill extractor still works. Falls back to empty when the
   model emits an action with no query, so the existing last-call path
   stays unchanged for that edge.

2. shell_call's tool_start was emitted from
   response.output_item.done for the call item, but tool_end lived in
   the separate response.output_item.done handler for shell_call_output.
   When OpenAI's Responses stream bundles the output array onto the
   shell_call item's own done event (no separate shell_call_output
   item), the previous handler emitted tool_start with no following
   tool_end. The card spun on "running" indefinitely and stored as
   NoneType in the thread DB. Fix: when the shell_call's done event
   carries an embedded output list, emit tool_end immediately from
   that. Track tool_end_emitted on the shell_calls map so a subsequent
   shell_call_output event (some streams ship both) is skipped instead
   of double-completing the card. A final flush at response.completed
   emits tool_end for any orphan shell_call that received neither
   bundled output nor a separate output event, so cards always finalise.

Tests (studio/backend/tests/test_openai_tool_result_fallbacks.py, 6
new):
- web_search: three calls, each card's result is its own Searching:
  query (no empties)
- web_search: last call still gets the aggregated citation block when
  url_citations arrive (pins the overwrite path)
- web_search: empty action.query falls back to result == "" (no junk
  Searching: placeholder)
- shell_call: bundled output on done emits a single tool_end with that
  output as the result text
- shell_call: bundled-then-separate output does not double-emit
  tool_end (subsequent shell_call_output is skipped)
- shell_call: orphan call with neither bundled nor separate output is
  flushed at response.completed so the card finalises

15/15 tests green when combined with the existing 9 in
test_openai_code_execution.py. Pre-commit + ruff format clean.

Scope: OpenAI Responses-API code path only. The Anthropic native
Messages-API path (_stream_anthropic) is untouched, as is the local
llama-server path. Local-model behaviour cannot regress because the
edited handlers only fire inside the OpenAI cloud branch.

* Studio: per-model external max_tokens cap + clamp on model switch

Two related external-provider issues that surfaced from the same
investigation as the per-card web_search / shell_call result bugs in
the previous commit:

A. Slider cap was a one-size-fits-all 32768 for every external model.

   provider-capabilities.ts kept a single EXTERNAL_MAX_OUTPUT_TOKENS
   constant (32k), well below what most providers actually accept. The
   docstring even called out the right per-provider numbers (Anthropic
   Opus 128k, GPT-5.x ~128k, Gemini 2.5 ~65k, DeepSeek 8k) but the
   code picked the lowest as a conservative floor. Effect: long
   generations from gpt-5.5 / claude-opus-4-7 silently truncated at
   32k even though the API would have served up to 128k.

   Fix: introduce getExternalMaxOutputTokens(providerType, modelId)
   returning the documented per-model cap. Patterns are checked
   longest-first so e.g. gpt-5.5-pro matches before gpt-5.5. Unknown
   provider/model combinations fall back to the existing 32k floor so
   no surprise increases for ids we don't know about.

   Per-model caps from the official docs:
   - OpenAI gpt-5.5 / gpt-5.5-pro: 128000
   - OpenAI gpt-5.4 / gpt-5.4-pro: 65536
   - OpenAI gpt-5.3: 16384
   - Anthropic claude-opus-4-7: 128000
   - Anthropic claude-opus-4-6 / sonnet-4-6 / opus-4-5 / sonnet-4-5 /
     haiku-4-5: 64000
   - Gemini 3.x family: 65535
   - DeepSeek: 8192
   - OpenRouter: strip provider/ prefix from the id and re-resolve

   The slider in chat-settings-sheet.tsx and the send-time clamp in
   chat-adapter.ts both call the new function so the slider's max=
   matches what the wire layer will accept.

B. Slider value lied after switching from a local model to external.

   When Studio auto-loads the helper Gemma-4-E2B-it on first chat,
   chat-adapter sets params.maxTokens to Gemma's context_length
   (262144 for Gemma 4). Switching the model picker to gpt-5.5 then
   flips the slider's max prop to the external cap, but the stored
   params.maxTokens is never reset. The numeric value next to the
   slider would render 262144 against a track that ended at the
   external cap. The send-time clamp brought the outbound max_tokens
   back down to the cap, so the API call was safe, but the displayed
   number had no relationship to what was actually being sent.

   Fix: chat-runtime-store.setCheckpoint now clamps params.maxTokens
   to getExternalMaxOutputTokens(...) on transitions into an external
   model. Looks up the provider via useExternalProvidersStore so we
   can derive providerType from the parsed external model id. No-op
   when the stored maxTokens is already at or below the new cap, so
   user-tuned values within range survive the switch.

Scope: pure frontend changes scoped to external-provider code paths.
Local model behaviour is untouched -- the ggufContextLength branch of
the slider's max= is unchanged, and setCheckpoint only mutates
maxTokens when isExternalModelId(modelId) is true. The send-time
clamp continues to be the safety net for any in-flight request that
crosses a model switch before the store-level clamp has applied.

Typecheck (tsc -b) clean; bun run build succeeds (2.13s).

Co-changes with the previous commit (7fe1adbf, per-card web_search +
shell_call output fallback) form a single PR: every empty-output and
silent-truncation issue surfaced from the same animal-popularity
prompt reproduction is now addressed in one branch.

* Studio: correct external max_tokens caps for Gemini and DeepSeek

Per-doc corrections to the per-model cap table added in 95da8d52:

- Gemini 3.x family: 65535 -> 65536, per
  https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro-preview
  (the published max_output_tokens is exactly 64K = 65536). The earlier
  65535 was an off-by-one rough cap.
- DeepSeek (deepseek-chat / deepseek-reasoner aliases): 8192 -> 384000,
  per https://api-docs.deepseek.com/quick_start/pricing. DeepSeek V4
  Flash / Pro both list MAX OUTPUT = 384K; the chat / reasoner ids are
  deprecated aliases for V4 Flash non-thinking / thinking modes. The
  8192 value was carried over from V3 and silently truncated V4 traffic
  at 2% of its actual ceiling.

Affects only the slider max and the send-time clamp for these provider
types. Other providers' caps unchanged. tsc -b clean.

* Studio: also flush orphan shell_calls on response.incomplete

Addresses gemini-code-assist[bot] high-priority inline review on PR
5785: the orphan-shell_call final flush added in 7fe1adbf landed only
in the response.completed branch. Truncated OpenAI Responses streams
emit response.incomplete instead (for example when the request hits
max_output_tokens), which left in-flight shell_call cards spinning
indefinitely in the UI.

Mirror the same flush block in the response.incomplete handler so the
truncated-stream path finalizes every pending tool card. The
tool_end_emitted guard keeps the path idempotent: if a shell_call
already completed via bundled output on its done event, the incomplete
flush is a no-op for it.

Two new tests in test_openai_tool_result_fallbacks.py:
- test_shell_call_flushed_on_response_incomplete_truncation pins the
  bug repro: an in-flight shell_call followed by response.incomplete
  must emit tool_end so the card finalizes.
- test_shell_call_incomplete_does_not_double_emit pins idempotency:
  a shell_call that completed via bundled output and is then followed
  by response.incomplete emits exactly one tool_end with the bundled
  result text.

17/17 tests green (8 fallback tests + 9 existing code-execution). Pre-
commit + ruff format clean.

* Studio: trim verbose comments across PR 5785 edits

Compress the in-code commentary added across this branch to one or two
lines per block; the verbose prose was easier as a PR description than
as inline noise. No behavioural changes: 17/17 tests still green, tsc -b
still clean.
2026-05-26 04:31:22 -07:00
Wasim Yousef Said
b1ef65c07a
Improve image generation UI (#5784)
* Improve image generation UI

* Polish generated image edit UI

* Tune generated image UI polish

* Soften generated image UI

* Refine generated image loading surface

* Align generated image caption clamp
2026-05-26 04:17:43 -07:00
Wasim Yousef Said
31ac558a73
Recipe Studio local model selector (#5769)
* feat(recipes): round-trip local model variants

* feat(recipes): add local model selector

* feat(recipes): wire selector into model editors

* fix(recipes): clear stale model state on relink

* feat(recipes): load selected local models for jobs

* chore(frontend): simplify biome scripts

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(recipes): handle local selector edge cases

* fix(recipes): polish local model selector behavior

* fix(recipes): delay local model restore until terminal runs

* fix(recipes): accept resolved default gguf variants

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-26 02:37:24 -07:00
Wasim Yousef Said
f364b08b6c
Support follow-up edits for generated images (#5712)
* Support follow-up edits for generated images

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix generated image edit references

* Fix image generation tool guard

* Replay OpenAI image reasoning refs

* Capture streamed OpenAI reasoning refs

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Trigger Gather Town test

* Use OpenAI response context for image edits

* Studio: improve generated image card UI

* Studio: refine generated image overlay and edit context

* Studio: animate generated image loading state

* Studio: bind generated-image edits to selected image

* Studio: drop empty external assistant payloads

* Studio: guard generated image clipboard MIME

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Simplify image generation edit wiring

* Remove temporary image generation test changes

* Preserve explicit image edit references

* Scope image edit references to threads

* Harden OpenAI image edit replay handling

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Update image generation tool event test

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-26 02:34:06 -07:00
Daniel Han
c43c48a7e0 Merge branch 'main' into feat/expose-sampling-params-core
Resolve 5 conflicts where main added Anthropic Opus 4.6/4.7 fast-mode
support that touches the same signatures PR 5711 extended:

  studio/backend/core/inference/external_provider.py
    Keep both PR 5711 sampling fields (frequency_penalty, seed, stop,
    service_tier, parallel_tool_calls) and main's fast_mode in the
    stream signature, docstring, and _stream_anthropic call site.
  studio/backend/models/inference.py
    Append fast_mode Field alongside PR 5711's new ChatCompletionRequest
    fields; both flow through the existing dispatch.
  studio/backend/routes/inference.py
    Forward fast_mode and the PR 5711 sampling fields to the stream
    generator together.
  studio/frontend/src/features/chat/types/api.ts
    Add fast_mode? to OpenAIChatCompletionsRequest after the PR 5711
    field block.
  studio/frontend/src/features/chat/utils/chat-settings-storage.ts
    Persist fastMode alongside seed / stop / serviceTier /
    parallelToolCalls.

No semantic changes to either feature surface. 161 backend routing
tests still pass; frontend tsc clean.
2026-05-26 07:22:06 +00:00
Daniel Han
542d74370f
Studio: pricing follow-up to #5690 (longest-prefix match + chat-style usage keys) (#5722)
* Studio: longest-prefix pricing match + accept chat-style usage keys

Two P1 / High follow-ups from PR 5690 review feedback:

1. Pricing prefix lookup returned the first key it iterated, so
   dated snapshots like ``gpt-5.4-mini-2026-04-23`` collided with
   the shorter ``gpt-5.4`` entry and overbilled by 3x+. Sort the
   table keys longest-first so the most specific entry wins.

2. ``calculate_cost`` only read ``input_tokens`` / ``output_tokens``,
   but Studio's OpenAI-Chat-style usage envelope re-emits
   ``prompt_tokens`` / ``completion_tokens`` (the OpenAI Chat
   Completions vocabulary). Callers handing in the chat-style
   shape silently got a zeroed bill. Accept either pair so the
   calculator works against both raw upstream usage and the
   Studio-translated envelope.

Tests (4 new in test_pricing.py): dated mini/pro snapshots inherit
the right rate; chat-style usage keys price correctly; raw key wins
when both shapes are present.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: dedupe cache buckets when costing chat-style Anthropic usage

When the caller hands in Studio's chat-style envelope (``prompt_tokens``
emitted by ``_build_usage_chunk``) for Anthropic, that value already
folds ``cache_creation_input_tokens`` + ``cache_read_input_tokens`` into
the total. The previous follow-up accepted the chat-style key but then
re-added both cache buckets in ``billable_input_tokens`` and ``input_usd``,
double-counting cache tokens on every Anthropic chat-style call.

Detect which envelope landed (``input_tokens`` present = raw upstream;
absent + ``prompt_tokens`` present = Studio chat-style) and peel the
cache buckets off for Anthropic before the downstream math so both
envelopes produce identical costs.

OpenAI: ``input_tokens`` and Studio's ``prompt_tokens`` both already
include ``cache_read`` and exclude any notional ``cache_creation``, so
the OpenAI path stays a straight passthrough.

Tests (2 new): both envelopes match for Anthropic on a triple
(uncached + cache_creation + cache_read); OpenAI envelopes match on a
cached-tokens fixture.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: prefer raw output_tokens over chat-style completion_tokens

Codex flagged that the previous fallback chain
'usage.get("output_tokens") or usage.get("completion_tokens")'
treats an explicit 0 as missing -- a mixed-envelope payload where
'output_tokens' is 0 but 'completion_tokens' is non-zero (or
stale) bills the wrong amount. Mirror the has_input_tokens
precedence pattern: when the raw key is present we use it even at
0; otherwise fall back to completion_tokens.

* Studio: read OpenAI cached tokens from prompt_tokens_details too

Codex flagged that the chat-style OpenAI envelope Studio re-emits
via _build_usage_chunk surfaces cached prompt tokens under
prompt_tokens_details.cached_tokens, not input_tokens_details. The
OpenAI branch only checked input_tokens_details, so a cache-heavy
chat-style turn billed every cached token at the full input rate
instead of the 0.1x cache_read discount.

Walk both keys when discovering the cached count. New regression
test pins that the two envelopes price identically for a turn with
80k of 100k tokens cached.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: tighten pricing prefix match + clamp corrupt usage

Three follow-ups on the longest-prefix pricing match landed in this PR:

- Prefix match now requires a dash boundary or end-of-string. The
  longest-key sort alone still falsely landed "claude-opus-4-15" on
  the "claude-opus-4-1" row, and "gpt-5.5-prod" on the "gpt-5.5-pro"
  row (a 6x overcharge). Demanding the next character be "-" rules
  out the lookalikes while keeping dated snapshots
  ("gpt-5.4-mini-2026-04-23", "claude-opus-4-7-20260414") landing on
  their canonical row.
- Clamp every token count to >= 0. A corrupted upstream payload
  (negative cached count, off-by-one in a fixture) could previously
  produce a negative bill that masked real spend in the session
  total tooltip.
- Tolerate a non-dict "cache_creation" (e.g. an upstream proxy
  folded the field down to a single int). The current code raised
  AttributeError mid-turn; now it falls back to the 5m-default
  bucket so the rest of the cost calculation still runs.

Adds tests/test_pricing_edge.py with 20 adversarial cases covering
the boundary check, negative / None / zero token values across both
envelopes, cache_read > prompt corruption, the OpenAI long-context
threshold crossover on cache-inflated billable input, malformed
sub-objects, and unknown-provider degradation. Combined suite is
51 tests, all green.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Surface Anthropic cache-read fallback and forward 1h breakdown

Two correctness gaps surfaced on the chat-style usage envelope:

1) Anthropic cache_read fell through to "uncached input" pricing when
   the envelope arrived without the native ``cache_read_input_tokens``
   key (e.g. via a proxy that only emits the mirrored
   ``prompt_tokens_details.cached_tokens`` block). Studio's canonical
   ``_build_usage_chunk`` always sets both so production traffic was
   never affected, but the calculator should accept either as a
   defense-in-depth measure. Add a fallback to read the mirrored
   field when the native one is missing or zero; the native key still
   wins when both are present so the math stays deterministic.

2) ``_build_usage_chunk`` dropped the ``cache_creation`` 5m / 1h
   breakdown. Downstream ``calculate_cost`` then could not apply the
   2x 1h premium and silently fell back to the 5m default,
   underbilling 1h cache writes by 2x on chat-style traffic. Forward
   the breakdown verbatim when the upstream usage carries it.

Tests grow by 4 (20 -> 24): two for the prompt_tokens_details
fallback (with native-precedence pin), one for the chunk shape, one
for the end-to-end pricing parity check at 1h.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Add Anthropic fast_mode pricing multiplier

PR 5715 wires the fast-mode-2026-02-01 beta header + speed:"fast"
field through to Anthropic, but the cost calculator never learnt
about the matching 6x premium documented at
https://platform.claude.com/docs/en/build-with-claude/fast-mode
(Opus 4.7 standard $5/$25 per MTok, fast $30/$150).

This adds:
- ANTHROPIC_FAST_MODE_MULT = 6.0 constant.
- calculate_cost(..., fast_mode=True) applies the 6x to base input
  AND output rates before any cache multipliers (cache mults stack
  on top of fast per Anthropic docs).
- Provider+model gate: silently no-op on every model that is not
  claude-opus-4-6 / claude-opus-4-7 so a stray fast_mode=True on
  Sonnet/Haiku can never over-charge.
- model_priced label tagged "(fast)" so the cost tooltip can
  surface which rate fired.
- pricing_snapshot now exposes fast_mode_mult so the frontend cost
  panel doesn't have to hard-code 6.

7 new edge tests pin the math; existing 55 still pass.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Honor explicit zero cache_read_input_tokens on Anthropic envelopes

The previous follow-up fell back to ``prompt_tokens_details.cached_tokens``
whenever the native ``cache_read_input_tokens`` was missing OR equal to 0,
even though the commit message stated the native key always wins when
present. A proxy that forwards a stale ``prompt_tokens_details`` block
alongside an authoritative ``cache_read_input_tokens: 0`` would then
inflate cache_read past the real native count, posting a false cache_read
line and bumping billable_input_tokens. Switch the gate to native-key
presence so an explicit zero stays authoritative; the mirror only kicks
in when the native key is absent. Add a regression test pinning the
explicit-zero precedence.

* Move fast_mode pricing back to #5715

The fast_mode 6x multiplier landed in two places at once -- here
(f66df7ba) and on #5715 (4f1afdb5) -- since both audits ran in
parallel. Drop the duplicate from this branch so the change lives
in its natural home (#5715, which introduces fast_mode itself);
this PR stays focused on the cache-read fallback + 1h breakdown.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Shorten pricing comments for PR #5722

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-25 23:39:58 -07:00
Daniel Han
4854d4579f
Studio: surface Anthropic document citations inline + in Sources panel (#5718)
* Studio: surface Anthropic document citations inline + in Sources panel

Anthropic's Messages API streams ``citations_delta`` events on
``content_block_delta`` when the request enables
``citations: {enabled: true}`` on document blocks. Each event carries
one citation pointing at the source document; previously they were
silently dropped, so reader-visible references never reached the chat
UI even when the model was citing properly.

The proxy now:
- dedupes by the type-specific anchor (char_location / page_location /
  content_block_location / search_result_location) so re-cites of the
  same span collapse onto a single footnote;
- injects ``[N]`` inline right after the matching text run;
- forwards the full list as a synthetic ``document_citations``
  tool_event at ``message_stop`` so the Sources panel can render
  per-document footnotes next to web_search / web_fetch citations.

Streams that never emit ``citations_delta`` stay byte-identical.

References:
- https://platform.claude.com/docs/en/build-with-claude/citations
- https://platform.claude.com/docs/en/build-with-claude/search-results

Tests (5 in test_anthropic_citations.py): passthrough, single
char_location, dedup of repeat citations, distinct sources get
distinct numbers, search_result_location supported.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: surface Anthropic document_citations in the Sources panel

The PR added a backend _toolEvent.type='document_citations' on
message_stop and an inline [N] marker in the assistant text, but the
chat-adapter only handles container_*/tool_*/sources from
web_search and web_fetch tool calls. Reviewers flagged that the
inline [N] markers had no matching footnote entries in the Sources
panel.

Capture the new event into a documentCitationParts buffer, convert
each citation dict into a Sources-panel source entry (using
document_title or search-result source URL plus cited_text as the
snippet), dedupe by id, and append to the final yield alongside
the existing web_search/web_fetch sourceParts.

* Studio: dedupe search_result_location citations by search_result_index

Anthropic's documented search_result_location citation shape carries
search_result_index, source, title, and start/end_block_index --
NOT document_index/document_title. The previous key keyed on
document_index + document_title + source + start_block_index, so
two distinct search results from the same source collapsed onto the
same footnote and the second [N] marker was lost.

Switch the search_result_location branch to key on the documented
fields, and pin the behaviour with a regression test asserting that
two citations sharing source/title but with different
search_result_index get distinct [1] [2] markers.

* Studio: keep each citation distinct across the end-anchor

Codex follow-ups on the citations PR:

  * Backend _anthropic_citation_key now includes the end anchor for
    every variant (end_char_index, end_page_number,
    end_block_index). Anthropic ranges are start-AND-end pairs, so
    a same-start / different-end pair is two distinct citations
    that previously collapsed onto one footnote.

  * Frontend documentCitationToSource ids include the position
    fields (search_result_index, start/end char/page/block) instead
    of being keyed on URL alone. Two citations from the same
    document or two search_result_locations with the same source
    now produce distinct Sources-panel entries, matching the
    inline [N] numbering.

* Studio: key Sources list by per-citation id instead of url

Codex flagged that the Sources renderer keys badges on source.url,
so two Anthropic document citations sharing the same source URL
collide as React keys and one badge gets dropped (or duplicated).

The chat-adapter already mints a per-citation id that folds the
position fields (search_result_index, start/end char/page/block)
into the URL, so the two citations have distinct ids even when
their URL matches. Plumb that id through SourceData and use it as
the React key for both the measurement badges and the visible
SourceBadge list. Falls back to the URL when no id is supplied
(web_search and web_fetch source parts).

* Studio: enable Anthropic doc citations on input_document blocks

Plumb citations: {enabled: true} onto the translated Anthropic document
block (both base64 and URL source branches) so the upstream actually
emits citations_delta events. Without this opt-in the inline [N] +
Sources panel plumbing added in this PR is a no-op for real user
PDF / doc uploads.

Refs https://platform.claude.com/docs/en/build-with-claude/citations

Also add edge-case coverage for the citations_delta path:
malformed citations, mixed types per document, reversed indices,
missing document_index, non-int block indices, unknown citation
type, internal _key never leaking, footnote numbering across
content blocks, and the input_document wire-through itself.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Reject unsafe citation sources, bound cited_text payload

Three follow-ups on top of #5718 surfaced by a deeper review pass:

1) javascript: / data: / vbscript: in citation source is XSS-able.
   ``documentCitationToSource`` was assigning ``cit.source`` straight
   into ``Source.url`` and rendering it as an <a href>. A hostile
   model emitting ``cit.source = "javascript:alert(document.domain)"``
   would execute on click (openLink only intercepts URLs that contain
   "://" or start with "mailto:", which both miss the javascript:
   scheme). Restrict the navigable path to http(s):// only; anything
   else falls back to the existing #anthropic-doc anchor and the
   source title still renders the raw identifier for context. Also
   reject CR/LF inside the URL string.

2) Frontend sources collapse distinct backend footnotes when the
   citation type differs but positions match. char_location(0,5) and
   page_location(0,5) over the same source previously deduped into
   one entry because the id only carried position. Fold citation
   type into the id anchor so the 1:1 mapping with inline [N]
   markers is preserved across every citation shape.

3) ``cited_text`` was forwarded unbounded inside the synthetic
   document_citations tool_event. The Sources panel trims to 240
   chars for display anyway; for large RAG / search_result spans
   (~10kB cited_text is plausible) this inflates SSE bytes 40x
   for no UI benefit. Truncate server-side at 512 chars with an
   ellipsis so the description-trim downstream still has room to
   work and the wire stays bounded.

Tests grow from 21 to 22; existing 7 + edge 15 still green. Frontend
typecheck clean.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: apply http(s) URL guard to all Sources-panel link sources

The previous round only filtered ``cit.source`` inside
``documentCitationToSource``. Two parallel code paths still copied
provider/tool-controlled ``URL:`` text directly into clickable
``<a href>`` Sources-panel links:

  * ``parseSourcesFromResult`` in chat-adapter.ts (legacy web_search /
    web_fetch tool result parser)
  * ``parseSearchResults`` in tool-ui-web-search.tsx (inline tool card)

A hostile tool response like ``URL: javascript:alert(1)`` or
``URL: data:text/html,...`` was therefore still rendered as a
navigable badge in the Sources panel.

Centralise the safe-URL test (``isSafeNavigableSourceUrl``,
``isSafeHttpUrl``) using ``new URL()`` + protocol allowlist + CR/LF
rejection, and apply it to both parsers. Unsafe blocks are dropped
rather than rewritten to a hash anchor because the web_search /
web_fetch parsers have no document-index fallback.

Citation conversion now uses the same helper so the in-place
http(s) regex and CR/LF check stay in one place.

* Shorten citation comments for PR #5718

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-25 23:39:02 -07:00
Daniel Han
7d3c472461
Studio: standalone Fetch pill for Anthropic web_fetch (#5742)
* Studio: surface Anthropic web_fetch as a standalone Fetch pill

web_fetch used to be silently bundled with the Search pill on the
assumption that "search returns URLs, fetch reads them" is the
typical workflow. Two problems with that:

- Anthropic bills each web_fetch invocation separately from
  web_search hits, so combining them made the per-message cost
  surface ambiguous.
- It blocked "just fetch this one URL" workflows where the user
  already knows the page they want read and does not want a search
  round-trip.

Adds:

- `webFetchToolsEnabled` to the chat-runtime-store, persisted to
  localStorage under `unsloth_chat_web_fetch_tools_enabled`, with a
  matching `supportsBuiltinWebFetch` capability flag and a
  `setWebFetchToolsEnabled` setter.
- A new Fetch pill in the chat composer, rendered next to Images and
  only when the active provider returns true from
  `providerSupportsBuiltinWebFetch` (Anthropic today). The pill
  defaults off so per-fetch billing is always a deliberate opt-in.
- chat-page bootstraps `webFetchToolsEnabled` from the same stored-
  preference fallback the other pills use.
- chat-adapter reads `webFetchToolsEnabled` directly when deciding
  whether to append "web_fetch" to `enabled_tools`, decoupling it
  from `toolsEnabled` (Search).

Backend translation is unchanged: when `enabled_tools` already
contains "web_fetch", `_stream_anthropic` appends the
`web_fetch_20250910` / `web_fetch_20260209` tool exactly as before
(test_anthropic_web_fetch.py pins the standalone-only path at
`test_web_fetch_tool_appended_to_request_body` and the combined
path at `test_web_fetch_combined_with_web_search_and_code_execution`).
Frontend tsc passes.

* ci: re-trigger after transient GitHub API HTTP flake (checkout + ggml-org release fetch)

* Studio: include web_fetch in the disabled-tool guard axis

Reviewer P1 / High on PR #5742 (codex + gemini): after introducing
the standalone Fetch pill, `disabledToolGuard` still only branched on
`webSearchEnabledForThisTurn`. With Fetch ON and Search OFF the
system prompt would tell Claude "you do not have web search or web
fetch tools in this conversation", which contradicts the actual tool
schema being sent and suppresses `web_fetch` tool calls, defeating
the standalone-fetch workflow this PR adds.

Treat search and fetch as a single "any web tool enabled" axis. The
guard only needs to warn the model when no web tool is wired in for
this turn; once either pill is on the model can pick the right one
from the tool schema. The existing `webLabel` already covers both
names, so the user-visible guard text stays accurate in every
combination.

tsc clean.

* ci: re-trigger after transient infra flake on Windows prebuilt / actions/checkout

* Studio: route web_fetch through per-model version dispatch

The web_fetch tool body in `_stream_anthropic` hardcoded
`web_fetch_20250910` instead of calling `_anthropic_web_fetch_version`,
so Opus 4.6 / 4.7 and Sonnet 4.6 missed the `web_fetch_20260209`
dynamic-filtering variant. The picker, the unit tests for it, and a
deliberate "follow-up" note in `test_anthropic_web_fetch.py` already
existed; this just threads it through the emission site.

Mirrors how web_search and code_execution are dispatched per model.
Old models still resolve to `web_fetch_20250910` and continue to work.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Shorten web_fetch comments for PR #5742

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-25 23:37:26 -07:00
Daniel Han
063e1e497b
Studio: rewrite OpenAI Responses citation markers to markdown links (#5713)
* Studio: rewrite OpenAI Responses citation markers to markdown links

OpenAI's /v1/responses stream interleaves text deltas with inline
citation markers built from private-use codepoints (U+E200 / U+E201 /
U+E202) shaped like `citeSOURCE_ID`. The codepoints render
as garbled "E202" glyphs or empty boxes in most fonts, and the
markdown layer further strips them, leaving run-on text like
"citeturn1view0turn1view1turn3view0...". The url list still arrived in
the Sources panel via url_citation annotations, but the inline cite
hand-off into the prose was unreadable.

Rewrite each marker into `[N](URL)` when the matching url_citation
has already been recorded on this stream, and drop the marker
silently otherwise. The lookup uses a new `source_id` field captured
on `_record_url_citation` (accepts source_id / id / locator across
Responses API revisions). Annotations are now applied BEFORE the
delta text is rewritten so that markers and their resolving
annotation arriving in the same SSE event still resolve.

Reference: https://developers.openai.com/api/docs/guides/citation-formatting

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Preserve every source_id alias for a deduplicated url_citation

OpenAI's Responses stream cites the same URL under multiple
source_id markers when the model references different spans of the
same page. The previous dedup-by-URL kept only the first alias and
dropped the rest, so subsequent markers for the same URL never
resolved and got stripped from the prose. Switch the citation
record to a ``source_ids`` list and append new aliases on every
duplicate. The rewriter resolves any alias back to the same
citation number so the inline markers all collapse onto one footnote
rather than fanning out into bogus repeats.

Also collapse the two passes over ``all_url_citations`` in
``_record_url_citation`` into a single loop for clarity. Adds two
regression tests covering the alias-collision and mixed-shape cases.

* ci: re-trigger after flake in Studio GGUF Tool calling (rebased on main #5741 already)

* ci: re-run after transient CodeQL Python checkout auth flake

* Fix split-marker buffer + multi-source ids for PR #5713

The original rewriter only handles markers that arrive whole inside a
single response.output_text.delta event. OpenAI's stream chunks text
on byte-buffer boundaries with no awareness of the marker grammar,
so a marker can straddle two deltas (delta-1 ends with
"citetu", delta-2 starts with "rn0view0"). Each delta
was rewritten in isolation, so the half-marker leaked as garbled
"E200/E202" glyphs in the rendered prose.

Buffer the unterminated tail across deltas and concatenate it onto
the front of the next one so the rewriter sees a complete marker.
Flush the held-over tail on response.completed / response.incomplete /
[DONE], stripping any leftover private-use bytes so a never-closed
marker (truncated stream, missing annotation) never leaks.

Also handle the multi-source marker shape from the OpenAI docs --
citeid1id2 should expand to one bracket
link per resolvable id. The previous regex captured only the first
source id and silently dropped id2/id3.

Reference: https://developers.openai.com/api/docs/guides/citation-formatting

Tests: 21 new cases covering multi-source, locator suffix, marker
split across two and three deltas, unterminated marker on truncation,
late annotation resolving a buffered marker, idempotency, and the
head/tail split helper directly.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Defer citation segments until url_citation annotation arrives

The split-marker buffer already concatenates a marker that straddles
two response.output_text.delta events. But when the annotation event
for a url_citation arrives AFTER the delta that contains its inline
marker (the typical OpenAI Responses ordering), the rewriter still
saw an empty lookup table at delta time and silently stripped the
marker. The URL kept showing up in the sources panel but the inline
link reference was permanently gone.

Add _rewrite_citation_markers_partial which leaves an unresolved
marker verbatim and reports has_unresolved=True. The streaming loop
buffers any closed segment that contains an unresolved marker into a
pending_citation_segments FIFO and drains the queue on every later
annotation event, on response.completed, on response.incomplete, and
on the [DONE] sentinel. Drain order is preserved so later clean text
does not leapfrog an earlier deferred segment. End-of-stream forces a
strip so no codepoint leaks if the annotation never arrived.

Add six regression tests covering single-pass resolution, the late-
annotation two-pass case, multi-source markers with partial
resolution, mixed known and pending markers in one segment, and
idempotency on marker-free input.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Drop unterminated citation tail to prevent cite-prefix plain-text leak

`_flush_pending_marker_tail` stripped the three private-use citation
codepoints from the held-over buffer, but left the literal ``cite``
keyword plus the source id behind as plain text. A stream ending
mid-marker therefore emitted user-visible garbage like
``Some text citeturn0view0`` instead of the intended clean prose.

``pending_marker_tail`` is by construction the suffix that starts at
an unclosed ``\\ue200`` opener -- the split helper guarantees there is
no closing ``\\ue201`` byte. Without that close the marker is
meaningless: the source id cannot be resolved to a URL and the user
prose before the opener was already emitted as ``head`` on the
originating delta. Bail out before the strip step and return the
empty string. As a belt-and-braces measure also drop any orphan
``cite<sid>`` literal at the head of the buffer in case a future
caller passes a partially-terminated tail.

Update the matching ``_simulate_delta_stream`` harness in the edge
tests so it mirrors the new flush logic, and add four regression
tests covering unterminated marker with surrounding prose, marker-
only inputs, prefix-only outputs, and the split-then-close path that
still must resolve to a link.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Defer multi-source markers until all ids resolve for PR #5713

`_rewrite_citation_markers_partial` previously treated a marker as
resolved when even one token in a multi-source marker resolved,
dropping any still-pending source ids. In streamed Responses events
the annotations for a multi-source marker can arrive across separate
`annotation.added` chunks, so the caller no longer buffered that
segment for retry and the late source id was lost from the inline
citation entirely.

Flag the marker unresolved whenever any token misses the lookup so
the streamer keeps the segment pending. End-of-stream force flush
still drops unresolved tokens through `_replace_openai_citation_markers`
so locator-style suffixes (which look like unresolved ids at the token
level but only appear at end-of-stream) render cleanly.

Updated the multi-source test to assert the new pending-then-flush
behavior; locator output now lands at force-flush rather than mid
stream.

* Shorten citation marker comments for PR #5713

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-25 23:37:16 -07:00
Daniel Han
7d1b68079e
Studio: Anthropic fast_mode toggle and streaming refusal handling (#5715)
* Studio: add Anthropic fast_mode toggle + surface streaming refusals

Fast mode (beta `fast-mode-2026-02-01`) lets Claude Opus 4.6 and 4.7
generate output tokens up to 2.5x faster at 6x standard Opus
pricing. The toggle lives in Configuration → Provider when the
selected Anthropic model is Opus 4.6 or 4.7 and is otherwise
hidden. Backend gates the same prefixes a second time so a stale
frontend cannot make Anthropic 400 the request, and the
`fast-mode-2026-02-01` beta header is merged onto whatever other
betas the request already needed (code-execution, compaction).

Streaming refusals (`message_delta.delta.stop_reason="refusal"` on
Claude 4 models) now surface a short user-facing notice in the
assistant message before the translated OpenAI chunk emits the
existing `finish_reason="content_filter"`. Previously the chat
bubble truncated silently because the SSE stopped mid-stream with
no visible explanation. Per the upstream docs the conversation
must be reset before continuing, so the notice tells the user
exactly that.

Reference:
- https://platform.claude.com/docs/en/build-with-claude/fast-mode
- https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/handle-streaming-refusals

Tests:
- studio/backend/tests/test_anthropic_fast_mode_and_refusal.py (8 cases
  pinning fast_mode pass-through on 4.6/4.7, silent drop on Sonnet /
  Haiku / older Opus / None / False, and the refusal notice + finish
  reason on a synthetic refusal stream).

* Studio: drop refused Anthropic turns from the next request

Anthropic's streaming-refusal guidance says the refused assistant
turn must be removed or updated before the next call -- otherwise
the safety classifier keeps refusing. The PR only added a
user-visible notice; the partial assistant output (plus the notice
itself) still rode the next request via toOpenAIMessage.

Tag the refusal turn with an HTML-comment sentinel emitted alongside
the notice. The chat-adapter checks for that sentinel in
toOpenAIMessage and returns null, so the refused turn is excluded
from outboundMessages. The notice still renders in the transcript
(HTML comments don't display), so users keep the explanation.

* Studio: filter None finish_reason entries in test helper

test_refusal_maps_to_content_filter expects only ['content_filter']
in the finish_reasons list, but the post-PR refusal path emits a
user-visible content notice chunk first. Every _content_chunk
carries 'finish_reason: None' by construction; the helper was
appending those, so the assertion saw [None, 'content_filter']
instead of ['content_filter'].

None is not a finish reason -- it's just mid-stream delta noise.
Skip None values in _finish_reasons so the helper reflects what
the test names actually claim to check. Same fix applies cleanly
to the other helper usages (pause_turn test expects [] and the
sibling stop test expects ['stop'], both unaffected).

* Studio: cover Anthropic fast-mode edge cases

Adds 19 cases on top of the 9 in test_anthropic_fast_mode_and_refusal.
The base file pins the happy path; this file fills in the cliffs:

* Dated-snapshot prefix matching: claude-opus-4-7-2026-02-01 and
  claude-opus-4-6-2026-02-01 still gate fast_mode through, while
  claude-opus-4-5-2025-08-01 and claude-sonnet-4-6-2026-02-01 do not.
* Strict opt-in: a future claude-opus-4-8 or claude-opus-5 does NOT
  auto-enable fast_mode -- the prefix tuple must be bumped explicitly
  when a new family is whitelisted upstream.
* Beta-header merge: fast_mode coexists with code-execution-2025-08-25
  and compact-2026-01-12 in one comma-separated anthropic-beta header
  with no duplicates and no truncation. Pins the value to the exact
  fast-mode-2026-02-01 docs token so a typo would fail CI.
* Non-destruction: fast_mode=None produces byte-identical outbound
  body and headers to the version that omits the argument entirely.
  Same for fast_mode=False. Guarantees the upgrade path is
  non-breaking on existing Anthropic streams.
* Refusal stream ordering: the user-visible notice precedes the
  finish_reason chunk so a streaming UI paints text before flipping
  to content_filter. Refusal sentinel emitted exactly once. Notice
  rides a normal content delta chunk with finish_reason still null.
  Partial assistant deltas survive before the notice.
* Provider-side refusal coverage: a refusal on Sonnet (not just Opus)
  still emits the notice + sentinel + content_filter mapping, since
  refusal handling is not gated on fast-mode capability.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Persist fastMode, drop refused user message on retry

Two follow-ups on #5715:

1) sanitizeInferenceParams stripped fastMode. fastMode is in
   PERSISTED_INFERENCE_PARAM_KEYS but the storage sanitizer only kept
   numeric fields plus systemPrompt and trustRemoteCode, so the new
   toggle was silently dropped on reload and on the
   /api/chat/settings round-trip. Save it the same way trustRemoteCode
   is saved.

2) Refusal recovery now also drops the triggering user turn.
   Returning null from toOpenAIMessage on the assistant side left the
   user prompt that caused the refusal in the outbound history, so
   the very next request would re-trigger the same classifier.
   Anthropic's refusal-handling guidance is explicit on this: remove
   the refused turn AND the user message that triggered it before
   the next call. Implemented via a pre-pass that pops the trailing
   user message when an assistant carries the refusal sentinel.

Typecheck clean.

* Studio: out-of-band refusal signal + fast-mode prefix/usage/pricing fixes

The text sentinel for the Anthropic refusal drop signal was spoofable:
any assistant message containing the literal
<!--studio:anthropic-refusal--> would prune the prior user + assistant
pair on the next request. Move the signal onto a separate _toolEvent
chunk that the chat adapter latches into
assistant.metadata.custom.anthropicRefusal; assistant text can no
longer control the pruner.

Tighten the fast-mode model gate (backend + frontend) to require a "-"
family boundary so claude-opus-4-70 / claude-opus-4-7b style IDs do
not get speed: "fast" on a naive startswith match.

Use survivingMessages for the image / audio attachment scan so a
refused user turn does not gate or mis-attribute the next non-refused
turn.

Propagate Anthropic usage.speed onto the OpenAI-style usage chunk and
apply the documented 6x fast-mode multiplier in the cost calculator
(stacks with prompt-cache multipliers per the docs); expose the new
multiplier on the pricing snapshot for the UI tooltip.

Tests cover the tool-event chunk shape, the prefix-collision rejects,
usage.speed propagation, the 6x pricing math, and that the visible
refusal text carries no embedded sentinel.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Shorten fast-mode and refusal comments for PR #5715

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-25 23:37:12 -07:00
Daniel Han
cc68720385
Studio: surface external-provider cache hits and writes in context bar (#5736)
* Studio: surface external-provider cache hits and writes in context bar

The Anthropic / OpenAI Responses streaming paths already emit an
include_usage-style SSE chunk carrying prompt_tokens_details.cached_tokens
and cache_creation_input_tokens / cache_read_input_tokens (see
_build_usage_chunk in external_provider.py), but the chat-adapter only
read the local llama-server timings.cache_n field. As a result, the
context-usage tooltip never showed cache hits or writes for external
providers, even though the backend was computing them.

Read the external usage envelope as a fallback when timings.cache_n is
absent, and surface Anthropic cache_creation_input_tokens as a separate
"Cache writes" line in the tooltip so users can tell a cache miss from a
cache hit on a turn that both reads and writes the cache.

- ServerUsage gains optional prompt_tokens_details.cached_tokens,
  cache_creation_input_tokens, cache_read_input_tokens.
- contextUsage store entry gains optional cacheWriteTokens.
- ContextUsageBar gains optional cacheWrites tooltip line.
- chat-page wires both fields through to the bar.

* Studio: render cache stats for external providers too

Reviewer round on the original PR caught three asymmetric-fix sites
where the producer side surfaced external prompt-cache stats but the
consumer side still gated on ggufContextLength (which is only ever set
for the local llama-server runtime). Result: the entire cache-stats
PR shipped invisible for Anthropic / OpenAI Responses / Gemini, which
is exactly the set of providers it was added for.

- chat-page.tsx: drop the ggufContextLength precondition on the
  ContextUsageBar mount. The bar already tracks usage; let it decide
  what to render based on what it knows.
- context-usage-bar.tsx: make `total` optional. When absent, drop the
  "/ total" ratio + percentage progress bar + "approaching limit"
  helper, and just show per-turn counters + cache stats. Bootstrap
  guard tightened so an all-zero, all-undefined state still renders
  nothing.
- runtime-provider.tsx: external-provider rehydration was rejected by
  the `store.ggufContextLength` check. Keep the "fits inside window"
  sanity check when a local context window IS known, drop it when
  it isn't.
- message-timing.tsx: the per-message timing popover used a separate
  "Cache hits" code path that only read llama-server's timings.cache_n.
  Fall through to custom.contextUsage for external providers, and add
  a parallel "Cache writes" line for Anthropic cache_creation events.

* Studio: tighten cache-stats comments

* Scope contextUsage to active checkpoint

Three follow-ups on #5736 so the relaxed external-provider render
gate does not show stale token / cache stats from a different model:

1) setCheckpoint now clears contextUsage on a real checkpoint
   change. setActiveThreadId and clearCheckpoint already did this;
   the most-traveled transition path (the user switching models from
   the picker) leaked the prior turn's counts because they were never
   cleared.

2) The external-selection branch in chat-page.tsx now also clears
   contextUsage at the same time it nulls ggufContextLength /
   activeNativePathToken. Without this an in-session switch from a
   local model to an external provider would visibly carry the
   previous local turn's counters into the new provider's bar.

3) exitCompare's rehydration is now scoped: restore the saved
   usage only when the message's modelId matches the active
   checkpoint AND, for local turns where a context window is known,
   when the saved total fits inside that window. Without this the
   bar could render a stale local-model usage on top of an external
   provider, or an oversized usage object that exceeds the now-
   active window.

Typecheck clean.

* Plug remaining stale-contextUsage paths

Follow-up to 042e0ac4 that catches four asymmetric-fix sites the
checkpoint-scoping pass missed:

1) setParams now also clears contextUsage on a real checkpoint
   change. The local model load path in use-chat-model-runtime calls
   setParams(mergeBackendRecommendedInference(...)) which mutates
   params.checkpoint before refresh() eventually fires setCheckpoint;
   the intermediate window rendered the previous model's counters
   under the new checkpoint.

2) chat-adapter.ts setContextUsage on stream completion now gates on
   the captured params.checkpoint still being active. A late
   completion from provider A used to clobber the context bar after
   the user switched to provider B mid-stream.

3) chat-page.tsx exitCompare rehydration no longer accepts a saved
   modelId-stamped usage when the active checkpoint is empty. A user
   who entered compare, cleared the model, and exited compare would
   otherwise see the cleared model's stats reappear.

4) runtime-provider.tsx thread-load no longer restores legacy
   unscoped usage (no modelId) unless a local context window is
   known. With the relaxed external-provider render gate, old
   pre-PR persisted messages without a modelId stamp could attach
   their counts to an unrelated active provider.

Also switches message-timing.tsx cache-hit fallback from || to ??
so an explicit cache_n=0 is not replaced by a stale cachedTokens.

Typecheck clean.

* Shorten cache-stats comments for PR #5736
2026-05-25 23:37:04 -07:00
Daniel Han
034ff512e7
Studio: stop seeded admin to cross-origin callers (#5739)
* Studio: stop leaking seeded admin pw to cross-origin callers

The "/" SPA fallback serves index.html with an inline
``window.__UNSLOTH_BOOTSTRAP__`` script containing the seeded admin
password while a password change is pending. Default web mode runs
``CORSMiddleware`` with ``allow_origins=["*"]`` + ``allow_credentials=
True``, which reflects an attacker-controlled ``Origin`` back on every
request and sets Access-Control-Allow-Credentials true. The combination
let any cross-origin page ``fetch('/')`` with credentials and read the
bootstrap admin password out of the HTML body. The API smoke
``CORS: GET / leaks bootstrap pw to cross-origin caller`` audit already
tracked this (tests/studio/studio_api_smoke.py:224) but did not gate CI.

Gate ``_inject_bootstrap`` on a same-origin check: legitimate top-level
navigations omit ``Origin`` on most engines, so the absence of the
header is treated as same-origin; when the header IS present and does
not match ``request.url.scheme://request.url.netloc`` exactly, we now
skip injecting the bootstrap tag. ``Vary: Origin`` is added so an
intermediary cache cannot serve a same-origin response (with bootstrap)
to a later cross-origin caller (and vice versa).

Coverage: ``test_index_bootstrap_origin.py`` exercises the helper with
missing / matching / evil / scheme-mismatch / port-mismatch origins.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: tighten comments on bootstrap cross-origin helper

* Studio: canonicalise Origin before same-origin gate

A plain string-compare between the Origin header and request.url.netloc
misclassifies legitimate same-origin requests as cross-origin in three
scenarios:

- Browser strips the default port from Origin (https://example.com)
  but Starlette's netloc keeps it (example.com:443). Per RFC 6454 the
  default port is dropped on the wire, so the strings will not match
  even though the requests share an origin.
- Host case differs (Origin: http://Example.com vs netloc:
  example.com). Per RFC 3986 host comparison is case-insensitive.
- Scheme case differs (HTTP:// vs http://). Per RFC 3986 the scheme
  is also case-insensitive.

These are usability degradations rather than security gaps (legitimate
user denied the bootstrap injection, no attacker gain), but worth
shipping so non-default Studio deployments keep the change-password
auto-fill.

Adds _canonical_origin(scheme, netloc) -> (scheme, host, port) and
compares the canonical tuples. Default-port lookup covers
http/https/ws/wss; userinfo (user:pass@) is stripped per RFC 3986
since Origin never carries credentials. Origin: "null" (sandboxed
iframes, file:// pages) and unparseable values collapse to cross-
origin so the bootstrap pw is never leaked through those paths either.

Tests: 14 cases (was 5). Covers the original same/missing/evil/
scheme/port matrix plus default-port stripping in both directions,
host + scheme case folding, Origin: null, garbage values, and
userinfo-in-netloc.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix IPv6 netloc parsing for PR #5739

The canonical-origin helper used ``netloc.partition(":")`` which
mis-parses bracketed IPv6 hosts (``[::1]:8902`` -> host=``[``,
port-str=``:1]:8902``). The int() then raises and the canonicaliser
returns None, so every IPv6 same-origin request is misclassified as
cross-origin and Studio refuses to inject the bootstrap pw on a
legitimate top-level nav when launched with ``unsloth studio -H ::1``.

Bracket-aware split per RFC 3986 §3.2.2, plus extra regression tests
for IPv6, opaque (data:/blob:/file:), comma-joined multi-Origin and
localhost-vs-127 cases.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Guard urlparse ValueError in same-origin gate

urlparse raises ValueError on malformed bracketed Origin values
(unclosed [, invalid IPv6 hex, text after ]) and on a few NFKC
edge cases since Py 3.8. Without a guard, a request carrying
Origin: http://[malformed surfaced as HTTP 500 from the SPA
handler rather than being treated as cross-origin per the
docstring's safer-default rule. Wrap both urlparse calls in
try/except ValueError and return False on parse failure.

Also distinguish a missing Origin header (top-level same-document
GET, treat as same-origin) from an explicit empty string (not a
valid serialised origin per RFC 6454 §6.1, treat as cross-origin).

Four new regression tests pinned down by the PR audit: malformed
IPv6 bracket, invalid IPv6 hex, bracket with trailing garbage,
and the empty Origin header.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Shorten origin-gate comments for PR #5739

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-25 23:36:51 -07:00
Daniel Han
4f9125a5f9 Pin Claude 4.7 sampling-removed regex with backend test (PR #5711)
Guards against drift between the backend _ANTHROPIC_4_7_SAMPLING_REMOVED
regex and the frontend ANTHROPIC_4_7_SAMPLING_REMOVED_REGEX added in the
previous commit. If a future patch widens one without the other the panel
will either silently strip a knob the user moved or 400 on a knob the UI
should have hidden — both bad UX.

Test pins the canonical 4.7 id shapes (opus/sonnet/haiku) including dated
snapshots, and the non-4.7 ids that must NOT match (4-6, 4-5, 4-71,
future 5, gpt-4o, etc.). Sampling-params suite now 60 passing
(previously 59).
2026-05-26 05:48:17 +00:00
Daniel Han
643c0a88f1 Per-model OpenAI / Anthropic 4.7 sampling gating (PR #5711)
OpenAI gating was per-provider — the restrictive reasoning-class capability
applied to gpt-4o too, even though gpt-4o on /v1/responses still accepts
temperature / top_p / seed / frequency_penalty / presence_penalty. Anthropic
4.7 was the inverse: backend stripped temperature / top_p / top_k per-model
but the UI still showed the sliders, so moving a knob silently did nothing.

Split openai capabilities into OPENAI_REASONING_CAPABILITIES (current
restrictive set, used for gpt-5.x / o1 / o3 / o4) and OPENAI_CHAT_CAPABILITIES
(full sampling minus top_k and stop, used for gpt-4o and any non-reasoning id
from the registry). Mirror the backend _ANTHROPIC_4_7_SAMPLING_REMOVED regex
on the frontend so claude-(opus|sonnet|haiku)-4-7 hides temperature / top_p /
top_k in the panel instead of relying on backend strip. getProviderCapabilities
now takes an optional modelId so chat-settings-sheet and chat-adapter both
resolve the same per-model variant; no behavior change for unspecified modelId.

Verified live OpenAI / Anthropic docs:
  - GPT-5 temperature must equal 1: platform.openai.com/docs/guides/reasoning,
    community.openai.com/t/temperature-in-gpt-5-models/1337133
  - GPT-4o accepts full sampling on Responses: docs.aimlapi.com gpt-4o ref,
    OpenAI cookbook seed example
  - Claude 4.7 sampling removed (400 on any non-default temperature/top_p/
    top_k): platform.claude.com/docs/en/about-claude/models/whats-new-claude-4-7

160 backend routing tests still pass; frontend tsc clean.
2026-05-26 05:46:14 +00:00
pre-commit-ci[bot]
b73480e554
[pre-commit.ci] pre-commit autoupdate (#5773)
updates:
- [github.com/astral-sh/ruff-pre-commit: v0.15.13 → v0.15.14](https://github.com/astral-sh/ruff-pre-commit/compare/v0.15.13...v0.15.14)

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-25 22:21:43 -07:00
Ricardo-M-L
af6504f900
fix(chat_templates): check find() return value before slicing on placeholders (#5763)
* fix(chat_templates): check find() return value before slicing on placeholders

Two places in `construct_chat_template()` use `str.find()` for sentinel
placeholders (`{INPUT}` / `{OUTPUT}`) without checking the -1 return:

1. The `except:` fallback (around line 2464) computes
   `chat_template[chat_template.find("{OUTPUT}") + len("{OUTPUT}"):]`.
   If the template has no `{OUTPUT}` marker, `find()` returns -1 and the
   slice starts at offset 7 (`-1 + len("{OUTPUT}")`), producing garbage
   that's then `re.escape`-d and fed back into the template-recovery
   regex. The user sees a confusing `IndexError` on
   `response_part = response_part[0]` instead of the real problem.

2. The final trim before returning (`input_part[:input_part.find("{INPUT}")]`
   and the matching `{OUTPUT}` line) silently drops the last character
   when the placeholder is missing — `find()` returns -1, and `[:-1]`
   slices everything except the last character, returning a corrupted
   template prefix to the caller.

Replace both with an explicit `-1` check that raises a clear
`RuntimeError` naming the missing placeholder, matching the existing
guard pattern from #4923 (`try_fix_tokenizer`).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(chat_templates): also guard {INPUT} and fallback regex/separator paths

Builds on the {OUTPUT} / final-trim guards in this branch by closing
the three remaining ways the except-block fallback in
construct_chat_template() can still raise a confusing IndexError or
AttributeError on malformed templates:

1. Validate both {INPUT} and {OUTPUT} before deriving `ending`. The
   regex two lines later (`{INPUT} + ending + ...`) still produced an
   empty list and crashed on `response_part[0]` if {INPUT} was missing.
2. Guard the regex no-match case. Some templates contain both
   placeholders but not in a recoverable two-example shape, in which
   case `re.findall` returns an empty list and `[0]` raises.
3. Initialize `found = None` before the separator-search loop and
   raise if the loop never sets it. Previously, if the first
   iteration's `re.finditer` was empty the loop broke without binding
   `found`, and `found.group(1)` raised AttributeError on the stale
   int left over from the outer rfind loop.

Rephrase the final-trim error messages from internal variable names
("input_part") to user-facing wording ("instruction section") and
include a bounded (200-char) excerpt of the offending content so the
error is debuggable without being unbounded.

Add tests/python/test_construct_chat_template_validation.py covering
each failure mode with a fake tokenizer (no HF_TOKEN, no model
download, CPU-only).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-25 06:19:01 -07:00
Long Yixing
748aa1c482
fix: repair mlx studio base export save_method (#5727) 2026-05-25 04:04:07 -07:00
Daniel Han
5d4ddd6b37 Revert "Surface OpenAI Responses service_tier='scale' (PR #5711)"
Round 19 added scale to /v1/responses based on the openai-python
SDK type, but round 20 reviewers (3/10 against) and the round 18
aggregator both noted that the live OpenAI Responses reference
limits Responses service tiers to auto/default/flex/priority. The
PR contract in the original description also lists scale only for
Chat Completions, not Responses. Studio routes OpenAI through
Responses, so forwarding scale risks a 400 from the upstream and
exposes a picker option the API does not accept.

Restore the conservative drop behaviour: only documented Responses
tiers reach the wire; legacy scale settings still validate (the
ServiceTier Literal and chat-settings storage allowlist keep it
for forward-compat).
2026-05-24 20:25:55 +00:00
pre-commit-ci[bot]
f7c11d8a0a [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 20:11:43 +00:00
Daniel Han
b961c6f2f7 Surface OpenAI Responses service_tier="scale" (PR #5711)
Round 19 reviewer consensus (3/10 plus an asymmetric-fix call-out
across rounds 8/9/12/14/17): the openai-python SDK ships
service_tier as Literal["auto","default","flex","scale","priority"]
on /v1/responses, and enterprise Scale Tier customers need to opt in
explicitly. Drop the defensive scale-filter on the backend and add
"scale" to the OpenAI picker option list so the field reaches the
wire when set. Other providers remain at auto/default per their docs.

Update the routing tests so `scale` lives in the forwarded-set fixture
and the dropped-set fixture only carries truly out-of-enum values
(Anthropic-only `standard_only`, typos, empty string).
2026-05-24 20:09:17 +00:00
Daniel Han
07831d4c0c Shut down asyncgens in routing tests to silence cleanup warnings (PR #5711)
Round 18 reviewers (and earlier) noted CI noise from the routing test
helper: `_drive(coro)` ran a fresh event loop but never closed it or
called `shutdown_asyncgens`, so the MockTransport-backed httpx async
generators in the providers were finalised by GC in a later task and
emitted "Response.aiter_text.aclose was never awaited" / "Task was
destroyed but it is pending" warnings.

Explicitly close the loop after `run_until_complete`, running
`shutdown_asyncgens` first so the iterators finalise in this task.
Tests still pass and the warnings are gone.
2026-05-24 19:54:08 +00:00
Daniel Han
0be1e09421 Clean local stop list + cap Responses bridge parallel tool calls (PR #5711)
Round 17 reviewer consensus on two extensions of the round 16 cap:

1. Direct GGUF stop forwarding (3/10 + sibling findings) — every
   llama_cpp.py payload builder and routes/inference.py direct-GGUF
   call site pass `stop` through unfiltered, while the external
   provider helper and `_build_passthrough_payload` already strip
   empty / non-string entries. Add a shared `_clean_local_stop_list`
   helper in the route layer for the two callers, mirror the same
   inline filter in `llama_cpp.py`'s three payload builders. Stops
   `stop=["", "END"]` from a stale client 400'ing llama-server.

2. Responses bridge tool-call cap (3/10 streaming + 1/10 non-streaming)
   — `_responses_stream` iterated every streamed `delta.tool_calls`
   index and `_responses_non_streaming` translated every returned
   `message.tool_calls` entry, even when `parallel_tool_calls=false`.
   Latch the first tool-call index in the streaming bridge and drop
   subsequent siblings; cap to one in the non-streaming bridge.
   Matches the GGUF agentic-loop / Anthropic-passthrough caps.

Local-passthrough OpenAI paths (verbatim SSE / verbatim JSON) are
left alone because the contract is "raw upstream forwarding"; clients
calling /v1/chat/completions through Studio directly should still see
llama-server's native output.
2026-05-24 19:39:37 +00:00
Daniel Han
5218a01d14 Apply parallel_tool_calls cap to Anthropic passthrough + safetensors path (PR #5711)
Round 16 reviewer consensus extended the round 12c asymmetric-fix:
the GGUF agentic loop capped tool_calls to one when the caller opted
out, but every other Studio-internal path that emits tool calls from
llama-server output skipped the same guard.

Mirror the cap in three places that have full ownership of the
emitted list (passthrough verbatim paths are out of scope):

1. `AnthropicPassthroughEmitter` now takes `parallel_tool_calls` and
   silently drops streamed `delta.tool_calls` entries beyond the
   first index. Wired from `_anthropic_passthrough_stream`.
2. `_anthropic_passthrough_non_streaming` truncates the upstream
   `message.tool_calls` list before producing `tool_use` blocks.
3. `run_safetensors_tool_loop` truncates the parsed tool_calls list
   before appending the assistant message and executing tools.
   `InferenceOrchestrator.generate_chat_completion_with_tools` and
   the safetensors route now thread `parallel_tool_calls` through.

Also harden `_build_passthrough_payload` to strip empty / non-string
`stop` entries before forwarding to llama-server, matching the
`_normalize_stop_for_provider` shape the external-provider helper
already enforces.

Test pins the AnthropicPassthroughEmitter serial-tool-call gate.
2026-05-24 18:46:07 +00:00
Daniel Han
4c3be18d00 Preserve explicit serviceTier="auto" through the settings picker (PR #5711)
Round 13 P1 finding: the Service tier picker rendered `null` as the
displayed `auto` and converted any explicit `auto` selection back to
`null`, so the chat-adapter's truthy guard then omitted `service_tier`
on the wire. For Anthropic the docs distinguish:

  - omitting `service_tier`     -> provider default
  - `service_tier="auto"`       -> opts into Priority Tier when available
  - `service_tier="standard_only"` -> opts out

Drop the auto -> null conversion so the user's explicit pick reaches
the adapter and the wire reflects it. `null` still means "never set"
and falls through to the provider default; the existing serviceTier
allowlist already includes "auto" everywhere it matters.
2026-05-24 17:46:33 +00:00
Daniel Han
67e371934c Enforce parallel_tool_calls=False client-side on local GGUF (PR #5711)
Two related findings from round 12 reviewers:

1. The local GGUF tool loop in `generate_chat_completion_with_tools`
   iterates every entry of `tool_calls` returned by llama-server, even
   when the caller explicitly opted out of parallel tool calls. The
   `parallel_tool_calls` flag is forwarded to llama-server, but llama
   .cpp does not enforce it on every jinja template
   (https://github.com/ggml-org/llama.cpp/issues/22043), so a model
   that ignores the flag still ran multiple tools per turn. Cap
   `tool_calls` to the first entry when the flag is False so the
   client-side contract holds regardless of upstream behavior.

2. llama-server documents `parallel_tool_calls` as defaulting to FALSE
   (https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md),
   so the previous chat-adapter shape (forward only on explicit false)
   meant the UI's default-on state could never enable parallel tool
   calls there. Always forward the user's preference on the local
   path so the toggle actually does what it says. External providers
   default to true everywhere, so the external branch is unchanged.

Test pins the GGUF tool-loop cap by source-level assertion (the loop
itself is integration-only).
2026-05-24 17:35:05 +00:00
Daniel Han
737c5ad0e0 Allow whitespace stop sequences from the chips editor (PR #5711)
Round 11 P2 finding: the stop-sequence chips input rejected any draft
that strip to empty, which silently dropped pasted whitespace stops
like `"\n\n"` for blank-line halts. Local llama-server and OpenAI-
compat backends accept those; the Anthropic helper independently
filters whitespace entries before they hit the wire, so allowing them
in the UI cannot turn into a 400.

Drop the .trim() gate; reject only the truly empty draft. Single-line
Input behaviour is unchanged for the common typed-letters path.
2026-05-24 17:19:50 +00:00
Daniel Han
48df6a98c8 Forward disable_parallel_tool_use through Anthropic client-tool passthrough (PR #5711)
Round 11 reviewer consensus (10/10): the `disable_parallel_tool_use`
translation added in round 11b reached the Anthropic-compat server-tool
GGUF loop but not the analogous client-tool passthrough branch. A
client sending `/v1/messages` with custom tools plus
`tool_choice: {"type":"auto","disable_parallel_tool_use":true}` therefore
took the passthrough branch with the opt-out silently dropped on the
llama-server `/v1/chat/completions` body.

Thread the translated `anthropic_parallel_tool_calls` value through
`_anthropic_passthrough_stream` and `_anthropic_passthrough_non_streaming`
into the shared `_build_passthrough_payload`, which already knows the
field. Test pins both helpers' signatures and that the field reaches
the body via the payload builder.
2026-05-24 17:18:33 +00:00
Daniel Han
0c68f79ebd Merge remote-tracking branch 'origin/main' into pr-5711-head 2026-05-24 16:55:29 +00:00
Daniel Han
d3ae9142a5 Gemini stop cap is 4, matching the OpenAI compat layer (PR #5711)
Gemini exposes its OpenAI-compatible endpoint at
https://generativelanguage.googleapis.com/v1beta/openai. Google's own
docs (https://ai.google.dev/gemini-api/docs/openai) list the supported
parameters and inherit OpenAI's 4-entry stop cap. Without an explicit
`stop_max=4` on the registry the default 16 leaks through and the
upstream silently drops the overflow.

Add the backend registry entry, mirror it in the frontend
`PROVIDER_STOP_MAX` map, and pin the cap with a focused unit test.
2026-05-24 16:39:49 +00:00
Daniel Han
9cd730130e Hide non-supported sampling controls for local safetensors (PR #5711)
The local non-external GGUF path (llama-server) accepts frequency_penalty,
seed, stop and parallel_tool_calls, but the safetensors / HF transformers
worker has no equivalent kwargs and silently drops them. Showing the
controls there has been confusing reviewers: the UI promises a knob that
does nothing.

Gate frequencyPenalty / seed / stop / parallelToolCalls on `isGguf` for
non-external local models so safetensors sessions only show controls the
backend actually honours. External-provider gating is unchanged.

Stale persisted values from a prior GGUF session are still sent on the
wire but the safetensors worker keeps absorbing them via **_unused, so
this is a presentation-only change with no behaviour difference.
2026-05-24 16:34:14 +00:00
pre-commit-ci[bot]
aee1b7b9c1 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 16:26:08 +00:00
Daniel Han
ad36aaa71d Local /v1/messages: invert disable_parallel_tool_use into parallel_tool_calls (PR #5711)
Anthropic Messages API nests `disable_parallel_tool_use` inside the
`tool_choice` object (per docs.claude.com/parallel-tool-use). The local
Anthropic-compat endpoint dropped that flag because the OpenAI shape it
translates into uses a different name and lives at the top level
instead. SDK clients (anthropic-python, anthropic-sdk-go, etc.) that
already speak this dialect therefore could not opt out of parallel
tool calls against the local GGUF model.

Extract `disable_parallel_tool_use` from the incoming tool_choice and
invert it to `parallel_tool_calls` on the agentic-loop call. Plain-chat
and existing tool_choice shapes are untouched. Added a focused unit
test that pins the dict/None/bool/string boundary cases.
2026-05-24 16:23:23 +00:00
Daniel Han
d7a09d975b Drop seed and parallel_tool_calls for Kimi too (PR #5711)
Kimi K2.5/K2.6 chat schema documents temperature, top_p and a small
fixed set of knobs; seed and parallel_tool_calls are not in it. The
frontend already hides those controls (provider-capabilities.ts), so
the only way they reach Kimi is a stale client or a direct API caller.
Add them to body_omit so the registry strips them on the wire instead
of relying on the upstream to 400.

Sync the Kimi web-search bypass test to assert both fields are dropped
alongside frequency_penalty/temperature/top_p.
2026-05-24 16:14:12 +00:00