Commit graph

32 commits

Author SHA1 Message Date
Daniel Han
64328962d0 Expand local-backend coverage further: 10 more knobs from vLLM + llama.cpp live docs (PR #5711)
Round 4 expansion driven by direct fetches of the canonical
SamplingParams + server README pages cited in the user's request.
Adds ten more knobs the docs explicitly support but the panel doesn't
surface yet:

Knob (wire name)              llama.cpp  vLLM   Ollama   Source
----------------------------- ---------- ------ -------- ----------------
skip_special_tokens           no         yes    no       vLLM SamplingParams
spaces_between_special_tokens no         yes    no       vLLM SamplingParams
include_stop_str_in_output    no         yes    no       vLLM SamplingParams
truncate_prompt_tokens        no         yes    no       vLLM SamplingParams
n_keep                        yes        no     no       llama.cpp README
n_probs                       yes        no     no       llama.cpp README
cache_prompt                  yes        no     no       llama.cpp README
return_tokens                 yes        no     no       llama.cpp README
timings_per_token             yes        no     no       llama.cpp README
post_sampling_probs           yes        no     no       llama.cpp README

Backend rationale:
  - vLLM's documented SamplingParams class at
    https://docs.vllm.ai/en/latest/api/vllm/sampling_params/ lists
    skip_special_tokens (default True), spaces_between_special_tokens
    (True), include_stop_str_in_output (False), truncate_prompt_tokens
    (None). All four are vLLM-only; llama-server's README does not
    document them and Ollama's openai/openai.go translator does not
    forward them.
  - llama-server's README at
    https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
    lists n_keep, n_probs, cache_prompt, return_tokens, timings_per_token
    and post_sampling_probs as documented per-request fields. vLLM's
    SamplingParams has no analog, and Ollama's OAI translator drops them.

Capability matrix:
  LLAMA_CPP_CAPABILITIES: 6 llama-only true + 4 vLLM-only false.
  VLLM_CAPABILITIES:       4 vLLM-only true + 6 llama-only false.
  OLLAMA_CAPABILITIES:     all 10 off (OAI translator drops all of them).
  Every other bucket:      all 10 off.

Skip-when-default rules (mirror upstream defaults):
  skip_special_tokens / spaces_between_special_tokens / cache_prompt:
    default true upstream — forward only when explicitly false.
  include_stop_str_in_output / return_tokens / timings_per_token /
    post_sampling_probs: default false — forward only when true.
  truncate_prompt_tokens / n_probs: 0 / null = unset — forward when > 0.
  n_keep: accepts -1 for "keep all", so the gate is value != 0.

Frontend:
  - ProviderCapabilities interface +10 flags.
  - InferenceParams +10 nullable fields (3 numeric + 7 boolean), all
    null in DEFAULT_INFERENCE_PARAMS.
  - OpenAIChatCompletionsRequest wire shape +10 optional fields.
  - chat-adapter forwards each in both the external (capability-aware)
    and local (capability-bypass) branches.
  - chat-settings-storage adds the 3 numeric keys to the existing
    nullable-number loop and 7 boolean keys to a new nullable-boolean
    loop (alongside ignoreEos).

Backend:
  - ChatCompletionRequest +10 Optional Fields with pydantic bounds
    (truncate_prompt_tokens ge=1, n_probs ge=0; booleans unbounded;
    n_keep accepts -1 so no lower bound).
  - llama_cpp.py three payload builders (generate_chat_stream + the
    tool-loop payload block + the final-pass stream_payload) each
    accept and forward the 10 new kwargs.
  - routes/inference.py _build_passthrough_payload accepts and forwards
    the 10; both per-request call sites (lines ~2591, ~2790) thread
    them from the request payload into the llama_cpp methods.

Test: test_local_passthrough_forwards_vllm_output_and_llama_cpp_
  instrumentation round-trips all 10 fields with explicit values
  matching each backend's upstream default and confirms each is absent
  from the body when unset.

65/65 sampling_params_routing tests pass; frontend tsc clean.

Total local-backend knob coverage now (this PR):
  Standard:    temperature, top_p, top_k, min_p, repetition_penalty,
               presence_penalty, frequency_penalty, seed, stop,
               parallel_tool_calls (10)
  llama.cpp:   typical_p, top_n_sigma, repeat_last_n, dynatemp_range,
               dynatemp_exponent, mirostat, mirostat_tau, mirostat_eta,
               dry_multiplier, dry_base, dry_allowed_length,
               dry_penalty_last_n, xtc_probability, xtc_threshold,
               min_keep, ignore_eos, min_tokens, n_keep, n_probs,
               cache_prompt, return_tokens, timings_per_token,
               post_sampling_probs (23)
  vLLM-extra:  ignore_eos, min_tokens, skip_special_tokens,
               spaces_between_special_tokens, include_stop_str_in_output,
               truncate_prompt_tokens (6)
  OpenRouter:  top_a (1)

Deferred for future PRs (require array / object field shape):
  - llama.cpp DRY sequence_breakers (string array)
  - llama.cpp samplers ordering (string array)
  - llama.cpp / vLLM logit_bias (dict)
  - llama.cpp grammar (string) + json_schema (object)
  - vLLM guided_json / guided_regex / guided_choice / guided_grammar
  - vLLM allowed_token_ids / bad_words / stop_token_ids (int / str arrays)
  - OpenAI / Ollama logprobs + top_logprobs (bool + int pairing)
  - n / best_of (need SSE multi-choice handling first)
2026-05-27 07:41:25 +00:00
Daniel Han
3674e11f07 Expand local-backend sampler coverage: DRY + XTC + min_keep + ignore_eos + min_tokens (PR #5711)
Round 3 expansion driven by direct fetches of the llama.cpp server README,
vLLM's SamplingParams source, and Ollama's openai.go OAI translator.
Adds nine new sampling/control knobs with per-backend capability gating:

Knob (wire name)       llama.cpp  vLLM   Ollama   Source
---------------------- ---------- ------ -------- -----------------------
dry_multiplier         yes        no     no       llama.cpp README
dry_base               yes        no     no       llama.cpp README
dry_allowed_length     yes        no     no       llama.cpp README
dry_penalty_last_n     yes        no     no       llama.cpp README
xtc_probability        yes        no     no       llama.cpp README
xtc_threshold          yes        no     no       llama.cpp README
min_keep               yes        no     no       llama.cpp README
ignore_eos             yes        yes    no       llama.cpp + vLLM SamplingParams
min_tokens             yes        yes    no       llama.cpp + vLLM SamplingParams

Backend-side rationale:
  - llama.cpp: full chain documented at
    https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
  - vLLM: SamplingParams source confirms ignore_eos + min_tokens; the
    other seven have no field in
    https://github.com/vllm-project/vllm/blob/main/vllm/sampling_params.py
  - Ollama: openai/openai.go FromChatRequest copies only the OpenAI
    subset (temp/top_p/seed/freq/pres/max_tokens/logprobs/topLogprobs/
    response_format/reasoning_effort) on the /v1/chat/completions path
    Studio uses. All nine new knobs are silently dropped, so the
    OLLAMA_CAPABILITIES bucket keeps them off.

Frontend:
  - ProviderCapabilities interface gains 9 boolean flags.
  - InferenceParams gains 9 nullable fields (8 numeric + ignoreEos
    boolean), all defaulting to null in DEFAULT_INFERENCE_PARAMS.
  - OpenAIChatCompletionsRequest wire shape gains 9 optional fields
    with doc comments.
  - LLAMA_CPP_CAPABILITIES: all 9 on. VLLM_CAPABILITIES: 2 on
    (ignoreEos + minTokens) via inheritance, 7 off via explicit
    override. OLLAMA_CAPABILITIES: all 9 off (inherits + overrides
    ignoreEos/minTokens). Every other bucket (openai cloud / chat,
    anthropic, gemini, mistral, kimi, deepseek, openrouter) gets all
    9 off explicitly.
  - chat-adapter.ts gates each knob in both the external (capability-
    aware) and local (unconditional-when-meaningful) branches.
    Skip-when-default rules:
      dry_multiplier > 0 unlocks the 4-field DRY chain
      xtc_probability > 0 unlocks the 2-field XTC chain
      min_keep > 0, min_tokens > 0 forward only when set higher than 0
      ignore_eos forwards only when explicitly true
  - chat-settings-storage.ts persists all 9 keys (8 numeric in the
    existing nullable-number loop, ignoreEos with its own boolean
    handler).

Backend:
  - ChatCompletionRequest gains 9 Optional Field declarations with
    pydantic ge/le bounds (dry_multiplier ge=0; dry_base ge=1; xtc_*
    ge=0 le=1; min_keep / min_tokens / dry_allowed_length ge=0).
  - llama_cpp.py: three payload builders (generate_chat_stream + the
    two payload-construction blocks inside the tool-loop stream) each
    accept the 9 new kwargs and forward via `if x is not None`.
  - routes/inference.py: _build_passthrough_payload accepts the 9 new
    kwargs and forwards into the body. Two call sites that thread
    sampler params from the request payload (lines 2581, 2771) are
    extended to forward the 9 new fields.

Test:
  - test_local_passthrough_forwards_dry_xtc_min_keep_eos_min_tokens
    round-trips all 9 fields through _build_passthrough_payload and
    confirms each is absent when unset (so llama-server / vLLM apply
    their own defaults).

64/64 sampling_params_routing tests pass; frontend tsc clean.

Deferred for future PRs (require array / object field shape):
  - llama.cpp DRY sequence_breakers (string array)
  - llama.cpp samplers ordering (string array)
  - llama.cpp / vLLM logit_bias (dict)
  - llama.cpp n_probs + OpenAI logprobs/top_logprobs
  - llama.cpp grammar (string) + json_schema (object)
  - vLLM guided_json / guided_regex / guided_choice / guided_grammar
  - vLLM allowed_token_ids / bad_words / stop_token_ids
2026-05-27 07:04:29 +00:00
Daniel Han
0234bef047 Apply 5-reviewer audit fixes to per-provider capability buckets (PR #5711)
Five independent reviewers cross-checked every provider's per-model
sampling-knob exposure against live docs (OpenAI, Anthropic, Gemini,
DeepSeek, Kimi, Mistral, OpenRouter, llama.cpp, vLLM, Ollama).
Applying the high-confidence drift fixes here; speculative items (pro
model effort restrictions, gpt-5.3 cap, OpenAI verbosity / o-series
output cap, Gemini topK / service_tier) are deferred to a follow-up
because they need backend wire changes or unverified doc claims.

Anthropic:
  - Move claude-opus-4-6 from the 64k group into the 128k group (live
    legacy table shows Opus 4.6 Max output = 128k tokens).
    https://platform.claude.com/docs/en/about-claude/models/overview
  - Add claude-sonnet-4 to the 64k group (was falling through to 32k
    default; live legacy table shows Sonnet 4 Max output = 64k tokens).
  - Extend ANTHROPIC_REASONING_MODELS with legacy claude-opus-4-1 /
    claude-opus-4 / claude-sonnet-4 at none/low/medium/high (live
    legacy table marks Extended thinking = Yes for all three).

OpenAI:
  - Split the gpt-5/gpt-5.1/gpt-5.2 reasoning bucket. Per Azure docs
    footnote ^7^, "minimal is only supported with the original GPT-5
    reasoning models. minimal is not supported with gpt-5.1 or greater".
    gpt-5.1 / gpt-5.2 now get none/low/medium/high/xhigh with
    supportsOff=true; bare gpt-5 keeps minimal/low/medium/high
    supportsOff=false. Ordering puts gpt-5.1 / gpt-5.2 before gpt-5 in
    the find() loop so the longer prefix matches first.
    https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/reasoning

DeepSeek:
  - Hide `seed` and `parallel_tool_calls` in the deepseek capability
    bucket. Neither field is in the current /chat/completions schema
    (body fields: messages, model, thinking, max_tokens, response_format,
    stop, stream, stream_options, temperature, top_p, tools, tool_choice,
    logprobs, top_logprobs, user_id). Surfacing them in the UI would be
    the silent-drop UX the file header warns against.
    https://api-docs.deepseek.com/api/create-chat-completion

Mistral:
  - magistral-medium-latest / magistral-small-latest are NATIVE
    always-on reasoning models; injecting reasoning_effort returns 422
    upstream. Switch both to withEnableThinkingStyle({reasoningAlwaysOn:
    true}) instead of the old none/medium/high effort ladder.
  - mistral-small-latest / mistral-medium-latest / mistral-vibe-cli-latest
    expose the documented three-tier adjustable ladder
    (none/low/medium/high), not the truncated none/high pair that was
    here before. mistral-medium-latest was not handled at all and now
    sits in the same bucket as small.
    https://docs.mistral.ai/studio-api/conversations/reasoning
    https://mistral.ai/news/magistral

OpenRouter:
  - Drop google/gemini-pro-latest from OPENROUTER_MANDATORY_REASONING_
    MODELS; the gateway 404s the id today
    (https://openrouter.ai/google/gemini-pro-latest). Removing rather
    than re-pinning to a versioned id that may rotate again.

Local backends:
  - Split LOCAL_LLAMA_CAPABILITIES into LLAMA_CPP_CAPABILITIES (full
    chain — for llama_cpp + custom) and VLLM_OLLAMA_CAPABILITIES (OpenAI
    subset + top_k/min_p/repetition_penalty/seed, no extended samplers).
    vLLM's SamplingParams has no typical_p / top_n_sigma / repeat_last_n
    / dynatemp_* / mirostat* fields, and Ollama's OpenAI translator
    (ollama/openai/openai.go FromChatRequest) only copies the OpenAI
    subset. Surfacing the eight extra sliders for vllm / ollama was
    silent-drop UX.

Tests:
  - test_deepseek_payload_omits_seed_and_parallel_tool_calls: read the
    TS file as text and assert the bucket has seed:false and
    parallelToolCalls:false. Backend has no JS engine; this is the
    cheapest way to lock the wire-drop invariant.
  - 63/63 sampling_params_routing tests pass; frontend tsc clean.
2026-05-27 06:31:40 +00:00
Daniel Han
22111744a4 Narrow Anthropic 4.7 sampling-removed gate to Opus only (PR #5711)
The 4.7 generation only shipped Claude Opus 4.7; Sonnet stops at 4.6
and Haiku at 4.5 per
https://platform.claude.com/docs/en/about-claude/models/overview.
The earlier `^claude-(?:opus|sonnet|haiku)-4-7` regex on both the
backend strip (_ANTHROPIC_4_7_SAMPLING_REMOVED in external_provider.py)
and the frontend mirror (ANTHROPIC_4_7_SAMPLING_REMOVED_REGEX in
provider-capabilities.ts) would have pre-emptively hidden temperature
/ top_p / top_k for any future claude-sonnet-4-7 or claude-haiku-4-7
id, even though Anthropic has explicitly not extended the sampling
removal beyond Opus. Tighten both regexes to `^claude-opus-4-7(?:[-.]|$)`
and update the routing-test pin so claude-sonnet-4-7 and claude-haiku-4-7
are in `should_not_match`. If those ids ever ship and adopt the same
removal, widening the regex is one-line.
2026-05-27 05:43:38 +00:00
Daniel Han
facdff9ad7 Add extended llama.cpp samplers + OpenRouter top_a (PR #5711)
Cross-checked every supported sampling field against each provider's
live docs + LiteLLM's drop_params surface + the llama.cpp server
README. Pulled in the most-asked-for samplers that the PR was missing.

New ProviderCapabilities flags (default false on every SaaS provider
since none accept these):
  - typicalP            (already shipped one commit prior)
  - topNSigma           llama.cpp `top_n_sigma`
  - repeatLastN         llama.cpp `repeat_last_n` (paired w/ repeat_penalty)
  - dynatempRange       llama.cpp `dynatemp_range`
  - dynatempExponent    llama.cpp `dynatemp_exponent`
  - mirostat            llama.cpp `mirostat` mode (0/1/2)
  - mirostatTau         llama.cpp `mirostat_tau`
  - mirostatEta         llama.cpp `mirostat_eta`
  - topA                OpenRouter `top_a` (alternate dynamic-top-P)

Capability bucketing split: ALL_SUPPORTED retired in favor of
  - LOCAL_LLAMA_CAPABILITIES  -> custom / vllm / ollama / llama_cpp
    (full llama.cpp sampler chain, top_a off — not a llama.cpp field)
  - OPENROUTER_CAPABILITIES   -> openrouter
    (gateway's documented set incl. top_a, llama.cpp-only knobs off
     because OpenRouter docs don't list them and they'd be silently
     dropped on most underlying routes)

InferenceParams gains 8 nullable-number fields (mirroring `seed`'s
"null = unset, finite-number = forwarded" shape). DEFAULT_INFERENCE_PARAMS
defaults each to null. Persistence handler in chat-settings-storage
mirrors typicalP's nullable-float handling for all 8.

Backend:
  - 8 new ChatCompletionRequest fields with appropriate `ge`/`le`
    validators (mirostat 0..2, ranges 0.0..1.0 where applicable).
  - llama_cpp.py: signatures + payload forwarding extended on all
    three builders (chat-completion, agentic tool-loop, final-pass)
    so the new fields survive the local tool-loop too. `is not None`
    gating so defaults (e.g. mirostat=0) reach the wire only when the
    caller explicitly opted in.
  - routes/inference.py: _build_passthrough_payload extends to the
    extended sampler chain; 3 call sites (generate_chat_completion,
    generate_chat_completion_with_tools, _build_passthrough_payload)
    forward each field from `payload.*`.

Frontend chat-adapter: external branch forwards only when capability
allows (so OpenRouter gets top_a but not mirostat, local gets mirostat
but not top_a); local branch forwards unconditionally when the value
is meaningful (e.g. mirostat != 0, dynatemp_range > 0).

Test pinning the new field round-trip through _build_passthrough_payload
added; full PR-touched suite now 163 passing (was 161).

References:
  - llama.cpp server params: https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
  - OpenRouter params:       https://openrouter.ai/docs/api/reference/parameters
  - LiteLLM provider params: https://docs.litellm.ai/docs/completion/input
2026-05-26 14:34:32 +00:00
Daniel Han
a02320aa7b Add typical_p sampler (local) + DeepSeek-reasoner per-model gating (PR #5711)
Two follow-ups from a closer reading of each provider's published
sampling surface against llama.cpp's own server README.

typical_p (locally typical sampling, `typ_p` in the llama.cpp sampler
chain):
  - New ProviderCapabilities.typicalP flag; defaults false on every
    SaaS provider (none accept the field) and true only on the
    permissive local buckets (custom, vllm, ollama, llama_cpp,
    openrouter via ALL_SUPPORTED). InferenceParams.typicalP is
    nullable number (null = unset; 1.0 = llama-server default, also
    treated as no-op when forwarding).
  - Backend: new ChatCompletionRequest.typical_p Field (0.0..1.0).
    Threaded through all three llama_cpp.py payload builders
    (chat-completion, agentic tool-loop, final-pass) so the field
    survives the local tool-loop too. _build_passthrough_payload in
    routes/inference.py picks it up and only writes the body when
    the caller set a value; left absent it falls back to llama-server
    default. Three route call sites (generate_chat_completion,
    generate_chat_completion_with_tools, _build_passthrough_payload)
    forward payload.typical_p.
  - Frontend: chat-adapter forwards on both branches (external opt-in
    only when capability allows + value != 1; local forwards
    unconditionally when set and != 1). OpenAIChatCompletionsRequest
    grows a `typical_p?` field. Persisted via chat-settings-storage
    mirroring the seed nullable-float handler.
  - Test: pin _build_passthrough_payload's forward + absent behavior.

DeepSeek per-model gating:
  - deepseek-reasoner / deepseek-r1 silently ignore temperature, top_p,
    presence_penalty, frequency_penalty per
    https://api-docs.deepseek.com/guides/reasoning_model — mirror the
    OpenAI / Claude 4.7 per-model approach: getProviderCapabilities
    downshifts these ids to a stripped capability set so the panel
    does not offer knobs the upstream silently drops.

161+1 sampling-routing tests pass; frontend tsc clean.

Refs:
  - llama.cpp server params: https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
  - DeepSeek reasoner restrictions: https://api-docs.deepseek.com/guides/reasoning_model
2026-05-26 14:12:40 +00:00
Daniel Han
4f9125a5f9 Pin Claude 4.7 sampling-removed regex with backend test (PR #5711)
Guards against drift between the backend _ANTHROPIC_4_7_SAMPLING_REMOVED
regex and the frontend ANTHROPIC_4_7_SAMPLING_REMOVED_REGEX added in the
previous commit. If a future patch widens one without the other the panel
will either silently strip a knob the user moved or 400 on a knob the UI
should have hidden — both bad UX.

Test pins the canonical 4.7 id shapes (opus/sonnet/haiku) including dated
snapshots, and the non-4.7 ids that must NOT match (4-6, 4-5, 4-71,
future 5, gpt-4o, etc.). Sampling-params suite now 60 passing
(previously 59).
2026-05-26 05:48:17 +00:00
Daniel Han
5d4ddd6b37 Revert "Surface OpenAI Responses service_tier='scale' (PR #5711)"
Round 19 added scale to /v1/responses based on the openai-python
SDK type, but round 20 reviewers (3/10 against) and the round 18
aggregator both noted that the live OpenAI Responses reference
limits Responses service tiers to auto/default/flex/priority. The
PR contract in the original description also lists scale only for
Chat Completions, not Responses. Studio routes OpenAI through
Responses, so forwarding scale risks a 400 from the upstream and
exposes a picker option the API does not accept.

Restore the conservative drop behaviour: only documented Responses
tiers reach the wire; legacy scale settings still validate (the
ServiceTier Literal and chat-settings storage allowlist keep it
for forward-compat).
2026-05-24 20:25:55 +00:00
pre-commit-ci[bot]
f7c11d8a0a [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 20:11:43 +00:00
Daniel Han
b961c6f2f7 Surface OpenAI Responses service_tier="scale" (PR #5711)
Round 19 reviewer consensus (3/10 plus an asymmetric-fix call-out
across rounds 8/9/12/14/17): the openai-python SDK ships
service_tier as Literal["auto","default","flex","scale","priority"]
on /v1/responses, and enterprise Scale Tier customers need to opt in
explicitly. Drop the defensive scale-filter on the backend and add
"scale" to the OpenAI picker option list so the field reaches the
wire when set. Other providers remain at auto/default per their docs.

Update the routing tests so `scale` lives in the forwarded-set fixture
and the dropped-set fixture only carries truly out-of-enum values
(Anthropic-only `standard_only`, typos, empty string).
2026-05-24 20:09:17 +00:00
Daniel Han
07831d4c0c Shut down asyncgens in routing tests to silence cleanup warnings (PR #5711)
Round 18 reviewers (and earlier) noted CI noise from the routing test
helper: `_drive(coro)` ran a fresh event loop but never closed it or
called `shutdown_asyncgens`, so the MockTransport-backed httpx async
generators in the providers were finalised by GC in a later task and
emitted "Response.aiter_text.aclose was never awaited" / "Task was
destroyed but it is pending" warnings.

Explicitly close the loop after `run_until_complete`, running
`shutdown_asyncgens` first so the iterators finalise in this task.
Tests still pass and the warnings are gone.
2026-05-24 19:54:08 +00:00
Daniel Han
5218a01d14 Apply parallel_tool_calls cap to Anthropic passthrough + safetensors path (PR #5711)
Round 16 reviewer consensus extended the round 12c asymmetric-fix:
the GGUF agentic loop capped tool_calls to one when the caller opted
out, but every other Studio-internal path that emits tool calls from
llama-server output skipped the same guard.

Mirror the cap in three places that have full ownership of the
emitted list (passthrough verbatim paths are out of scope):

1. `AnthropicPassthroughEmitter` now takes `parallel_tool_calls` and
   silently drops streamed `delta.tool_calls` entries beyond the
   first index. Wired from `_anthropic_passthrough_stream`.
2. `_anthropic_passthrough_non_streaming` truncates the upstream
   `message.tool_calls` list before producing `tool_use` blocks.
3. `run_safetensors_tool_loop` truncates the parsed tool_calls list
   before appending the assistant message and executing tools.
   `InferenceOrchestrator.generate_chat_completion_with_tools` and
   the safetensors route now thread `parallel_tool_calls` through.

Also harden `_build_passthrough_payload` to strip empty / non-string
`stop` entries before forwarding to llama-server, matching the
`_normalize_stop_for_provider` shape the external-provider helper
already enforces.

Test pins the AnthropicPassthroughEmitter serial-tool-call gate.
2026-05-24 18:46:07 +00:00
Daniel Han
67e371934c Enforce parallel_tool_calls=False client-side on local GGUF (PR #5711)
Two related findings from round 12 reviewers:

1. The local GGUF tool loop in `generate_chat_completion_with_tools`
   iterates every entry of `tool_calls` returned by llama-server, even
   when the caller explicitly opted out of parallel tool calls. The
   `parallel_tool_calls` flag is forwarded to llama-server, but llama
   .cpp does not enforce it on every jinja template
   (https://github.com/ggml-org/llama.cpp/issues/22043), so a model
   that ignores the flag still ran multiple tools per turn. Cap
   `tool_calls` to the first entry when the flag is False so the
   client-side contract holds regardless of upstream behavior.

2. llama-server documents `parallel_tool_calls` as defaulting to FALSE
   (https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md),
   so the previous chat-adapter shape (forward only on explicit false)
   meant the UI's default-on state could never enable parallel tool
   calls there. Always forward the user's preference on the local
   path so the toggle actually does what it says. External providers
   default to true everywhere, so the external branch is unchanged.

Test pins the GGUF tool-loop cap by source-level assertion (the loop
itself is integration-only).
2026-05-24 17:35:05 +00:00
Daniel Han
48df6a98c8 Forward disable_parallel_tool_use through Anthropic client-tool passthrough (PR #5711)
Round 11 reviewer consensus (10/10): the `disable_parallel_tool_use`
translation added in round 11b reached the Anthropic-compat server-tool
GGUF loop but not the analogous client-tool passthrough branch. A
client sending `/v1/messages` with custom tools plus
`tool_choice: {"type":"auto","disable_parallel_tool_use":true}` therefore
took the passthrough branch with the opt-out silently dropped on the
llama-server `/v1/chat/completions` body.

Thread the translated `anthropic_parallel_tool_calls` value through
`_anthropic_passthrough_stream` and `_anthropic_passthrough_non_streaming`
into the shared `_build_passthrough_payload`, which already knows the
field. Test pins both helpers' signatures and that the field reaches
the body via the payload builder.
2026-05-24 17:18:33 +00:00
Daniel Han
d3ae9142a5 Gemini stop cap is 4, matching the OpenAI compat layer (PR #5711)
Gemini exposes its OpenAI-compatible endpoint at
https://generativelanguage.googleapis.com/v1beta/openai. Google's own
docs (https://ai.google.dev/gemini-api/docs/openai) list the supported
parameters and inherit OpenAI's 4-entry stop cap. Without an explicit
`stop_max=4` on the registry the default 16 leaks through and the
upstream silently drops the overflow.

Add the backend registry entry, mirror it in the frontend
`PROVIDER_STOP_MAX` map, and pin the cap with a focused unit test.
2026-05-24 16:39:49 +00:00
pre-commit-ci[bot]
aee1b7b9c1 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 16:26:08 +00:00
Daniel Han
ad36aaa71d Local /v1/messages: invert disable_parallel_tool_use into parallel_tool_calls (PR #5711)
Anthropic Messages API nests `disable_parallel_tool_use` inside the
`tool_choice` object (per docs.claude.com/parallel-tool-use). The local
Anthropic-compat endpoint dropped that flag because the OpenAI shape it
translates into uses a different name and lives at the top level
instead. SDK clients (anthropic-python, anthropic-sdk-go, etc.) that
already speak this dialect therefore could not opt out of parallel
tool calls against the local GGUF model.

Extract `disable_parallel_tool_use` from the incoming tool_choice and
invert it to `parallel_tool_calls` on the agentic-loop call. Plain-chat
and existing tool_choice shapes are untouched. Added a focused unit
test that pins the dict/None/bool/string boundary cases.
2026-05-24 16:23:23 +00:00
Daniel Han
d7a09d975b Drop seed and parallel_tool_calls for Kimi too (PR #5711)
Kimi K2.5/K2.6 chat schema documents temperature, top_p and a small
fixed set of knobs; seed and parallel_tool_calls are not in it. The
frontend already hides those controls (provider-capabilities.ts), so
the only way they reach Kimi is a stale client or a direct API caller.
Add them to body_omit so the registry strips them on the wire instead
of relying on the upstream to 400.

Sync the Kimi web-search bypass test to assert both fields are dropped
alongside frequency_penalty/temperature/top_p.
2026-05-24 16:14:12 +00:00
Daniel Han
e0a9b1d76a Drop Kimi frequency_penalty and gate generic service_tier on opt-in
5/10 reviewers in the last round flagged Kimi forwarding non-default frequency_penalty as a 400 risk for K2.5 / K2.6, mirroring the existing lock on temperature and top_p. Hide the slider on the frontend and add frequency_penalty to Kimi's body_omit so even stale clients have the field stripped before the request hits the wire.

service_tier on the generic OpenAI-compatible branch was forwarding whatever value the dispatcher received, so a stale frontend could send standard_only (Anthropic) or scale to providers like Mistral that do not document the field, producing 400s. Gate the forward on an explicit accepts_service_tier=True provider registry opt-in; Anthropic and OpenAI Responses already handle service_tier inside their own helpers.
2026-05-24 15:54:26 +00:00
Daniel Han
1d1a205a19 OpenRouter stop cap is 4, GGUF tool-loop final pass forwards new fields
OpenRouter normalises to OpenAI's chat schema and inherits the 4-entry stop cap. The default 16-cap was too permissive; add stop_max=4 on both the backend provider registry and the frontend PROVIDER_STOP_MAX map.

The GGUF tool-iteration final-answer pass at llama_cpp.py:5182 was carrying only the legacy sampling fields. Forward frequency_penalty, seed, and parallel_tool_calls there too so the cap-exhausted path matches the per-iteration loop.

Test pins the OpenRouter 4-cap.
2026-05-24 15:33:42 +00:00
Daniel Han
1cc52465f3 Kimi 32-byte per-stop cap; extract _normalize_stop_for_provider helper
Kimi documents max 5 stop strings AND <= 32 bytes per string at
https://platform.kimi.ai/docs/api/chat. The previous code capped
count but forwarded oversize entries, which can produce upstream
400s. Add stop_max_bytes=32 on the Kimi registry entry and apply
both checks in a new _normalize_stop_for_provider helper shared
between the default OAI-compat path and the Kimi web-search bypass.

Tests pin the byte-cap drop on both Kimi paths.
2026-05-24 15:06:15 +00:00
Daniel Han
95e143545f Per-provider stop cap on Kimi web-search bypass and frontend sheet
Round 5 review flagged two asymmetries:

1. Kimi web-search bypass hard-capped stops at 4 while the default OAI-compat path honours provider_info["stop_max"]. Apply the same provider-aware logic in _stream_kimi_web_search so kimi-with-search and kimi-without-search match. Also add Kimi's documented 5-stop max (https://platform.kimi.ai/docs/api/chat) to the provider registry so the cap actually fires.

2. chat-settings-sheet.tsx caps every non-Anthropic external provider at 4 stops. Replace with a per-provider getProviderStopMax helper in provider-capabilities.ts so DeepSeek, Mistral, and local backends are not artificially restricted while OpenAI Chat still hits its 4-entry hard limit and Kimi hits its documented 5-entry cap.

Tests pin the Kimi 5-cap on both Kimi paths.
2026-05-24 14:50:04 +00:00
Daniel Han
b48d68f8bf Fix Mistral seed mapping, raise default OAI-compat stop cap, thread sampling through GGUF direct path
Mistral chat completions uses random_seed not seed; map the field via a new seed_field on the provider registry so the new seed control actually works on Mistral. Default for other providers stays seed.

DeepSeek and Mistral both accept up to 16 stop sequences but the default OAI-compat branch was hard-capping at 4 (the OpenAI Chat limit). Studio routes the openai provider through /v1/responses not /v1/chat/completions so the 4-cap only applies if we explicitly added an openai entry. Raise the default to 16 and let per-provider stop_max overrides tighten if needed.

The local GGUF direct chat path (gguf_generate / gguf_generate_with_tools) bypassed _build_openai_passthrough_body and therefore dropped frequency_penalty, seed, stop, and parallel_tool_calls on the floor for users on the default no-tools and with-tools paths. Thread the new fields through LlamaCppBackend.generate_chat_completion and generate_chat_completion_with_tools and the two callsites that invoke them.

Also tighten comments to drop review-process narration that crept in and to remove the em dashes I had introduced in this PR's earlier commits.

Tests pin the Mistral random_seed rename, the DeepSeek 16-cap, and confirm the openai-compat default cap is 16.
2026-05-24 14:32:36 +00:00
Daniel Han
30d6ce201e Studio: drop service_tier=scale on OpenAI Responses path
Round 4 reviewer consensus (~9/20 independent reviewers) flagged
service_tier=scale as a 400 risk on /v1/responses. The earlier commit
added scale based on the openai-python SDK literal, but the live
OpenAI Responses API reference, the PR's own provider matrix, and the
9-reviewer round-4 consensus all agree the documented Responses enum
is auto|default|flex|priority only. Drop scale on this path to remove
the risk.

Keeps scale on the Chat Completions / OAI-compat path where the SDK
enum is honored and where users who want Scale Tier can still select
it. The widened TypeScript ServiceTier / ServiceTierOption / api.ts
union and the storage sanitizer allowlist remain permissive so legacy
persisted "scale" values do not get silently dropped on reload; the
runtime per-provider gate makes the routing decision.

Tests are updated to pin the restricted Responses enum and the
explicit drop of scale + standard_only + bogus values.
2026-05-24 14:12:13 +00:00
Daniel Han
b8cef29b50 Studio: forward parallel_tool_calls through /v1/responses bridge
Round 3 reviewer feedback:

- studio/backend/routes/inference.py: _build_chat_request (the
  /v1/responses → /v1/chat/completions translator) was dropping
  parallel_tool_calls on the floor. A Responses-API caller that set
  `parallel_tool_calls=false` saw the flag accepted at the schema
  layer but never reach llama-server because the translated
  ChatCompletionRequest had no first-class field for it. Now that
  parallel_tool_calls IS a first-class field on ChatCompletionRequest
  (added by this PR's earlier commits), translate it through the
  bridge so the preference actually fires.

- studio/frontend/src/features/chat/utils/chat-settings-storage.ts:
  the stop sanitizer silently dropped `stop: []` instead of persisting
  the empty array. That meant a user could not clear the last chip —
  on reload, the previously-persisted stops came back. Persist empty
  arrays explicitly so the cleared state round-trips.

- studio/backend/tests/test_sampling_params_routing.py: pin both with
  the raw reproductions reviewers cited.
2026-05-24 13:57:46 +00:00
pre-commit-ci[bot]
fdf0be484e [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 12:45:36 +00:00
Daniel Han
d6765fddce Studio: thread sampling extensions through local + Kimi-search paths
Round-2 round of review-feedback fixes for the sampling-knobs PR:

- studio/backend/routes/chat_history.py: ChatInferenceSettings still had
  the pre-PR field list with extra="forbid", so every settings save the
  new frontend issued would 422 on the new keys (frequencyPenalty,
  seed, stop, serviceTier, parallelToolCalls). Add the fields with the
  same range / enum constraints the chat-completions schema uses, so
  the settings-persistence path round-trips cleanly.

- studio/backend/routes/inference.py: _build_passthrough_payload and
  _build_openai_passthrough_body now thread frequency_penalty, seed,
  and parallel_tool_calls through to llama-server. The frontend exposes
  these knobs for local backends; without the forwarding the UI was a
  decoration. Each field is gated on `is not None` so 0 / False / "0"
  still reach the body.

- studio/backend/core/inference/external_provider.py: the Kimi
  $web_search bypass takes an early return into _stream_kimi_web_search
  before the default OAI-compat body builder runs, so the new sampling
  fields never landed on Kimi-with-search. Forward them through the
  helper, with the same dedupe / truncate behavior the main path
  applies to `stop`. Also extend the OpenAI Responses service_tier
  allowlist to include `scale` per the live openai-python SDK
  (response_create_params.py declares
  Literal["auto","default","flex","scale","priority"]).

- studio/frontend/src/features/chat/provider-capabilities.ts +
  types/runtime.ts: add `scale` to ServiceTier / ServiceTierOption and
  surface it on the OpenAI Responses options so the UI matches the
  upstream enum.

- studio/backend/tests/test_sampling_params_routing.py: add tests for
  every gap above: Kimi web-search bypass forwarding, local OpenAI
  passthrough forwarding, ChatSettingsPayload round-trip, and the full
  Responses service_tier enum (parametrized over the five accepted
  values plus a drop check for the Anthropic-only standard_only).
2026-05-24 12:45:11 +00:00
Daniel Han
d6b4c36e0a Studio: nest disable_parallel_tool_use, drop ws-only stops, fix persistence
Anthropic Messages API rejects `disable_parallel_tool_use` as a
top-level field; it is only accepted as a property on the `tool_choice`
object. Move the inversion into a tool_choice merge that defaults to
`{type:"auto"}` when no choice is supplied, and skip the field entirely
when no tools are defined (it is a no-op without tools).

The same path also dropped `stop` chips that contain only whitespace,
because Anthropic 400s with `each stop sequence must contain
non-whitespace` on entries like " ", "\n", and "\n\n". The previous
filter only dropped truly empty strings; switch to `s.strip()` so the
common newline-stop defaults are also filtered out client-side.

Frontend persistence had three round-trip data-loss bugs:

  - `VALID_SERVICE_TIERS` was missing `standard_only`, so any Anthropic
    user who picked that tier lost it on the next reload.
  - The settings sanitizer truncated `stop` to 4 entries on save,
    which defeated the Anthropic UI cap of 16. Use 16 here and let the
    per-provider stream helper cap to the wire's allowed length.
  - The chat-settings sheet's `stopMaxEntries` capped local backends
    (llama.cpp / vLLM / ollama / generic OpenAI-compat) at 4 even
    though those backends happily accept more. Match Anthropic's 16
    for the local path.

Preset policy now carries `frequencyPenalty` and `stop` so a saved
preset can fix a user's preferred decoding style. `seed`,
`serviceTier`, and `parallelToolCalls` stay out of presets because
they are per-request determinism / per-provider account / per-tool
state, not reusable preset values.

Drops the test that pinned the buggy top-level placement of
`disable_parallel_tool_use` and adds two tests for the nested shape
plus the without-tools skip path, plus a test pinning the
whitespace-stop filter against the documented Anthropic error.
2026-05-24 12:37:28 +00:00
pre-commit-ci[bot]
093f465620 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-23 15:33:14 +00:00
Daniel Han
807165810f Address review feedback on sampling-params knobs
- Drop `scale` from the OpenAI service-tier picker (frontend types and
  picker option list). OpenAI in Studio routes through `/v1/responses`,
  which does not accept `scale`; offering it in the UI silently
  dropped the value at the backend and misled users into thinking
  their selection was applied. Backend Literal still accepts it on
  input so stale clients are not 422'd, and `_stream_openai_responses`
  continues to drop it from the wire body.
- Dedupe + drop empty entries for OpenAI Chat `stop` and Anthropic
  `stop_sequences` before forwarding so whitespace chips or accidental
  repeats do not waste the 4-entry OpenAI cap or the 16-entry
  Anthropic cap. Anthropic over-cap now logs and truncates, matching
  the OpenAI path.
- Static `aria-label="Parallel tool calls"` on the Switch; screen
  readers already announce checked / unchecked state, so the dynamic
  Enable/Disable label was redundant.
- Forward an `aria-label` onto the inner Input inside
  `StopSequencesInput` so screen-reader users can identify the field.
- Regression tests covering the new dedup, truncation, and the
  preserved silent-drop of `scale` on Responses.
2026-05-23 15:33:14 +00:00
pre-commit-ci[bot]
3cd3a64088 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-23 15:33:13 +00:00
Daniel Han
91d04741ff Studio: expose Anthropic / OpenAI sampling knobs per provider
Adds the missing sampling parameters that the upstream APIs accept and
that Studio's chat UI previously hid. Each knob is gated per provider
so the picker never offers a field the upstream would 400 on, and the
per-provider stream functions translate / drop fields to match each
API's naming.

New `InferenceParams` fields (round-trip through PersistedInferenceParams
and the chat-settings server store automatically):

- frequencyPenalty (-2..2): OpenAI Chat Completions only.
- seed (int | null): OpenAI Chat + OpenAI-compat local backends.
- stop (string[]): all OpenAI Chat + Anthropic Messages. Backend
  truncates to 4 entries on OpenAI Chat per docs and renames to
  `stop_sequences` on Anthropic.
- serviceTier (auto|default|flex|priority|scale|standard_only):
  per-provider enum sets resolved by getServiceTierOptions.
- parallelToolCalls (bool, default true): forwarded as
  `parallel_tool_calls` on both OpenAI APIs and inverted into
  `disable_parallel_tool_use` on Anthropic.

OpenAI Responses (gpt-5.x / o3) explicitly drops frequencyPenalty /
seed / stop alongside the existing temperature / top_p drop, since
the upstream 400s on all of them. service_tier on Responses accepts a
subset (no `scale`) which the dispatch already enforces.

UI rows land in the existing Sampling section of the chat settings
sheet using ParamSlider (frequency penalty), a numeric Input (seed),
a new chips editor `StopSequencesInput` (stop), Select (service tier),
and Switch (parallel tool calls). Each row's visibility follows the
new ProviderCapabilities flag.

Tests pin the gating contract: stop_sequences renamed on Anthropic,
4-entry truncation on OpenAI Chat, every Responses-rejected field
dropped, schema-level validation for the service_tier Literal and
frequency_penalty range.

Plan: plans/hashed-riding-porcupine.md
2026-05-23 15:33:13 +00:00