Cross-checked every supported sampling field against each provider's
live docs + LiteLLM's drop_params surface + the llama.cpp server
README. Pulled in the most-asked-for samplers that the PR was missing.
New ProviderCapabilities flags (default false on every SaaS provider
since none accept these):
- typicalP (already shipped one commit prior)
- topNSigma llama.cpp `top_n_sigma`
- repeatLastN llama.cpp `repeat_last_n` (paired w/ repeat_penalty)
- dynatempRange llama.cpp `dynatemp_range`
- dynatempExponent llama.cpp `dynatemp_exponent`
- mirostat llama.cpp `mirostat` mode (0/1/2)
- mirostatTau llama.cpp `mirostat_tau`
- mirostatEta llama.cpp `mirostat_eta`
- topA OpenRouter `top_a` (alternate dynamic-top-P)
Capability bucketing split: ALL_SUPPORTED retired in favor of
- LOCAL_LLAMA_CAPABILITIES -> custom / vllm / ollama / llama_cpp
(full llama.cpp sampler chain, top_a off — not a llama.cpp field)
- OPENROUTER_CAPABILITIES -> openrouter
(gateway's documented set incl. top_a, llama.cpp-only knobs off
because OpenRouter docs don't list them and they'd be silently
dropped on most underlying routes)
InferenceParams gains 8 nullable-number fields (mirroring `seed`'s
"null = unset, finite-number = forwarded" shape). DEFAULT_INFERENCE_PARAMS
defaults each to null. Persistence handler in chat-settings-storage
mirrors typicalP's nullable-float handling for all 8.
Backend:
- 8 new ChatCompletionRequest fields with appropriate `ge`/`le`
validators (mirostat 0..2, ranges 0.0..1.0 where applicable).
- llama_cpp.py: signatures + payload forwarding extended on all
three builders (chat-completion, agentic tool-loop, final-pass)
so the new fields survive the local tool-loop too. `is not None`
gating so defaults (e.g. mirostat=0) reach the wire only when the
caller explicitly opted in.
- routes/inference.py: _build_passthrough_payload extends to the
extended sampler chain; 3 call sites (generate_chat_completion,
generate_chat_completion_with_tools, _build_passthrough_payload)
forward each field from `payload.*`.
Frontend chat-adapter: external branch forwards only when capability
allows (so OpenRouter gets top_a but not mirostat, local gets mirostat
but not top_a); local branch forwards unconditionally when the value
is meaningful (e.g. mirostat != 0, dynatemp_range > 0).
Test pinning the new field round-trip through _build_passthrough_payload
added; full PR-touched suite now 163 passing (was 161).
References:
- llama.cpp server params: https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
- OpenRouter params: https://openrouter.ai/docs/api/reference/parameters
- LiteLLM provider params: https://docs.litellm.ai/docs/completion/input
Two follow-ups from a closer reading of each provider's published
sampling surface against llama.cpp's own server README.
typical_p (locally typical sampling, `typ_p` in the llama.cpp sampler
chain):
- New ProviderCapabilities.typicalP flag; defaults false on every
SaaS provider (none accept the field) and true only on the
permissive local buckets (custom, vllm, ollama, llama_cpp,
openrouter via ALL_SUPPORTED). InferenceParams.typicalP is
nullable number (null = unset; 1.0 = llama-server default, also
treated as no-op when forwarding).
- Backend: new ChatCompletionRequest.typical_p Field (0.0..1.0).
Threaded through all three llama_cpp.py payload builders
(chat-completion, agentic tool-loop, final-pass) so the field
survives the local tool-loop too. _build_passthrough_payload in
routes/inference.py picks it up and only writes the body when
the caller set a value; left absent it falls back to llama-server
default. Three route call sites (generate_chat_completion,
generate_chat_completion_with_tools, _build_passthrough_payload)
forward payload.typical_p.
- Frontend: chat-adapter forwards on both branches (external opt-in
only when capability allows + value != 1; local forwards
unconditionally when set and != 1). OpenAIChatCompletionsRequest
grows a `typical_p?` field. Persisted via chat-settings-storage
mirroring the seed nullable-float handler.
- Test: pin _build_passthrough_payload's forward + absent behavior.
DeepSeek per-model gating:
- deepseek-reasoner / deepseek-r1 silently ignore temperature, top_p,
presence_penalty, frequency_penalty per
https://api-docs.deepseek.com/guides/reasoning_model — mirror the
OpenAI / Claude 4.7 per-model approach: getProviderCapabilities
downshifts these ids to a stripped capability set so the panel
does not offer knobs the upstream silently drops.
161+1 sampling-routing tests pass; frontend tsc clean.
Refs:
- llama.cpp server params: https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
- DeepSeek reasoner restrictions: https://api-docs.deepseek.com/guides/reasoning_model
Resolve 5 conflicts where main added Anthropic Opus 4.6/4.7 fast-mode
support that touches the same signatures PR 5711 extended:
studio/backend/core/inference/external_provider.py
Keep both PR 5711 sampling fields (frequency_penalty, seed, stop,
service_tier, parallel_tool_calls) and main's fast_mode in the
stream signature, docstring, and _stream_anthropic call site.
studio/backend/models/inference.py
Append fast_mode Field alongside PR 5711's new ChatCompletionRequest
fields; both flow through the existing dispatch.
studio/backend/routes/inference.py
Forward fast_mode and the PR 5711 sampling fields to the stream
generator together.
studio/frontend/src/features/chat/types/api.ts
Add fast_mode? to OpenAIChatCompletionsRequest after the PR 5711
field block.
studio/frontend/src/features/chat/utils/chat-settings-storage.ts
Persist fastMode alongside seed / stop / serviceTier /
parallelToolCalls.
No semantic changes to either feature surface. 161 backend routing
tests still pass; frontend tsc clean.
* Studio: longest-prefix pricing match + accept chat-style usage keys
Two P1 / High follow-ups from PR 5690 review feedback:
1. Pricing prefix lookup returned the first key it iterated, so
dated snapshots like ``gpt-5.4-mini-2026-04-23`` collided with
the shorter ``gpt-5.4`` entry and overbilled by 3x+. Sort the
table keys longest-first so the most specific entry wins.
2. ``calculate_cost`` only read ``input_tokens`` / ``output_tokens``,
but Studio's OpenAI-Chat-style usage envelope re-emits
``prompt_tokens`` / ``completion_tokens`` (the OpenAI Chat
Completions vocabulary). Callers handing in the chat-style
shape silently got a zeroed bill. Accept either pair so the
calculator works against both raw upstream usage and the
Studio-translated envelope.
Tests (4 new in test_pricing.py): dated mini/pro snapshots inherit
the right rate; chat-style usage keys price correctly; raw key wins
when both shapes are present.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Studio: dedupe cache buckets when costing chat-style Anthropic usage
When the caller hands in Studio's chat-style envelope (``prompt_tokens``
emitted by ``_build_usage_chunk``) for Anthropic, that value already
folds ``cache_creation_input_tokens`` + ``cache_read_input_tokens`` into
the total. The previous follow-up accepted the chat-style key but then
re-added both cache buckets in ``billable_input_tokens`` and ``input_usd``,
double-counting cache tokens on every Anthropic chat-style call.
Detect which envelope landed (``input_tokens`` present = raw upstream;
absent + ``prompt_tokens`` present = Studio chat-style) and peel the
cache buckets off for Anthropic before the downstream math so both
envelopes produce identical costs.
OpenAI: ``input_tokens`` and Studio's ``prompt_tokens`` both already
include ``cache_read`` and exclude any notional ``cache_creation``, so
the OpenAI path stays a straight passthrough.
Tests (2 new): both envelopes match for Anthropic on a triple
(uncached + cache_creation + cache_read); OpenAI envelopes match on a
cached-tokens fixture.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Studio: prefer raw output_tokens over chat-style completion_tokens
Codex flagged that the previous fallback chain
'usage.get("output_tokens") or usage.get("completion_tokens")'
treats an explicit 0 as missing -- a mixed-envelope payload where
'output_tokens' is 0 but 'completion_tokens' is non-zero (or
stale) bills the wrong amount. Mirror the has_input_tokens
precedence pattern: when the raw key is present we use it even at
0; otherwise fall back to completion_tokens.
* Studio: read OpenAI cached tokens from prompt_tokens_details too
Codex flagged that the chat-style OpenAI envelope Studio re-emits
via _build_usage_chunk surfaces cached prompt tokens under
prompt_tokens_details.cached_tokens, not input_tokens_details. The
OpenAI branch only checked input_tokens_details, so a cache-heavy
chat-style turn billed every cached token at the full input rate
instead of the 0.1x cache_read discount.
Walk both keys when discovering the cached count. New regression
test pins that the two envelopes price identically for a turn with
80k of 100k tokens cached.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Studio: tighten pricing prefix match + clamp corrupt usage
Three follow-ups on the longest-prefix pricing match landed in this PR:
- Prefix match now requires a dash boundary or end-of-string. The
longest-key sort alone still falsely landed "claude-opus-4-15" on
the "claude-opus-4-1" row, and "gpt-5.5-prod" on the "gpt-5.5-pro"
row (a 6x overcharge). Demanding the next character be "-" rules
out the lookalikes while keeping dated snapshots
("gpt-5.4-mini-2026-04-23", "claude-opus-4-7-20260414") landing on
their canonical row.
- Clamp every token count to >= 0. A corrupted upstream payload
(negative cached count, off-by-one in a fixture) could previously
produce a negative bill that masked real spend in the session
total tooltip.
- Tolerate a non-dict "cache_creation" (e.g. an upstream proxy
folded the field down to a single int). The current code raised
AttributeError mid-turn; now it falls back to the 5m-default
bucket so the rest of the cost calculation still runs.
Adds tests/test_pricing_edge.py with 20 adversarial cases covering
the boundary check, negative / None / zero token values across both
envelopes, cache_read > prompt corruption, the OpenAI long-context
threshold crossover on cache-inflated billable input, malformed
sub-objects, and unknown-provider degradation. Combined suite is
51 tests, all green.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Surface Anthropic cache-read fallback and forward 1h breakdown
Two correctness gaps surfaced on the chat-style usage envelope:
1) Anthropic cache_read fell through to "uncached input" pricing when
the envelope arrived without the native ``cache_read_input_tokens``
key (e.g. via a proxy that only emits the mirrored
``prompt_tokens_details.cached_tokens`` block). Studio's canonical
``_build_usage_chunk`` always sets both so production traffic was
never affected, but the calculator should accept either as a
defense-in-depth measure. Add a fallback to read the mirrored
field when the native one is missing or zero; the native key still
wins when both are present so the math stays deterministic.
2) ``_build_usage_chunk`` dropped the ``cache_creation`` 5m / 1h
breakdown. Downstream ``calculate_cost`` then could not apply the
2x 1h premium and silently fell back to the 5m default,
underbilling 1h cache writes by 2x on chat-style traffic. Forward
the breakdown verbatim when the upstream usage carries it.
Tests grow by 4 (20 -> 24): two for the prompt_tokens_details
fallback (with native-precedence pin), one for the chunk shape, one
for the end-to-end pricing parity check at 1h.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Add Anthropic fast_mode pricing multiplier
PR 5715 wires the fast-mode-2026-02-01 beta header + speed:"fast"
field through to Anthropic, but the cost calculator never learnt
about the matching 6x premium documented at
https://platform.claude.com/docs/en/build-with-claude/fast-mode
(Opus 4.7 standard $5/$25 per MTok, fast $30/$150).
This adds:
- ANTHROPIC_FAST_MODE_MULT = 6.0 constant.
- calculate_cost(..., fast_mode=True) applies the 6x to base input
AND output rates before any cache multipliers (cache mults stack
on top of fast per Anthropic docs).
- Provider+model gate: silently no-op on every model that is not
claude-opus-4-6 / claude-opus-4-7 so a stray fast_mode=True on
Sonnet/Haiku can never over-charge.
- model_priced label tagged "(fast)" so the cost tooltip can
surface which rate fired.
- pricing_snapshot now exposes fast_mode_mult so the frontend cost
panel doesn't have to hard-code 6.
7 new edge tests pin the math; existing 55 still pass.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Honor explicit zero cache_read_input_tokens on Anthropic envelopes
The previous follow-up fell back to ``prompt_tokens_details.cached_tokens``
whenever the native ``cache_read_input_tokens`` was missing OR equal to 0,
even though the commit message stated the native key always wins when
present. A proxy that forwards a stale ``prompt_tokens_details`` block
alongside an authoritative ``cache_read_input_tokens: 0`` would then
inflate cache_read past the real native count, posting a false cache_read
line and bumping billable_input_tokens. Switch the gate to native-key
presence so an explicit zero stays authoritative; the mirror only kicks
in when the native key is absent. Add a regression test pinning the
explicit-zero precedence.
* Move fast_mode pricing back to #5715
The fast_mode 6x multiplier landed in two places at once -- here
(f66df7ba) and on #5715 (4f1afdb5) -- since both audits ran in
parallel. Drop the duplicate from this branch so the change lives
in its natural home (#5715, which introduces fast_mode itself);
this PR stays focused on the cache-read fallback + 1h breakdown.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Shorten pricing comments for PR #5722
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* Studio: surface Anthropic document citations inline + in Sources panel
Anthropic's Messages API streams ``citations_delta`` events on
``content_block_delta`` when the request enables
``citations: {enabled: true}`` on document blocks. Each event carries
one citation pointing at the source document; previously they were
silently dropped, so reader-visible references never reached the chat
UI even when the model was citing properly.
The proxy now:
- dedupes by the type-specific anchor (char_location / page_location /
content_block_location / search_result_location) so re-cites of the
same span collapse onto a single footnote;
- injects ``[N]`` inline right after the matching text run;
- forwards the full list as a synthetic ``document_citations``
tool_event at ``message_stop`` so the Sources panel can render
per-document footnotes next to web_search / web_fetch citations.
Streams that never emit ``citations_delta`` stay byte-identical.
References:
- https://platform.claude.com/docs/en/build-with-claude/citations
- https://platform.claude.com/docs/en/build-with-claude/search-results
Tests (5 in test_anthropic_citations.py): passthrough, single
char_location, dedup of repeat citations, distinct sources get
distinct numbers, search_result_location supported.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Studio: surface Anthropic document_citations in the Sources panel
The PR added a backend _toolEvent.type='document_citations' on
message_stop and an inline [N] marker in the assistant text, but the
chat-adapter only handles container_*/tool_*/sources from
web_search and web_fetch tool calls. Reviewers flagged that the
inline [N] markers had no matching footnote entries in the Sources
panel.
Capture the new event into a documentCitationParts buffer, convert
each citation dict into a Sources-panel source entry (using
document_title or search-result source URL plus cited_text as the
snippet), dedupe by id, and append to the final yield alongside
the existing web_search/web_fetch sourceParts.
* Studio: dedupe search_result_location citations by search_result_index
Anthropic's documented search_result_location citation shape carries
search_result_index, source, title, and start/end_block_index --
NOT document_index/document_title. The previous key keyed on
document_index + document_title + source + start_block_index, so
two distinct search results from the same source collapsed onto the
same footnote and the second [N] marker was lost.
Switch the search_result_location branch to key on the documented
fields, and pin the behaviour with a regression test asserting that
two citations sharing source/title but with different
search_result_index get distinct [1] [2] markers.
* Studio: keep each citation distinct across the end-anchor
Codex follow-ups on the citations PR:
* Backend _anthropic_citation_key now includes the end anchor for
every variant (end_char_index, end_page_number,
end_block_index). Anthropic ranges are start-AND-end pairs, so
a same-start / different-end pair is two distinct citations
that previously collapsed onto one footnote.
* Frontend documentCitationToSource ids include the position
fields (search_result_index, start/end char/page/block) instead
of being keyed on URL alone. Two citations from the same
document or two search_result_locations with the same source
now produce distinct Sources-panel entries, matching the
inline [N] numbering.
* Studio: key Sources list by per-citation id instead of url
Codex flagged that the Sources renderer keys badges on source.url,
so two Anthropic document citations sharing the same source URL
collide as React keys and one badge gets dropped (or duplicated).
The chat-adapter already mints a per-citation id that folds the
position fields (search_result_index, start/end char/page/block)
into the URL, so the two citations have distinct ids even when
their URL matches. Plumb that id through SourceData and use it as
the React key for both the measurement badges and the visible
SourceBadge list. Falls back to the URL when no id is supplied
(web_search and web_fetch source parts).
* Studio: enable Anthropic doc citations on input_document blocks
Plumb citations: {enabled: true} onto the translated Anthropic document
block (both base64 and URL source branches) so the upstream actually
emits citations_delta events. Without this opt-in the inline [N] +
Sources panel plumbing added in this PR is a no-op for real user
PDF / doc uploads.
Refs https://platform.claude.com/docs/en/build-with-claude/citations
Also add edge-case coverage for the citations_delta path:
malformed citations, mixed types per document, reversed indices,
missing document_index, non-int block indices, unknown citation
type, internal _key never leaking, footnote numbering across
content blocks, and the input_document wire-through itself.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Reject unsafe citation sources, bound cited_text payload
Three follow-ups on top of #5718 surfaced by a deeper review pass:
1) javascript: / data: / vbscript: in citation source is XSS-able.
``documentCitationToSource`` was assigning ``cit.source`` straight
into ``Source.url`` and rendering it as an <a href>. A hostile
model emitting ``cit.source = "javascript:alert(document.domain)"``
would execute on click (openLink only intercepts URLs that contain
"://" or start with "mailto:", which both miss the javascript:
scheme). Restrict the navigable path to http(s):// only; anything
else falls back to the existing #anthropic-doc anchor and the
source title still renders the raw identifier for context. Also
reject CR/LF inside the URL string.
2) Frontend sources collapse distinct backend footnotes when the
citation type differs but positions match. char_location(0,5) and
page_location(0,5) over the same source previously deduped into
one entry because the id only carried position. Fold citation
type into the id anchor so the 1:1 mapping with inline [N]
markers is preserved across every citation shape.
3) ``cited_text`` was forwarded unbounded inside the synthetic
document_citations tool_event. The Sources panel trims to 240
chars for display anyway; for large RAG / search_result spans
(~10kB cited_text is plausible) this inflates SSE bytes 40x
for no UI benefit. Truncate server-side at 512 chars with an
ellipsis so the description-trim downstream still has room to
work and the wire stays bounded.
Tests grow from 21 to 22; existing 7 + edge 15 still green. Frontend
typecheck clean.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Studio: apply http(s) URL guard to all Sources-panel link sources
The previous round only filtered ``cit.source`` inside
``documentCitationToSource``. Two parallel code paths still copied
provider/tool-controlled ``URL:`` text directly into clickable
``<a href>`` Sources-panel links:
* ``parseSourcesFromResult`` in chat-adapter.ts (legacy web_search /
web_fetch tool result parser)
* ``parseSearchResults`` in tool-ui-web-search.tsx (inline tool card)
A hostile tool response like ``URL: javascript:alert(1)`` or
``URL: data:text/html,...`` was therefore still rendered as a
navigable badge in the Sources panel.
Centralise the safe-URL test (``isSafeNavigableSourceUrl``,
``isSafeHttpUrl``) using ``new URL()`` + protocol allowlist + CR/LF
rejection, and apply it to both parsers. Unsafe blocks are dropped
rather than rewritten to a hash anchor because the web_search /
web_fetch parsers have no document-index fallback.
Citation conversion now uses the same helper so the in-place
http(s) regex and CR/LF check stay in one place.
* Shorten citation comments for PR #5718
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* Studio: surface Anthropic web_fetch as a standalone Fetch pill
web_fetch used to be silently bundled with the Search pill on the
assumption that "search returns URLs, fetch reads them" is the
typical workflow. Two problems with that:
- Anthropic bills each web_fetch invocation separately from
web_search hits, so combining them made the per-message cost
surface ambiguous.
- It blocked "just fetch this one URL" workflows where the user
already knows the page they want read and does not want a search
round-trip.
Adds:
- `webFetchToolsEnabled` to the chat-runtime-store, persisted to
localStorage under `unsloth_chat_web_fetch_tools_enabled`, with a
matching `supportsBuiltinWebFetch` capability flag and a
`setWebFetchToolsEnabled` setter.
- A new Fetch pill in the chat composer, rendered next to Images and
only when the active provider returns true from
`providerSupportsBuiltinWebFetch` (Anthropic today). The pill
defaults off so per-fetch billing is always a deliberate opt-in.
- chat-page bootstraps `webFetchToolsEnabled` from the same stored-
preference fallback the other pills use.
- chat-adapter reads `webFetchToolsEnabled` directly when deciding
whether to append "web_fetch" to `enabled_tools`, decoupling it
from `toolsEnabled` (Search).
Backend translation is unchanged: when `enabled_tools` already
contains "web_fetch", `_stream_anthropic` appends the
`web_fetch_20250910` / `web_fetch_20260209` tool exactly as before
(test_anthropic_web_fetch.py pins the standalone-only path at
`test_web_fetch_tool_appended_to_request_body` and the combined
path at `test_web_fetch_combined_with_web_search_and_code_execution`).
Frontend tsc passes.
* ci: re-trigger after transient GitHub API HTTP flake (checkout + ggml-org release fetch)
* Studio: include web_fetch in the disabled-tool guard axis
Reviewer P1 / High on PR #5742 (codex + gemini): after introducing
the standalone Fetch pill, `disabledToolGuard` still only branched on
`webSearchEnabledForThisTurn`. With Fetch ON and Search OFF the
system prompt would tell Claude "you do not have web search or web
fetch tools in this conversation", which contradicts the actual tool
schema being sent and suppresses `web_fetch` tool calls, defeating
the standalone-fetch workflow this PR adds.
Treat search and fetch as a single "any web tool enabled" axis. The
guard only needs to warn the model when no web tool is wired in for
this turn; once either pill is on the model can pick the right one
from the tool schema. The existing `webLabel` already covers both
names, so the user-visible guard text stays accurate in every
combination.
tsc clean.
* ci: re-trigger after transient infra flake on Windows prebuilt / actions/checkout
* Studio: route web_fetch through per-model version dispatch
The web_fetch tool body in `_stream_anthropic` hardcoded
`web_fetch_20250910` instead of calling `_anthropic_web_fetch_version`,
so Opus 4.6 / 4.7 and Sonnet 4.6 missed the `web_fetch_20260209`
dynamic-filtering variant. The picker, the unit tests for it, and a
deliberate "follow-up" note in `test_anthropic_web_fetch.py` already
existed; this just threads it through the emission site.
Mirrors how web_search and code_execution are dispatched per model.
Old models still resolve to `web_fetch_20250910` and continue to work.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Shorten web_fetch comments for PR #5742
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* Studio: rewrite OpenAI Responses citation markers to markdown links
OpenAI's /v1/responses stream interleaves text deltas with inline
citation markers built from private-use codepoints (U+E200 / U+E201 /
U+E202) shaped like `citeSOURCE_ID`. The codepoints render
as garbled "E202" glyphs or empty boxes in most fonts, and the
markdown layer further strips them, leaving run-on text like
"citeturn1view0turn1view1turn3view0...". The url list still arrived in
the Sources panel via url_citation annotations, but the inline cite
hand-off into the prose was unreadable.
Rewrite each marker into `[N](URL)` when the matching url_citation
has already been recorded on this stream, and drop the marker
silently otherwise. The lookup uses a new `source_id` field captured
on `_record_url_citation` (accepts source_id / id / locator across
Responses API revisions). Annotations are now applied BEFORE the
delta text is rewritten so that markers and their resolving
annotation arriving in the same SSE event still resolve.
Reference: https://developers.openai.com/api/docs/guides/citation-formatting
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Preserve every source_id alias for a deduplicated url_citation
OpenAI's Responses stream cites the same URL under multiple
source_id markers when the model references different spans of the
same page. The previous dedup-by-URL kept only the first alias and
dropped the rest, so subsequent markers for the same URL never
resolved and got stripped from the prose. Switch the citation
record to a ``source_ids`` list and append new aliases on every
duplicate. The rewriter resolves any alias back to the same
citation number so the inline markers all collapse onto one footnote
rather than fanning out into bogus repeats.
Also collapse the two passes over ``all_url_citations`` in
``_record_url_citation`` into a single loop for clarity. Adds two
regression tests covering the alias-collision and mixed-shape cases.
* ci: re-trigger after flake in Studio GGUF Tool calling (rebased on main #5741 already)
* ci: re-run after transient CodeQL Python checkout auth flake
* Fix split-marker buffer + multi-source ids for PR #5713
The original rewriter only handles markers that arrive whole inside a
single response.output_text.delta event. OpenAI's stream chunks text
on byte-buffer boundaries with no awareness of the marker grammar,
so a marker can straddle two deltas (delta-1 ends with
"citetu", delta-2 starts with "rn0view0"). Each delta
was rewritten in isolation, so the half-marker leaked as garbled
"E200/E202" glyphs in the rendered prose.
Buffer the unterminated tail across deltas and concatenate it onto
the front of the next one so the rewriter sees a complete marker.
Flush the held-over tail on response.completed / response.incomplete /
[DONE], stripping any leftover private-use bytes so a never-closed
marker (truncated stream, missing annotation) never leaks.
Also handle the multi-source marker shape from the OpenAI docs --
citeid1id2 should expand to one bracket
link per resolvable id. The previous regex captured only the first
source id and silently dropped id2/id3.
Reference: https://developers.openai.com/api/docs/guides/citation-formatting
Tests: 21 new cases covering multi-source, locator suffix, marker
split across two and three deltas, unterminated marker on truncation,
late annotation resolving a buffered marker, idempotency, and the
head/tail split helper directly.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Defer citation segments until url_citation annotation arrives
The split-marker buffer already concatenates a marker that straddles
two response.output_text.delta events. But when the annotation event
for a url_citation arrives AFTER the delta that contains its inline
marker (the typical OpenAI Responses ordering), the rewriter still
saw an empty lookup table at delta time and silently stripped the
marker. The URL kept showing up in the sources panel but the inline
link reference was permanently gone.
Add _rewrite_citation_markers_partial which leaves an unresolved
marker verbatim and reports has_unresolved=True. The streaming loop
buffers any closed segment that contains an unresolved marker into a
pending_citation_segments FIFO and drains the queue on every later
annotation event, on response.completed, on response.incomplete, and
on the [DONE] sentinel. Drain order is preserved so later clean text
does not leapfrog an earlier deferred segment. End-of-stream forces a
strip so no codepoint leaks if the annotation never arrived.
Add six regression tests covering single-pass resolution, the late-
annotation two-pass case, multi-source markers with partial
resolution, mixed known and pending markers in one segment, and
idempotency on marker-free input.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Drop unterminated citation tail to prevent cite-prefix plain-text leak
`_flush_pending_marker_tail` stripped the three private-use citation
codepoints from the held-over buffer, but left the literal ``cite``
keyword plus the source id behind as plain text. A stream ending
mid-marker therefore emitted user-visible garbage like
``Some text citeturn0view0`` instead of the intended clean prose.
``pending_marker_tail`` is by construction the suffix that starts at
an unclosed ``\\ue200`` opener -- the split helper guarantees there is
no closing ``\\ue201`` byte. Without that close the marker is
meaningless: the source id cannot be resolved to a URL and the user
prose before the opener was already emitted as ``head`` on the
originating delta. Bail out before the strip step and return the
empty string. As a belt-and-braces measure also drop any orphan
``cite<sid>`` literal at the head of the buffer in case a future
caller passes a partially-terminated tail.
Update the matching ``_simulate_delta_stream`` harness in the edge
tests so it mirrors the new flush logic, and add four regression
tests covering unterminated marker with surrounding prose, marker-
only inputs, prefix-only outputs, and the split-then-close path that
still must resolve to a link.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Defer multi-source markers until all ids resolve for PR #5713
`_rewrite_citation_markers_partial` previously treated a marker as
resolved when even one token in a multi-source marker resolved,
dropping any still-pending source ids. In streamed Responses events
the annotations for a multi-source marker can arrive across separate
`annotation.added` chunks, so the caller no longer buffered that
segment for retry and the late source id was lost from the inline
citation entirely.
Flag the marker unresolved whenever any token misses the lookup so
the streamer keeps the segment pending. End-of-stream force flush
still drops unresolved tokens through `_replace_openai_citation_markers`
so locator-style suffixes (which look like unresolved ids at the token
level but only appear at end-of-stream) render cleanly.
Updated the multi-source test to assert the new pending-then-flush
behavior; locator output now lands at force-flush rather than mid
stream.
* Shorten citation marker comments for PR #5713
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* Studio: add Anthropic fast_mode toggle + surface streaming refusals
Fast mode (beta `fast-mode-2026-02-01`) lets Claude Opus 4.6 and 4.7
generate output tokens up to 2.5x faster at 6x standard Opus
pricing. The toggle lives in Configuration → Provider when the
selected Anthropic model is Opus 4.6 or 4.7 and is otherwise
hidden. Backend gates the same prefixes a second time so a stale
frontend cannot make Anthropic 400 the request, and the
`fast-mode-2026-02-01` beta header is merged onto whatever other
betas the request already needed (code-execution, compaction).
Streaming refusals (`message_delta.delta.stop_reason="refusal"` on
Claude 4 models) now surface a short user-facing notice in the
assistant message before the translated OpenAI chunk emits the
existing `finish_reason="content_filter"`. Previously the chat
bubble truncated silently because the SSE stopped mid-stream with
no visible explanation. Per the upstream docs the conversation
must be reset before continuing, so the notice tells the user
exactly that.
Reference:
- https://platform.claude.com/docs/en/build-with-claude/fast-mode
- https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/handle-streaming-refusals
Tests:
- studio/backend/tests/test_anthropic_fast_mode_and_refusal.py (8 cases
pinning fast_mode pass-through on 4.6/4.7, silent drop on Sonnet /
Haiku / older Opus / None / False, and the refusal notice + finish
reason on a synthetic refusal stream).
* Studio: drop refused Anthropic turns from the next request
Anthropic's streaming-refusal guidance says the refused assistant
turn must be removed or updated before the next call -- otherwise
the safety classifier keeps refusing. The PR only added a
user-visible notice; the partial assistant output (plus the notice
itself) still rode the next request via toOpenAIMessage.
Tag the refusal turn with an HTML-comment sentinel emitted alongside
the notice. The chat-adapter checks for that sentinel in
toOpenAIMessage and returns null, so the refused turn is excluded
from outboundMessages. The notice still renders in the transcript
(HTML comments don't display), so users keep the explanation.
* Studio: filter None finish_reason entries in test helper
test_refusal_maps_to_content_filter expects only ['content_filter']
in the finish_reasons list, but the post-PR refusal path emits a
user-visible content notice chunk first. Every _content_chunk
carries 'finish_reason: None' by construction; the helper was
appending those, so the assertion saw [None, 'content_filter']
instead of ['content_filter'].
None is not a finish reason -- it's just mid-stream delta noise.
Skip None values in _finish_reasons so the helper reflects what
the test names actually claim to check. Same fix applies cleanly
to the other helper usages (pause_turn test expects [] and the
sibling stop test expects ['stop'], both unaffected).
* Studio: cover Anthropic fast-mode edge cases
Adds 19 cases on top of the 9 in test_anthropic_fast_mode_and_refusal.
The base file pins the happy path; this file fills in the cliffs:
* Dated-snapshot prefix matching: claude-opus-4-7-2026-02-01 and
claude-opus-4-6-2026-02-01 still gate fast_mode through, while
claude-opus-4-5-2025-08-01 and claude-sonnet-4-6-2026-02-01 do not.
* Strict opt-in: a future claude-opus-4-8 or claude-opus-5 does NOT
auto-enable fast_mode -- the prefix tuple must be bumped explicitly
when a new family is whitelisted upstream.
* Beta-header merge: fast_mode coexists with code-execution-2025-08-25
and compact-2026-01-12 in one comma-separated anthropic-beta header
with no duplicates and no truncation. Pins the value to the exact
fast-mode-2026-02-01 docs token so a typo would fail CI.
* Non-destruction: fast_mode=None produces byte-identical outbound
body and headers to the version that omits the argument entirely.
Same for fast_mode=False. Guarantees the upgrade path is
non-breaking on existing Anthropic streams.
* Refusal stream ordering: the user-visible notice precedes the
finish_reason chunk so a streaming UI paints text before flipping
to content_filter. Refusal sentinel emitted exactly once. Notice
rides a normal content delta chunk with finish_reason still null.
Partial assistant deltas survive before the notice.
* Provider-side refusal coverage: a refusal on Sonnet (not just Opus)
still emits the notice + sentinel + content_filter mapping, since
refusal handling is not gated on fast-mode capability.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Persist fastMode, drop refused user message on retry
Two follow-ups on #5715:
1) sanitizeInferenceParams stripped fastMode. fastMode is in
PERSISTED_INFERENCE_PARAM_KEYS but the storage sanitizer only kept
numeric fields plus systemPrompt and trustRemoteCode, so the new
toggle was silently dropped on reload and on the
/api/chat/settings round-trip. Save it the same way trustRemoteCode
is saved.
2) Refusal recovery now also drops the triggering user turn.
Returning null from toOpenAIMessage on the assistant side left the
user prompt that caused the refusal in the outbound history, so
the very next request would re-trigger the same classifier.
Anthropic's refusal-handling guidance is explicit on this: remove
the refused turn AND the user message that triggered it before
the next call. Implemented via a pre-pass that pops the trailing
user message when an assistant carries the refusal sentinel.
Typecheck clean.
* Studio: out-of-band refusal signal + fast-mode prefix/usage/pricing fixes
The text sentinel for the Anthropic refusal drop signal was spoofable:
any assistant message containing the literal
<!--studio:anthropic-refusal--> would prune the prior user + assistant
pair on the next request. Move the signal onto a separate _toolEvent
chunk that the chat adapter latches into
assistant.metadata.custom.anthropicRefusal; assistant text can no
longer control the pruner.
Tighten the fast-mode model gate (backend + frontend) to require a "-"
family boundary so claude-opus-4-70 / claude-opus-4-7b style IDs do
not get speed: "fast" on a naive startswith match.
Use survivingMessages for the image / audio attachment scan so a
refused user turn does not gate or mis-attribute the next non-refused
turn.
Propagate Anthropic usage.speed onto the OpenAI-style usage chunk and
apply the documented 6x fast-mode multiplier in the cost calculator
(stacks with prompt-cache multipliers per the docs); expose the new
multiplier on the pricing snapshot for the UI tooltip.
Tests cover the tool-event chunk shape, the prefix-collision rejects,
usage.speed propagation, the 6x pricing math, and that the visible
refusal text carries no embedded sentinel.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Shorten fast-mode and refusal comments for PR #5715
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* Studio: surface external-provider cache hits and writes in context bar
The Anthropic / OpenAI Responses streaming paths already emit an
include_usage-style SSE chunk carrying prompt_tokens_details.cached_tokens
and cache_creation_input_tokens / cache_read_input_tokens (see
_build_usage_chunk in external_provider.py), but the chat-adapter only
read the local llama-server timings.cache_n field. As a result, the
context-usage tooltip never showed cache hits or writes for external
providers, even though the backend was computing them.
Read the external usage envelope as a fallback when timings.cache_n is
absent, and surface Anthropic cache_creation_input_tokens as a separate
"Cache writes" line in the tooltip so users can tell a cache miss from a
cache hit on a turn that both reads and writes the cache.
- ServerUsage gains optional prompt_tokens_details.cached_tokens,
cache_creation_input_tokens, cache_read_input_tokens.
- contextUsage store entry gains optional cacheWriteTokens.
- ContextUsageBar gains optional cacheWrites tooltip line.
- chat-page wires both fields through to the bar.
* Studio: render cache stats for external providers too
Reviewer round on the original PR caught three asymmetric-fix sites
where the producer side surfaced external prompt-cache stats but the
consumer side still gated on ggufContextLength (which is only ever set
for the local llama-server runtime). Result: the entire cache-stats
PR shipped invisible for Anthropic / OpenAI Responses / Gemini, which
is exactly the set of providers it was added for.
- chat-page.tsx: drop the ggufContextLength precondition on the
ContextUsageBar mount. The bar already tracks usage; let it decide
what to render based on what it knows.
- context-usage-bar.tsx: make `total` optional. When absent, drop the
"/ total" ratio + percentage progress bar + "approaching limit"
helper, and just show per-turn counters + cache stats. Bootstrap
guard tightened so an all-zero, all-undefined state still renders
nothing.
- runtime-provider.tsx: external-provider rehydration was rejected by
the `store.ggufContextLength` check. Keep the "fits inside window"
sanity check when a local context window IS known, drop it when
it isn't.
- message-timing.tsx: the per-message timing popover used a separate
"Cache hits" code path that only read llama-server's timings.cache_n.
Fall through to custom.contextUsage for external providers, and add
a parallel "Cache writes" line for Anthropic cache_creation events.
* Studio: tighten cache-stats comments
* Scope contextUsage to active checkpoint
Three follow-ups on #5736 so the relaxed external-provider render
gate does not show stale token / cache stats from a different model:
1) setCheckpoint now clears contextUsage on a real checkpoint
change. setActiveThreadId and clearCheckpoint already did this;
the most-traveled transition path (the user switching models from
the picker) leaked the prior turn's counts because they were never
cleared.
2) The external-selection branch in chat-page.tsx now also clears
contextUsage at the same time it nulls ggufContextLength /
activeNativePathToken. Without this an in-session switch from a
local model to an external provider would visibly carry the
previous local turn's counters into the new provider's bar.
3) exitCompare's rehydration is now scoped: restore the saved
usage only when the message's modelId matches the active
checkpoint AND, for local turns where a context window is known,
when the saved total fits inside that window. Without this the
bar could render a stale local-model usage on top of an external
provider, or an oversized usage object that exceeds the now-
active window.
Typecheck clean.
* Plug remaining stale-contextUsage paths
Follow-up to 042e0ac4 that catches four asymmetric-fix sites the
checkpoint-scoping pass missed:
1) setParams now also clears contextUsage on a real checkpoint
change. The local model load path in use-chat-model-runtime calls
setParams(mergeBackendRecommendedInference(...)) which mutates
params.checkpoint before refresh() eventually fires setCheckpoint;
the intermediate window rendered the previous model's counters
under the new checkpoint.
2) chat-adapter.ts setContextUsage on stream completion now gates on
the captured params.checkpoint still being active. A late
completion from provider A used to clobber the context bar after
the user switched to provider B mid-stream.
3) chat-page.tsx exitCompare rehydration no longer accepts a saved
modelId-stamped usage when the active checkpoint is empty. A user
who entered compare, cleared the model, and exited compare would
otherwise see the cleared model's stats reappear.
4) runtime-provider.tsx thread-load no longer restores legacy
unscoped usage (no modelId) unless a local context window is
known. With the relaxed external-provider render gate, old
pre-PR persisted messages without a modelId stamp could attach
their counts to an unrelated active provider.
Also switches message-timing.tsx cache-hit fallback from || to ??
so an explicit cache_n=0 is not replaced by a stale cachedTokens.
Typecheck clean.
* Shorten cache-stats comments for PR #5736
* Studio: stop leaking seeded admin pw to cross-origin callers
The "/" SPA fallback serves index.html with an inline
``window.__UNSLOTH_BOOTSTRAP__`` script containing the seeded admin
password while a password change is pending. Default web mode runs
``CORSMiddleware`` with ``allow_origins=["*"]`` + ``allow_credentials=
True``, which reflects an attacker-controlled ``Origin`` back on every
request and sets Access-Control-Allow-Credentials true. The combination
let any cross-origin page ``fetch('/')`` with credentials and read the
bootstrap admin password out of the HTML body. The API smoke
``CORS: GET / leaks bootstrap pw to cross-origin caller`` audit already
tracked this (tests/studio/studio_api_smoke.py:224) but did not gate CI.
Gate ``_inject_bootstrap`` on a same-origin check: legitimate top-level
navigations omit ``Origin`` on most engines, so the absence of the
header is treated as same-origin; when the header IS present and does
not match ``request.url.scheme://request.url.netloc`` exactly, we now
skip injecting the bootstrap tag. ``Vary: Origin`` is added so an
intermediary cache cannot serve a same-origin response (with bootstrap)
to a later cross-origin caller (and vice versa).
Coverage: ``test_index_bootstrap_origin.py`` exercises the helper with
missing / matching / evil / scheme-mismatch / port-mismatch origins.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Studio: tighten comments on bootstrap cross-origin helper
* Studio: canonicalise Origin before same-origin gate
A plain string-compare between the Origin header and request.url.netloc
misclassifies legitimate same-origin requests as cross-origin in three
scenarios:
- Browser strips the default port from Origin (https://example.com)
but Starlette's netloc keeps it (example.com:443). Per RFC 6454 the
default port is dropped on the wire, so the strings will not match
even though the requests share an origin.
- Host case differs (Origin: http://Example.com vs netloc:
example.com). Per RFC 3986 host comparison is case-insensitive.
- Scheme case differs (HTTP:// vs http://). Per RFC 3986 the scheme
is also case-insensitive.
These are usability degradations rather than security gaps (legitimate
user denied the bootstrap injection, no attacker gain), but worth
shipping so non-default Studio deployments keep the change-password
auto-fill.
Adds _canonical_origin(scheme, netloc) -> (scheme, host, port) and
compares the canonical tuples. Default-port lookup covers
http/https/ws/wss; userinfo (user:pass@) is stripped per RFC 3986
since Origin never carries credentials. Origin: "null" (sandboxed
iframes, file:// pages) and unparseable values collapse to cross-
origin so the bootstrap pw is never leaked through those paths either.
Tests: 14 cases (was 5). Covers the original same/missing/evil/
scheme/port matrix plus default-port stripping in both directions,
host + scheme case folding, Origin: null, garbage values, and
userinfo-in-netloc.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Fix IPv6 netloc parsing for PR #5739
The canonical-origin helper used ``netloc.partition(":")`` which
mis-parses bracketed IPv6 hosts (``[::1]:8902`` -> host=``[``,
port-str=``:1]:8902``). The int() then raises and the canonicaliser
returns None, so every IPv6 same-origin request is misclassified as
cross-origin and Studio refuses to inject the bootstrap pw on a
legitimate top-level nav when launched with ``unsloth studio -H ::1``.
Bracket-aware split per RFC 3986 §3.2.2, plus extra regression tests
for IPv6, opaque (data:/blob:/file:), comma-joined multi-Origin and
localhost-vs-127 cases.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Guard urlparse ValueError in same-origin gate
urlparse raises ValueError on malformed bracketed Origin values
(unclosed [, invalid IPv6 hex, text after ]) and on a few NFKC
edge cases since Py 3.8. Without a guard, a request carrying
Origin: http://[malformed surfaced as HTTP 500 from the SPA
handler rather than being treated as cross-origin per the
docstring's safer-default rule. Wrap both urlparse calls in
try/except ValueError and return False on parse failure.
Also distinguish a missing Origin header (top-level same-document
GET, treat as same-origin) from an explicit empty string (not a
valid serialised origin per RFC 6454 §6.1, treat as cross-origin).
Four new regression tests pinned down by the PR audit: malformed
IPv6 bracket, invalid IPv6 hex, bracket with trailing garbage,
and the empty Origin header.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Shorten origin-gate comments for PR #5739
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Guards against drift between the backend _ANTHROPIC_4_7_SAMPLING_REMOVED
regex and the frontend ANTHROPIC_4_7_SAMPLING_REMOVED_REGEX added in the
previous commit. If a future patch widens one without the other the panel
will either silently strip a knob the user moved or 400 on a knob the UI
should have hidden — both bad UX.
Test pins the canonical 4.7 id shapes (opus/sonnet/haiku) including dated
snapshots, and the non-4.7 ids that must NOT match (4-6, 4-5, 4-71,
future 5, gpt-4o, etc.). Sampling-params suite now 60 passing
(previously 59).
OpenAI gating was per-provider — the restrictive reasoning-class capability
applied to gpt-4o too, even though gpt-4o on /v1/responses still accepts
temperature / top_p / seed / frequency_penalty / presence_penalty. Anthropic
4.7 was the inverse: backend stripped temperature / top_p / top_k per-model
but the UI still showed the sliders, so moving a knob silently did nothing.
Split openai capabilities into OPENAI_REASONING_CAPABILITIES (current
restrictive set, used for gpt-5.x / o1 / o3 / o4) and OPENAI_CHAT_CAPABILITIES
(full sampling minus top_k and stop, used for gpt-4o and any non-reasoning id
from the registry). Mirror the backend _ANTHROPIC_4_7_SAMPLING_REMOVED regex
on the frontend so claude-(opus|sonnet|haiku)-4-7 hides temperature / top_p /
top_k in the panel instead of relying on backend strip. getProviderCapabilities
now takes an optional modelId so chat-settings-sheet and chat-adapter both
resolve the same per-model variant; no behavior change for unspecified modelId.
Verified live OpenAI / Anthropic docs:
- GPT-5 temperature must equal 1: platform.openai.com/docs/guides/reasoning,
community.openai.com/t/temperature-in-gpt-5-models/1337133
- GPT-4o accepts full sampling on Responses: docs.aimlapi.com gpt-4o ref,
OpenAI cookbook seed example
- Claude 4.7 sampling removed (400 on any non-default temperature/top_p/
top_k): platform.claude.com/docs/en/about-claude/models/whats-new-claude-4-7
160 backend routing tests still pass; frontend tsc clean.
* fix(chat_templates): check find() return value before slicing on placeholders
Two places in `construct_chat_template()` use `str.find()` for sentinel
placeholders (`{INPUT}` / `{OUTPUT}`) without checking the -1 return:
1. The `except:` fallback (around line 2464) computes
`chat_template[chat_template.find("{OUTPUT}") + len("{OUTPUT}"):]`.
If the template has no `{OUTPUT}` marker, `find()` returns -1 and the
slice starts at offset 7 (`-1 + len("{OUTPUT}")`), producing garbage
that's then `re.escape`-d and fed back into the template-recovery
regex. The user sees a confusing `IndexError` on
`response_part = response_part[0]` instead of the real problem.
2. The final trim before returning (`input_part[:input_part.find("{INPUT}")]`
and the matching `{OUTPUT}` line) silently drops the last character
when the placeholder is missing — `find()` returns -1, and `[:-1]`
slices everything except the last character, returning a corrupted
template prefix to the caller.
Replace both with an explicit `-1` check that raises a clear
`RuntimeError` naming the missing placeholder, matching the existing
guard pattern from #4923 (`try_fix_tokenizer`).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(chat_templates): also guard {INPUT} and fallback regex/separator paths
Builds on the {OUTPUT} / final-trim guards in this branch by closing
the three remaining ways the except-block fallback in
construct_chat_template() can still raise a confusing IndexError or
AttributeError on malformed templates:
1. Validate both {INPUT} and {OUTPUT} before deriving `ending`. The
regex two lines later (`{INPUT} + ending + ...`) still produced an
empty list and crashed on `response_part[0]` if {INPUT} was missing.
2. Guard the regex no-match case. Some templates contain both
placeholders but not in a recoverable two-example shape, in which
case `re.findall` returns an empty list and `[0]` raises.
3. Initialize `found = None` before the separator-search loop and
raise if the loop never sets it. Previously, if the first
iteration's `re.finditer` was empty the loop broke without binding
`found`, and `found.group(1)` raised AttributeError on the stale
int left over from the outer rfind loop.
Rephrase the final-trim error messages from internal variable names
("input_part") to user-facing wording ("instruction section") and
include a bounded (200-char) excerpt of the offending content so the
error is debuggable without being unbounded.
Add tests/python/test_construct_chat_template_validation.py covering
each failure mode with a fake tokenizer (no HF_TOKEN, no model
download, CPU-only).
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Round 19 added scale to /v1/responses based on the openai-python
SDK type, but round 20 reviewers (3/10 against) and the round 18
aggregator both noted that the live OpenAI Responses reference
limits Responses service tiers to auto/default/flex/priority. The
PR contract in the original description also lists scale only for
Chat Completions, not Responses. Studio routes OpenAI through
Responses, so forwarding scale risks a 400 from the upstream and
exposes a picker option the API does not accept.
Restore the conservative drop behaviour: only documented Responses
tiers reach the wire; legacy scale settings still validate (the
ServiceTier Literal and chat-settings storage allowlist keep it
for forward-compat).
Round 19 reviewer consensus (3/10 plus an asymmetric-fix call-out
across rounds 8/9/12/14/17): the openai-python SDK ships
service_tier as Literal["auto","default","flex","scale","priority"]
on /v1/responses, and enterprise Scale Tier customers need to opt in
explicitly. Drop the defensive scale-filter on the backend and add
"scale" to the OpenAI picker option list so the field reaches the
wire when set. Other providers remain at auto/default per their docs.
Update the routing tests so `scale` lives in the forwarded-set fixture
and the dropped-set fixture only carries truly out-of-enum values
(Anthropic-only `standard_only`, typos, empty string).
Round 18 reviewers (and earlier) noted CI noise from the routing test
helper: `_drive(coro)` ran a fresh event loop but never closed it or
called `shutdown_asyncgens`, so the MockTransport-backed httpx async
generators in the providers were finalised by GC in a later task and
emitted "Response.aiter_text.aclose was never awaited" / "Task was
destroyed but it is pending" warnings.
Explicitly close the loop after `run_until_complete`, running
`shutdown_asyncgens` first so the iterators finalise in this task.
Tests still pass and the warnings are gone.
Round 17 reviewer consensus on two extensions of the round 16 cap:
1. Direct GGUF stop forwarding (3/10 + sibling findings) — every
llama_cpp.py payload builder and routes/inference.py direct-GGUF
call site pass `stop` through unfiltered, while the external
provider helper and `_build_passthrough_payload` already strip
empty / non-string entries. Add a shared `_clean_local_stop_list`
helper in the route layer for the two callers, mirror the same
inline filter in `llama_cpp.py`'s three payload builders. Stops
`stop=["", "END"]` from a stale client 400'ing llama-server.
2. Responses bridge tool-call cap (3/10 streaming + 1/10 non-streaming)
— `_responses_stream` iterated every streamed `delta.tool_calls`
index and `_responses_non_streaming` translated every returned
`message.tool_calls` entry, even when `parallel_tool_calls=false`.
Latch the first tool-call index in the streaming bridge and drop
subsequent siblings; cap to one in the non-streaming bridge.
Matches the GGUF agentic-loop / Anthropic-passthrough caps.
Local-passthrough OpenAI paths (verbatim SSE / verbatim JSON) are
left alone because the contract is "raw upstream forwarding"; clients
calling /v1/chat/completions through Studio directly should still see
llama-server's native output.
Round 16 reviewer consensus extended the round 12c asymmetric-fix:
the GGUF agentic loop capped tool_calls to one when the caller opted
out, but every other Studio-internal path that emits tool calls from
llama-server output skipped the same guard.
Mirror the cap in three places that have full ownership of the
emitted list (passthrough verbatim paths are out of scope):
1. `AnthropicPassthroughEmitter` now takes `parallel_tool_calls` and
silently drops streamed `delta.tool_calls` entries beyond the
first index. Wired from `_anthropic_passthrough_stream`.
2. `_anthropic_passthrough_non_streaming` truncates the upstream
`message.tool_calls` list before producing `tool_use` blocks.
3. `run_safetensors_tool_loop` truncates the parsed tool_calls list
before appending the assistant message and executing tools.
`InferenceOrchestrator.generate_chat_completion_with_tools` and
the safetensors route now thread `parallel_tool_calls` through.
Also harden `_build_passthrough_payload` to strip empty / non-string
`stop` entries before forwarding to llama-server, matching the
`_normalize_stop_for_provider` shape the external-provider helper
already enforces.
Test pins the AnthropicPassthroughEmitter serial-tool-call gate.
Round 13 P1 finding: the Service tier picker rendered `null` as the
displayed `auto` and converted any explicit `auto` selection back to
`null`, so the chat-adapter's truthy guard then omitted `service_tier`
on the wire. For Anthropic the docs distinguish:
- omitting `service_tier` -> provider default
- `service_tier="auto"` -> opts into Priority Tier when available
- `service_tier="standard_only"` -> opts out
Drop the auto -> null conversion so the user's explicit pick reaches
the adapter and the wire reflects it. `null` still means "never set"
and falls through to the provider default; the existing serviceTier
allowlist already includes "auto" everywhere it matters.
Two related findings from round 12 reviewers:
1. The local GGUF tool loop in `generate_chat_completion_with_tools`
iterates every entry of `tool_calls` returned by llama-server, even
when the caller explicitly opted out of parallel tool calls. The
`parallel_tool_calls` flag is forwarded to llama-server, but llama
.cpp does not enforce it on every jinja template
(https://github.com/ggml-org/llama.cpp/issues/22043), so a model
that ignores the flag still ran multiple tools per turn. Cap
`tool_calls` to the first entry when the flag is False so the
client-side contract holds regardless of upstream behavior.
2. llama-server documents `parallel_tool_calls` as defaulting to FALSE
(https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md),
so the previous chat-adapter shape (forward only on explicit false)
meant the UI's default-on state could never enable parallel tool
calls there. Always forward the user's preference on the local
path so the toggle actually does what it says. External providers
default to true everywhere, so the external branch is unchanged.
Test pins the GGUF tool-loop cap by source-level assertion (the loop
itself is integration-only).
Round 11 P2 finding: the stop-sequence chips input rejected any draft
that strip to empty, which silently dropped pasted whitespace stops
like `"\n\n"` for blank-line halts. Local llama-server and OpenAI-
compat backends accept those; the Anthropic helper independently
filters whitespace entries before they hit the wire, so allowing them
in the UI cannot turn into a 400.
Drop the .trim() gate; reject only the truly empty draft. Single-line
Input behaviour is unchanged for the common typed-letters path.
Round 11 reviewer consensus (10/10): the `disable_parallel_tool_use`
translation added in round 11b reached the Anthropic-compat server-tool
GGUF loop but not the analogous client-tool passthrough branch. A
client sending `/v1/messages` with custom tools plus
`tool_choice: {"type":"auto","disable_parallel_tool_use":true}` therefore
took the passthrough branch with the opt-out silently dropped on the
llama-server `/v1/chat/completions` body.
Thread the translated `anthropic_parallel_tool_calls` value through
`_anthropic_passthrough_stream` and `_anthropic_passthrough_non_streaming`
into the shared `_build_passthrough_payload`, which already knows the
field. Test pins both helpers' signatures and that the field reaches
the body via the payload builder.
Gemini exposes its OpenAI-compatible endpoint at
https://generativelanguage.googleapis.com/v1beta/openai. Google's own
docs (https://ai.google.dev/gemini-api/docs/openai) list the supported
parameters and inherit OpenAI's 4-entry stop cap. Without an explicit
`stop_max=4` on the registry the default 16 leaks through and the
upstream silently drops the overflow.
Add the backend registry entry, mirror it in the frontend
`PROVIDER_STOP_MAX` map, and pin the cap with a focused unit test.
The local non-external GGUF path (llama-server) accepts frequency_penalty,
seed, stop and parallel_tool_calls, but the safetensors / HF transformers
worker has no equivalent kwargs and silently drops them. Showing the
controls there has been confusing reviewers: the UI promises a knob that
does nothing.
Gate frequencyPenalty / seed / stop / parallelToolCalls on `isGguf` for
non-external local models so safetensors sessions only show controls the
backend actually honours. External-provider gating is unchanged.
Stale persisted values from a prior GGUF session are still sent on the
wire but the safetensors worker keeps absorbing them via **_unused, so
this is a presentation-only change with no behaviour difference.
Anthropic Messages API nests `disable_parallel_tool_use` inside the
`tool_choice` object (per docs.claude.com/parallel-tool-use). The local
Anthropic-compat endpoint dropped that flag because the OpenAI shape it
translates into uses a different name and lives at the top level
instead. SDK clients (anthropic-python, anthropic-sdk-go, etc.) that
already speak this dialect therefore could not opt out of parallel
tool calls against the local GGUF model.
Extract `disable_parallel_tool_use` from the incoming tool_choice and
invert it to `parallel_tool_calls` on the agentic-loop call. Plain-chat
and existing tool_choice shapes are untouched. Added a focused unit
test that pins the dict/None/bool/string boundary cases.
Kimi K2.5/K2.6 chat schema documents temperature, top_p and a small
fixed set of knobs; seed and parallel_tool_calls are not in it. The
frontend already hides those controls (provider-capabilities.ts), so
the only way they reach Kimi is a stale client or a direct API caller.
Add them to body_omit so the registry strips them on the wire instead
of relying on the upstream to 400.
Sync the Kimi web-search bypass test to assert both fields are dropped
alongside frequency_penalty/temperature/top_p.
5/10 reviewers in the last round flagged Kimi forwarding non-default frequency_penalty as a 400 risk for K2.5 / K2.6, mirroring the existing lock on temperature and top_p. Hide the slider on the frontend and add frequency_penalty to Kimi's body_omit so even stale clients have the field stripped before the request hits the wire.
service_tier on the generic OpenAI-compatible branch was forwarding whatever value the dispatcher received, so a stale frontend could send standard_only (Anthropic) or scale to providers like Mistral that do not document the field, producing 400s. Gate the forward on an explicit accepts_service_tier=True provider registry opt-in; Anthropic and OpenAI Responses already handle service_tier inside their own helpers.
OpenRouter normalises to OpenAI's chat schema and inherits the 4-entry stop cap. The default 16-cap was too permissive; add stop_max=4 on both the backend provider registry and the frontend PROVIDER_STOP_MAX map.
The GGUF tool-iteration final-answer pass at llama_cpp.py:5182 was carrying only the legacy sampling fields. Forward frequency_penalty, seed, and parallel_tool_calls there too so the cap-exhausted path matches the per-iteration loop.
Test pins the OpenRouter 4-cap.
Kimi's official Chat Completion schema at https://platform.kimi.ai/docs/api/chat does not list seed or parallel_tool_calls. Hide both controls so users are not offered settings the upstream may silently drop or 400 on. Frequency penalty, presence penalty, and stop sequences remain exposed because Kimi documents them with full ranges.
Kimi documents max 5 stop strings AND <= 32 bytes per string at
https://platform.kimi.ai/docs/api/chat. The previous code capped
count but forwarded oversize entries, which can produce upstream
400s. Add stop_max_bytes=32 on the Kimi registry entry and apply
both checks in a new _normalize_stop_for_provider helper shared
between the default OAI-compat path and the Kimi web-search bypass.
Tests pin the byte-cap drop on both Kimi paths.
Round 5 review flagged two asymmetries:
1. Kimi web-search bypass hard-capped stops at 4 while the default OAI-compat path honours provider_info["stop_max"]. Apply the same provider-aware logic in _stream_kimi_web_search so kimi-with-search and kimi-without-search match. Also add Kimi's documented 5-stop max (https://platform.kimi.ai/docs/api/chat) to the provider registry so the cap actually fires.
2. chat-settings-sheet.tsx caps every non-Anthropic external provider at 4 stops. Replace with a per-provider getProviderStopMax helper in provider-capabilities.ts so DeepSeek, Mistral, and local backends are not artificially restricted while OpenAI Chat still hits its 4-entry hard limit and Kimi hits its documented 5-entry cap.
Tests pin the Kimi 5-cap on both Kimi paths.
Mistral chat completions uses random_seed not seed; map the field via a new seed_field on the provider registry so the new seed control actually works on Mistral. Default for other providers stays seed.
DeepSeek and Mistral both accept up to 16 stop sequences but the default OAI-compat branch was hard-capping at 4 (the OpenAI Chat limit). Studio routes the openai provider through /v1/responses not /v1/chat/completions so the 4-cap only applies if we explicitly added an openai entry. Raise the default to 16 and let per-provider stop_max overrides tighten if needed.
The local GGUF direct chat path (gguf_generate / gguf_generate_with_tools) bypassed _build_openai_passthrough_body and therefore dropped frequency_penalty, seed, stop, and parallel_tool_calls on the floor for users on the default no-tools and with-tools paths. Thread the new fields through LlamaCppBackend.generate_chat_completion and generate_chat_completion_with_tools and the two callsites that invoke them.
Also tighten comments to drop review-process narration that crept in and to remove the em dashes I had introduced in this PR's earlier commits.
Tests pin the Mistral random_seed rename, the DeepSeek 16-cap, and confirm the openai-compat default cap is 16.
PyPI release unsloth 2026.5.7 is now live. Bumps the pinned floor in
install.sh and install.ps1 from unsloth>=2026.5.6 to unsloth>=2026.5.7
so fresh installs resolve to the new wheel.
Tagged on main as v0.1.416-beta.
Round 4 reviewer consensus (~9/20 independent reviewers) flagged
service_tier=scale as a 400 risk on /v1/responses. The earlier commit
added scale based on the openai-python SDK literal, but the live
OpenAI Responses API reference, the PR's own provider matrix, and the
9-reviewer round-4 consensus all agree the documented Responses enum
is auto|default|flex|priority only. Drop scale on this path to remove
the risk.
Keeps scale on the Chat Completions / OAI-compat path where the SDK
enum is honored and where users who want Scale Tier can still select
it. The widened TypeScript ServiceTier / ServiceTierOption / api.ts
union and the storage sanitizer allowlist remain permissive so legacy
persisted "scale" values do not get silently dropped on reload; the
runtime per-provider gate makes the routing decision.
Tests are updated to pin the restricted Responses enum and the
explicit drop of scale + standard_only + bogus values.
Round 3 reviewer feedback:
- studio/backend/routes/inference.py: _build_chat_request (the
/v1/responses → /v1/chat/completions translator) was dropping
parallel_tool_calls on the floor. A Responses-API caller that set
`parallel_tool_calls=false` saw the flag accepted at the schema
layer but never reach llama-server because the translated
ChatCompletionRequest had no first-class field for it. Now that
parallel_tool_calls IS a first-class field on ChatCompletionRequest
(added by this PR's earlier commits), translate it through the
bridge so the preference actually fires.
- studio/frontend/src/features/chat/utils/chat-settings-storage.ts:
the stop sanitizer silently dropped `stop: []` instead of persisting
the empty array. That meant a user could not clear the last chip —
on reload, the previously-persisted stops came back. Persist empty
arrays explicitly so the cleared state round-trips.
- studio/backend/tests/test_sampling_params_routing.py: pin both with
the raw reproductions reviewers cited.
Round 2 reviewer feedback:
- studio/frontend/src/features/chat/types/api.ts: `OpenAIChatCompletionsRequest.service_tier` did not include `"scale"`, so the request builder in chat-adapter.ts failed typecheck after the runtime ServiceTier union widened (`Type 'ServiceTier | undefined' is not assignable...`). Widen the type to match the SDK and keep the typecheck green.
- studio/frontend/src/components/ui/stop-sequences-input.tsx: the chip editor used `draft.trim()` for storage, which silently mutated semantically meaningful stops like " End", "### ", and "\n\n". Keep the whitespace-only rejection (Anthropic 400s on those, OpenAI silently drops them) but persist the raw draft so leading/trailing whitespace inside otherwise-meaningful stops survives.
- studio/backend/core/inference/external_provider.py: the Kimi web-search bypass dropped a single string `stop="\n\n"` via `stop.strip()` while the normal default OAI-compat path forwards it verbatim. Mirror the default path's behavior here so kimi-with-search and kimi-without-search apply the same rules (asymmetric provider-path fix flagged in round-2 review).
Round-2 round of review-feedback fixes for the sampling-knobs PR:
- studio/backend/routes/chat_history.py: ChatInferenceSettings still had
the pre-PR field list with extra="forbid", so every settings save the
new frontend issued would 422 on the new keys (frequencyPenalty,
seed, stop, serviceTier, parallelToolCalls). Add the fields with the
same range / enum constraints the chat-completions schema uses, so
the settings-persistence path round-trips cleanly.
- studio/backend/routes/inference.py: _build_passthrough_payload and
_build_openai_passthrough_body now thread frequency_penalty, seed,
and parallel_tool_calls through to llama-server. The frontend exposes
these knobs for local backends; without the forwarding the UI was a
decoration. Each field is gated on `is not None` so 0 / False / "0"
still reach the body.
- studio/backend/core/inference/external_provider.py: the Kimi
$web_search bypass takes an early return into _stream_kimi_web_search
before the default OAI-compat body builder runs, so the new sampling
fields never landed on Kimi-with-search. Forward them through the
helper, with the same dedupe / truncate behavior the main path
applies to `stop`. Also extend the OpenAI Responses service_tier
allowlist to include `scale` per the live openai-python SDK
(response_create_params.py declares
Literal["auto","default","flex","scale","priority"]).
- studio/frontend/src/features/chat/provider-capabilities.ts +
types/runtime.ts: add `scale` to ServiceTier / ServiceTierOption and
surface it on the OpenAI Responses options so the UI matches the
upstream enum.
- studio/backend/tests/test_sampling_params_routing.py: add tests for
every gap above: Kimi web-search bypass forwarding, local OpenAI
passthrough forwarding, ChatSettingsPayload round-trip, and the full
Responses service_tier enum (parametrized over the five accepted
values plus a drop check for the Anthropic-only standard_only).
Anthropic Messages API rejects `disable_parallel_tool_use` as a
top-level field; it is only accepted as a property on the `tool_choice`
object. Move the inversion into a tool_choice merge that defaults to
`{type:"auto"}` when no choice is supplied, and skip the field entirely
when no tools are defined (it is a no-op without tools).
The same path also dropped `stop` chips that contain only whitespace,
because Anthropic 400s with `each stop sequence must contain
non-whitespace` on entries like " ", "\n", and "\n\n". The previous
filter only dropped truly empty strings; switch to `s.strip()` so the
common newline-stop defaults are also filtered out client-side.
Frontend persistence had three round-trip data-loss bugs:
- `VALID_SERVICE_TIERS` was missing `standard_only`, so any Anthropic
user who picked that tier lost it on the next reload.
- The settings sanitizer truncated `stop` to 4 entries on save,
which defeated the Anthropic UI cap of 16. Use 16 here and let the
per-provider stream helper cap to the wire's allowed length.
- The chat-settings sheet's `stopMaxEntries` capped local backends
(llama.cpp / vLLM / ollama / generic OpenAI-compat) at 4 even
though those backends happily accept more. Match Anthropic's 16
for the local path.
Preset policy now carries `frequencyPenalty` and `stop` so a saved
preset can fix a user's preferred decoding style. `seed`,
`serviceTier`, and `parallelToolCalls` stay out of presets because
they are per-request determinism / per-provider account / per-tool
state, not reusable preset values.
Drops the test that pinned the buggy top-level placement of
`disable_parallel_tool_use` and adds two tests for the nested shape
plus the without-tools skip path, plus a test pinning the
whitespace-stop filter against the documented Anthropic error.
* Studio: strip orphan tool_call XML from streamed visible content
The speculative-buffer state machine in
`studio/backend/core/inference/llama_cpp.py` can slice a tool_call XML
block between the silent DRAINING path and the user-visible
content_accum, depending on when in the model's emission the BUFFERING
-> STREAMING -> DRAINING transitions fire. Three leak shapes were
observed in a 2026-05-22 sweep of 900 Qwen3.5 / Qwen3.6 GGUF runs:
Pre-fix XML leak rate: 20/900 (2.22%), concentrated 6.7% on the
larger Q8 / MTP configs:
Qwen3.6-35B-A3B Q8_0 4/60 (6.7%)
Qwen3.6-35B-A3B-MTP Q4 4/60 (6.7%)
Qwen3.5-35B-A3B Q8_0 3/60 (5.0%)
Qwen3.6-27B Q8_0 3/60 (5.0%)
The existing `_TOOL_XML_RE` only matched well-formed
`<tool_call>...</tool_call>` and `<function=...></function>` pairs, so
unterminated openings (close was DRAINED) and orphan closes (opening
was DRAINED) survived the strip and reached the user.
Fix relaxes the regex to also strip:
1. Orphan opening up to end-of-string: `(?:</tool_call>|\Z)`
2. Orphan closing tag: bare `</tool_call>` / `</function>`
Verified on the full sweep: 20/900 -> 0/900 (100% of detected leaks
eliminated). 16 unit tests in `test_tool_xml_strip.py` pin all three
leak shapes plus the well-formed cases, plus parametrised checks on
the 5 actual real-world leak samples from the sweep data.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Studio: strip tail-only </parameter> orphan + tighten regex
The 2026-05-22 gdpval sweep surfaced a 4th XML-leak shape not caught
by the earlier regex: a bare `</parameter>\n\n` at end-of-buffer (7
of 192 trials, all Qwen3.5-27B + a few Qwen3.6-27B). The model emits
the full `<tool_call><function=...><parameter=...>...content...
</parameter></function></tool_call>` envelope, the speculative buffer
DRAINS the opening tags as intended, but EOS (max_tokens cutoff)
truncates the outer `</function></tool_call>` close, leaving just
`</parameter>` as the visible tail.
We strip this ONLY when end-anchored (`\s*\Z`) so legitimate
mid-text uses (user code samples, documentation discussing the
Qwen tool-call XML shape) survive. Verified on the 192-trial
gdpval corpus: before=7, after=0.
While at it, fold the five top-level alternations into three by
sharing tag-name and prefix subgroups:
<tool_call>... + <function=\w+>... + --> <(?:tool_call|function=\w+)>...
</tool_call> | </function> --> </(?:tool_call|function)>
Semantically identical (verified by replay over the 192-trial
corpus + adversarial inputs, 0 diffs) and 1.34x faster on real
workloads. Backtracking-safety pinned by two new perf guards
(256KB '<' spam, 1000x orphan opens).
Tests: 16 -> 28 (6 new functional + 4 well-formed-vs-orphan +
2 perf guards).
* Tighten comments in XML-strip regex and tests
Code says what it does; comments were repeating it. Strip the verbose
explanations down to the WHY-only bits (engine quirk, tail-anchor
rationale, real-world source of each test sample). No code changes.
inference.py: 21 -> 12 lines around _TOOL_XML_RE
test_tool_xml_strip.py: 343 -> 259 lines (-84)
Tests: 28/28 still pass.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>