unsloth/studio/backend/tests/test_sampling_params_routing.py
Daniel Han 5ec1208206 Fix consensus review findings on PR 5711
Round 3 of 3-Opus parallel review (2 reviewers HIGH on the persistence
chain, 2 HIGH on the routing chain, 1 HIGH on the test coverage gap).

HIGH fixes:
1. chat-runtime-store.ts: PERSISTED_INFERENCE_PARAM_KEYS extended from
   16 to 44 keys (28 extended samplers added). Before this, any value
   set on the Advanced Sampling sliders was lost on page reload because
   getChangedInferenceParams / getHydratedSettingsState iterate this
   list.
2. routes/chat_history.py: ChatInferenceSettings (extra="forbid")
   extended to mirror InferenceParams including fastMode + 28 new
   samplers. Without this every settings PUT containing any of those
   fields would 422.
3. routes/inference.py _build_openai_passthrough_body: was forwarding
   typical_p / mirostat / dynatemp but silently dropping dry_*, xtc_*,
   min_keep, ignore_eos, min_tokens, vLLM output knobs (skip/spaces
   special-tokens, include_stop_str_in_output, truncate_prompt_tokens),
   and llama.cpp instrumentation flags (n_keep, n_probs, cache_prompt,
   return_tokens, timings_per_token, post_sampling_probs). Now forwards
   all 18 to _build_passthrough_payload.
4. routes/inference.py _proxy_to_external_provider + external_provider.py
   stream_chat_completion: 20 extended kwargs are now plumbed through
   the route -> client -> OAI-compat body builder. Before this fix the
   chat-adapter computed top_a / vLLM output knobs / llama.cpp samplers
   on the frontend, sent them on the wire, and the route layer dropped
   them on the floor.
5. test_sampling_params_routing.py: extended
   test_chat_settings_payload_accepts_new_sampling_keys to round-trip
   every persisted field (was only 5). Added
   test_openrouter_forwards_top_a and test_vllm_forwards_output_shape_knobs
   to lock in the new wire forwarding.

MEDIUM fixes:
- providers.py: Mistral stop_max=4 (matches third-party shims; OAI docs
  publish no max but every consumer caps at 4).
- providers.py: Kimi body_omit now includes "presence_penalty" (Kimi
  k2.5/k2.6 chat schema lists temperature/top_p/max_tokens/stream/tools/
  tool_choice/thinking but not presence_penalty).
- external_provider.py: body_omit loop also pops the seed_field
  rename so a future provider with both `seed_field="random_seed"` and
  `body_omit=("seed",)` strips correctly. No current provider has both;
  defensive only.
- chat-adapter.ts: local-path parallel_tool_calls forwards only on
  explicit opt-out (matches the external-path stanza). Before this the
  field was sent on every chat from every existing local user.
- chat-settings-sheet.tsx: service tier Select now clamps the displayed
  value to a legal option for the active provider (e.g. "priority"
  saved on OpenAI, then user switches to Anthropic which only allows
  auto/standard_only -> Radix Select was showing a blank trigger).

LOW fixes:
- Em-dash cleanup: 7 em-dashes removed from provider-capabilities.ts /
  runtime.ts / chat-settings-sheet.tsx / test_sampling_params_routing.py
  per project rules.

Tests: 397/397 backend pass (sampling routing 69 plus anthropic /
openai / gemini / llama-server suites). Frontend tsc + vite build clean.
2026-05-27 16:32:21 +00:00

1430 lines
52 KiB
Python

# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""End-to-end routing tests for the new sampling parameters.
Pins the per-provider gating contract added by the
expose-sampling-params PR: each of `frequency_penalty`, `seed`, `stop`
/ `stop_sequences`, `service_tier`, `parallel_tool_calls` only appears
on the outbound body when the upstream provider actually accepts it.
The provider matrix is captured per docs:
- Anthropic Messages: accepts stop_sequences, service_tier
(auto|standard_only), disable_parallel_tool_use (inverted). REJECTS
frequency_penalty, seed, logprobs (silently dropped client-side).
- OpenAI Chat Completions (default OAI-compat branch): accepts every
field; OpenAI cloud uses `max_completion_tokens` rather than
`max_tokens`.
- OpenAI Responses (gpt-5.x / o3): rejects temperature, top_p,
frequency_penalty, seed, stop, logprobs. Accepts service_tier
(auto|default|flex|priority) and parallel_tool_calls.
"""
import asyncio
import json
import httpx
import pytest
from core.inference import external_provider as ep_mod
from core.inference.external_provider import ExternalProviderClient
def _drive(coro):
# Explicit loop lifecycle + asyncgen shutdown so the httpx /
# MockTransport-backed async generators in the providers are
# finalised in this task instead of being collected later (which
# triggers the "aiter_text aclose was never awaited" warning the
# reviewer round noticed).
loop = asyncio.new_event_loop()
try:
result = loop.run_until_complete(coro)
loop.run_until_complete(loop.shutdown_asyncgens())
return result
finally:
loop.close()
def _install_mock(monkeypatch, *, sse_payload: bytes | None = None) -> dict:
captured: dict = {}
def handler(request: httpx.Request) -> httpx.Response:
try:
captured["body"] = json.loads(request.content.decode("utf-8"))
except json.JSONDecodeError:
captured["body"] = None
captured["url"] = str(request.url)
captured["headers"] = dict(request.headers)
return httpx.Response(
200,
content = sse_payload
or (b'event: message_stop\ndata: {"type":"message_stop"}\n\n'),
headers = {"content-type": "text/event-stream"},
)
monkeypatch.setattr(
ep_mod,
"_http_client",
httpx.AsyncClient(transport = httpx.MockTransport(handler)),
)
return captured
# ── Anthropic ──────────────────────────────────────────────────────────
def _drive_anthropic(captured, **kwargs) -> dict:
async def run():
client = ExternalProviderClient(
provider_type = "anthropic",
base_url = "https://api.anthropic.com/v1",
api_key = "sk-ant-test",
)
async for _ in client.stream_chat_completion(
messages = [{"role": "user", "content": "hi"}],
model = "claude-opus-4-7",
temperature = 0.7,
top_p = 0.95,
max_tokens = 64,
**kwargs,
):
pass
await client.close()
_drive(run())
return captured["body"]
def test_anthropic_stop_sequences_forwarded_as_renamed_field(monkeypatch):
captured = _install_mock(monkeypatch)
body = _drive_anthropic(captured, stop = ["END", "DONE"])
assert body.get("stop_sequences") == ["END", "DONE"], body
# Anthropic does not have a `stop` field; the unrenamed key must not appear.
assert "stop" not in body, body
def test_anthropic_single_string_stop_is_wrapped(monkeypatch):
captured = _install_mock(monkeypatch)
body = _drive_anthropic(captured, stop = "STOPHERE")
assert body.get("stop_sequences") == ["STOPHERE"], body
def test_anthropic_empty_stop_omitted(monkeypatch):
captured = _install_mock(monkeypatch)
body = _drive_anthropic(captured, stop = [])
assert "stop_sequences" not in body, body
assert "stop" not in body, body
def test_anthropic_stop_sequences_dedup_and_drop_whitespace(monkeypatch):
"""Anthropic 400s on any stop sequence that contains no non-
whitespace character (`stop_sequences: each stop sequence must
contain non-whitespace`). Empty strings, " ", "\\n", "\\n\\n", and
other whitespace-only chips are filtered out client-side so the
request reaches the wire. Duplicates are deduped to avoid wasting
slots against the cap.
"""
captured = _install_mock(monkeypatch)
body = _drive_anthropic(
captured,
stop = ["END", "", "END", "DONE", " ", "END", "\n\n", "\t"],
)
# Order preserved on first sight, duplicates + every whitespace-only
# entry dropped.
assert body.get("stop_sequences") == ["END", "DONE"], body
def test_anthropic_single_whitespace_stop_string_dropped(monkeypatch):
"""Single-string stop="\\n\\n" must not reach the wire either."""
captured = _install_mock(monkeypatch)
body = _drive_anthropic(captured, stop = "\n\n")
assert "stop_sequences" not in body, body
def test_anthropic_stop_sequences_truncated_to_16(monkeypatch):
captured = _install_mock(monkeypatch)
body = _drive_anthropic(captured, stop = [f"S{i}" for i in range(20)])
assert len(body.get("stop_sequences", [])) == 16, body
assert body["stop_sequences"][0] == "S0"
assert body["stop_sequences"][-1] == "S15"
def test_anthropic_service_tier_forwarded_when_valid(monkeypatch):
captured = _install_mock(monkeypatch)
body = _drive_anthropic(captured, service_tier = "standard_only")
assert body.get("service_tier") == "standard_only", body
@pytest.mark.parametrize(
"bogus", ["flex", "priority", "scale", "default", "", "auto-foo"]
)
def test_anthropic_service_tier_unsupported_values_dropped(monkeypatch, bogus):
captured = _install_mock(monkeypatch)
body = _drive_anthropic(captured, service_tier = bogus)
assert "service_tier" not in body, body
def _drive_anthropic_with_tools(captured, **kwargs) -> dict:
"""Same as `_drive_anthropic` but enables a server-side tool
(`web_search`) so the request body carries `tools`. Needed to
exercise the `disable_parallel_tool_use` nesting path, which only
fires when there is at least one tool defined.
"""
enabled_tools = kwargs.pop("enabled_tools", None) or ["web_search"]
return _drive_anthropic(captured, enabled_tools = enabled_tools, **kwargs)
def test_anthropic_disable_parallel_tool_use_nested_under_tool_choice(monkeypatch):
"""`disable_parallel_tool_use` must be a property of `tool_choice`,
NOT a top-level body field. Top-level placement is rejected with
`extraneous key [disable_parallel_tool_use] is not permitted`. See
https://platform.claude.com/docs/en/agents-and-tools/tool-use/implement-tool-use.
"""
captured = _install_mock(monkeypatch)
body = _drive_anthropic_with_tools(captured, parallel_tool_calls = False)
# Top-level placement is rejected with 400.
assert "disable_parallel_tool_use" not in body, body
assert "parallel_tool_calls" not in body, body
# Flag lives on tool_choice; default type is "auto".
tc = body.get("tool_choice")
assert isinstance(tc, dict), body
assert tc.get("disable_parallel_tool_use") is True, body
assert tc.get("type") == "auto", body
def test_anthropic_disable_parallel_tool_use_skipped_without_tools(monkeypatch):
"""Without tools the flag is a no-op upstream; keep the body
minimal and never emit it at top level either.
"""
captured = _install_mock(monkeypatch)
body = _drive_anthropic(captured, parallel_tool_calls = False)
assert "disable_parallel_tool_use" not in body, body
assert "parallel_tool_calls" not in body, body
assert "tool_choice" not in body, body
def test_anthropic_parallel_tool_calls_default_not_sent(monkeypatch):
captured = _install_mock(monkeypatch)
body = _drive_anthropic_with_tools(captured, parallel_tool_calls = True)
# True is the upstream default; do not surface a tool_choice we
# would otherwise not have set, and definitely no top-level
# `disable_parallel_tool_use`.
assert "disable_parallel_tool_use" not in body, body
assert "parallel_tool_calls" not in body, body
tc = body.get("tool_choice")
if isinstance(tc, dict):
assert "disable_parallel_tool_use" not in tc, body
def test_anthropic_rejects_openai_only_knobs(monkeypatch):
"""frequency_penalty / seed are dropped at the dispatch layer.
Anthropic has no equivalent; the keyword args are not even forwarded
from stream_chat_completion to _stream_anthropic. This test pins
that no such field reaches the Messages body.
"""
captured = _install_mock(monkeypatch)
body = _drive_anthropic(
captured,
frequency_penalty = 1.5,
seed = 42,
)
assert "frequency_penalty" not in body, body
assert "seed" not in body, body
# ── OpenAI Chat Completions (default OAI-compat) ─────────────────────────
def _drive_openai_compat(captured, **kwargs) -> dict:
"""Send through the default OAI-compat branch (NOT /v1/responses).
Use qwen so the dispatcher takes the default branch and the provider
inherits the default 16-stop cap (Mistral now caps at 4 per its
third-party shims).
"""
async def run():
client = ExternalProviderClient(
provider_type = "qwen",
base_url = "https://dashscope-intl.aliyuncs.com/compatible-mode/v1",
api_key = "test-key",
)
async for _ in client.stream_chat_completion(
messages = [{"role": "user", "content": "hi"}],
model = "qwen-plus",
temperature = 0.5,
top_p = 0.9,
max_tokens = 64,
**kwargs,
):
pass
await client.close()
_drive(run())
return captured["body"]
def _oai_done_payload() -> bytes:
return b"data: [DONE]\n\n"
def test_openai_compat_forwards_frequency_penalty(monkeypatch):
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
body = _drive_openai_compat(captured, frequency_penalty = 1.25)
assert body.get("frequency_penalty") == 1.25, body
def test_openai_compat_forwards_seed(monkeypatch):
"""qwen has no seed_field override so seed forwards as `seed`."""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
body = _drive_openai_compat(captured, seed = 12345)
assert body.get("seed") == 12345, body
def test_mistral_renames_seed_to_random_seed(monkeypatch):
"""Mistral's registry sets seed_field="random_seed" so the OAI seed
is renamed on the wire. https://docs.mistral.ai/api/endpoint/chat"""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
async def run():
client = ExternalProviderClient(
provider_type = "mistral",
base_url = "https://api.mistral.ai/v1",
api_key = "mistral-test",
)
async for _ in client.stream_chat_completion(
messages = [{"role": "user", "content": "hi"}],
model = "mistral-small-latest",
temperature = 0.5,
top_p = 0.9,
max_tokens = 64,
seed = 12345,
):
pass
await client.close()
_drive(run())
body = captured["body"]
assert body.get("random_seed") == 12345, body
assert "seed" not in body, body
def test_openai_compat_seed_field_default_is_seed(monkeypatch):
"""Providers without a seed_field override get the OpenAI default."""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
async def run():
client = ExternalProviderClient(
provider_type = "deepseek",
base_url = "https://api.deepseek.com/v1",
api_key = "ds-test",
)
async for _ in client.stream_chat_completion(
messages = [{"role": "user", "content": "hi"}],
model = "deepseek-chat",
temperature = 0.5,
top_p = 0.9,
max_tokens = 64,
seed = 7,
):
pass
await client.close()
_drive(run())
body = captured["body"]
assert body.get("seed") == 7, body
assert "random_seed" not in body, body
def test_openai_compat_deepseek_stop_cap_is_16(monkeypatch):
"""DeepSeek docs allow up to 16 stop sequences; the previous
4-cap silently truncated valid configs."""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
async def run():
client = ExternalProviderClient(
provider_type = "deepseek",
base_url = "https://api.deepseek.com/v1",
api_key = "ds-test",
)
async for _ in client.stream_chat_completion(
messages = [{"role": "user", "content": "hi"}],
model = "deepseek-chat",
temperature = 0.5,
top_p = 0.9,
max_tokens = 64,
stop = [f"S{i}" for i in range(20)],
):
pass
await client.close()
_drive(run())
body = captured["body"]
assert len(body.get("stop", [])) == 16, body
def test_openai_compat_forwards_stop_array(monkeypatch):
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
body = _drive_openai_compat(captured, stop = ["END", "DONE"])
assert body.get("stop") == ["END", "DONE"], body
# The default OAI-compat branch does not rename to stop_sequences.
assert "stop_sequences" not in body, body
def test_openai_compat_truncates_stop_to_default_cap(monkeypatch):
"""Default OAI-compat cap is 16 (Qwen, DeepSeek, HuggingFace, custom);
OpenAI Chat / OpenRouter / Gemini / Mistral have tighter 4-entry caps."""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
body = _drive_openai_compat(captured, stop = [f"s{i}" for i in range(20)])
assert len(body.get("stop", [])) == 16, body
assert body["stop"][0] == "s0"
assert body["stop"][-1] == "s15"
def test_mistral_stop_cap_is_4(monkeypatch):
"""Mistral's docs publish no max but third-party shims cap at 4;
match OpenAI Chat's cap to avoid silent upstream truncation."""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
async def run():
client = ExternalProviderClient(
provider_type = "mistral",
base_url = "https://api.mistral.ai/v1",
api_key = "mistral-test",
)
async for _ in client.stream_chat_completion(
messages = [{"role": "user", "content": "hi"}],
model = "mistral-small-latest",
temperature = 0.5,
top_p = 0.9,
max_tokens = 64,
stop = [f"S{i}" for i in range(8)],
):
pass
await client.close()
_drive(run())
body = captured["body"]
assert body.get("stop") == ["S0", "S1", "S2", "S3"], body
def test_openai_compat_stop_dedup_and_drop_empties(monkeypatch):
"""Duplicates and empties shouldn't eat into the cap."""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
body = _drive_openai_compat(captured, stop = ["END", "", "END", "DONE", "FIN", "END"])
assert body.get("stop") == ["END", "DONE", "FIN"], body
def test_openai_compat_empty_stop_omitted(monkeypatch):
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
body = _drive_openai_compat(captured, stop = [])
assert "stop" not in body, body
def test_openai_compat_drops_service_tier_by_default(monkeypatch):
"""Generic OAI-compat providers (mistral, deepseek, openrouter, ...)
do not document a `service_tier` field. The dispatcher must drop
it unless the provider registry explicitly opts in with
`accepts_service_tier=True`; otherwise a stale frontend could
smuggle Anthropic/OpenAI-Responses-only values onto unrelated
providers."""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
body = _drive_openai_compat(captured, service_tier = "flex")
assert "service_tier" not in body, body
def test_openai_compat_forwards_parallel_tool_calls(monkeypatch):
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
body = _drive_openai_compat(captured, parallel_tool_calls = False)
assert body.get("parallel_tool_calls") is False, body
def test_openai_compat_omits_unset_optionals(monkeypatch):
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
body = _drive_openai_compat(captured)
# Optional knobs default to None / unset -> never appear.
assert "frequency_penalty" not in body, body
assert "seed" not in body, body
assert "stop" not in body, body
assert "service_tier" not in body, body
assert "parallel_tool_calls" not in body, body
# ── OpenAI Responses (gpt-5.x via /v1/responses) ─────────────────────────
def _responses_done_payload() -> bytes:
return (
b"event: response.completed\n"
b'data: {"type":"response.completed","response":{"usage":{}}}\n\n'
)
def _drive_openai_responses(captured, **kwargs) -> dict:
async def run():
client = ExternalProviderClient(
provider_type = "openai",
base_url = "https://api.openai.com/v1",
api_key = "sk-test",
)
async for _ in client.stream_chat_completion(
messages = [{"role": "user", "content": "hi"}],
model = "gpt-5.5",
temperature = 1.0,
top_p = 1.0,
max_tokens = 64,
**kwargs,
):
pass
await client.close()
_drive(run())
return captured["body"]
def test_openai_responses_drops_temperature_top_p(monkeypatch):
captured = _install_mock(monkeypatch, sse_payload = _responses_done_payload())
body = _drive_openai_responses(captured)
assert "temperature" not in body, body
assert "top_p" not in body, body
def test_openai_responses_drops_frequency_penalty_seed_stop(monkeypatch):
captured = _install_mock(monkeypatch, sse_payload = _responses_done_payload())
body = _drive_openai_responses(
captured,
frequency_penalty = 1.5,
seed = 99,
stop = ["END"],
)
# Responses 400s on any of these; the dispatch must drop them
# before they hit the wire.
assert "frequency_penalty" not in body, body
assert "seed" not in body, body
assert "stop" not in body, body
assert "stop_sequences" not in body, body
def test_openai_responses_forwards_service_tier(monkeypatch):
captured = _install_mock(monkeypatch, sse_payload = _responses_done_payload())
body = _drive_openai_responses(captured, service_tier = "priority")
assert body.get("service_tier") == "priority", body
@pytest.mark.parametrize("value", ["auto", "default", "flex", "priority"])
def test_openai_responses_forwards_documented_service_tiers(monkeypatch, value):
"""The live OpenAI Responses API reference lists `service_tier` as
`auto|default|flex|priority` for /v1/responses. Pin that every value
in the documented enum forwards untouched."""
captured = _install_mock(monkeypatch, sse_payload = _responses_done_payload())
body = _drive_openai_responses(captured, service_tier = value)
assert body.get("service_tier") == value, body
@pytest.mark.parametrize("bogus", ["scale", "standard_only", "bogus", ""])
def test_openai_responses_drops_undocumented_service_tier(monkeypatch, bogus):
"""`scale` and `standard_only` are not in the documented Responses
request enum; drop them client-side so a stale frontend never
sends an upstream-rejected value."""
captured = _install_mock(monkeypatch, sse_payload = _responses_done_payload())
body = _drive_openai_responses(captured, service_tier = bogus)
assert "service_tier" not in body, body
def test_openai_responses_forwards_parallel_tool_calls(monkeypatch):
captured = _install_mock(monkeypatch, sse_payload = _responses_done_payload())
body = _drive_openai_responses(captured, parallel_tool_calls = False)
assert body.get("parallel_tool_calls") is False, body
def test_openai_responses_omits_unset_optionals(monkeypatch):
captured = _install_mock(monkeypatch, sse_payload = _responses_done_payload())
body = _drive_openai_responses(captured)
assert "service_tier" not in body, body
assert "parallel_tool_calls" not in body, body
# ── Schema-level smoke tests ─────────────────────────────────────────────
def test_chat_completion_request_accepts_new_sampling_fields():
from models.inference import ChatCompletionRequest
payload = ChatCompletionRequest.model_validate(
{
"messages": [{"role": "user", "content": "hi"}],
"frequency_penalty": -1.0,
"seed": 0,
"stop": ["END"],
"service_tier": "auto",
"parallel_tool_calls": True,
}
)
assert payload.frequency_penalty == -1.0
assert payload.seed == 0
assert payload.stop == ["END"]
assert payload.service_tier == "auto"
assert payload.parallel_tool_calls is True
def test_chat_completion_request_rejects_bad_service_tier():
import pydantic
from models.inference import ChatCompletionRequest
with pytest.raises(pydantic.ValidationError):
ChatCompletionRequest.model_validate(
{
"messages": [{"role": "user", "content": "hi"}],
"service_tier": "bogus",
}
)
def test_chat_completion_request_clamps_frequency_penalty_range():
import pydantic
from models.inference import ChatCompletionRequest
with pytest.raises(pydantic.ValidationError):
ChatCompletionRequest.model_validate(
{
"messages": [{"role": "user", "content": "hi"}],
"frequency_penalty": 3.0,
}
)
with pytest.raises(pydantic.ValidationError):
ChatCompletionRequest.model_validate(
{
"messages": [{"role": "user", "content": "hi"}],
"frequency_penalty": -3.0,
}
)
# ── Kimi web-search bypass forwards new sampling fields ────────────────
def test_kimi_web_search_bypass_forwards_new_sampling_fields(monkeypatch):
"""The Kimi $web_search path takes an early return into
`_stream_kimi_web_search` before the default OAI-compat body
builder runs; forwarding here keeps Kimi-with-search and
Kimi-without-search in lockstep."""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
async def run():
client = ExternalProviderClient(
provider_type = "kimi",
base_url = "https://api.moonshot.ai/v1",
api_key = "kimi-test",
)
async for _ in client.stream_chat_completion(
messages = [{"role": "user", "content": "hi"}],
model = "kimi-k2.6",
temperature = 1.0,
top_p = 1.0,
max_tokens = 256,
enabled_tools = ["web_search"],
presence_penalty = 0.5,
frequency_penalty = 1.25,
seed = 7,
stop = ["END"],
parallel_tool_calls = False,
):
pass
await client.close()
_drive(run())
body = captured["body"]
# Kimi locks these: stripped by body_omit in providers.py.
assert "frequency_penalty" not in body, body
assert "temperature" not in body, body
assert "top_p" not in body, body
assert "seed" not in body, body
assert "parallel_tool_calls" not in body, body
assert "presence_penalty" not in body, body
# Knobs not on Kimi's drop-list forward through the bypass.
assert body.get("stop") == ["END"], body
def test_kimi_web_search_uses_kimi_stop_cap_5(monkeypatch):
"""Kimi documents a 5-stop max; the web-search bypass must honour
`provider_info["stop_max"]` rather than the OpenAI 4-cap or the
permissive default."""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
async def run():
client = ExternalProviderClient(
provider_type = "kimi",
base_url = "https://api.moonshot.ai/v1",
api_key = "kimi-test",
)
async for _ in client.stream_chat_completion(
messages = [{"role": "user", "content": "hi"}],
model = "kimi-k2.6",
temperature = 1.0,
top_p = 1.0,
max_tokens = 256,
enabled_tools = ["web_search"],
stop = [f"S{i}" for i in range(10)],
):
pass
await client.close()
_drive(run())
body = captured["body"]
assert len(body.get("stop", [])) == 5, body
assert body["stop"] == ["S0", "S1", "S2", "S3", "S4"], body
def test_openrouter_stop_cap_is_4(monkeypatch):
"""OpenRouter normalises to OpenAI's chat schema and inherits the
4-entry stop cap; the default 16-cap is too permissive for it."""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
async def run():
client = ExternalProviderClient(
provider_type = "openrouter",
base_url = "https://openrouter.ai/api/v1",
api_key = "or-test",
)
async for _ in client.stream_chat_completion(
messages = [{"role": "user", "content": "hi"}],
model = "openai/gpt-4o",
temperature = 0.5,
top_p = 0.9,
max_tokens = 64,
stop = [f"S{i}" for i in range(10)],
):
pass
await client.close()
_drive(run())
body = captured["body"]
assert len(body.get("stop", [])) == 4, body
assert body["stop"] == ["S0", "S1", "S2", "S3"], body
def test_gemini_stop_sequences_capped_to_5(monkeypatch):
"""Native Gemini API forwards `stop` as generationConfig.stopSequences,
capped at 5 per https://ai.google.dev/api/generate-content#generationconfig."""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
async def run():
client = ExternalProviderClient(
provider_type = "gemini",
base_url = "https://generativelanguage.googleapis.com/v1beta/openai",
api_key = "gemini-test",
)
async for _ in client.stream_chat_completion(
messages = [{"role": "user", "content": "hi"}],
model = "gemini-3.1-pro-preview",
temperature = 0.5,
top_p = 0.9,
max_tokens = 64,
stop = [f"S{i}" for i in range(10)],
):
pass
await client.close()
_drive(run())
body = captured["body"]
gen_config = body.get("generationConfig", {})
assert gen_config.get("stopSequences") == ["S0", "S1", "S2", "S3", "S4"], body
def test_openrouter_forwards_top_a(monkeypatch):
"""OpenRouter exposes `top_a` (the tail-cut sampler). The frontend
capability map lets the value through; verify the OAI-compat body
builder actually carries it to the wire."""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
async def run():
client = ExternalProviderClient(
provider_type = "openrouter",
base_url = "https://openrouter.ai/api/v1",
api_key = "or-test",
)
async for _ in client.stream_chat_completion(
messages = [{"role": "user", "content": "hi"}],
model = "openai/gpt-4o",
temperature = 0.5,
top_p = 0.9,
max_tokens = 64,
top_a = 0.25,
):
pass
await client.close()
_drive(run())
body = captured["body"]
assert body.get("top_a") == 0.25, body
def test_vllm_forwards_output_shape_knobs(monkeypatch):
"""vLLM accepts skip_special_tokens / spaces_between_special_tokens /
include_stop_str_in_output / truncate_prompt_tokens. Verify the
OAI-compat external proxy forwards them to the wire body."""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
async def run():
client = ExternalProviderClient(
provider_type = "vllm",
base_url = "https://vllm.example.com/v1",
api_key = "vllm-test",
)
async for _ in client.stream_chat_completion(
messages = [{"role": "user", "content": "hi"}],
model = "Qwen/Qwen3-4B",
temperature = 0.5,
top_p = 0.9,
max_tokens = 64,
skip_special_tokens = False,
spaces_between_special_tokens = False,
include_stop_str_in_output = True,
truncate_prompt_tokens = 2048,
min_tokens = 8,
ignore_eos = True,
):
pass
await client.close()
_drive(run())
body = captured["body"]
assert body.get("skip_special_tokens") is False, body
assert body.get("spaces_between_special_tokens") is False, body
assert body.get("include_stop_str_in_output") is True, body
assert body.get("truncate_prompt_tokens") == 2048, body
assert body.get("min_tokens") == 8, body
assert body.get("ignore_eos") is True, body
def test_kimi_drops_stop_strings_over_32_bytes(monkeypatch):
"""Kimi limits each stop string to <= 32 bytes per
https://platform.kimi.ai/docs/api/chat. Drop overlong entries
client-side so a stale UI cannot 400 the request."""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
async def run():
client = ExternalProviderClient(
provider_type = "kimi",
base_url = "https://api.moonshot.ai/v1",
api_key = "kimi-test",
)
async for _ in client.stream_chat_completion(
messages = [{"role": "user", "content": "hi"}],
model = "kimi-k2.6",
temperature = 1.0,
top_p = 1.0,
max_tokens = 256,
stop = ["END", "x" * 33, "DONE"],
):
pass
await client.close()
_drive(run())
body = captured["body"]
assert body.get("stop") == ["END", "DONE"], body
def test_kimi_web_search_drops_stop_strings_over_32_bytes(monkeypatch):
"""Same byte cap applies to the Kimi web-search bypass."""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
async def run():
client = ExternalProviderClient(
provider_type = "kimi",
base_url = "https://api.moonshot.ai/v1",
api_key = "kimi-test",
)
async for _ in client.stream_chat_completion(
messages = [{"role": "user", "content": "hi"}],
model = "kimi-k2.6",
temperature = 1.0,
top_p = 1.0,
max_tokens = 256,
enabled_tools = ["web_search"],
stop = ["END", "x" * 40, "DONE"],
):
pass
await client.close()
_drive(run())
body = captured["body"]
assert body.get("stop") == ["END", "DONE"], body
def test_kimi_default_path_uses_kimi_stop_cap_5(monkeypatch):
"""The normal Kimi path must also honour the documented 5-cap."""
captured = _install_mock(monkeypatch, sse_payload = _oai_done_payload())
async def run():
client = ExternalProviderClient(
provider_type = "kimi",
base_url = "https://api.moonshot.ai/v1",
api_key = "kimi-test",
)
async for _ in client.stream_chat_completion(
messages = [{"role": "user", "content": "hi"}],
model = "kimi-k2.6",
temperature = 1.0,
top_p = 1.0,
max_tokens = 256,
stop = [f"S{i}" for i in range(10)],
):
pass
await client.close()
_drive(run())
body = captured["body"]
assert len(body.get("stop", [])) == 5, body
# ── Local OpenAI passthrough forwards new sampling fields ──────────────
def test_local_openai_passthrough_forwards_new_sampling_fields():
"""`_build_openai_passthrough_body` forwards frequency_penalty,
seed, stop, and parallel_tool_calls to llama-server."""
from models.inference import ChatCompletionRequest
from routes.inference import _build_openai_passthrough_body
payload = ChatCompletionRequest.model_validate(
{
"messages": [{"role": "user", "content": "hi"}],
"stream": True,
"frequency_penalty": 1.25,
"seed": 123,
"stop": ["END"],
"parallel_tool_calls": False,
}
)
body = _build_openai_passthrough_body(payload, backend_ctx = 4096)
assert body["frequency_penalty"] == 1.25, body
assert body["seed"] == 123, body
assert body["stop"] == ["END"], body
assert body["parallel_tool_calls"] is False, body
# ── Responses → ChatCompletions bridge preserves parallel_tool_calls ──
def test_responses_to_chat_bridge_preserves_parallel_tool_calls():
"""`_build_chat_request` (the /v1/responses to /v1/chat/completions
translator) must forward parallel_tool_calls so a Responses-API
caller's preference reaches llama-server."""
from models.inference import ChatMessage, ResponsesRequest
from routes.inference import _build_chat_request, _build_openai_passthrough_body
payload = ResponsesRequest(
input = "hi",
stream = True,
parallel_tool_calls = False,
)
chat_req = _build_chat_request(
payload,
[ChatMessage(role = "user", content = "hi")],
stream = True,
)
assert chat_req.parallel_tool_calls is False, chat_req
body = _build_openai_passthrough_body(chat_req, backend_ctx = 4096)
assert body["parallel_tool_calls"] is False, body
def test_responses_to_chat_bridge_omits_unset_parallel_tool_calls():
"""Unset parallel_tool_calls (None) must not appear on the
translated body; the upstream default is true everywhere so
forwarding None would over-specify."""
from models.inference import ChatMessage, ResponsesRequest
from routes.inference import _build_chat_request, _build_openai_passthrough_body
payload = ResponsesRequest(input = "hi", stream = True)
chat_req = _build_chat_request(
payload,
[ChatMessage(role = "user", content = "hi")],
stream = True,
)
assert chat_req.parallel_tool_calls is None, chat_req
body = _build_openai_passthrough_body(chat_req, backend_ctx = 4096)
assert "parallel_tool_calls" not in body, body
# ── Backend ChatInferenceSettings schema accepts new fields ────────────
def test_chat_settings_payload_accepts_new_sampling_keys():
"""ChatSettingsPayload has extra="forbid" so the new keys must be
listed explicitly; otherwise every settings save with any of them
422s. Pin the round-trip for every persisted sampler the frontend
can emit (see PERSISTED_INFERENCE_PARAM_KEYS in chat-runtime-store.ts)."""
from routes.chat_history import ChatSettingsPayload
inference = {
"frequencyPenalty": 0.7,
"seed": 42,
"stop": ["END"],
"serviceTier": "standard_only",
"parallelToolCalls": False,
"fastMode": True,
# Extended llama.cpp / vLLM / OpenRouter samplers.
"typicalP": 0.85,
"topNSigma": 2.5,
"repeatLastN": 64,
"dynatempRange": 0.3,
"dynatempExponent": 1.2,
"mirostat": 2,
"mirostatTau": 4.0,
"mirostatEta": 0.15,
"topA": 0.2,
"dryMultiplier": 0.8,
"dryBase": 1.75,
"dryAllowedLength": 2,
"dryPenaltyLastN": -1,
"xtcProbability": 0.5,
"xtcThreshold": 0.1,
"minKeep": 5,
"ignoreEos": True,
"minTokens": 10,
"skipSpecialTokens": False,
"spacesBetweenSpecialTokens": False,
"includeStopStrInOutput": True,
"truncatePromptTokens": 1024,
"nKeep": -1,
"nProbs": 5,
"cachePrompt": False,
"returnTokens": True,
"timingsPerToken": True,
"postSamplingProbs": True,
}
parsed = ChatSettingsPayload.model_validate({"inferenceParams": inference})
ip = parsed.inferenceParams
assert ip is not None
for key, expected in inference.items():
assert getattr(ip, key) == expected, f"{key} did not round-trip"
# ── Local /v1/messages: disable_parallel_tool_use translation ──────────
def test_local_anthropic_disable_parallel_tool_use_translation():
"""Anthropic nests `disable_parallel_tool_use` under `tool_choice`
(per docs.claude.com). The local /v1/messages GGUF tool path must
invert it into OpenAI-shaped `parallel_tool_calls` so third-party
clients (Claude SDK, LiteLLM in passthrough mode) opt out of
parallel calls successfully even on the local model."""
# Mirror the extraction logic in routes/inference.py:anthropic_messages.
def _extract(tc):
if isinstance(tc, dict):
v = tc.get("disable_parallel_tool_use")
if isinstance(v, bool):
return not v
return None
assert _extract({"type": "auto", "disable_parallel_tool_use": True}) is False
assert _extract({"type": "any", "disable_parallel_tool_use": False}) is True
assert _extract({"type": "auto"}) is None
assert _extract(None) is None
assert _extract("auto") is None # string form (non-dict) → no opinion
assert _extract({"type": "auto", "disable_parallel_tool_use": "yes"}) is None
def test_anthropic_passthrough_emitter_serialises_tool_calls_on_opt_out():
"""When the Anthropic-compat passthrough is asked to disable
parallel tool calls, `AnthropicPassthroughEmitter.feed_chunk()`
must drop every streamed `delta.tool_calls` entry beyond the
first index, matching the GGUF agentic-loop client-side cap and
keeping the wire-side `disable_parallel_tool_use=true` honest
even when llama-server's jinja template ignores it."""
from core.inference.anthropic_compat import AnthropicPassthroughEmitter
emitter = AnthropicPassthroughEmitter(parallel_tool_calls = False)
emitter.start("msg_x", "test-model")
events = emitter.feed_chunk(
{
"choices": [
{
"delta": {
"tool_calls": [
{
"index": 0,
"id": "call_a",
"function": {"name": "first", "arguments": "{"},
},
{
"index": 1,
"id": "call_b",
"function": {"name": "second", "arguments": "{"},
},
]
}
}
]
}
)
joined = "\n".join(events)
assert "first" in joined, joined
assert "second" not in joined, joined
emitter_open = AnthropicPassthroughEmitter(parallel_tool_calls = True)
emitter_open.start("msg_y", "test-model")
events_open = emitter_open.feed_chunk(
{
"choices": [
{
"delta": {
"tool_calls": [
{"index": 0, "id": "a", "function": {"name": "x"}},
{"index": 1, "id": "b", "function": {"name": "y"}},
]
}
}
]
}
)
joined_open = "\n".join(events_open)
assert "x" in joined_open and "y" in joined_open, joined_open
def test_gguf_tool_loop_enforces_parallel_tool_calls_false():
"""llama.cpp's `parallel_tool_calls` flag is not enforced by every
jinja template (see ggml-org/llama.cpp#22043), so when the caller
opted out we must cap tool_calls to the first entry before the
agentic loop executes them. The cap is a single-line slice in
`generate_chat_completion_with_tools`; pin the contract."""
from pathlib import Path
src = Path(__file__).resolve().parent.parent / "core/inference/llama_cpp.py"
text = src.read_text()
assert "if parallel_tool_calls is False and tool_calls" in text, (
"GGUF tool loop must enforce parallel_tool_calls=False by "
"truncating tool_calls before assistant_msg is built; that "
"is the client-side guarantee llama-server's flag does not "
"give us. See routes/inference.py and chat-adapter.ts for "
"the wire-side forwarding of the same flag."
)
assert "tool_calls = tool_calls[:1]" in text
def test_local_anthropic_passthrough_helpers_accept_parallel_tool_calls():
"""The Anthropic-compat client-tool passthrough helpers
(`_anthropic_passthrough_stream` /
`_anthropic_passthrough_non_streaming`) must accept and forward
`parallel_tool_calls` through `_build_passthrough_payload` so the
`disable_parallel_tool_use` translation works on the client-tool
branch the same way it does on the server-tool loop. Verified by
introspecting the signatures and confirming the field reaches the
body via the shared payload builder."""
import inspect
from routes import inference as route_mod
for fn in (
route_mod._anthropic_passthrough_stream,
route_mod._anthropic_passthrough_non_streaming,
):
params = inspect.signature(fn).parameters
assert "parallel_tool_calls" in params, (
f"{fn.__name__} must accept parallel_tool_calls so the "
"Anthropic disable_parallel_tool_use translation reaches "
"the llama-server body on the client-tool branch"
)
body = route_mod._build_passthrough_payload(
openai_messages = [{"role": "user", "content": "hi"}],
openai_tools = [{"type": "function", "function": {"name": "x"}}],
temperature = 0.7,
top_p = 0.95,
top_k = 20,
max_tokens = 64,
stream = True,
parallel_tool_calls = False,
)
assert body.get("parallel_tool_calls") is False, body
def test_local_passthrough_forwards_extended_llama_cpp_samplers():
"""The local llama.cpp passthrough payload builder must forward
the extended sampler chain when set: top_n_sigma, repeat_last_n,
dynatemp_range/exponent, mirostat/mirostat_tau/mirostat_eta. Each
is gated `is not None` so a default-off value (e.g. mirostat=0) is
still forwarded explicitly when the caller opted in.
"""
from routes import inference as route_mod
body = route_mod._build_passthrough_payload(
openai_messages = [{"role": "user", "content": "hi"}],
openai_tools = None,
temperature = 0.6,
top_p = 0.95,
top_k = 20,
max_tokens = 64,
stream = True,
top_n_sigma = 1.5,
repeat_last_n = 128,
dynatemp_range = 0.2,
dynatemp_exponent = 1.5,
mirostat = 2,
mirostat_tau = 5.0,
mirostat_eta = 0.1,
)
assert body.get("top_n_sigma") == 1.5
assert body.get("repeat_last_n") == 128
assert body.get("dynatemp_range") == 0.2
assert body.get("dynatemp_exponent") == 1.5
assert body.get("mirostat") == 2
assert body.get("mirostat_tau") == 5.0
assert body.get("mirostat_eta") == 0.1
# Unset = absent from body so llama-server falls back to defaults.
body2 = route_mod._build_passthrough_payload(
openai_messages = [{"role": "user", "content": "hi"}],
openai_tools = None,
temperature = 0.6,
top_p = 0.95,
top_k = 20,
max_tokens = 64,
stream = True,
)
for key in (
"top_n_sigma",
"repeat_last_n",
"dynatemp_range",
"dynatemp_exponent",
"mirostat",
"mirostat_tau",
"mirostat_eta",
):
assert key not in body2, body2
def test_local_passthrough_forwards_dry_xtc_min_keep_eos_min_tokens():
"""The local llama.cpp passthrough payload builder must forward the
DRY (4-field) + XTC (2-field) + min_keep + ignore_eos + min_tokens
chain when set. Each is gated `is not None` so an explicit
upstream-default value (e.g. min_keep=0, ignore_eos=False) still
reaches the wire when the caller opted in.
"""
from routes import inference as route_mod
body = route_mod._build_passthrough_payload(
openai_messages = [{"role": "user", "content": "hi"}],
openai_tools = None,
temperature = 0.6,
top_p = 0.95,
top_k = 20,
max_tokens = 64,
stream = True,
dry_multiplier = 0.8,
dry_base = 1.75,
dry_allowed_length = 3,
dry_penalty_last_n = -1,
xtc_probability = 0.5,
xtc_threshold = 0.1,
min_keep = 1,
ignore_eos = True,
min_tokens = 16,
)
assert body.get("dry_multiplier") == 0.8
assert body.get("dry_base") == 1.75
assert body.get("dry_allowed_length") == 3
assert body.get("dry_penalty_last_n") == -1
assert body.get("xtc_probability") == 0.5
assert body.get("xtc_threshold") == 0.1
assert body.get("min_keep") == 1
assert body.get("ignore_eos") is True
assert body.get("min_tokens") == 16
# Unset = absent from body. Matches the upstream "use default"
# contract: llama-server / vLLM apply their own defaults instead.
body2 = route_mod._build_passthrough_payload(
openai_messages = [{"role": "user", "content": "hi"}],
openai_tools = None,
temperature = 0.6,
top_p = 0.95,
top_k = 20,
max_tokens = 64,
stream = True,
)
for key in (
"dry_multiplier",
"dry_base",
"dry_allowed_length",
"dry_penalty_last_n",
"xtc_probability",
"xtc_threshold",
"min_keep",
"ignore_eos",
"min_tokens",
):
assert key not in body2, body2
def test_local_passthrough_forwards_vllm_output_and_llama_cpp_instrumentation():
"""Round-trip the 10 extra knobs added in round 4:
skip_special_tokens / spaces_between_special_tokens /
include_stop_str_in_output / truncate_prompt_tokens (vLLM
SamplingParams) + n_keep / n_probs / cache_prompt / return_tokens /
timings_per_token / post_sampling_probs (llama-server README).
Each is gated `is not None` so explicit upstream-default values
(skip_special_tokens=True, cache_prompt=True, etc) still reach the
wire when the caller opted in.
"""
from routes import inference as route_mod
body = route_mod._build_passthrough_payload(
openai_messages = [{"role": "user", "content": "hi"}],
openai_tools = None,
temperature = 0.6,
top_p = 0.95,
top_k = 20,
max_tokens = 64,
stream = True,
skip_special_tokens = False,
spaces_between_special_tokens = False,
include_stop_str_in_output = True,
truncate_prompt_tokens = 4096,
n_keep = -1,
n_probs = 5,
cache_prompt = False,
return_tokens = True,
timings_per_token = True,
post_sampling_probs = True,
)
assert body.get("skip_special_tokens") is False
assert body.get("spaces_between_special_tokens") is False
assert body.get("include_stop_str_in_output") is True
assert body.get("truncate_prompt_tokens") == 4096
assert body.get("n_keep") == -1
assert body.get("n_probs") == 5
assert body.get("cache_prompt") is False
assert body.get("return_tokens") is True
assert body.get("timings_per_token") is True
assert body.get("post_sampling_probs") is True
# Unset = absent from body so each backend applies its own default.
body2 = route_mod._build_passthrough_payload(
openai_messages = [{"role": "user", "content": "hi"}],
openai_tools = None,
temperature = 0.6,
top_p = 0.95,
top_k = 20,
max_tokens = 64,
stream = True,
)
for key in (
"skip_special_tokens",
"spaces_between_special_tokens",
"include_stop_str_in_output",
"truncate_prompt_tokens",
"n_keep",
"n_probs",
"cache_prompt",
"return_tokens",
"timings_per_token",
"post_sampling_probs",
):
assert key not in body2, body2
def test_local_passthrough_forwards_typical_p_when_set():
"""`typical_p` is a llama.cpp-specific sampler (`typ_p` in the
sampler chain). The local-llama-cpp passthrough payload builder must
forward it when set so the chat-adapter can opt in for local
backends without the field bleeding into external providers (whose
capability map gates it off)."""
from routes import inference as route_mod
body = route_mod._build_passthrough_payload(
openai_messages = [{"role": "user", "content": "hi"}],
openai_tools = None,
temperature = 0.6,
top_p = 0.95,
top_k = 20,
max_tokens = 64,
stream = True,
typical_p = 0.7,
)
assert body.get("typical_p") == 0.7, body
# When unset, the field is omitted entirely so llama-server falls
# back to its 1.0 default.
body2 = route_mod._build_passthrough_payload(
openai_messages = [{"role": "user", "content": "hi"}],
openai_tools = None,
temperature = 0.6,
top_p = 0.95,
top_k = 20,
max_tokens = 64,
stream = True,
)
assert "typical_p" not in body2, body2
def test_anthropic_4_7_sampling_removed_regex_matches_expected_ids():
"""Pin the canonical Claude 4.7 model-id shape so the frontend
ANTHROPIC_4_7_SAMPLING_REMOVED_REGEX in
studio/frontend/src/features/chat/provider-capabilities.ts stays
in lockstep with the backend strip in
external_provider._stream_anthropic.
Drift would mean the panel either silently strips a knob the user
moved (UI shows, wire drops) or the wire 400s after the user moved
a knob the UI should have hidden. Both are user-visible bugs.
"""
from core.inference.external_provider import (
_ANTHROPIC_4_7_SAMPLING_REMOVED as RX,
)
# Only Opus shipped in the 4.7 generation per
# platform.claude.com/docs/en/about-claude/models/overview; Sonnet
# stops at 4.6 and Haiku at 4.5. Pin both directions explicitly so
# the regex never widens by accident.
should_match = [
"claude-opus-4-7",
"claude-opus-4-7-20260418",
"claude-opus-4-7.1",
]
should_not_match = [
"claude-sonnet-4-7",
"claude-haiku-4-7",
"claude-opus-4-6",
"claude-sonnet-4-6",
"claude-haiku-4-5",
"claude-opus-4-71",
"claude-opus-5",
"claude-3-opus",
"gpt-4o",
]
for mid in should_match:
assert RX.match(mid), f"{mid!r} should match 4.7 sampling-removed regex"
for mid in should_not_match:
assert not RX.match(mid), f"{mid!r} should NOT match 4.7 regex"
def test_deepseek_payload_omits_seed_and_parallel_tool_calls():
"""DeepSeek's published /chat/completions schema lists
messages/model/thinking/max_tokens/response_format/stop/stream/
temperature/top_p/tools/tool_choice/logprobs/top_logprobs/user_id
only. `seed` and `parallel_tool_calls` are not in the schema; the
capability bucket hides them so the chat-adapter never sends them.
Source:
https://api-docs.deepseek.com/api/create-chat-completion
"""
# Frontend capability flags are the source of truth. Re-derive them
# by reading the TS file as text (the backend has no JS engine) and
# confirm the deepseek bucket has seed:false + parallelToolCalls:false.
from pathlib import Path
src = (
Path(__file__).resolve().parents[2]
/ "frontend"
/ "src"
/ "features"
/ "chat"
/ "provider-capabilities.ts"
).read_text(encoding = "utf-8")
deepseek_idx = src.index(" deepseek: {")
end = src.index("},", deepseek_idx)
bucket = src[deepseek_idx:end]
assert "seed: false" in bucket, bucket
assert "parallelToolCalls: false" in bucket, bucket