Adds the missing sampling parameters that the upstream APIs accept and that Studio's chat UI previously hid. Each knob is gated per provider so the picker never offers a field the upstream would 400 on, and the per-provider stream functions translate / drop fields to match each API's naming. New `InferenceParams` fields (round-trip through PersistedInferenceParams and the chat-settings server store automatically): - frequencyPenalty (-2..2): OpenAI Chat Completions only. - seed (int | null): OpenAI Chat + OpenAI-compat local backends. - stop (string[]): all OpenAI Chat + Anthropic Messages. Backend truncates to 4 entries on OpenAI Chat per docs and renames to `stop_sequences` on Anthropic. - serviceTier (auto|default|flex|priority|scale|standard_only): per-provider enum sets resolved by getServiceTierOptions. - parallelToolCalls (bool, default true): forwarded as `parallel_tool_calls` on both OpenAI APIs and inverted into `disable_parallel_tool_use` on Anthropic. OpenAI Responses (gpt-5.x / o3) explicitly drops frequencyPenalty / seed / stop alongside the existing temperature / top_p drop, since the upstream 400s on all of them. service_tier on Responses accepts a subset (no `scale`) which the dispatch already enforces. UI rows land in the existing Sampling section of the chat settings sheet using ParamSlider (frequency penalty), a numeric Input (seed), a new chips editor `StopSequencesInput` (stop), Select (service tier), and Switch (parallel tool calls). Each row's visibility follows the new ProviderCapabilities flag. Tests pin the gating contract: stop_sequences renamed on Anthropic, 4-entry truncation on OpenAI Chat, every Responses-rejected field dropped, schema-level validation for the service_tier Literal and frequency_penalty range. Plan: plans/hashed-riding-porcupine.md |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| _html_to_md.py | ||
| anthropic_compat.py | ||
| audio_codecs.py | ||
| chat_template_helpers.py | ||
| defaults.py | ||
| external_provider.py | ||
| inference.py | ||
| key_exchange.py | ||
| llama_cpp.py | ||
| llama_server_args.py | ||
| mlx_inference.py | ||
| orchestrator.py | ||
| pricing.py | ||
| providers.py | ||
| safetensors_agentic.py | ||
| tool_call_parser.py | ||
| tools.py | ||
| worker.py | ||