- Drop `scale` from the OpenAI service-tier picker (frontend types and
picker option list). OpenAI in Studio routes through `/v1/responses`,
which does not accept `scale`; offering it in the UI silently
dropped the value at the backend and misled users into thinking
their selection was applied. Backend Literal still accepts it on
input so stale clients are not 422'd, and `_stream_openai_responses`
continues to drop it from the wire body.
- Dedupe + drop empty entries for OpenAI Chat `stop` and Anthropic
`stop_sequences` before forwarding so whitespace chips or accidental
repeats do not waste the 4-entry OpenAI cap or the 16-entry
Anthropic cap. Anthropic over-cap now logs and truncates, matching
the OpenAI path.
- Static `aria-label="Parallel tool calls"` on the Switch; screen
readers already announce checked / unchecked state, so the dynamic
Enable/Disable label was redundant.
- Forward an `aria-label` onto the inner Input inside
`StopSequencesInput` so screen-reader users can identify the field.
- Regression tests covering the new dedup, truncation, and the
preserved silent-drop of `scale` on Responses.
Adds the missing sampling parameters that the upstream APIs accept and
that Studio's chat UI previously hid. Each knob is gated per provider
so the picker never offers a field the upstream would 400 on, and the
per-provider stream functions translate / drop fields to match each
API's naming.
New `InferenceParams` fields (round-trip through PersistedInferenceParams
and the chat-settings server store automatically):
- frequencyPenalty (-2..2): OpenAI Chat Completions only.
- seed (int | null): OpenAI Chat + OpenAI-compat local backends.
- stop (string[]): all OpenAI Chat + Anthropic Messages. Backend
truncates to 4 entries on OpenAI Chat per docs and renames to
`stop_sequences` on Anthropic.
- serviceTier (auto|default|flex|priority|scale|standard_only):
per-provider enum sets resolved by getServiceTierOptions.
- parallelToolCalls (bool, default true): forwarded as
`parallel_tool_calls` on both OpenAI APIs and inverted into
`disable_parallel_tool_use` on Anthropic.
OpenAI Responses (gpt-5.x / o3) explicitly drops frequencyPenalty /
seed / stop alongside the existing temperature / top_p drop, since
the upstream 400s on all of them. service_tier on Responses accepts a
subset (no `scale`) which the dispatch already enforces.
UI rows land in the existing Sampling section of the chat settings
sheet using ParamSlider (frequency penalty), a numeric Input (seed),
a new chips editor `StopSequencesInput` (stop), Select (service tier),
and Switch (parallel tool calls). Each row's visibility follows the
new ProviderCapabilities flag.
Tests pin the gating contract: stop_sequences renamed on Anthropic,
4-entry truncation on OpenAI Chat, every Responses-rejected field
dropped, schema-level validation for the service_tier Literal and
frequency_penalty range.
Plan: plans/hashed-riding-porcupine.md