unsloth/studio/backend/core/inference
Daniel Han 91d04741ff Studio: expose Anthropic / OpenAI sampling knobs per provider
Adds the missing sampling parameters that the upstream APIs accept and
that Studio's chat UI previously hid. Each knob is gated per provider
so the picker never offers a field the upstream would 400 on, and the
per-provider stream functions translate / drop fields to match each
API's naming.

New `InferenceParams` fields (round-trip through PersistedInferenceParams
and the chat-settings server store automatically):

- frequencyPenalty (-2..2): OpenAI Chat Completions only.
- seed (int | null): OpenAI Chat + OpenAI-compat local backends.
- stop (string[]): all OpenAI Chat + Anthropic Messages. Backend
  truncates to 4 entries on OpenAI Chat per docs and renames to
  `stop_sequences` on Anthropic.
- serviceTier (auto|default|flex|priority|scale|standard_only):
  per-provider enum sets resolved by getServiceTierOptions.
- parallelToolCalls (bool, default true): forwarded as
  `parallel_tool_calls` on both OpenAI APIs and inverted into
  `disable_parallel_tool_use` on Anthropic.

OpenAI Responses (gpt-5.x / o3) explicitly drops frequencyPenalty /
seed / stop alongside the existing temperature / top_p drop, since
the upstream 400s on all of them. service_tier on Responses accepts a
subset (no `scale`) which the dispatch already enforces.

UI rows land in the existing Sampling section of the chat settings
sheet using ParamSlider (frequency penalty), a numeric Input (seed),
a new chips editor `StopSequencesInput` (stop), Select (service tier),
and Switch (parallel tool calls). Each row's visibility follows the
new ProviderCapabilities flag.

Tests pin the gating contract: stop_sequences renamed on Anthropic,
4-entry truncation on OpenAI Chat, every Responses-rejected field
dropped, schema-level validation for the service_tier Literal and
frequency_penalty range.

Plan: plans/hashed-riding-porcupine.md
2026-05-23 15:33:13 +00:00
..
__init__.py Final cleanup 2026-03-12 18:28:04 +00:00
_html_to_md.py fix: studio web search SSL failures and empty page content (#4754) 2026-04-01 06:12:02 -07:00
anthropic_compat.py Studio: Claude Code Anthropic API tool compatibility (#5390) 2026-05-21 16:45:05 +04:00
audio_codecs.py Add native GGUF intake to Studio (#5246) 2026-05-04 11:46:18 +02:00
chat_template_helpers.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00
defaults.py studio: engage draft-mtp on vision MTP GGUFs (drop incorrect vision gate) (#5560) 2026-05-18 08:42:55 -07:00
external_provider.py Studio: expose Anthropic / OpenAI sampling knobs per provider 2026-05-23 15:33:13 +00:00
inference.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00
key_exchange.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
llama_cpp.py studio: settle GPU VRAM after killing llama-server before the next reload (#5693) 2026-05-22 05:50:39 -07:00
llama_server_args.py studio: add --spec-draft-n-max toggle for MTP speculative decoding (#5582) 2026-05-19 06:17:04 -07:00
mlx_inference.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00
orchestrator.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00
pricing.py Studio: per-session cost calculator + /api/providers/pricing endpoint (#5690) 2026-05-22 06:03:43 -07:00
providers.py Studio: expand Connections model picker for local inference server (#5643) 2026-05-20 15:06:06 +04:00
safetensors_agentic.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00
tool_call_parser.py Revert "studio: tool calling for Llama-3, Mistral, Gemma 4 on safetensors + MLX (#5615)" (#5619) 2026-05-19 07:26:39 -07:00
tools.py studio: tighten sandbox blocklist precision (bash, hf upload, NOFILE) (#5487) 2026-05-18 00:01:17 -07:00
worker.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00