unsloth/studio/backend/models
Daniel Han a02320aa7b Add typical_p sampler (local) + DeepSeek-reasoner per-model gating (PR #5711)
Two follow-ups from a closer reading of each provider's published
sampling surface against llama.cpp's own server README.

typical_p (locally typical sampling, `typ_p` in the llama.cpp sampler
chain):
  - New ProviderCapabilities.typicalP flag; defaults false on every
    SaaS provider (none accept the field) and true only on the
    permissive local buckets (custom, vllm, ollama, llama_cpp,
    openrouter via ALL_SUPPORTED). InferenceParams.typicalP is
    nullable number (null = unset; 1.0 = llama-server default, also
    treated as no-op when forwarding).
  - Backend: new ChatCompletionRequest.typical_p Field (0.0..1.0).
    Threaded through all three llama_cpp.py payload builders
    (chat-completion, agentic tool-loop, final-pass) so the field
    survives the local tool-loop too. _build_passthrough_payload in
    routes/inference.py picks it up and only writes the body when
    the caller set a value; left absent it falls back to llama-server
    default. Three route call sites (generate_chat_completion,
    generate_chat_completion_with_tools, _build_passthrough_payload)
    forward payload.typical_p.
  - Frontend: chat-adapter forwards on both branches (external opt-in
    only when capability allows + value != 1; local forwards
    unconditionally when set and != 1). OpenAIChatCompletionsRequest
    grows a `typical_p?` field. Persisted via chat-settings-storage
    mirroring the seed nullable-float handler.
  - Test: pin _build_passthrough_payload's forward + absent behavior.

DeepSeek per-model gating:
  - deepseek-reasoner / deepseek-r1 silently ignore temperature, top_p,
    presence_penalty, frequency_penalty per
    https://api-docs.deepseek.com/guides/reasoning_model — mirror the
    OpenAI / Claude 4.7 per-model approach: getProviderCapabilities
    downshifts these ids to a stripped capability set so the panel
    does not offer knobs the upstream silently drops.

161+1 sampling-routing tests pass; frontend tsc clean.

Refs:
  - llama.cpp server params: https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
  - DeepSeek reasoner restrictions: https://api-docs.deepseek.com/guides/reasoning_model
2026-05-26 14:12:40 +00:00
..
.gitkeep fix: restore models directory files deleted during restructure 2026-02-02 19:36:30 +00:00
__init__.py Studio: Dark theme refactor, right sidebar redesign, and chat UI polish (#5150) 2026-05-07 14:33:31 +04:00
auth.py studio: security and hardening pass (auth rate-limit, sandbox, path containment, schema validation, headers) (#5375) 2026-05-13 06:12:18 -07:00
data_recipe.py feat(studio): multi-file unstructured seed upload with better backend extraction (#4468) 2026-03-20 13:22:42 -07:00
datasets.py Final cleanup 2026-03-12 18:28:04 +00:00
export.py studio: security and hardening pass (auth rate-limit, sandbox, path containment, schema validation, headers) (#5375) 2026-05-13 06:12:18 -07:00
inference.py Add typical_p sampler (local) + DeepSeek-reasoner per-model gating (PR #5711) 2026-05-26 14:12:40 +00:00
models.py Studio: add folder browser modal for Custom Folders (#5035) 2026-04-15 08:04:33 -07:00
providers.py Studio: make API key optional for local providers (llama.cpp/vLLM/Ollama) (#5457) 2026-05-15 23:33:22 +04:00
responses.py Final cleanup 2026-03-12 18:28:04 +00:00
training.py studio: drop unused max_grad_value schema + route plumbing (#5424) 2026-05-14 05:43:58 -07:00
users.py fix: remove old comments (#4292) 2026-03-14 16:50:13 +04:00