Two follow-ups from a closer reading of each provider's published
sampling surface against llama.cpp's own server README.
typical_p (locally typical sampling, `typ_p` in the llama.cpp sampler
chain):
- New ProviderCapabilities.typicalP flag; defaults false on every
SaaS provider (none accept the field) and true only on the
permissive local buckets (custom, vllm, ollama, llama_cpp,
openrouter via ALL_SUPPORTED). InferenceParams.typicalP is
nullable number (null = unset; 1.0 = llama-server default, also
treated as no-op when forwarding).
- Backend: new ChatCompletionRequest.typical_p Field (0.0..1.0).
Threaded through all three llama_cpp.py payload builders
(chat-completion, agentic tool-loop, final-pass) so the field
survives the local tool-loop too. _build_passthrough_payload in
routes/inference.py picks it up and only writes the body when
the caller set a value; left absent it falls back to llama-server
default. Three route call sites (generate_chat_completion,
generate_chat_completion_with_tools, _build_passthrough_payload)
forward payload.typical_p.
- Frontend: chat-adapter forwards on both branches (external opt-in
only when capability allows + value != 1; local forwards
unconditionally when set and != 1). OpenAIChatCompletionsRequest
grows a `typical_p?` field. Persisted via chat-settings-storage
mirroring the seed nullable-float handler.
- Test: pin _build_passthrough_payload's forward + absent behavior.
DeepSeek per-model gating:
- deepseek-reasoner / deepseek-r1 silently ignore temperature, top_p,
presence_penalty, frequency_penalty per
https://api-docs.deepseek.com/guides/reasoning_model — mirror the
OpenAI / Claude 4.7 per-model approach: getProviderCapabilities
downshifts these ids to a stripped capability set so the panel
does not offer knobs the upstream silently drops.
161+1 sampling-routing tests pass; frontend tsc clean.
Refs:
- llama.cpp server params: https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
- DeepSeek reasoner restrictions: https://api-docs.deepseek.com/guides/reasoning_model