Seed each request with the model's recommended sampling (matching the Chat UI), add per-field override flags, ignore oversized overrides, warn when sampling pins cannot apply to a reused server, and apply pins to the completions endpoint. |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| inference_config.py | ||