Every passthrough request through `/v1/chat/completions` and the streaming `/v1/responses` path calls `load_inference_config(model)` to fold family-default `chat_template_kwargs` (e.g. gpt-oss `reasoning_effort=medium`) into the outbound body. The function walks the model_defaults directory recursively via `rglob` inside `_has_specific_yaml` plus reads two YAML files on every call, so the hot path was paying the full lookup cost for every token-budget poll, every tool turn, every reasoning sub-step. The defaults directory is shipped with the package and does not mutate at runtime, so wrap both `_has_specific_yaml` and the expensive inner work of `load_inference_config` in `lru_cache`. The public entry point still returns a fresh deepcopy of the cached snapshot so the existing `test_load_returns_a_fresh_dict_per_call` contract (callers may safely mutate the dict) is preserved. Adds a regression test pinning the cache-hit count over repeated calls for the same identifier. |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| src-tauri | ||
| __init__.py | ||
| install_llama_prebuilt.py | ||
| install_python_stack.py | ||
| LICENSE.AGPL-3.0 | ||
| package-lock.json | ||
| package.json | ||
| setup.bat | ||
| setup.ps1 | ||
| setup.sh | ||
| Unsloth_Studio_Colab.ipynb | ||