Every passthrough request through `/v1/chat/completions` and the streaming `/v1/responses` path calls `load_inference_config(model)` to fold family-default `chat_template_kwargs` (e.g. gpt-oss `reasoning_effort=medium`) into the outbound body. The function walks the model_defaults directory recursively via `rglob` inside `_has_specific_yaml` plus reads two YAML files on every call, so the hot path was paying the full lookup cost for every token-budget poll, every tool turn, every reasoning sub-step. The defaults directory is shipped with the package and does not mutate at runtime, so wrap both `_has_specific_yaml` and the expensive inner work of `load_inference_config` in `lru_cache`. The public entry point still returns a fresh deepcopy of the cached snapshot so the existing `test_load_returns_a_fresh_dict_per_call` contract (callers may safely mutate the dict) is preserved. Adds a regression test pinning the cache-hit count over repeated calls for the same identifier. |
||
|---|---|---|
| .. | ||
| assets | ||
| auth | ||
| core | ||
| loggers | ||
| models | ||
| plugins | ||
| requirements | ||
| routes | ||
| state | ||
| storage | ||
| tests | ||
| utils | ||
| __init__.py | ||
| _platform_compat.py | ||
| colab.py | ||
| main.py | ||
| run.py | ||
| startup_banner.py | ||