* studio: reserve VRAM headroom for the MTP draft cache in auto-fit When MTP is going to engage on this load, _fit_context_to_vram now budgets 0.85 of available VRAM instead of 0.90, leaving room for llama.cpp's secondary MTP draft KV cache + compute graph buffers. Motivation: a user report on RTX 5090 (32 GB) showed Qwen3.6-27B-MTP-GGUF UD-Q4_K_XL at native auto-context running roughly half the speed of the same model with a slightly smaller context. The most parsimonious explanation is a VRAM cliff: at native context the target's KV already eats the 90% budget, then llama-server allocates the draft cache + draft graph on top and spills into a slower partial-offload path. Reducing the budget by 5% on MTP loads avoids the spill without penalising non-MTP loads. On hardware with abundant VRAM (B200, etc.) the fit is unchanged because the requested context already fits in the tighter budget too. MTP detection mirrors the auto-promotion logic in load_model: the GGUF advertises nextn_predict_layers, or the model identifier / local path matches the -MTP marker, and the user has not explicitly opted out via speculative_type="off" or --spec-type extra args. Tests: two new cases in test_kv_cache_estimation.py verify that mtp_engaged=True yields a context less-than-or-equal-to the non-MTP path on a tight budget, and that kv_on_gpu=False still short-circuits regardless of mtp_engaged. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * studio: gate _mtp_will_engage on canonical-mode resolver After PR #5582 introduced the 5-mode Speculative Decoding dropdown plus _canonicalize_spec_mode, the auto-fit MTP-engaged predicate becomes: * forced mtp / mtp+ngram -> always engage MTP (extra VRAM needed) * auto + MTP GGUF (>= 3B) -> engages MTP via auto-promotion * auto + MTP GGUF (sub-3B) -> falls back to ngram-mod (no extra VRAM) * ngram / ngram-simple / off -> never engage MTP * user --spec-type in extra_args -> resolver suppressed; no headroom The old gate triggered on "anything but off", so it over-reserved the 0.85 budget when the user explicitly picked Ngram (no MTP) or when Auto fell back to ngram-mod on a sub-3B MTP model. The 5% headroom cost was minor but unnecessary. Mirrors the same logic already encoded in _build_speculative_flags so the auto-fit budget and the actual emission agree on whether MTP is running. All 361 backend tests pass. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| src-tauri | ||
| __init__.py | ||
| install_llama_prebuilt.py | ||
| install_python_stack.py | ||
| LICENSE.AGPL-3.0 | ||
| package-lock.json | ||
| package.json | ||
| setup.bat | ||
| setup.ps1 | ||
| setup.sh | ||
| Unsloth_Studio_Colab.ipynb | ||