Compare commits

...
Sign in to create a new pull request.

2 commits

Author SHA1 Message Date
pre-commit-ci[bot]
6b19a8984a [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-06-27 09:04:58 +00:00
Daniel Han
e6ab7f6952 Studio: size separate-drafter MTP draft KV at full context (swa_full)
For a separate drafter (Gemma mtp-*.gguf) the draft caches against the full
target window, so its sliding-window layers are not window-capped at runtime
and the draft KV grows with n_ctx. _estimate_kv_cache_bytes defaulted to the
SWA window cap, which sized the drafter at a near-flat ~16 MiB and under-
reserved on long contexts.

Pass swa_full = True so the draft KV is sized as a safe upper bound. Measured
draft overhead on a 12b drafter is 580 MiB at 32768, 868 MiB at 131072 and
1252 MiB at 262144, all under-reserved before. Embedded MTP heads (Qwen) are
unaffected.

All fit, KV, MTP and context tests pass.
2026-06-27 09:04:06 +00:00

View file

@ -2907,7 +2907,9 @@ class LlamaCppBackend:
# The drafter is served under the same --parallel slot count as the
# main model, so price its KV per slot too: a sliding-window drafter
# (Gemma) grows KV with slots and would otherwise be under-reserved.
kv = db._estimate_kv_cache_bytes(n_ctx, heavier, n_parallel = n_parallel)
# swa_full: the draft caches against the full target window, so its SWA
# layers are not window-capped and grow with n_ctx (safe upper bound).
kv = db._estimate_kv_cache_bytes(n_ctx, heavier, n_parallel = n_parallel, swa_full = True)
return kv or None
nextn = self._nextn_predict_layers or 0
n_kv = self._n_kv_heads or self._n_heads