Tighten _can_estimate_kv gate and treat sliding_window=0 as disabled
Two additional fixes from review round 1 (5/8 and 4/8 reviewer consensus): 1. _can_estimate_kv now requires BOTH key_length AND value_length for the explicit-dims path. Previously key_length alone was enough, which could cause silent fallthrough to the legacy formula with fabricated defaults (n_kv=1, head_dim=128) when value_length was absent from the GGUF. 2. SWA path now requires sliding_window > 0. Some GGUFs use 0 as a disabled sentinel. Without this guard, min(ctx, 0) would zero out all SWA layer contributions, severely underestimating KV cache.
This commit is contained in:
parent
ae6fb93b6f
commit
434dee9618
1 changed files with 5 additions and 4 deletions
|
|
@ -358,12 +358,12 @@ class LlamaCppBackend:
|
|||
"""True if we have enough GGUF metadata to estimate KV cache size."""
|
||||
if self._n_layers is None:
|
||||
return False
|
||||
# New-style: explicit key/value dimensions from GGUF
|
||||
if self._kv_key_length is not None:
|
||||
return True
|
||||
# MLA: kv_lora_rank is sufficient
|
||||
# MLA: kv_lora_rank is sufficient (K-only cache)
|
||||
if self._kv_lora_rank is not None:
|
||||
return True
|
||||
# New-style: need both explicit key AND value dimensions
|
||||
if self._kv_key_length is not None and self._kv_value_length is not None:
|
||||
return True
|
||||
# Legacy: need embedding_length + head count
|
||||
return self._embedding_length is not None and (
|
||||
self._n_kv_heads is not None or self._n_heads is not None
|
||||
|
|
@ -434,6 +434,7 @@ class LlamaCppBackend:
|
|||
# which is still far more accurate than the legacy formula (which ignores SWA).
|
||||
if (
|
||||
self._sliding_window is not None
|
||||
and self._sliding_window > 0
|
||||
and key_len is not None
|
||||
and val_len is not None
|
||||
):
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue