unsloth/studio
Daniel Han 79d9bf0c9a
Fix GGUF GPU fit check to account for KV cache VRAM (#4623)
* fix: account for KV cache in GGUF GPU fit check and auto-cap context length

The GPU fit check only compared GGUF file size against free VRAM,
ignoring KV cache memory. Models with large native context lengths
(e.g. Qwen3.5-9B at 262k) would pass the fit check since the GGUF
is only 5.6 GB, but the KV cache at 262k context needs ~40 GB at
f16. This caused llama-server to silently fall back to CPU inference.

Changes:
- Parse block_count, head_count_kv, head_count, and embedding_length
  from GGUF metadata alongside context_length
- Add KV cache VRAM estimation based on architecture params and the
  selected cache quantization type (f16, q8_0, q4_0, etc.)
- Auto-reduce context length to the maximum that fits in available
  GPU VRAM when the native context would exceed it
- Include estimated KV cache size in the _select_gpus total so the
  fit decision reflects actual runtime memory, not just file size

For the reported scenario (Qwen3.5-9B on RTX 3090 with 22415 MiB
free), context is auto-reduced from 262144 to ~63k with f16 KV cache,
keeping the model fully on GPU. With q4_0 KV cache quantization the
context can reach ~226k.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: resolve 6 bugs in KV cache VRAM estimation and add test harness

- Fix q8_0 BPE constant: 1.125 -> 34/32 (1.0625) to match llama.cpp block size
- Fix _fit_context_to_vram returning min_ctx when weights exceed budget
  (should return requested_ctx unchanged, let --fit handle it)
- Fix binary search inflating below-2048 requests (lo=min_ctx=2048 > hi)
- Fix n_ctx=0 regressing to 4096 when metadata unavailable (preserve sentinel)
- Fix multi-GPU auto-cap using single-GPU budget instead of aggregate
- Fix _context_length being overwritten with capped effective value

Add tests/test_gguf_kv_vram.py: 43 cross-platform pytest tests covering
pure logic, integration (monkeypatched load_model), and real GGUF parsing.
Runs in an isolated uv venv with only pytest -- no GPU/torch/structlog needed.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: complete _effective_context_length lifecycle

- Initialize _effective_context_length in __init__ (prevents AttributeError)
- Reset _effective_context_length in unload_model (prevents stale values)
- Update context_length property to return effective (capped) value for
  the UI/API, falling back to native _context_length if not set

* fix: multi-GPU selection tries smallest subset first

The previous approach summed all GPUs' memory to cap context, then
selected GPUs afterward. This was overly optimistic for heterogeneous
setups (e.g., 48 GiB + 4 GiB): the context was inflated by the tiny
GPU's contribution, then both GPUs were dragged in.

Now we try GPU subsets from smallest (1 GPU) to largest, capping
context for each. We pick the smallest subset where the model+KV
fits. This prefers single-GPU when possible (simpler, no tensor
split overhead) and avoids pulling in GPUs that barely help.

Add tests: test_multi_gpu_prefers_fewer_gpus,
test_multi_gpu_heterogeneous.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: prefer fewer GPUs over higher context in GPU selection

Multi-GPU inference is slower due to tensor-split overhead, so we
should prefer fewer GPUs with reduced context over more GPUs with
full context. Now the loop stops at the first GPU subset where the
model fits, rather than continuing to find subsets that allow higher
context. Only if the model can't fit on N GPUs do we try N+1.

This preserves the original behavior: use multi-GPU only when the
model doesn't fit on a single GPU.

* fix: make _kill_orphaned_servers cross-platform via psutil

Replace pgrep + os.kill(SIGKILL) with psutil.process_iter() and
proc.kill(), which work on Linux, macOS, and Windows. Build an
allowlist of install roots matching _find_llama_server_binary so
only studio-managed servers are killed.

* fix: skip KV estimation loop when effective context is unknown

When n_ctx=0 and GGUF metadata lacks context_length, effective_ctx
stays 0. _estimate_kv_cache_bytes(0) returns 0, so a GPU could be
selected with no KV headroom. Guard the loop with effective_ctx > 0
to fall back to file-size-only GPU selection in this case.

* chore: temporarily remove test harness (will add back separately)

* refactor: deduplicate UINT32/UINT64 handling in GGUF parser

Replace duplicated if/elif chains for vtype 4 and 10 with a single
block using setattr. No behavioral change.

* fix: honor explicit n_ctx by using multi-GPU before capping

When the user explicitly sets n_ctx, try to fit the full requested
context using _select_gpus (which adds GPUs as needed). Only cap
context if it doesn't fit on any GPU combination.

When n_ctx=0 (auto/native context), keep the existing behavior:
prefer fewer GPUs with reduced context, since multi-GPU is slower
and the user didn't ask for a specific context length.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: context_length property returns native value for frontend slider

The frontend uses context_length as the slider max. Returning the
capped effective value prevented users from requesting higher context
on reload (e.g., after switching to q4_0 KV cache). Revert to
returning the native GGUF metadata value -- the backend auto-caps
at load time regardless.

* revert: context_length returns effective (capped) value

The UI slider should show what the server is actually running at,
not the theoretical maximum. Revert to returning the effective
context length.

* fix: raise minimum context floor from 2048 to 4096

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-03-27 03:14:42 -07:00
..
backend Fix GGUF GPU fit check to account for KV cache VRAM (#4623) 2026-03-27 03:14:42 -07:00
frontend studio: fix chat CPU spike (#4632) 2026-03-27 06:20:26 +00:00
__init__.py Final cleanup 2026-03-12 18:28:04 +00:00
install_llama_prebuilt.py fix(studio): source-build fallback prefers Unsloth's tested tag over upstream latest (#4593) 2026-03-25 07:25:47 -07:00
install_python_stack.py studio: setup log styling (#4494) 2026-03-27 03:12:48 -07:00
LICENSE.AGPL-3.0 Add AGPL-3.0 license to studio folder 2026-03-09 19:36:25 +00:00
setup.bat Final cleanup 2026-03-12 18:28:04 +00:00
setup.ps1 studio: setup log styling (#4494) 2026-03-27 03:12:48 -07:00
setup.sh studio: setup log styling (#4494) 2026-03-27 03:12:48 -07:00
Unsloth_Studio_Colab.ipynb Allow install_python_stack to run on Colab (#4633) 2026-03-27 00:29:27 +04:00