* fix(studio): Windows GGUF cancel hang + CPU spinlock overhead (#5692)
Two fixes for Windows-native GGUF inference via llama-server:
**Issue 1 — GPU/CUDA Hang on Stream Cancellation:**
- Add `Connection: close` header to all httpx requests proxying to
llama-server, preventing Keep-Alive from masking downstream socket
closure.
- Introduce `_await_disconnect_then_close` background watcher that
polls `request.is_disconnected()` every 100ms and calls
`resp.aclose()` immediately when the client disconnects. This runs
alongside the existing cancel-POST watcher and covers client aborts
that never reach the /cancel endpoint (tab close, proxy aborts,
Colab, mobile navigation, etc.).
- Change all StreamingResponse `Connection: keep-alive` headers to
`Connection: close`.
**Issue 2 — High CPU Spinlock & KV Cache Backup Overhead:**
- Set OMP_WAIT_POLICY=PASSIVE and OMP_NUM_THREADS=2 in the
llama-server subprocess environment on Windows to prevent OpenMP
from spin-waiting on all logical cores while the GPU decodes.
- Limit `--threads` to 2 on Windows when the model is fully
GPU-offloaded (`-ngl -1`). Auto-detect otherwise.
- Pass `--cache-ram 0 --ctx-checkpoints 0 --no-cache-prompt
--checkpoint-every-n-tokens -1` on Windows to disable prompt-cache
snapshots that copy KV cache to system RAM over the WDDM/PCI-E bus.
Closes#5692.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix: use local import to avoid ruff F823 (sys used before assignment)
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* review: address gemini review feedback
- Simplify _fully_gpu_offloaded init: default to False, only set True
in the gpu_indices branch, drop redundant else.
- Log exceptions in _await_disconnect_then_close at debug level instead
of silent pass, per review suggestion.
* Adjust review feedback for PR #5749
- _await_disconnect_then_close: set cancel_event before resp.aclose() so
the streamer's RemoteProtocolError handler treats the watcher-driven
close as cancellation, not an upstream error. Both call sites pass
cancel_event through.
- Windows --cache-ram / --no-cache-prompt / --ctx-checkpoints block: gate
on _fully_gpu_offloaded so CPU and partial-offload Windows runs keep
prompt-cache reuse across turns.
- Windows OMP_WAIT_POLICY / OMP_NUM_THREADS env: same gate so CPU and
partial-offload Windows runs keep default OpenMP parallelism.
* Shorten code comments touched by PR #5749
* Clean up local imports and rename underscore locals in PR #5749
- Drop the function-local `import sys as _sys` introduced as an F823
workaround; remove the redundant in-function `import os`/`import sys`
block so module-level imports resolve sys/os instead. F823 no longer
triggers because no shadowing import remains inside load_model.
- Rename `_fully_gpu_offloaded` and `_t` to `fully_gpu_offloaded` and
`threads_arg`. Underscore-prefixed names usually mean private/module-
level; plain locals match Python style for in-function temporaries.
No behavior change. ruff clean, py_compile clean, 35 studio cancel-
infra tests + 13 launch-gating AST locks + 6 disconnect-watcher locks
+ 4 spoof live-import tests all pass.
* Fix Windows GGUF follow-ups for PR #5749
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Fix cache flag gating for PR #5749
* Fix Python 3.9 annotations for PR #5749
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: Anmol Mishra <anmolx.work@gmail.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Co-authored-by: wasimysaid <wasimysdev@gmail.com>
Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com>