unsloth/tests/studio
Anmol Mishra 9f694ab750
fix(studio): Windows GGUF cancel hang + CPU spinlock overhead (#5692) (#5749)
* fix(studio): Windows GGUF cancel hang + CPU spinlock overhead (#5692)

Two fixes for Windows-native GGUF inference via llama-server:

**Issue 1 — GPU/CUDA Hang on Stream Cancellation:**
- Add `Connection: close` header to all httpx requests proxying to
  llama-server, preventing Keep-Alive from masking downstream socket
  closure.
- Introduce `_await_disconnect_then_close` background watcher that
  polls `request.is_disconnected()` every 100ms and calls
  `resp.aclose()` immediately when the client disconnects. This runs
  alongside the existing cancel-POST watcher and covers client aborts
  that never reach the /cancel endpoint (tab close, proxy aborts,
  Colab, mobile navigation, etc.).
- Change all StreamingResponse `Connection: keep-alive` headers to
  `Connection: close`.

**Issue 2 — High CPU Spinlock & KV Cache Backup Overhead:**
- Set OMP_WAIT_POLICY=PASSIVE and OMP_NUM_THREADS=2 in the
  llama-server subprocess environment on Windows to prevent OpenMP
  from spin-waiting on all logical cores while the GPU decodes.
- Limit `--threads` to 2 on Windows when the model is fully
  GPU-offloaded (`-ngl -1`). Auto-detect otherwise.
- Pass `--cache-ram 0 --ctx-checkpoints 0 --no-cache-prompt
  --checkpoint-every-n-tokens -1` on Windows to disable prompt-cache
  snapshots that copy KV cache to system RAM over the WDDM/PCI-E bus.

Closes #5692.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: use local import to avoid ruff F823 (sys used before assignment)

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* review: address gemini review feedback

- Simplify _fully_gpu_offloaded init: default to False, only set True
  in the gpu_indices branch, drop redundant else.
- Log exceptions in _await_disconnect_then_close at debug level instead
  of silent pass, per review suggestion.

* Adjust review feedback for PR #5749

- _await_disconnect_then_close: set cancel_event before resp.aclose() so
  the streamer's RemoteProtocolError handler treats the watcher-driven
  close as cancellation, not an upstream error. Both call sites pass
  cancel_event through.
- Windows --cache-ram / --no-cache-prompt / --ctx-checkpoints block: gate
  on _fully_gpu_offloaded so CPU and partial-offload Windows runs keep
  prompt-cache reuse across turns.
- Windows OMP_WAIT_POLICY / OMP_NUM_THREADS env: same gate so CPU and
  partial-offload Windows runs keep default OpenMP parallelism.

* Shorten code comments touched by PR #5749

* Clean up local imports and rename underscore locals in PR #5749

- Drop the function-local `import sys as _sys` introduced as an F823
  workaround; remove the redundant in-function `import os`/`import sys`
  block so module-level imports resolve sys/os instead. F823 no longer
  triggers because no shadowing import remains inside load_model.
- Rename `_fully_gpu_offloaded` and `_t` to `fully_gpu_offloaded` and
  `threads_arg`. Underscore-prefixed names usually mean private/module-
  level; plain locals match Python style for in-function temporaries.

No behavior change. ruff clean, py_compile clean, 35 studio cancel-
infra tests + 13 launch-gating AST locks + 6 disconnect-watcher locks
+ 4 spoof live-import tests all pass.

* Fix Windows GGUF follow-ups for PR #5749

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix cache flag gating for PR #5749

* Fix Python 3.9 annotations for PR #5749

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: Anmol Mishra <anmolx.work@gmail.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Co-authored-by: wasimysaid <wasimysdev@gmail.com>
Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com>
2026-06-15 10:32:10 +01:00
..
install Studio: don't silently fall back to a CPU prebuilt on NVIDIA Linux GPU hosts (#6310) 2026-06-14 02:06:31 -03:00
load_freeze Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
_playwright_robust.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
playwright_chat_ime_i18n.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
playwright_chat_ui.py Tests: follow Compare chat into the More submenu in the Playwright chat UI driver (#6153) 2026-06-10 08:05:24 -07:00
playwright_extra_ui.py Tests: follow Compare chat into the More submenu in the extra UI driver (#6177) 2026-06-10 21:51:16 -07:00
run_real_mlx_smoke.py MLX Training updates (#5656) 2026-06-14 04:58:50 -07:00
studio_api_smoke.py Formatting: ruff line-length 100, kwarg-spacing passes, drop blank after short local imports (#6079) 2026-06-08 04:24:13 -07:00
test_auth_form_input_count.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
test_cancel_atomicity.py Formatting: ruff line-length 100, kwarg-spacing passes, drop blank after short local imports (#6079) 2026-06-08 04:24:13 -07:00
test_cancel_id_wiring.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
test_chat_preset_builtin_invariants.py Formatting: ruff line-length 100, kwarg-spacing passes, drop blank after short local imports (#6079) 2026-06-08 04:24:13 -07:00
test_cli_repo_variant.py Formatting: ruff line-length 100, kwarg-spacing passes, drop blank after short local imports (#6079) 2026-06-08 04:24:13 -07:00
test_cli_run_alias.py Formatting: ruff line-length 100, kwarg-spacing passes, drop blank after short local imports (#6079) 2026-06-08 04:24:13 -07:00
test_cli_studio_defaults.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
test_cli_studio_stop_windows.py Windows/WSL installer: fix winget msstore cert failure, amd-smi DiskPart prompt, and enable AMD GPU (Strix Halo gfx1151) (#5940) 2026-06-10 04:24:49 -07:00
test_composer_rtl_bidi_attribute.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
test_export_output_path_contract.py Formatting: ruff line-length 100, kwarg-spacing passes, drop blank after short local imports (#6079) 2026-06-08 04:24:13 -07:00
test_frontend_dep_removal.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
test_hardware_dispatch_matrix.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
test_is_mlx_dispatch_gate.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
test_llama_cpp_wall_clock_cap.py Studio: extend llama.cpp first-token timeout (#5841) 2026-06-12 18:41:38 +02:00
test_mlx_training_worker_behaviors.py MLX training support for Studio on Apple Silicon (#5340) 2026-05-14 05:24:20 -07:00
test_resolve_cuda_toolkit.ps1 fix: clearer Studio setup error when GPU driver is too old for the installed CUDA toolkit (#5993) 2026-06-10 02:17:21 -07:00
test_stream_cancel_registration_timing.py fix(studio): Windows GGUF cancel hang + CPU spinlock overhead (#5692) (#5749) 2026-06-15 10:32:10 +01:00
test_studio_gguf_export_script_pin.py Formatting: ruff line-length 100, kwarg-spacing passes, drop blank after short local imports (#6079) 2026-06-08 04:24:13 -07:00
test_studio_text_descender_clipping.py Fix stale sidebar regression test to match the gap-px markup (#6232) 2026-06-12 00:53:13 -07:00
test_sync_allow_scripts_pins.py Studio: auto-sync allowScripts pins after dependency bumps (#6136) 2026-06-10 02:35:37 -07:00