unsloth/studio/backend/tests
Daniel Han 379f5a5aa6
Studio: add torch's pip nvidia DLL dirs to PATH on Windows (#5324)
* Studio: add torch's pip nvidia DLL dirs to PATH on Windows

Studio's install_python_stack bundles torch with matching CUDA
wheels (nvidia-cuda-runtime-cu13, nvidia-cublas-cu13, etc.) which
ship cudart64_X.dll, cublas64_X.dll, and cublasLt64_X.dll under
the prefix's Lib/site-packages/nvidia/<pkg>/(bin|Library/bin)/
tree. The Linux runtime env block in start_llama_server already
pulls the equivalent nvidia/cu*/lib paths into LD_LIBRARY_PATH,
but the Windows block did not do this, so the prebuilt
llama-server.exe could not resolve cudart64_X.dll at runtime
unless the user had a matching system CUDA toolkit on PATH. That
is the root cause of the Windows reports in
unslothai/unsloth#5106 ("GPU detected but model loaded entirely
on RAM/CPU"), and matches Roland's repeated workaround in that
issue: install matching CUDA toolkit version.

Brings the Windows env block in line with the Linux pattern:

* New LlamaCppBackend._windows_pip_nvidia_dll_dirs resolver
  globs <prefix>/Lib/site-packages/nvidia/<pkg>/bin and
  <prefix>/Lib/site-packages/nvidia/<pkg>/Library/bin. Both
  layouts are seen in the wild across cuda_runtime / cublas /
  cudnn / nvjitlink wheels.

* The Windows env block now extends path_dirs with the
  resolver's output before falling back to CUDA_PATH/bin, so
  pip-installed wheels are the canonical source (mirroring the
  Linux LD_LIBRARY_PATH ordering). System CUDA toolkit remains a
  valid fallback.

Tests: 7 new cases in
studio/backend/tests/test_llama_cpp_windows_nvidia_path.py:

* empty resolver when no nvidia wheels installed
* nvidia/<pkg>/bin layout resolved
* nvidia/<pkg>/Library/bin layout resolved
* mixed bin and Library/bin layouts both resolved
* unrelated site-packages contents not walked
* non-directory entries skipped
* missing prefix does not raise

110 backend tests pass. No regressions.

Refs #5106

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: also scan torch/lib in Windows pip nvidia DLL resolver

PyTorch's Windows CUDA wheels frequently bundle cudart64_X.dll and
cublas64_X.dll directly under Lib/site-packages/torch/lib/ instead of
shipping separate nvidia-cuda-runtime-cuXX / nvidia-cublas-cuXX wheels.
On those installs _windows_pip_nvidia_dll_dirs previously returned
nothing useful, and llama-server.exe fell back to needing a system CUDA
toolkit on PATH -- the original #5106 failure mode.

The install-side equivalent python_runtime_dirs in
install_llama_prebuilt.py already treats torch/lib as a Python runtime
DLL source for the same reason. Bring the runtime resolver in parity
so torch-bundled-CUDA installs find their cudart at llama-server start.

Updates the existing test that codified the bug (asserted torch/lib was
excluded), and adds three new cases: pickup, combined-with-nvidia, and
the must-be-a-directory guard.

* Studio: cover cu13 bin/x86_64 layout in Windows DLL resolver

Three follow-ups from a 12-reviewer batch over c1c8a074 (PR #5324):

1. The current nvidia-cuda-runtime (unsuffixed) 13.2.75 and
   nvidia-cublas 13.4.0.1 Windows wheels on PyPI ship under
   nvidia/cu13/bin/x86_64/cudart64_13.dll etc, not under
   nvidia/PKG/bin/. The previous resolver matched only one
   directory level past nvidia/PKG/ and silently missed the
   actual cu13 DLL location, leaving CUDA 13 users on the same
   failure mode as before #5106. Verified against:
       pip download nvidia-cuda-runtime --platform win_amd64
   which produces nvidia/cu13/bin/x86_64/cudart64_13.dll.

2. glob.glob over sys.prefix interprets [ and ] as a
   character class. Valid Windows usernames / install paths can
   contain those characters (for example C:\Users\alice[work]\studio),
   so the previous resolver silently returned an empty list for such
   prefixes even when DLL dirs were present.

3. The resolver only ever returned nvidia/PKG/bin -- if both
   bin and bin/x86_64 exist (current wheels do), Windows
   DLL search should land on the arch-specific subdir first so the
   explicit cudart64_X.dll location wins.

Rewritten as a pathlib.Path.iterdir walk to fix all three:
no glob escaping needed, arch-specific subdirs added explicitly,
and ordering puts bin/x86_64 before bin. Conda-style
Library/bin/x86_64 and Library/bin/x64 are also covered for
parity. A seen set dedupes when wheels happen to expose the
same directory through multiple layouts.

New tests:
 - test_picks_up_cu13_bin_x86_64_layout (the actual real-world cu13 case)
 - test_picks_up_bin_x64_layout
 - test_mixed_cu12_and_cu13_layouts
 - test_glob_meta_in_prefix_is_safe (bracket repro)
 - test_arch_subdir_listed_before_parent_bin (ordering)

Verified empirically against PyPI:
       nvidia-cuda-runtime 13.2.75 -> nvidia/cu13/bin/x86_64/cudart64_13.dll
       nvidia-cublas       13.4.0.1 -> nvidia/cu13/bin/x86_64/cublas64_13.dll
                                       nvidia/cu13/bin/x86_64/cublasLt64_13.dll
       nvidia-cudnn-cu13   9.22.0.52 -> nvidia/cudnn/bin/cudnn64_9.dll (already covered)

Refs #5106

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-11 05:42:09 -07:00
..
__init__.py Final cleanup 2026-03-12 18:28:04 +00:00
conftest.py Studio: Expose openai and anthropic compatible external API end points (#4956) 2026-04-13 21:08:11 +04:00
test_anthropic_messages.py Studio: support images on /v1/messages (Anthropic-compat) (#5128) 2026-04-22 03:25:07 +04:00
test_browse_folders_route.py Studio: add folder browser modal for Custom Folders (#5035) 2026-04-15 08:04:33 -07:00
test_cache_case_resolution.py Add tests for cache case resolution (from PR #4822) (#4823) 2026-04-03 13:58:26 -07:00
test_cached_gguf_routes.py Studio: support GGUF variant selection for non-suffixed repos (#5023) 2026-04-15 15:32:01 +04:00
test_data_recipe_github_progress.py Studio: add github_repo seed reader and GitHub Support Bot recipe (#5169) 2026-04-24 12:02:03 -07:00
test_data_recipe_seed.py fix(seed): disable remote code execution in seed inspect dataset loads (#4275) 2026-03-13 19:37:43 +04:00
test_desktop_auth.py Studio: fix 7 failing studio_unit_tests on main (#5216) 2026-04-28 22:43:44 -07:00
test_export_log_cursor.py studio: stream export worker output into the export dialog (#4897) 2026-04-14 08:55:43 -07:00
test_gpu_selection.py Update VRAM estimator to cater to broader model configs (#5175) 2026-05-05 04:12:36 -07:00
test_gpu_selection_sandbox.py [Studio] multi gpu finetuning/inference via "balanced_low0/sequential" device_map (#4602) 2026-03-30 02:33:15 -07:00
test_host_defaults.py Default Studio host to 127.0.0.1 and prompt before auto-start (#5267) 2026-05-04 13:03:16 +04:00
test_inference_model_validation.py Studio: Dark theme refactor, right sidebar redesign, and chat UI polish (#5150) 2026-05-07 14:33:31 +04:00
test_kv_cache_estimation.py fix KVCache estimates for gemma4 style sliding window models (#5225) 2026-05-05 04:06:46 -07:00
test_llama_cpp_cache_aware_disk_check.py Studio: make GGUF disk-space preflight cache-aware (#5012) 2026-04-14 08:53:37 -07:00
test_llama_cpp_context_fit.py fix KVCache estimates for gemma4 style sliding window models (#5225) 2026-05-05 04:06:46 -07:00
test_llama_cpp_load_progress.py Studio: live model-load progress + rate/ETA on download and load (#5017) 2026-04-14 09:46:22 -07:00
test_llama_cpp_load_progress_live.py Studio: split model-load progress label across two rows (#5020) 2026-04-14 10:58:16 -07:00
test_llama_cpp_load_progress_matrix.py Studio: split model-load progress label across two rows (#5020) 2026-04-14 10:58:16 -07:00
test_llama_cpp_max_context_threshold.py fix KVCache estimates for gemma4 style sliding window models (#5225) 2026-05-05 04:06:46 -07:00
test_llama_cpp_no_context_shift.py Studio: hard-stop at n_ctx with a 'Context limit reached' toast (#5021) 2026-04-14 10:58:20 -07:00
test_llama_cpp_windows_nvidia_path.py Studio: add torch's pip nvidia DLL dirs to PATH on Windows (#5324) 2026-05-11 05:42:09 -07:00
test_llama_server_args.py Studio: forward llama-server args from unsloth studio run , activate unsloth run , and allow passing model:quant to load models (#5271) 2026-05-04 17:08:04 +04:00
test_log_filter_no_truncation.py Studio: stop truncating long log lines as suspected base64 (#5335) 2026-05-08 13:07:18 +04:00
test_mlx_inference_backend.py feat(studio): MLX training tab on Apple Silicon (LoRA / full FT, VLM, export) (#5265) 2026-05-05 23:54:58 -07:00
test_mlx_training_worker_config.py feat(studio): MLX training tab on Apple Silicon (LoRA / full FT, VLM, export) (#5265) 2026-05-05 23:54:58 -07:00
test_models_get_model_config_case_resolution.py Add tests for cache case resolution (from PR #4822) (#4823) 2026-04-03 13:58:26 -07:00
test_native_context_length.py Studio: Fix chat template disappearing after browser refresh (#5209) 2026-05-01 08:19:09 -07:00
test_openai_tool_passthrough.py Studio: Fix image-only chat requests failing validation (#5212) 2026-04-28 14:49:13 -07:00
test_pytorch_mirror.py Add configurable PyTorch mirror via UNSLOTH_PYTORCH_MIRROR env var (#5024) 2026-04-15 11:39:11 +04:00
test_responses_api.py Studio: Expose openai and anthropic compatible external API end points (#4956) 2026-04-13 21:08:11 +04:00
test_responses_tool_passthrough.py Studio: forward standard OpenAI tools / tool_choice on /v1/responses (Codex compat) (#5122) 2026-04-21 13:17:20 +04:00
test_studio_api.py Studio: forward standard OpenAI tools / tool_choice to llama-server (#5099) 2026-04-18 12:53:23 +04:00
test_tool_policy_gates.py unsloth run: add --enable-tools/--disable-tools server-side tool policy (#5277) 2026-05-05 12:45:15 +04:00
test_tool_policy_state.py unsloth run: add --enable-tools/--disable-tools server-side tool policy (#5277) 2026-05-05 12:45:15 +04:00
test_trained_model_scan.py [Studio] Show non exported models in chat UI (#4892) 2026-04-14 15:03:58 +04:00
test_training_history_update.py Studio: Dark theme refactor, right sidebar redesign, and chat UI polish (#5150) 2026-05-07 14:33:31 +04:00
test_training_raw_support.py feat(studio): add Continued Pretraining (CPT) as a training method (#4677) 2026-05-06 13:38:35 +04:00
test_training_worker_flash_attn.py Add Qwen3.6 support (#5257) 2026-05-02 23:30:57 +04:00
test_transformers_version.py split venv_t5 into tiered 5.3.0/5.5.0 and fix trust_remote_code (#4878) 2026-04-07 20:05:01 +04:00
test_utils.py Add AMD ROCm/HIP support across installer and hardware detection (#4720) 2026-04-10 01:56:12 -07:00
test_vision_cache.py Studio: split vision-cache exception test to match transient vs permanent (#5145) 2026-04-23 00:22:40 -07:00
test_vram_estimation.py Update VRAM estimator to cater to broader model configs (#5175) 2026-05-05 04:12:36 -07:00