unsloth/studio/backend/tests
Leo Borcherding 61df3aaef0
feature/use lemonade-sdk llamacpp-rocm binaries (#5303)
* feat(studio): use lemonade-sdk/llamacpp-rocm per-GPU prebuilts for ROCm hosts

For AMD GPUs that rocminfo/hipinfo reports a recognised gfx target
(gfx103X / gfx110X / gfx1150 / gfx1151 / gfx120X), resolve_lemonade_rocm_choice()
now fetches the latest lemonade-sdk/llamacpp-rocm release and returns the
matching per-architecture zip, bundling all required ROCm runtime libs.

This runs before the existing upstream ggml-org combined-ROCm tarball fallback
on Linux and before the upstream HIP zip on Windows, so both platforms benefit
from the more targeted build when available.

Changes:
- Add LEMONADE_ROCM_REPO / LEMONADE_ROCM_RELEASES_API constants
- Add HostInfo.rocm_gfx_target populated from rocminfo (Linux) / hipinfo (Windows)
- Add _LEMONADE_GFX_FAMILIES prefix map and _lemonade_gfx_family() helper
- Add resolve_lemonade_rocm_choice() that fetches latest lemonade release and
  constructs the llama-{tag}-{os}-rocm-{gfxFamily}-x64.zip asset URL
- Wire into resolve_upstream_asset_choice() for both Linux (ubuntu) and Windows paths

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Honor pinned llama.cpp tag in lemonade ROCm resolver

The resolver always fetched lemonade-sdk/llamacpp-rocm's /releases/latest,
ignoring the upstream llama.cpp tag the caller had pinned. On a reproducible
install where the user requested 'b1260' that meant we would silently pick
up whatever lemonade had published as latest at install time, with no way
to roll back to the matching tag.

Lemonade tags llama.cpp upstream tags 1:1, so:
  - When llama_tag is unset or 'latest', keep hitting /releases/latest.
  - When llama_tag is pinned (e.g. 'b1260'), hit /releases/tags/b1260.
  - When the pinned tag is not published by lemonade (404), skip silently
    and let the caller fall through to the upstream tarball -- this keeps
    pinned installs reproducible instead of drifting.

resolve_lemonade_rocm_choice now takes llama_tag (default 'latest' for
backward compatibility) and both call sites in resolve_upstream_asset_choice
forward the upstream llama_tag to it.

Note: this PR still has open integration concerns flagged in review --
the simple-policy planner and approved-checksum manifest don't yet route
or accept lemonade assets. Those are larger changes and not in scope for
this commit; addressing the pinned-tag drift independently because it is
small, localized, and self-contained.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Add mock test for lemonade ROCm prebuilt asset resolution

Validates GPU family mapping and that resolve_lemonade_rocm_choice
returns real lemonade release URLs for all supported gfx targets on
both Linux and Windows, without requiring AMD hardware.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Wire lemonade ROCm prebuilts into the simple-policy install path

setup.sh invokes install_llama_prebuilt.py with --simple-policy, which
dispatches through resolve_simple_install_release_plans -> direct_linux_release_plan
(or direct_upstream_release_plan on Windows). Those planners only
handled CUDA + CPU attempts, so ROCm-only hosts (e.g. gfx1151 Strix
Halo) had no compatible prebuilt asset and silently fell through to
source build, even though resolve_lemonade_rocm_choice already knew
how to fetch a per-GPU lemonade-sdk binary.

Add a lemonade ROCm/HIP attempt to both simple-policy planners for
ROCm-only hosts, and add regression tests that drive the dispatchers
end-to-end so this can't be skipped silently again.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Document that lemonade ROCm prebuilts work on any glibc Linux

The lemonade-sdk asset filename uses "ubuntu" as a label, but the binary
is a manylinux-style glibc build with no Ubuntu-specific dependencies.
It runs on Arch, Fedora, openSUSE, Debian, etc. as long as the host
glibc is recent enough.

No behavior change -- the dispatch already runs for any Linux ROCm
host. This commit only clarifies the comment, docstring, and log
message so users on non-Ubuntu distros (e.g. Strix Halo on Arch) don't
mistake the asset name for distro gating.

* Pattern matching fix for libggml-cpu*.so*

* fix(lemonade): pass resolved tag to lemonade resolver; add upstream HIP fallback; stub API in tests

- direct_linux_release_plan: pass bundle.upstream_tag (not requested_tag)
  to resolve_lemonade_rocm_choice so a "latest" request doesn't mix a
  newer lemonade binary with an older planned unsloth release (Codex P2)

- direct_upstream_release_plan: same fix on the Windows path (release_tag
  instead of requested_tag); also add the upstream HIP asset
  (llama-<tag>-bin-win-hip-radeon-x64.zip) as a fallback between lemonade
  and CPU so unsupported GPUs or transient lemonade failures don't silently
  downgrade to CPU when an upstream ROCm prebuilt exists (Codex P2)

- test file: stub fetch_json with a synthetic lemonade release payload so
  the suite is hermetic and not subject to GitHub API rate limits (Codex P1);
  add test_simple_policy_windows_hip_falls_back_to_upstream_when_lemonade_unavailable
  to cover the new HIP fallback path

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(lemonade): revert bad tag fix; keep upstream HIP fallback + hermetic tests

The previous commit wrongly passed bundle.upstream_tag / release_tag to
resolve_lemonade_rocm_choice. Lemonade uses its own versioning (b1262,
b1264, …) completely independent of unslothai's tags (b9186, …), so
passing a resolved unslothai tag caused a 404 and silently skipped the
lemonade binary entirely. Revert both call sites to requested_tag.

Keep the two valid fixes from the prior commit:
- Upstream HIP fallback (llama-<tag>-bin-win-hip-radeon-x64.zip) between
  lemonade and CPU in direct_upstream_release_plan, so unsupported GPUs
  or transient lemonade failures don't silently downgrade to CPU (Codex P2)
- Stub fetch_json in tests so the suite is hermetic (Codex P1)

* fix(studio/rocm): respect HIP_VISIBLE_DEVICES when picking lemonade gfx target

The rocminfo / hipinfo regex took the first gfx match in the agent listing.
On mixed APU + dGPU hosts (e.g. Strix Halo gfx1151 + discrete RX 7900 gfx1100)
this picked whichever GPU appeared first in the tool's stdout, not the one
HIP actually runs on. The downloaded lemonade asset could then be a binary
for a different arch than the active device.

Extracted a module-level _pick_rocm_gfx_target() helper that:
- collects every gfx token in order via re.findall (skips gfx000 / generic ISAs)
- if HIP_VISIBLE_DEVICES or ROCR_VISIBLE_DEVICES is set, parses the first
  comma-separated entry as an integer index into that list
- falls back to the first GPU for non-integer (UUID-style) or out-of-range
  values, matching the previous default behaviour

Both Linux (rocminfo) and Windows (hipinfo) branches use the helper.
Existing 18 lemonade tests pass; no behavioural change for single-GPU hosts.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(studio/rocm): dedup rocminfo gfx tokens and honor disabled visibility

Follow-up to 8793aef0. rocminfo / hipinfo emit each gfx target multiple
times per GPU (Name, ISA triple, marketing name), so the prior re.findall
indexing returned the wrong device when HIP_VISIBLE_DEVICES picked GPU 1
on a mixed-arch host -- the helper picked the second occurrence of GPU 0
instead. Collapse to unique tokens (insertion-ordered) before indexing.

Also handle HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES values of '' and
'-1' as "no AMD visible" (matches the rest of Studio's visibility code),
returning None so the planner does not pick a Lemonade asset for a
hidden GPU.

* fix(studio/rocm): memoise lemonade release lookup + HIP_PATH hipinfo fallback

Robustness pass on top of 25d4ab63:

1. resolve_lemonade_rocm_choice() is called twice per install (planner +
   resolve_upstream_asset_choice), so every install previously hit
   api.github.com twice with identical args -- doubling the 403/rate-limit
   failure surface on busy CI runners. Extract the fetch into a
   functools.lru_cache(maxsize=8) helper keyed on (api_url, llama_tag).
   The cached helper also owns the error-path logging the resolver was
   doing inline. Tests that need to vary fetch_json output across
   invocations should call _fetch_lemonade_release_cached.cache_clear().

2. Windows detect_host probe for hipinfo / amd-smi only used
   shutil.which(), so HIP SDK installs that set HIP_PATH but do not put
   %HIP_PATH%\bin on system PATH classified the host as non-ROCm. The
   PowerShell installer and studio/install_python_stack.py already
   resolve HIP_PATH\bin\hipinfo.exe as a fallback; mirror that here so
   the install planner agrees with the rest of Studio on what counts as
   a ROCm host.

Tests: 18 passed (shipped); 52 extra sim cases pass (added 2 for the
lru_cache deduplication path).

* fix(studio/rocm): lemonade URL trust pinning, runtime overlay covers HIP libs, opt-out env

Robustness pass driven by 5 parallel reviewers of head c08b15e6:

1. URL trust pinning. AssetChoice.url comes from the GitHub API response's
   browser_download_url field. Lemonade attempts are not in the approved-hash
   manifest, so a compromised API response could redirect the download to an
   attacker-chosen host without the integrity gate catching it. New
   _is_trusted_github_release_url() validates https + github.com/<expected_repo>
   release path OR objects.githubusercontent.com (GitHub's CDN). Resolver
   refuses to download otherwise.

2. UNSLOTH_DISABLE_LEMONADE_ROCM opt-out. Users who prefer the upstream HIP
   build path can set this env var to skip lemonade outright (e.g. for
   air-gapped installs or stricter trust requirements). The install log
   already prints a NOTE explaining that lemonade lacks approved-hash
   coverage so users know the trust model.

3. linux-rocm runtime overlay patterns extended to cover lemonade's bundled
   HIP/ROCm runtime libs (libamdhip64.so*, libhsa-runtime64.so*, libhipblas*,
   librocblas*, librocsolver*, librocsparse*, librocrand*, libMIOpen*,
   libmagma*). The upstream tarball does not ship these (links against
   system /opt/rocm), so the new glob entries are no-op for upstream and
   load-bearing for lemonade. Without this, install_from_archives would
   drop the bundled runtime libs from the lemonade ZIP and llama-server's
   RPATH would fail to load amdhip64 at first inference.

4. _lemonade_release_api_for now URL-encodes llama_tag with quote(safe="").
   Defence in depth: a tag containing /, ?, #, or whitespace cannot reshape
   the request URL. Tags come from internal resolution today, but this
   removes the risk if a future caller passes user-controlled input.

5. Empty browser_download_url skipped explicitly in the resolver. The
   release_asset_map helper defaults missing URLs to "". Previously this
   would have been passed to download_file("") which raises a less obvious
   error than the new clean log + return None.

6. Docstring on _lemonade_release_api_for clarified: lemonade tags match
   ggml-org/llama.cpp tags 1:1, NOT unslothai/llama.cpp fork tags. The
   earlier P1 review report misread this and re-asked for the tag-drift
   fix that the author intentionally reverted in f256cea950.

7. Autouse pytest fixture in test_lemonade_llamacpp_rocm_bins_mock.py
   clears _fetch_lemonade_release_cached between tests. Today's tests all
   mock the same payload so no pollution surfaces, but the lru_cache
   becomes a footgun the moment any future test parametrises return values.

Tests: 28 passed (was 18, added 10 covering URL pinning, opt-out env,
pinned-tag helper, URL encoding, empty URL, runtime patterns).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(studio/rocm): complete lemonade runtime overlay + honor CUDA_VISIBLE_DEVICES

runtime_patterns_for_choice("linux-rocm") was missing libamd_comgr.so*,
librocm_kpack.so*, and librocm_sysdeps_*.so*, which are direct NEEDED
entries of libamdhip64.so.7 in every lemonade bundle. install_from_archives
did not copy them, so the preflight ldd walk failed on llama-server and every
lemonade attempt fell back to source build. Confirmed on gfx1151 b1272 by
h34v3nzc0dex; manually staging the full bundle with the three missing
patterns restored the install and preserved bench parity.

Also extend _pick_rocm_gfx_target to check CUDA_VISIBLE_DEVICES after
HIP_VISIBLE_DEVICES and ROCR_VISIBLE_DEVICES: AMD's HIP runtime honours
all three with identical semantics, so mixed-arch hosts where users set
only CUDA_VISIBLE_DEVICES were still picking the first gfx token instead
of the user-selected device.

Tests: 30 passed (was 28); added 2 for CUDA_VISIBLE_DEVICES (multi-GPU
pick + -1 opt-out) and extended the runtime-patterns test to assert the
three newly added lib globs are present.

* fix: tighten CDN trust check and fix multi-GPU arch selection

- _is_trusted_github_release_url: require /github-production-release-asset-
  path prefix so only real GitHub release CDN URLs are accepted
- _pick_rocm_gfx_target: parse rocminfo Agent N section boundaries to build
  a per-physical-GPU arch list; same arch across multiple GPUs no longer
  collapses to a single token, so HIP_VISIBLE_DEVICES indexing works correctly
- tests: update CDN test to use realistic path prefix, add rejection test for
  arbitrary CDN path, add regression test for same-arch multi-GPU case

* fix: use broad lib*.so* glob for linux-rocm runtime overlay

Replace the explicit lib allowlist in runtime_patterns_for_choice with
lib*.so* for the linux-rocm path.

The lemonade ROCm ZIPs carry a full HIP/ROCm runtime including transitive
deps like libLLVM.so.23.0git and libclang-cpp.so.23.0git (pulled in by
libamd_comgr.so.3). These names change across ROCm releases and were not
in the allowlist, so preflight would see them as unresolved NEEDED entries
and fall back to a source build. The broad glob catches everything in the
bundle now and in future releases without needing to enumerate each library.

* fix: show lemonade binary tag in install summary log line

Store binary_repo and binary_release_tag in UNSLOTH_PREBUILT_INFO.json
so the setup.sh summary can distinguish the source tree (unslothai/llama.cpp)
from the actual binary origin (lemonade-sdk/llamacpp-rocm).

Before: 'installed release: unslothai/llama.cpp@b9334'
After:  'installed release: unslothai/llama.cpp@b9334 + lemonade@b1280'

Upstream installs are unchanged (binary_repo == published_repo).

* fix(merge-compat): align Windows ROCm guard and helper name with strix branch

- elif host.has_rocm instead of if...not to match fix/rocm-strix-halo-unified-memory
- _resolve_exe instead of _resolve_amd_exe with identical body/docstring

Eliminates the two conflict hunks that would otherwise block a clean bot merge
of feature/lemonade-rocm-prebuilts on top of fix/rocm-strix-halo-unified-memory.

* fix: three PR review corrections for lemonade ROCm prebuilt integration

- _lemonade_release_api_for: fix docstring claiming lemonade tags match
  ggml-org 1:1 -- lemonade may be several builds behind ggml-org (noted
  by oobabooga, confirmed: lemonade b1281 vs ggml-org b9370)
- direct_linux_release_plan: move cpu_choice into else-branch so ROCm-only
  hosts never get a CPU fallback in the attempts list -- a failed lemonade
  binary now raises PrebuiltFallback and triggers the HIP source build
  instead of silently installing a CPU-only binary
- apply_approved_hashes: pass lemonade attempts through without requiring
  a manifest entry; lemonade is explicitly documented as relying on
  functional validation only, so rejecting it here caused PrebuiltFallback
  on the non-simple-policy path before the upstream ROCm/HIP fallback
  could be considered

* fix: copy hipblaslt/rocblas library subdirs from lemonade ROCm archives

copy_globs matches filenames only and copies flat, so it cannot
preserve the hipblaslt/library/<gfx>/ and rocblas/library/<gfx>/
Tensile kernel catalog trees that lemonade ROCm zips ship alongside
the .so files. Add runtime_subdirs_for_choice() and call shutil.copytree
for each named subdir after the copy_globs pass in install_from_archives.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Co-authored-by: Jaeic Lee <jaeiclee@users.noreply.github.com>
2026-05-29 22:46:03 -07:00
..
__init__.py Final cleanup 2026-03-12 18:28:04 +00:00
conftest.py Studio: Expose openai and anthropic compatible external API end points (#4956) 2026-04-13 21:08:11 +04:00
test_amd_apu_unified_memory.py fix/strix halo and windows AMD ROCm support (#5301) 2026-05-29 22:29:56 -07:00
test_anthropic_cache_ttl.py Studio: support Anthropic 1h cache TTL via prompt_cache_ttl (#5685) 2026-05-22 06:03:32 -07:00
test_anthropic_citations.py Studio: surface Anthropic document citations inline + in Sources panel (#5718) 2026-05-25 23:39:02 -07:00
test_anthropic_citations_edge.py Studio: surface Anthropic document citations inline + in Sources panel (#5718) 2026-05-25 23:39:02 -07:00
test_anthropic_code_execution.py Studio: add Gemini provider with web_search, code_execution, prompt caching, and Nano Banana image generation (#5720) 2026-05-27 06:01:24 -07:00
test_anthropic_compaction.py Studio: wire Anthropic server-side context compaction (#5686) 2026-05-22 06:19:09 -07:00
test_anthropic_fast_mode_and_refusal.py Studio: Anthropic fast_mode toggle and streaming refusal handling (#5715) 2026-05-25 23:37:12 -07:00
test_anthropic_fast_mode_edge.py Studio: Anthropic fast_mode toggle and streaming refusal handling (#5715) 2026-05-25 23:37:12 -07:00
test_anthropic_messages.py Studio: Claude Code Anthropic API tool compatibility (#5390) 2026-05-21 16:45:05 +04:00
test_anthropic_thinking_translation.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
test_anthropic_tool_versions.py Studio: per-model Anthropic server-side tool versions (#5679) 2026-05-22 06:03:27 -07:00
test_anthropic_web_fetch.py Studio: add Gemini provider with web_search, code_execution, prompt caching, and Nano Banana image generation (#5720) 2026-05-27 06:01:24 -07:00
test_browse_folders_route.py Studio: add folder browser modal for Custom Folders (#5035) 2026-04-15 08:04:33 -07:00
test_cache_case_resolution.py Add tests for cache case resolution (from PR #4822) (#4823) 2026-04-03 13:58:26 -07:00
test_cached_gguf_routes.py Studio: support GGUF variant selection for non-suffixed repos (#5023) 2026-04-15 15:32:01 +04:00
test_chat_history_routes.py Studio: persist chat history in backend storage (#5272) 2026-05-22 06:18:05 -07:00
test_chat_history_storage.py Studio: persist chat history in backend storage (#5272) 2026-05-22 06:18:05 -07:00
test_cleanup_cancelled_checkpoints.py studio: scope cancel-cleanup to in-flight tmp dirs; walk back tool_call_id (#5488) 2026-05-18 00:01:48 -07:00
test_cpu_threads.py Clear MRoPE after generation for GRPO (#5683) 2026-05-27 07:32:20 -07:00
test_data_recipe_github_progress.py Studio: add github_repo seed reader and GitHub Support Bot recipe (#5169) 2026-04-24 12:02:03 -07:00
test_data_recipe_seed.py fix(seed): disable remote code execution in seed inspect dataset loads (#4275) 2026-03-13 19:37:43 +04:00
test_desktop_auth.py Studio: add remote MCP server support (#5750) 2026-05-27 07:01:11 -07:00
test_detect_mmproj_file.py fix(studio/mmproj): block cross-family projectors in flat local GGUF dirs (#5347) (#5350) 2026-05-14 20:31:20 -07:00
test_export_log_cursor.py studio: stream export worker output into the export dialog (#4897) 2026-04-14 08:55:43 -07:00
test_external_provider_usage_chunk.py Studio: surface prompt-cache token counts in /v1/chat/completions usage chunk (#5670) 2026-05-22 06:02:52 -07:00
test_frontend_resolution.py Studio: auto-recover when shadowed 'unsloth' on PATH hides the frontend dist (#5782) 2026-05-26 05:29:42 -07:00
test_gemini_provider.py Studio: add Gemini provider with web_search, code_execution, prompt caching, and Nano Banana image generation (#5720) 2026-05-27 06:01:24 -07:00
test_gguf_completion_usage.py Fix non-streaming GGUF chat completion usage (#5781) 2026-05-28 13:28:52 +04:00
test_gguf_metadata.py fix(studio/mmproj): block cross-family projectors in flat local GGUF dirs (#5347) (#5350) 2026-05-14 20:31:20 -07:00
test_gguf_reload_inheritance.py studio: add --spec-draft-n-max toggle for MTP speculative decoding (#5582) 2026-05-19 06:17:04 -07:00
test_gpu_selection.py Update VRAM estimator to cater to broader model configs (#5175) 2026-05-05 04:12:36 -07:00
test_gpu_selection_sandbox.py [Studio] multi gpu finetuning/inference via "balanced_low0/sequential" device_map (#4602) 2026-03-30 02:33:15 -07:00
test_host_defaults.py Default Studio host to 127.0.0.1 and prompt before auto-start (#5267) 2026-05-04 13:03:16 +04:00
test_index_bootstrap_origin.py Studio: stop seeded admin to cross-origin callers (#5739) 2026-05-25 23:36:51 -07:00
test_index_bootstrap_origin_extra.py Studio: stop seeded admin to cross-origin callers (#5739) 2026-05-25 23:36:51 -07:00
test_inference_model_validation.py studio: scope cancel-cleanup to in-flight tmp dirs; walk back tool_call_id (#5488) 2026-05-18 00:01:48 -07:00
test_kv_cache_estimation.py studio: reserve VRAM headroom for the MTP draft cache in auto-fit (#5585) 2026-05-19 06:19:02 -07:00
test_lemonade_llamacpp_rocm_bins_mock.py feature/use lemonade-sdk llamacpp-rocm binaries (#5303) 2026-05-29 22:46:03 -07:00
test_llama_cpp_cache_aware_disk_check.py Studio: make GGUF disk-space preflight cache-aware (#5012) 2026-04-14 08:53:37 -07:00
test_llama_cpp_context_fit.py fix: honor --ctx-size and other forwarded args from unsloth studio run in Studio's context-fit logic (#5815) 2026-05-28 11:34:35 +04:00
test_llama_cpp_freshness.py Studio: warn when llama.cpp prebuilt is at least 3 days behind (#5529) 2026-05-18 00:21:50 -07:00
test_llama_cpp_load_progress.py Studio: live model-load progress + rate/ETA on download and load (#5017) 2026-04-14 09:46:22 -07:00
test_llama_cpp_load_progress_live.py Studio: split model-load progress label across two rows (#5020) 2026-04-14 10:58:16 -07:00
test_llama_cpp_load_progress_matrix.py Studio: split model-load progress label across two rows (#5020) 2026-04-14 10:58:16 -07:00
test_llama_cpp_max_context_threshold.py fix KVCache estimates for gemma4 style sliding window models (#5225) 2026-05-05 04:06:46 -07:00
test_llama_cpp_mtp_detection.py studio: add --spec-draft-n-max toggle for MTP speculative decoding (#5582) 2026-05-19 06:17:04 -07:00
test_llama_cpp_no_context_shift.py Studio: hard-stop at n_ctx with a 'Context limit reached' toast (#5021) 2026-04-14 10:58:20 -07:00
test_llama_cpp_wait_for_health.py tests/studio: lock in Windows GPU detection fix (#5106) with a synthetic CI test (#5376) 2026-05-18 00:06:01 -07:00
test_llama_cpp_wait_for_vram_settle.py studio: settle GPU VRAM after killing llama-server before the next reload (#5693) 2026-05-22 05:50:39 -07:00
test_llama_cpp_windows_nvidia_path.py Studio: add torch's pip nvidia DLL dirs to PATH on Windows (#5324) 2026-05-11 05:42:09 -07:00
test_llama_server_args.py fix: honor --ctx-size and other forwarded args from unsloth studio run in Studio's context-fit logic (#5815) 2026-05-28 11:34:35 +04:00
test_log_filter_no_truncation.py fix/strix halo and windows AMD ROCm support (#5301) 2026-05-29 22:29:56 -07:00
test_login_rate_limit.py studio: proxy-aware login rate-limit; allow google favicons in CSP (#5489) 2026-05-18 00:02:15 -07:00
test_mcp_servers.py Studio: add remote MCP server support (#5750) 2026-05-27 07:01:11 -07:00
test_middleware.py studio: proxy-aware login rate-limit; allow google favicons in CSP (#5489) 2026-05-18 00:02:15 -07:00
test_mlx_inference_backend.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00
test_mlx_training_worker_config.py Studio: expose image size setting in training UI (#5743) 2026-05-27 05:01:24 -07:00
test_models_get_model_config_case_resolution.py Add tests for cache case resolution (from PR #4822) (#4823) 2026-04-03 13:58:26 -07:00
test_multimodal_document.py Studio: surface Anthropic document citations inline + in Sources panel (#5718) 2026-05-25 23:39:02 -07:00
test_native_context_length.py Studio: Fix chat template disappearing after browser refresh (#5209) 2026-05-01 08:19:09 -07:00
test_offline_gguf_cache_fallback.py studio: load cached GGUF models when fully offline (#5505) 2026-05-17 21:25:39 -07:00
test_offline_inference_parent.py studio: extend offline DNS auto-detect to inference parent + training (#5512) 2026-05-18 00:31:33 -07:00
test_openai_citation_markers.py Studio: rewrite OpenAI Responses citation markers to markdown links (#5713) 2026-05-25 23:37:16 -07:00
test_openai_citation_markers_edge.py Studio: rewrite OpenAI Responses citation markers to markdown links (#5713) 2026-05-25 23:37:16 -07:00
test_openai_code_execution.py Studio: add Gemini provider with web_search, code_execution, prompt caching, and Nano Banana image generation (#5720) 2026-05-27 06:01:24 -07:00
test_openai_compaction.py Studio: wire OpenAI Responses server-side context compaction (#5687) 2026-05-22 06:20:45 -07:00
test_openai_container_crud.py tests/openai: patch httpx.AsyncClient ctor so delete tests hit mock (#5469) 2026-05-15 15:53:54 -07:00
test_openai_image_generation.py Studio: add Gemini provider with web_search, code_execution, prompt caching, and Nano Banana image generation (#5720) 2026-05-27 06:01:24 -07:00
test_openai_responses_translation.py Studio: add Gemini provider with web_search, code_execution, prompt caching, and Nano Banana image generation (#5720) 2026-05-27 06:01:24 -07:00
test_openai_tool_passthrough.py Fix GGUF multi-image chat handling (#5508) 2026-05-19 04:36:20 -07:00
test_openai_tool_result_fallbacks.py Studio: per-card web_search result + shell_call output fallback (OpenAI) (#5785) 2026-05-26 04:31:22 -07:00
test_pricing.py Studio: pricing follow-up to #5690 (longest-prefix match + chat-style usage keys) (#5722) 2026-05-25 23:39:58 -07:00
test_pricing_edge.py Studio: pricing follow-up to #5690 (longest-prefix match + chat-style usage keys) (#5722) 2026-05-25 23:39:58 -07:00
test_providers_api.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
test_pytorch_mirror.py Add configurable PyTorch mirror via UNSLOTH_PYTORCH_MIRROR env var (#5024) 2026-04-15 11:39:11 +04:00
test_recommended_folders_permission.py Fix /recommended-folders 500 on unreadable model directories (Python 3.12+) (#5523) 2026-05-18 00:16:14 +04:00
test_responses_api.py Studio: Expose openai and anthropic compatible external API end points (#4956) 2026-04-13 21:08:11 +04:00
test_responses_tool_passthrough.py Studio: forward standard OpenAI tools / tool_choice on /v1/responses (Codex compat) (#5122) 2026-04-21 13:17:20 +04:00
test_rocm_oom_guard.py fix/strix halo and windows AMD ROCm support (#5301) 2026-05-29 22:29:56 -07:00
test_safetensors_capability_advertise.py Revert "studio: tool calling for Llama-3, Mistral, Gemma 4 on safetensors + MLX (#5615)" (#5619) 2026-05-19 07:26:39 -07:00
test_safetensors_tool_loop.py Revert "studio: tool calling for Llama-3, Mistral, Gemma 4 on safetensors + MLX (#5615)" (#5619) 2026-05-19 07:26:39 -07:00
test_sandbox_tools.py studio: tighten sandbox blocklist precision (bash, hf upload, NOFILE) (#5487) 2026-05-18 00:01:17 -07:00
test_studio_api.py Studio: forward standard OpenAI tools / tool_choice to llama-server (#5099) 2026-04-18 12:53:23 +04:00
test_studio_train_validation.py Studio: expose image size setting in training UI (#5743) 2026-05-27 05:01:24 -07:00
test_tool_policy_gates.py unsloth run: add --enable-tools/--disable-tools server-side tool policy (#5277) 2026-05-05 12:45:15 +04:00
test_tool_policy_state.py unsloth run: add --enable-tools/--disable-tools server-side tool policy (#5277) 2026-05-05 12:45:15 +04:00
test_tool_xml_strip.py Studio: strip orphan tool_call XML leaking into visible content (#5735) 2026-05-24 05:00:08 -07:00
test_trained_model_scan.py studio: security and hardening pass (auth rate-limit, sandbox, path containment, schema validation, headers) (#5375) 2026-05-13 06:12:18 -07:00
test_training_history_update.py Studio: Dark theme refactor, right sidebar redesign, and chat UI polish (#5150) 2026-05-07 14:33:31 +04:00
test_training_raw_support.py studio: drop unused max_grad_value schema + route plumbing (#5424) 2026-05-14 05:43:58 -07:00
test_training_worker_flash_attn.py studio: install flash-linear-attention and tilelang for Qwen3.5 family (#5434) 2026-05-18 03:49:06 -07:00
test_transformers_version.py split venv_t5 into tiered 5.3.0/5.5.0 and fix trust_remote_code (#4878) 2026-04-07 20:05:01 +04:00
test_utils.py Add AMD ROCm/HIP support across installer and hardware detection (#4720) 2026-04-10 01:56:12 -07:00
test_vision_cache.py Studio: split vision-cache exception test to match transient vs permanent (#5145) 2026-04-23 00:22:40 -07:00
test_vram_estimation.py Update VRAM estimator to cater to broader model configs (#5175) 2026-05-05 04:12:36 -07:00
test_windows_gpu_detection_mock.py tests/studio: lock in Windows GPU detection fix (#5106) with a synthetic CI test (#5376) 2026-05-18 00:06:01 -07:00