Installer: opt-in Vulkan llama.cpp backend (and fallback when no AMD card is HIP-supported) (#7373)
* feat(install): opt-in Vulkan llama.cpp backend and HIP gfx fallback (#7357) Add UNSLOTH_LLAMA_BACKEND=vulkan and --llama-backend vulkan to force the upstream Vulkan prebuilt on any host, persist llama_backend in the install marker, and re-assert it during Studio updates. On Windows AMD, auto-fallback to Vulkan when no detected gfx arch is in the upstream win-hip-radeon GPU_TARGETS set (e.g. gfx803 / RX 480). Mixed setups where at least one card is HIP-supported still default to HIP unless opted in. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix(install): address Codex P2s on Vulkan gfx routing (#7357) Honor ROCm family tokens (gfx110X), include fork-supported gfx1103, require a known active gfx before auto-Vulkan, and base the HIP floor check on the visible-device target instead of every physical GPU in hipinfo. * Address Codex review: env namespace, physical-NVIDIA guard, test kwarg - llama_backend_from_env: stop reading UNSLOTH_LLAMA_CPP_BACKEND. That is a separate pre-existing setup variable meaning auto/cpu; setup.sh/setup.ps1 warn and ignore other values, so reading it here forced Vulkan behind that warning. Vulkan opt-in stays on UNSLOTH_LLAMA_BACKEND / UNSLOTH_FORCE_VULKAN. - _should_auto_vulkan_for_amd_windows: gate on not has_physical_nvidia (not merely has_usable_nvidia). A CUDA-masked NVIDIA card keeps has_physical_nvidia while has_usable_nvidia goes False; Vulkan ignores CUDA_VISIBLE_DEVICES and could enumerate the reserved card. Mirrors the Intel auto path. Explicit opt-in still overrides. - test fakes: validate_prebuilt_attempts/validate_prebuilt_choice gained a llama_backend kwarg; the four fake signatures in the fallback tests now accept it, clearing the TypeError that reddened Backend CI / Repo tests (CPU). Tests: UNSLOTH_LLAMA_CPP_BACKEND=vulkan no longer triggers Vulkan; hidden physical NVIDIA suppresses AMD auto-Vulkan while explicit opt-in overrides. * Keep gfx1034 on the ROCm path (fork gfx103X bundle covers it) The WINDOWS_HIP_PREBUILT_GFX_TARGETS allow-list omitted gfx1034, so _route_to_vulkan_prebuilt downgraded RX 6500/6400-class hosts to the upstream Vulkan prebuilt before published_rocm_choice_for_host could match the fork windows-rocm gfx103X bundle (whose members include gfx1034). Add gfx1034 to the allow-list and a regression test asserting it stays on the fork ROCm asset. * Fix auto-Vulkan stealing fork windows-rocm gfx908/gfx90a hosts for PR #7373 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix Vulkan marker claiming a backend that was never installed for PR #7373 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Tighten the Vulkan backend routing comments for PR #7373 * Keep the visible-device-aware gfx when setup forwards --rocm-gfx setup.ps1 resolves the gfx arch from its own probe, and that pick is not fully visible-device aware: neither the hipinfo nor the amd-smi branch reads CUDA_VISIBLE_DEVICES, and the amd-smi branch matches a bare integer only, so a comma-separated HIP/ROCR mask such as 1,0 also falls back to GPU 0. The resulting arch was then forwarded through --rocm-gfx and replaced the arch detect_host() had already resolved for the runtime-visible GPU. On a mixed-AMD Windows host that flipped the auto-Vulkan decision: with GPU 0 gfx1100 and a masked-in gfx1010, the forward reinstated gfx1100, _should_auto_vulkan_for_amd_windows() saw a HIP-supported arch and the HIP bundle was installed for a GPU that cannot run it. Fold the forward in as a fill rather than a replacement: it still supplies the arch on amd-smi-only, driver-only and name-inferred hosts where the probe reports none, which is what --rocm-gfx exists for, but no longer overwrites a successfully detected active arch. An explicit UNSLOTH_ROCM_GFX_ARCH stays authoritative, since it is the documented manual override for hosts whose arch the probes get wrong. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Scope the Windows AMD Vulkan fallback per device and per repo Three follow-ups on the auto-Vulkan routing for #7357. Keep an explicit --rocm-gfx authoritative. The previous round stopped a forwarded gfx from replacing an arch detect_host() had already resolved, but --rocm-gfx is also the documented operator override for hosts whose probe is wrong or stale, and both arrive as the same argv. Narrow the advisory case to the two shapes setup can actually be describing: an arch the probe saw on this host (setup picked a different physical GPU of the same box), or a family label such as gfx110X, which is a bundle name the update path derives from the marker asset rather than a real GPU arch. Any other value is an override for an arch no probe reported and stays authoritative. Keeping family labels advisory also preserves the rule that an in-generation-but-unbuilt arch (gfx1033) is never upgraded into the gfx103X bundle. Do not auto-route to Vulkan from a HIP-only device mask. HIP_VISIBLE_DEVICES, ROCR_VISIBLE_DEVICES and CUDA_VISIBLE_DEVICES select the active arch, but the Vulkan runtime honours none of them: it enumerates through GGML_VK_VISIBLE_DEVICES and Vulkan ordinals in LlamaCppBackend._get_gpu_free_memory_vulkan. Masking down to a below-floor card therefore used to install a backend that could still enumerate the HIP-capable card the user deliberately hid, possibly one reserved for another workload. Require every physical AMD gfx to be below the floor, matching the has_physical_nvidia gate right above it. So the per-GPU list survives to that check, a forward that agrees with the probe no longer collapses rocm_gfx_targets to a single entry. Make the HIP support predicate repository-specific. The floor constant is a union of ggml-org's windows-hip gpu_targets and the fork's windows-rocm bundles, so it only answers "is this arch served" for the fork. With --published-repo ggml-org/llama.cpp, direct_upstream_release_plan() offers win-hip-radeon then CPU and never Vulkan, so the four fork-only archs (gfx908, gfx90a, gfx1034, gfx1103) were declared supported and fell through to CPU instead of the Vulkan bundle that would actually run. Add UPSTREAM_WINDOWS_HIP_GFX_TARGETS and select the set from the planned repo. * Keep probe-confirmed AMD GPUs in the physical list when a gfx is forwarded rocm_gfx_targets is the physical inventory _should_auto_vulkan_for_amd_windows() reads, so a forwarded --rocm-gfx that the probe never reported was deleting cards the probe had confirmed. On a mixed Windows AMD box whose active device is masked down to a below-floor card, a stale UNSLOTH_ROCM_GFX_ARCH or a name-inferred arch for the other GPU collapsed the list to that one arch, the floor check concluded no AMD GPU on the host reaches the Windows HIP prebuilt, and the install auto-fell back to Vulkan, which honours no HIP mask and would enumerate the reserved HIP-capable card. Add the forwarded arch to the list instead of replacing it: it selects the HIP target, it does not redefine what hardware is present. An empty probe still yields a single-entry list, so the driver-only Windows AMD host the forward exists for keeps its automatic Vulkan fallback, and an explicit --llama-backend vulkan is unaffected. * Do not auto-fall back to Vulkan when a HIP device mask filtered the probe hipinfo is itself a HIP application, and AMD documents HIP_VISIBLE_DEVICES as "only devices whose index is present in the sequence are visible to HIP", with that spelling recommended on Windows. Under a mask the Windows probe therefore enumerates the visible devices, so rocm_gfx_targets is what survived the mask rather than the physical inventory the auto-Vulkan floor check assumes. A masked-out gfx1100 next to a visible gfx803 made the check conclude that no AMD GPU on the box reaches the Windows HIP prebuilt and route the install to Vulkan, which honours none of these masks and would enumerate the reserved card. Decline to guess when a mask is set: the physical inventory is unknowable from a masked probe, so keep the HIP / fork / source path. This only ever turns the automatic fallback off, never on. The driver-only single-GPU host the fallback exists for sets no mask, an all-hiding "" / -1 mask is still handled as no active target rather than a partial view, and an explicit --llama-backend vulkan or UNSLOTH_LLAMA_BACKEND=vulkan is unaffected. Reading the physical inventory through an unmasked re-probe would also correct _pick_rocm_gfx_target, which indexes the token list by the mask value and so already assumes an unmasked probe. That is pre-existing behaviour on main and is left alone here. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Treat an all-hiding HIP device mask as suppressing the Vulkan fallback too The mask guard exempted an empty or -1 value on the grounds that the probe reports no active target under it, but that only holds for the probe: a forwarded --rocm-gfx still reconstructs an active arch, and setup infers that arch from the display-adapter name, which no HIP mask touches. A user who hid every AMD GPU from HIP could therefore still be auto-routed to Vulkan, which honours none of these masks and would then use all of them. That is the strongest form of the hazard the guard exists for, not an exemption from it. Presence of any of the three variables is now the whole test, which also removes the value parsing. An explicit --llama-backend vulkan or UNSLOTH_LLAMA_BACKEND is still unaffected. * Grant the fork-only Windows HIP coverage to the fork, not to every mirror The floor set is a union of the fork's windows-rocm bundles and only the fork is planned from its manifest: resolve_simple_install_release_plans() compares == DEFAULT_PUBLISHED_REPO and sends every other --published-repo through direct_upstream_release_plan(), whose AMD branch offers win-hip-radeon then CPU and never Vulkan. Exempting only the exact ggml-org spelling therefore told a mirror carrying upstream-standard assets that fork-only archs such as gfx1034, gfx1103 and gfx908 were HIP-served, landing them on HIP or CPU instead of the Vulkan bundle that would actually run. Gate on the fork instead. Matching the dispatch exactly, spelling included, also fixes a differently cased repo: that really does take the upstream path, so it must be answered with upstream coverage rather than the fork superset. An empty repo still defaults to the fork, as the resolver does. * Derive the Windows HIP gfx floor guard from the published manifest The guard compared WINDOWS_HIP_PREBUILT_GFX_TARGETS against a second hardcoded tuple in the same test file, so a windows-rocm arch newly published by the fork passed both. Affected hosts would then be routed off the hash-approved fork ROCm bundle onto an unhashed upstream Vulkan build with nothing failing. Read the fork's llama-prebuilt-manifest.json through the installer's own resolver instead, and assert the floor, the family labels, and the routing tuple all still cover what it publishes. The manifest ships only as a release asset, so an unreachable release skips with an explicit reason rather than flaking. Both literals match the manifest as published today. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Compress the Vulkan backend routing comments and docstrings for PR #7373 * Correct the family-label rationale in the Windows HIP coverage check The comment justified serving gfx103X / gfx110X against any repository by claiming upstream's windows-hip targets build every member of those families. The fork manifest maps gfx103X to gfx1030..1032 plus gfx1034 and gfx110X to gfx1100..1102 plus gfx1103, and UPSTREAM_WINDOWS_HIP_GFX_TARGETS carries neither gfx1034 nor gfx1103, so the stated reason is wrong even though the answer is right. State the real reason instead. A family label is a bundle name, not an arch, so the concrete GPU is unknown at this point; answering unsupported to cover the two uncovered members would move gfx1030..1032 and gfx1100..1102 off a working HIP build onto Vulkan for a card the label cannot identify. Those two archs still reach Vulkan through the concrete-arch branch below, which does answer per repository. Comment only. No behaviour change: the 5850-combination override sweep still reports 0 rocm_gfx_target changes, 0 auto_vulkan False to True flips and 680 True to False flips all backed by a probe-confirmed HIP GPU, and both the feature and override profile matrices are byte-identical. * Pin that a deliberate CPU install outranks Vulkan for PR #7373 UNSLOTH_LLAMA_CPP_BACKEND (setup.sh / setup.ps1, "auto" or "cpu") and UNSLOTH_LLAMA_BACKEND (this module, a backend name) are separate variables at separate layers, and both accept "cpu". setup translates its own =cpu into --force-cpu, which is what pins the CPU-only bundle on a GPU host and keeps Intel iGPU Vulkan crashes away (#7213), so no trigger this PR adds may outrank it. _route_to_vulkan_prebuilt already gets this right, since force_cpu short-circuits ahead of the forced, auto-Intel and auto-no-HIP triggers. Cover it so it stays that way: the matrix runs [Linux, Windows, macOS] x [NVIDIA, AMD, Intel, CPU only] x [unset, vulkan, hip, rocm, cpu] with the legacy UNSLOTH_FORCE_VULKAN set as well, and asserts the published bundle survives every one. WSL presents as Linux to this resolver, so it rides the Linux row. Also assert the guard is not vacuous: the same host still takes Vulkan once the CPU pin is gone, so the matrix cannot pass on a resolver that had simply stopped routing to Vulkan. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: LeoBorcherding <borchborchmail@gmail.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com>
This commit is contained in:
parent
4b3809a2f4
commit
7917c7828c
6 changed files with 1309 additions and 84 deletions
|
|
@ -6,7 +6,9 @@ by default; --published-repo overrides).
|
|||
|
||||
These back the in-app update for source-build (markerless) installs: the backend
|
||||
asks the installer whether an official prebuilt exists for this host without
|
||||
downloading. Network and host detection are stubbed; no GPU or internet needed.
|
||||
downloading. Network and host detection are stubbed; no GPU or internet needed. The one
|
||||
exception is the windows-rocm floor guard, which reads the fork's published manifest
|
||||
because nothing in-tree mirrors it, and skips when that release is unreachable.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
|
@ -32,6 +34,18 @@ FORK = ilp.DEFAULT_PUBLISHED_REPO # unslothai/llama.cpp
|
|||
UPSTREAM = ilp.UPSTREAM_REPO # ggml-org/llama.cpp
|
||||
|
||||
|
||||
@pytest.fixture(autouse = True)
|
||||
def _no_ambient_hip_device_mask(monkeypatch):
|
||||
"""These tests describe hosts through HostInfo, not through the environment.
|
||||
|
||||
A mask inherited from the shell (ML boxes commonly export CUDA_VISIBLE_DEVICES) means
|
||||
the arch probe saw only part of the GPUs, which the Windows auto-Vulkan guard treats as
|
||||
an unknown physical inventory. Clear all three so a host is described by its fields
|
||||
alone; the tests that are about the mask set it explicitly."""
|
||||
for _env in ("HIP_VISIBLE_DEVICES", "ROCR_VISIBLE_DEVICES", "CUDA_VISIBLE_DEVICES"):
|
||||
monkeypatch.delenv(_env, raising = False)
|
||||
|
||||
|
||||
def _host(**kw):
|
||||
base = dict(
|
||||
system = "Linux",
|
||||
|
|
@ -407,7 +421,9 @@ def test_route_to_vulkan_prebuilt_auto_intel_goes_upstream_and_drops_fork_pin():
|
|||
# Routing fork -> upstream also drops the fork release pin, which is in a
|
||||
# different tag namespace and would make the upstream resolver miss.
|
||||
host = _host(is_linux = True, is_x86_64 = True, has_intel_gpu = True)
|
||||
routed, repo, tag = ilp._route_to_vulkan_prebuilt(host, FORK, "b9596-mix-abc", force_cpu = False)
|
||||
routed, repo, tag, _persist = ilp._route_to_vulkan_prebuilt(
|
||||
host, FORK, "b9596-mix-abc", force_cpu = False
|
||||
)
|
||||
assert repo == UPSTREAM
|
||||
assert tag == ""
|
||||
assert routed.has_intel_gpu is True
|
||||
|
|
@ -416,7 +432,9 @@ def test_route_to_vulkan_prebuilt_auto_intel_goes_upstream_and_drops_fork_pin():
|
|||
def test_route_to_vulkan_prebuilt_preserves_explicit_upstream_pin():
|
||||
# A pin set WITH an explicit upstream repo is already on upstream -> kept.
|
||||
host = _host(is_linux = True, is_x86_64 = True, has_intel_gpu = True)
|
||||
_routed, repo, tag = ilp._route_to_vulkan_prebuilt(host, UPSTREAM, "b9596", force_cpu = False)
|
||||
_routed, repo, tag, _persist = ilp._route_to_vulkan_prebuilt(
|
||||
host, UPSTREAM, "b9596", force_cpu = False
|
||||
)
|
||||
assert repo == UPSTREAM
|
||||
assert tag == "b9596"
|
||||
|
||||
|
|
@ -424,7 +442,9 @@ def test_route_to_vulkan_prebuilt_preserves_explicit_upstream_pin():
|
|||
def test_route_to_vulkan_prebuilt_cpu_fallback_wins():
|
||||
# --cpu-fallback suppresses Vulkan routing even for an Intel host.
|
||||
host = _host(is_linux = True, is_x86_64 = True, has_intel_gpu = True)
|
||||
routed, repo, tag = ilp._route_to_vulkan_prebuilt(host, FORK, "b9596-mix-abc", force_cpu = True)
|
||||
routed, repo, tag, _persist = ilp._route_to_vulkan_prebuilt(
|
||||
host, FORK, "b9596-mix-abc", force_cpu = True
|
||||
)
|
||||
assert repo == FORK
|
||||
assert tag == "b9596-mix-abc"
|
||||
assert routed is host
|
||||
|
|
@ -536,20 +556,20 @@ def test_route_to_vulkan_prebuilt_hidden_nvidia_not_rerouted():
|
|||
has_physical_nvidia = True,
|
||||
has_usable_nvidia = False,
|
||||
)
|
||||
_routed, repo, _tag = ilp._route_to_vulkan_prebuilt(host, FORK, "", force_cpu = False)
|
||||
_routed, repo, _tag, _persist = ilp._route_to_vulkan_prebuilt(host, FORK, "", force_cpu = False)
|
||||
assert repo == FORK
|
||||
|
||||
|
||||
def test_route_to_vulkan_prebuilt_rocm_host_not_rerouted():
|
||||
# An Intel iGPU alongside a usable ROCm GPU stays on its ROCm/fork path.
|
||||
host = _host(is_linux = True, is_x86_64 = True, has_intel_gpu = True, has_rocm = True)
|
||||
_routed, repo, _tag = ilp._route_to_vulkan_prebuilt(host, FORK, "", force_cpu = False)
|
||||
_routed, repo, _tag, _persist = ilp._route_to_vulkan_prebuilt(host, FORK, "", force_cpu = False)
|
||||
assert repo == FORK
|
||||
|
||||
|
||||
def test_route_to_vulkan_prebuilt_non_intel_unchanged():
|
||||
host = _host(is_linux = True, is_x86_64 = True)
|
||||
routed, repo, _tag = ilp._route_to_vulkan_prebuilt(host, FORK, "", force_cpu = False)
|
||||
routed, repo, _tag, _persist = ilp._route_to_vulkan_prebuilt(host, FORK, "", force_cpu = False)
|
||||
assert repo == FORK
|
||||
assert routed is host
|
||||
|
||||
|
|
@ -797,3 +817,800 @@ def test_detect_host_cim_rescues_exploding_registry(monkeypatch):
|
|||
)
|
||||
assert host.has_intel_gpu is True
|
||||
assert "powershell" in captured
|
||||
|
||||
|
||||
def _windows_amd_host(**overrides):
|
||||
defaults = dict(
|
||||
system = "Windows",
|
||||
machine = "amd64",
|
||||
is_windows = True,
|
||||
is_linux = False,
|
||||
is_macos = False,
|
||||
is_x86_64 = True,
|
||||
is_arm64 = False,
|
||||
nvidia_smi = None,
|
||||
driver_cuda_version = None,
|
||||
compute_caps = [],
|
||||
visible_cuda_devices = None,
|
||||
has_physical_nvidia = False,
|
||||
has_usable_nvidia = False,
|
||||
has_rocm = True,
|
||||
has_intel_gpu = False,
|
||||
)
|
||||
defaults.update(overrides)
|
||||
return ilp.HostInfo(**defaults)
|
||||
|
||||
|
||||
def test_route_to_vulkan_prebuilt_auto_fallback_for_legacy_amd_gfx():
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx803", rocm_gfx_targets = ["gfx803"])
|
||||
routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert repo == UPSTREAM
|
||||
assert persist == "vulkan"
|
||||
assert routed.has_intel_gpu is True
|
||||
assert routed.has_rocm is False
|
||||
|
||||
|
||||
def test_route_to_vulkan_prebuilt_keeps_hip_when_one_gpu_is_supported():
|
||||
host = _windows_amd_host(
|
||||
rocm_gfx_target = "gfx1201",
|
||||
rocm_gfx_targets = ["gfx1201", "gfx803"],
|
||||
)
|
||||
routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert routed is host
|
||||
assert repo == FORK
|
||||
assert persist is None
|
||||
|
||||
|
||||
def test_route_to_vulkan_prebuilt_auto_fallback_skips_hip_masked_hosts():
|
||||
# A HIP mask can hide a HIP-capable dGPU, but the Vulkan runtime honours none of them,
|
||||
# so auto-routing would let the installed backend grab the gfx1201 the user masked
|
||||
# off.
|
||||
host = _windows_amd_host(
|
||||
rocm_gfx_target = "gfx803",
|
||||
rocm_gfx_targets = ["gfx1201", "gfx803"],
|
||||
)
|
||||
routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert repo == FORK
|
||||
assert persist is None
|
||||
assert routed is host
|
||||
|
||||
|
||||
def test_route_to_vulkan_prebuilt_auto_fallback_when_no_amd_gpu_reaches_floor():
|
||||
# Every physical AMD device is below the floor, so no card can be exposed to HIP and
|
||||
# the #7357 auto-Vulkan fallback still fires.
|
||||
host = _windows_amd_host(
|
||||
rocm_gfx_target = "gfx900",
|
||||
rocm_gfx_targets = ["gfx803", "gfx900"],
|
||||
)
|
||||
routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert repo == UPSTREAM
|
||||
assert persist == "vulkan"
|
||||
assert routed.has_rocm is False
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"mask_env", ["HIP_VISIBLE_DEVICES", "ROCR_VISIBLE_DEVICES", "CUDA_VISIBLE_DEVICES"]
|
||||
)
|
||||
def test_auto_vulkan_declines_when_a_hip_device_mask_filtered_the_probe(mask_env, monkeypatch):
|
||||
# hipinfo is a HIP application, so under a mask rocm_gfx_targets is the VISIBLE set and
|
||||
# a HIP-capable card can be hidden entirely. "No AMD GPU here reaches the floor" is then
|
||||
# unprovable, and Vulkan honours none of these masks, so the auto fallback must decline
|
||||
# rather than hand it the reserved card.
|
||||
monkeypatch.setenv(mask_env, "1")
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx803", rocm_gfx_targets = ["gfx803"])
|
||||
assert ilp._should_auto_vulkan_for_amd_windows(host, FORK) is False
|
||||
routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert routed is host
|
||||
assert repo == FORK
|
||||
assert persist is None
|
||||
|
||||
|
||||
@pytest.mark.parametrize("mask_value", ["", " ", "-1"])
|
||||
def test_auto_vulkan_declines_when_the_mask_hides_every_amd_gpu(mask_value, monkeypatch):
|
||||
# An all-hiding mask is the strongest form of the same signal, not an exemption:
|
||||
# detect_host() resolves no arch under it, but a forwarded --rocm-gfx still reconstructs
|
||||
# one (setup infers it from the display-adapter name, which no HIP mask touches), so
|
||||
# auto-routing would hand Vulkan every AMD GPU the user hid from HIP.
|
||||
monkeypatch.setenv("HIP_VISIBLE_DEVICES", mask_value)
|
||||
host = _windows_amd_host(rocm_gfx_target = None, rocm_gfx_targets = [])
|
||||
host = ilp._apply_host_overrides(host, override_rocm_gfx = "gfx803")
|
||||
assert ilp._active_rocm_gfx_target(host) == "gfx803"
|
||||
assert ilp._should_auto_vulkan_for_amd_windows(host, FORK) is False
|
||||
routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert routed is host
|
||||
assert repo == FORK
|
||||
assert persist is None
|
||||
|
||||
|
||||
def test_hip_device_mask_check_is_presence_not_value(monkeypatch):
|
||||
# Presence is the whole test: any value means the HIP view is not the physical one, and
|
||||
# no value can be read as "the probe saw everything".
|
||||
assert ilp._hip_visible_device_mask_set() is False
|
||||
for value in ("", " ", "-1", "0", "1", "0,1"):
|
||||
monkeypatch.setenv("HIP_VISIBLE_DEVICES", value)
|
||||
assert ilp._hip_visible_device_mask_set() is True, value
|
||||
monkeypatch.delenv("HIP_VISIBLE_DEVICES")
|
||||
monkeypatch.setenv("ROCR_VISIBLE_DEVICES", "0")
|
||||
assert ilp._hip_visible_device_mask_set() is True
|
||||
monkeypatch.delenv("ROCR_VISIBLE_DEVICES")
|
||||
monkeypatch.setenv("CUDA_VISIBLE_DEVICES", "0")
|
||||
assert ilp._hip_visible_device_mask_set() is True
|
||||
|
||||
|
||||
def test_masked_probe_suppression_does_not_touch_non_amd_auto_paths(monkeypatch):
|
||||
# The mask says nothing about an Intel iGPU, whose Vulkan auto path is unrelated.
|
||||
monkeypatch.setenv("HIP_VISIBLE_DEVICES", "1")
|
||||
host = _host(
|
||||
system = "Windows",
|
||||
is_windows = True,
|
||||
has_intel_gpu = True,
|
||||
has_rocm = False,
|
||||
has_physical_nvidia = False,
|
||||
has_usable_nvidia = False,
|
||||
)
|
||||
_routed, repo, _tag, _persist = ilp._route_to_vulkan_prebuilt(
|
||||
host, FORK, "pin", force_cpu = False
|
||||
)
|
||||
assert repo == UPSTREAM
|
||||
|
||||
|
||||
def test_route_to_vulkan_prebuilt_hip_masked_host_still_honours_explicit_optin(monkeypatch):
|
||||
# The mask guard only suppresses the AUTOMATIC fallback; an explicit opt-in is the user
|
||||
# taking responsibility for the Vulkan device mask themselves.
|
||||
monkeypatch.delenv("UNSLOTH_FORCE_VULKAN", raising = False)
|
||||
host = _windows_amd_host(
|
||||
rocm_gfx_target = "gfx803",
|
||||
rocm_gfx_targets = ["gfx1201", "gfx803"],
|
||||
)
|
||||
_routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(
|
||||
host, FORK, "pin", force_cpu = False, llama_backend = "vulkan"
|
||||
)
|
||||
assert repo == UPSTREAM
|
||||
assert persist == "vulkan"
|
||||
|
||||
|
||||
def test_auto_vulkan_is_repository_specific_for_fork_only_gfx():
|
||||
# gfx1034 is served only by the fork's gfx103X bundle: ggml-org's windows-hip radeon
|
||||
# build does not target it and direct_upstream_release_plan() offers win-hip then CPU
|
||||
# with no Vulkan branch, so the predicate must answer per repo.
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx1034", rocm_gfx_targets = ["gfx1034"])
|
||||
assert ilp._should_auto_vulkan_for_amd_windows(host, FORK) is False
|
||||
assert ilp._should_auto_vulkan_for_amd_windows(host, UPSTREAM) is True
|
||||
# An arch upstream really does build stays on HIP for both repos.
|
||||
supported = _windows_amd_host(rocm_gfx_target = "gfx1100", rocm_gfx_targets = ["gfx1100"])
|
||||
assert ilp._should_auto_vulkan_for_amd_windows(supported, FORK) is False
|
||||
assert ilp._should_auto_vulkan_for_amd_windows(supported, UPSTREAM) is False
|
||||
# A family label is a bundle name, not an arch: upstream builds every member but
|
||||
# gfx1034 / gfx1103, and the label cannot say which card this is, so it stays on HIP
|
||||
# rather than moving the covered members onto Vulkan.
|
||||
family = _windows_amd_host(rocm_gfx_target = "gfx110X", rocm_gfx_targets = ["gfx110X"])
|
||||
assert ilp._should_auto_vulkan_for_amd_windows(family, UPSTREAM) is False
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"repo", ["acme/llama.cpp-mirror", "GGML-ORG/llama.cpp", "unslothAI/llama.cpp"]
|
||||
)
|
||||
def test_fork_only_gfx_coverage_is_not_granted_to_other_repos(repo):
|
||||
# Only the fork is planned from a manifest: resolve_simple_install_release_plans()
|
||||
# compares == DEFAULT_PUBLISHED_REPO and sends everything else, mirrors and differently
|
||||
# cased spellings alike, to direct_upstream_release_plan(). Granting a fork-only arch
|
||||
# coverage there lands it on win-hip-radeon or CPU instead of Vulkan, so the predicate
|
||||
# must gate on the fork rather than exempt one name.
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx1034", rocm_gfx_targets = ["gfx1034"])
|
||||
assert ilp._should_auto_vulkan_for_amd_windows(host, repo) is True
|
||||
supported = _windows_amd_host(rocm_gfx_target = "gfx1100", rocm_gfx_targets = ["gfx1100"])
|
||||
assert ilp._should_auto_vulkan_for_amd_windows(supported, repo) is False
|
||||
|
||||
|
||||
@pytest.mark.parametrize("repo", [None, ""])
|
||||
def test_empty_published_repo_gets_fork_coverage(repo):
|
||||
# Negative control: the resolver defaults an empty repo to the fork, so the predicate
|
||||
# must too, or the default install path loses its fork-only archs.
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx1034", rocm_gfx_targets = ["gfx1034"])
|
||||
assert ilp._should_auto_vulkan_for_amd_windows(host, repo) is False
|
||||
|
||||
|
||||
def test_upstream_windows_hip_targets_are_a_subset_of_the_combined_floor():
|
||||
# The floor must stay a superset, else auto-Vulkan steals a host upstream builds for.
|
||||
assert ilp.UPSTREAM_WINDOWS_HIP_GFX_TARGETS <= ilp.WINDOWS_HIP_PREBUILT_GFX_TARGETS
|
||||
# The fork-only extras are exactly the archs that must route to Vulkan upstream.
|
||||
assert ilp.WINDOWS_HIP_PREBUILT_GFX_TARGETS - ilp.UPSTREAM_WINDOWS_HIP_GFX_TARGETS == {
|
||||
"gfx908",
|
||||
"gfx90a",
|
||||
"gfx1034",
|
||||
"gfx1103",
|
||||
}
|
||||
|
||||
|
||||
def test_route_to_vulkan_prebuilt_unknown_gfx_does_not_auto_fallback():
|
||||
host = _windows_amd_host(
|
||||
has_rocm = True,
|
||||
rocm_gfx_target = None,
|
||||
rocm_gfx_targets = [],
|
||||
)
|
||||
routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert routed is host
|
||||
assert repo == FORK
|
||||
assert persist is None
|
||||
|
||||
|
||||
def test_route_to_vulkan_prebuilt_family_gfx_token_keeps_rocm():
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx110X", rocm_gfx_targets = ["gfx110X"])
|
||||
routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert routed is host
|
||||
assert repo == FORK
|
||||
assert persist is None
|
||||
|
||||
|
||||
def test_route_to_vulkan_prebuilt_gfx1103_keeps_rocm():
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx1103", rocm_gfx_targets = ["gfx1103"])
|
||||
routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert routed is host
|
||||
assert repo == FORK
|
||||
assert persist is None
|
||||
|
||||
|
||||
def test_route_to_vulkan_prebuilt_gfx1034_keeps_rocm():
|
||||
# gfx1034 (RX 6500/6400-class) is covered by the fork's gfx103X bundle.
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx1034", rocm_gfx_targets = ["gfx1034"])
|
||||
routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert routed is host
|
||||
assert repo == FORK
|
||||
assert persist is None
|
||||
|
||||
|
||||
def test_route_to_vulkan_prebuilt_explicit_opt_in_on_mixed_amd(monkeypatch):
|
||||
monkeypatch.setenv("UNSLOTH_LLAMA_BACKEND", "vulkan")
|
||||
host = _windows_amd_host(
|
||||
rocm_gfx_target = "gfx1201",
|
||||
rocm_gfx_targets = ["gfx1201", "gfx803"],
|
||||
)
|
||||
routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert repo == UPSTREAM
|
||||
assert persist == "vulkan"
|
||||
assert routed.has_rocm is False
|
||||
|
||||
|
||||
def test_direct_upstream_windows_amd_legacy_gfx_routes_to_vulkan():
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx803", rocm_gfx_targets = ["gfx803"])
|
||||
routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
rel = _upstream_release(
|
||||
"b9925",
|
||||
[
|
||||
"llama-b9925-bin-win-hip-radeon-x64.zip",
|
||||
"llama-b9925-bin-win-vulkan-x64.zip",
|
||||
"llama-b9925-bin-win-cpu-x64.zip",
|
||||
],
|
||||
)
|
||||
plan = ilp.direct_upstream_release_plan(rel, routed, repo, "latest")
|
||||
assert persist == "vulkan"
|
||||
assert plan.attempts[0].install_kind == "windows-vulkan"
|
||||
|
||||
|
||||
def test_llama_backend_env_requests_vulkan(monkeypatch):
|
||||
assert ilp.llama_backend_from_env() is None
|
||||
monkeypatch.setenv("UNSLOTH_LLAMA_BACKEND", "vulkan")
|
||||
assert ilp.llama_backend_from_env() == "vulkan"
|
||||
assert ilp.force_vulkan_requested() is True
|
||||
|
||||
|
||||
def test_llama_cpp_backend_env_does_not_trigger_vulkan(monkeypatch):
|
||||
# UNSLOTH_LLAMA_CPP_BACKEND is a separate setup variable (auto/cpu) whose other values
|
||||
# setup warns about and ignores, so reading it here would opt in behind that warning.
|
||||
monkeypatch.delenv("UNSLOTH_LLAMA_BACKEND", raising = False)
|
||||
monkeypatch.delenv("UNSLOTH_FORCE_VULKAN", raising = False)
|
||||
monkeypatch.setenv("UNSLOTH_LLAMA_CPP_BACKEND", "vulkan")
|
||||
assert ilp.llama_backend_from_env() is None
|
||||
assert ilp.force_vulkan_requested() is False
|
||||
|
||||
|
||||
def test_route_to_vulkan_prebuilt_hidden_physical_nvidia_amd_not_rerouted():
|
||||
# Vulkan ignores CUDA_VISIBLE_DEVICES, so a CUDA-masked NVIDIA card next to a legacy
|
||||
# AMD gfx must not auto-route: Vulkan could grab the reserved NVIDIA GPU.
|
||||
host = _windows_amd_host(
|
||||
rocm_gfx_target = "gfx803",
|
||||
rocm_gfx_targets = ["gfx803"],
|
||||
has_physical_nvidia = True,
|
||||
has_usable_nvidia = False,
|
||||
)
|
||||
routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert routed is host
|
||||
assert repo == FORK
|
||||
assert persist is None
|
||||
|
||||
|
||||
def test_route_to_vulkan_prebuilt_explicit_opt_in_overrides_hidden_nvidia(monkeypatch):
|
||||
# The physical-NVIDIA guard only gates the AMD auto path; an explicit opt-in wins.
|
||||
monkeypatch.setenv("UNSLOTH_LLAMA_BACKEND", "vulkan")
|
||||
host = _windows_amd_host(
|
||||
rocm_gfx_target = "gfx803",
|
||||
rocm_gfx_targets = ["gfx803"],
|
||||
has_physical_nvidia = True,
|
||||
has_usable_nvidia = False,
|
||||
)
|
||||
routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert repo == UPSTREAM
|
||||
assert persist == "vulkan"
|
||||
|
||||
|
||||
# The gfx archs the fork's llama-prebuilt-manifest.json maps to a windows-rocm bundle.
|
||||
# Static because parametrisation happens at import time and the routing tests below must
|
||||
# stay offline; the guard further down re-derives it from the published manifest and fails
|
||||
# on drift, so this is a checked mirror, not a second source of truth.
|
||||
_FORK_WINDOWS_ROCM_GFX = (
|
||||
"gfx908",
|
||||
"gfx90a",
|
||||
"gfx1030",
|
||||
"gfx1031",
|
||||
"gfx1032",
|
||||
"gfx1034",
|
||||
"gfx1100",
|
||||
"gfx1101",
|
||||
"gfx1102",
|
||||
"gfx1103",
|
||||
"gfx1150",
|
||||
"gfx1151",
|
||||
"gfx1200",
|
||||
"gfx1201",
|
||||
)
|
||||
|
||||
|
||||
def _published_fork_windows_rocm_artifacts():
|
||||
"""The fork's windows-rocm artifact records, read the way an install reads them.
|
||||
|
||||
_download_host_resolved_release is the path a default fork install takes first: it
|
||||
resolves the latest release off the download host and hands llama-prebuilt-manifest.json
|
||||
to parse_published_release_bundle, so these are the very records
|
||||
published_rocm_choice_for_host later matches a host gfx against. No api.github.com call,
|
||||
hence no shared rate-limit bucket to exhaust.
|
||||
|
||||
The manifest ships only as a release asset and nothing in-tree mirrors it, so this is
|
||||
the one honest source. Only OSError and the release-side PrebuiltFallback become a skip,
|
||||
so an offline run stays quiet while a manifest that fetches but no longer parses still
|
||||
fails loudly."""
|
||||
try:
|
||||
resolved = ilp._download_host_resolved_release(FORK)
|
||||
except OSError as exc:
|
||||
pytest.skip(f"{FORK} release manifest unreachable: {exc}")
|
||||
except ilp.PrebuiltFallback as exc:
|
||||
pytest.skip(f"{FORK} latest release was rejected before its manifest parsed: {exc}")
|
||||
if resolved is None:
|
||||
pytest.skip(f"{FORK} published no resolvable latest release")
|
||||
tag = resolved.bundle.release_tag
|
||||
artifacts = [
|
||||
artifact
|
||||
for artifact in resolved.bundle.artifacts
|
||||
if artifact.install_kind == "windows-rocm"
|
||||
]
|
||||
assert artifacts, f"{FORK}@{tag} manifest listed no windows-rocm artifacts"
|
||||
return tag, artifacts
|
||||
|
||||
|
||||
def test_windows_hip_gfx_floor_covers_every_fork_windows_rocm_bundle():
|
||||
# Derived from the published manifest, not a second literal: a gfx the fork builds but
|
||||
# the floor omits bypasses the fork manifest, downgrading a hash-approved windows-rocm
|
||||
# bundle to an unhashed upstream Vulkan build. A newly published arch must redden here.
|
||||
tag, artifacts = _published_fork_windows_rocm_artifacts()
|
||||
# published_rocm_choice_for_host serves a bundle on a concrete mapped_targets entry or on
|
||||
# the umbrella gfx_target itself, so both spellings must clear a floor. A gfx_target
|
||||
# absent from its own mapped_targets is the family label (gfx110X); one present in it is
|
||||
# a standalone bundle (gfx908) already counted as concrete.
|
||||
concrete = {target.lower() for artifact in artifacts for target in artifact.mapped_targets}
|
||||
labels = {
|
||||
artifact.gfx_target.lower()
|
||||
for artifact in artifacts
|
||||
if artifact.gfx_target and artifact.gfx_target.lower() not in concrete
|
||||
}
|
||||
unfloored = sorted(concrete - ilp.WINDOWS_HIP_PREBUILT_GFX_TARGETS)
|
||||
assert (
|
||||
not unfloored
|
||||
), f"auto-Vulkan would steal windows-rocm archs published in {FORK}@{tag}: {unfloored}"
|
||||
unlabelled = sorted(labels - ilp.WINDOWS_ROCM_FAMILY_GFX_LABELS)
|
||||
assert not unlabelled, (
|
||||
f"update markers forward family labels {FORK}@{tag} publishes but "
|
||||
f"WINDOWS_ROCM_FAMILY_GFX_LABELS omits: {unlabelled}"
|
||||
)
|
||||
# Keep the import-time tuple the offline routing tests parametrise on an exact mirror.
|
||||
assert set(_FORK_WINDOWS_ROCM_GFX) == concrete, (
|
||||
f"_FORK_WINDOWS_ROCM_GFX drifted from {FORK}@{tag}: "
|
||||
f"gained {sorted(concrete - set(_FORK_WINDOWS_ROCM_GFX))}, "
|
||||
f"lost {sorted(set(_FORK_WINDOWS_ROCM_GFX) - concrete)}"
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("gfx", _FORK_WINDOWS_ROCM_GFX)
|
||||
def test_route_to_vulkan_prebuilt_keeps_every_fork_windows_rocm_arch(gfx, monkeypatch):
|
||||
# No ambient opt-in: this asserts the AUTO path leaves covered archs alone.
|
||||
monkeypatch.delenv("UNSLOTH_LLAMA_BACKEND", raising = False)
|
||||
monkeypatch.delenv("UNSLOTH_FORCE_VULKAN", raising = False)
|
||||
host = _windows_amd_host(rocm_gfx_target = gfx, rocm_gfx_targets = [gfx])
|
||||
routed, repo, tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert routed is host
|
||||
assert (repo, tag) == (FORK, "pin")
|
||||
assert persist is None
|
||||
|
||||
|
||||
def test_forwarded_gfx_does_not_undo_visible_device_auto_vulkan(monkeypatch):
|
||||
# Mixed-AMD Windows host: GPU 0 = gfx1100 (HIP prebuilt exists), GPU 1 = gfx1010 (none).
|
||||
# Under CUDA_VISIBLE_DEVICES=1 setup.ps1 still resolves GPU 0 and forwards gfx1100, but
|
||||
# detect_host() resolved the visible gfx1010, so folding the forward in must not
|
||||
# reinstate gfx1100 and install a HIP bundle the visible GPU cannot run.
|
||||
monkeypatch.delenv("UNSLOTH_LLAMA_BACKEND", raising = False)
|
||||
monkeypatch.delenv("UNSLOTH_FORCE_VULKAN", raising = False)
|
||||
monkeypatch.delenv("UNSLOTH_ROCM_GFX_ARCH", raising = False)
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx1010", rocm_gfx_targets = ["gfx1100", "gfx1010"])
|
||||
host = ilp._apply_host_overrides(host, override_rocm_gfx = "gfx1100")
|
||||
assert ilp._active_rocm_gfx_target(host) == "gfx1010"
|
||||
assert host.rocm_gfx_targets == ["gfx1100", "gfx1010"]
|
||||
# gfx1100 is masked off, not absent, and Vulkan does not honour the HIP mask, so the
|
||||
# automatic fallback stays off and the HIP / fork path is kept.
|
||||
assert ilp._should_auto_vulkan_for_amd_windows(host, FORK) is False
|
||||
_routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert repo == FORK
|
||||
assert persist is None
|
||||
|
||||
|
||||
def test_forwarded_gfx_absent_from_probe_keeps_the_physical_hip_card(monkeypatch):
|
||||
# Mixed-AMD Windows host: GPU 0 = gfx1100 (HIP prebuilt exists), GPU 1 = gfx803 (below
|
||||
# the floor). CUDA_VISIBLE_DEVICES=1 reserves the gfx1100, so detect_host() picks gfx803
|
||||
# as active but still reports both cards, and setup forwards a third arch the probe never
|
||||
# saw (a stale env var, or name inference reading the other card). That forward selects
|
||||
# the HIP target but must not delete the probe's inventory, or the floor check concludes
|
||||
# no AMD GPU here reaches HIP and auto-routes to Vulkan, which ignores the HIP mask and
|
||||
# enumerates the reserved gfx1100.
|
||||
monkeypatch.delenv("UNSLOTH_LLAMA_BACKEND", raising = False)
|
||||
monkeypatch.delenv("UNSLOTH_FORCE_VULKAN", raising = False)
|
||||
monkeypatch.delenv("UNSLOTH_ROCM_GFX_ARCH", raising = False)
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx803", rocm_gfx_targets = ["gfx1100", "gfx803"])
|
||||
host = ilp._apply_host_overrides(host, override_rocm_gfx = "gfx900")
|
||||
assert ilp._active_rocm_gfx_target(host) == "gfx900"
|
||||
assert host.rocm_gfx_targets == ["gfx1100", "gfx803", "gfx900"]
|
||||
assert ilp._should_auto_vulkan_for_amd_windows(host, FORK) is False
|
||||
_routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert repo == FORK
|
||||
assert persist is None
|
||||
|
||||
|
||||
def test_forwarded_gfx_absent_from_probe_keeps_a_single_probed_hip_card(monkeypatch):
|
||||
# Same rule on a single-GPU box: a stale below-floor forward over a probe-confirmed
|
||||
# gfx1100 must not auto-route that machine to Vulkan.
|
||||
monkeypatch.delenv("UNSLOTH_LLAMA_BACKEND", raising = False)
|
||||
monkeypatch.delenv("UNSLOTH_FORCE_VULKAN", raising = False)
|
||||
monkeypatch.delenv("UNSLOTH_ROCM_GFX_ARCH", raising = False)
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx1100", rocm_gfx_targets = ["gfx1100"])
|
||||
host = ilp._apply_host_overrides(host, override_rocm_gfx = "gfx803")
|
||||
assert ilp._active_rocm_gfx_target(host) == "gfx803"
|
||||
assert host.rocm_gfx_targets == ["gfx1100", "gfx803"]
|
||||
assert ilp._should_auto_vulkan_for_amd_windows(host, FORK) is False
|
||||
|
||||
|
||||
def test_forwarded_gfx_absent_from_probe_still_allows_explicit_vulkan(monkeypatch):
|
||||
# The physical-inventory rule gates the AUTO path only; naming the backend wins.
|
||||
monkeypatch.setenv("UNSLOTH_LLAMA_BACKEND", "vulkan")
|
||||
monkeypatch.delenv("UNSLOTH_ROCM_GFX_ARCH", raising = False)
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx1100", rocm_gfx_targets = ["gfx1100"])
|
||||
host = ilp._apply_host_overrides(host, override_rocm_gfx = "gfx803")
|
||||
_routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert repo == UPSTREAM
|
||||
assert persist == "vulkan"
|
||||
|
||||
|
||||
def test_forwarded_gfx_on_unprobed_host_still_auto_vulkans(monkeypatch):
|
||||
# Negative control: a driver-only AMD host runs no successful probe (no hipinfo, amd-smi
|
||||
# suppressed), so --rocm-gfx is the ONLY source of the arch and there is no inventory to
|
||||
# preserve. This is the #7357 path the feature exists for; it must still reach Vulkan.
|
||||
monkeypatch.delenv("UNSLOTH_LLAMA_BACKEND", raising = False)
|
||||
monkeypatch.delenv("UNSLOTH_FORCE_VULKAN", raising = False)
|
||||
monkeypatch.delenv("UNSLOTH_ROCM_GFX_ARCH", raising = False)
|
||||
host = _windows_amd_host(rocm_gfx_target = None, rocm_gfx_targets = [])
|
||||
host = ilp._apply_host_overrides(host, override_rocm_gfx = "gfx803")
|
||||
assert host.rocm_gfx_targets == ["gfx803"]
|
||||
assert ilp._should_auto_vulkan_for_amd_windows(host, FORK) is True
|
||||
_routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert repo == UPSTREAM
|
||||
assert persist == "vulkan"
|
||||
|
||||
|
||||
def test_forwarded_gfx_still_fills_an_unprobed_arch(monkeypatch):
|
||||
# Negative control: on an amd-smi-only host detect_host() reports no arch, so the
|
||||
# forward is the only source and must still apply.
|
||||
monkeypatch.delenv("UNSLOTH_LLAMA_BACKEND", raising = False)
|
||||
monkeypatch.delenv("UNSLOTH_FORCE_VULKAN", raising = False)
|
||||
monkeypatch.delenv("UNSLOTH_ROCM_GFX_ARCH", raising = False)
|
||||
host = _windows_amd_host(rocm_gfx_target = None, rocm_gfx_targets = [])
|
||||
host = ilp._apply_host_overrides(host, override_rocm_gfx = "gfx1151")
|
||||
assert ilp._active_rocm_gfx_target(host) == "gfx1151"
|
||||
assert ilp._should_auto_vulkan_for_amd_windows(host) is False
|
||||
_routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert repo == FORK
|
||||
assert persist is None
|
||||
|
||||
|
||||
def test_llama_backend_hip_opts_out_of_auto_vulkan(monkeypatch):
|
||||
# hip names a backend, so it keeps the fork path even on an auto-fallback arch.
|
||||
monkeypatch.setenv("UNSLOTH_LLAMA_BACKEND", "hip")
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx803", rocm_gfx_targets = ["gfx803"])
|
||||
routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert routed is host
|
||||
assert repo == FORK
|
||||
assert persist is None
|
||||
assert ilp.force_vulkan_requested() is False
|
||||
|
||||
|
||||
def test_explicit_backend_beats_legacy_force_vulkan(monkeypatch):
|
||||
# A stale UNSLOTH_FORCE_VULKAN must not overrule UNSLOTH_LLAMA_BACKEND=rocm (== hip).
|
||||
monkeypatch.setenv("UNSLOTH_FORCE_VULKAN", "1")
|
||||
monkeypatch.setenv("UNSLOTH_LLAMA_BACKEND", "rocm")
|
||||
assert ilp.resolved_llama_backend() == "hip"
|
||||
assert ilp.force_vulkan_requested() is False
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx1100", rocm_gfx_targets = ["gfx1100"])
|
||||
_routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert repo == FORK
|
||||
assert persist is None
|
||||
|
||||
|
||||
def test_unknown_llama_backend_value_falls_through_to_legacy_flag(monkeypatch):
|
||||
# An unrecognised value is ignored, not an error, so the legacy flag still works.
|
||||
monkeypatch.setenv("UNSLOTH_LLAMA_BACKEND", "banana")
|
||||
assert ilp.resolved_llama_backend() is None
|
||||
assert ilp.force_vulkan_requested() is False
|
||||
monkeypatch.setenv("UNSLOTH_FORCE_VULKAN", "1")
|
||||
assert ilp.force_vulkan_requested() is True
|
||||
|
||||
|
||||
def test_llama_backend_flag_beats_conflicting_env(monkeypatch):
|
||||
# --llama-backend is the caller's explicit request and outranks the env.
|
||||
monkeypatch.setenv("UNSLOTH_LLAMA_BACKEND", "hip")
|
||||
assert ilp.force_vulkan_requested("vulkan") is True
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx1100", rocm_gfx_targets = ["gfx1100"])
|
||||
_routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(
|
||||
host, FORK, "pin", force_cpu = False, llama_backend = "vulkan"
|
||||
)
|
||||
assert repo == UPSTREAM
|
||||
assert persist == "vulkan"
|
||||
|
||||
|
||||
def _windows_arm64_host(**overrides):
|
||||
defaults = dict(
|
||||
system = "Windows",
|
||||
machine = "ARM64",
|
||||
is_windows = True,
|
||||
is_linux = False,
|
||||
is_macos = False,
|
||||
is_x86_64 = False,
|
||||
is_arm64 = True,
|
||||
nvidia_smi = None,
|
||||
driver_cuda_version = None,
|
||||
compute_caps = [],
|
||||
visible_cuda_devices = None,
|
||||
has_physical_nvidia = False,
|
||||
has_usable_nvidia = False,
|
||||
has_rocm = False,
|
||||
has_intel_gpu = False,
|
||||
)
|
||||
defaults.update(overrides)
|
||||
return ilp.HostInfo(**defaults)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"env, flag",
|
||||
[
|
||||
({"UNSLOTH_LLAMA_BACKEND": "vulkan"}, None),
|
||||
({"UNSLOTH_FORCE_VULKAN": "1"}, None),
|
||||
({}, "vulkan"),
|
||||
],
|
||||
)
|
||||
def test_vulkan_opt_in_ignored_on_windows_arm64(monkeypatch, env, flag):
|
||||
# Upstream builds win-vulkan for x64 only (arm64 gets CPU + opencl-adreno), so rewriting
|
||||
# the host would only swap the published arm64 bundle for the upstream CPU one.
|
||||
for name, value in env.items():
|
||||
monkeypatch.setenv(name, value)
|
||||
host = _windows_arm64_host()
|
||||
routed, repo, tag, persist = ilp._route_to_vulkan_prebuilt(
|
||||
host, FORK, "pin", force_cpu = False, llama_backend = flag
|
||||
)
|
||||
assert routed is host
|
||||
assert (repo, tag) == (FORK, "pin")
|
||||
assert persist is None
|
||||
|
||||
|
||||
def test_vulkan_opt_in_still_routes_on_windows_x64(monkeypatch):
|
||||
# Negative control for the arm64 guard: x64 keeps its Vulkan routing.
|
||||
monkeypatch.setenv("UNSLOTH_LLAMA_BACKEND", "vulkan")
|
||||
host = _windows_amd_host(rocm_gfx_target = "gfx1100", rocm_gfx_targets = ["gfx1100"])
|
||||
_routed, repo, _tag, persist = ilp._route_to_vulkan_prebuilt(host, FORK, "pin", force_cpu = False)
|
||||
assert repo == UPSTREAM
|
||||
assert persist == "vulkan"
|
||||
|
||||
|
||||
def _choice(install_kind, name = "asset.zip"):
|
||||
return ilp.AssetChoice(
|
||||
repo = UPSTREAM,
|
||||
tag = "b9925",
|
||||
name = name,
|
||||
url = f"https://example/{name}",
|
||||
source_label = "upstream",
|
||||
install_kind = install_kind,
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("kind", ["windows-vulkan", "linux-vulkan"])
|
||||
def test_persisted_llama_backend_keeps_vulkan_for_a_vulkan_bundle(kind):
|
||||
assert ilp.persisted_llama_backend("vulkan", _choice(kind)) == "vulkan"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("kind", ["windows-arm64", "windows-cpu", "linux-cpu", "windows-rocm"])
|
||||
def test_persisted_llama_backend_drops_vulkan_for_a_non_vulkan_bundle(kind):
|
||||
# _plan_llama_phase re-asserts the marker's backend on every later update, so a Vulkan
|
||||
# request that fell through to CPU must not leave a marker claiming Vulkan.
|
||||
assert ilp.persisted_llama_backend("vulkan", _choice(kind)) is None
|
||||
|
||||
|
||||
def test_persisted_llama_backend_passes_none_through():
|
||||
assert ilp.persisted_llama_backend(None, _choice("windows-vulkan")) is None
|
||||
|
||||
|
||||
def test_marker_records_no_backend_when_vulkan_fell_back_to_cpu(tmp_path):
|
||||
# End to end over write_prebuilt_metadata: describe the CPU attempt that actually won,
|
||||
# so the next update re-detects instead of re-asserting Vulkan forever.
|
||||
checksums = ilp.ApprovedReleaseChecksums(
|
||||
repo = UPSTREAM,
|
||||
release_tag = "b9925",
|
||||
upstream_tag = "b9925",
|
||||
source_repo = UPSTREAM,
|
||||
source_repo_url = f"https://github.com/{UPSTREAM}",
|
||||
)
|
||||
cpu = _choice("windows-arm64", "llama-b9925-bin-win-cpu-arm64.zip")
|
||||
ilp.write_prebuilt_metadata(
|
||||
tmp_path,
|
||||
requested_tag = "latest",
|
||||
llama_tag = "b9925",
|
||||
release_tag = "b9925",
|
||||
choice = cpu,
|
||||
approved_checksums = checksums,
|
||||
prebuilt_fallback_used = False,
|
||||
llama_backend = "vulkan",
|
||||
)
|
||||
marker = json.loads((tmp_path / "UNSLOTH_PREBUILT_INFO.json").read_text())
|
||||
assert marker["asset"] == "llama-b9925-bin-win-cpu-arm64.zip"
|
||||
assert marker["llama_backend"] is None
|
||||
|
||||
vulkan = _choice("windows-vulkan", "llama-b9925-bin-win-vulkan-x64.zip")
|
||||
ilp.write_prebuilt_metadata(
|
||||
tmp_path,
|
||||
requested_tag = "latest",
|
||||
llama_tag = "b9925",
|
||||
release_tag = "b9925",
|
||||
choice = vulkan,
|
||||
approved_checksums = checksums,
|
||||
prebuilt_fallback_used = False,
|
||||
llama_backend = "vulkan",
|
||||
)
|
||||
marker = json.loads((tmp_path / "UNSLOTH_PREBUILT_INFO.json").read_text())
|
||||
assert marker["llama_backend"] == "vulkan"
|
||||
|
||||
|
||||
# UNSLOTH_LLAMA_CPP_BACKEND (setup.sh/setup.ps1, "auto"|"cpu") and
|
||||
# UNSLOTH_LLAMA_BACKEND (this module, a backend name) are different variables at
|
||||
# different layers, and both accept "cpu". setup translates its own =cpu into
|
||||
# --force-cpu to pin the CPU-only bundle on a GPU host, which is what keeps Intel
|
||||
# iGPU Vulkan crashes away (#7213). Vulkan is opt-in here, so no trigger it adds
|
||||
# may outrank that flag on any host.
|
||||
_SIM_PLATFORMS = {
|
||||
# WSL presents as Linux to this resolver, so it rides the Linux row.
|
||||
"Linux": dict(
|
||||
system = "Linux",
|
||||
is_windows = False,
|
||||
is_linux = True,
|
||||
is_macos = False,
|
||||
machine = "x86_64",
|
||||
is_x86_64 = True,
|
||||
is_arm64 = False,
|
||||
),
|
||||
"Windows": dict(
|
||||
system = "Windows",
|
||||
is_windows = True,
|
||||
is_linux = False,
|
||||
is_macos = False,
|
||||
machine = "amd64",
|
||||
is_x86_64 = True,
|
||||
is_arm64 = False,
|
||||
),
|
||||
"macOS": dict(
|
||||
system = "Darwin",
|
||||
is_windows = False,
|
||||
is_linux = False,
|
||||
is_macos = True,
|
||||
machine = "arm64",
|
||||
is_x86_64 = False,
|
||||
is_arm64 = True,
|
||||
),
|
||||
}
|
||||
_SIM_GPUS = {
|
||||
"nvidia": dict(
|
||||
has_physical_nvidia = True,
|
||||
has_usable_nvidia = True,
|
||||
has_rocm = False,
|
||||
has_intel_gpu = False,
|
||||
nvidia_smi = "/usr/bin/nvidia-smi",
|
||||
driver_cuda_version = "12.4",
|
||||
compute_caps = ["8.9"],
|
||||
),
|
||||
"amd": dict(
|
||||
has_physical_nvidia = False,
|
||||
has_usable_nvidia = False,
|
||||
has_rocm = True,
|
||||
has_intel_gpu = False,
|
||||
nvidia_smi = None,
|
||||
driver_cuda_version = None,
|
||||
compute_caps = [],
|
||||
rocm_gfx_target = "gfx803",
|
||||
rocm_gfx_targets = ["gfx803"],
|
||||
),
|
||||
"intel": dict(
|
||||
has_physical_nvidia = False,
|
||||
has_usable_nvidia = False,
|
||||
has_rocm = False,
|
||||
has_intel_gpu = True,
|
||||
nvidia_smi = None,
|
||||
driver_cuda_version = None,
|
||||
compute_caps = [],
|
||||
),
|
||||
"cpu_only": dict(
|
||||
has_physical_nvidia = False,
|
||||
has_usable_nvidia = False,
|
||||
has_rocm = False,
|
||||
has_intel_gpu = False,
|
||||
nvidia_smi = None,
|
||||
driver_cuda_version = None,
|
||||
compute_caps = [],
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def _sim_host(platform_name, gpu_name):
|
||||
base = dict(visible_cuda_devices = None)
|
||||
base.update(_SIM_PLATFORMS[platform_name])
|
||||
base.update(_SIM_GPUS[gpu_name])
|
||||
return ilp.HostInfo(**base)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("platform_name", sorted(_SIM_PLATFORMS))
|
||||
@pytest.mark.parametrize("gpu_name", sorted(_SIM_GPUS))
|
||||
@pytest.mark.parametrize("backend_env", [None, "vulkan", "hip", "rocm", "cpu"])
|
||||
def test_forced_cpu_outranks_every_vulkan_trigger(
|
||||
monkeypatch, platform_name, gpu_name, backend_env
|
||||
):
|
||||
"""A deliberate CPU install stays CPU on every host, whatever asks for Vulkan."""
|
||||
monkeypatch.delenv("UNSLOTH_FORCE_VULKAN", raising = False)
|
||||
if backend_env is None:
|
||||
monkeypatch.delenv("UNSLOTH_LLAMA_BACKEND", raising = False)
|
||||
else:
|
||||
monkeypatch.setenv("UNSLOTH_LLAMA_BACKEND", backend_env)
|
||||
# The legacy switch too, so a stale one cannot smuggle Vulkan past --force-cpu.
|
||||
monkeypatch.setenv("UNSLOTH_FORCE_VULKAN", "1")
|
||||
|
||||
repo, tag = "unslothai/llama.cpp-prebuilt", "latest"
|
||||
_, out_repo, _, persist = ilp._route_to_vulkan_prebuilt(
|
||||
_sim_host(platform_name, gpu_name),
|
||||
repo,
|
||||
tag,
|
||||
force_cpu = True,
|
||||
llama_backend = "vulkan",
|
||||
)
|
||||
assert out_repo == repo, (platform_name, gpu_name, backend_env)
|
||||
assert persist is None, (platform_name, gpu_name, backend_env)
|
||||
|
||||
|
||||
def test_the_forced_cpu_guard_is_not_vacuous():
|
||||
"""The same host DOES take Vulkan once the CPU pin is gone, or the check above
|
||||
would pass on a resolver that had stopped routing to Vulkan entirely."""
|
||||
repo, tag = "unslothai/llama.cpp-prebuilt", "latest"
|
||||
_, out_repo, _, persist = ilp._route_to_vulkan_prebuilt(
|
||||
_sim_host("Linux", "amd"),
|
||||
repo,
|
||||
tag,
|
||||
force_cpu = False,
|
||||
llama_backend = "vulkan",
|
||||
)
|
||||
assert out_repo != repo or persist == "vulkan"
|
||||
|
|
|
|||
|
|
@ -473,6 +473,7 @@ def test_start_update_preserves_vulkan_via_env(monkeypatch, tmp_path):
|
|||
monkeypatch.setattr(freshness, "_fetch_latest_release_tag", lambda repo, timeout = 5.0: "b9518")
|
||||
|
||||
def _on_start(cmd):
|
||||
captured["cmd"] = cmd
|
||||
_write_install(
|
||||
install_dir,
|
||||
"b9518",
|
||||
|
|
@ -480,6 +481,7 @@ def test_start_update_preserves_vulkan_via_env(monkeypatch, tmp_path):
|
|||
asset = "llama-b9518-bin-ubuntu-vulkan-x64.tar.gz",
|
||||
)
|
||||
|
||||
captured: dict = {}
|
||||
popen_kwargs: dict = {}
|
||||
_patch_installer_popen(
|
||||
monkeypatch,
|
||||
|
|
@ -497,6 +499,8 @@ def test_start_update_preserves_vulkan_via_env(monkeypatch, tmp_path):
|
|||
time.sleep(0.05)
|
||||
assert job["state"] == "success", job
|
||||
assert popen_kwargs["env"]["UNSLOTH_FORCE_VULKAN"] == "1"
|
||||
assert popen_kwargs["env"]["UNSLOTH_LLAMA_BACKEND"] == "vulkan"
|
||||
assert "--llama-backend" in captured["cmd"] and "vulkan" in captured["cmd"]
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
|
|
|
|||
|
|
@ -403,6 +403,7 @@ def _run_llama_phase(
|
|||
pin_release_tag: Optional[str],
|
||||
set_progress,
|
||||
force_cpu: bool = False,
|
||||
llama_backend: Optional[str] = None,
|
||||
) -> dict:
|
||||
"""The llama phase of a chained update: put the backend into a maintenance
|
||||
state, run the installer for the latest prebuilt, then refresh caches so the
|
||||
|
|
@ -454,14 +455,15 @@ def _run_llama_phase(
|
|||
# updates. A natural fallback (or a legacy marker without the flag) heals to GPU (#6097).
|
||||
if force_cpu:
|
||||
cmd.append("--force-cpu")
|
||||
if llama_backend == "vulkan":
|
||||
cmd.extend(["--llama-backend", "vulkan"])
|
||||
logger.info("llama update: installing", cmd = " ".join(cmd))
|
||||
env = dict(os.environ, UNSLOTH_PROGRESS_PERCENT_STEP = "5")
|
||||
# Preserve a Vulkan install across updates: detect_host on a CUDA/ROCm
|
||||
# box would otherwise re-route and silently replace the Vulkan build.
|
||||
# Re-assert it via the same env flag setup uses (mirrors
|
||||
# _rocm_install_args).
|
||||
if asset and "vulkan" in asset.lower():
|
||||
# Preserve a Vulkan install across updates: detect_host on a CUDA/ROCm box would
|
||||
# otherwise re-route and silently replace it. Re-assert via setup's env/CLI flags.
|
||||
if llama_backend == "vulkan" or (asset and "vulkan" in asset.lower()):
|
||||
env["UNSLOTH_FORCE_VULKAN"] = "1"
|
||||
env["UNSLOTH_LLAMA_BACKEND"] = "vulkan"
|
||||
_flow.stream_installer(
|
||||
cmd,
|
||||
env,
|
||||
|
|
@ -578,6 +580,9 @@ def _plan_llama_phase() -> dict:
|
|||
from_tag = marker.get("tag") or marker.get("release_tag")
|
||||
asset = marker.get("asset")
|
||||
force_cpu = bool(marker.get("force_cpu"))
|
||||
llama_backend = marker.get("llama_backend")
|
||||
if llama_backend == "vulkan" or (asset and "vulkan" in str(asset).lower()):
|
||||
llama_backend = "vulkan"
|
||||
# Install exactly the release the banner offered: the installer's own
|
||||
# "latest" is commit-date ordered and can lag the published_at pick
|
||||
# above, reinstalling the current build in a loop (the #6219 class).
|
||||
|
|
@ -621,6 +626,7 @@ def _plan_llama_phase() -> dict:
|
|||
asset = (res or {}).get("asset")
|
||||
# Source builds carry no forced-CPU marker, so nothing to preserve here.
|
||||
force_cpu = False
|
||||
llama_backend = None
|
||||
# No pin: source-build detection resolves via --resolve-prebuilt latest,
|
||||
# the same resolver the unpinned apply uses, so the two already agree.
|
||||
pin_release_tag = None
|
||||
|
|
@ -643,6 +649,7 @@ def _plan_llama_phase() -> dict:
|
|||
"pin_release_tag": pin_release_tag,
|
||||
"from_tag": from_tag,
|
||||
"force_cpu": force_cpu,
|
||||
"llama_backend": llama_backend,
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -695,6 +702,7 @@ def start_update() -> dict:
|
|||
llama_spec["pin_release_tag"],
|
||||
set_progress,
|
||||
force_cpu = llama_spec.get("force_cpu", False),
|
||||
llama_backend = llama_spec.get("llama_backend"),
|
||||
)
|
||||
)
|
||||
if llama_spec
|
||||
|
|
|
|||
|
|
@ -65,6 +65,54 @@ EXIT_ERROR = 1
|
|||
EXIT_BUSY = 3
|
||||
EXIT_NO_SPACE = 4
|
||||
|
||||
# Every gfx a Windows AMD host can be served: ggml-org release.yml windows-hip GPU_TARGETS
|
||||
# plus the fork's windows-rocm bundles. Must stay a superset of the manifest's windows-rocm
|
||||
# mapped_targets, else auto-Vulkan steals a host the fork already builds for. Below this
|
||||
# floor (e.g. gfx803 / RX 480) HIP has no prebuilt and Vulkan is the practical Windows
|
||||
# llama-server backend (#7357).
|
||||
WINDOWS_HIP_PREBUILT_GFX_TARGETS = frozenset(
|
||||
{
|
||||
"gfx908",
|
||||
"gfx90a",
|
||||
"gfx1030",
|
||||
"gfx1031",
|
||||
"gfx1032",
|
||||
"gfx1034",
|
||||
"gfx1100",
|
||||
"gfx1101",
|
||||
"gfx1102",
|
||||
"gfx1103",
|
||||
"gfx1150",
|
||||
"gfx1151",
|
||||
"gfx1200",
|
||||
"gfx1201",
|
||||
}
|
||||
)
|
||||
# Family labels forwarded by update markers / --rocm-gfx (gfx110X.zip assets).
|
||||
WINDOWS_ROCM_FAMILY_GFX_LABELS = frozenset({"gfx103x", "gfx110x", "gfx120x"})
|
||||
|
||||
# Exactly ggml-org release.yml's windows-hip "radeon" gpu_targets. The set above adds the
|
||||
# fork-only bundles (gfx1034, gfx1103, gfx908, gfx90a), served only against the fork.
|
||||
UPSTREAM_WINDOWS_HIP_GFX_TARGETS = frozenset(
|
||||
{
|
||||
"gfx1030",
|
||||
"gfx1031",
|
||||
"gfx1032",
|
||||
"gfx1100",
|
||||
"gfx1101",
|
||||
"gfx1102",
|
||||
"gfx1150",
|
||||
"gfx1151",
|
||||
"gfx1200",
|
||||
"gfx1201",
|
||||
}
|
||||
)
|
||||
|
||||
# install_kinds that really are a Vulkan bundle. A Vulkan request can still end on a CPU
|
||||
# bundle (no Vulkan archive on Windows arm64; x64 falls through when it is missing or fails
|
||||
# validation), so check against this to keep the marker honest (#7357).
|
||||
VULKAN_INSTALL_KINDS = frozenset({"linux-vulkan", "windows-vulkan"})
|
||||
|
||||
# DiskPart-prompt suppression. RunAsInvoker does NOT stop amd-smi's runtime
|
||||
# elevation (its manifest is asInvoker), so this is just harmless belt-and-
|
||||
# suspenders for manifest-elevating tools. The real guard is _amd_smi_allowed():
|
||||
|
|
@ -282,6 +330,7 @@ class HostInfo:
|
|||
has_rocm: bool = False
|
||||
has_intel_gpu: bool = False
|
||||
rocm_gfx_target: str | None = None
|
||||
rocm_gfx_targets: list[str] = field(default_factory = list)
|
||||
# (major, minor) from platform.mac_ver(); None off macOS or if unparseable.
|
||||
# Skips a macos prebuilt whose minimum-OS exceeds this host.
|
||||
macos_version: tuple[int, int] | None = None
|
||||
|
|
@ -2159,47 +2208,37 @@ def run_capture(
|
|||
return result
|
||||
|
||||
|
||||
def _pick_rocm_gfx_target(out: str) -> str | None:
|
||||
"""Choose the gfx target rocminfo / hipinfo report for the active GPU.
|
||||
def _list_rocm_gfx_targets(out: str) -> list[str]:
|
||||
"""List gfx targets rocminfo / hipinfo report, one entry per physical GPU.
|
||||
|
||||
A bare first-match picked the wrong device on mixed APU + dGPU hosts
|
||||
(e.g. Strix Halo gfx1151 + discrete RX 7900 gfx1100). Respect
|
||||
HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES / CUDA_VISIBLE_DEVICES so the
|
||||
asset matches what HIP actually runs on. Falls back to the first GPU when
|
||||
no env var is set.
|
||||
|
||||
rocminfo / hipinfo print the same gfx token multiple times per GPU (Name,
|
||||
ISA, marketing-name). We first try to split the output on per-GPU section
|
||||
headers (rocminfo: "Agent N" blocks, hipinfo: "device#N" entries) and take
|
||||
exactly one gfx token per section. This gives the correct per-GPU list even
|
||||
on same-arch multi-GPU hosts (e.g. two RX 7900 XTX cards) where global
|
||||
dict.fromkeys dedup would collapse both cards to a single entry and make
|
||||
HIP_VISIBLE_DEVICES=1 point out of range.
|
||||
|
||||
Falls back to insertion-order dedup when the output has no recognisable
|
||||
section markers (flat gfx-string inputs, unit-test stubs, etc.).
|
||||
|
||||
Empty / "-1" env values mean no AMD GPU is visible to HIP: return None.
|
||||
Both repeat the same gfx token per GPU (Name, ISA, marketing-name), so split on per-GPU
|
||||
section headers to keep two entries on a dual same-arch host; flat strings and test stubs
|
||||
fall back to insertion-order dedup.
|
||||
"""
|
||||
# Try to build a per-GPU token list by splitting on section boundaries.
|
||||
# rocminfo sections are introduced by "Agent N" lines (optionally between
|
||||
# rows of asterisks). hipinfo sections start with "device#N".
|
||||
_sections = re.split(
|
||||
r"(?mi)^\s*\*+\s*$\s*agent\s+\d+\s*$|\bdevice\s*#\s*\d+\b",
|
||||
out,
|
||||
)
|
||||
if len(_sections) > 1:
|
||||
# Section-based: one gfx token per GPU section preserves physical order.
|
||||
_tokens: list[str] = []
|
||||
for _sec in _sections[1:]:
|
||||
_m = re.search(r"gfx[1-9][0-9a-z]{2,3}", _sec.lower())
|
||||
if _m:
|
||||
_tokens.append(_m.group(0))
|
||||
else:
|
||||
# Fallback: insertion-order dedup (handles flat strings / unknown formats).
|
||||
_raw = re.findall(r"gfx[1-9][0-9a-z]{2,3}", out.lower())
|
||||
_tokens = list(dict.fromkeys(_raw))
|
||||
return _tokens
|
||||
|
||||
|
||||
def _pick_rocm_gfx_target(out: str) -> str | None:
|
||||
"""Choose the gfx target rocminfo / hipinfo report for the active GPU.
|
||||
|
||||
A bare first-match picked the wrong device on mixed APU + dGPU hosts (Strix Halo gfx1151
|
||||
+ RX 7900 gfx1100), so honour HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES /
|
||||
CUDA_VISIBLE_DEVICES; no env var means the first GPU, empty / "-1" means none (None).
|
||||
"""
|
||||
_tokens = _list_rocm_gfx_targets(out)
|
||||
if not _tokens:
|
||||
return None
|
||||
|
||||
|
|
@ -2407,6 +2446,7 @@ def detect_host() -> HostInfo:
|
|||
|
||||
has_rocm = False
|
||||
rocm_gfx_target: str | None = None
|
||||
rocm_gfx_targets: list[str] = []
|
||||
if is_linux and not has_usable_nvidia:
|
||||
# WSL2 ROCDXG: the system rocminfo enumerates the GPU over /dev/dxg
|
||||
# only when HSA_ENABLE_DXG_DETECTION=1 (a no-op on bare metal), and
|
||||
|
|
@ -2444,6 +2484,7 @@ def detect_host() -> HostInfo:
|
|||
if _result.returncode == 0 and _result.stdout.strip():
|
||||
if _check(_result.stdout):
|
||||
has_rocm = True
|
||||
rocm_gfx_targets = _list_rocm_gfx_targets(_result.stdout)
|
||||
rocm_gfx_target = _pick_rocm_gfx_target(_result.stdout)
|
||||
break
|
||||
elif is_windows and not has_usable_nvidia:
|
||||
|
|
@ -2489,6 +2530,7 @@ def detect_host() -> HostInfo:
|
|||
if _check(_result.stdout):
|
||||
has_rocm = True
|
||||
# hipinfo reports "gcnArchName: gfx1100" -- extract if present
|
||||
rocm_gfx_targets = _list_rocm_gfx_targets(_result.stdout)
|
||||
rocm_gfx_target = _pick_rocm_gfx_target(_result.stdout)
|
||||
break
|
||||
# Note: amdhip64.dll presence alone is NOT treated as GPU evidence
|
||||
|
|
@ -2551,6 +2593,7 @@ def detect_host() -> HostInfo:
|
|||
has_rocm = has_rocm,
|
||||
has_intel_gpu = has_intel_gpu,
|
||||
rocm_gfx_target = rocm_gfx_target,
|
||||
rocm_gfx_targets = rocm_gfx_targets,
|
||||
macos_version = macos_version,
|
||||
)
|
||||
|
||||
|
|
@ -2573,12 +2616,12 @@ def _apply_host_overrides(
|
|||
force_cpu: bool = False,
|
||||
) -> HostInfo:
|
||||
"""Fold setup.sh/setup.ps1's forwarded detection into the host profile.
|
||||
A forwarded gfx (--rocm-gfx or UNSLOTH_ROCM_GFX_ARCH) is authoritative and
|
||||
implies ROCm: the installer's own hipinfo/amd-smi probe can miss the arch on
|
||||
amd-smi-only hosts or when setup inferred it from the GPU name, leaving
|
||||
rocm_gfx_target None and no per-gfx ROCm prebuilt selected. force_cpu is the
|
||||
opposite explicit signal (arm64 Linux GPU host whose source build failed):
|
||||
drop GPU attributes so the CPU prebuilt for this OS/arch is selected."""
|
||||
A forwarded gfx (--rocm-gfx or UNSLOTH_ROCM_GFX_ARCH) implies ROCm and fills the gap
|
||||
where our own hipinfo/amd-smi probe misses the arch (amd-smi-only hosts, or setup
|
||||
inferring it from the GPU name), leaving no per-gfx ROCm prebuilt selected; it stays
|
||||
authoritative except for the two advisory shapes narrowed below. force_cpu is the
|
||||
opposite explicit signal (arm64 Linux GPU host whose source build failed): drop GPU
|
||||
attributes so the CPU prebuilt for this OS/arch is selected."""
|
||||
if force_cpu:
|
||||
return dataclasses_replace(
|
||||
host,
|
||||
|
|
@ -2586,11 +2629,43 @@ def _apply_host_overrides(
|
|||
has_physical_nvidia = False,
|
||||
has_rocm = False,
|
||||
rocm_gfx_target = None,
|
||||
rocm_gfx_targets = [],
|
||||
has_intel_gpu = False,
|
||||
)
|
||||
gfx = _normalize_forwarded_gfx(override_rocm_gfx)
|
||||
if gfx:
|
||||
return dataclasses_replace(host, has_rocm = True, rocm_gfx_target = gfx)
|
||||
# setup.ps1's pick is not fully visible-device aware (neither branch reads
|
||||
# CUDA_VISIBLE_DEVICES; amd-smi matches a bare integer only, so "1,0" falls back to
|
||||
# GPU 0), while _pick_rocm_gfx_target() honours all three vars with HIP's semantics.
|
||||
# So keep a probed active arch when the forward is only advisory, else
|
||||
# _should_auto_vulkan_for_amd_windows() reads a HIP-supported GPU the user masked
|
||||
# off and installs an unusable HIP bundle instead of Vulkan. Advisory means:
|
||||
# * another GPU the probe saw ON THIS HOST, i.e. setup picked a different card;
|
||||
# * a family label (gfx110X), a bundle name the update path emits from the marker
|
||||
# asset and never a real arch -- it would upgrade an in-generation-but-unbuilt
|
||||
# GPU (gfx1033) into a bundle it must not be served.
|
||||
# Anything else the probe never reported is an operator override for a host whose
|
||||
# arch the probe gets wrong or stale, exactly what --rocm-gfx documents, so it stays
|
||||
# authoritative. UNSLOTH_ROCM_GFX_ARCH also still wins.
|
||||
_manual = _normalize_forwarded_gfx(os.environ.get("UNSLOTH_ROCM_GFX_ARCH"))
|
||||
_physical = _host_rocm_gfx_targets(host)
|
||||
_active = _active_rocm_gfx_target(host)
|
||||
_advisory = gfx in _physical or gfx in WINDOWS_ROCM_FAMILY_GFX_LABELS
|
||||
if gfx != _manual and _active and gfx != _active and _advisory:
|
||||
return dataclasses_replace(host, has_rocm = True)
|
||||
return dataclasses_replace(
|
||||
host,
|
||||
has_rocm = True,
|
||||
rocm_gfx_target = gfx,
|
||||
# ADD the forwarded arch to the probe's per-GPU list, never replace it: that
|
||||
# list is the PHYSICAL inventory _should_auto_vulkan_for_amd_windows() reads,
|
||||
# and a forward says which GPU HIP should target, not which cards exist.
|
||||
# Dropping a probe-confirmed GPU would let a stale below-floor forward
|
||||
# auto-route a box with a HIP-capable card to Vulkan, which then enumerates
|
||||
# that card regardless of HIP_VISIBLE_DEVICES. An empty probe still yields
|
||||
# [gfx], so the driver-only host the forward exists for keeps auto-Vulkan.
|
||||
rocm_gfx_targets = list(dict.fromkeys([*_physical, gfx])),
|
||||
)
|
||||
if override_has_rocm and not host.has_rocm:
|
||||
return dataclasses_replace(host, has_rocm = True)
|
||||
return host
|
||||
|
|
@ -5464,6 +5539,18 @@ def _fork_manifest_release_plans(
|
|||
raise PrebuiltFallback("no installable published llama.cpp releases were found")
|
||||
|
||||
|
||||
def persisted_llama_backend(llama_backend: str | None, choice: AssetChoice) -> str | None:
|
||||
"""The backend to record for an install that actually landed ``choice``.
|
||||
|
||||
A Vulkan request can end on a non-Vulkan bundle (no upstream Vulkan archive for Windows
|
||||
arm64; x64 falls through to win-cpu-x64 when it is missing or fails validation), and
|
||||
recording "vulkan" there would make the updater re-assert a backend that was never
|
||||
installed. Mirrors force_cpu: persist the real outcome only."""
|
||||
if llama_backend == "vulkan" and choice.install_kind not in VULKAN_INSTALL_KINDS:
|
||||
return None
|
||||
return llama_backend
|
||||
|
||||
|
||||
def write_prebuilt_metadata(
|
||||
install_dir: Path,
|
||||
*,
|
||||
|
|
@ -5474,6 +5561,7 @@ def write_prebuilt_metadata(
|
|||
approved_checksums: ApprovedReleaseChecksums,
|
||||
prebuilt_fallback_used: bool,
|
||||
force_cpu: bool = False,
|
||||
llama_backend: str | None = None,
|
||||
) -> None:
|
||||
source_asset_name, source_sha256 = selected_source_archive_metadata(
|
||||
approved_checksums,
|
||||
|
|
@ -5502,6 +5590,9 @@ def write_prebuilt_metadata(
|
|||
# so a forced CPU install is not re-routed to a GPU bundle (#7213). An automatic
|
||||
# --cpu-fallback (e.g. arm64 GPU-build recovery) stays False so it can heal to GPU.
|
||||
"force_cpu": force_cpu,
|
||||
# Deliberate or auto-selected Vulkan backend (#7357); the updater re-asserts it so
|
||||
# AMD hosts are not swapped back to HIP. Dropped if the winning attempt was not Vulkan.
|
||||
"llama_backend": persisted_llama_backend(llama_backend, choice),
|
||||
"asset_sha256": choice.expected_sha256,
|
||||
"source": choice.source_label,
|
||||
# Binary-side repo/tag for non-fork sources (e.g. the ggml-org upstream
|
||||
|
|
@ -5549,6 +5640,23 @@ def sync_marker_force_cpu(install_dir: Path, persist_force_cpu: bool) -> None:
|
|||
log(f"existing install reused; recorded force_cpu={persist_force_cpu} from this run")
|
||||
|
||||
|
||||
def sync_marker_llama_backend(install_dir: Path, llama_backend: str | None) -> None:
|
||||
"""Sync the persisted llama.cpp backend when the bundle is reused unchanged."""
|
||||
marker_path = install_dir / "UNSLOTH_PREBUILT_INFO.json"
|
||||
try:
|
||||
marker = json.loads(marker_path.read_text())
|
||||
except (OSError, ValueError):
|
||||
return
|
||||
if not isinstance(marker, dict) or marker.get("llama_backend") == llama_backend:
|
||||
return
|
||||
if llama_backend is None:
|
||||
marker.pop("llama_backend", None)
|
||||
else:
|
||||
marker["llama_backend"] = llama_backend
|
||||
marker_path.write_text(json.dumps(marker, indent = 2) + "\n")
|
||||
log(f"existing install reused; recorded llama_backend={llama_backend!r} from this run")
|
||||
|
||||
|
||||
def expected_install_fingerprint(
|
||||
*,
|
||||
llama_tag: str,
|
||||
|
|
@ -5841,6 +5949,7 @@ def validate_prebuilt_choice(
|
|||
prebuilt_fallback_used: bool,
|
||||
quantized_path: Path,
|
||||
force_cpu: bool = False,
|
||||
llama_backend: str | None = None,
|
||||
) -> tuple[Path, Path]:
|
||||
source_repo, source_ref, source_archive, exact_source = preferred_source_archive(
|
||||
approved_checksums, llama_tag
|
||||
|
|
@ -5882,6 +5991,7 @@ def validate_prebuilt_choice(
|
|||
approved_checksums = approved_checksums,
|
||||
prebuilt_fallback_used = prebuilt_fallback_used,
|
||||
force_cpu = force_cpu,
|
||||
llama_backend = llama_backend,
|
||||
)
|
||||
# Hashless external prebuilts are not in the approved-sha256
|
||||
# manifest and rely on the functional smoke test as their only integrity gate,
|
||||
|
|
@ -5969,6 +6079,7 @@ def validate_prebuilt_attempts(
|
|||
initial_fallback_used: bool = False,
|
||||
existing_install_dir: Path | None = None,
|
||||
force_cpu: bool = False,
|
||||
llama_backend: str | None = None,
|
||||
) -> tuple[AssetChoice, Path, bool]:
|
||||
attempt_list = list(attempts)
|
||||
if not attempt_list:
|
||||
|
|
@ -6030,6 +6141,7 @@ def validate_prebuilt_attempts(
|
|||
prebuilt_fallback_used = tried_fallback,
|
||||
quantized_path = quantized_path,
|
||||
force_cpu = force_cpu,
|
||||
llama_backend = llama_backend,
|
||||
)
|
||||
except Exception as exc:
|
||||
remove_tree(staging_dir)
|
||||
|
|
@ -6055,12 +6167,43 @@ def validate_prebuilt_attempts(
|
|||
raise PrebuiltFallback("no prebuilt bundle passed validation")
|
||||
|
||||
|
||||
def force_vulkan_requested() -> bool:
|
||||
"""Whether UNSLOTH_FORCE_VULKAN opts this host into the Vulkan llama.cpp
|
||||
prebuilt instead of its detected CUDA/ROCm backend (e.g. so an AMD user can
|
||||
run the Vulkan build for inference). Scoped to the llama.cpp backend; the
|
||||
torch/training stack installs separately and still sees the real GPU.
|
||||
def _normalized_llama_backend(value: str | None) -> str | None:
|
||||
if not value:
|
||||
return None
|
||||
backend = value.strip().lower()
|
||||
if backend in {"vulkan", "hip", "rocm", "cpu"}:
|
||||
return "hip" if backend == "rocm" else backend
|
||||
return None
|
||||
|
||||
|
||||
def llama_backend_from_env() -> str | None:
|
||||
"""Read an explicit llama.cpp backend preference from the environment.
|
||||
|
||||
Only ``UNSLOTH_LLAMA_BACKEND`` is honored. ``UNSLOTH_LLAMA_CPP_BACKEND`` is a separate
|
||||
setup variable meaning ``auto``/``cpu`` (not a backend name) that setup warns about and
|
||||
otherwise ignores, so reading it here would force Vulkan behind that warning.
|
||||
"""
|
||||
return _normalized_llama_backend(os.environ.get("UNSLOTH_LLAMA_BACKEND"))
|
||||
|
||||
|
||||
def resolved_llama_backend(llama_backend: str | None = None) -> str | None:
|
||||
"""The explicit backend for this run: --llama-backend, else the env var. None when
|
||||
neither is set or the value is not a backend name we know."""
|
||||
return _normalized_llama_backend(llama_backend) or llama_backend_from_env()
|
||||
|
||||
|
||||
def force_vulkan_requested(llama_backend: str | None = None) -> bool:
|
||||
"""Whether this run should install the upstream Vulkan llama.cpp prebuilt.
|
||||
|
||||
Triggered by ``UNSLOTH_LLAMA_BACKEND=vulkan``, legacy ``UNSLOTH_FORCE_VULKAN``, or
|
||||
``--llama-backend vulkan``. Scoped to the llama.cpp backend; the torch/training stack
|
||||
installs separately and still sees the real GPU.
|
||||
"""
|
||||
backend = resolved_llama_backend(llama_backend)
|
||||
if backend is not None:
|
||||
# Authoritative, so =hip is a real opt-out a stale UNSLOTH_FORCE_VULKAN cannot
|
||||
# overrule.
|
||||
return backend == "vulkan"
|
||||
return os.environ.get("UNSLOTH_FORCE_VULKAN", "").strip().lower() in (
|
||||
"1",
|
||||
"true",
|
||||
|
|
@ -6068,6 +6211,117 @@ def force_vulkan_requested() -> bool:
|
|||
)
|
||||
|
||||
|
||||
def _host_rocm_gfx_targets(host: HostInfo) -> list[str]:
|
||||
if host.rocm_gfx_targets:
|
||||
return [target.lower() for target in host.rocm_gfx_targets]
|
||||
if host.rocm_gfx_target:
|
||||
return [host.rocm_gfx_target.lower()]
|
||||
return []
|
||||
|
||||
|
||||
def _active_rocm_gfx_target(host: HostInfo) -> str | None:
|
||||
"""The gfx HIP will run on (visible-device aware), not every physical GPU."""
|
||||
if host.rocm_gfx_target:
|
||||
return host.rocm_gfx_target.lower().strip()
|
||||
return None
|
||||
|
||||
|
||||
def _hip_visible_device_mask_set() -> bool:
|
||||
"""Whether a HIP visible-device mask is in force for this process.
|
||||
|
||||
The Windows arch probe is hipinfo, itself a HIP application, so under a mask it
|
||||
enumerates the VISIBLE devices, not the physical ones. Presence is the whole test: a
|
||||
partial mask leaves the inventory unknowable, and an all-hiding "" / "-1" is the
|
||||
strongest form of that, not an exemption, since --rocm-gfx can still supply an arch
|
||||
(setup infers it from the display adapter, which no HIP mask touches) and would
|
||||
auto-route a host on which the user hid every AMD GPU. Reads the same three vars as
|
||||
_pick_rocm_gfx_target, so the two cannot disagree about the host."""
|
||||
return any(
|
||||
os.environ.get(_env) is not None
|
||||
for _env in ("HIP_VISIBLE_DEVICES", "ROCR_VISIBLE_DEVICES", "CUDA_VISIBLE_DEVICES")
|
||||
)
|
||||
|
||||
|
||||
def _windows_hip_gfx_targets(published_repo: str | None) -> frozenset[str]:
|
||||
"""gfx targets the Windows HIP bundle of ``published_repo`` is actually built for.
|
||||
|
||||
The combined floor above includes the FORK's windows-rocm bundles, and only the fork is
|
||||
planned from a manifest: resolve_simple_install_release_plans() sends every other
|
||||
--published-repo (ggml-org, but equally any mirror of upstream-standard assets) to
|
||||
direct_upstream_release_plan(), whose AMD branch offers win-hip-radeon then CPU and never
|
||||
Vulkan. Answering "supported" there for a fork-only arch would silently land it on
|
||||
HIP/CPU instead of the Vulkan bundle that would actually run, so gate on the fork rather
|
||||
than exempting one repo, mirroring that dispatch exactly, spelling included: an empty
|
||||
value defaults to the fork, and a differently cased repo really does take the upstream
|
||||
path and must be answered with upstream coverage."""
|
||||
if (published_repo or DEFAULT_PUBLISHED_REPO) == DEFAULT_PUBLISHED_REPO:
|
||||
return WINDOWS_HIP_PREBUILT_GFX_TARGETS
|
||||
return UPSTREAM_WINDOWS_HIP_GFX_TARGETS
|
||||
|
||||
|
||||
def _gfx_is_windows_hip_supported(gfx: str, published_repo: str | None = None) -> bool:
|
||||
token = gfx.lower().strip()
|
||||
if token in WINDOWS_ROCM_FAMILY_GFX_LABELS:
|
||||
# A bundle name, not an arch, so the concrete GPU is unknown here. The fork builds
|
||||
# every member; upstream builds all but gfx1034 / gfx1103. Answering "unsupported"
|
||||
# to cover that pair would move gfx1030..1032 / gfx1100..1102 off a working HIP
|
||||
# build onto Vulkan for a member the label cannot identify, so HIP serves it either
|
||||
# way. A concrete arch below still answers per repo, which is where gfx1034 and
|
||||
# gfx1103 do reach Vulkan.
|
||||
return True
|
||||
return token in _windows_hip_gfx_targets(published_repo)
|
||||
|
||||
|
||||
def _host_has_windows_hip_prebuilt_gfx(host: HostInfo, published_repo: str | None = None) -> bool:
|
||||
active = _active_rocm_gfx_target(host)
|
||||
if not active:
|
||||
return False
|
||||
return _gfx_is_windows_hip_supported(active, published_repo)
|
||||
|
||||
|
||||
def _should_auto_vulkan_for_amd_windows(host: HostInfo, published_repo: str | None = None) -> bool:
|
||||
"""True when NO AMD GPU on the host reaches the Windows HIP prebuilt floor."""
|
||||
active = _active_rocm_gfx_target(host)
|
||||
if not active:
|
||||
# ROCm confirmed but gfx unknown (--has-rocm only): keep the HIP / fork / source path.
|
||||
return False
|
||||
if not (
|
||||
host.is_windows
|
||||
and host.has_rocm
|
||||
# PHYSICAL, not merely usable: Vulkan ignores CUDA_VISIBLE_DEVICES and would
|
||||
# enumerate a card hidden by it. Same gate as the Intel auto path below.
|
||||
and not host.has_physical_nvidia
|
||||
):
|
||||
return False
|
||||
# Judge every PHYSICAL AMD gfx, not just the active one. The visible-device vars choose
|
||||
# `active`, but the Vulkan runtime honours none of them (it enumerates through
|
||||
# GGML_VK_VISIBLE_DEVICES and Vulkan ordinals), so masking down to a below-floor card
|
||||
# must not route the install to Vulkan: the installed backend would happily enumerate
|
||||
# the HIP-capable card the user deliberately hid, possibly one reserved for another
|
||||
# workload. Auto-fall back only when no AMD device on the box can be exposed to HIP; an
|
||||
# explicit vulkan opt-in is unaffected.
|
||||
#
|
||||
# Under a mask the probe cannot supply that inventory at all (hipinfo sees only visible
|
||||
# devices), so "no AMD GPU here reaches the floor" is unprovable and guessing wrong is
|
||||
# the same reserved-card handover. An all-hiding "" / "-1" is included: a forwarded
|
||||
# --rocm-gfx still reconstructs an arch there, and auto-routing would then hand Vulkan
|
||||
# every AMD GPU the user hid. A mask is only ever set deliberately, and the driver-only
|
||||
# single-GPU host this fallback exists for does not set one.
|
||||
if _hip_visible_device_mask_set():
|
||||
return False
|
||||
targets = list(dict.fromkeys([*_host_rocm_gfx_targets(host), active]))
|
||||
return not any(_gfx_is_windows_hip_supported(target, published_repo) for target in targets)
|
||||
|
||||
|
||||
def _has_no_vulkan_prebuilt(host: HostInfo) -> bool:
|
||||
"""Platforms that ship no Vulkan prebuilt at all, so routing there is pointless.
|
||||
|
||||
Upstream builds win-vulkan for x64 only; Windows arm64 gets CPU plus opencl-adreno, so
|
||||
rewriting it to Vulkan-only would just swap the published bundle for the upstream CPU
|
||||
one. macOS is handled separately (Metal)."""
|
||||
return host.is_windows and host.is_arm64
|
||||
|
||||
|
||||
def _vulkan_only_host(host: HostInfo) -> HostInfo:
|
||||
"""Rewrite ``host`` so the asset selectors take their Vulkan branch.
|
||||
|
||||
|
|
@ -6081,52 +6335,79 @@ def _vulkan_only_host(host: HostInfo) -> HostInfo:
|
|||
has_usable_nvidia = False,
|
||||
has_physical_nvidia = False,
|
||||
has_rocm = False,
|
||||
rocm_gfx_target = None,
|
||||
rocm_gfx_targets = [],
|
||||
has_intel_gpu = True,
|
||||
)
|
||||
|
||||
|
||||
def _route_to_vulkan_prebuilt(
|
||||
host: HostInfo, published_repo: str, published_release_tag: str, *, force_cpu: bool
|
||||
) -> tuple[HostInfo, str, str]:
|
||||
host: HostInfo,
|
||||
published_repo: str,
|
||||
published_release_tag: str,
|
||||
*,
|
||||
force_cpu: bool,
|
||||
llama_backend: str | None = None,
|
||||
) -> tuple[HostInfo, str, str, str | None]:
|
||||
"""Point a Vulkan-capable host at the upstream ggml-org Vulkan prebuilt.
|
||||
|
||||
The unsloth published repo ships only CUDA/ROCm/CPU assets, so Vulkan comes
|
||||
from UPSTREAM_REPO. Two triggers route here, both suppressed when a CPU flag
|
||||
(--cpu-fallback or --force-cpu, folded into force_cpu) wins:
|
||||
* UNSLOTH_FORCE_VULKAN forces Vulkan over the detected CUDA/ROCm backend;
|
||||
* an auto-detected Intel GPU with NO physical NVIDIA/ROCm -- the purpose
|
||||
of the has_intel_gpu probe, since the fork manifest ships no Vulkan asset.
|
||||
Applied by BOTH the install path and the --resolve-prebuilt probe so the
|
||||
"is a prebuilt available" answer matches what actually gets installed.
|
||||
The unsloth published repo ships only CUDA/ROCm/CPU assets, so Vulkan comes from
|
||||
UPSTREAM_REPO. Three triggers route here, all suppressed when a CPU flag (--cpu-fallback
|
||||
or --force-cpu, folded into force_cpu) wins:
|
||||
* ``UNSLOTH_LLAMA_BACKEND=vulkan`` / ``UNSLOTH_FORCE_VULKAN`` / ``--llama-backend
|
||||
vulkan`` forces Vulkan over the detected CUDA/ROCm backend;
|
||||
* Windows AMD with no HIP-prebuilt gfx arch auto-falls back to Vulkan (#7357);
|
||||
* an auto-detected Intel GPU with NO physical NVIDIA/ROCm, the purpose of the
|
||||
has_intel_gpu probe, since the fork manifest ships no Vulkan asset.
|
||||
Applied by BOTH the install path and the --resolve-prebuilt probe so the "is a prebuilt
|
||||
available" answer matches what actually gets installed.
|
||||
|
||||
Returns the (possibly rewritten) host, repo, and release tag.
|
||||
Returns the (possibly rewritten) host, repo, release tag, and a backend to persist in
|
||||
the install marker when updates must re-assert Vulkan.
|
||||
"""
|
||||
forced = force_vulkan_requested()
|
||||
# Gate auto-routing on no PHYSICAL NVIDIA, not merely no usable one: a mixed
|
||||
# NVIDIA+Intel host that hides NVIDIA with CUDA_VISIBLE_DEVICES=""/-1 keeps
|
||||
# has_physical_nvidia=True while has_usable_nvidia goes False. Vulkan ignores
|
||||
# CUDA_VISIBLE_DEVICES, so auto-routing such a host would let it grab the
|
||||
# reserved NVIDIA GPU. An explicit UNSLOTH_FORCE_VULKAN still overrides.
|
||||
forced = force_vulkan_requested(llama_backend)
|
||||
# Auto-fall back only when the run named no backend: an explicit hip/cpu is the opt-out.
|
||||
explicit_backend = resolved_llama_backend(llama_backend)
|
||||
auto_no_hip = explicit_backend is None and _should_auto_vulkan_for_amd_windows(
|
||||
host, published_repo
|
||||
)
|
||||
# No PHYSICAL NVIDIA, not merely no usable one: Vulkan ignores CUDA_VISIBLE_DEVICES, so
|
||||
# auto-routing a host that hides its NVIDIA card would let it grab the reserved GPU.
|
||||
auto_intel = host.has_intel_gpu and not host.has_physical_nvidia and not host.has_rocm
|
||||
if force_cpu or not (forced or auto_intel):
|
||||
return host, published_repo, published_release_tag
|
||||
if force_cpu or not (forced or auto_intel or auto_no_hip):
|
||||
return host, published_repo, published_release_tag, None
|
||||
if host.is_macos:
|
||||
if forced:
|
||||
log(
|
||||
"UNSLOTH_FORCE_VULKAN is set but ignored on macOS "
|
||||
"UNSLOTH_LLAMA_BACKEND=vulkan is set but ignored on macOS "
|
||||
"(Metal is used; there is no Vulkan prebuilt)"
|
||||
)
|
||||
return host, published_repo, published_release_tag
|
||||
if forced:
|
||||
return host, published_repo, published_release_tag, None
|
||||
if _has_no_vulkan_prebuilt(host):
|
||||
if forced:
|
||||
log(
|
||||
"Vulkan llama.cpp backend requested but ignored on Windows arm64 "
|
||||
"(upstream ships no Vulkan arm64 prebuilt); keeping the published bundle"
|
||||
)
|
||||
return host, published_repo, published_release_tag, None
|
||||
if auto_no_hip:
|
||||
active = _active_rocm_gfx_target(host) or "unknown"
|
||||
log(
|
||||
"UNSLOTH_FORCE_VULKAN is set; installing the upstream Vulkan "
|
||||
"llama.cpp prebuilt instead of the detected GPU backend"
|
||||
"Active AMD GPU arch is not supported by the Windows HIP prebuilt "
|
||||
f"({active}); installing the upstream Vulkan llama.cpp prebuilt instead"
|
||||
)
|
||||
# Forcing may override a detected NVIDIA/ROCm host, so normalize it to
|
||||
# Vulkan-only; an auto-detected Intel host already is.
|
||||
host = _vulkan_only_host(host)
|
||||
persist_backend = "vulkan"
|
||||
elif forced:
|
||||
log(
|
||||
"Vulkan llama.cpp backend requested; installing the upstream Vulkan "
|
||||
"prebuilt instead of the detected GPU backend"
|
||||
)
|
||||
host = _vulkan_only_host(host)
|
||||
persist_backend = "vulkan"
|
||||
else:
|
||||
log("Intel GPU detected; installing the upstream Vulkan llama.cpp prebuilt")
|
||||
persist_backend = None
|
||||
# Swapping the fork for upstream invalidates a fork release pin: the two use
|
||||
# different tag namespaces (fork b9596-mix-<sha> vs upstream b9596), so a
|
||||
# pinned fork tag would make the upstream resolver query a nonexistent
|
||||
|
|
@ -6135,7 +6416,7 @@ def _route_to_vulkan_prebuilt(
|
|||
# (repo unchanged here) is preserved.
|
||||
if published_repo != UPSTREAM_REPO:
|
||||
published_release_tag = ""
|
||||
return host, UPSTREAM_REPO, published_release_tag
|
||||
return host, UPSTREAM_REPO, published_release_tag, persist_backend
|
||||
|
||||
|
||||
def diffusion_visual_server_backfill_needed(
|
||||
|
|
@ -6299,6 +6580,7 @@ def install_prebuilt(
|
|||
override_rocm_gfx: str | None = None,
|
||||
force_cpu: bool = False,
|
||||
persist_force_cpu: bool = False,
|
||||
llama_backend: str | None = None,
|
||||
instruction_cleanup_root: Path | None = None,
|
||||
) -> None:
|
||||
# force_cpu drops GPU detection (mechanism, both --cpu-fallback and --force-cpu);
|
||||
|
|
@ -6310,8 +6592,12 @@ def install_prebuilt(
|
|||
override_rocm_gfx = override_rocm_gfx,
|
||||
force_cpu = force_cpu,
|
||||
)
|
||||
host, published_repo, published_release_tag = _route_to_vulkan_prebuilt(
|
||||
host, published_repo, published_release_tag, force_cpu = force_cpu
|
||||
host, published_repo, published_release_tag, persist_llama_backend = _route_to_vulkan_prebuilt(
|
||||
host,
|
||||
published_repo,
|
||||
published_release_tag,
|
||||
force_cpu = force_cpu,
|
||||
llama_backend = llama_backend,
|
||||
)
|
||||
choice: AssetChoice | None = None
|
||||
cleanup_root = install_dir if instruction_cleanup_root is None else instruction_cleanup_root
|
||||
|
|
@ -6356,6 +6642,10 @@ def install_prebuilt(
|
|||
# Reused bundle is unchanged, but a fresh --force-cpu still must be
|
||||
# recorded so the updater re-asserts it (#7213).
|
||||
sync_marker_force_cpu(install_dir, persist_force_cpu)
|
||||
sync_marker_llama_backend(
|
||||
install_dir,
|
||||
persisted_llama_backend(persist_llama_backend, current.attempts[0]),
|
||||
)
|
||||
return
|
||||
with scratch_dir("unsloth-llama-prebuilt-") as work_dir:
|
||||
probe_path = work_dir / "stories260K.gguf"
|
||||
|
|
@ -6376,6 +6666,10 @@ def install_prebuilt(
|
|||
f"{plan.release_tag} upstream_tag={plan.llama_tag}; skipping reinstall"
|
||||
)
|
||||
sync_marker_force_cpu(install_dir, persist_force_cpu)
|
||||
sync_marker_llama_backend(
|
||||
install_dir,
|
||||
persisted_llama_backend(persist_llama_backend, choice),
|
||||
)
|
||||
return
|
||||
log(
|
||||
"selected "
|
||||
|
|
@ -6398,6 +6692,7 @@ def install_prebuilt(
|
|||
existing_install_dir = install_dir,
|
||||
# Persist only the deliberate choice, not a transient fallback.
|
||||
force_cpu = persist_force_cpu,
|
||||
llama_backend = persist_llama_backend,
|
||||
)
|
||||
except ExistingInstallSatisfied:
|
||||
return
|
||||
|
|
@ -6519,6 +6814,16 @@ def parse_args() -> argparse.Namespace:
|
|||
"bundle that would revive the Intel iGPU crash (#7213)."
|
||||
),
|
||||
)
|
||||
parser.add_argument(
|
||||
"--llama-backend",
|
||||
choices = ("vulkan",),
|
||||
help = (
|
||||
"Force the llama.cpp prebuilt backend. vulkan installs the upstream Vulkan "
|
||||
"bundle and records the choice so Studio updates keep it; ignored on hosts "
|
||||
"with no Vulkan prebuilt (macOS, Windows arm64). "
|
||||
"Same effect as UNSLOTH_LLAMA_BACKEND=vulkan / UNSLOTH_FORCE_VULKAN=1."
|
||||
),
|
||||
)
|
||||
resolve_group = parser.add_mutually_exclusive_group()
|
||||
resolve_group.add_argument(
|
||||
"--resolve-llama-tag",
|
||||
|
|
@ -6683,8 +6988,12 @@ def main() -> int:
|
|||
)
|
||||
# Same Vulkan routing the install path applies, so the probe's answer
|
||||
# matches what would install (an Intel/forced-Vulkan host -> upstream).
|
||||
host, repo, release_tag = _route_to_vulkan_prebuilt(
|
||||
host, args.published_repo, args.published_release_tag or "", force_cpu = _cpu_mechanism
|
||||
host, repo, release_tag, _persist_llama_backend = _route_to_vulkan_prebuilt(
|
||||
host,
|
||||
args.published_repo,
|
||||
args.published_release_tag or "",
|
||||
force_cpu = _cpu_mechanism,
|
||||
llama_backend = args.llama_backend,
|
||||
)
|
||||
try:
|
||||
_requested, plans = resolve_simple_install_release_plans(
|
||||
|
|
@ -6726,6 +7035,7 @@ def main() -> int:
|
|||
# updater re-asserts it. --cpu-fallback stays transient and heals to GPU.
|
||||
force_cpu = args.cpu_fallback or args.force_cpu,
|
||||
persist_force_cpu = args.force_cpu,
|
||||
llama_backend = args.llama_backend,
|
||||
instruction_cleanup_root = install_arg.absolute(),
|
||||
)
|
||||
return EXIT_SUCCESS
|
||||
|
|
|
|||
|
|
@ -1253,6 +1253,7 @@ def test_install_prebuilt_falls_back_to_older_release_plan(
|
|||
initial_fallback_used = False,
|
||||
existing_install_dir = None,
|
||||
force_cpu = False,
|
||||
llama_backend = None,
|
||||
):
|
||||
call_log.append((llama_tag, initial_fallback_used))
|
||||
if llama_tag == "b9002":
|
||||
|
|
@ -2457,6 +2458,7 @@ def test_install_prebuilt_skips_when_older_release_fallback_matches_existing_ins
|
|||
initial_fallback_used = False,
|
||||
existing_install_dir = None,
|
||||
force_cpu = False,
|
||||
llama_backend = None,
|
||||
):
|
||||
call_log.append(llama_tag)
|
||||
raise PrebuiltFallback("validation failed for latest release")
|
||||
|
|
@ -2605,6 +2607,7 @@ def test_install_prebuilt_skips_same_release_fallback_attempt_when_installed(
|
|||
prebuilt_fallback_used,
|
||||
quantized_path,
|
||||
force_cpu = False,
|
||||
llama_backend = None,
|
||||
):
|
||||
attempted_names.append(choice.name)
|
||||
if choice.name == first_choice.name:
|
||||
|
|
@ -2732,6 +2735,7 @@ def test_install_prebuilt_same_tag_upstream_failure_uses_older_unsloth_release_p
|
|||
initial_fallback_used = False,
|
||||
existing_install_dir = None,
|
||||
force_cpu = False,
|
||||
llama_backend = None,
|
||||
):
|
||||
attempted.append((llama_tag, release_tag, attempts[0].source_label))
|
||||
if llama_tag == "b9002":
|
||||
|
|
|
|||
|
|
@ -4823,11 +4823,93 @@ class TestApplyHostOverrides:
|
|||
assert out.has_rocm is True
|
||||
assert out.rocm_gfx_target == "gfx1200"
|
||||
|
||||
def test_forwarded_gfx_is_authoritative(self):
|
||||
# setup already applied visible-device selection; its value wins.
|
||||
host = rocm_host(rocm_gfx_target = "gfx1100")
|
||||
def test_forwarded_gfx_does_not_clobber_probed_arch(self, monkeypatch):
|
||||
# setup.ps1's pick is not fully visible-device aware (ignores CUDA_VISIBLE_DEVICES,
|
||||
# amd-smi branch drops comma masks), so when it resolved the host's OTHER physical
|
||||
# GPU it must not replace the arch detect_host() picked for the visible one.
|
||||
monkeypatch.delenv("UNSLOTH_ROCM_GFX_ARCH", raising = False)
|
||||
host = rocm_host(rocm_gfx_target = "gfx1010", rocm_gfx_targets = ["gfx1100", "gfx1010"])
|
||||
out = _apply_host_overrides(host, override_rocm_gfx = "gfx1100")
|
||||
assert out.rocm_gfx_target == "gfx1010"
|
||||
assert out.rocm_gfx_targets == ["gfx1100", "gfx1010"]
|
||||
assert out.has_rocm is True
|
||||
|
||||
def test_forwarded_gfx_absent_from_host_stays_authoritative(self, monkeypatch):
|
||||
# An arch no probe here ever reported is not a setup mispick: it is an explicit
|
||||
# --rocm-gfx for a host whose probe is wrong or stale, so it must still win.
|
||||
monkeypatch.delenv("UNSLOTH_ROCM_GFX_ARCH", raising = False)
|
||||
host = rocm_host(rocm_gfx_target = "gfx1100", rocm_gfx_targets = ["gfx1100"])
|
||||
out = _apply_host_overrides(host, override_rocm_gfx = "gfx1151")
|
||||
assert out.rocm_gfx_target == "gfx1151"
|
||||
# ... but it says which arch HIP targets, not which cards exist, so the probed
|
||||
# gfx1100 is still in the box and stays in the per-GPU list.
|
||||
assert out.rocm_gfx_targets == ["gfx1100", "gfx1151"]
|
||||
assert out.has_rocm is True
|
||||
|
||||
def test_forwarded_family_label_never_overrides_a_probed_arch(self, monkeypatch):
|
||||
# The update path re-derives --rocm-gfx from the marker's family-named asset, so a
|
||||
# family label is a bundle name, not a real arch, and must stay advisory: gfx1033 is
|
||||
# in-generation but unbuilt, so gfx103X winning would serve a bundle it cannot run.
|
||||
monkeypatch.delenv("UNSLOTH_ROCM_GFX_ARCH", raising = False)
|
||||
host = rocm_host(rocm_gfx_target = "gfx1033", rocm_gfx_targets = ["gfx1033"])
|
||||
out = _apply_host_overrides(host, override_rocm_gfx = "gfx103X")
|
||||
assert out.rocm_gfx_target == "gfx1033"
|
||||
assert out.has_rocm is True
|
||||
|
||||
def test_forwarded_family_label_still_fills_an_unprobed_arch(self, monkeypatch):
|
||||
# Negative control: with no probed arch the forward is the only source, so it
|
||||
# applies.
|
||||
monkeypatch.delenv("UNSLOTH_ROCM_GFX_ARCH", raising = False)
|
||||
out = _apply_host_overrides(cpu_host(), override_rocm_gfx = "gfx110X")
|
||||
assert out.rocm_gfx_target == "gfx110x"
|
||||
assert out.has_rocm is True
|
||||
|
||||
def test_forwarded_gfx_matching_active_keeps_physical_gfx_list(self, monkeypatch):
|
||||
# When the forward agrees with the probe the per-GPU list must survive: collapsing
|
||||
# it would hide the host's other AMD cards from the Windows auto-Vulkan floor
|
||||
# check.
|
||||
monkeypatch.delenv("UNSLOTH_ROCM_GFX_ARCH", raising = False)
|
||||
host = rocm_host(rocm_gfx_target = "gfx1010", rocm_gfx_targets = ["gfx1100", "gfx1010"])
|
||||
out = _apply_host_overrides(host, override_rocm_gfx = "gfx1010")
|
||||
assert out.rocm_gfx_target == "gfx1010"
|
||||
assert out.rocm_gfx_targets == ["gfx1100", "gfx1010"]
|
||||
|
||||
def test_forwarded_gfx_never_drops_a_probed_physical_gpu(self, monkeypatch):
|
||||
# The per-GPU list is the PHYSICAL inventory the Windows auto-Vulkan floor check
|
||||
# reads, so a forwarded arch the probe never saw must be ADDED, not replace it:
|
||||
# dropping the probe-confirmed gfx1100 would tell that check no AMD GPU on the box
|
||||
# reaches the HIP floor when one plainly does.
|
||||
monkeypatch.delenv("UNSLOTH_ROCM_GFX_ARCH", raising = False)
|
||||
host = rocm_host(rocm_gfx_target = "gfx803", rocm_gfx_targets = ["gfx1100", "gfx803"])
|
||||
out = _apply_host_overrides(host, override_rocm_gfx = "gfx900")
|
||||
assert out.rocm_gfx_target == "gfx900"
|
||||
assert out.rocm_gfx_targets == ["gfx1100", "gfx803", "gfx900"]
|
||||
|
||||
def test_forwarded_gfx_not_duplicated_when_already_probed(self, monkeypatch):
|
||||
# UNSLOTH_ROCM_GFX_ARCH makes the forward win over the probe's visible-device
|
||||
# pick, so this reaches the same branch; the list must stay deduplicated.
|
||||
monkeypatch.setenv("UNSLOTH_ROCM_GFX_ARCH", "gfx803")
|
||||
host = rocm_host(rocm_gfx_target = "gfx1100", rocm_gfx_targets = ["gfx1100", "gfx803"])
|
||||
out = _apply_host_overrides(host, override_rocm_gfx = "gfx803")
|
||||
assert out.rocm_gfx_target == "gfx803"
|
||||
assert out.rocm_gfx_targets == ["gfx1100", "gfx803"]
|
||||
|
||||
def test_forwarded_gfx_on_an_unprobed_host_lists_only_itself(self, monkeypatch):
|
||||
# Negative control for the two above: nothing probed means no inventory to
|
||||
# preserve, so the driver-only host keeps a single-entry list and auto-Vulkan.
|
||||
monkeypatch.delenv("UNSLOTH_ROCM_GFX_ARCH", raising = False)
|
||||
out = _apply_host_overrides(cpu_host(), override_rocm_gfx = "gfx803")
|
||||
assert out.rocm_gfx_target == "gfx803"
|
||||
assert out.rocm_gfx_targets == ["gfx803"]
|
||||
|
||||
def test_manual_env_override_still_wins_over_probe(self, monkeypatch):
|
||||
# UNSLOTH_ROCM_GFX_ARCH is the manual escape hatch for hosts whose arch the probes
|
||||
# get wrong, so it stays authoritative.
|
||||
monkeypatch.setenv("UNSLOTH_ROCM_GFX_ARCH", "gfx1151")
|
||||
host = rocm_host(rocm_gfx_target = "gfx1100", rocm_gfx_targets = ["gfx1100"])
|
||||
out = _apply_host_overrides(host, override_rocm_gfx = "gfx1151")
|
||||
assert out.rocm_gfx_target == "gfx1151"
|
||||
assert out.rocm_gfx_targets == ["gfx1100", "gfx1151"]
|
||||
|
||||
def test_has_rocm_only_keeps_probe_gfx(self):
|
||||
out = _apply_host_overrides(cpu_host(), override_has_rocm = True)
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue