unsloth/studio/install_llama_prebuilt.py
Daniel Han 62191c4765
Windows/WSL installer: fix winget msstore cert failure, amd-smi DiskPart prompt, and enable AMD GPU (Strix Halo gfx1151) (#5940)
* Fix Windows installer winget msstore certificate failure

`winget install` was invoked without `--source winget`, so winget also
queried the msstore source. When msstore fails certificate pinning
(error 0x8a15005e, "The server certificate did not match any of the
expected values") winget aborts and demands `--source`, so the Python
(and uv) install fails even though the package exists in the winget
source.

- Pass `--source winget` to all winget install calls (Python x2, uv).
  Both packages live in the winget source, so this is strictly correct
  and skips the failing msstore round-trip entirely.
- Add a python.org fallback (Install-PythonFromPythonOrg) that downloads
  the official installer and runs it silently per-user (no admin/UAC)
  when winget is unavailable or fails for any reason. Mirrors the
  existing uv -> astral.sh fallback so Python installs without manual
  steps. Resolves the latest 3.13.x from python.org with a pinned
  fallback, and selects the amd64/arm64/x86 installer per architecture.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Pin remaining setup.ps1 winget calls to --source winget

Two winget invocations in studio/setup.ps1 still queried all sources and
could hit the same msstore certificate-pinning failure (0x8a15005e) that
broke the Python install in install.ps1:

- `winget show Nvidia.CUDA --versions` (CUDA Toolkit version probe)
- `winget install ... ShiningLight.OpenSSL.Dev` (OpenSSL dev for llama-server)

Every other winget call in this file already passes `--source winget`
(Git, CMake, VS Build Tools, CUDA install, Node.js, and setup.ps1's own
Python 3.12 install), so these two were stragglers. Both packages live in
the winget source; pinning it makes setup robust to an unhealthy msstore
source, matching the rest of the file.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Stop amd-smi GPU probe from popping a DiskPart UAC prompt

On Windows, AMD GPU detection in install.ps1 and studio/setup.ps1 runs
`amd-smi list` / `static --asic` / `version`. amd-smi (shipped in
System32 by the Adrenalin driver) auto-elevates to read GPU/APU memory
details, surfacing a confusing DiskPart UAC prompt mid-install. The
Studio backend already documents and circuit-breaks on this in
studio/backend/utils/hardware/amd.py, but the installers did not.

Add an Invoke-AmdSmiNoElevate helper (both scripts) that runs amd-smi via
Start-Process under __COMPAT_LAYER=RunAsInvoker so it cannot auto-elevate
(no prompt), with a 30s timeout (matching amd.py) so a flaky amd-smi
cannot stall the install for minutes. On failure/timeout the existing WMI
name -> gfx fallback still resolves the arch, so detection is unchanged on
working hosts.

Verified on a Strix Halo (Radeon 8060S / gfx1151) box: the prompt is gone
and the probe is bounded.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Add experimental ROCm-on-WSL setup helper for Strix Halo (gfx1151)

install.sh already routes gfx1151 (Radeon 8060S / Strix Halo) to the
repo.amd.com/rocm/whl/gfx1151 wheels once a ROCm runtime is present, but
it does not install AMD's driver/ROCm stack -- a large, admin-gated
prerequisite. scripts/install_rocm_wsl_strixhalo.sh automates the Linux
side on a dedicated Ubuntu 24.04 WSL2 distro: ROCm 7.2 (wsl usecase), the
rocr4wsl HSA runtime, a librocdxg build, env setup, and a PyTorch gfx1151
GPU smoke test. A hard preflight refuses to run until the Adrenalin
>=26.3.1 driver is actually present, so it cannot half-install.

Procedure adapted from AMD's ROCm-on-WSL docs and community gfx1151 notes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Detect AMD GPUs by name so native Windows gets a GPU llama.cpp

The gfx-arch inference from the WMI GPU name was gated behind $HasROCm,
which the hipinfo/amd-smi probe leaves false on the common Windows case
(Adrenalin driver only, no HIP SDK -- and amd-smi often cannot read the
arch without elevation). So an AMD GPU was detected by name but never
mapped to a gfx target, --rocm-gfx was not forwarded, and studio setup
fell back to a CPU llama.cpp build.

Un-gate the inference (install.ps1 + studio/setup.ps1) so it runs whenever
an AMD GPU name is available. The inferred gfx is forwarded as --rocm-gfx,
which makes install_llama_prebuilt.py download the matching lemonade-sdk
ROCm prebuilt (e.g. llama-bNNNN-windows-rocm-gfx1151-x64.zip) -- a
GPU-accelerated llama.cpp that bundles its own ROCm runtime, so it runs
with just the Adrenalin driver. PyTorch's ROCm wheels still require a
confirmed HIP SDK ($HasROCm), so this only affects llama.cpp / inference
and never pulls broken ROCm torch.

Also broaden the name->arch table to every family lemonade ships Windows
assets for: gfx120X (RDNA 4), gfx110X (RDNA 3), gfx1151/gfx1150
(RDNA 3.5), and gfx103X (RDNA 2). Unknown names still fall back to CPU.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Suppress amd-smi DiskPart UAC prompt in the Python install/runtime paths

The earlier PowerShell guard covered install.ps1 / setup.ps1, but the
Python installer (install_llama_prebuilt.py detect_host,
install_python_stack.py ROCm probes) and the Studio backend monitor
(amd.py) also shell out to amd-smi on Windows, where it auto-elevates and
pops the same DiskPart UAC prompt mid-install / at runtime.

Inject __COMPAT_LAYER=RunAsInvoker into the amd-smi subprocess env on
Windows so it runs un-elevated (no prompt). Callers already tolerate an
empty/failed result and fall back to WMI / name detection (installer) or
the existing circuit breaker (amd.py). Gated to Windows so Linux/macOS
amd-smi behaviour is unchanged.

- install_llama_prebuilt.py: handled centrally in run_capture (covers
  detect_host's `amd-smi list` and the version probe).
- install_python_stack.py: new _amd_smi_env() helper on its 3 raw
  subprocess.run amd-smi calls.
- amd.py: merge RunAsInvoker into the existing child env.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Tighten AMD GPU name->arch patterns to avoid mismatches

The W9[0-9]{3} and RX 90[0-9]{2} patterns added for RDNA 4 were
speculative and over-broad: W9xxx would also match old GCN FirePro
W9100/W9000 cards (wrong gfx1201 -> a lemonade gfx120X download that
fails validation), and RX 90[0-9]{2} was redundant with the explicit
9070/9060 entries. Drop both; keep only confirmed RDNA 4 SKUs. Unmatched
AMD names still fall back cleanly to CPU.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Fetch the llama.cpp validation model via huggingface_hub

The prebuilt validation downloads a tiny GGUF test model from huggingface
via bare urllib. On Windows / proxy setups where the server sends an
incomplete TLS chain, urllib cannot complete the Amazon CA chain (it does
no AIA intermediate fetching) and fails with CERTIFICATE_VERIFY_FAILED, so
a perfectly good GPU prebuilt is rejected and the installer falls back to a
CPU source build.

Route the validation-model download through huggingface_hub
(hf_hub_download) -- the same mechanism Studio uses for model downloads,
which completes the chain where urllib cannot -- keeping the direct URL as
a fallback. This lets the lemonade ROCm prebuilt validate and install on
cert-restricted machines (verified: hf_hub_download succeeds where urllib
returns CERTIFICATE_VERIFY_FAILED).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Guard the remaining raw amd-smi version probe via run_capture

A ROCm-version detector in install_llama_prebuilt.py called amd-smi version through a raw subprocess.run that bypassed run_capture's Windows RunAsInvoker guard, so it still triggered the DiskPart UAC prompt during setup. Route it through run_capture like the other amd-smi calls.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Forward --rocm-gfx even when the ROCm runtime is unconfirmed

setup.ps1 forwarded --rocm-gfx (and picked the windows-hip llama.cpp
prebuilt) only inside `if ($HasROCm)`. On Adrenalin-only hosts (amd-smi
present but no HIP SDK, so $HasROCm stays false) the gfx arch was
name-inferred but never forwarded, so install_llama_prebuilt.py saw
has_rocm=False and installed the CPU build -- even though the lemonade
gfx1151 GPU prebuilt runs fine there (it bundles its own ROCm runtime;
verified: llama-cli --list-devices -> ROCm0: AMD Radeon 8060S, 69 GB).

Forward --rocm-gfx whenever a gfx arch is known (it is authoritative and
implies ROCm in install_llama_prebuilt.py), and treat a known gfx arch as
windows-hip in the existing-install mismatch check. --has-rocm stays gated
on the confirmed-runtime signal.

Verified on Radeon 8060S / gfx1151: the installer now selects, validates,
and installs llama-b1286-windows-rocm-gfx1151-x64.zip (ROCm DLLs present)
instead of the CPU build.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Install AMD ROCm PyTorch on name-inferred gfx hosts (enables Train/Export)

setup.ps1 picked the AMD ROCm PyTorch wheels only inside `if ($HasROCm ...)`.
On Adrenalin-only hosts (amd-smi present but no HIP SDK, so $HasROCm is
false) the gfx arch was name-inferred but the ROCm-wheel branch never ran,
so the host got torch+cpu. With CPU torch, torch.cuda.is_available() is
False, so the Studio backend sets CHAT_ONLY=True and hides Train/Export.

Un-gate the ROCm PyTorch index resolution on a known gfx arch (mirrors the
llama.cpp --rocm-gfx fix). AMD's per-arch Windows wheels
(repo.amd.com/rocm/whl/<gfx>) bundle the ROCm runtime, so they work without
a HIP SDK; a failed install still falls back to CPU.

Verified on Radeon 8060S / gfx1151: torch 2.11.0+rocm7.13.0 installs and
torch.cuda.is_available() -> True, device "AMD Radeon(TM) 8060S Graphics",
GPU matmul OK -> CHAT_ONLY=False -> Train/Export enabled.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Force amd-smi un-elevated process-wide in the Python installers

Guarding individual amd-smi call sites kept missing some (install_python_stack.py's probe loop and its Windows GPU re-check), so the DiskPart UAC prompt kept reappearing. Set __COMPAT_LAYER=RunAsInvoker process-wide at the top of install_python_stack.py and install_llama_prebuilt.py on Windows so every amd-smi subprocess (current and future) runs un-elevated with no per-call guard. Safe: these scripts only spawn amd-smi/rocminfo/hipinfo probes and pip/uv. setup.ps1 keeps per-call guards because it also spawns winget installers that need elevation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Fix Invoke-AmdSmiNoElevate exit code on PS 5.1 + RX 7700S arch match

Start-Process -PassThru leaves the returned process object's .ExitCode
$null after WaitForExit on Windows PowerShell 5.1, so the helper set
$LASTEXITCODE to $null and every caller's `if ($LASTEXITCODE -eq 0 ...)`
was always false -- the amd-smi GPU / gfx-token / ROCm-version detection
branch was effectively dead (masked only because the un-gated WMI
name->gfx inference still ran). Reproduced on PS 5.1.26100.

Rewrite the helper to use [System.Diagnostics.Process]::Start with a
ProcessStartInfo (UseShellExecute=false), whose .ExitCode is reliable,
with async stream reads (ReadToEndAsync) to avoid a pipe-buffer deadlock
and WaitForExit(timeout) to bound a flaky amd-smi. __COMPAT_LAYER=
RunAsInvoker (inherited via the process env) still suppresses the
auto-elevation / DiskPart prompt. Also drops the temp files and the
empty-ArgumentList edge case. Verified: exit code propagates
(7 -> $LASTEXITCODE=7), output captured, env restored.

Also fix the gfx1100 name pattern `RX 7700(?! S)` -> `RX 7700(?!S)` so the
spaceless retail name "RX 7700S" is correctly excluded (it belongs to the
gfx1102 row). Both found by PR review.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Address PR review follow-ups (install.sh table, update path, tests, WSL)

From the multi-agent PR review:

- install.sh: sync the AMD name->arch table with install.ps1 / setup.ps1
  (the bash table had drifted to the old narrow patterns). Adds RDNA 2
  (gfx103X), workstation PRO W SKUs, and more Strix Halo/Point names, and
  orders gfx1102 before gfx1100 so the spaceless retail name "RX 7700S"
  resolves correctly (bash case has no negative lookahead). AMD-ROCm-only:
  the name inference stays gated behind _has_amd_rocm_gpu(), so NVIDIA /
  CPU / macOS are unaffected.

- setup.ps1: the "dependencies up to date" fast path skipped the torch
  reinstall, so an existing user who had CPU torch (installed before
  ROCm-wheel support) stayed stuck in CHAT_ONLY. Now, when an AMD gfx arch
  is known AND the installed torch is CPU-only, don't skip -- force the
  dependency pass so the ROCm wheels install.

- scripts/install_rocm_wsl_strixhalo.sh: resolve the real /opt/rocm dir
  instead of hardcoding ROCM_VER for LD_LIBRARY_PATH / the librocdxg
  symlink (breaks if amdgpu-install lays ROCm under a patch-version dir);
  add a LIBROCDXG_REF pin knob and a "verified against" freshness header.

- tests/studio/install/test_pr5940_followups.py: cover _hf_resolve_url_parts,
  _fetch_validation_model_bytes (hf path + urllib fallback), run_capture's
  Windows-only amd-smi RunAsInvoker injection, and install.ps1 vs setup.ps1
  name-table parity (catches future drift). 14 tests, all passing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix DiskPart UAC prompt: skip amd-smi on Windows without a HIP SDK

On Windows, amd-smi re-initialises the ROCm runtime on every invocation
(even `amd-smi version`) and, on hosts without a working HIP runtime
(consumer APUs/dGPUs with only the Adrenalin driver), elevates a child
process at runtime -- popping a UAC/DiskPart prompt. amd-smi's own
manifest is asInvoker, so __COMPAT_LAYER=RunAsInvoker cannot suppress
that runtime elevation (verified: even `amd-smi version` hangs and
times out with RunAsInvoker set).

Replace the ineffective RunAsInvoker-only approach with a real gate:
only spawn amd-smi on Windows when a HIP SDK is detectable (hipinfo
present, so amd-smi runs un-elevated) or the user opts in with
UNSLOTH_ENABLE_AMD_SMI=1. The gfx arch is already resolved from WMI
name inference (forwarded via --rocm-gfx), so ROCm wheel + lemonade
llama.cpp selection is unaffected. Linux/macOS amd-smi never elevates
and is untouched (no regression). RunAsInvoker is kept as harmless
belt-and-suspenders for tools that DO use manifest elevation.

Applied consistently across:
  - studio/backend/utils/hardware/amd.py  (runtime GPU polling)
  - install.ps1, studio/setup.ps1         (install-time detection)
  - studio/install_llama_prebuilt.py      (prebuilt arch probe + version)
  - studio/install_python_stack.py        (ROCm version + arch probe)

Verified live on AMD Radeon 8060S (gfx1151), native Windows: fresh
install detects the GPU, installs ROCm torch (torch.cuda.is_available()
True), launches Studio with no DiskPart prompt, and inference, tool
calling, web search, LoRA finetuning, and GGUF export all run on the GPU.

Tests: add 6 _amd_smi_allowed() gating tests + PowerShell-installer gate
assertions; update the three amd-smi monitoring tests to opt in (they
mock amd-smi as available). Full suite: 267 passed, 2 skipped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install.sh: helpful WSL message when the GPU isn't exposed to ROCm

In WSL, an AMD GPU's ROCm-on-WSL runtime is only available with a recent
Adrenalin driver AND a distro AMD supports (currently Ubuntu 24.04). When
neither is in place, GPU detection (rocminfo/_has_amd_rocm_gpu) finds
nothing and we silently fall back to CPU.

Add an actionable hint in the CPU-fallback path, shown only on WSL and
only AFTER detection has already failed -- so it is forward-compatible:
the moment a driver/distro DOES expose the GPU (e.g. if AMD later adds
Ubuntu 26.04 support), detection succeeds and the hint never fires. The
message:
  - notes a GPU is plumbed in (/dev/dxg) but no ROCm runtime is exposed,
  - lists the two prerequisites (Adrenalin driver + Ubuntu 24.04),
  - if the distro is not 24.04, says AMD may not support it yet,
  - tells the user to `wsl --install Ubuntu-24.04` and re-run,
  - links AMD's ROCm-on-WSL guide + the experimental Strix Halo helper.

Verified live: on Ubuntu-24.04 the hint shows (version-warning omitted)
and the CPU install completes; on Ubuntu-26.04 the extra "this distro may
not be supported" line appears and points to 24.04.

Also fix the experimental scripts/install_rocm_wsl_strixhalo.sh: AMD's
repo.radeon.com/amdgpu-install/ is indexed by unified installer version
(30.30, 31.30, ...), NOT ROCm version, so the hard-coded
amdgpu-install/7.2.0/ path 404'd. Scan the installer dirs newest-first
for a noble .deb matching the target ROCm major.minor (ROCm 7.2 ->
30.30.x/amdgpu-install_7.2.x), falling back to the newest available.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* WSL: fix shortcut collision + pin ROCm-on-WSL driver reqs from AMD docs

Two WSL-related fixes informed by AMD's official ROCm-on-WSL docs and
field reports for Strix Halo / Ryzen AI Max+ (Radeon 8060S, gfx1151):

1. Shortcut collision (real bug). install.sh's WSL branch wrote
   "Unsloth Studio.lnk" to the SAME Desktop / Start Menu folder as the
   native-Windows installer (install.ps1 New-StudioShortcuts). Running
   install.sh in WSL therefore silently retargeted the native shortcut at
   the WSL launcher (wt.exe -> wsl.exe), so the desktop/start-menu icon
   stopped launching native GPU Studio. Now the WSL shortcut uses a
   DISTINCT name -- "Unsloth Studio (WSL - <distro>).lnk" -- and fetches
   the Unsloth .ico to %LOCALAPPDATA%\Unsloth Studio so it shows the
   proper icon. Native and WSL shortcuts now coexist.

2. Precise ROCm-on-WSL prerequisites. Research (AMD radeon-ryzen WSL
   compatibility matrix, gianni.rosagallina.com Feb-2026 guide,
   ROCm/ROCm#4952/#5509/#6022) confirms WSL GPU on Strix Halo requires
   AMD Adrenalin Edition >= 26.1.1 (26.2.2+ is the first production
   ROCDXG/WSL release) + ROCm 7.2.1 + Ubuntu 24.04; an older driver does
   not inject the ROCm/DXG runtime into /usr/lib/wsl/lib, so rocminfo sees
   only the CPU. install.sh's WSL hint and the experimental
   install_rocm_wsl_strixhalo.sh header/preflight now state the exact
   driver version (was a guessed ">=26.3.1"), bump ROCM_VER to 7.2.1, link
   AMD's radeon-ryzen docs, and document the known librocdxg caveat that
   usable VRAM is currently capped at the .wslconfig memory setting.

bash -n clean; install test suite 267 passed, 2 skipped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* installer: hint when the AMD driver is too old for ROCm-on-WSL

Adds a detect-and-guide hook for the optional WSL-GPU path. An AMD GPU on
native Windows can also be used inside WSL2, but only with AMD Adrenalin
Edition >= 26.2.2 (the first production ROCDXG/WSL release). Native Windows
GPU works with any recent driver, so this is purely about enabling the WSL
path.

We intentionally do NOT auto-install the driver: AMD referrer-gates driver
downloads (scripted curl/Invoke-WebRequest are blocked) and does not publish
Adrenalin via winget, so no installer can reliably fetch it -- and silently
swapping a live display driver is risky. Instead we point the user at AMD's
official download page (one click), after which the existing WSL detection
lights up automatically.

- install.ps1: new Show-AmdWslDriverHint -- when an AMD GPU is present and the
  installed driver predates the 26.2.2 release (DriverDate < 2026-02-01),
  print a concise tip with the AMD download URL. Handles DriverDate as either
  a CIM DateTime or a WMI string. Suppress with UNSLOTH_SKIP_AMD_DRIVER_HINT=1.
- install.sh (WSL hint): add the direct Adrenalin 26.2.2 download URL and note
  that AMD downloads are referrer-gated (open in a browser).

Verified: hint fires on a Sept-2025 driver, auto-suppresses on >= 2026-02-01;
install.ps1 parses; install.sh bash -n clean; suite 267 passed, 2 skipped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* install.ps1: refresh shell icon cache after creating the shortcut

After writing the Desktop / Start Menu .lnk, nudge Explorer to refresh
its icon (ie4uinit.exe -show). Without this, a stale icon cache can show
a blank shortcut icon until the next explorer restart -- most visible
when a shortcut of the same name was rewritten (e.g. a native install
followed by a WSL install, which previously shared the name; now they use
distinct names, but the cache nudge makes the icon appear immediately
regardless). Best-effort and wrapped in try/catch so it never fails the
install. The bundled unsloth.ico itself is valid (verified it renders).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* setup.ps1: don't silently CPU-build llama.cpp on an AMD GPU

For AMD, GPU acceleration comes from the lemonade ROCm prebuilt (it bundles
the ROCm runtime, no HIP SDK needed) and is the preferred/default path. The
source-build fallback is CPU-only -- a HIP/ROCm *source* build would need the
full HIP SDK + ROCm clang toolchain, which the prebuilt exists to avoid.

Previously, if an AMD-GPU host ever fell through to the source build (e.g. the
prebuilt could not be downloaded), it printed "building llama.cpp (CPU-only,
no NVIDIA GPU detected)" and quietly produced a CPU binary -- masking the lost
GPU acceleration. Now that case emits a loud [WARN] explaining the GPU prebuilt
is the AMD path and how to restore it (re-run / check network / set
UNSLOTH_LLAMA_RELEASE_TAG), so AMD never silently degrades to CPU.

No behavior change on the happy path: AMD still gets the GPU prebuilt (verified
on gfx1151: ggml-hip.dll bundled, ~80% GPU compute during inference). NVIDIA
(CUDA source build) and CPU-only hosts are unchanged.

setup.ps1 parses; install suite 267 passed, 2 skipped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* uninstall: remove shared llama.cpp build, kill lock-holders, match WSL shortcut

Three gaps found by running a real uninstall on a native-Windows + WSL host;
all fixes are scoped to Unsloth-owned paths and no-op on the other pathways
(env/custom-root, NVIDIA/AMD/CPU, Mac) so nothing else regresses.

uninstall.ps1:
  - Remove the default-mode SHARED llama.cpp build + cache. setup.ps1 installs
    them at ~/.unsloth/llama.cpp and ~/.unsloth/.cache -- SIBLINGS of studio,
    not under it -- so deleting <studio> left hundreds of MB behind. Now removed
    explicitly, then ~/.unsloth is dropped ONLY if empty (never nukes unrelated
    content). No-op in env/custom mode (llama.cpp nests under the custom root,
    removed already) and when absent. UNSLOTH_LLAMA_CPP_PATH (user-owned) is kept.
  - New _StopProcessesLockingRoots: _StopStudioProcesses only matched the venv
    unsloth/python/studio exe, so it missed (a) llama-server.exe under llama.cpp
    and (b) an orphaned multiprocessing python fork that ran from the SYSTEM
    python but loaded a venv DLL (bitsandbytes) -- on Windows an open DLL handle
    blocks the directory delete, leaving a half-removed install. The new helper
    kills any process whose image path OR loaded module is under a target root
    (module scan scoped to python/unsloth/llama-server names; vendor-agnostic).
  - _RemovePath now retries (transient post-kill handle release).

uninstall.sh:
  - Remove the default-mode ~/.unsloth/llama.cpp + ~/.unsloth/.cache; rmdir
    ~/.unsloth only if empty.
  - WSL Windows-side shortcut cleanup now matches by TARGET (any
    "Unsloth Studio*.lnk" whose target launches wsl.exe), covering both the
    legacy "Unsloth Studio.lnk" and the new "Unsloth Studio (WSL - <distro>).lnk"
    -- and never removes a native-Windows shortcut (which launches wscript.exe).

uninstall.ps1 parses; uninstall.sh passes sh -n and bash -n.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* install.ps1: invalidate Win11 Start Menu tile cache after creating shortcut

The Start Menu shortcut kept showing a blank/generic icon even after the
Explorer icon-cache rebuild, because Windows 11's StartMenuExperienceHost
keeps its OWN pre-rendered tile-icon cache
(%LOCALAPPDATA%\Packages\Microsoft.Windows.StartMenuExperienceHost_cw5n1h2txyewy\
TempState\TileCache_*.bin + StartUnifiedTileModelCache.dat), separate from
Explorer's iconcache_*.db. ie4uinit and an explorer.exe restart do not touch
it, and they don't recycle the host -- so a rewritten same-name shortcut keeps
showing the first-rendered (often the generic wscript ">") tile until the host
restarts on its own.

Fix: after creating the shortcut, drop only the Start Menu RENDER caches
(TileCache_* + StartUnifiedTileModelCache.dat) and stop StartMenuExperienceHost
(Windows auto-relaunches it), so the tile re-resolves the real icon via the
shell image factory. start2.bin (the user's pinned layout) is deliberately
preserved. Guarded by Test-Path (Windows 10 has no such host -> skipped) and
wrapped in try/catch so it can never fail the install. Windows-only
(install.ps1); no effect on Linux/macOS/Studio.

Verified live: rendering the shortcut via IShellItemImageFactory::GetImage (the
API StartMenuExperienceHost uses) returns the Unsloth sloth icon, color-matched,
after this invalidation -- previously it returned the generic script tile.

install.ps1 parses; install suite 267 passed, 2 skipped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* ROCm-on-WSL for AMD Strix Halo (gfx1151): auto-setup + runtime enablement

Make Unsloth Studio set up ROCm-on-WSL automatically for AMD Strix Halo
(Radeon 8060S / gfx1151) and use the GPU at runtime, validated end-to-end
on a Ryzen AI Max+ PRO 395 (ROCm 7.2.1 + librocdxg + Adrenalin Apr-2026):
rocminfo enumerates gfx1151, torch.cuda True, ~85.8 GB UMA pool.

Every change is a strict no-op for all other configs (NVIDIA/CUDA,
discrete + native-Linux AMD ROCm, macOS/MLX, Windows, CPU-only, non-Strix
WSL) and can never abort the installer.

- scripts/install_rocm_wsl_strixhalo.sh: rewrite to the validated recipe.
  Fixes that would have broken a working box: drop the /usr/lib/wsl/lib
  preflight (a working ROCDXG host has only d3d12/dxcore there); remove the
  obsolete rocr4wsl step (gone from the 7.2.1 repo; would hard-fail and also
  rips out the standard hsa-rocr ROCDXG needs); dynamic librocdxg soname
  (was hardcoded 1.1.0; build is 1.2.0); direct apt-repo install; Windows
  SDK auto-discovery; persist env to /etc/profile.d + ~/.bashrc; idempotent.
- install.sh: _maybe_bootstrap_rocm_wsl auto-offers/runs the helper when it
  detects a Strix Halo APU in WSL (/dev/dxg) with no ROCm runtime, then
  loads the env so detection routes to the gfx1151 wheels. Fast-path when
  already configured. Fix an inaccurate WSL hint line.
- studio/backend/main.py + worker.py: set HSA_ENABLE_DXG_DETECTION=1
  in-process before torch (gated on /dev/dxg AND librocdxg.so), so the
  worker uses the GPU even when launched outside a login shell. Mirrors the
  existing BNB_ROCM_VERSION injection.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* uninstall: clean up ROCm-on-WSL artifacts + Start Menu tile cache

- uninstall.sh: remove the ROCm-on-WSL helper artifacts -- the librocdxg
  build clone (~/.unsloth/librocdxg, which otherwise blocks the empty-dir
  rmdir of ~/.unsloth), the throwaway smoke-test venv, the persisted env
  (/etc/profile.d/unsloth-rocm-wsl.sh) and the ~/.bashrc block. The system
  ROCm userspace is a shared prereq like CUDA and is kept by default;
  UNSLOTH_UNINSTALL_ROCM=1 removes it too. No-ops on macOS / non-Strix Linux.
- uninstall.ps1: invalidate the Win11 Start Menu tile cache after removing
  the shortcut so its tile disappears promptly (mirrors install.ps1),
  preserving start2.bin.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* installer: accurate AMD ROCm messaging (HIP SDK optional, not required)

The Windows installer printed "HIP SDK not found - GPU-accelerated training
unavailable" / "ROCm wheels require the HIP SDK" whenever the HIP SDK was
absent. That is misleading: for a detected AMD GPU arch (gfx1151 etc.),
setup.ps1 installs AMD's bundled-runtime ROCm PyTorch wheels (repo.amd.com)
which ship their own ROCm runtime and do NOT need the HIP SDK -- verified
end-to-end (torch 2.11.0+rocm7.13.0, cuda True, QLoRA training on GPU) on a
Radeon 8060S with no HIP SDK installed.

Gate the GPU-detection + rocm-step messages on a detected gfx arch: when one
is known, state that GPU PyTorch uses bundled-runtime wheels and the HIP SDK
is optional; only when the arch is unknown fall back to the HIP-SDK hint.
Behavior (torch routing) is unchanged; this is messaging only. No-op for
NVIDIA/CUDA, HIP-SDK-present, and CPU paths (they hit earlier branches).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* installer: fix /opt/rocm data-loss + make WSL shortcut create/remove interop-robust

Two fixes from the 3-reviewer regression audit + live testing on a
systemd-enabled WSL distro (interop disabled):

F1 (data-loss, install_rocm_wsl_strixhalo.sh): the /opt/rocm symlink-repair
could force-delete a pre-existing REAL ROCm install. The guard only checked
that /opt/rocm is a real directory, not that it is the stray librocdxg stub.
Now it only touches /opt/rocm when it is NOT a real install (no bin/rocminfo,
bin/hipcc, or .info/version present), and MOVES it aside (rocm.unsloth-stub-bak)
instead of deleting it, so a wrong guess can never lose data.

WSL interop robustness (install.sh + uninstall.sh): both relied on
`command -v powershell.exe`, which is true even when WSL interop cannot EXECUTE
it (on systemd distros powershell.exe fails with "Exec format error"). Result:
the WSL shortcut silently failed to create (install) and to remove (uninstall).
- uninstall.sh: test that powershell.exe actually runs; if not, remove the
  "Unsloth Studio (WSL...).lnk" files directly via drvfs (/mnt/<drive>), which
  works without interop. The name is WSL-install-specific, so a native install's
  "Unsloth Studio.lnk" is never touched.
- install.sh: when the shortcut cannot be created, warn with the manual launch
  command + how to re-enable interop, instead of failing silently.

No behavior change on the interop-on path. The regression audit otherwise found
no regressions on Linux/Mac/Windows/CPU/NVIDIA install paths.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* install.sh: fast-path fully restores ROCm-on-WSL env when the drop-in is gone

Reinstall regression found by uninstall->reinstall testing: after a Studio
uninstall that removed /etc/profile.d/unsloth-rocm-wsl.sh but KEPT the shared
ROCm (the default), a non-login reinstall hit the bootstrap fast-path
(librocdxg present) and its else-branch only set HSA_ENABLE_DXG_DETECTION --
NOT PATH/LD_LIBRARY_PATH. So rocminfo was not on PATH, GPU detection failed,
and the installer fell back to CPU-only PyTorch.

Fix: when librocdxg is present but the env drop-in is missing, restore the
FULL env inline (HSA + TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL + PATH +
LD_LIBRARY_PATH) so rocminfo is found and detection routes to the GPU, and
recreate /etc/profile.d/unsloth-rocm-wsl.sh so future shells and the Studio
worker get it too. No change to the env-present fast-path or any other host.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* installer: clear Explorer icon cache so shortcut icons aren't blank

Root cause of the persistent blank Desktop + Start Menu icons: Explorer caches
each shortcut's icon in iconcache_*.db and does NOT re-read the .ico when a
same-name .lnk is recreated across reinstalls. The .ico and .lnk are correct
(the shell renders them non-blank via IShellItemImageFactory; the .ico has real
image data at 16/32/48/128 px), but the stale cache entry wins. The previous
fix only ran a weak `ie4uinit -show` + the Start Menu tile-cache clear -- it
never invalidated Explorer's icon cache, so the desktop icon stayed blank.

Fix (native install.ps1 New-StudioShortcuts AND the WSL shortcut path in
install.sh):
- ie4uinit -ClearIconCache (thorough; replaces -show as the primary refresh)
- SHChangeNotify(SHCNE_ASSOCCHANGED) to force a live desktop/taskbar refresh
  WITHOUT restarting explorer
- keep the Win11 Start Menu tile-cache invalidation (and add it to the WSL
  shortcut path too, preserving start2.bin)

Non-disruptive (no explorer restart). install.ps1 parses clean; install.sh
passes bash -n + dash -n; the heredoc-generated WSL PowerShell parses clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* installer: per-item SHChangeNotify(UPDATEITEM) reliably fixes blank icons

The blank Desktop/Start Menu shortcut icons are a stale Explorer PER-ITEM icon
cache: when a same-name .lnk is recreated across reinstalls, Explorer caches the
previously-resolved (often generic "white page") icon for that item and won't
re-extract the .ico on its own. The .ico and the .lnk's IconLocation are correct
(every icon API renders the sloth) -- only Explorer's cached display is stale.

The previous refresh (ie4uinit -ClearIconCache + a GLOBAL SHCNE_ASSOCCHANGED
broadcast) does NOT recover a stale item -- confirmed by reproduction. The
reliable, NON-disruptive fix (no explorer restart) is a PER-ITEM
SHChangeNotify(SHCNE_UPDATEITEM, SHCNF_PATHW, <lnk path>) for each created
shortcut, which forces Explorer to re-read that exact item's icon.

Verified end-to-end: deliberately staled a shortcut to the generic icon, ran the
installer's exact new refresh code, and the sloth icon recovered with NO explorer
restart (confirmed by capturing the live desktop via PrintWindow).

Applied to both native install.ps1 (New-StudioShortcuts) and the WSL shortcut
path in install.sh. Still clears the on-disk icon cache (ie4uinit) and the Win11
Start Menu tile cache (preserving start2.bin).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* uninstall: remove leftover llama.cpp .staging root so ~/.unsloth is cleaned

The llama.cpp atomic-install staging root (install_llama_prebuilt.py
INSTALL_STAGING_ROOT_NAME=.staging) is a sibling of the llama.cpp install
dir (~/.unsloth/.staging in default mode). It is normally pruned after a
successful activate, but an interrupted or retained build can leave a
<name>.staging-XXXX tree behind. The uninstallers removed llama.cpp and
.cache but not .staging, so the final empty-dir cleanup of ~/.unsloth failed
and the directory lingered. Reproduced on WSL (Ubuntu-24.04) where an empty
llama.cpp.staging-XXXX dir kept ~/.unsloth alive after uninstall.

Remove ~/.unsloth/.staging in both uninstall.sh and uninstall.ps1. No-op in
env/custom mode (staging nests under the custom root removed already) and
when absent. Cross-platform fix (the staging logic is platform-agnostic).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* installer: WSL-absent hint + fix here-string lint false positive

install.ps1: in the AMD WSL-ROCm driver hint, detect when wsl.exe is absent
and add a one-line "wsl --install -d Ubuntu-24.04" pointer so a Strix Halo
user with no WSL yet gets an actionable next step (the hint previously assumed
an Ubuntu-24.04 distro already existed). Best-effort, informational only.

test_rocm_support.py: test_no_here_strings did a crude substring check that
false-positived on the conda-style block marker
printf '# <<< Unsloth ROCm-on-WSL (gfx1151) <<<' -- a string literal written
into the /etc/profile.d drop-in, also used as a sed delimiter pair by
uninstall.sh, not a here-string. Strip quoted spans before the check so the
lint still catches a real here-string operator but ignores quoted literals.
install.sh remains POSIX-clean (sh -n / dash -n / bash -n all pass).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* installer: address PR review comments (gfx1150 mapping, amd-smi opt-out, WSL bootstrap, SDK path, make)

Apply the valid bot review findings on #5940; reject the ones that don't hold.

Fixed:
- AMD name->gfx table (setup.ps1 + install.ps1): Radeon 890M and Ryzen AI 9 HX
  370/375 are Strix POINT (gfx1150), not Strix Halo (gfx1151). Move 890M / HX 37x
  / AI 9 HX to the gfx1150 row and drop the bogus HX 38x pattern (no such Strix
  Halo SKU). Matches the runtime classifier in worker.py (890M/880M -> gfx1150;
  8060S/8050S -> gfx1151). Prevents Strix Point hosts from getting the wrong ROCm
  prebuilt/wheels.
- amd-smi opt-out (setup.ps1 + install.ps1): an explicit UNSLOTH_ENABLE_AMD_SMI=
  0/false/no/off now wins over the HIP-SDK heuristic, so a host with a HIP SDK
  binary but a broken runtime no longer gets the DiskPart/UAC prompt the opt-out
  exists to avoid.
- amd-smi warning probes (install_python_stack.py): _has_rocm_gpu and
  _detect_amd_gfx_codes now gate amd-smi behind _amd_smi_allowed() (and pass
  _amd_smi_env()), closing the last unguarded amd-smi spawn on Windows.
- WSL ROCm bootstrap (install.sh): the "already-usable ROCm?" early return now
  requires rocminfo to enumerate the real gfx1151 agent instead of the generic
  _has_amd_rocm_gpu (whose broad gfx[1-9][0-9] match accepts a fallback
  "gfx11-generic" ISA), so a Strix Halo box missing the ROCDXG bridge is no longer
  skipped. The shared helper is untouched (no gfx90a regression).
- install_rocm_wsl_strixhalo.sh:
  * Quote-safe Windows SDK discovery: the old for-in-$(ls -d "...Program Files
    (x86)/...") word-split on the space and never matched; use find + read loop.
  * Add `make` to apt prereqs (cmake only recommends it; minimal images lacked it
    and the librocdxg `make -j` build failed).
  * Verification requires gfx1151 exactly (not gfx1[0-9]) so a generic ISA or an
    unrelated RDNA GPU can't pass while the real GPU is absent.

Reviewed but NOT changed:
- "Forward inferred ROCm arch without HasROCm" (setup.ps1): already correct --
  --rocm-gfx is forwarded under `if ($script:ROCmGfxArch)`, not `if ($HasROCm)`.
- "Route inferred arch into install.ps1 torch path": not a bug -- install.ps1
  installs CPU torch as a base by design and setup.ps1 swaps in the ROCm wheel for
  the inferred arch (gate `($HasROCm -or $ROCmGfxArch) -and cpu`); verified live
  the native install ends on torch 2.11.0+rocm7.13.0.
- "$p null guard after Start-Process" (install.ps1/setup.ps1): redundant -- the
  amd-smi runner uses [Process]::Start wrapped in try/catch, so a null process
  already returns "" with LASTEXITCODE=1 (no uncaught exception).
- "ls -> find for /usr/lib/wsl/lib" (gemini): stale -- that heuristic was removed;
  only a comment about it remains.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* installer(rocm-wsl): auto-install the Windows 11 SDK via winget (fewer manual steps)

librocdxg's build needs the Windows SDK 'shared' headers on the Windows host.
Previously the helper just die()d with "install the Windows 11 SDK and re-run" if
they were missing -- a manual prerequisite that broke the otherwise-seamless
`curl ... install.sh | sh` one-liner on Strix Halo.

Now, when the headers aren't found, the helper installs the Windows 11 SDK on the
Windows host from inside WSL via winget (powershell.exe interop), then
re-discovers them. The SDK installer elevates -> ONE UAC prompt on the Windows
desktop; the headers appear under /mnt/c immediately (drvfs is live, no reboot).
The user already consented to the ROCm-on-WSL setup, so no extra prompt is added
beyond the OS UAC gate.

- New _find_win_sdk (space-safe find of the newest installed SDK 'shared' dir)
  and _install_windows_sdk_via_winget helpers.
- winget IDs tried newest-stable first: Microsoft.WindowsSDK.10.0.26100, then
  .22621. The presence of the headers (re-check) is the source of truth, not
  winget's exit code. </dev/null so winget never consumes a piped `curl|sh` stdin.
- Best-effort + non-fatal: interop-off / no-winget / declined-UAC all fall
  through to the existing clear manual-install die(). Opt out with
  UNSLOTH_SKIP_WIN_SDK_INSTALL=1.

Removes the last avoidable manual step from the WSL Strix Halo path; only the AMD
Adrenalin driver (AMD referrer-gates the download) remains manual. Verified
_find_win_sdk resolves the spaced "Program Files (x86)" path; bash -n clean; all
winget flags validated against `winget install --help`.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* installer(amd): gate install-time amd-smi probe to fix DiskPart UAC prompt

install_python_stack.py's Windows "AMD GPU detected but ROCm torch missing"
warning probe ran `amd-smi list` whenever amd-smi was on PATH -- and amd-smi
ships in C:\Windows\System32 with the AMD Adrenalin driver -- without the
_amd_smi_allowed() gate that every other amd-smi call site in the file uses.
On Adrenalin-only hosts (no HIP SDK) amd-smi elevates a child at runtime and
pops a UAC/DiskPart prompt that __COMPAT_LAYER=RunAsInvoker cannot suppress
(amd-smi's manifest is asInvoker). The probe also ran before the
ROCm-torch-installed check, so it fired on every Windows AMD install.

Gate it behind _amd_smi_allowed() and pass _amd_smi_env(), matching
_has_rocm_gpu()/_detect_amd_gfx_codes(). When skipped, the only loss is the
best-effort "AMD GPU detected" note on HIP-SDK-less hosts.

Adds a per-function AST regression test asserting every function in
install_python_stack.py that names the amd-smi command and spawns a subprocess
also references _amd_smi_allowed() (flags the pre-fix code; passes after).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* studio(cli): fix `unsloth studio stop` crashing on Windows

`stop` used the POSIX `os.kill(pid, 0)` liveness probe, but on Windows
CPython raises OSError (WinError 87, "The parameter is incorrect") for
*every* pid -- alive or dead. `stop` only catches ProcessLookupError /
PermissionError, so the OSError propagated and the command crashed with
a traceback before ever reaching its (correct) `taskkill /F` path.

Add a cross-platform `_pid_alive(pid)` helper (tasklist on Windows,
signal-0 elsewhere) and use it for both the pre-check and the post-kill
wait loop. The actual kill path is unchanged.

Verified on Windows (Python 3.13): os.kill(pid,0) raises WinError 87 for
both a live and a dead pid; `_pid_alive` returns True/False correctly and
the full stop() flow (alive -> taskkill -> dead -> "stopped") passes
end-to-end against a throwaway process.

Adds tests/studio/test_cli_studio_stop_windows.py (AST guard against a
bare os.kill(pid,0) liveness probe + mock-only _pid_alive behaviour for
the win32 tasklist branch and the POSIX signal-0 branch).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* installer(amd): fix install.sh name->arch table misrouting Strix Point to gfx1151

The bash name->arch inference table in install.sh placed Strix Point
identifiers (Radeon 890M, "Ryzen AI 9 HX 370/375", "AI 9 HX") in the
gfx1151 (Strix Halo) row, diverging from the install.ps1 / setup.ps1
PowerShell tables which correctly map them to gfx1150. It also carried a
stray "HX 38" token absent from the PowerShell source-of-truth.

Align install.sh with the PowerShell tables:
  gfx1151 row: 8060S|8050S|8040S|Strix Halo|Ryzen AI Max|AI Max
  gfx1150 row: 890M|880M|860M|840M|Strix Point|Krackan|HX 37|AI 9 HX|...

Impact is low (the bash table only feeds the display label _gpu_disp_gfx
and the "set UNSLOTH_ROCM_GFX_ARCH=..." hint; wheel selection is driven
by the detected ROCm version, not this name string) but a Strix Point
user would otherwise see/copy the wrong gfx arch.

Add a parity test (test_install_sh_name_arch_agrees_with_ps_for_strix_and_non_amd)
that parses install.sh's case table and asserts Strix Halo->gfx1151,
Strix Point->gfx1150, RX 7700S->gfx1102, and NVIDIA/Intel->no match,
cross-checking against install.ps1 (the previous parity test only
compared install.ps1 <-> setup.ps1, missing install.sh).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* setup.ps1: keep prebuilt-llama ownership guard within the test's block window

The AMD additions to the prebuilt-llama.cpp block (the windows-hip vs
windows-cpu existing-install kind validation) pushed the
install_llama_prebuilt.py invocation to ~1999 chars after the
"installing prebuilt llama.cpp bundle (preferred path)" anchor, right at
the edge of the 2000-char window that
test_setup_ps1_prebuilt_llama_cpp_has_ownership_guard slices -- so the
helper string was truncated and the test failed with "substring not
found" (CI: Repo tests (CPU)).

The ownership-guard invariant (Assert-StudioOwnedOrAbsent precedes the
install_llama_prebuilt.py call) was already satisfied; only the proximity
to the anchor regressed. Move the "installing prebuilt..." substep to
immediately before the install (after the existing-install pre-cleanup),
which also reads better (validate/clean existing -> then "installing"),
shrinking anchor->helper from 1999 to 413 chars. Behaviour is unchanged
(console message ordering only).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* install.sh: auto-run Strix Halo ROCm-on-WSL setup by default

`curl -fsSL https://unsloth.ai/install.sh | sh` should make a Strix Halo
(gfx1151) GPU usable inside WSL with no extra commands. Previously the
ROCm-on-WSL bootstrap was opt-in: it required UNSLOTH_ROCM_WSL_AUTO=1 or an
interactive [Y/n] at a TTY, and silently skipped under a pipe (no /dev/tty),
so the piped one-liner never set the GPU up automatically.

Flip it to auto-by-default for the single narrow case the existing guards
allow (WSL + Strix Halo + /dev/dxg + no usable ROCm yet) -- exactly the GPU
setup the user ran the installer for. Opt out with
UNSLOTH_SKIP_ROCM_WSL_SETUP=1. The Tauri desktop app keeps its own consent UI
(only auto-runs when it passes UNSLOTH_ROCM_WSL_AUTO=1). All hardware/OS
guards are unchanged, so non-Strix / non-WSL / NVIDIA / native-Linux / macOS /
CPU paths are unaffected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* PR comments: condense to be succinct (comments/docstrings only)

Shorten the verbose explanatory comments and docstrings this PR added across
the installer, scripts, backend shims, CLI, and tests -- tighter, fewer lines,
while preserving every non-obvious "why" (os.kill WinError 87, amd-smi
RunAsInvoker/UAC, /dev/dxg + librocdxg gating, the ROCm-on-WSL bootstrap guard
chain, ownership guards, etc.). No executable code, string literals, messages,
or behavior changed.

Verified comments-only: docstring-normalized AST equality (Python, 9 files),
non-comment token equality (PowerShell, 3 files), comment-stripped diff +
sh -n / bash -n (shell, 3 files). Behavior re-confirmed: get_torch_index_url +
gfx name->arch table 44/44 under dash & bash; rocm_support / pr5940_followups /
cli_studio_stop tests green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Installer: address PR review (amd-smi opt-out, pipefail, multi-distro, non-root)

Fixes valid findings from the Codex/Gemini PR review:
- install.ps1 / setup.ps1: gate the `amd-smi version` ROCm-version fallback with
  $amdSmiAllowed so UNSLOTH_ENABLE_AMD_SMI=0 opt-out is honored (the device
  probe was gated but this fallback wasn't), avoiding the DiskPart/UAC prompt.
- install_rocm_wsl_strixhalo.sh: make the post-verification rocminfo summary
  best-effort (|| true) so head's early pipe-close under `set -o pipefail` can't
  fail the bootstrap after gfx1151 was already enumerated; pin the Windows SDK
  `winget install` to --source winget (matches the msstore-cert fix rationale).
- install.ps1: python.org fallback installs the py launcher per-user
  (InstallLauncherAllUsers=0, avoids admin), and derives the fallback full
  version from the requested minor so a non-default UNSLOTH_PYTHON (e.g. 3.12)
  isn't silently replaced with 3.13 when the listing is unreachable.
- install.sh: recreate /etc/profile.d/unsloth-rocm-wsl.sh via `sudo tee` for a
  non-root reinstall (a plain redirect failed silently, dropping the ROCm env).
- uninstall.sh: scope WSL Windows-side shortcut removal to the current
  WSL_DISTRO_NAME (per-distro name or -d "<distro>" arg) so uninstalling one
  distro no longer deletes other distros' launchers.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Studio ROCm Windows: fix field-reported issues from Strix Halo testers

Four fixes from PR #5940 field reports (Win11 native, gfx1151):

1. bitsandbytes arch-probe spam: bnb's get_rocm_gpu_arch() runs
   hipinfo.exe via subprocess PATH at import; the AMD torch wheel ships
   hipInfo.exe in the venv Scripts dir, which is only on PATH for
   activated venvs. Every bnb import logged "Could not detect ROCm GPU
   architecture: [WinError 2]" ERROR + WARNING (even with the HIP SDK
   installed, whose bin dir is not on PATH either). Prepend the Scripts
   dir to PATH before bnb imports in main.py, worker.py, and
   install_python_stack.py, gated on the file existing (only AMD wheels
   ship it). Verified on gfx1151: ROCM_GPU_ARCH now resolves to gfx1151
   with zero errors.

2. OOM-guard double-tax on native Windows unified APUs: mem_get_info's
   total is the WDDM budget the driver grants HIP (BIOS carve + ~half
   of remaining RAM) -- the OS share is already outside it. The 0.80
   unified cap on top denied loads that fit (field report: 48.49 GiB
   budget -> "38.79 GiB allowed" OOM for a 47.29 GiB load with 48.08
   free). Use 1.0 on win32 unified; Linux keeps 0.80, discrete 0.90.

3. "Missing VRAM" confusion: log the WDDM budget vs physical RAM with
   the fix (BIOS UMA frame buffer / AMD Software Variable Graphics
   Memory) when the grant is under 75% of RAM, so a 48 GiB cap on a
   96 GiB box reads as policy, not a Studio bug.

4. llama-server fit-step crash (Qwen3.6-27B-MTP + mmproj, lemonade
   gfx1151): --fit defaults to 'on' upstream, so the fit step runs even
   when Studio already placed the model via -ngl -1, and aborts in
   ggml-cuda.cu on some ROCm hosts. Retry the spawn once with --fit off
   when the server crashes during startup and Studio's own VRAM math
   had placed the model (never when use_fit or an explicit fit flag was
   passed). Also keep the TAIL of crash output in the error log (the
   diagnostic line prints last; head-truncation cut exactly that) and
   reference the full on-disk log.

Verified live on Radeon 8060S: bnb import clean, Qwen3.5-4B-MTP loads
and generates through the new spawn loop, stub-crash retry appends
--fit off and recovers, fraction probes confirm WDDM overcommit and
sub-1.0-only enforcement on current AMD wheels.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio ROCm Windows: GPU-name fallbacks so nothing depends on amd-smi

amd-smi does not reliably exist on Windows: the HIP SDK never ships a
CLI, inbox Windows Update drivers do not, and only some full Adrenalin
packages drop amd-smi.exe into System32 (field report: fresh Win11 +
Adrenalin + HIP SDK, still no amd-smi anywhere). Make every consumer
work without it:

- install_python_stack._detect_windows_gfx_arch: two new probes after
  hipinfo/amd-smi -- (2b) the venv Scripts hipInfo.exe shipped by AMD
  torch wheels (drives `studio update` on driver-only hosts), and (4) a
  last-resort GPU marketing-name -> gfx table via WMI
  (Win32_VideoController), mirroring setup.ps1's $nameArchTable so a
  standalone repair resolves the arch with zero AMD tooling installed.

- install_llama_prebuilt._resolve_exe: also probe the venv Scripts dir
  so a standalone rerun finds hipInfo.exe without HIP_PATH.

- hardware/amd.py _run_amd_smi: which() guard before spawning --
  absence now disables the poller in one step instead of burning the
  3-strike circuit breaker on FileNotFoundError; corrected the stale
  comment claiming Adrenalin ships amd-smi.

Simulated against the real detection functions on gfx1151: amd-smi
absent, present-but-crashing (exit 1), present-but-hanging (60s sleep
vs 5-10s probe timeouts), and hard opt-out -- all resolve gfx1151, no
exceptions, bounded time. Full adversarial install (broken amd-smi
stub first on PATH + UNSLOTH_ENABLE_AMD_SMI=1, fresh uninstall first):
exit 0, name-table arch inference, lemonade gfx1151 b1292 prebuilt,
torch 2.11.0+rocm7.13.0 cuda_avail=True on the 8060S, Studio boots
healthy and stops cleanly.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: per-attempt llama-server log names + amd-smi test portability

Found by cross-platform simulation of the --fit off retry (Windows +
Linux sandboxes, real load_model with stub servers):

- llama-server log filename now carries the spawn-attempt index. The
  retry can respawn within the same epoch second; reusing the name
  opened the same file with "w" and truncated the crash log the retry
  warning had just pointed the user at (proven with a frozen
  time.time: one file, crash evidence gone; with the suffix both
  attempts keep their logs). Regression-pinned in
  test_llama_cpp_wait_for_health.py.

- test_amd_primary_gpu_with_mock now mocks shutil.which alongside
  subprocess.run: the amd-smi absence guard which()-checks before
  spawning, so on hosts without a real amd-smi (Linux CI, driver-only
  Windows) the subprocess mock was never reached and the test failed.
  Surfaced by running the suite in a clean Linux sandbox.

Simulation coverage on both OSes: 67-case platform/edge matrix
(real shipped code blocks under win32/linux/darwin spoofs: OOM-guard
fractions + VGM-hint boundary, bnb PATH-prepend gates, retry
eligibility incl. equals-forms and decoy tokens, GPU-name table
adversarial set, WMI fallback without powershell, monitor absence
semantics), 6-scenario live retry matrix (crash-once/crash-always/
exit-zero/explicit-fit/hang/log-collision) against real llama-server
spawns on Windows and WSL (GPU success legs on the 8060S), and a
3-engine browser matrix (chromium/firefox/webkit) driving the live
backend's health + authed /v1 chat completion.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: classify unified-memory via props.is_integrated first

Align the ROCm OOM-guard classifier with PR #5988's UMA gate: consult
hipDeviceProp_t.integrated (props.is_integrated) before the hardcoded
arch set. Strictly additive -- truthy upgrades to unified; 0/absent
falls through to the existing gfx1150/gfx1151 + device-name logic, so
wheels that omit or zero the field cannot downgrade the known APU set.
Extends correct unified-cap treatment to APUs outside that set (e.g.
gfx1103 Phoenix iGPUs) and keeps Studio's two unified-memory consumers
on one driver signal. Verified live on gfx1151 (is_integrated == 1 on
the AMD Windows wheel -> ('gfx1151', True) via the new path).

* AMD detection: probe rocminfo with HSA_ENABLE_DXG_DETECTION and sync setup.sh gfx table

Fleet validation on a Strix Halo WSL2 box showed the system rocminfo
(HSA 1.18, ROCm 7.2.1) only enumerates the GPU over /dev/dxg when
HSA_ENABLE_DXG_DETECTION=1, and that rocminfo can sit at /opt/rocm/bin
off PATH outside login shells. Detection probes that miss either of
these report no GPU on a working ROCDXG host and select the CPU build
even though the lemonade bundle offloads fine (95.7 tok/s measured vs
64.5 CPU on the same laptop). Seed the env (a no-op on bare metal) and
the PATH fallback in install.sh, studio/setup.sh, and the installer's
Linux rocm probe, mirroring what main.py/worker.py already do for the
runtime.

Also sync studio/setup.sh's name->gfx table with install.sh: 890M and
the HX 37/AI 9 HX SKUs are Strix Point (gfx1150, not gfx1151), RX 7700S
must match gfx1102 before the gfx1100 row, and the RDNA2/workstation
rows were missing. New parity test pins the two bash tables together so
they cannot drift again.

* Studio: persist server session logs + native-crash stacks to disk

Field report (Strix Halo, 96 GB UMA carve, WSL and native Windows):
"the studio just terminates without a warning". A native crash in the
GPU runtime kills the process with no Python traceback, and a desktop-
shortcut console closes before anything can be read. The server only
ever logged to the console, so there was nothing to send back.

run_server now tees stdout/stderr to
~/.unsloth/studio/logs/server/server-<ts>-pid<n>.log (console behavior
unchanged; file copy is best-effort), arms faulthandler at the same
file so access violations / SIGSEGV leave a stack trace on disk, and
exports PYTHONFAULTHANDLER=1 so training workers inherit crash dumps
on their captured stderr. Armed before `from main import app` so even
import-time failures leave evidence. Keeps the newest 20 session logs;
opt out with UNSLOTH_STUDIO_NO_FILE_LOG=1. Prints "Session log: <path>"
at startup so users know what to attach.

Verified on this box: a forced real segfault (faulthandler._sigsegv)
leaves the full session output plus "Fatal Python error: Segmentation
fault" and the thread stack in the file while the console shows
nothing; a normal server boot captures the startup banner and serves
health as before.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* AMD probe: honor a pre-set HSA_ENABLE_DXG_DETECTION value

Match the shell helpers, which use the parameter-default form: a user
who exports HSA_ENABLE_DXG_DETECTION=0 to deliberately hide the GPU
from DXG detection should not have the probe override it.

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: danielhanchen <michaelhan2050@gmail.com>
2026-06-10 04:24:49 -07:00

6607 lines
254 KiB
Python

#!/usr/bin/env python3
# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""Cross platform llama.cpp prebuilt installer for Unsloth Studio"""
from __future__ import annotations
import argparse
import errno
import fnmatch
import functools
import hashlib
import json
import os
import platform
import random
import re
import shutil
import site
import socket
import struct
import subprocess
import sys
import tarfile
import tempfile
import textwrap
import time
import urllib.error
import urllib.parse
import urllib.request
import zipfile
from contextlib import contextmanager
from dataclasses import dataclass, field, replace as dataclasses_replace
try:
from filelock import FileLock, Timeout as FileLockTimeout
except ImportError:
FileLock = None
FileLockTimeout = None
from pathlib import Path
from typing import Any, Iterable, Iterator
EXIT_SUCCESS = 0
EXIT_FALLBACK = 2
EXIT_ERROR = 1
EXIT_BUSY = 3
# DiskPart-prompt suppression. RunAsInvoker does NOT stop amd-smi's runtime
# elevation (its manifest is asInvoker), so this is just harmless belt-and-
# suspenders for manifest-elevating tools. The real guard is _amd_smi_allowed():
# we don't spawn amd-smi on Windows w/o a HIP SDK (or opt-in).
if platform.system() == "Windows":
os.environ.setdefault("__COMPAT_LAYER", "RunAsInvoker")
def _amd_smi_allowed() -> bool:
"""Whether it is safe to spawn amd-smi here.
On Windows w/o a working HIP runtime, amd-smi elevates a child and pops a
UAC/DiskPart prompt RunAsInvoker can't suppress. Only call it on Windows
when a HIP SDK is detectable (hipinfo present) or UNSLOTH_ENABLE_AMD_SMI=1;
Linux/macOS always allowed. When skipped, the gfx arch still arrives via the
forwarded --rocm-gfx, so prebuilt selection is unaffected.
"""
if platform.system() != "Windows":
return True
flag = os.environ.get("UNSLOTH_ENABLE_AMD_SMI", "").strip().lower()
if flag in ("1", "true", "yes", "on"):
return True
if flag in ("0", "false", "no", "off"):
return False
if shutil.which("hipinfo"):
return True
for _var in ("HIP_PATH", "HIP_PATH_57", "ROCM_PATH"):
_root = os.environ.get(_var)
if _root and os.path.isfile(os.path.join(_root, "bin", "hipinfo.exe")):
return True
return False
def windows_hidden_subprocess_kwargs() -> dict[str, object]:
"""Return Windows-only subprocess kwargs that suppress console windows."""
if sys.platform != "win32":
return {}
kwargs: dict[str, object] = {}
create_no_window = getattr(subprocess, "CREATE_NO_WINDOW", 0)
if create_no_window:
kwargs["creationflags"] = create_no_window
startupinfo_factory = getattr(subprocess, "STARTUPINFO", None)
startf_use_showwindow = getattr(subprocess, "STARTF_USESHOWWINDOW", 0)
sw_hide = getattr(subprocess, "SW_HIDE", 0)
if startupinfo_factory is not None and startf_use_showwindow:
startupinfo = startupinfo_factory()
startupinfo.dwFlags |= startf_use_showwindow
startupinfo.wShowWindow = sw_hide
kwargs["startupinfo"] = startupinfo
return kwargs
def env_int(
name: str,
default: int,
*,
minimum: int | None = None,
) -> int:
raw = os.environ.get(name)
if raw is None:
value = default
else:
try:
value = int(str(raw).strip())
except (TypeError, ValueError):
value = default
if minimum is not None:
value = max(minimum, value)
return value
# Prefer "latest" over "master": "master" bypasses the prebuilt resolver
# (no matching GitHub release), forces a source build, and causes HTTP 422
# errors. Use "master" only temporarily when the latest release lacks
# support for a new model architecture.
DEFAULT_LLAMA_TAG = os.environ.get("UNSLOTH_LLAMA_TAG", "latest")
# Default published repo for prebuilt release resolution. Linux uses
# Unsloth prebuilts; setup.sh/setup.ps1 pass --published-repo explicitly
# for macOS/Windows to override with ggml-org/llama.cpp when needed.
DEFAULT_PUBLISHED_REPO = "unslothai/llama.cpp"
DEFAULT_PUBLISHED_TAG = os.environ.get("UNSLOTH_LLAMA_RELEASE_TAG")
DEFAULT_PUBLISHED_MANIFEST_ASSET = os.environ.get(
"UNSLOTH_LLAMA_RELEASE_MANIFEST_ASSET", "llama-prebuilt-manifest.json"
)
DEFAULT_PUBLISHED_SHA256_ASSET = os.environ.get(
"UNSLOTH_LLAMA_RELEASE_SHA256_ASSET", "llama-prebuilt-sha256.json"
)
UPSTREAM_REPO = "ggml-org/llama.cpp"
UPSTREAM_RELEASES_API = f"https://api.github.com/repos/{UPSTREAM_REPO}/releases/latest"
LEMONADE_ROCM_REPO = "lemonade-sdk/llamacpp-rocm"
LEMONADE_ROCM_RELEASES_API = f"https://api.github.com/repos/{LEMONADE_ROCM_REPO}/releases/latest"
def _lemonade_release_api_for(llama_tag: str) -> str:
"""GitHub API URL for the lemonade release matching a llama.cpp tag.
"latest"/unset -> /releases/latest; a pinned tag -> the same lemonade tag.
Lemonade tracks ggml-org build tags but may lag, so a tag it skipped 404s
and the caller falls back to the upstream tarball (keeps pinned installs
reproducible). Do NOT pass a fork tag -- the fork namespace always 404s.
Tag is URL-encoded (safe="") to prevent URL injection.
"""
normalized = (llama_tag or "").strip()
if not normalized or normalized.lower() == "latest":
return LEMONADE_ROCM_RELEASES_API
return (
f"https://api.github.com/repos/{LEMONADE_ROCM_REPO}/releases/tags/"
f"{urllib.parse.quote(normalized, safe = '')}"
)
TEST_MODEL_URL = "https://huggingface.co/ggml-org/models/resolve/main/tinyllamas/stories260K.gguf"
TEST_MODEL_SHA256 = "270cba1bd5109f42d03350f60406024560464db173c0e387d91f0426d3bd256d"
VALIDATION_MODEL_CACHE_DIRNAME = ".cache"
VALIDATION_MODEL_CACHE_FILENAME = "stories260K.gguf"
INSTALL_LOCK_TIMEOUT_SECONDS = 300
INSTALL_STAGING_ROOT_NAME = ".staging"
GITHUB_AUTH_HOSTS = {"api.github.com", "github.com"}
RETRYABLE_HTTP_STATUS = {408, 429, 500, 502, 503, 504}
HTTP_FETCH_ATTEMPTS = 4
HTTP_FETCH_BASE_DELAY_SECONDS = 0.75
JSON_FETCH_ATTEMPTS = 3
DEFAULT_GITHUB_RELEASE_SCAN_MAX_PAGES = env_int(
"UNSLOTH_LLAMA_GITHUB_RELEASE_SCAN_MAX_PAGES",
5,
minimum = 1,
)
SERVER_PORT_BIND_ATTEMPTS = 3
SERVER_BIND_RETRY_WINDOW_SECONDS = 5.0
TTY_PROGRESS_START_DELAY_SECONDS = 0.5
DEFAULT_MAX_PREBUILT_RELEASE_FALLBACKS = env_int(
"UNSLOTH_LLAMA_MAX_PREBUILT_RELEASE_FALLBACKS",
2,
minimum = 1,
)
# Deeper macOS-only walk-back: upstream can ship a run of prebuilts built for
# a newer macOS than the host, caught only at validate time, so an older host
# must skip the whole run. Free on new hosts (first plan validates).
DEFAULT_MAX_MACOS_RELEASE_FALLBACKS = env_int(
"UNSLOTH_LLAMA_MAX_MACOS_RELEASE_FALLBACKS",
16,
minimum = 1,
)
# Deterministic macOS pin. At b9428 ggml-org's macOS runner moved to macOS 26
# (Tahoe), so b9428+ prebuilts load only on macOS 26+. b9415 is the last build
# stamped below 26 (arm64 minos 14, x64 minos 13.3); loads on 13.3/14/15/26.
_PINNED_MACOS_FALLBACK_TAG = "b9415"
_PINNED_MACOS_LATEST_FLOOR = (26, 0)
FORCE_COMPILE_DEFAULT_REF = os.environ.get("UNSLOTH_LLAMA_FORCE_COMPILE_REF", "master")
# sm_103 (B300 / GB300 Blackwell Ultra) is not built natively but runs on the
# bundled base compute_100 PTX, which the driver JIT-compiles forward to sm_103.
# Listed in every bundle that ships the sm_100 build (the "newer" and
# "portable" classes) so those hosts get a prebuilt, not a source compile.
DIRECT_LINUX_BUNDLE_PROFILES: dict[str, dict[str, Any]] = {
"cuda12-older": {
"runtime_line": "cuda12",
"coverage_class": "older",
"supported_sms": ["70", "75", "80", "86", "89"],
"min_sm": 70,
"max_sm": 89,
"rank": 10,
},
"cuda12-newer": {
"runtime_line": "cuda12",
"coverage_class": "newer",
"supported_sms": ["86", "89", "90", "100", "103", "120"],
"min_sm": 86,
"max_sm": 120,
"rank": 20,
},
"cuda12-portable": {
"runtime_line": "cuda12",
"coverage_class": "portable",
"supported_sms": ["70", "75", "80", "86", "89", "90", "100", "103", "120"],
"min_sm": 70,
"max_sm": 120,
"rank": 30,
},
"cuda13-older": {
"runtime_line": "cuda13",
"coverage_class": "older",
"supported_sms": ["75", "80", "86", "89"],
"min_sm": 75,
"max_sm": 89,
"rank": 40,
},
"cuda13-newer": {
"runtime_line": "cuda13",
"coverage_class": "newer",
"supported_sms": ["86", "89", "90", "100", "103", "120"],
"min_sm": 86,
"max_sm": 120,
"rank": 50,
},
"cuda13-portable": {
"runtime_line": "cuda13",
"coverage_class": "portable",
"supported_sms": ["75", "80", "86", "89", "90", "100", "103", "120"],
"min_sm": 75,
"max_sm": 120,
"rank": 60,
},
}
# Lowest CUDA major we ship prebuilts for, and the highest major we probe for
# installed runtime libraries. Detection and runtime-line derivation are
# generated per major so a new toolkit (cuda14, ...) needs no code change as
# long as llama.cpp keeps the cudart64_<major>.dll / libcudart.so.<major> naming.
_MIN_CUDA_MAJOR = 12
_MAX_PROBE_CUDA_MAJOR = 19
# Last ggml-org release whose Windows win-cuda-13 build is still sub-13.3
# (cuda-13.1, b9360, 2026-05-27). Upstream bumped win-cuda-13 to 13.3 at b9365
# and now ships only cuda-12.4 + cuda-13.3. cuda-12.4 predates Blackwell (ggml
# compiles sm_120 only at toolkit >= 12.8), so a Blackwell host on a
# 13.0/13.1/13.2 driver is gated off 13.3 and would drop to CPU-only 12.4.
# b9360 is immutable, so we pin its cuda-13.1 build (plus paired cudart) as a
# GPU fallback for exactly those hosts. See unslothai/unsloth#5887.
_PINNED_BLACKWELL_FALLBACK_TAG = "b9360"
_PINNED_BLACKWELL_FALLBACK_RUNTIME = "13.1"
# Floor at 13.0: b9360 ships native sm_120a SASS (no PTX/JIT) and a bundled
# cuda-13.1 cudart, both of which run on a CUDA 13.0 r580+ driver via CUDA
# minor-version compatibility, so the mainstream 13.0 Blackwell branch is covered.
_PINNED_BLACKWELL_DRIVER_FLOOR = (13, 0)
_BLACKWELL_MIN_SM = 120
# ggml compiles Blackwell sm_120 only at toolkit >= 12.8, so an in-release
# windows-cuda build at or above this already covers Blackwell and makes the
# pinned 13.1 fallback unnecessary (cuda-12.4 is below it).
_BLACKWELL_MIN_TOOLKIT = (12, 8)
_PINNED_BLACKWELL_LLAMA_SHA256 = "31ddb8b42d7ab4a47cab8c48c397519f580ca502df7e73f3ab396eacc16c8e8d"
_PINNED_BLACKWELL_CUDART_SHA256 = "f96935e7e385e3b2d0189239077c10fe8fd7e95690fea4afec455b1b6c7e3f18"
def _cuda_runtime_lines_for_major(major: int) -> list[str]:
"""Runtime lines a driver of this CUDA major can use, newest first down to
the minimum we ship. A driver runs its own major and any older one
(backward compatibility)."""
return [f"cuda{m}" for m in range(major, _MIN_CUDA_MAJOR - 1, -1)]
def _resolve_linux_bundle_profile(bundle_profile: str) -> "dict[str, Any] | None":
"""Profile (runtime line + sm coverage) for a linux-x64-cuda<major>-<class>
bundle. Known majors use their published coverage; an unknown future major
reuses the newest known major's coverage for the same class as a forward
default, with the post-build GPU smoke test as backstop."""
known = DIRECT_LINUX_BUNDLE_PROFILES.get(bundle_profile)
if known is not None:
return known
m = re.fullmatch(r"cuda(?P<major>\d+)-(?P<klass>older|newer|portable)", bundle_profile)
if not m:
return None
base_key = max(
(
k
for k, v in DIRECT_LINUX_BUNDLE_PROFILES.items()
if v["coverage_class"] == m.group("klass")
),
key = lambda k: int(re.match(r"cuda(\d+)-", k).group(1)),
default = None,
)
if base_key is None:
return None
profile = dict(DIRECT_LINUX_BUNDLE_PROFILES[base_key])
profile["runtime_line"] = f"cuda{m.group('major')}"
return profile
@dataclass
class HostInfo:
system: str
machine: str
is_windows: bool
is_linux: bool
is_macos: bool
is_x86_64: bool
is_arm64: bool
nvidia_smi: str | None
driver_cuda_version: tuple[int, int] | None
compute_caps: list[str]
visible_cuda_devices: str | None
has_physical_nvidia: bool
has_usable_nvidia: bool
has_rocm: bool = False
rocm_gfx_target: str | None = None
# (major, minor) from platform.mac_ver(); None off macOS or if unparseable.
# Skips a macos prebuilt whose minimum-OS exceeds this host.
macos_version: tuple[int, int] | None = None
@dataclass
class AssetChoice:
repo: str
tag: str
name: str
url: str
source_label: str
# Paired runtime archive (Windows CUDA cudart bundle). When set,
# install_from_archives also downloads it and overlays its DLLs on top
# of the main install. See unslothai/unsloth#5106.
runtime_name: str | None = None
runtime_url: str | None = None
runtime_sha256: str | None = None
is_ready_bundle: bool = False
install_kind: str = ""
bundle_profile: str | None = None
runtime_line: str | None = None
coverage_class: str | None = None
supported_sms: list[str] | None = None
min_sm: int | None = None
max_sm: int | None = None
selection_log: list[str] | None = None
expected_sha256: str | None = None
@dataclass(frozen = True)
class PublishedLlamaArtifact:
asset_name: str
install_kind: str
runtime_line: str | None
coverage_class: str | None
supported_sms: list[str]
min_sm: int | None
max_sm: int | None
bundle_profile: str | None
rank: int
@dataclass
class PublishedReleaseBundle:
repo: str
release_tag: str
upstream_tag: str
manifest_sha256: str | None = None
source_repo: str | None = None
source_repo_url: str | None = None
source_ref_kind: str | None = None
requested_source_ref: str | None = None
resolved_source_ref: str | None = None
source_commit: str | None = None
source_commit_short: str | None = None
assets: dict[str, str] = field(default_factory = dict)
manifest_asset_name: str = DEFAULT_PUBLISHED_MANIFEST_ASSET
artifacts: list[PublishedLlamaArtifact] = field(default_factory = list)
selection_log: list[str] = field(default_factory = list)
@dataclass
class LinuxCudaSelection:
attempts: list[AssetChoice]
selection_log: list[str]
@property
def primary(self) -> AssetChoice:
if not self.attempts:
raise RuntimeError("linux CUDA selection unexpectedly had no attempts")
return self.attempts[0]
@dataclass
class CudaRuntimePreference:
runtime_line: str | None
selection_log: list[str]
@dataclass(frozen = True)
class ApprovedArtifactHash:
asset_name: str
sha256: str
repo: str | None
kind: str | None
@dataclass
class ApprovedReleaseChecksums:
repo: str
release_tag: str
upstream_tag: str
source_repo: str | None = None
source_repo_url: str | None = None
source_ref_kind: str | None = None
requested_source_ref: str | None = None
resolved_source_ref: str | None = None
source_commit: str | None = None
source_commit_short: str | None = None
artifacts: dict[str, ApprovedArtifactHash] = field(default_factory = dict)
@dataclass(frozen = True)
class ResolvedPublishedRelease:
bundle: PublishedReleaseBundle
checksums: ApprovedReleaseChecksums
@dataclass(frozen = True)
class SourceBuildPlan:
source_url: str
source_ref: str
source_ref_kind: str
compatibility_upstream_tag: str
source_repo: str | None = None
source_repo_url: str | None = None
requested_source_ref: str | None = None
resolved_source_ref: str | None = None
source_commit: str | None = None
@dataclass(frozen = True)
class InstallReleasePlan:
requested_tag: str
llama_tag: str
release_tag: str
attempts: list[AssetChoice]
approved_checksums: ApprovedReleaseChecksums
class PrebuiltFallback(RuntimeError):
pass
class BusyInstallConflict(RuntimeError):
pass
class ExistingInstallSatisfied(RuntimeError):
def __init__(self, choice: AssetChoice, used_fallback: bool):
super().__init__(f"existing install already matches candidate {choice.name}")
self.choice = choice
self.used_fallback = used_fallback
def _os_error_messages(exc: BaseException) -> list[str]:
messages: list[str] = []
if isinstance(exc, OSError):
for value in (
getattr(exc, "strerror", None),
getattr(exc, "filename", None),
getattr(exc, "filename2", None),
):
if isinstance(value, str) and value:
messages.append(value)
text = str(exc)
if text:
messages.append(text)
return [message.lower() for message in messages if message]
def is_busy_lock_error(exc: BaseException) -> bool:
if isinstance(exc, BusyInstallConflict):
return True
if isinstance(exc, OSError):
if exc.errno in {
errno.EACCES,
errno.EBUSY,
errno.EPERM,
errno.ETXTBSY,
}:
return True
if getattr(exc, "winerror", None) in {5, 32, 145}:
return True
for message in _os_error_messages(exc):
if any(
needle in message
for needle in (
"access is denied",
"being used by another process",
"device or resource busy",
"permission denied",
"text file busy",
"file is in use",
"process cannot access the file",
"cannot create a file when that file already exists",
)
):
return True
return False
def log(message: str) -> None:
print(f"[llama-prebuilt] {message}", file = sys.stderr)
def log_lines(lines: Iterable[str]) -> None:
for line in lines:
log(line)
def parsed_hostname(url: str | None) -> str | None:
if not url:
return None
try:
hostname = urllib.parse.urlparse(url).hostname
except Exception:
return None
if not hostname:
return None
return hostname.lower()
def should_send_github_auth(url: str | None) -> bool:
return parsed_hostname(url) in GITHUB_AUTH_HOSTS
def auth_headers(url: str | None = None) -> dict[str, str]:
headers = {
"User-Agent": "unsloth-studio-llama-prebuilt",
}
token = os.environ.get("GH_TOKEN") or os.environ.get("GITHUB_TOKEN")
if token and should_send_github_auth(url):
headers["Authorization"] = f"Bearer {token}"
return headers
def github_api_headers(url: str | None = None) -> dict[str, str]:
return {
"Accept": "application/vnd.github+json",
**auth_headers(url),
}
def is_github_api_url(url: str | None) -> bool:
return parsed_hostname(url) == "api.github.com"
def is_retryable_url_error(exc: Exception) -> bool:
if isinstance(exc, urllib.error.HTTPError):
# GitHub returns 403 (not 429) when the API rate limit is hit.
# Anonymous calls share a 60-req/hour bucket per runner IP, which
# CI fleets exhaust trivially. Treat 403 against api.github.com as
# retryable so we get a backoff cycle or two before the source-build
# fallback fires; sleep_backoff honours Retry-After /
# X-RateLimit-Reset for accurate waits. Real 403s on other hosts
# (private artefact downloads, auth failures) stay non-retryable.
if exc.code == 403:
return is_github_api_url(getattr(exc, "url", None))
return exc.code in RETRYABLE_HTTP_STATUS
if isinstance(exc, urllib.error.URLError):
return True
if isinstance(exc, TimeoutError):
return True
if isinstance(exc, socket.timeout):
return True
return False
_RATE_LIMIT_WAIT_CAP_SECONDS = 60.0
def _http_error_retry_delay(exc: Exception) -> float | None:
"""Extract a recommended wait from rate-limit headers on a 403/429.
Returns None when no header is present or the wait exceeds
_RATE_LIMIT_WAIT_CAP_SECONDS (the caller should not block then -- the
source-build fallback is faster).
"""
if not isinstance(exc, urllib.error.HTTPError):
return None
headers = getattr(exc, "headers", None)
if headers is None:
return None
retry_after = headers.get("Retry-After")
if retry_after and retry_after.strip().isdigit():
wait = float(retry_after.strip())
return wait if wait <= _RATE_LIMIT_WAIT_CAP_SECONDS else None
rate_reset = headers.get("X-RateLimit-Reset")
if rate_reset and rate_reset.strip().isdigit():
wait = float(rate_reset.strip()) - time.time()
if 0.0 < wait <= _RATE_LIMIT_WAIT_CAP_SECONDS:
return wait + 1.0 # +1s of slack so the bucket is fresh
return None
def sleep_backoff(
attempt: int,
*,
base_delay: float = HTTP_FETCH_BASE_DELAY_SECONDS,
exc: Exception | None = None,
) -> None:
delay = base_delay * (2 ** max(attempt - 1, 0))
header_delay = _http_error_retry_delay(exc) if exc is not None else None
if header_delay is not None:
delay = max(delay, header_delay)
delay += random.uniform(0.0, 0.2)
time.sleep(delay)
def atomic_write_bytes(destination: Path, data: bytes) -> None:
destination.parent.mkdir(parents = True, exist_ok = True)
with tempfile.NamedTemporaryFile(
prefix = destination.name + ".tmp-",
dir = destination.parent,
delete = False,
) as handle:
tmp_path = Path(handle.name)
handle.write(data)
handle.flush()
os.fsync(handle.fileno())
os.replace(tmp_path, destination)
def atomic_replace_from_tempfile(tmp_path: Path, destination: Path) -> None:
destination.parent.mkdir(parents = True, exist_ok = True)
os.replace(tmp_path, destination)
def source_archive_logical_name(upstream_tag: str) -> str:
return f"llama.cpp-source-{upstream_tag}.tar.gz"
def exact_source_archive_logical_name(source_commit: str) -> str:
return f"llama.cpp-source-commit-{source_commit}.tar.gz"
def sha256_file(path: Path) -> str:
digest = hashlib.sha256()
with path.open("rb") as handle:
for chunk in iter(lambda: handle.read(1024 * 1024), b""):
digest.update(chunk)
return digest.hexdigest()
def sha256_bytes(data: bytes) -> str:
return hashlib.sha256(data).hexdigest()
def normalize_sha256_digest(value: str | None) -> str | None:
if not isinstance(value, str) or not value:
return None
lowered = value.lower()
if lowered.startswith("sha256:"):
lowered = lowered.split(":", 1)[1]
if len(lowered) != 64 or any(ch not in "0123456789abcdef" for ch in lowered):
return None
return lowered
def normalize_source_ref_kind(value: str | None) -> str | None:
if not isinstance(value, str):
return None
normalized = value.strip().lower()
if normalized in {"tag", "branch", "pull", "commit", "custom"}:
return normalized
return None
def normalize_source_commit(value: str | None) -> str | None:
if not isinstance(value, str):
return None
normalized = value.strip().lower()
if len(normalized) < 7 or len(normalized) > 40:
return None
if any(ch not in "0123456789abcdef" for ch in normalized):
return None
return normalized
def validate_schema_version(payload: dict[str, Any], *, label: str) -> None:
schema_version = payload.get("schema_version")
if schema_version is None:
return
try:
normalized = int(schema_version)
except (TypeError, ValueError) as exc:
raise RuntimeError(f"{label} schema_version was not an integer") from exc
if normalized != 1:
raise RuntimeError(f"{label} schema_version={normalized} is unsupported")
def repo_slug_from_source(value: str | None) -> str | None:
if not isinstance(value, str):
return None
normalized = value.strip()
if not normalized:
return None
normalized = normalized.removesuffix(".git")
if normalized.startswith("https://github.com/"):
slug = normalized[len("https://github.com/") :]
elif normalized.startswith("http://github.com/"):
slug = normalized[len("http://github.com/") :]
elif normalized.startswith("git@github.com:"):
slug = normalized[len("git@github.com:") :]
else:
slug = normalized
slug = slug.strip("/")
parts = slug.split("/")
if len(parts) != 2 or not all(parts):
return None
return f"{parts[0]}/{parts[1]}"
def source_url_from_repo_slug(repo_slug: str | None) -> str | None:
if not isinstance(repo_slug, str) or not repo_slug:
return None
return f"https://github.com/{repo_slug}"
def source_repo_clone_url(repo: str | None, repo_url: str | None) -> str | None:
if isinstance(repo_url, str) and repo_url.strip():
return repo_url.strip().removesuffix(".git")
return source_url_from_repo_slug(repo_slug_from_source(repo))
def infer_source_ref_kind(ref: str | None) -> str:
if not isinstance(ref, str):
return "tag"
normalized = ref.strip()
lowered = normalized.lower()
if not normalized:
return "tag"
if lowered.startswith("refs/pull/") or lowered.startswith("pull/"):
return "pull"
if (
lowered.startswith("refs/heads/")
or lowered in {"main", "master", "head"}
or lowered.startswith("origin/")
):
return "branch"
normalized_commit = normalize_source_commit(normalized)
if normalized_commit is not None:
return "commit"
return "tag"
def normalized_ref_aliases(ref: str | None) -> set[str]:
if not isinstance(ref, str):
return set()
normalized = ref.strip()
if not normalized:
return set()
aliases = {normalized}
lowered = normalized.lower()
commit = normalize_source_commit(normalized)
if commit is not None:
aliases.add(commit)
if lowered.startswith("refs/heads/"):
aliases.add(normalized.split("/", 2)[2])
elif "/" not in normalized and infer_source_ref_kind(normalized) == "branch":
aliases.add(f"refs/heads/{normalized}")
if lowered.startswith("refs/pull/"):
aliases.add(normalized.removeprefix("refs/"))
elif lowered.startswith("pull/"):
aliases.add(f"refs/{normalized}")
return aliases
def refs_match(candidate_ref: str | None, requested_ref: str | None) -> bool:
candidate_aliases = normalized_ref_aliases(candidate_ref)
requested_aliases = normalized_ref_aliases(requested_ref)
if not candidate_aliases or not requested_aliases:
return False
if candidate_aliases & requested_aliases:
return True
candidate_commit = normalize_source_commit(candidate_ref)
requested_commit = normalize_source_commit(requested_ref)
if candidate_commit and requested_commit:
return candidate_commit.startswith(requested_commit) or requested_commit.startswith(
candidate_commit
)
return False
def checkout_friendly_ref(ref_kind: str | None, ref: str | None) -> str | None:
"""Normalize a source ref to a form ``git clone --branch`` accepts.
Fully qualified branch refs (``refs/heads/main``) are stripped to
``main``; tag refs (``refs/tags/b8508``) to ``b8508``. Pull refs
(``refs/pull/123/head``) are left as-is since they are fetched
explicitly rather than cloned with ``--branch``.
"""
if not isinstance(ref, str) or not ref:
return ref
lowered = ref.lower()
if ref_kind == "branch" and lowered.startswith("refs/heads/"):
return ref.split("/", 2)[2]
if ref_kind == "tag" and lowered.startswith("refs/tags/"):
return ref.split("/", 2)[2]
return ref
def windows_cuda_upstream_asset_names(llama_tag: str, runtime: str) -> list[str]:
return [
f"llama-{llama_tag}-bin-win-cuda-{runtime}-x64.zip",
f"cudart-llama-bin-win-cuda-{runtime}-x64.zip",
]
def windows_cuda_asset_aliases(
asset_name: str, *, compatibility_tag: str | None = None
) -> list[str]:
aliases: list[str] = []
legacy_match = re.fullmatch(
r"llama-(?P<tag>[^/]+)-bin-win-cuda-(?P<runtime>\d+\.\d+)-x64\.zip",
asset_name,
)
if legacy_match:
runtime = legacy_match.group("runtime")
aliases.append(f"cudart-llama-bin-win-cuda-{runtime}-x64.zip")
if compatibility_tag:
aliases.append(f"llama-{compatibility_tag}-bin-win-cuda-{runtime}-x64.zip")
return aliases
current_match = re.fullmatch(
r"cudart-llama-bin-win-cuda-(?P<runtime>\d+\.\d+)-x64\.zip",
asset_name,
)
if current_match and compatibility_tag:
runtime = current_match.group("runtime")
aliases.append(f"llama-{compatibility_tag}-bin-win-cuda-{runtime}-x64.zip")
return aliases
def _published_windows_cuda_runtime(
upstream_assets: dict[str, str], major: int, driver: tuple[int, int] | None
) -> str | None:
"""Highest cuda-<major>.<minor> published upstream that `driver` can run by
default CUDA compatibility, i.e. (major, minor) <= driver. None if nothing
qualifies. Gating on the driver (not just the major) keeps a 13.3 build off
a 13.1-only driver, where it would rely on the unguaranteed
minor-version-compatibility path."""
if driver is None:
return None
best: int | None = None
for name in upstream_assets:
m = re.search(r"-bin-win-cuda-(\d+)\.(\d+)-x64\.zip$", name)
if m and int(m.group(1)) == major:
minor = int(m.group(2))
if (major, minor) <= driver and (best is None or minor > best):
best = minor
return f"{major}.{best}" if best is not None else None
def format_byte_count(num_bytes: float) -> str:
units = ["B", "KiB", "MiB", "GiB", "TiB"]
value = float(num_bytes)
for unit in units:
if abs(value) < 1024.0 or unit == units[-1]:
if unit == "B":
return f"{int(value)} {unit}"
return f"{value:.1f} {unit}"
value /= 1024.0
return f"{num_bytes:.1f} B"
class DownloadProgress:
def __init__(self, label: str, total_bytes: int | None) -> None:
self.label = label
self.total_bytes = total_bytes if total_bytes and total_bytes > 0 else None
self.start_time = time.monotonic()
self.last_emit = 0.0
term_ok = os.environ.get("TERM", "").lower() != "dumb"
self.stream = (
sys.stderr if sys.stderr.isatty() else sys.stdout if sys.stdout.isatty() else sys.stderr
)
self.is_tty = term_ok and self.stream.isatty()
self.completed = False
self.last_milestone_percent = -1
self.last_milestone_bytes = 0
self.has_rendered_tty_progress = False
def _render(
self,
downloaded_bytes: int,
*,
final: bool = False,
) -> str:
elapsed = max(time.monotonic() - self.start_time, 1e-6)
speed = downloaded_bytes / elapsed
speed_text = f"{format_byte_count(speed)}/s"
if self.total_bytes is not None:
percent = min(100.0, (downloaded_bytes / self.total_bytes) * 100.0)
return (
f"{self.label}: {percent:5.1f}% "
f"({format_byte_count(downloaded_bytes)}/{format_byte_count(self.total_bytes)}) "
f"at {speed_text}"
)
if final:
return f"{self.label}: {format_byte_count(downloaded_bytes)} downloaded at {speed_text}"
return f"{self.label}: {format_byte_count(downloaded_bytes)} downloaded at {speed_text}"
def update(self, downloaded_bytes: int) -> None:
now = time.monotonic()
if self.is_tty:
elapsed = now - self.start_time
if not self.has_rendered_tty_progress:
if self.total_bytes is not None and downloaded_bytes >= self.total_bytes:
return
if elapsed < TTY_PROGRESS_START_DELAY_SECONDS:
return
min_interval = 0.2
if (
self.has_rendered_tty_progress
and not self.completed
and (now - self.last_emit) < min_interval
):
return
self.last_emit = now
line = self._render(downloaded_bytes)
self.stream.write("\r\033[K" + line)
self.stream.flush()
self.has_rendered_tty_progress = True
return
should_emit = False
if self.total_bytes is not None:
percent = int((downloaded_bytes * 100) / max(self.total_bytes, 1))
milestone_percent = min((percent // 25) * 25, 100)
if milestone_percent > self.last_milestone_percent and milestone_percent < 100:
self.last_milestone_percent = milestone_percent
should_emit = True
else:
byte_step = 25 * 1024 * 1024
if (
downloaded_bytes - self.last_milestone_bytes >= byte_step
and (now - self.last_emit) >= 5.0
):
self.last_milestone_bytes = downloaded_bytes
should_emit = True
if not should_emit:
return
self.last_emit = now
self.stream.write(self._render(downloaded_bytes) + "\n")
self.stream.flush()
def finish(self, downloaded_bytes: int) -> None:
self.completed = True
line = self._render(downloaded_bytes, final = True)
if self.is_tty:
if not self.has_rendered_tty_progress:
return
self.stream.write("\r\033[K")
else:
self.stream.write(line + "\n")
self.stream.flush()
def download_label_from_url(url: str) -> str:
name = Path(urllib.parse.urlparse(url).path).name
return name or url
def download_bytes(
url: str,
*,
timeout: int = 120,
attempts: int = HTTP_FETCH_ATTEMPTS,
headers: dict[str, str] | None = None,
progress_label: str | None = None,
) -> bytes:
last_exc: Exception | None = None
for attempt in range(1, attempts + 1):
try:
request = urllib.request.Request(url, headers = headers or auth_headers(url))
with urllib.request.urlopen(request, timeout = timeout) as response:
total_bytes: int | None = None
content_length = response.headers.get("Content-Length")
if content_length and content_length.isdigit():
total_bytes = int(content_length)
progress = DownloadProgress(progress_label, total_bytes) if progress_label else None
data = bytearray()
while True:
chunk = response.read(1024 * 1024)
if not chunk:
break
data.extend(chunk)
if progress is not None:
progress.update(len(data))
if progress is not None:
progress.finish(len(data))
return bytes(data)
except Exception as exc:
last_exc = exc
if attempt >= attempts or not is_retryable_url_error(exc):
raise
log(f"fetch failed ({attempt}/{attempts}) for {url}: {exc}; retrying")
sleep_backoff(attempt, exc = exc)
assert last_exc is not None
raise last_exc
def fetch_json(url: str) -> Any:
attempts = JSON_FETCH_ATTEMPTS if is_github_api_url(url) else 1
last_decode_exc: Exception | None = None
for attempt in range(1, attempts + 1):
try:
data = download_bytes(
url,
timeout = 30,
headers = github_api_headers(url) if is_github_api_url(url) else auth_headers(url),
)
except urllib.error.HTTPError as exc:
if exc.code == 403 and is_github_api_url(url):
hint = ""
if not (os.environ.get("GH_TOKEN") or os.environ.get("GITHUB_TOKEN")):
hint = "; set GH_TOKEN or GITHUB_TOKEN to avoid GitHub API rate limits"
raise RuntimeError(f"GitHub API returned 403 for {url}{hint}") from exc
raise
if not data:
last_decode_exc = RuntimeError(f"downloaded empty JSON payload from {url}")
else:
try:
payload = json.loads(data.decode("utf-8"))
except (UnicodeDecodeError, json.JSONDecodeError) as exc:
last_decode_exc = RuntimeError(f"downloaded invalid JSON from {url}: {exc}")
else:
if not isinstance(payload, dict) and not isinstance(payload, list):
raise RuntimeError(
f"downloaded unexpected JSON type from {url}: {type(payload).__name__}"
)
return payload
if attempt >= attempts:
assert last_decode_exc is not None
raise last_decode_exc
log(f"json fetch failed ({attempt}/{attempts}) for {url}; retrying")
sleep_backoff(attempt)
assert last_decode_exc is not None
raise last_decode_exc
def download_file(url: str, destination: Path) -> None:
destination.parent.mkdir(parents = True, exist_ok = True)
last_exc: Exception | None = None
for attempt in range(1, HTTP_FETCH_ATTEMPTS + 1):
tmp_path: Path | None = None
try:
request = urllib.request.Request(url, headers = auth_headers(url))
with tempfile.NamedTemporaryFile(
prefix = destination.name + ".tmp-",
dir = destination.parent,
delete = False,
) as handle:
tmp_path = Path(handle.name)
with urllib.request.urlopen(request, timeout = 120) as response:
total_bytes: int | None = None
content_length = response.headers.get("Content-Length")
if content_length and content_length.isdigit():
total_bytes = int(content_length)
progress = DownloadProgress(f"Downloading {destination.name}", total_bytes)
downloaded_bytes = 0
while True:
chunk = response.read(1024 * 1024)
if not chunk:
break
handle.write(chunk)
downloaded_bytes += len(chunk)
progress.update(downloaded_bytes)
progress.finish(downloaded_bytes)
handle.flush()
os.fsync(handle.fileno())
if not tmp_path.exists() or tmp_path.stat().st_size == 0:
raise RuntimeError(f"downloaded empty file from {url}")
atomic_replace_from_tempfile(tmp_path, destination)
return
except Exception as exc:
last_exc = exc
if tmp_path is not None:
try:
tmp_path.unlink(missing_ok = True)
except Exception:
pass
if attempt >= HTTP_FETCH_ATTEMPTS or not is_retryable_url_error(exc):
raise
log(f"download failed ({attempt}/{HTTP_FETCH_ATTEMPTS}) for {url}: {exc}; retrying")
sleep_backoff(attempt, exc = exc)
assert last_exc is not None
raise last_exc
def download_file_verified(
url: str, destination: Path, *, expected_sha256: str | None, label: str
) -> None:
normalized_expected = normalize_sha256_digest(expected_sha256)
if not normalized_expected:
download_file(url, destination)
log(f"downloaded {label} without a published sha256; relying on install validation")
return
for attempt in range(1, 3):
download_file(url, destination)
actual_sha256 = sha256_file(destination)
if actual_sha256 == normalized_expected:
log(f"verified {label} sha256={actual_sha256}")
return
log(
f"{label} checksum mismatch on attempt {attempt}/2: "
f"expected={normalized_expected} actual={actual_sha256}"
)
destination.unlink(missing_ok = True)
if attempt == 2:
raise PrebuiltFallback(
f"{label} checksum mismatch after retry: expected={normalized_expected} actual={actual_sha256}"
)
log(f"retrying {label} download after checksum mismatch")
def upstream_source_archive_urls(tag: str) -> list[str]:
encoded_tag = urllib.parse.quote(tag, safe = "")
return [
f"https://codeload.github.com/{UPSTREAM_REPO}/tar.gz/refs/tags/{encoded_tag}",
f"https://github.com/{UPSTREAM_REPO}/archive/refs/tags/{encoded_tag}.tar.gz",
]
def commit_source_archive_urls(repo: str, source_commit: str) -> list[str]:
encoded_commit = urllib.parse.quote(source_commit, safe = "")
return [
f"https://codeload.github.com/{repo}/tar.gz/{encoded_commit}",
f"https://github.com/{repo}/archive/{encoded_commit}.tar.gz",
]
def github_release_assets(repo: str, tag: str) -> dict[str, str]:
payload = fetch_json(
f"https://api.github.com/repos/{repo}/releases/tags/{urllib.parse.quote(tag, safe = '')}"
)
if not isinstance(payload, dict):
raise RuntimeError(f"unexpected release payload for {repo}@{tag}")
return release_asset_map(payload)
def github_release(repo: str, tag: str) -> dict[str, Any]:
payload = fetch_json(
f"https://api.github.com/repos/{repo}/releases/tags/{urllib.parse.quote(tag, safe = '')}"
)
if not isinstance(payload, dict):
raise RuntimeError(f"unexpected release payload for {repo}@{tag}")
return payload
def github_releases(
repo: str,
*,
per_page: int = 100,
max_pages: int = 0,
) -> list[dict[str, Any]]:
releases: list[dict[str, Any]] = []
page = 1
while True:
payload = fetch_json(
f"https://api.github.com/repos/{repo}/releases?per_page={per_page}&page={page}"
)
if not isinstance(payload, list):
raise RuntimeError(f"unexpected releases payload for {repo}")
page_items = [item for item in payload if isinstance(item, dict)]
releases.extend(page_items)
if len(payload) < per_page:
break
page += 1
if max_pages > 0 and page > max_pages:
break
return releases
def latest_upstream_release_tag() -> str:
payload = fetch_json(UPSTREAM_RELEASES_API)
tag = payload.get("tag_name")
if not isinstance(tag, str) or not tag:
raise RuntimeError(f"latest release tag was missing from {UPSTREAM_RELEASES_API}")
return tag
def is_release_tag_like(value: str | None) -> bool:
return isinstance(value, str) and bool(re.fullmatch(r"b\d+", value.strip()))
def release_time_sort_key(release: dict[str, Any]) -> tuple[str, int]:
published_at = release.get("published_at")
created_at = release.get("created_at")
release_id = release.get("id")
timestamp = (
published_at
if isinstance(published_at, str) and published_at
else created_at
if isinstance(created_at, str) and created_at
else ""
)
try:
normalized_id = int(release_id)
except (TypeError, ValueError):
normalized_id = 0
return (timestamp, normalized_id)
def iter_release_payloads_by_time(
repo: str,
published_release_tag: str = "",
requested_tag: str = "",
) -> Iterable[dict[str, Any]]:
if published_release_tag:
yield github_release(repo, published_release_tag)
return
if requested_tag and requested_tag != "latest" and is_release_tag_like(requested_tag):
try:
yield github_release(repo, requested_tag)
return
except urllib.error.HTTPError as exc:
if exc.code == 404:
log(f"release tag {requested_tag} not found in {repo}; scanning recent releases")
else:
raise
except Exception:
raise
releases = [
release
for release in github_releases(repo, max_pages = DEFAULT_GITHUB_RELEASE_SCAN_MAX_PAGES)
if isinstance(release, dict) and not release.get("draft") and not release.get("prerelease")
]
releases.sort(key = release_time_sort_key, reverse = True)
for release in releases:
yield release
def direct_release_matches_request(*, release_tag: str, llama_tag: str, requested_tag: str) -> bool:
if requested_tag == "latest":
return True
for candidate in (release_tag, llama_tag):
if refs_match(candidate, requested_tag):
return True
return False
def synthetic_checksums_for_release(
repo: str, release_tag: str, upstream_tag: str
) -> ApprovedReleaseChecksums:
return ApprovedReleaseChecksums(
repo = repo,
release_tag = release_tag,
upstream_tag = upstream_tag,
artifacts = {},
)
def parse_direct_linux_release_bundle(
repo: str, release: dict[str, Any]
) -> PublishedReleaseBundle | None:
release_tag = release.get("tag_name")
if not isinstance(release_tag, str) or not release_tag:
return None
assets = release_asset_map(release)
artifacts: list[PublishedLlamaArtifact] = []
inferred_labels: list[str] = []
linux_asset_re = re.compile(
r"^app-(?P<label>.+)-(?P<target>linux-x64(?:-cpu)?|linux-x64-cuda\d+-(?:older|newer|portable))\.tar\.gz$"
)
for asset_name in sorted(assets):
match = linux_asset_re.fullmatch(asset_name)
if not match:
continue
inferred_labels.append(match.group("label"))
target = match.group("target")
if target in {"linux-x64", "linux-x64-cpu"}:
artifacts.append(
PublishedLlamaArtifact(
asset_name = asset_name,
install_kind = "linux-cpu",
runtime_line = None,
coverage_class = None,
supported_sms = [],
min_sm = None,
max_sm = None,
bundle_profile = None,
rank = 1000,
)
)
continue
bundle_profile = target.removeprefix("linux-x64-")
profile = _resolve_linux_bundle_profile(bundle_profile)
if profile is None:
continue
artifacts.append(
PublishedLlamaArtifact(
asset_name = asset_name,
install_kind = "linux-cuda",
runtime_line = str(profile["runtime_line"]),
coverage_class = str(profile["coverage_class"]),
supported_sms = [str(value) for value in profile["supported_sms"]],
min_sm = int(profile["min_sm"]),
max_sm = int(profile["max_sm"]),
bundle_profile = bundle_profile,
rank = int(profile["rank"]),
)
)
if not artifacts:
return None
upstream_tag = (
release_tag
if is_release_tag_like(release_tag)
else inferred_labels[0]
if len(set(inferred_labels)) == 1 and inferred_labels
else release_tag
)
selection_log = [
f"published_release: repo={repo}",
f"published_release: tag={release_tag}",
f"published_release: upstream_tag={upstream_tag}",
"published_release: direct_asset_scan=linux",
]
return PublishedReleaseBundle(
repo = repo,
release_tag = release_tag,
upstream_tag = upstream_tag,
assets = assets,
manifest_asset_name = DEFAULT_PUBLISHED_MANIFEST_ASSET,
artifacts = artifacts,
selection_log = selection_log,
)
def direct_linux_release_plan(
release: dict[str, Any], host: HostInfo, repo: str, requested_tag: str
) -> InstallReleasePlan | None:
bundle = parse_direct_linux_release_bundle(repo, release)
if bundle is None:
return None
if not direct_release_matches_request(
release_tag = bundle.release_tag,
llama_tag = bundle.upstream_tag,
requested_tag = requested_tag,
):
return None
attempts: list[AssetChoice] = []
if host.has_usable_nvidia:
# Prefer the cudart major Studio loads at runtime (torch's bundled
# libcudart), not the newest on disk. Otherwise a stray cuda13
# runtime outranks the torch cuda12 the binary links against.
torch_preference = detect_torch_cuda_runtime_preference(host)
selection = linux_cuda_choice_from_release(
host,
bundle,
preferred_runtime_line = torch_preference.runtime_line,
selection_preamble = torch_preference.selection_log,
)
if selection is not None:
attempts.extend(selection.attempts)
if host.has_rocm and not host.has_usable_nvidia:
# Per-GPU lemonade prebuilts ship the ROCm runtime libs alongside
# llama.cpp, so they install cleanly even on hosts (e.g. gfx1151
# Strix Halo) the upstream combined-ROCm tarball doesn't cover.
# "ubuntu" is lemonade's asset naming convention only -- the binary
# is a manylinux-style glibc build that runs on Arch, Fedora,
# openSUSE, etc. with a recent-enough glibc. Do NOT append the CPU
# asset for ROCm-only hosts: if lemonade fails validation we want
# validate_prebuilt_attempts to raise PrebuiltFallback so the caller
# triggers the HIP source build, not silently install a CPU binary.
lemonade_choice = resolve_lemonade_rocm_choice(
host, "ubuntu", "linux-rocm", llama_tag = requested_tag
)
if lemonade_choice is not None:
attempts.append(lemonade_choice)
else:
cpu_choice = published_asset_choice_for_kind(bundle, "linux-cpu")
if cpu_choice is not None:
attempts.append(cpu_choice)
if not attempts:
raise PrebuiltFallback("no compatible Linux prebuilt asset was found")
approved_checksums = synthetic_checksums_for_release(
repo,
bundle.release_tag,
bundle.upstream_tag,
)
resolved_upstream_tag = bundle.upstream_tag
if DEFAULT_PUBLISHED_SHA256_ASSET in bundle.assets and not is_release_tag_like(
bundle.upstream_tag
):
approved_checksums = load_approved_release_checksums(repo, bundle.release_tag)
# Require exact source provenance for branch/pull/commit releases.
# Mirrors validated_checksums_for_bundle so incomplete metadata fails
# closed instead of degrading to the legacy branch-as-tag source
# hydration path this PR eliminates.
if (
not approved_checksums.source_commit
or exact_source_archive_hash(approved_checksums) is None
or source_clone_url_from_checksums(approved_checksums) is None
):
raise PrebuiltFallback(
f"approved checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET} for "
f"{repo}@{bundle.release_tag} did not contain exact source provenance"
)
attempts = apply_approved_hashes(attempts, approved_checksums)
return InstallReleasePlan(
requested_tag = requested_tag,
llama_tag = resolved_upstream_tag,
release_tag = bundle.release_tag,
attempts = attempts,
approved_checksums = approved_checksums,
)
def direct_upstream_release_plan(
release: dict[str, Any], host: HostInfo, repo: str, requested_tag: str
) -> InstallReleasePlan | None:
release_tag = release.get("tag_name")
if not isinstance(release_tag, str) or not release_tag:
return None
if not direct_release_matches_request(
release_tag = release_tag,
llama_tag = release_tag,
requested_tag = requested_tag,
):
return None
assets = release_asset_map(release)
attempts: list[AssetChoice] = []
if host.is_windows and host.is_x86_64:
if host.has_usable_nvidia:
torch_preference = detect_torch_cuda_runtime_preference(host)
attempts.extend(
windows_cuda_attempts(
host,
release_tag,
assets,
torch_preference.runtime_line,
torch_preference.selection_log,
)
)
# Blackwell on a 13.1/13.2 driver: prefer the pinned cuda-13.1 GPU
# build over the CPU-only cuda-12.4 left by in-release gating.
pinned = _pinned_windows_cuda_fallback(host, attempts)
if pinned is not None:
attempts.insert(0, pinned)
elif host.has_rocm:
lemonade_choice = resolve_lemonade_rocm_choice(
host, "windows", "windows-hip", llama_tag = requested_tag
)
if lemonade_choice is not None:
attempts.append(lemonade_choice)
hip_asset = f"llama-{release_tag}-bin-win-hip-radeon-x64.zip"
hip_url = assets.get(hip_asset)
if hip_url:
attempts.append(
AssetChoice(
repo = repo,
tag = release_tag,
name = hip_asset,
url = hip_url,
source_label = "upstream",
install_kind = "windows-hip",
)
)
cpu_asset = f"llama-{release_tag}-bin-win-cpu-x64.zip"
cpu_url = assets.get(cpu_asset)
if cpu_url:
attempts.append(
AssetChoice(
repo = repo,
tag = release_tag,
name = cpu_asset,
url = cpu_url,
source_label = "upstream",
install_kind = "windows-cpu",
)
)
elif host.is_windows and host.is_arm64:
# Upstream ggml-org/llama.cpp ships llama-bNNNN-bin-win-cpu-arm64.zip
# (in the b9334 release manifest). Without this branch the selector
# returned 0 attempts and fell back to a source build on every
# Windows ARM64 host.
cpu_asset = f"llama-{release_tag}-bin-win-cpu-arm64.zip"
cpu_url = assets.get(cpu_asset)
if cpu_url:
attempts.append(
AssetChoice(
repo = repo,
tag = release_tag,
name = cpu_asset,
url = cpu_url,
source_label = "upstream",
install_kind = "windows-arm64",
)
)
elif host.is_macos and host.is_arm64:
asset_name = f"llama-{release_tag}-bin-macos-arm64.tar.gz"
asset_url = assets.get(asset_name)
if asset_url:
attempts.append(
AssetChoice(
repo = repo,
tag = release_tag,
name = asset_name,
url = asset_url,
source_label = "upstream",
install_kind = "macos-arm64",
)
)
elif host.is_macos and host.is_x86_64:
asset_name = f"llama-{release_tag}-bin-macos-x64.tar.gz"
asset_url = assets.get(asset_name)
if asset_url:
attempts.append(
AssetChoice(
repo = repo,
tag = release_tag,
name = asset_name,
url = asset_url,
source_label = "upstream",
install_kind = "macos-x64",
)
)
elif host.is_linux and host.is_x86_64 and not host.has_usable_nvidia:
asset_name = f"llama-{release_tag}-bin-ubuntu-x64.tar.gz"
asset_url = assets.get(asset_name)
if asset_url:
attempts.append(
AssetChoice(
repo = repo,
tag = release_tag,
name = asset_name,
url = asset_url,
source_label = "upstream",
install_kind = "linux-cpu",
)
)
elif host.is_linux and host.is_arm64 and not host.has_usable_nvidia:
# Upstream ggml-org/llama.cpp ships llama-bNNNN-bin-ubuntu-arm64.tar.gz
# (in the b9334 release manifest). Without this branch the selector
# returned 0 attempts and fell back to a source build on every Linux
# ARM64 host (DGX Spark, Ampere Altra, GitHub ubuntu-24.04-arm
# runners, etc.).
asset_name = f"llama-{release_tag}-bin-ubuntu-arm64.tar.gz"
asset_url = assets.get(asset_name)
if asset_url:
attempts.append(
AssetChoice(
repo = repo,
tag = release_tag,
name = asset_name,
url = asset_url,
source_label = "upstream",
install_kind = "linux-arm64",
)
)
if not attempts:
raise PrebuiltFallback("no compatible upstream prebuilt asset was found")
return InstallReleasePlan(
requested_tag = requested_tag,
llama_tag = release_tag,
release_tag = release_tag,
attempts = attempts,
approved_checksums = synthetic_checksums_for_release(
repo,
release_tag,
release_tag,
),
)
def pinned_macos_release_tag(host: HostInfo, repo: str) -> str | None:
"""Pin b9415 (last upstream macOS build that loads below macOS 26) for a
known pre-26 host on ggml-org upstream; None keeps latest selection. The
unslothai/llama.cpp fork ships its own prebuilts (arm64 minos 14, x64
minos 13.3) and needs no pin, so this is a no-op there and for macOS 26+,
unknown version, non-macOS."""
if repo != UPSTREAM_REPO:
return None
if not host.is_macos:
return None
version = host.macos_version
if version is None:
return None
if version >= _PINNED_MACOS_LATEST_FLOOR:
return None
return _PINNED_MACOS_FALLBACK_TAG
def resolve_simple_install_release_plans(
llama_tag: str,
host: HostInfo,
published_repo: str,
published_release_tag: str,
*,
max_release_fallbacks: int = DEFAULT_MAX_PREBUILT_RELEASE_FALLBACKS,
) -> tuple[str, list[InstallReleasePlan]]:
repo = published_repo or DEFAULT_PUBLISHED_REPO
requested_tag = normalized_requested_llama_tag(llama_tag)
# The unslothai/llama.cpp fork ships only linux-x64 bundles. An arm64 Linux
# host with a GPU (GH200/GB200/DGX Spark) routes here; it must not install
# an x64 binary, so fall back to a GPU-targeting source build rather than
# the wrong arch (or silently dropping to a CPU arm64 build).
if host.is_linux and not host.is_x86_64 and repo == DEFAULT_PUBLISHED_REPO:
raise PrebuiltFallback(
f"{repo} ships only linux-x64 prebuilts; "
f"{host.machine or 'non-x64'} Linux falls back to source build"
)
allow_older_release_fallback = requested_tag == "latest" and not published_release_tag
# macOS: pin the last upstream build that loads on a pre-26 host instead of
# fetching the latest (macOS 26 only) build and walking back release by
# release. No-op on macOS 26+, unknown version, non-macOS, the fork.
if allow_older_release_fallback:
pinned_macos = pinned_macos_release_tag(host, repo)
if pinned_macos is not None:
requested_tag = pinned_macos
allow_older_release_fallback = False
release_limit = max(1, max_release_fallbacks)
plans: list[InstallReleasePlan] = []
last_error: PrebuiltFallback | None = None
try:
releases = iter_release_payloads_by_time(repo, published_release_tag, requested_tag)
for release in releases:
try:
if host.is_linux and repo == "unslothai/llama.cpp":
plan = direct_linux_release_plan(release, host, repo, requested_tag)
else:
plan = direct_upstream_release_plan(release, host, repo, requested_tag)
if plan is None:
continue
except PrebuiltFallback as exc:
last_error = exc
if not allow_older_release_fallback:
raise
release_tag = release.get("tag_name") or "unknown"
log(
"published release skipped for install planning: "
f"{repo}@{release_tag} ({exc})"
)
continue
plans.append(plan)
if not allow_older_release_fallback or len(plans) >= release_limit:
break
except PrebuiltFallback:
raise
except Exception as exc:
raise PrebuiltFallback(f"failed to inspect published releases in {repo}: {exc}") from exc
if plans:
return requested_tag, plans
if last_error is not None:
raise last_error
raise PrebuiltFallback(f"no installable published llama.cpp releases were found in {repo}")
def normalized_requested_llama_tag(requested_tag: str | None) -> str:
if isinstance(requested_tag, str):
normalized = requested_tag.strip()
if normalized:
return normalized
return "latest"
def normalize_compute_cap(value: Any) -> str | None:
raw = str(value).strip()
if not raw:
return None
if "." in raw:
parts = raw.split(".", 1)
if len(parts) != 2:
return None
major, minor = parts
if not major.isdigit() or not minor.isdigit():
return None
return f"{int(major)}{int(minor)}"
if raw.isdigit():
return str(int(raw))
return None
def normalize_compute_caps(compute_caps: Iterable[str]) -> list[str]:
normalized: list[str] = []
seen: set[str] = set()
for raw in compute_caps:
normalized_value = normalize_compute_cap(raw)
if normalized_value is None:
continue
if normalized_value in seen:
continue
seen.add(normalized_value)
normalized.append(normalized_value)
normalized.sort(key = int)
return normalized
def parse_cuda_visible_devices(value: str | None) -> list[str] | None:
if value is None:
return None
raw = value.strip()
if not raw or raw == "-1":
return []
return [token.strip() for token in raw.split(",") if token.strip()]
def supports_explicit_visible_device_matching(visible_devices: list[str] | None) -> bool:
if not visible_devices:
return False
for token in visible_devices:
lowered = token.lower()
if token.isdigit() or lowered.startswith("gpu-"):
continue
return False
return True
def select_visible_gpu_rows(
gpu_rows: Iterable[tuple[str, str, str]], visible_devices: list[str] | None
) -> list[tuple[str, str, str]]:
rows = list(gpu_rows)
if visible_devices is None:
return rows
if not visible_devices:
return []
by_index = {index: (index, uuid, cap) for index, uuid, cap in rows}
by_uuid = {uuid.lower(): (index, uuid, cap) for index, uuid, cap in rows}
selected: list[tuple[str, str, str]] = []
seen_indices: set[str] = set()
for token in visible_devices:
row = by_index.get(token)
if row is None:
normalized_token = token.lower()
row = by_uuid.get(normalized_token)
if row is None and normalized_token.startswith("gpu-"):
row = by_uuid.get(normalized_token)
if row is None and not normalized_token.startswith("gpu-"):
row = by_uuid.get("gpu-" + normalized_token)
if row is None:
continue
index = row[0]
if index in seen_indices:
continue
seen_indices.add(index)
selected.append(row)
return selected
def dir_provides_exact_library(directory: str | Path, library: str) -> bool:
if not library:
return False
candidate = Path(directory) / library
return candidate.exists() and (candidate.is_file() or candidate.is_symlink())
def linux_runtime_dirs_for_required_libraries(required_libraries: Iterable[str]) -> list[str]:
required = [library for library in required_libraries if library]
candidates: list[str | Path] = []
env_dirs = os.environ.get("CUDA_RUNTIME_LIB_DIR", "")
if env_dirs:
candidates.extend(part for part in env_dirs.split(os.pathsep) if part)
ld_library_path = os.environ.get("LD_LIBRARY_PATH", "")
if ld_library_path:
candidates.extend(part for part in ld_library_path.split(os.pathsep) if part)
cuda_roots: list[Path] = []
for name in ("CUDA_HOME", "CUDA_PATH", "CUDA_ROOT"):
value = os.environ.get(name)
if value:
cuda_roots.append(Path(value))
cuda_roots.extend(Path(path) for path in glob_paths("/usr/local/cuda", "/usr/local/cuda-*"))
for root in cuda_roots:
candidates.extend(
[
root / "lib",
root / "lib64",
root / "targets" / "x86_64-linux" / "lib",
]
)
candidates.extend(
Path(path)
for path in glob_paths(
"/lib",
"/lib64",
"/usr/lib",
"/usr/lib64",
"/usr/local/lib",
"/usr/local/lib64",
"/lib/x86_64-linux-gnu",
"/usr/lib/x86_64-linux-gnu",
)
)
candidates.extend(
Path(path) for path in glob_paths("/usr/local/lib/ollama/cuda_v*", "/usr/lib/wsl/lib")
)
candidates.extend(Path(path) for path in python_runtime_dirs())
candidates.extend(Path(path) for path in ldconfig_runtime_dirs(required))
resolved = dedupe_existing_dirs(candidates)
if not required:
return resolved
matched: list[tuple[int, str]] = []
for directory in resolved:
base = Path(directory)
provided = sum(1 for library in required if dir_provides_exact_library(directory, library))
if provided:
matched.append((provided, directory))
matched.sort(key = lambda item: item[0], reverse = True)
return [directory for _, directory in matched]
def detected_linux_runtime_lines() -> tuple[list[str], dict[str, list[str]]]:
line_requirements = {
f"cuda{m}": [f"libcudart.so.{m}", f"libcublas.so.{m}"]
for m in range(_MAX_PROBE_CUDA_MAJOR, _MIN_CUDA_MAJOR - 1, -1)
}
detected: list[str] = []
runtime_dirs: dict[str, list[str]] = {}
for line, required in line_requirements.items():
dirs = linux_runtime_dirs_for_required_libraries(required)
library_matches: dict[str, list[str]] = {}
matching_dirs: list[str] = []
for library in required:
matched_dirs = [
directory for directory in dirs if any(Path(directory).glob(f"{library}*"))
]
if not matched_dirs:
library_matches = {}
matching_dirs = []
break
library_matches[library] = matched_dirs
for directory in matched_dirs:
if directory not in matching_dirs:
matching_dirs.append(directory)
if library_matches:
detected.append(line)
runtime_dirs[line] = matching_dirs
return detected, runtime_dirs
def release_asset_map(release: dict[str, Any]) -> dict[str, str]:
assets = release.get("assets")
if not isinstance(assets, list):
return {}
return {
asset["name"]: asset.get("browser_download_url", "")
for asset in assets
if isinstance(asset, dict)
and isinstance(asset.get("name"), str)
and isinstance(asset.get("browser_download_url"), str)
}
def parse_published_artifact(raw: Any) -> PublishedLlamaArtifact | None:
if not isinstance(raw, dict):
raise ValueError("artifact entry was not an object")
asset_name = raw.get("asset_name")
install_kind = raw.get("install_kind")
if not isinstance(asset_name, str) or not asset_name:
raise ValueError("artifact.asset_name was missing or not a string")
if not isinstance(install_kind, str) or not install_kind:
raise ValueError(f"artifact {asset_name} install_kind was missing or not a string")
supported_sms_raw = raw.get("supported_sms", [])
if not isinstance(supported_sms_raw, (list, tuple)):
raise ValueError(f"artifact {asset_name} supported_sms must be a list or tuple")
if any(not isinstance(value, (int, str)) for value in supported_sms_raw):
raise ValueError(f"artifact {asset_name} supported_sms entries must be ints or strings")
supported_sms = normalize_compute_caps(supported_sms_raw)
min_sm_raw = raw.get("min_sm")
max_sm_raw = raw.get("max_sm")
try:
min_sm = int(min_sm_raw) if min_sm_raw is not None else None
max_sm = int(max_sm_raw) if max_sm_raw is not None else None
except (TypeError, ValueError) as exc:
raise ValueError(f"artifact {asset_name} min_sm/max_sm were not integers") from exc
runtime_line = raw.get("runtime_line")
coverage_class = raw.get("coverage_class")
bundle_profile = raw.get("bundle_profile")
rank_raw = raw.get("rank", 1000)
if runtime_line is not None and not isinstance(runtime_line, str):
raise ValueError(f"artifact {asset_name} runtime_line was not a string")
if coverage_class is not None and not isinstance(coverage_class, str):
raise ValueError(f"artifact {asset_name} coverage_class was not a string")
if bundle_profile is not None and not isinstance(bundle_profile, str):
raise ValueError(f"artifact {asset_name} bundle_profile was not a string")
try:
rank = int(rank_raw)
except (TypeError, ValueError):
raise ValueError(f"artifact {asset_name} rank was not an integer")
return PublishedLlamaArtifact(
asset_name = asset_name,
install_kind = install_kind,
runtime_line = runtime_line if isinstance(runtime_line, str) and runtime_line else None,
coverage_class = coverage_class
if isinstance(coverage_class, str) and coverage_class
else None,
supported_sms = supported_sms,
min_sm = min_sm,
max_sm = max_sm,
bundle_profile = bundle_profile
if isinstance(bundle_profile, str) and bundle_profile
else None,
rank = rank,
)
def parse_published_release_bundle(
repo: str, release: dict[str, Any]
) -> PublishedReleaseBundle | None:
release_tag = release.get("tag_name")
if not isinstance(release_tag, str) or not release_tag:
return None
assets = release_asset_map(release)
manifest_url = assets.get(DEFAULT_PUBLISHED_MANIFEST_ASSET)
if not manifest_url:
return None
# Mixed repos are filtered by an explicit release-side manifest, not by
# release tag or asset filename conventions.
manifest_bytes = download_bytes(
manifest_url,
timeout = 30,
headers = auth_headers(manifest_url),
)
manifest_sha256 = sha256_bytes(manifest_bytes)
try:
manifest_payload = json.loads(manifest_bytes.decode("utf-8"))
except (UnicodeDecodeError, json.JSONDecodeError) as exc:
raise RuntimeError(
f"published manifest {DEFAULT_PUBLISHED_MANIFEST_ASSET} was not valid JSON"
) from exc
if not isinstance(manifest_payload, dict):
raise RuntimeError(
f"published manifest {DEFAULT_PUBLISHED_MANIFEST_ASSET} was not a JSON object"
)
validate_schema_version(
manifest_payload,
label = f"published manifest {DEFAULT_PUBLISHED_MANIFEST_ASSET} in {repo}@{release_tag}",
)
component = manifest_payload.get("component")
upstream_tag = manifest_payload.get("upstream_tag")
source_repo = manifest_payload.get("source_repo")
source_repo_url = manifest_payload.get("source_repo_url")
source_ref_kind = normalize_source_ref_kind(manifest_payload.get("source_ref_kind"))
requested_source_ref = manifest_payload.get("requested_source_ref")
resolved_source_ref = manifest_payload.get("resolved_source_ref")
source_commit = normalize_source_commit(manifest_payload.get("source_commit"))
source_commit_short = manifest_payload.get("source_commit_short")
if component != "llama.cpp":
return None
if not isinstance(upstream_tag, str) or not upstream_tag:
raise RuntimeError(
f"published manifest {DEFAULT_PUBLISHED_MANIFEST_ASSET} in {repo}@{release_tag} omitted upstream_tag"
)
artifacts_payload = manifest_payload.get("artifacts")
if not isinstance(artifacts_payload, list):
raise RuntimeError(
f"published manifest {DEFAULT_PUBLISHED_MANIFEST_ASSET} in {repo}@{release_tag} omitted artifacts"
)
artifacts: list[PublishedLlamaArtifact] = []
for index, raw_artifact in enumerate(artifacts_payload):
try:
artifact = parse_published_artifact(raw_artifact)
except ValueError as exc:
log(f"published artifact ignored for {repo}@{release_tag} artifact[{index}]: {exc}")
continue
if artifact is not None:
artifacts.append(artifact)
selection_log = [
f"published_release: repo={repo}",
f"published_release: tag={release_tag}",
f"published_release: manifest={DEFAULT_PUBLISHED_MANIFEST_ASSET}",
f"published_release: upstream_tag={upstream_tag}",
]
if isinstance(source_repo, str) and source_repo:
selection_log.append(f"published_release: source_repo={source_repo}")
if source_commit:
selection_log.append(f"published_release: source_commit={source_commit}")
return PublishedReleaseBundle(
repo = repo,
release_tag = release_tag,
upstream_tag = upstream_tag,
manifest_sha256 = manifest_sha256,
source_repo = source_repo if isinstance(source_repo, str) and source_repo else None,
source_repo_url = source_repo_url
if isinstance(source_repo_url, str) and source_repo_url
else None,
source_ref_kind = source_ref_kind,
requested_source_ref = requested_source_ref
if isinstance(requested_source_ref, str) and requested_source_ref
else None,
resolved_source_ref = resolved_source_ref
if isinstance(resolved_source_ref, str) and resolved_source_ref
else None,
source_commit = source_commit,
source_commit_short = source_commit_short
if isinstance(source_commit_short, str) and source_commit_short
else None,
assets = assets,
manifest_asset_name = DEFAULT_PUBLISHED_MANIFEST_ASSET,
artifacts = artifacts,
selection_log = selection_log,
)
def parse_approved_release_checksums(
repo: str, release_tag: str, payload: Any
) -> ApprovedReleaseChecksums:
if not isinstance(payload, dict):
raise RuntimeError(
f"published checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET} was not a JSON object"
)
validate_schema_version(
payload,
label = f"published checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET}",
)
if payload.get("component") != "llama.cpp":
raise RuntimeError(
f"published checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET} did not describe llama.cpp"
)
payload_release_tag = payload.get("release_tag")
if not isinstance(payload_release_tag, str) or not payload_release_tag:
raise RuntimeError(
f"published checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET} omitted release_tag"
)
if payload_release_tag != release_tag:
raise RuntimeError(
f"published checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET} release_tag={payload_release_tag} "
f"did not match pinned release tag {release_tag}"
)
upstream_tag = payload.get("upstream_tag")
if not isinstance(upstream_tag, str) or not upstream_tag:
raise RuntimeError(
f"published checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET} omitted upstream_tag"
)
artifacts_payload = payload.get("artifacts")
if not isinstance(artifacts_payload, dict):
raise RuntimeError(
f"published checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET} omitted artifacts"
)
artifacts: dict[str, ApprovedArtifactHash] = {}
for asset_name, raw_entry in artifacts_payload.items():
if not isinstance(asset_name, str) or not asset_name:
raise RuntimeError("published checksum asset used a non-string artifact key")
if not isinstance(raw_entry, dict):
raise RuntimeError(f"published checksum entry for {asset_name} was not an object")
digest = normalize_sha256_digest(raw_entry.get("sha256"))
if not digest:
raise RuntimeError(f"published checksum entry for {asset_name} omitted a valid sha256")
repo_value = raw_entry.get("repo")
kind_value = raw_entry.get("kind")
artifacts[asset_name] = ApprovedArtifactHash(
asset_name = asset_name,
sha256 = digest,
repo = repo_value if isinstance(repo_value, str) and repo_value else None,
kind = kind_value if isinstance(kind_value, str) and kind_value else None,
)
source_commit = normalize_source_commit(payload.get("source_commit"))
source_commit_short = payload.get("source_commit_short")
source_repo = payload.get("source_repo")
source_repo_url = payload.get("source_repo_url")
source_ref_kind = normalize_source_ref_kind(payload.get("source_ref_kind"))
requested_source_ref = payload.get("requested_source_ref")
resolved_source_ref = payload.get("resolved_source_ref")
return ApprovedReleaseChecksums(
repo = repo,
release_tag = release_tag,
upstream_tag = upstream_tag,
source_repo = source_repo if isinstance(source_repo, str) and source_repo else None,
source_repo_url = source_repo_url
if isinstance(source_repo_url, str) and source_repo_url
else None,
source_ref_kind = source_ref_kind,
requested_source_ref = requested_source_ref
if isinstance(requested_source_ref, str) and requested_source_ref
else None,
resolved_source_ref = resolved_source_ref
if isinstance(resolved_source_ref, str) and resolved_source_ref
else None,
source_commit = source_commit,
source_commit_short = source_commit_short
if isinstance(source_commit_short, str) and source_commit_short
else None,
artifacts = artifacts,
)
def load_approved_release_checksums(repo: str, release_tag: str) -> ApprovedReleaseChecksums:
try:
release = github_release(repo, release_tag)
except Exception as exc:
raise PrebuiltFallback(
f"approved prebuilt release {repo}@{release_tag} was not available"
) from exc
assets = release_asset_map(release)
checksum_url = assets.get(DEFAULT_PUBLISHED_SHA256_ASSET)
if not checksum_url:
raise PrebuiltFallback(
f"approved prebuilt release {repo}@{release_tag} did not expose {DEFAULT_PUBLISHED_SHA256_ASSET}"
)
try:
payload = fetch_json(checksum_url)
checksums = parse_approved_release_checksums(repo, release_tag, payload)
except PrebuiltFallback:
raise
except Exception as exc:
raise PrebuiltFallback(
f"approved checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET} in {repo}@{release_tag} was invalid"
) from exc
return checksums
def iter_published_release_bundles(
repo: str, published_release_tag: str = ""
) -> Iterable[PublishedReleaseBundle]:
releases = (
[github_release(repo, published_release_tag)]
if published_release_tag
else github_releases(repo, max_pages = DEFAULT_GITHUB_RELEASE_SCAN_MAX_PAGES)
)
for release in releases:
if not published_release_tag and (release.get("draft") or release.get("prerelease")):
continue
try:
bundle = parse_published_release_bundle(repo, release)
except Exception as exc:
release_tag = release.get("tag_name", "unknown")
log(f"published release metadata ignored for {repo}@{release_tag}: {exc}")
continue
if bundle is None:
continue
yield bundle
def linux_cuda_choice_from_release(
host: HostInfo,
release: PublishedReleaseBundle,
preferred_runtime_line: str | None = None,
selection_preamble: Iterable[str] = (),
) -> LinuxCudaSelection | None:
host_sms = normalize_compute_caps(host.compute_caps)
detected_runtime_lines, runtime_dirs = detected_linux_runtime_lines()
driver_runtime_lines = compatible_linux_runtime_lines(host)
runtime_lines = [
runtime_line
for runtime_line in detected_runtime_lines
if runtime_line in driver_runtime_lines
]
ordered_runtime_lines = list(runtime_lines)
selection_log = (
list(release.selection_log)
+ list(selection_preamble)
+ [
f"linux_cuda_selection: release={release.release_tag}",
f"linux_cuda_selection: detected_sms={','.join(host_sms) if host_sms else 'unknown'}",
"linux_cuda_selection: detected_runtime_lines="
+ (",".join(detected_runtime_lines) if detected_runtime_lines else "none"),
"linux_cuda_selection: driver_runtime_lines="
+ (",".join(driver_runtime_lines) if driver_runtime_lines else "none"),
"linux_cuda_selection: compatible_runtime_lines="
+ (",".join(runtime_lines) if runtime_lines else "none"),
]
)
for runtime_line in ("cuda13", "cuda12"):
selection_log.append(
"linux_cuda_selection: runtime_dirs "
f"{runtime_line}="
+ (
",".join(runtime_dirs.get(runtime_line, []))
if runtime_dirs.get(runtime_line)
else "none"
)
)
published_artifacts = [
artifact for artifact in release.artifacts if artifact.install_kind == "linux-cuda"
]
published_asset_names = sorted(artifact.asset_name for artifact in published_artifacts)
selection_log.append(
"linux_cuda_selection: published_assets="
+ (",".join(published_asset_names) if published_asset_names else "none")
)
if not host_sms:
selection_log.append(
"linux_cuda_selection: compute capability detection unavailable; prefer portable by runtime line"
)
if not runtime_lines:
selection_log.append(
"linux_cuda_selection: no Linux CUDA runtime line satisfied both runtime libraries and driver compatibility"
)
return None
if preferred_runtime_line:
if preferred_runtime_line in ordered_runtime_lines:
ordered_runtime_lines = [preferred_runtime_line] + [
runtime_line
for runtime_line in ordered_runtime_lines
if runtime_line != preferred_runtime_line
]
selection_log.append(
"linux_cuda_selection: torch_preferred_runtime_line="
f"{preferred_runtime_line} reordered_attempts={','.join(ordered_runtime_lines)}"
)
else:
selection_log.append(
"linux_cuda_selection: torch_preferred_runtime_line="
f"{preferred_runtime_line} unavailable_on_host"
)
attempts: list[AssetChoice] = []
seen_attempts: set[str] = set()
def add_attempt(artifact: PublishedLlamaArtifact, asset_url: str, reason: str) -> None:
asset_name = artifact.asset_name
if asset_name in seen_attempts:
return
seen_attempts.add(asset_name)
attempts.append(
AssetChoice(
repo = release.repo,
tag = release.release_tag,
name = asset_name,
url = asset_url,
source_label = "published",
is_ready_bundle = True,
install_kind = "linux-cuda",
bundle_profile = artifact.bundle_profile,
runtime_line = artifact.runtime_line,
coverage_class = artifact.coverage_class,
supported_sms = artifact.supported_sms,
min_sm = artifact.min_sm,
max_sm = artifact.max_sm,
selection_log = list(selection_log)
+ [
"linux_cuda_selection: selected "
f"{asset_name} runtime_line={artifact.runtime_line} coverage_class={artifact.coverage_class} reason={reason}"
],
)
)
for runtime_line in ordered_runtime_lines:
coverage_candidates: list[tuple[PublishedLlamaArtifact, str]] = []
portable_candidate: tuple[PublishedLlamaArtifact, str] | None = None
for artifact in published_artifacts:
if artifact.runtime_line != runtime_line:
continue
asset_name = artifact.asset_name
asset_url = release.assets.get(asset_name)
if not asset_url:
selection_log.append(f"linux_cuda_selection: reject {asset_name} missing asset")
continue
if not host_sms and artifact.coverage_class != "portable":
selection_log.append(
"linux_cuda_selection: reject "
f"{asset_name} runtime_line={runtime_line} coverage_class={artifact.coverage_class} "
"reason=unknown_compute_caps_prefer_portable"
)
continue
if not artifact.supported_sms:
selection_log.append(
"linux_cuda_selection: reject "
f"{asset_name} runtime_line={runtime_line} coverage_class={artifact.coverage_class} "
"reason=artifact_missing_supported_sms"
)
continue
if artifact.min_sm is None or artifact.max_sm is None:
selection_log.append(
"linux_cuda_selection: reject "
f"{asset_name} runtime_line={runtime_line} coverage_class={artifact.coverage_class} "
"reason=artifact_missing_sm_bounds"
)
continue
supported_sms = {str(value) for value in artifact.supported_sms}
missing_sms = [sm for sm in host_sms if sm not in supported_sms]
out_of_range_sms = [
sm for sm in host_sms if not (artifact.min_sm <= int(sm) <= artifact.max_sm)
]
reasons: list[str] = []
if missing_sms:
reasons.append(f"missing_sms={','.join(missing_sms)}")
if out_of_range_sms:
reasons.append(f"out_of_range_sms={','.join(out_of_range_sms)}")
if reasons:
selection_log.append(
"linux_cuda_selection: reject "
f"{asset_name} runtime_line={runtime_line} coverage_class={artifact.coverage_class} "
f"coverage={artifact.min_sm}-{artifact.max_sm} supported={','.join(artifact.supported_sms)} "
f"reasons={' '.join(reasons)}"
)
continue
selection_log.append(
"linux_cuda_selection: accept "
f"{asset_name} runtime_line={runtime_line} coverage_class={artifact.coverage_class} "
f"coverage={artifact.min_sm}-{artifact.max_sm} supported={','.join(artifact.supported_sms)}"
)
if artifact.coverage_class == "portable":
portable_candidate = (artifact, asset_url)
else:
coverage_candidates.append((artifact, asset_url))
if coverage_candidates:
artifact, url = sorted(
coverage_candidates,
key = lambda item: (
(item[0].max_sm or 0) - (item[0].min_sm or 0),
item[0].rank,
item[0].max_sm or 0,
),
)[0]
add_attempt(artifact, url, "best coverage for runtime line")
if portable_candidate:
artifact, url = portable_candidate
add_attempt(artifact, url, "portable fallback for runtime line")
if not attempts:
return None
selection_log.append(
"linux_cuda_selection: attempt_order=" + ",".join(choice.name for choice in attempts)
)
for attempt in attempts:
attempt.selection_log = list(selection_log) + [
"linux_cuda_selection: attempt "
f"{attempt.name} runtime_line={attempt.runtime_line} coverage_class={attempt.coverage_class}"
]
return LinuxCudaSelection(attempts = attempts, selection_log = selection_log)
def latest_published_linux_cuda_tag(host: HostInfo, published_repo: str) -> str | None:
for release in iter_published_release_bundles(published_repo):
if linux_cuda_choice_from_release(host, release):
return release.upstream_tag
return None
def iter_upstream_releases() -> Iterable[dict[str, Any]]:
for release in github_releases(UPSTREAM_REPO, max_pages = DEFAULT_GITHUB_RELEASE_SCAN_MAX_PAGES):
if release.get("draft") or release.get("prerelease"):
continue
yield release
def pinned_published_release_bundle(
repo: str, published_release_tag: str
) -> PublishedReleaseBundle:
bundle = next(iter_published_release_bundles(repo, published_release_tag), None)
if bundle is None:
raise PrebuiltFallback(
f"published release {repo}@{published_release_tag} did not expose a usable llama.cpp manifest"
)
return bundle
def validated_checksums_for_bundle(
repo: str, bundle: PublishedReleaseBundle
) -> ApprovedReleaseChecksums:
checksums = load_approved_release_checksums(repo, bundle.release_tag)
manifest_hash = checksums.artifacts.get(bundle.manifest_asset_name)
if manifest_hash is not None and bundle.manifest_sha256 is not None:
if manifest_hash.sha256 != bundle.manifest_sha256:
raise PrebuiltFallback(
"published manifest checksum did not match the approved checksum asset"
)
# Accept bundles carrying only an exact-commit source archive
# (llama.cpp-source-commit-<sha>.tar.gz) without requiring the legacy
# llama.cpp-source-<upstream_tag>.tar.gz entry.
if exact_source_archive_hash(checksums) is None:
require_approved_source_hash(checksums, bundle.upstream_tag)
return checksums
def published_release_matches_request(bundle: PublishedReleaseBundle, requested_ref: str) -> bool:
if requested_ref == "latest":
return True
for candidate in (
bundle.upstream_tag,
bundle.requested_source_ref,
bundle.resolved_source_ref,
bundle.source_commit,
):
if refs_match(candidate, requested_ref):
return True
return False
def resolve_published_release(
requested_tag: str | None,
published_repo: str,
published_release_tag: str = "",
) -> ResolvedPublishedRelease:
repo = published_repo or DEFAULT_PUBLISHED_REPO
normalized_requested = normalized_requested_llama_tag(requested_tag)
if published_release_tag:
bundle = pinned_published_release_bundle(repo, published_release_tag)
if not published_release_matches_request(bundle, normalized_requested):
raise PrebuiltFallback(
"published release "
f"{repo}@{published_release_tag} targeted upstream tag {bundle.upstream_tag}, "
f"but requested {normalized_requested}"
)
return ResolvedPublishedRelease(
bundle = bundle,
checksums = validated_checksums_for_bundle(repo, bundle),
)
skipped_invalid = 0
for bundle in iter_published_release_bundles(repo):
if not published_release_matches_request(bundle, normalized_requested):
continue
try:
checksums = validated_checksums_for_bundle(repo, bundle)
except PrebuiltFallback as exc:
skipped_invalid += 1
log(
"published release ignored for install resolution: "
f"{repo}@{bundle.release_tag} ({exc})"
)
continue
return ResolvedPublishedRelease(bundle = bundle, checksums = checksums)
if normalized_requested == "latest":
if skipped_invalid:
raise PrebuiltFallback(
f"no usable published llama.cpp releases were available in {repo}"
)
raise PrebuiltFallback(f"no published llama.cpp releases were available in {repo}")
raise PrebuiltFallback(
f"no published prebuilt release in {repo} matched upstream tag {normalized_requested}"
)
def iter_resolved_published_releases(
requested_tag: str | None,
published_repo: str,
published_release_tag: str = "",
) -> Iterable[ResolvedPublishedRelease]:
repo = published_repo or DEFAULT_PUBLISHED_REPO
normalized_requested = normalized_requested_llama_tag(requested_tag)
if published_release_tag:
bundle = pinned_published_release_bundle(repo, published_release_tag)
if not published_release_matches_request(bundle, normalized_requested):
raise PrebuiltFallback(
"published release "
f"{repo}@{published_release_tag} targeted upstream tag {bundle.upstream_tag}, "
f"but requested {normalized_requested}"
)
yield ResolvedPublishedRelease(
bundle = bundle,
checksums = validated_checksums_for_bundle(repo, bundle),
)
return
matched_any = False
skipped_invalid = 0
yielded_valid = False
for bundle in iter_published_release_bundles(repo):
if not published_release_matches_request(bundle, normalized_requested):
continue
matched_any = True
try:
checksums = validated_checksums_for_bundle(repo, bundle)
except PrebuiltFallback as exc:
skipped_invalid += 1
log(
"published release ignored for install resolution: "
f"{repo}@{bundle.release_tag} ({exc})"
)
continue
yielded_valid = True
yield ResolvedPublishedRelease(bundle = bundle, checksums = checksums)
if yielded_valid:
return
if matched_any:
if skipped_invalid:
raise PrebuiltFallback(
f"no usable published llama.cpp releases were available in {repo}"
)
return
if normalized_requested == "latest":
raise PrebuiltFallback(f"no published llama.cpp releases were available in {repo}")
raise PrebuiltFallback(
f"no published prebuilt release in {repo} matched upstream tag {normalized_requested}"
)
def resolve_requested_llama_tag(
requested_tag: str | None,
published_repo: str = "",
published_release_tag: str = "",
) -> str:
"""Resolve a llama.cpp tag for source-build fallback.
Resolution order:
1. Concrete tag (e.g. "b8508") -- returned as-is.
2. "latest" with published_repo -- the latest usable Unsloth published
bundle's upstream_tag (matches the published prebuilt metadata).
3. "latest" without published_repo, or if (2) fails -- query upstream
ggml-org/llama.cpp. May return a newer, untested tag.
The Unsloth repo is preferred because its releases are pinned to upstream
tags validated with Unsloth Studio; the upstream bleeding-edge tag risks
API/ABI incompatibilities.
"""
normalized_requested = normalized_requested_llama_tag(requested_tag)
if normalized_requested != "latest":
return normalized_requested
# Prefer the Unsloth release repo tag (tested/approved) over bleeding-edge
# upstream. E.g. unslothai/llama.cpp may publish b8508 while ggml-org
# latest is b8514. The source-build fallback should compile the same
# version the prebuilt path would have installed.
if published_repo:
try:
return resolve_published_release(
"latest",
published_repo,
published_release_tag,
).bundle.upstream_tag
except Exception:
pass
# Fall back to the upstream ggml-org latest release tag
return latest_upstream_release_tag()
def resolve_requested_install_tag(
requested_tag: str | None,
published_release_tag: str = "",
published_repo: str = DEFAULT_PUBLISHED_REPO,
) -> str:
return resolve_published_release(
requested_tag,
published_repo,
published_release_tag,
).bundle.upstream_tag
def exact_source_archive_hash(checksums: ApprovedReleaseChecksums) -> ApprovedArtifactHash | None:
if not checksums.source_commit:
return None
return checksums.artifacts.get(exact_source_archive_logical_name(checksums.source_commit))
def source_clone_url_from_checksums(checksums: ApprovedReleaseChecksums) -> str | None:
return source_repo_clone_url(checksums.source_repo, checksums.source_repo_url)
def source_build_plan_for_release(release: ResolvedPublishedRelease) -> SourceBuildPlan:
checksums = release.checksums
exact_source = exact_source_archive_hash(checksums)
source_repo = checksums.source_repo or release.bundle.source_repo
source_repo_url = checksums.source_repo_url or release.bundle.source_repo_url
requested_source_ref = checksums.requested_source_ref or release.bundle.requested_source_ref
resolved_source_ref = checksums.resolved_source_ref or release.bundle.resolved_source_ref
source_commit = checksums.source_commit or release.bundle.source_commit
source_ref_kind = checksums.source_ref_kind or release.bundle.source_ref_kind
source_url = source_repo_clone_url(source_repo, source_repo_url)
if exact_source is not None and source_url and source_commit:
return SourceBuildPlan(
source_url = source_url,
source_ref = source_commit,
source_ref_kind = "commit",
compatibility_upstream_tag = release.bundle.upstream_tag,
source_repo = source_repo,
source_repo_url = source_repo_url,
requested_source_ref = requested_source_ref,
resolved_source_ref = resolved_source_ref,
source_commit = source_commit,
)
source_ref = checkout_friendly_ref(source_ref_kind, resolved_source_ref or requested_source_ref)
if source_url and source_ref and source_ref_kind in {"tag", "branch", "pull", "commit"}:
return SourceBuildPlan(
source_url = source_url,
source_ref = source_ref,
source_ref_kind = source_ref_kind,
compatibility_upstream_tag = release.bundle.upstream_tag,
source_repo = source_repo,
source_repo_url = source_repo_url,
requested_source_ref = requested_source_ref,
resolved_source_ref = resolved_source_ref,
source_commit = source_commit,
)
return SourceBuildPlan(
source_url = source_url_from_repo_slug(UPSTREAM_REPO)
or "https://github.com/ggml-org/llama.cpp",
source_ref = release.bundle.upstream_tag,
source_ref_kind = "tag",
compatibility_upstream_tag = release.bundle.upstream_tag,
source_repo = source_repo,
source_repo_url = source_repo_url,
requested_source_ref = requested_source_ref,
resolved_source_ref = resolved_source_ref,
source_commit = source_commit,
)
def resolve_source_build_plan(
requested_tag: str | None,
published_repo: str,
published_release_tag: str = "",
) -> SourceBuildPlan:
normalized_requested = normalized_requested_llama_tag(requested_tag)
if normalized_requested != "latest":
try:
release = resolve_published_release(
normalized_requested,
published_repo,
published_release_tag,
)
return source_build_plan_for_release(release)
except Exception:
pass
inferred_kind = infer_source_ref_kind(normalized_requested)
return SourceBuildPlan(
source_url = "https://github.com/ggml-org/llama.cpp",
source_ref = checkout_friendly_ref(inferred_kind, normalized_requested)
or normalized_requested,
source_ref_kind = inferred_kind,
compatibility_upstream_tag = normalized_requested,
)
if published_repo:
try:
release = resolve_published_release(
"latest",
published_repo,
published_release_tag,
)
return source_build_plan_for_release(release)
except Exception:
pass
latest_tag = latest_upstream_release_tag()
return SourceBuildPlan(
source_url = "https://github.com/ggml-org/llama.cpp",
source_ref = latest_tag,
source_ref_kind = "tag",
compatibility_upstream_tag = latest_tag,
)
def run_capture(
command: list[str],
*,
timeout: int = 30,
check: bool = False,
env: dict[str, str] | None = None,
) -> subprocess.CompletedProcess[str]:
# amd-smi on Windows auto-elevates and pops a UAC/DiskPart prompt mid-install;
# RunAsInvoker forces it un-elevated. Callers already fall back to WMI/name
# detection. Mirrors install.ps1's Invoke-AmdSmiNoElevate; Windows-only.
if (
command
and platform.system() == "Windows"
and os.path.basename(command[0]).lower().startswith("amd-smi")
):
env = {**(os.environ if env is None else env), "__COMPAT_LAYER": "RunAsInvoker"}
result = subprocess.run(
command,
capture_output = True,
text = True,
timeout = timeout,
env = env,
**windows_hidden_subprocess_kwargs(),
)
if check and result.returncode != 0:
raise subprocess.CalledProcessError(
result.returncode, command, result.stdout, result.stderr
)
return result
def _pick_rocm_gfx_target(out: str) -> str | None:
"""Choose the gfx target rocminfo / hipinfo report for the active GPU.
A bare first-match picked the wrong device on mixed APU + dGPU hosts
(e.g. Strix Halo gfx1151 + discrete RX 7900 gfx1100). Respect
HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES / CUDA_VISIBLE_DEVICES so the
asset matches what HIP runs on; falls back to the first GPU with no env
var set.
rocminfo / hipinfo print the same gfx token multiple times per GPU (Name,
ISA, marketing-name). We first split the output on per-GPU section headers
(rocminfo: "Agent N" blocks, hipinfo: "device#N" entries) and take exactly
one gfx token per section -- this gives the correct per-GPU list even on
same-arch multi-GPU hosts (e.g. two RX 7900 XTX) where global dict.fromkeys
dedup would collapse both to one entry and make HIP_VISIBLE_DEVICES=1 point
out of range.
Falls back to insertion-order dedup when the output has no recognisable
section markers (flat gfx-string inputs, unit-test stubs, etc.).
Empty / "-1" env values mean no AMD GPU is visible to HIP: return None.
"""
# Build a per-GPU token list by splitting on section boundaries. rocminfo
# sections start with "Agent N" lines (optionally between rows of
# asterisks); hipinfo sections with "device#N".
_sections = re.split(
r"(?mi)^\s*\*+\s*$\s*agent\s+\d+\s*$|\bdevice\s*#\s*\d+\b",
out,
)
if len(_sections) > 1:
# One gfx token per GPU section preserves physical order.
_tokens: list[str] = []
for _sec in _sections[1:]:
_m = re.search(r"gfx[1-9][0-9a-z]{2,3}", _sec.lower())
if _m:
_tokens.append(_m.group(0))
else:
# Fallback: insertion-order dedup (flat strings / unknown formats).
_raw = re.findall(r"gfx[1-9][0-9a-z]{2,3}", out.lower())
_tokens = list(dict.fromkeys(_raw))
if not _tokens:
return None
_vis_raw = None
# AMD's HIP runtime honours all three env vars with identical semantics.
for _env in ("HIP_VISIBLE_DEVICES", "ROCR_VISIBLE_DEVICES", "CUDA_VISIBLE_DEVICES"):
_val = os.environ.get(_env)
if _val is not None:
_vis_raw = _val
break
if _vis_raw is not None:
_vis = _vis_raw.strip()
# Empty or "-1" means "no AMD GPU visible" (matches the rest of Studio)
if _vis == "" or _vis == "-1":
return None
_first = _vis.split(",")[0].strip()
try:
_idx = int(_first)
if 0 <= _idx < len(_tokens):
return _tokens[_idx]
except ValueError:
pass
return _tokens[0]
def detect_host() -> HostInfo:
system = platform.system()
machine = platform.machine().lower()
is_windows = system == "Windows"
is_linux = system == "Linux"
is_macos = system == "Darwin"
is_x86_64 = machine in {"x86_64", "amd64"}
is_arm64 = machine in {"arm64", "aarch64"}
macos_version = parse_macos_version(platform.mac_ver()[0]) if is_macos else None
nvidia_smi = shutil.which("nvidia-smi")
driver_cuda_version = None
compute_caps: list[str] = []
visible_cuda_devices = os.environ.get("CUDA_VISIBLE_DEVICES")
visible_device_tokens = parse_cuda_visible_devices(visible_cuda_devices)
has_physical_nvidia = False
has_usable_nvidia = False
if nvidia_smi:
# Require `nvidia-smi -L` to list a GPU before treating the host as
# NVIDIA. The "NVIDIA-SMI ..." banner prints even when the command
# can't reach the driver (e.g. stale container leftovers), which
# would misclassify an AMD ROCm host as NVIDIA and skip the ROCm path.
try:
listing = run_capture([nvidia_smi, "-L"], timeout = 20)
gpu_lines = [line for line in listing.stdout.splitlines() if line.startswith("GPU ")]
if gpu_lines:
has_physical_nvidia = True
has_usable_nvidia = visible_device_tokens != []
except Exception:
pass
try:
result = run_capture([nvidia_smi], timeout = 20)
merged = "\n".join(part for part in (result.stdout, result.stderr) if part)
# Newer NVIDIA drivers (e.g. 610.x on Windows) print "CUDA UMD
# Version: X.Y" instead of the legacy "CUDA Version: X.Y"; accept
# both spellings.
cuda_match = re.search(
r"CUDA(?: UMD)? Version:\s*(\d+)\.(\d+)",
merged,
)
if cuda_match is not None:
driver_cuda_version = (
int(cuda_match.group(1)),
int(cuda_match.group(2)),
)
except Exception:
pass
try:
caps = run_capture(
[
nvidia_smi,
"--query-gpu=index,uuid,compute_cap",
"--format=csv,noheader",
],
timeout = 20,
)
visible_gpu_rows: list[tuple[str, str, str]] = []
for raw in caps.stdout.splitlines():
parts = [part.strip() for part in raw.split(",")]
if len(parts) != 3:
continue
index, uuid, cap = parts
visible_gpu_row = select_visible_gpu_rows(
[(index, uuid, cap)],
visible_device_tokens,
)
if not visible_gpu_row:
continue
visible_gpu_rows.extend(visible_gpu_row)
normalized_cap = normalize_compute_cap(cap)
if normalized_cap is None:
continue
if normalized_cap not in compute_caps:
compute_caps.append(normalized_cap)
if visible_gpu_rows:
has_usable_nvidia = True
# Older nvidia-smi (pre -L support) hits the except in the
# first try block but still succeeds here, leaving
# has_physical_nvidia unset. Mirror the -L path so downstream
# diagnostics on line ~4390 still run.
if not has_physical_nvidia:
has_physical_nvidia = True
elif visible_device_tokens == []:
has_usable_nvidia = False
elif supports_explicit_visible_device_matching(visible_device_tokens):
has_usable_nvidia = False
elif has_physical_nvidia:
has_usable_nvidia = True
except Exception:
pass
# Detect AMD ROCm (HIP) -- require actual GPU, not just tools installed
def _amd_smi_has_gpu(stdout: str) -> bool:
"""Check for 'GPU: <number>' data rows, not just a table header."""
return bool(re.search(r"(?im)^gpu\s*[:\[]\s*\d", stdout))
has_rocm = False
rocm_gfx_target: str | None = None
if is_linux:
# WSL2 ROCDXG: the system rocminfo enumerates the GPU over /dev/dxg
# only when HSA_ENABLE_DXG_DETECTION=1 (a no-op on bare metal), and
# rocminfo can live only under /opt/rocm/bin (the profile.d PATH
# drop-in reaches login shells only). Probe accordingly or a ROCDXG
# WSL host is misdetected as CPU-only.
_dxg_probe_env = {**os.environ}
_dxg_probe_env.setdefault("HSA_ENABLE_DXG_DETECTION", "1")
for _cmd, _check in (
# rocminfo: a real gfx GPU id (3-4 chars, nonzero first digit).
# gfx000 is the CPU agent; ROCm 6.1+ also emits generic ISA lines
# ("gfx11-generic", "gfx9-4-generic") with only 1-2 digits before
# the dash, which must not be treated as a real GPU.
(
["rocminfo"],
lambda out: bool(re.search(r"gfx[1-9][0-9a-z]{2,3}", out.lower())),
),
(["amd-smi", "list"], _amd_smi_has_gpu),
):
_exe = shutil.which(_cmd[0])
if not _exe and _cmd[0] == "rocminfo":
_opt_rocminfo = "/opt/rocm/bin/rocminfo"
if os.access(_opt_rocminfo, os.X_OK):
_exe = _opt_rocminfo
if not _exe:
continue
try:
_result = run_capture(
[_exe, *_cmd[1:]],
timeout = 10,
env = _dxg_probe_env if _cmd[0] == "rocminfo" else None,
)
except Exception:
continue
if _result.returncode == 0 and _result.stdout.strip():
if _check(_result.stdout):
has_rocm = True
rocm_gfx_target = _pick_rocm_gfx_target(_result.stdout)
break
elif is_windows:
# Windows: prefer active probes that validate GPU presence. hipinfo /
# amd-smi are often NOT on PATH -- the HIP SDK installer sets HIP_PATH
# / ROCM_PATH but doesn't always add the bin dir to PATH. Mirror
# setup.ps1's fallback: check the env-var bin dirs before giving up so
# `has_rocm` isn't silently False when PATH isn't updated yet.
def _resolve_exe(name: str) -> str | None:
"""Full path to `name`, checking PATH then HIP_PATH/ROCM_PATH bin."""
found = shutil.which(name)
if found:
return found
for _env in ("HIP_PATH", "ROCM_PATH"):
_root = os.environ.get(_env)
if _root:
_candidate = os.path.join(_root, "bin", f"{name}.exe")
if os.path.isfile(_candidate):
return _candidate
# AMD torch wheels ship hipInfo.exe into the venv Scripts dir
# (next to python.exe) -- resolvable on driver-only hosts where no
# SDK dir exists, so a standalone rerun can still detect the GPU.
_venv_candidate = os.path.join(os.path.dirname(sys.executable), f"{name}.exe")
if os.path.isfile(_venv_candidate):
return _venv_candidate
return None
_win_probes = [(["hipinfo"], lambda out: "gcnarchname" in out.lower())]
if _amd_smi_allowed():
# Skipped on Windows w/o a HIP SDK (avoids the UAC/DiskPart prompt);
# gfx arch still arrives via --rocm-gfx, so has_rocm is set by override.
_win_probes.append((["amd-smi", "list"], _amd_smi_has_gpu))
for _cmd, _check in _win_probes:
_exe = _resolve_exe(_cmd[0])
if not _exe:
continue
try:
_result = run_capture([_exe, *_cmd[1:]], timeout = 10)
except Exception:
continue
if _result.returncode == 0 and _result.stdout.strip():
if _check(_result.stdout):
has_rocm = True
# hipinfo reports "gcnArchName: gfx1100" -- extract if present
rocm_gfx_target = _pick_rocm_gfx_target(_result.stdout)
break
# Note: amdhip64.dll presence alone is NOT GPU evidence -- the HIP SDK
# can be installed without an AMD GPU.
return HostInfo(
system = system,
machine = machine,
is_windows = is_windows,
is_linux = is_linux,
is_macos = is_macos,
is_x86_64 = is_x86_64,
is_arm64 = is_arm64,
nvidia_smi = nvidia_smi,
driver_cuda_version = driver_cuda_version,
compute_caps = compute_caps,
visible_cuda_devices = visible_cuda_devices,
has_physical_nvidia = has_physical_nvidia,
has_usable_nvidia = has_usable_nvidia,
has_rocm = has_rocm,
rocm_gfx_target = rocm_gfx_target,
macos_version = macos_version,
)
def _normalize_forwarded_gfx(value: str | None) -> str | None:
"""Extract a single gfx token from a forwarded --rocm-gfx / env value.
setup.sh/setup.ps1 already picked the active GPU, so take the token as-is
without re-applying visible-device selection. Ignore malformed input."""
if not value:
return None
m = re.search(r"gfx[1-9][0-9a-z]{2,3}", value.lower())
return m.group(0) if m else None
def _apply_host_overrides(
host: HostInfo,
*,
override_has_rocm: bool = False,
override_rocm_gfx: str | None = None,
force_cpu: bool = False,
) -> HostInfo:
"""Fold setup.sh/setup.ps1's forwarded detection into the host profile.
A forwarded gfx (--rocm-gfx or UNSLOTH_ROCM_GFX_ARCH) is authoritative and
implies ROCm: the installer's own hipinfo/amd-smi probe can miss the arch
on amd-smi-only hosts or when setup inferred it from the GPU name, leaving
rocm_gfx_target None and no lemonade prebuilt selected. force_cpu is the
opposite explicit signal (arm64 Linux GPU host whose source build failed):
drop GPU attributes so the CPU prebuilt for this OS/arch is selected."""
if force_cpu:
return dataclasses_replace(
host,
has_usable_nvidia = False,
has_physical_nvidia = False,
has_rocm = False,
rocm_gfx_target = None,
)
gfx = _normalize_forwarded_gfx(override_rocm_gfx)
if gfx:
return dataclasses_replace(host, has_rocm = True, rocm_gfx_target = gfx)
if override_has_rocm and not host.has_rocm:
return dataclasses_replace(host, has_rocm = True)
return host
def pick_windows_cuda_runtime(host: HostInfo) -> str | None:
if not host.driver_cuda_version:
return None
major, minor = host.driver_cuda_version
if major > 13 or (major == 13): # and minor >= 1):
return "13.1"
if major > 12 or (major == 12 and minor >= 4):
return "12.4"
return None
def compatible_linux_runtime_lines(host: HostInfo) -> list[str]:
if not host.driver_cuda_version:
return []
major, _minor = host.driver_cuda_version
if major < _MIN_CUDA_MAJOR:
return []
return _cuda_runtime_lines_for_major(major)
def windows_runtime_line_info() -> dict[str, tuple[str, ...]]:
# Generated per CUDA major (newest first) so a new toolkit is detected
# without code changes while the cudart64_<major>.dll naming holds.
return {
f"cuda{m}": (
f"cudart64_{m}*.dll",
f"cublas64_{m}*.dll",
f"cublasLt64_{m}*.dll",
)
for m in range(_MAX_PROBE_CUDA_MAJOR, _MIN_CUDA_MAJOR - 1, -1)
}
def detected_windows_runtime_lines() -> tuple[list[str], dict[str, list[str]]]:
dirs = windows_runtime_dirs()
detected: list[str] = []
runtime_dirs: dict[str, list[str]] = {}
for runtime_line, required_patterns in windows_runtime_line_info().items():
matching_dirs = windows_runtime_dirs_for_patterns(required_patterns, dirs)
if matching_dirs:
detected.append(runtime_line)
runtime_dirs[runtime_line] = matching_dirs
return detected, runtime_dirs
def compatible_windows_runtime_lines(host: HostInfo) -> list[str]:
if not host.driver_cuda_version:
return []
major, minor = host.driver_cuda_version
# cuda12 prebuilts need a 12.4+ driver; cuda13+ any minor.
if major < _MIN_CUDA_MAJOR or (major == _MIN_CUDA_MAJOR and minor < 4):
return []
return _cuda_runtime_lines_for_major(major)
def runtime_line_from_cuda_version(cuda_version: str | None) -> str | None:
if not cuda_version:
return None
raw = str(cuda_version).strip()
if not raw:
return None
major, _, _ = raw.partition(".")
if major == "12":
return "cuda12"
if major == "13":
return "cuda13"
return None
def detect_torch_cuda_runtime_preference(host: HostInfo) -> CudaRuntimePreference:
selection_log: list[str] = []
if host.is_macos:
selection_log.append("torch_cuda_preference: skipped on macOS")
return CudaRuntimePreference(runtime_line = None, selection_log = selection_log)
if not (host.has_usable_nvidia and (host.is_linux or host.is_windows)):
selection_log.append(
"torch_cuda_preference: skipped because CUDA host prerequisites were not met"
)
return CudaRuntimePreference(runtime_line = None, selection_log = selection_log)
try:
import torch
except Exception as exc:
selection_log.append(f"torch_cuda_preference: import failed: {exc}")
return CudaRuntimePreference(runtime_line = None, selection_log = selection_log)
cuda_version = getattr(getattr(torch, "version", None), "cuda", None)
if not isinstance(cuda_version, str) or not cuda_version.strip():
selection_log.append(
"torch_cuda_preference: torch.version.cuda missing; skipping Torch shortcut"
)
return CudaRuntimePreference(runtime_line = None, selection_log = selection_log)
try:
cuda_available = bool(torch.cuda.is_available())
except Exception as exc:
selection_log.append(f"torch_cuda_preference: torch.cuda.is_available() failed: {exc}")
return CudaRuntimePreference(runtime_line = None, selection_log = selection_log)
if not cuda_available:
selection_log.append(
"torch_cuda_preference: torch.cuda.is_available() returned False; falling back to normal selection"
)
return CudaRuntimePreference(runtime_line = None, selection_log = selection_log)
runtime_line = runtime_line_from_cuda_version(cuda_version)
if runtime_line is None:
selection_log.append(
f"torch_cuda_preference: unsupported torch.version.cuda={cuda_version}; falling back to normal selection"
)
return CudaRuntimePreference(runtime_line = None, selection_log = selection_log)
selection_log.append(
"torch_cuda_preference: selected runtime_line="
f"{runtime_line} from torch.version.cuda={cuda_version}"
)
return CudaRuntimePreference(runtime_line = runtime_line, selection_log = selection_log)
def windows_cuda_attempts(
host: HostInfo,
llama_tag: str,
upstream_assets: dict[str, str],
preferred_runtime_line: str | None,
selection_preamble: Iterable[str] = (),
) -> list[AssetChoice]:
selection_log = list(selection_preamble)
driver_runtime = pick_windows_cuda_runtime(host)
detected_runtime_lines, runtime_dirs = detected_windows_runtime_lines()
compatible_runtime_lines = compatible_windows_runtime_lines(host)
normal_runtime_lines: list[str]
if detected_runtime_lines:
normal_runtime_lines = [
line for line in compatible_runtime_lines if line in detected_runtime_lines
]
else:
normal_runtime_lines = compatible_runtime_lines
selection_log.append(
"windows_cuda_selection: driver_runtime="
+ (driver_runtime if driver_runtime else "unknown")
)
selection_log.append(
"windows_cuda_selection: detected_runtime_lines="
+ (",".join(detected_runtime_lines) if detected_runtime_lines else "none")
)
for runtime_line in ("cuda13", "cuda12"):
selection_log.append(
"windows_cuda_selection: runtime_dirs "
f"{runtime_line}="
+ (
",".join(runtime_dirs.get(runtime_line, []))
if runtime_dirs.get(runtime_line)
else "none"
)
)
if detected_runtime_lines:
selection_log.append(
"windows_cuda_selection: host_runtime_order="
+ (",".join(normal_runtime_lines) if normal_runtime_lines else "none")
)
else:
selection_log.append(
"windows_cuda_selection: no CUDA runtime DLL line detected; falling back to driver order"
)
if not normal_runtime_lines:
if detected_runtime_lines:
selection_log.append(
"windows_cuda_selection: detected CUDA runtime DLLs were incompatible with the reported driver"
)
normal_runtime_lines = compatible_runtime_lines
runtime_order: list[str] = []
if preferred_runtime_line and preferred_runtime_line in normal_runtime_lines:
runtime_order.append(preferred_runtime_line)
selection_log.append(
"windows_cuda_selection: torch_preferred_runtime_line="
f"{preferred_runtime_line} reordered_attempts"
)
elif preferred_runtime_line:
selection_log.append(
"windows_cuda_selection: torch_preferred_runtime_line="
f"{preferred_runtime_line} unavailable_or_incompatible"
)
else:
selection_log.append("windows_cuda_selection: no Torch runtime preference available")
runtime_order.extend(
runtime_line for runtime_line in normal_runtime_lines if runtime_line not in runtime_order
)
# Keep every driver-compatible line reachable as a fallback, so a line
# gated out by driver version still drops to an older major (cuda13->cuda12).
runtime_order.extend(
runtime_line
for runtime_line in compatible_runtime_lines
if runtime_line not in runtime_order
)
selection_log.append(
"windows_cuda_selection: normal_runtime_order="
+ (",".join(normal_runtime_lines) if normal_runtime_lines else "none")
)
selection_log.append(
"windows_cuda_selection: attempt_runtime_order="
+ (",".join(runtime_order) if runtime_order else "none")
)
attempts: list[AssetChoice] = []
for runtime_line in runtime_order:
major = int(runtime_line.removeprefix("cuda"))
# Track whatever minor llama.cpp actually ships for this major
# (cuda13 -> 13.1, 13.3, ...). Skip the line when the release lacks a
# matching asset instead of guessing a now-missing name.
runtime = _published_windows_cuda_runtime(upstream_assets, major, host.driver_cuda_version)
if runtime is None:
selection_log.append(
f"windows_cuda_selection: no driver-supported asset for {runtime_line}"
)
continue
selected_name = None
asset_url = None
for candidate_name in windows_cuda_upstream_asset_names(llama_tag, runtime):
asset_url = upstream_assets.get(candidate_name)
if asset_url:
selected_name = candidate_name
break
if not asset_url or not selected_name:
selection_log.append(
"windows_cuda_selection: skip missing assets "
+ ",".join(windows_cuda_upstream_asset_names(llama_tag, runtime))
)
continue
# Pair the cudart bundle when upstream ships it; otherwise the binary
# needs a system CUDA toolkit on PATH at runtime (#5106). Only pair
# when the selected main archive is the binary archive, not the cudart
# archive itself.
runtime_archive_name: str | None = None
runtime_archive_url: str | None = None
if selected_name.startswith("llama-"):
cudart_name = f"cudart-llama-bin-win-cuda-{runtime}-x64.zip"
cudart_url = upstream_assets.get(cudart_name)
if cudart_url and cudart_url != asset_url:
runtime_archive_name = cudart_name
runtime_archive_url = cudart_url
attempt_log = list(selection_log) + [
f"windows_cuda_selection: selected {selected_name} runtime={runtime}"
]
if runtime_archive_name:
attempt_log.append(
f"windows_cuda_selection: paired runtime archive {runtime_archive_name}"
)
else:
attempt_log.append(
"windows_cuda_selection: no paired runtime archive found; "
"binary will rely on a system CUDA toolkit at runtime"
)
attempts.append(
AssetChoice(
repo = UPSTREAM_REPO,
tag = llama_tag,
name = selected_name,
url = asset_url,
source_label = "upstream",
install_kind = "windows-cuda",
runtime_line = runtime_line,
runtime_name = runtime_archive_name,
runtime_url = runtime_archive_url,
selection_log = attempt_log,
)
)
return attempts
def _windows_cuda_attempt_covers_blackwell(attempt: AssetChoice) -> bool:
"""True if an in-release windows-cuda attempt's toolkit covers Blackwell
sm_120 (>= 12.8), read from its asset name's CUDA minor."""
if attempt.install_kind != "windows-cuda":
return False
m = re.search(r"-bin-win-cuda-(\d+)\.(\d+)-x64\.zip$", attempt.name)
return m is not None and (int(m.group(1)), int(m.group(2))) >= _BLACKWELL_MIN_TOOLKIT
def _pinned_windows_cuda_fallback(
host: HostInfo, existing_cuda_attempts: list[AssetChoice]
) -> AssetChoice | None:
"""Pinned GPU fallback for a Blackwell host the in-release build gates off.
Upstream stopped publishing a sub-13.3 Windows cuda13 build after b9360,
and cuda-12.4 cannot offload sm_120, so a 13.1/13.2 driver would land on
CPU. b9360's cuda-13.1 build is immutable and runs on those drivers.
Returns None (dormant) whenever the in-release selection already offers a
Blackwell-capable build (toolkit >= 12.8, e.g. a runnable cuda13/cuda14),
so it self-disables once upstream ships a driver-runnable build again.
The b9360 binary reuses the current release's source tree and convert
scripts and is recorded via binary_release_tag, the same binary/source
split used for the lemonade prebuilt."""
if not (host.is_windows and host.is_x86_64 and host.has_usable_nvidia):
return None
driver = host.driver_cuda_version
if driver is None or driver < _PINNED_BLACKWELL_DRIVER_FLOOR:
return None
caps = normalize_compute_caps(host.compute_caps)
if not caps or int(caps[-1]) < _BLACKWELL_MIN_SM:
return None
if any(_windows_cuda_attempt_covers_blackwell(attempt) for attempt in existing_cuda_attempts):
return None
tag = _PINNED_BLACKWELL_FALLBACK_TAG
runtime = _PINNED_BLACKWELL_FALLBACK_RUNTIME
base = (
f"https://github.com/{UPSTREAM_REPO}/releases/download/"
f"{urllib.parse.quote(tag, safe = '')}"
)
name = f"llama-{tag}-bin-win-cuda-{runtime}-x64.zip"
cudart_name = f"cudart-llama-bin-win-cuda-{runtime}-x64.zip"
return AssetChoice(
repo = UPSTREAM_REPO,
tag = tag,
name = name,
url = f"{base}/{name}",
source_label = "upstream",
install_kind = "windows-cuda",
runtime_line = "cuda13",
runtime_name = cudart_name,
runtime_url = f"{base}/{cudart_name}",
expected_sha256 = _PINNED_BLACKWELL_LLAMA_SHA256,
runtime_sha256 = _PINNED_BLACKWELL_CUDART_SHA256,
selection_log = [
f"windows_cuda_selection: pinned {tag} cuda-{runtime} Blackwell GPU "
f"fallback (in-release cuda13 gated off by driver "
f"{driver[0]}.{driver[1]})"
],
)
def _augment_checksums_with_pin(
checksums: ApprovedReleaseChecksums, pin: AssetChoice
) -> ApprovedReleaseChecksums:
"""Add the pin's verified hashes to a copy of the approved checksums so
apply_approved_hashes keeps it on the published path (b9360 isn't in the
release manifest)."""
artifacts = dict(checksums.artifacts)
if pin.expected_sha256:
artifacts[pin.name] = ApprovedArtifactHash(
asset_name = pin.name,
sha256 = pin.expected_sha256,
repo = pin.repo,
kind = "prebuilt",
)
if pin.runtime_name and pin.runtime_sha256:
artifacts[pin.runtime_name] = ApprovedArtifactHash(
asset_name = pin.runtime_name,
sha256 = pin.runtime_sha256,
repo = pin.repo,
kind = "prebuilt",
)
return dataclasses_replace(checksums, artifacts = artifacts)
def _with_pinned_windows_cuda_fallback(
host: HostInfo, attempts: list[AssetChoice], checksums: ApprovedReleaseChecksums
) -> tuple[list[AssetChoice], ApprovedReleaseChecksums]:
"""Insert the Blackwell pin ahead of the Windows CUDA attempts and keep it
through apply_approved_hashes, or return inputs unchanged when dormant.
Gives the published install path the same GPU fallback as the simple path."""
pin = _pinned_windows_cuda_fallback(host, attempts)
if pin is None:
return attempts, checksums
return [pin, *attempts], _augment_checksums_with_pin(checksums, pin)
def published_windows_cuda_attempts(
host: HostInfo,
release: PublishedReleaseBundle,
preferred_runtime_line: str | None,
selection_preamble: Iterable[str] = (),
) -> list[AssetChoice]:
selection_log = list(release.selection_log) + list(selection_preamble)
# Seed the runtime-line ordering from the real published windows-cuda minors
# (their names encode the minor), so a future CUDA major published here is
# ordered too rather than a hardcoded cuda12/cuda13 pair. Keys mirror the
# upstream naming so windows_cuda_attempts can match them; fall back to the
# long-standing default when the release lists no windows-cuda asset.
published_minors: list[str] = []
for artifact in release.artifacts:
if artifact.install_kind != "windows-cuda":
continue
m = re.search(r"-bin-win-cuda-(\d+\.\d+)-x64\.zip$", artifact.asset_name)
if m:
published_minors.append(m.group(1))
if not published_minors:
published_minors = ["12.4", "13.1"]
runtime_order = windows_cuda_attempts(
host,
release.upstream_tag,
{
f"llama-{release.upstream_tag}-bin-win-cuda-{minor}-x64.zip": "published"
for minor in published_minors
},
preferred_runtime_line,
selection_log,
)
published_artifacts = [
artifact for artifact in release.artifacts if artifact.install_kind == "windows-cuda"
]
artifacts_by_runtime: dict[str, list[PublishedLlamaArtifact]] = {}
for artifact in published_artifacts:
if not artifact.runtime_line:
continue
artifacts_by_runtime.setdefault(artifact.runtime_line, []).append(artifact)
attempts: list[AssetChoice] = []
for ordered_attempt in runtime_order:
runtime_line = ordered_attempt.runtime_line
if not runtime_line:
continue
candidates = sorted(
artifacts_by_runtime.get(runtime_line, []),
key = lambda artifact: (artifact.rank, artifact.asset_name),
)
for artifact in candidates:
asset_url = release.assets.get(artifact.asset_name)
if not asset_url:
continue
am = re.search(r"-bin-win-cuda-(\d+)\.(\d+)-x64\.zip$", artifact.asset_name)
# Gate the published minor against the driver so it can never
# bypass the driver-version gate.
if (
am is not None
and host.driver_cuda_version is not None
and (int(am.group(1)), int(am.group(2))) > host.driver_cuda_version
):
continue
# See windows_cuda_attempts: pair the cudart bundle for the real minor.
runtime_archive_name: str | None = None
runtime_archive_url: str | None = None
if am is not None and artifact.asset_name.startswith("llama-"):
runtime = f"{am.group(1)}.{am.group(2)}"
cudart_name = f"cudart-llama-bin-win-cuda-{runtime}-x64.zip"
cudart_url = release.assets.get(cudart_name)
if cudart_url and cudart_url != asset_url:
runtime_archive_name = cudart_name
runtime_archive_url = cudart_url
attempt_log = list(ordered_attempt.selection_log or []) + [
"windows_cuda_selection: selected published asset "
f"{artifact.asset_name} for runtime_line={runtime_line}"
]
if runtime_archive_name:
attempt_log.append(
f"windows_cuda_selection: paired published runtime archive {runtime_archive_name}"
)
attempts.append(
AssetChoice(
repo = release.repo,
tag = release.release_tag,
name = artifact.asset_name,
url = asset_url,
source_label = "published",
install_kind = "windows-cuda",
runtime_line = runtime_line,
runtime_name = runtime_archive_name,
runtime_url = runtime_archive_url,
selection_log = attempt_log,
)
)
break
return attempts
def resolve_windows_cuda_choices(
host: HostInfo, llama_tag: str, upstream_assets: dict[str, str]
) -> list[AssetChoice]:
torch_preference = detect_torch_cuda_runtime_preference(host)
attempts = windows_cuda_attempts(
host,
llama_tag,
upstream_assets,
torch_preference.runtime_line,
torch_preference.selection_log,
)
return attempts
def resolve_linux_cuda_choice(
host: HostInfo, release: PublishedReleaseBundle
) -> LinuxCudaSelection:
torch_preference = detect_torch_cuda_runtime_preference(host)
selection = linux_cuda_choice_from_release(
host,
release,
preferred_runtime_line = torch_preference.runtime_line,
selection_preamble = torch_preference.selection_log,
)
if selection is not None:
return selection
raise PrebuiltFallback("no compatible published Linux CUDA bundle was found")
def published_asset_choice_for_kind(
release: PublishedReleaseBundle, install_kind: str
) -> AssetChoice | None:
candidates = sorted(
(artifact for artifact in release.artifacts if artifact.install_kind == install_kind),
key = lambda artifact: (artifact.rank, artifact.asset_name),
)
for artifact in candidates:
asset_url = release.assets.get(artifact.asset_name)
if not asset_url:
continue
return AssetChoice(
repo = release.repo,
tag = release.release_tag,
name = artifact.asset_name,
url = asset_url,
source_label = "published",
install_kind = install_kind,
runtime_line = artifact.runtime_line,
selection_log = list(release.selection_log)
+ [f"published_selection: selected {artifact.asset_name} install_kind={install_kind}"],
)
return None
def _detect_host_rocm_version() -> tuple[int, int] | None:
"""Return (major, minor) of the installed ROCm runtime, or None.
Best-effort read from /opt/rocm/.info/version, amd-smi version, and
hipconfig --version. Used to pick a compatible upstream llama.cpp ROCm
prebuilt rather than the numerically newest one (which can be newer than
the host runtime).
"""
rocm_root = os.environ.get("ROCM_PATH") or "/opt/rocm"
for path in (
os.path.join(rocm_root, ".info", "version"),
os.path.join(rocm_root, "lib", "rocm_version"),
):
try:
with open(path) as fh:
parts = fh.read().strip().split("-")[0].split(".")
# Explicit length guard so we don't rely on the broad except
# below to swallow IndexError when the version file has a single
# component (e.g. "6\n" on a partial install).
if len(parts) >= 2:
return int(parts[0]), int(parts[1])
except Exception:
pass
amd_smi = shutil.which("amd-smi") if _amd_smi_allowed() else None
if amd_smi:
try:
# Off on Windows w/o a HIP SDK (avoids the UAC/DiskPart prompt);
# hipconfig below and the version-file reads above cover that case.
result = run_capture([amd_smi, "version"], timeout = 5)
if result.returncode == 0:
m = re.search(r"ROCm version:\s*(\d+)\.(\d+)", result.stdout)
if m:
return int(m.group(1)), int(m.group(2))
except Exception:
pass
hipconfig = shutil.which("hipconfig")
if hipconfig:
try:
result = subprocess.run(
[hipconfig, "--version"],
stdout = subprocess.PIPE,
stderr = subprocess.DEVNULL,
text = True,
timeout = 5,
)
if result.returncode == 0:
raw = (result.stdout or "").strip().split("\n")[0]
parts = raw.split(".")
if len(parts) >= 2 and parts[0].isdigit() and parts[1].split("-")[0].isdigit():
return int(parts[0]), int(parts[1].split("-")[0])
except Exception:
pass
# Distro package-manager fallbacks. Mirrors install.sh::get_torch_index_url
# and _detect_rocm_version() in install_python_stack.py so package-managed
# ROCm hosts without /opt/rocm/.info/version still report a usable version,
# letting the <= host version filter in resolve_upstream_asset_choice pick
# the correct upstream prebuilt instead of the newest-regardless fallback.
for _cmd in (
["dpkg-query", "-W", "-f=${Version}\n", "rocm-core"],
["rpm", "-q", "--qf", "%{VERSION}\n", "rocm-core"],
):
_exe = shutil.which(_cmd[0])
if not _exe:
continue
try:
_result = subprocess.run(
[_exe, *_cmd[1:]],
stdout = subprocess.PIPE,
stderr = subprocess.DEVNULL,
text = True,
timeout = 5,
)
except Exception:
continue
if _result.returncode != 0 or not _result.stdout.strip():
continue
_raw = _result.stdout.strip()
# dpkg can prepend an epoch ("1:6.3.0-1"); strip it first.
_raw = re.sub(r"^\d+:", "", _raw)
_m = re.match(r"(\d+)[.-](\d+)", _raw)
if _m:
return int(_m.group(1)), int(_m.group(2))
return None
# Map detected gfx IDs to lemonade-sdk asset family suffixes.
# More-specific prefixes must precede shorter ones (e.g. gfx1151 before gfx110).
_LEMONADE_GFX_FAMILIES: list[tuple[str, str]] = [
("gfx1151", "gfx1151"),
("gfx1150", "gfx1150"),
("gfx120", "gfx120X"),
("gfx110", "gfx110X"),
("gfx103", "gfx103X"),
]
def _lemonade_gfx_family(gfx_id: str) -> str | None:
gfx_id = gfx_id.lower().strip()
for prefix, family in _LEMONADE_GFX_FAMILIES:
if gfx_id.startswith(prefix):
return family
return None
def _is_trusted_github_release_url(url: str, expected_repo: str) -> bool:
"""Validate a release asset URL points at GitHub's expected hosts.
Accepts:
https://github.com/{expected_repo}/releases/download/...
https://objects.githubusercontent.com/... (GitHub's release CDN)
Anything else (http://, raw.githubusercontent.com, gist, etc.) is rejected
so a malicious API response can't redirect downloads to an attacker host.
"""
if not isinstance(url, str) or not url:
return False
try:
parsed = urllib.parse.urlparse(url)
except Exception:
return False
if parsed.scheme != "https":
return False
host = (parsed.netloc or "").lower()
if host == "objects.githubusercontent.com":
# GitHub's release CDN. Restrict to release-asset paths so a tampered
# API response pointing at an arbitrary CDN object is rejected. Real
# release asset URLs carry the "/github-production-release-asset-"
# prefix; gist / raw / avatar CDN paths do not.
return parsed.path.startswith("/github-production-release-asset-")
if host == "github.com":
return parsed.path.startswith(f"/{expected_repo}/releases/download/")
return False
@functools.lru_cache(maxsize = 8)
def _fetch_lemonade_release_cached(api_url: str, llama_tag: str) -> "dict | None":
"""Cached wrapper around fetch_json for lemonade release lookups.
resolve_lemonade_rocm_choice() is called twice per install (direct planner
+ resolve_upstream_asset_choice) with identical arguments. Without
memoisation each install hits api.github.com twice, doubling the
rate-limit failure surface on busy CI runners. Cache is process-scoped;
tests that vary fetch_json's return value across calls should call
cache_clear().
"""
try:
return fetch_json(api_url)
except Exception as exc:
normalized = (llama_tag or "").strip().lower()
if normalized and normalized != "latest":
log(
f"Could not fetch {LEMONADE_ROCM_REPO} release for "
f"llama_tag={llama_tag!r} ({exc}); skipping lemonade prebuilt"
)
else:
log(f"Could not fetch {LEMONADE_ROCM_REPO} latest release: {exc}")
return None
def resolve_lemonade_rocm_choice(
host: HostInfo,
os_prefix: str,
install_kind: str,
llama_tag: str = "latest",
) -> "AssetChoice | None":
"""Return an AssetChoice from lemonade-sdk/llamacpp-rocm for the detected GPU, or None.
os_prefix: lemonade's asset filename label, NOT a host-distro filter.
Pass "ubuntu" for any Linux host (Arch, Fedora, openSUSE,
Debian, ...) -- lemonade publishes one Linux variant, a
manylinux-style glibc build that runs on any distro with a
recent-enough glibc. Pass "windows" for Windows hosts.
install_kind: "linux-rocm" or "windows-hip"
llama_tag: the requested upstream llama.cpp tag ("latest" or a pinned
release like "b1260"). When pinned, fetch the matching
lemonade release; if lemonade hasn't published that tag, skip
silently (caller falls through to upstream) rather than drift
to whatever lemonade ships as latest.
"""
if not host.rocm_gfx_target:
return None
# Opt-out for users who want the upstream HIP build path only -- lemonade
# binaries lack approved-hash manifest entries, so their integrity gate is
# functional validation only.
if os.environ.get("UNSLOTH_DISABLE_LEMONADE_ROCM", "").strip().lower() in (
"1",
"true",
"yes",
):
log("UNSLOTH_DISABLE_LEMONADE_ROCM is set; skipping lemonade-sdk prebuilt")
return None
gfx_family = _lemonade_gfx_family(host.rocm_gfx_target)
if gfx_family is None:
log(
f"AMD GPU {host.rocm_gfx_target!r} is not covered by lemonade-sdk ROCm prebuilts; "
"skipping lemonade prebuilt"
)
return None
api_url = _lemonade_release_api_for(llama_tag)
release = _fetch_lemonade_release_cached(api_url, llama_tag)
if release is None:
return None
release_tag = release.get("tag_name") if isinstance(release, dict) else None
if not isinstance(release_tag, str) or not release_tag:
log(f"Unexpected {LEMONADE_ROCM_REPO} release payload; skipping lemonade prebuilt")
return None
assets = release_asset_map(release)
asset_name = f"llama-{release_tag}-{os_prefix}-rocm-{gfx_family}-x64.zip"
if asset_name not in assets:
log(
f"{LEMONADE_ROCM_REPO}@{release_tag} has no asset {asset_name!r}; "
"skipping lemonade prebuilt"
)
return None
asset_url = assets[asset_name]
if not asset_url:
# release_asset_map defaults to "" when an asset row lacks
# browser_download_url; skip cleanly instead of letting
# download_file("") raise a less obvious error downstream.
log(
f"{LEMONADE_ROCM_REPO}@{release_tag} asset {asset_name!r} has no "
"browser_download_url; skipping lemonade prebuilt"
)
return None
# Defence in depth: lemonade browser_download_url should be on github.com
# or githubusercontent.com. A compromised GitHub API response redirecting
# to an attacker host would otherwise be honoured silently (lemonade
# assets are not in the approved-hash manifest).
if not _is_trusted_github_release_url(asset_url, LEMONADE_ROCM_REPO):
log(
f"{LEMONADE_ROCM_REPO}@{release_tag} asset {asset_name!r} points "
f"to an unexpected host ({asset_url!r}); refusing to download "
"lemonade prebuilt"
)
return None
# Note: lemonade tags Linux assets "ubuntu" but the binary is a generic
# glibc build that runs on any distro (Arch, Fedora, ...), so this attempt
# is selected for all Linux ROCm hosts, not just Ubuntu.
log(
f"AMD GPU {host.rocm_gfx_target!r} ({gfx_family}) -- "
f"trying lemonade-sdk ROCm prebuilt {asset_name} "
f"(works on any glibc Linux, not just Ubuntu)"
)
log(
f"NOTE: lemonade-sdk/llamacpp-rocm releases are not covered by the "
f"Unsloth approved-hash manifest; download integrity relies on "
f"functional validation (llama-bench / llama-server smoke tests) "
f"after extraction. Set UNSLOTH_DISABLE_LEMONADE_ROCM=1 to skip "
f"lemonade and fall back to the upstream HIP build path."
)
return AssetChoice(
repo = LEMONADE_ROCM_REPO,
tag = release_tag,
name = asset_name,
url = asset_url,
source_label = "lemonade",
install_kind = install_kind,
)
def resolve_upstream_asset_choice(host: HostInfo, llama_tag: str) -> AssetChoice:
upstream_assets = github_release_assets(UPSTREAM_REPO, llama_tag)
if host.is_linux and host.is_x86_64:
# AMD ROCm: try upstream ROCm prebuilt first, then a source build. The
# source build (via setup.sh) compiles with -DGGML_HIP=ON and
# auto-detects the exact GPU target via rocminfo, more reliable for
# consumer GPUs (e.g. gfx1151) that may not be in the prebuilt.
if host.has_rocm and not host.has_usable_nvidia:
# Try lemonade-sdk per-GPU prebuilt first: built against specific
# gfx targets and bundle all required ROCm runtime libs.
lemonade_choice = resolve_lemonade_rocm_choice(
host, "ubuntu", "linux-rocm", llama_tag = llama_tag
)
if lemonade_choice is not None:
return lemonade_choice
# Fall back to the upstream combined ROCm tarball. Scan for any
# rocm-<version> prebuilt; when the host ROCm version is known,
# pick the newest candidate whose major.minor is <= host version
# -- otherwise a ROCm 6.4 host downloads the rocm-7.2 tarball,
# fails preflight, and source-builds even though a 6.4 prebuilt
# exists. If none is compatible (host older than every published
# prebuilt), fall back to the numerically newest so we try
# something.
_rocm_pattern = re.compile(
rf"llama-{re.escape(llama_tag)}-bin-ubuntu-rocm-([0-9]+(?:\.[0-9]+)*)-x64\.tar\.gz"
)
rocm_candidates: list[tuple[tuple[int, ...], str]] = []
for _name in upstream_assets:
_m = _rocm_pattern.match(_name)
if _m is None:
continue
_parts = tuple(int(p) for p in _m.group(1).split("."))
rocm_candidates.append((_parts, _name))
rocm_candidates.sort(reverse = True)
_host_rocm_version = _detect_host_rocm_version()
_compatible: list[tuple[tuple[int, ...], str]] = rocm_candidates
if _host_rocm_version is not None:
_compatible = [
item for item in rocm_candidates if item[0][:2] <= _host_rocm_version
]
if rocm_candidates and not _compatible:
# Fall back to the newest candidate so we don't force a source
# build when the host runtime is older than every published
# prebuilt: preflight still catches a true incompatibility and
# triggers a fallback.
_compatible = rocm_candidates[:1]
if _compatible:
rocm_name = _compatible[0][1]
if _host_rocm_version is not None:
log(
f"AMD ROCm {_host_rocm_version[0]}.{_host_rocm_version[1]} "
f"detected -- trying upstream prebuilt {rocm_name}"
)
else:
log(f"AMD ROCm detected -- trying upstream prebuilt {rocm_name}")
log(
"Note: if your ROCm runtime version differs significantly, "
"this may fail preflight and fall back to a source build (safe)"
)
return AssetChoice(
repo = UPSTREAM_REPO,
tag = llama_tag,
name = rocm_name,
url = upstream_assets[rocm_name],
source_label = "upstream",
install_kind = "linux-rocm",
)
# No ROCm prebuilt available -- fall back to a source build
raise PrebuiltFallback(
"AMD ROCm detected but no upstream ROCm prebuilt found; "
"falling back to source build with HIP support"
)
upstream_name = f"llama-{llama_tag}-bin-ubuntu-x64.tar.gz"
if upstream_name not in upstream_assets:
raise PrebuiltFallback("upstream Linux CPU asset was not found")
return AssetChoice(
repo = UPSTREAM_REPO,
tag = llama_tag,
name = upstream_name,
url = upstream_assets[upstream_name],
source_label = "upstream",
install_kind = "linux-cpu",
)
if host.is_windows and host.is_x86_64:
if host.has_usable_nvidia:
attempts = resolve_windows_cuda_choices(host, llama_tag, upstream_assets)
if attempts:
return attempts[0]
raise PrebuiltFallback("no compatible Windows CUDA asset was found")
# AMD ROCm on Windows: try lemonade per-GPU prebuilt first, then upstream HIP
if host.has_rocm:
lemonade_choice = resolve_lemonade_rocm_choice(
host, "windows", "windows-hip", llama_tag = llama_tag
)
if lemonade_choice is not None:
return lemonade_choice
hip_name = f"llama-{llama_tag}-bin-win-hip-radeon-x64.zip"
if hip_name in upstream_assets:
log(f"AMD ROCm detected on Windows -- trying upstream HIP prebuilt {hip_name}")
return AssetChoice(
repo = UPSTREAM_REPO,
tag = llama_tag,
name = hip_name,
url = upstream_assets[hip_name],
source_label = "upstream",
install_kind = "windows-hip",
)
log("AMD ROCm detected on Windows but no HIP prebuilt found -- falling back to CPU")
upstream_name = f"llama-{llama_tag}-bin-win-cpu-x64.zip"
if upstream_name not in upstream_assets:
raise PrebuiltFallback("upstream Windows CPU asset was not found")
return AssetChoice(
repo = UPSTREAM_REPO,
tag = llama_tag,
name = upstream_name,
url = upstream_assets[upstream_name],
source_label = "upstream",
install_kind = "windows-cpu",
)
if host.is_macos and host.is_arm64:
upstream_name = f"llama-{llama_tag}-bin-macos-arm64.tar.gz"
if upstream_name not in upstream_assets:
raise PrebuiltFallback("upstream macOS arm64 asset was not found")
return AssetChoice(
repo = UPSTREAM_REPO,
tag = llama_tag,
name = upstream_name,
url = upstream_assets[upstream_name],
source_label = "upstream",
install_kind = "macos-arm64",
)
if host.is_macos and host.is_x86_64:
upstream_name = f"llama-{llama_tag}-bin-macos-x64.tar.gz"
if upstream_name not in upstream_assets:
raise PrebuiltFallback("upstream macOS x64 asset was not found")
return AssetChoice(
repo = UPSTREAM_REPO,
tag = llama_tag,
name = upstream_name,
url = upstream_assets[upstream_name],
source_label = "upstream",
install_kind = "macos-x64",
)
raise PrebuiltFallback(f"no prebuilt policy exists for {host.system} {host.machine}")
def resolve_asset_choice(host: HostInfo, llama_tag: str) -> AssetChoice:
if host.is_linux and host.is_x86_64 and host.has_usable_nvidia:
raise PrebuiltFallback(
"Linux CUDA installs require a compatible published bundle; upstream fallback is not available"
)
return resolve_upstream_asset_choice(host, llama_tag)
def resolve_release_asset_choice(
host: HostInfo,
llama_tag: str,
release: PublishedReleaseBundle,
checksums: ApprovedReleaseChecksums,
) -> list[AssetChoice]:
if host.is_windows and host.is_x86_64 and host.has_usable_nvidia:
torch_preference = detect_torch_cuda_runtime_preference(host)
published_attempts = published_windows_cuda_attempts(
host,
release,
torch_preference.runtime_line,
torch_preference.selection_log,
)
if published_attempts:
pin_attempts, pin_checksums = _with_pinned_windows_cuda_fallback(
host, published_attempts, checksums
)
try:
return apply_approved_hashes(pin_attempts, pin_checksums)
except PrebuiltFallback as exc:
log(
"published Windows CUDA assets ignored for install planning: "
f"{release.repo}@{release.release_tag} ({exc})"
)
upstream_assets = github_release_assets(UPSTREAM_REPO, llama_tag)
upstream_attempts, upstream_checksums = _with_pinned_windows_cuda_fallback(
host,
resolve_windows_cuda_choices(host, llama_tag, upstream_assets),
checksums,
)
return apply_approved_hashes(upstream_attempts, upstream_checksums)
published_choice: AssetChoice | None = None
if host.is_windows and host.is_x86_64:
# AMD Windows hosts prefer a hash-approved published Windows HIP
# bundle when one exists, otherwise fall through to
# resolve_asset_choice() so the upstream HIP prebuilt is tried before
# the CPU fallback. Hard-pinning the published windows-cpu bundle here
# would make the HIP path unreachable.
if host.has_rocm:
published_choice = published_asset_choice_for_kind(release, "windows-hip")
else:
published_choice = published_asset_choice_for_kind(release, "windows-cpu")
elif host.is_macos and host.is_arm64:
published_choice = published_asset_choice_for_kind(release, "macos-arm64")
elif host.is_macos and host.is_x86_64:
published_choice = published_asset_choice_for_kind(release, "macos-x64")
if published_choice is not None:
try:
return apply_approved_hashes([published_choice], checksums)
except PrebuiltFallback as exc:
log(
"published platform asset ignored for install planning: "
f"{release.repo}@{release.release_tag} {published_choice.name} ({exc})"
)
return apply_approved_hashes([resolve_asset_choice(host, llama_tag)], checksums)
def extract_archive(archive_path: Path, destination: Path) -> None:
def safe_extract_path(base: Path, member_name: str) -> Path:
normalized = member_name.replace("\\", "/")
member_path = Path(normalized)
if member_path.is_absolute():
raise PrebuiltFallback(f"archive member used an absolute path: {member_name}")
target = (base / member_path).resolve()
base_resolved = base.resolve()
try:
target.relative_to(base_resolved)
except ValueError as exc:
raise PrebuiltFallback(f"archive member escaped destination: {member_name}") from exc
return target
def _try_repair_missing_slash(
member_name: str, link_name: str, archive_names: set[str]
) -> str | None:
"""Some upstream llama.cpp Mac releases (e.g. b9165, b9169) ship
symlinks whose linkname is missing the directory separator AND
the leading character of the file basename between the
top-level dir and the rest of the path:
llama-b9165/libggml-rpc.0.dylib -> llama-b9165ibggml-rpc.0.11.1.dylib
That cannot be resolved as written. Detect the pattern
(linkname starts with the top-level dir name but no following
slash) and search archive entries under that dir for a real
file whose basename ends with the mangled suffix. Only accept
when the suffix uniquely identifies a real archive entry.
Returns the corrected linkname expressed relative to the
member's parent directory -- callers join it with
`target.parent`, so a full `top/file` path would double the
prefix into `top/top/file`."""
if "/" not in member_name or "/" in link_name:
return None
top, _, _ = member_name.partition("/")
if not link_name.startswith(top) or len(link_name) <= len(top):
return None
bad_suffix = link_name[len(top) :]
if not bad_suffix or bad_suffix.startswith("/"):
return None
prefix = f"{top}/"
candidates = [
name
for name in archive_names
if name.startswith(prefix)
and "/" not in name[len(prefix) :]
and name[len(prefix) :].endswith(bad_suffix)
]
if len(candidates) != 1:
return None
# Strip the top-level dir so the caller's `target.parent / Path(...)`
# composition resolves inside the staging dir, not into a duplicate
# `top/top/...` path.
return candidates[0][len(prefix) :]
def safe_link_target(
base: Path, member_name: str, link_name: str, target: Path, archive_names: set[str]
) -> tuple[str, Path]:
normalized = link_name.replace("\\", "/")
repaired = _try_repair_missing_slash(member_name, normalized, archive_names)
if repaired is not None:
normalized = repaired
link_path = Path(normalized)
if link_path.is_absolute():
raise PrebuiltFallback(
f"archive link used an absolute target: {member_name} -> {link_name}"
)
if not normalized:
raise PrebuiltFallback(f"archive link used an empty target: {member_name}")
resolved = (target.parent / link_path).resolve()
base_resolved = base.resolve()
try:
resolved.relative_to(base_resolved)
except ValueError as exc:
raise PrebuiltFallback(
f"archive link escaped destination: {member_name} -> {link_name}"
) from exc
return normalized, resolved
def extract_zip_safely(source: Path, base: Path) -> None:
with zipfile.ZipFile(source) as archive:
for member in archive.infolist():
target = safe_extract_path(base, member.filename)
mode = (member.external_attr >> 16) & 0o170000
if mode == 0o120000:
raise PrebuiltFallback(
f"zip archive contained a symlink entry: {member.filename}"
)
if member.is_dir():
target.mkdir(parents = True, exist_ok = True)
continue
target.parent.mkdir(parents = True, exist_ok = True)
with archive.open(member, "r") as src, target.open("wb") as dst:
shutil.copyfileobj(src, dst)
def extract_tar_safely(source: Path, base: Path) -> None:
pending_links: list[tuple[tarfile.TarInfo, Path]] = []
archive_names: set[str] = set()
with tarfile.open(source, "r:gz") as archive:
for member in archive.getmembers():
archive_names.add(member.name)
target = safe_extract_path(base, member.name)
if member.isdir():
target.mkdir(parents = True, exist_ok = True)
continue
if member.islnk() or member.issym():
pending_links.append((member, target))
continue
if not member.isfile():
raise PrebuiltFallback(
f"tar archive contained an unsupported entry: {member.name}"
)
target.parent.mkdir(parents = True, exist_ok = True)
extracted = archive.extractfile(member)
if extracted is None:
raise PrebuiltFallback(f"tar archive entry could not be read: {member.name}")
with extracted, target.open("wb") as dst:
shutil.copyfileobj(extracted, dst)
unresolved = list(pending_links)
while unresolved:
next_round: list[tuple[tarfile.TarInfo, Path]] = []
progressed = False
for member, target in unresolved:
normalized_link, resolved_target = safe_link_target(
base, member.name, member.linkname, target, archive_names
)
if not resolved_target.exists() and not resolved_target.is_symlink():
next_round.append((member, target))
continue
if resolved_target.is_dir():
raise PrebuiltFallback(
f"archive link targeted a directory: {member.name} -> {member.linkname}"
)
target.parent.mkdir(parents = True, exist_ok = True)
if target.exists() or target.is_symlink():
target.unlink()
if member.issym():
target.symlink_to(normalized_link)
else:
shutil.copy2(resolved_target, target)
progressed = True
if not progressed:
details = ", ".join(
f"{member.name} -> {member.linkname}" for member, _ in next_round
)
raise PrebuiltFallback(f"tar archive contained unresolved link entries: {details}")
unresolved = next_round
destination.mkdir(parents = True, exist_ok = True)
if archive_path.name.endswith(".zip"):
extract_zip_safely(archive_path, destination)
return
if archive_path.name.endswith(".tar.gz"):
extract_tar_safely(archive_path, destination)
return
raise PrebuiltFallback(f"unsupported archive format: {archive_path.name}")
def copy_globs(
source_dir: Path,
destination: Path,
patterns: list[str],
*,
required: bool = True,
) -> None:
destination.mkdir(parents = True, exist_ok = True)
matched_sources: dict[str, Path] = {}
for path in sorted(
(candidate for candidate in source_dir.rglob("*") if candidate.is_file()),
key = lambda candidate: (
len(candidate.relative_to(source_dir).parts),
str(candidate),
),
):
for pattern in patterns:
if fnmatch.fnmatch(path.name, pattern):
previous = matched_sources.get(path.name)
if previous is not None and previous != path:
raise PrebuiltFallback(
f"ambiguous archive layout for {path.name}: "
f"{previous.relative_to(source_dir)} and {path.relative_to(source_dir)}"
)
matched_sources[path.name] = path
break
if required and not matched_sources:
raise PrebuiltFallback(f"required files missing from {source_dir}: {patterns}")
for name, path in matched_sources.items():
shutil.copy2(path, destination / name)
def ensure_converter_scripts(install_dir: Path, llama_tag: str) -> None:
canonical = install_dir / "convert_hf_to_gguf.py"
if not canonical.exists():
# Hydrated source tree should have placed this file already.
# Fall back to a network fetch so the install is not blocked.
raw_base = f"https://raw.githubusercontent.com/ggml-org/llama.cpp/{llama_tag}"
source_url = f"{raw_base}/convert_hf_to_gguf.py"
data = download_bytes(
source_url,
progress_label = f"Downloading {download_label_from_url(source_url)}",
)
if not data:
raise RuntimeError(f"downloaded empty converter script from {source_url}")
if b"import " not in data and b"def " not in data and b"#!/" not in data:
raise RuntimeError(
f"downloaded converter script did not look like Python source: {source_url}"
)
atomic_write_bytes(canonical, data)
legacy = install_dir / "convert-hf-to-gguf.py"
if legacy.exists() or legacy.is_symlink():
legacy.unlink()
try:
legacy.symlink_to("convert_hf_to_gguf.py")
except OSError:
shutil.copy2(canonical, legacy)
def extracted_archive_root(extract_dir: Path) -> Path:
children = [path for path in extract_dir.iterdir()]
if len(children) == 1 and children[0].is_dir():
return children[0]
return extract_dir
def copy_directory_contents(source_dir: Path, destination: Path) -> None:
destination.mkdir(parents = True, exist_ok = True)
for item in source_dir.iterdir():
target = destination / item.name
if item.is_dir():
shutil.copytree(item, target, dirs_exist_ok = True)
else:
shutil.copy2(item, target)
def hydrate_source_tree(
source_ref: str,
install_dir: Path,
work_dir: Path,
*,
source_repo: str = UPSTREAM_REPO,
expected_sha256: str | None,
source_label: str | None = None,
exact_source: bool = False,
) -> None:
archive_path = work_dir / f"llama.cpp-source-{source_ref}.tar.gz"
source_urls = (
commit_source_archive_urls(source_repo, source_ref)
if exact_source
else upstream_source_archive_urls(source_ref)
)
label = source_label or f"llama.cpp source tree for {source_ref}"
extract_dir = Path(tempfile.mkdtemp(prefix = "source-extract-", dir = work_dir))
try:
log(f"downloading {label}")
last_exc: Exception | None = None
downloaded = False
for index, source_url in enumerate(source_urls):
try:
if index > 0:
log(f"retrying source tree download from fallback URL: {source_url}")
download_file_verified(
source_url,
archive_path,
expected_sha256 = expected_sha256,
label = label,
)
downloaded = True
break
except Exception as exc:
last_exc = exc
if index == len(source_urls) - 1:
raise
log(f"source tree download failed from {source_url}: {exc}")
if not downloaded:
assert last_exc is not None
raise last_exc
extract_archive(archive_path, extract_dir)
source_root = extracted_archive_root(extract_dir)
required_paths = [
source_root / "CMakeLists.txt",
source_root / "convert_hf_to_gguf.py",
source_root / "gguf-py",
]
missing = [
str(path.relative_to(source_root)) for path in required_paths if not path.exists()
]
if missing:
raise PrebuiltFallback(
"upstream source archive was missing required repo files: " + ", ".join(missing)
)
copy_directory_contents(source_root, install_dir)
except PrebuiltFallback:
raise
except Exception as exc:
raise PrebuiltFallback(f"failed to hydrate {label}: {exc}") from exc
finally:
remove_tree(extract_dir)
def normalize_install_layout(install_dir: Path, host: HostInfo) -> tuple[Path, Path]:
build_bin = install_dir / "build" / "bin"
if host.is_windows:
exec_dir = build_bin / "Release"
exec_dir.mkdir(parents = True, exist_ok = True)
return exec_dir / "llama-server.exe", exec_dir / "llama-quantize.exe"
install_dir.mkdir(parents = True, exist_ok = True)
build_bin.mkdir(parents = True, exist_ok = True)
return install_dir / "llama-server", install_dir / "llama-quantize"
def discover_installed_executable(install_dir: Path, executable_name: str) -> Path:
direct = install_dir / executable_name
if direct.exists() and direct.is_file():
return direct
candidate = next((path for path in install_dir.rglob(executable_name) if path.is_file()), None)
if candidate is None:
raise PrebuiltFallback(f"{executable_name} was not installed")
return candidate
def write_exec_wrapper(entrypoint: Path, target: Path) -> None:
relative_target = os.path.relpath(target, entrypoint.parent)
script = "\n".join(
[
"#!/bin/sh",
f'exec "$(dirname "$0")/{relative_target}" "$@"',
"",
]
)
atomic_write_bytes(entrypoint, script.encode("utf-8"))
os.chmod(entrypoint, 0o755)
def create_exec_entrypoint(entrypoint: Path, target: Path) -> None:
if entrypoint == target:
return
if entrypoint.exists() or entrypoint.is_symlink():
entrypoint.unlink()
try:
entrypoint.symlink_to(os.path.relpath(target, entrypoint.parent))
except Exception:
write_exec_wrapper(entrypoint, target)
def overlay_directory_for_choice(install_dir: Path, choice: AssetChoice, host: HostInfo) -> Path:
if host.is_windows or choice.install_kind.startswith("windows"):
path = install_dir / "build" / "bin" / "Release"
else:
path = install_dir / "build" / "bin"
path.mkdir(parents = True, exist_ok = True)
return path
def paired_runtime_dll_patterns(choice: AssetChoice) -> list[str]:
"""Filename patterns the paired runtime archive is allowed to drop
into the install. Used for the second copy_globs pass in
install_from_archives, narrower than runtime_patterns_for_choice so
the runtime archive cannot overwrite main-archive payload like
llama-server.exe. Only Windows CUDA has paired runtimes today."""
if choice.install_kind == "windows-cuda":
return ["cudart64_*.dll", "cublas64_*.dll", "cublasLt64_*.dll"]
return []
def runtime_patterns_for_choice(choice: AssetChoice) -> list[str]:
# Broad shared-library glob + explicit binary names. Lets upstream
# repackage the SO/DLL set (e.g. ggml-org/llama.cpp#23462 split the
# per-binary entry code into paired ``lib<binary>-impl.so`` shared
# libraries between b9279 and b9283) without us re-enumerating
# every new file. Studio only invokes llama-server and llama-quantize;
# other CLIs upstream ships (llama-cli, llama-bench, ...) are skipped.
if choice.install_kind in {"linux-cpu", "linux-cuda", "linux-rocm", "linux-arm64"}:
return ["llama-server", "llama-quantize", "lib*.so*"]
if choice.install_kind in {"macos-arm64", "macos-x64"}:
return ["llama-server", "llama-quantize", "lib*.dylib"]
if choice.install_kind in {
"windows-cpu",
"windows-cuda",
"windows-hip",
"windows-arm64",
}:
return ["llama-server.exe", "llama-quantize.exe", "*.dll"]
raise PrebuiltFallback(f"unsupported install kind for runtime overlay: {choice.install_kind}")
def runtime_subdirs_for_choice(choice: AssetChoice) -> list[str]:
"""Subdirectory names within the archive root that must be copied into
the overlay directory alongside the flat shared libraries.
hipBLASLt and rocBLAS expect their Tensile kernel catalog trees
(hipblaslt/library/<gfx>/ and rocblas/library/<gfx>/) to sit next to
their shared libraries at runtime. These trees are multi-level and
cannot be handled by copy_globs (filename-only matching, flat copy)."""
if choice.source_label == "lemonade" and choice.install_kind in {
"linux-rocm",
"windows-hip",
}:
return ["hipblaslt", "rocblas"]
return []
def metadata_patterns_for_choice(choice: AssetChoice) -> list[str]:
patterns = ["BUILD_INFO.txt", "THIRD_PARTY_LICENSES.txt"]
if choice.install_kind.startswith("windows"):
patterns.append("LICENSE.txt")
else:
patterns.append("LICENSE")
return patterns
@contextmanager
def install_lock(lock_path: Path) -> Iterator[None]:
lock_path.parent.mkdir(parents = True, exist_ok = True)
if FileLock is None:
# Fallback: exclusive file creation as a simple lock.
# Write our PID so stale locks from crashed processes can be detected.
fd: int | None = None
deadline = time.monotonic() + INSTALL_LOCK_TIMEOUT_SECONDS
while True:
try:
fd = os.open(str(lock_path), os.O_CREAT | os.O_EXCL | os.O_RDWR)
try:
os.write(fd, f"{os.getpid()}\n".encode())
os.fsync(fd)
except Exception:
os.close(fd)
fd = None
lock_path.unlink(missing_ok = True)
raise
break
except FileExistsError:
# Check if the holder process is still alive
stale = False
try:
raw = lock_path.read_text().strip()
except FileNotFoundError:
# Lock vanished between our open attempt and read -- retry
continue
if not raw:
# File exists but PID not yet written -- another process
# just created it. Wait briefly for the write to land.
if time.monotonic() >= deadline:
raise BusyInstallConflict(
f"timed out after {INSTALL_LOCK_TIMEOUT_SECONDS}s waiting for concurrent install lock: {lock_path}"
)
time.sleep(0.1)
continue
try:
holder_pid = int(raw)
os.kill(holder_pid, 0) # signal 0 = existence check
except ValueError:
# PID unreadable (corrupted file)
stale = True
except ProcessLookupError:
# Process is dead
stale = True
except PermissionError:
# Process is alive but owned by another user -- not stale
pass
if stale:
lock_path.unlink(missing_ok = True)
continue
if time.monotonic() >= deadline:
raise BusyInstallConflict(
f"timed out after {INSTALL_LOCK_TIMEOUT_SECONDS}s waiting for concurrent install lock: {lock_path}"
)
time.sleep(0.5)
try:
yield
finally:
if fd is not None:
os.close(fd)
lock_path.unlink(missing_ok = True)
return
try:
with FileLock(lock_path, timeout = INSTALL_LOCK_TIMEOUT_SECONDS):
yield
except FileLockTimeout as exc:
raise BusyInstallConflict(
f"timed out after {INSTALL_LOCK_TIMEOUT_SECONDS}s waiting for concurrent install lock: {lock_path}"
) from exc
def install_lock_path(install_dir: Path) -> Path:
return install_dir.parent / f".{install_dir.name}.install.lock"
def install_staging_root(install_dir: Path) -> Path:
root = install_dir.parent / INSTALL_STAGING_ROOT_NAME
root.mkdir(parents = True, exist_ok = True)
return root
def prune_install_staging_root(install_dir: Path) -> None:
root = install_dir.parent / INSTALL_STAGING_ROOT_NAME
try:
root.rmdir()
except OSError:
pass
def create_install_staging_dir(install_dir: Path) -> Path:
staging_dir = Path(
tempfile.mkdtemp(
prefix = f"{install_dir.name}.staging-", dir = install_staging_root(install_dir)
)
)
log(f"created install staging dir {staging_dir}")
return staging_dir
def unique_install_side_path(install_dir: Path, label: str) -> Path:
root = install_staging_root(install_dir)
timestamp = time.strftime("%Y%m%d%H%M%S", time.gmtime())
prefix = f"{install_dir.name}.{label}-{timestamp}-{os.getpid()}"
candidate = root / prefix
counter = 0
while candidate.exists():
counter += 1
candidate = root / f"{prefix}-{counter}"
return candidate
def remove_tree(path: Path | None) -> None:
if path and path.exists():
shutil.rmtree(path, ignore_errors = True)
def remove_tree_logged(path: Path | None, label: str) -> None:
if not path:
return
if not path.exists():
log(f"{label} already absent at {path}")
return
log(f"removing {label} at {path}")
try:
shutil.rmtree(path)
except Exception as exc:
log(f"failed to remove {label} at {path}: {exc}")
raise
def cleanup_install_side_paths(
install_dir: Path,
*,
staging_dir: Path | None = None,
rollback_dir: Path | None = None,
failed_dir: Path | None = None,
active_dir: Path | None = None,
) -> None:
cleanup_failures: list[str] = []
for label, path in (
("failed install path", failed_dir),
("rollback path", rollback_dir),
("active install path", active_dir),
("staging dir", staging_dir),
):
if not path:
continue
try:
remove_tree_logged(path, label)
except Exception as exc:
cleanup_failures.append(f"{label} ({path}): {exc}")
prune_install_staging_root(install_dir)
if cleanup_failures:
raise RuntimeError("cleanup failed for " + "; ".join(cleanup_failures))
def confirm_install_tree(install_dir: Path, host: HostInfo) -> None:
if host.is_windows:
expected = [
install_dir / "build" / "bin" / "Release" / "llama-server.exe",
install_dir / "build" / "bin" / "Release" / "llama-quantize.exe",
install_dir / "convert_hf_to_gguf.py",
install_dir / "gguf-py",
]
else:
expected = [
install_dir / "llama-server",
install_dir / "llama-quantize",
install_dir / "build" / "bin" / "llama-server",
install_dir / "build" / "bin" / "llama-quantize",
install_dir / "convert_hf_to_gguf.py",
install_dir / "gguf-py",
]
expected.append(install_dir / "UNSLOTH_PREBUILT_INFO.json")
missing = [str(path) for path in expected if not path.exists()]
if missing:
raise RuntimeError("activated install was missing expected files: " + ", ".join(missing))
def activate_install_tree(staging_dir: Path, install_dir: Path, host: HostInfo) -> None:
rollback_dir: Path | None = None
failed_dir: Path | None = None
try:
if install_dir.exists():
rollback_dir = unique_install_side_path(install_dir, "rollback")
log(f"moving existing install to rollback path {rollback_dir}")
os.replace(install_dir, rollback_dir)
log(f"moved existing install to rollback path {rollback_dir.name}")
log(f"activating staged install {staging_dir} -> {install_dir}")
os.replace(staging_dir, install_dir)
log(f"activated staged install at {install_dir}")
log(f"confirming activated install tree at {install_dir}")
confirm_install_tree(install_dir, host)
log(f"activated install tree confirmed at {install_dir}")
except Exception as exc:
log(f"activation failed for staged install: {exc}")
try:
if install_dir.exists():
failed_dir = unique_install_side_path(install_dir, "failed")
log(f"moving failed active install to {failed_dir}")
os.replace(install_dir, failed_dir)
elif staging_dir.exists():
failed_dir = staging_dir
staging_dir = None
log(f"retaining failed staging tree at {failed_dir}")
if rollback_dir and rollback_dir.exists():
log(f"restoring rollback path {rollback_dir} -> {install_dir}")
os.replace(rollback_dir, install_dir)
log(f"restored previous install from rollback path {rollback_dir.name}")
if is_busy_lock_error(exc):
raise BusyInstallConflict(
"staged prebuilt validation passed but the existing install could not be replaced "
"because llama.cpp appears to still be in use; restored previous install "
f"({textwrap.shorten(str(exc), width = 200, placeholder = '...')})"
) from exc
raise PrebuiltFallback(
"staged prebuilt validation passed but activation failed; restored previous install "
f"({textwrap.shorten(str(exc), width = 200, placeholder = '...')})"
) from exc
except (BusyInstallConflict, PrebuiltFallback):
raise
except Exception as rollback_exc:
log(f"rollback after failed activation also failed: {rollback_exc}")
log(
"rollback restoration failed; cleaning staging, install, and rollback paths before source build fallback"
)
cleanup_error: Exception | None = None
try:
cleanup_install_side_paths(
install_dir,
staging_dir = staging_dir,
rollback_dir = rollback_dir,
failed_dir = failed_dir,
active_dir = install_dir,
)
except Exception as cleanup_exc:
cleanup_error = cleanup_exc
log(f"cleanup after rollback failure also failed: {cleanup_exc}")
details = textwrap.shorten(str(exc), width = 200, placeholder = "...")
if cleanup_error is not None:
raise PrebuiltFallback(
"staged prebuilt validation passed but activation and rollback failed; "
f"cleanup also reported errors ({details}; cleanup={cleanup_error})"
) from exc
raise PrebuiltFallback(
"staged prebuilt validation passed but activation and rollback failed; "
f"cleaned install state for fresh source build ({details})"
) from exc
else:
if rollback_dir:
try:
remove_tree_logged(rollback_dir, "rollback path")
except Exception as cleanup_exc:
log(
f"non-fatal: rollback cleanup failed after successful activation: {cleanup_exc}"
)
finally:
remove_tree(failed_dir)
remove_tree(staging_dir)
prune_install_staging_root(install_dir)
def install_from_archives(
choice: AssetChoice, host: HostInfo, install_dir: Path, work_dir: Path
) -> tuple[Path, Path]:
main_archive = work_dir / choice.name
log(f"downloading {choice.name} from {choice.source_label} release")
download_file_verified(
choice.url,
main_archive,
expected_sha256 = choice.expected_sha256,
label = f"prebuilt archive {choice.name}",
)
install_dir.mkdir(parents = True, exist_ok = True)
extract_dir = Path(tempfile.mkdtemp(prefix = "extract-", dir = work_dir))
runtime_extract_dir: Path | None = None
try:
extract_archive(main_archive, extract_dir)
# Download the paired runtime archive into its own temp dir to
# avoid copy_globs's ambiguous-layout guard on shared names
# like LICENSE.txt. Two passes of copy_globs land both archives
# in the same overlay dir. Fixes #5106.
if choice.runtime_url and choice.runtime_name:
runtime_archive = work_dir / choice.runtime_name
log(
f"downloading paired runtime archive {choice.runtime_name} "
f"from {choice.source_label} release"
)
download_file_verified(
choice.runtime_url,
runtime_archive,
expected_sha256 = choice.runtime_sha256,
label = f"prebuilt runtime archive {choice.runtime_name}",
)
runtime_extract_dir = Path(tempfile.mkdtemp(prefix = "extract-runtime-", dir = work_dir))
extract_archive(runtime_archive, runtime_extract_dir)
source_dir = extract_dir
overlay_dir = overlay_directory_for_choice(install_dir, choice, host)
copy_globs(source_dir, overlay_dir, runtime_patterns_for_choice(choice), required = True)
for _subdir in runtime_subdirs_for_choice(choice):
_src_subdir = source_dir / _subdir
if _src_subdir.is_dir():
shutil.copytree(_src_subdir, overlay_dir / _subdir, dirs_exist_ok = True)
if runtime_extract_dir is not None:
# The runtime archive only contributes the CUDA DLLs.
# Restrict the overlay to the cudart bundle's known
# filenames (cudart64_X.dll / cublas64_X.dll /
# cublasLt64_X.dll) rather than the broad ``*.exe`` /
# ``*.dll`` set from runtime_patterns_for_choice, so a
# malformed runtime archive can never overwrite
# llama-server.exe or other main-archive payload. The
# upstream cudart-llama-bin-win-cuda-X.Y-x64.zip currently
# ships exactly these three DLLs (verified against b9103
# cuda-12.4 and cuda-13.1 bundles).
copy_globs(
runtime_extract_dir,
overlay_dir,
paired_runtime_dll_patterns(choice),
required = False,
)
copy_globs(
source_dir,
install_dir,
metadata_patterns_for_choice(choice),
required = False,
)
finally:
remove_tree(extract_dir)
if runtime_extract_dir is not None:
remove_tree(runtime_extract_dir)
if host.is_windows:
exec_dir = install_dir / "build" / "bin" / "Release"
server_src = next(exec_dir.glob("llama-server.exe"), None)
quantize_src = next(exec_dir.glob("llama-quantize.exe"), None)
if server_src is None or quantize_src is None:
raise PrebuiltFallback("windows executables were not installed correctly")
return server_src, quantize_src
build_bin = install_dir / "build" / "bin"
source_server = build_bin / "llama-server"
source_quantize = build_bin / "llama-quantize"
if not source_server.exists() or not source_quantize.exists():
raise PrebuiltFallback("unix executables were not installed correctly into build/bin")
os.chmod(source_server, 0o755)
os.chmod(source_quantize, 0o755)
root_server = install_dir / "llama-server"
root_quantize = install_dir / "llama-quantize"
if source_server != root_server:
create_exec_entrypoint(root_server, source_server)
if source_quantize != root_quantize:
create_exec_entrypoint(root_quantize, source_quantize)
build_server = build_bin / "llama-server"
build_quantize = build_bin / "llama-quantize"
if source_server != build_server:
create_exec_entrypoint(build_server, source_server)
if source_quantize != build_quantize:
create_exec_entrypoint(build_quantize, source_quantize)
return source_server, source_quantize
def ensure_repo_shape(install_dir: Path) -> None:
required = [
install_dir / "CMakeLists.txt",
install_dir / "convert_hf_to_gguf.py",
install_dir / "gguf-py",
]
missing = [str(path.relative_to(install_dir)) for path in required if not path.exists()]
if missing:
raise PrebuiltFallback("hydrated llama.cpp source tree was missing: " + ", ".join(missing))
def validation_model_cache_path(install_dir: Path) -> Path:
cache_dir = install_dir.parent / VALIDATION_MODEL_CACHE_DIRNAME
cache_dir.mkdir(parents = True, exist_ok = True)
return cache_dir / VALIDATION_MODEL_CACHE_FILENAME
def validated_validation_model_bytes(data: bytes) -> bytes:
if not data:
raise RuntimeError(f"downloaded empty validation model from {TEST_MODEL_URL}")
digest = hashlib.sha256(data).hexdigest()
if digest != TEST_MODEL_SHA256:
raise RuntimeError(
f"validation model checksum mismatch: expected={TEST_MODEL_SHA256} actual={digest}"
)
return data
def _hf_resolve_url_parts(url: str) -> tuple[str, str, str] | None:
"""Parse a huggingface.co .../resolve/<rev>/<path> URL into
(repo_id, revision, filename); None if it is not such a URL."""
try:
parsed = urllib.parse.urlparse(url)
except Exception:
return None
if (parsed.netloc or "").lower() not in ("huggingface.co", "www.huggingface.co"):
return None
parts = parsed.path.strip("/").split("/")
# <owner>/<name>/resolve/<rev>/<path...>
if len(parts) >= 5 and parts[2] == "resolve":
return f"{parts[0]}/{parts[1]}", parts[3], "/".join(parts[4:])
return None
def _fetch_validation_model_bytes() -> bytes:
"""Fetch the tiny GGUF validation model. Prefer huggingface_hub (completes
TLS chains via AIA fetching that bare urllib can't on some Windows/proxy
setups); fall back to the direct URL when hf_hub is unavailable or fails."""
parts = _hf_resolve_url_parts(TEST_MODEL_URL)
if parts is not None:
repo_id, revision, filename = parts
try:
from huggingface_hub import hf_hub_download
local = hf_hub_download(repo_id = repo_id, filename = filename, revision = revision)
return validated_validation_model_bytes(Path(local).read_bytes())
except Exception as exc:
log(
f"huggingface_hub fetch of validation model failed ({exc}); "
"falling back to direct URL"
)
return validated_validation_model_bytes(
download_bytes(
TEST_MODEL_URL,
progress_label = f"Downloading {download_label_from_url(TEST_MODEL_URL)}",
)
)
def download_validation_model(path: Path, cache_path: Path | None = None) -> None:
try:
data: bytes | None = None
if cache_path and cache_path.exists():
try:
data = validated_validation_model_bytes(cache_path.read_bytes())
log(f"using cached tiny GGUF validation model from {cache_path}")
except Exception as exc:
log(f"cached tiny GGUF validation model was invalid; refreshing cache ({exc})")
data = None
if data is None:
log("downloading tiny GGUF validation model")
data = _fetch_validation_model_bytes()
if cache_path is not None:
atomic_write_bytes(cache_path, data)
atomic_write_bytes(path, data)
except Exception as exc:
raise PrebuiltFallback(f"validation model unavailable: {exc}") from exc
def free_local_port() -> int:
sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
sock.bind(("127.0.0.1", 0))
_, port = sock.getsockname()
sock.close()
return int(port)
def read_log_excerpt(log_path: Path, *, max_lines: int = 60) -> str:
try:
content = log_path.read_text(encoding = "utf-8", errors = "replace")
except FileNotFoundError:
return ""
return "\n".join(content.splitlines()[-max_lines:])
def is_retryable_server_bind_error(
exc: Exception | None,
output: str = "",
*,
exited_quickly: bool = False,
) -> bool:
haystack = output.lower()
bind_markers = (
"address already in use",
"only one usage of each socket address",
"failed to bind",
"bind failed",
"failed to listen",
"errno 98",
"errno 10048",
)
if any(marker in haystack for marker in bind_markers):
return True
if isinstance(exc, urllib.error.URLError):
reason = exc.reason
if exited_quickly and isinstance(reason, ConnectionRefusedError):
return True
if isinstance(reason, OSError) and reason.errno in {
98,
99,
111,
10048,
10049,
10061,
}:
return exited_quickly
if exited_quickly and isinstance(exc, ConnectionRefusedError):
return True
if isinstance(exc, OSError) and exc.errno in {98, 99, 111, 10048, 10049, 10061}:
return exited_quickly
return False
def dedupe_existing_dirs(paths: Iterable[str | Path]) -> list[str]:
unique: list[str] = []
seen: set[str] = set()
for raw in paths:
if not raw:
continue
path = Path(raw).expanduser()
if not path.is_dir():
continue
resolved = str(path.resolve())
if resolved in seen:
continue
seen.add(resolved)
unique.append(resolved)
return unique
def linux_missing_libraries(binary_path: Path, *, env: dict[str, str] | None = None) -> list[str]:
try:
result = run_capture(["ldd", str(binary_path)], timeout = 20, env = env)
except Exception:
return []
missing: list[str] = []
for line in (result.stdout + result.stderr).splitlines():
line = line.strip()
if "=> not found" not in line:
continue
library = line.split("=>", 1)[0].strip()
if library and library not in missing:
missing.append(library)
return missing
def python_runtime_dirs() -> list[str]:
candidates: list[Path] = []
search_roots = [Path(entry) for entry in sys.path if entry]
try:
search_roots.extend(Path(path) for path in site.getsitepackages())
except Exception:
pass
try:
user_site = site.getusersitepackages()
if user_site:
search_roots.append(Path(user_site))
except Exception:
pass
for root in search_roots:
if not root.is_dir():
continue
# ``nvidia/<pkg>/lib`` -- Linux convention; harmless on Windows
# where the directory simply does not exist on real wheels.
candidates.extend(root.glob("nvidia/*/lib"))
# ``nvidia/<pkg>/bin`` -- legacy modular Windows wheels
# (``nvidia-cuda-runtime-cu12``, ``nvidia-cublas-cu12``).
candidates.extend(root.glob("nvidia/*/bin"))
# ``nvidia/<pkg>/bin/x86_64`` and ``.../bin/x64`` -- current
# CUDA 13 Windows wheel layout (the unsuffixed
# ``nvidia-cuda-runtime`` 13.x and ``nvidia-cublas`` 13.x
# packages ship under ``nvidia/cu13/bin/x86_64/cudart64_13.dll``).
# Without these, Windows preflight CUDA detection misses cu13
# installs and falls back to the upstream cudart bundle path
# even when usable DLLs are already on disk (#5106). Kept in
# sync with the backend resolver
# ``llama_cpp.LlamaCppBackend._windows_pip_nvidia_dll_dirs``.
candidates.extend(root.glob("nvidia/*/bin/x86_64"))
candidates.extend(root.glob("nvidia/*/bin/x64"))
# ``nvidia/<pkg>/Library/bin`` -- conda-style wheel repacks.
candidates.extend(root.glob("nvidia/*/Library/bin"))
candidates.extend(root.glob("nvidia/*/Library/bin/x86_64"))
candidates.extend(root.glob("nvidia/*/Library/bin/x64"))
candidates.extend(root.glob("torch/lib"))
return dedupe_existing_dirs(candidates)
def ldconfig_runtime_dirs(required_libraries: Iterable[str]) -> list[str]:
try:
result = run_capture(["ldconfig", "-p"], timeout = 20)
except Exception:
return []
required = set(required_libraries)
candidates: list[str] = []
for line in result.stdout.splitlines():
if "=>" not in line:
continue
library, _, location = line.partition("=>")
library = library.strip().split()[0]
if required and library not in required:
continue
path = Path(location.strip()).parent
candidates.append(str(path))
return dedupe_existing_dirs(candidates)
def linux_runtime_dirs(binary_path: Path) -> list[str]:
missing = linux_missing_libraries(binary_path)
if not missing:
return []
return linux_runtime_dirs_for_required_libraries(missing)
# macOS prebuilt compatibility. Upstream macos prebuilts built on a newer macOS
# (e.g. minos=26, referencing Metal-4 symbols) fail dyld load on macOS 14/15. We
# read the host macOS version and each binary's minimum-OS so selection can skip
# a too-new prebuilt and walk back to the newest release that runs on this host.
# Mach-O constants (Apple mach-o/fat.h, mach-o/loader.h, mach/machine.h).
_MACHO_FAT_MAGICS = {0xCAFEBABE, 0xCAFEBABF} # universal binary (32/64-bit fat)
_LC_VERSION_MIN_MACOSX = 0x24 # legacy min-macOS load command
_LC_BUILD_VERSION = 0x32 # modern platform+minos+sdk load command
_MACHO_PLATFORM_MACOS = 1 # LC_BUILD_VERSION platform id for macOS (iOS=2, ...)
# CPU types (base | ABI64); used to pick the host slice in a fat binary.
_CPU_TYPE_X86_64 = 0x01000007
_CPU_TYPE_ARM64 = 0x0100000C
def parse_macos_version(value: str | None) -> tuple[int, int] | None:
"""Parse a macOS product version string into (major, minor).
Handles "14.7.1", "15.5", "26.0" and bare "26". Returns None when the
value is empty or cannot be parsed (callers then defer to runtime
validation rather than rejecting every prebuilt)."""
if not value:
return None
match = re.match(r"\s*(\d+)(?:\.(\d+))?", str(value))
if not match:
return None
return int(match.group(1)), int(match.group(2) or 0)
def host_supports_macos_minos(host: HostInfo, minos: tuple[int, int] | None) -> bool:
"""True if a prebuilt requiring `minos` can load on this host. Unknown host
version or unknown minos -> True: let runtime validation decide instead of
rejecting a binary we cannot reason about."""
if minos is None or host.macos_version is None:
return True
return host.macos_version >= minos
def _macho_slice_minos(data: bytes, offset: int) -> tuple[int, int] | None:
"""Minimum macOS for a single thin Mach-O at `offset`, via LC_BUILD_VERSION
(platform macOS) or the legacy LC_VERSION_MIN_MACOSX. None if absent."""
if offset + 4 > len(data):
return None
magic = struct.unpack_from(">I", data, offset)[0]
if magic in (0xFEEDFACE, 0xFEEDFACF):
endian, is64 = ">", magic == 0xFEEDFACF
elif magic in (0xCEFAEDFE, 0xCFFAEDFE):
endian, is64 = "<", magic == 0xCFFAEDFE
else:
return None
header_size = 32 if is64 else 28
if offset + header_size > len(data):
return None
ncmds = struct.unpack_from(endian + "I", data, offset + 16)[0]
cursor = offset + header_size
for _ in range(ncmds):
if cursor + 8 > len(data):
break
cmd, cmdsize = struct.unpack_from(endian + "II", data, cursor)
if cmdsize < 8:
break
if cmd == _LC_BUILD_VERSION and cursor + 16 <= len(data):
platform_id, minos = struct.unpack_from(endian + "II", data, cursor + 8)
if platform_id == _MACHO_PLATFORM_MACOS:
return (minos >> 16) & 0xFFFF, (minos >> 8) & 0xFF
elif cmd == _LC_VERSION_MIN_MACOSX and cursor + 12 <= len(data):
version = struct.unpack_from(endian + "I", data, cursor + 8)[0]
return (version >> 16) & 0xFFFF, (version >> 8) & 0xFF
cursor += cmdsize
return None
def macho_minimum_macos(path: Path, host: HostInfo | None = None) -> tuple[int, int] | None:
"""Minimum macOS (major, minor) a Mach-O binary or dylib requires.
Pure-Python so it works on consumer Macs without the Xcode command line
tools (otool/vtool). For universal binaries it prefers the host-arch slice,
else the highest minos found. Returns None for non-Mach-O files or when no
version load command is present."""
try:
data = path.read_bytes()
except Exception:
return None
if len(data) < 8:
return None
magic = struct.unpack_from(">I", data, 0)[0]
if magic in _MACHO_FAT_MAGICS:
is64 = magic == 0xCAFEBABF
nfat = struct.unpack_from(">I", data, 4)[0]
entry = 8
slices: list[tuple[int, tuple[int, int]]] = []
for _ in range(nfat):
if is64:
if entry + 32 > len(data):
break
cputype = struct.unpack_from(">I", data, entry)[0]
slice_offset = struct.unpack_from(">Q", data, entry + 8)[0]
entry += 32
else:
if entry + 20 > len(data):
break
cputype = struct.unpack_from(">I", data, entry)[0]
slice_offset = struct.unpack_from(">I", data, entry + 8)[0]
entry += 20
minos = _macho_slice_minos(data, slice_offset)
if minos is not None:
slices.append((cputype, minos))
if not slices:
return None
if host is not None:
want = (
_CPU_TYPE_ARM64 if host.is_arm64 else (_CPU_TYPE_X86_64 if host.is_x86_64 else None)
)
for cputype, minos in slices:
if cputype == want:
return minos
return max(minos for _cputype, minos in slices)
return _macho_slice_minos(data, 0)
def looks_like_macos_incompatibility(text: str) -> bool:
"""True when dyld output means a prebuilt needs a newer macOS than the host
(the runtime backstop for cases the static minos scan cannot read)."""
if not text:
return False
if "built for macOS" in text and "newer than running OS" in text:
return True
return "Symbol not found" in text and "MTLResidency" in text
def macos_binary_minos_issues(
binaries: Iterable[Path], install_dir: Path, host: HostInfo
) -> list[str]:
"""Issue strings for every installed Mach-O whose minimum macOS exceeds the
host. Scans the given executables plus every bundled .dylib next to them --
the dyld failure originates in libggml-metal.dylib, not the executable."""
candidates: list[Path] = list(binaries)
bin_dir = install_dir / "build" / "bin"
if bin_dir.is_dir():
candidates.extend(sorted(bin_dir.rglob("*.dylib")))
issues: list[str] = []
seen: set[Path] = set()
for path in candidates:
try:
resolved = path.resolve()
except Exception:
resolved = path
if resolved in seen or not path.exists():
continue
seen.add(resolved)
minos = macho_minimum_macos(path, host)
if minos is not None and not host_supports_macos_minos(host, minos):
issues.append(
f"{path.name}: built for macOS {minos[0]}.{minos[1]} > "
f"host macOS {host.macos_version[0]}.{host.macos_version[1]}"
)
return issues
def preflight_macos_installed_binaries(
binaries: Iterable[Path], install_dir: Path, host: HostInfo
) -> None:
"""Reject a macos prebuilt whose minimum-OS is newer than the host. The
upstream selector pins a loadable release up front, so here this is the
post-download backstop; the published/fork path also uses it to advance the
walk-back. No-op when the host macOS version is unknown (runtime validates)."""
if not host.is_macos or host.macos_version is None:
return
issues = macos_binary_minos_issues(binaries, install_dir, host)
if issues:
raise PrebuiltFallback(
"macos prebuilt requires a newer macOS than this host:\n" + "\n".join(issues)
)
def preflight_linux_installed_binaries(
binaries: Iterable[Path], install_dir: Path, host: HostInfo
) -> None:
if not host.is_linux:
return
issues: list[str] = []
for binary_path in binaries:
env = binary_env(binary_path, install_dir, host)
missing = linux_missing_libraries(binary_path, env = env)
if not missing:
continue
runtime_dirs = [part for part in env.get("LD_LIBRARY_PATH", "").split(os.pathsep) if part]
issues.append(
f"{binary_path.name}: missing={','.join(missing)} "
f"ld_library_path={','.join(runtime_dirs) if runtime_dirs else 'none'}"
)
if issues:
raise PrebuiltFallback("linux extracted binary preflight failed:\n" + "\n".join(issues))
def glob_paths(*patterns: str) -> list[str]:
matches: list[str] = []
for pattern in patterns:
if any(char in pattern for char in "*?[]"):
matches.extend(str(path) for path in Path("/").glob(pattern.lstrip("/")))
else:
matches.append(pattern)
return matches
def windows_runtime_dirs() -> list[str]:
candidates: list[str | Path] = []
env_dirs = os.environ.get("CUDA_RUNTIME_DLL_DIR", "")
if env_dirs:
candidates.extend(part for part in env_dirs.split(os.pathsep) if part)
path_dirs = os.environ.get("PATH", "")
if path_dirs:
candidates.extend(part for part in path_dirs.split(os.pathsep) if part)
cuda_roots: list[Path] = []
for name in ("CUDA_PATH", "CUDA_HOME", "CUDA_ROOT"):
value = os.environ.get(name)
if value:
cuda_roots.append(Path(value))
for root in cuda_roots:
candidates.extend([root / "bin", root / "lib" / "x64"])
program_files = os.environ.get("ProgramFiles", r"C:\Program Files")
toolkit_base = Path(program_files) / "NVIDIA GPU Computing Toolkit" / "CUDA"
if toolkit_base.is_dir():
candidates.extend(toolkit_base.glob("v*/bin"))
candidates.extend(toolkit_base.glob("v*/lib/x64"))
candidates.extend(Path(path) for path in python_runtime_dirs())
return dedupe_existing_dirs(candidates)
def windows_runtime_dirs_for_patterns(
required_patterns: Iterable[str], candidate_dirs: Iterable[str] | None = None
) -> list[str]:
directories = list(candidate_dirs) if candidate_dirs is not None else windows_runtime_dirs()
matching_dirs: list[str] = []
for pattern in required_patterns:
matched_dirs = [
directory for directory in directories if any(Path(directory).glob(pattern))
]
if not matched_dirs:
return []
for directory in matched_dirs:
if directory not in matching_dirs:
matching_dirs.append(directory)
return matching_dirs
def windows_runtime_dirs_for_runtime_line(runtime_line: str | None) -> list[str]:
if not runtime_line:
return []
patterns = windows_runtime_line_info().get(runtime_line)
if not patterns:
return []
return windows_runtime_dirs_for_patterns(patterns)
def binary_env(
binary_path: Path,
install_dir: Path,
host: HostInfo,
*,
runtime_line: str | None = None,
) -> dict[str, str]:
env = os.environ.copy()
if host.is_windows:
path_dirs = [
str(binary_path.parent),
*windows_runtime_dirs_for_runtime_line(runtime_line),
]
existing = [part for part in env.get("PATH", "").split(os.pathsep) if part]
env["PATH"] = os.pathsep.join(dedupe_existing_dirs([*path_dirs, *existing]))
elif host.is_linux:
ld_dirs = [
str(binary_path.parent),
str(install_dir),
*linux_runtime_dirs(binary_path),
]
existing = [part for part in env.get("LD_LIBRARY_PATH", "").split(os.pathsep) if part]
env["LD_LIBRARY_PATH"] = os.pathsep.join(dedupe_existing_dirs([*ld_dirs, *existing]))
elif host.is_macos:
dyld_dirs = [str(binary_path.parent), str(install_dir)]
existing = [part for part in env.get("DYLD_LIBRARY_PATH", "").split(os.pathsep) if part]
env["DYLD_LIBRARY_PATH"] = os.pathsep.join(dedupe_existing_dirs([*dyld_dirs, *existing]))
return env
def validate_quantize(
quantize_path: Path,
probe_path: Path,
quantized_path: Path,
install_dir: Path,
host: HostInfo,
*,
runtime_line: str | None = None,
) -> None:
command = [str(quantize_path), str(probe_path), str(quantized_path), "Q6_K", "2"]
result = subprocess.run(
command,
capture_output = True,
text = True,
timeout = 120,
env = binary_env(quantize_path, install_dir, host, runtime_line = runtime_line),
**windows_hidden_subprocess_kwargs(),
)
if result.returncode != 0 or not quantized_path.exists() or quantized_path.stat().st_size == 0:
combined = result.stdout + ("\n" + result.stderr if result.stderr else "")
# Backstop for prebuilts the static minos scan could not read: a dyld
# "built for macOS N" / missing Metal symbol failure means this binary
# needs a newer macOS than the host, so fall back to an older release.
prefix = (
"macos prebuilt requires a newer macOS than this host: "
if looks_like_macos_incompatibility(combined)
else ""
)
raise PrebuiltFallback(prefix + "llama-quantize validation failed:\n" + combined)
def validate_server(
server_path: Path,
probe_path: Path,
host: HostInfo,
install_dir: Path,
*,
runtime_line: str | None = None,
install_kind: str | None = None,
) -> None:
last_failure: PrebuiltFallback | None = None
for port_attempt in range(1, SERVER_PORT_BIND_ATTEMPTS + 1):
port = free_local_port()
command = [
str(server_path),
"-m",
str(probe_path),
"--host",
"127.0.0.1",
"--port",
str(port),
"-c",
"32",
"--parallel",
"1",
"--threads",
"1",
"--ubatch-size",
"32",
"--batch-size",
"32",
]
# Only enable GPU offload for assets that actually ship GPU code.
# Gating on `host.has_rocm` alone breaks the intentional CPU
# fallback on AMD Windows hosts without a HIP prebuilt: the CPU
# binary would be launched with `--n-gpu-layers 1` and fail
# validation. Use the resolved install_kind as the source of
# truth and fall back to host detection when the caller did not
# pass one (keeps backwards compatibility with older call sites).
_gpu_kinds = {
"linux-cuda",
"linux-rocm",
"windows-cuda",
"windows-hip",
"macos-arm64",
}
if install_kind is not None:
_enable_gpu_layers = install_kind in _gpu_kinds
else:
# Older call sites that don't pass install_kind: keep ROCm
# hosts in the GPU-validation path so an AMD-only Linux host
# is exercised against the actual hardware rather than the
# CPU fallback. NVIDIA and macOS-arm64 are already covered.
_enable_gpu_layers = (
host.has_usable_nvidia or host.has_rocm or (host.is_macos and host.is_arm64)
)
if _enable_gpu_layers:
command.extend(["--n-gpu-layers", "1"])
log_fd, log_name = tempfile.mkstemp(prefix = "llama-server-", suffix = ".log")
os.close(log_fd)
log_path = Path(log_name)
process: subprocess.Popen[str] | None = None
try:
with log_path.open("w", encoding = "utf-8", errors = "replace") as log_handle:
process = subprocess.Popen(
command,
stdout = log_handle,
stderr = subprocess.STDOUT,
text = True,
env = binary_env(server_path, install_dir, host, runtime_line = runtime_line),
**windows_hidden_subprocess_kwargs(),
)
deadline = time.time() + 60
startup_started = time.time()
response_body = ""
last_error: Exception | None = None
while time.time() < deadline:
if process.poll() is not None:
process.wait(timeout = 5)
log_handle.flush()
output = read_log_excerpt(log_path)
exited_quickly = (
time.time() - startup_started
) <= SERVER_BIND_RETRY_WINDOW_SECONDS
failure = PrebuiltFallback("llama-server exited during startup:\n" + output)
if (
port_attempt < SERVER_PORT_BIND_ATTEMPTS
and is_retryable_server_bind_error(
last_error,
output,
exited_quickly = exited_quickly,
)
):
log(
f"llama-server startup hit a port race on {port}; retrying with a fresh port "
f"({port_attempt}/{SERVER_PORT_BIND_ATTEMPTS})"
)
last_failure = failure
break
raise failure
payload = json.dumps({"prompt": "a", "n_predict": 1}).encode("utf-8")
request = urllib.request.Request(
f"http://127.0.0.1:{port}/completion",
data = payload,
headers = {"Content-Type": "application/json"},
)
try:
with urllib.request.urlopen(request, timeout = 5) as response:
status_code = response.status
response_body = response.read().decode("utf-8", "replace")
if status_code == 200:
return
last_error = RuntimeError(f"unexpected HTTP status {status_code}")
except urllib.error.HTTPError as exc:
response_body = exc.read().decode("utf-8", "replace")
last_error = exc
except Exception as exc:
last_error = exc
time.sleep(0.5)
else:
log_handle.flush()
output = read_log_excerpt(log_path)
raise PrebuiltFallback(
"llama-server completion validation timed out"
+ (f" ({last_error})" if last_error else "")
+ ":\n"
+ output
+ ("\n" + response_body if response_body else "")
)
finally:
if process is not None and process.poll() is None:
process.terminate()
try:
process.wait(timeout = 5)
except subprocess.TimeoutExpired:
process.kill()
process.wait(timeout = 5)
try:
log_path.unlink(missing_ok = True)
except Exception:
pass
if last_failure is not None:
raise last_failure
raise PrebuiltFallback("llama-server validation failed unexpectedly")
def collect_system_report(host: HostInfo, choice: AssetChoice | None, install_dir: Path) -> str:
lines = [
f"platform={host.system} machine={host.machine}",
f"driver_cuda_version={host.driver_cuda_version}",
f"compute_caps={','.join(host.compute_caps) if host.compute_caps else 'unknown'}",
f"cuda_visible_devices={host.visible_cuda_devices if host.visible_cuda_devices is not None else 'unset'}",
f"has_physical_nvidia={host.has_physical_nvidia}",
f"has_usable_nvidia={host.has_usable_nvidia}",
f"chosen_asset={(choice.name if choice else 'none')}",
f"asset_source={(choice.source_label if choice else 'none')}",
]
if host.is_linux and host.has_physical_nvidia:
runtime_lines, runtime_dirs = detected_linux_runtime_lines()
lines.append(
"linux_runtime_lines=" + (",".join(runtime_lines) if runtime_lines else "none")
)
for runtime_line in ("cuda13", "cuda12"):
lines.append(
f"linux_runtime_dirs_{runtime_line}="
+ (
",".join(runtime_dirs.get(runtime_line, []))
if runtime_dirs.get(runtime_line)
else "none"
)
)
if choice and choice.selection_log:
lines.append("selection_log:")
lines.extend(choice.selection_log)
if host.nvidia_smi:
try:
smi = run_capture([host.nvidia_smi], timeout = 20)
excerpt = "\n".join((smi.stdout + smi.stderr).splitlines()[:20])
lines.append("nvidia-smi:")
lines.append(excerpt)
except Exception as exc:
lines.append(f"nvidia-smi error: {exc}")
if host.is_linux:
server_binary = install_dir / "llama-server"
if server_binary.exists():
server_env = binary_env(server_binary, install_dir, host)
lines.append(
"linux_missing_libs="
+ (",".join(linux_missing_libraries(server_binary, env = server_env)) or "none")
)
lines.append(
"linux_runtime_dirs="
+ (
",".join(
[
part
for part in server_env.get("LD_LIBRARY_PATH", "").split(os.pathsep)
if part
]
)
or "none"
)
)
try:
ldd = run_capture(["ldd", str(server_binary)], timeout = 20, env = server_env)
lines.append("ldd llama-server:")
lines.append((ldd.stdout + ldd.stderr).strip())
except Exception as exc:
lines.append(f"ldd error: {exc}")
elif host.is_windows:
lines.append("windows_runtime_dirs=" + (",".join(windows_runtime_dirs()) or "none"))
runtime_lines, runtime_dirs = detected_windows_runtime_lines()
lines.append(
"windows_runtime_lines=" + (",".join(runtime_lines) if runtime_lines else "none")
)
for runtime_line in ("cuda13", "cuda12"):
lines.append(
f"windows_runtime_dirs_{runtime_line}="
+ (
",".join(runtime_dirs.get(runtime_line, []))
if runtime_dirs.get(runtime_line)
else "none"
)
)
elif host.is_macos:
server_binary = install_dir / "llama-server"
if server_binary.exists():
try:
otool = run_capture(["otool", "-L", str(server_binary)], timeout = 20)
lines.append("otool -L llama-server:")
lines.append((otool.stdout + otool.stderr).strip())
except Exception as exc:
lines.append(f"otool error: {exc}")
return "\n".join(lines)
def apply_approved_hashes(
attempts: Iterable[AssetChoice], checksums: ApprovedReleaseChecksums
) -> list[AssetChoice]:
def approved_hash_for_attempt(attempt: AssetChoice) -> ApprovedArtifactHash | None:
candidate_names = [attempt.name]
if (
isinstance(attempt.tag, str)
and attempt.tag
and attempt.tag != checksums.upstream_tag
and attempt.name.startswith("llama-")
):
legacy_prefix = f"llama-{attempt.tag}-"
compatibility_prefix = f"llama-{checksums.upstream_tag}-"
compatibility_name = (
attempt.name.replace(legacy_prefix, compatibility_prefix, 1)
if attempt.name.startswith(legacy_prefix)
else attempt.name
)
candidate_names.append(compatibility_name)
candidate_names.extend(
windows_cuda_asset_aliases(
attempt.name,
compatibility_tag = checksums.upstream_tag,
)
)
seen_names: set[str] = set()
for candidate_name in candidate_names:
if candidate_name in seen_names:
continue
seen_names.add(candidate_name)
approved = checksums.artifacts.get(candidate_name)
if approved is not None:
return approved
return None
approved_attempts: list[AssetChoice] = []
missing_assets: list[str] = []
for attempt in attempts:
# External prebuilts (e.g. lemonade-sdk) are not listed in the
# approved-hash manifest; they are explicitly documented as relying
# on functional validation only (llama-bench / smoke tests).
# Passing them through here lets the caller include both a lemonade
# attempt and a hash-approved upstream fallback in the same list
# without apply_approved_hashes discarding the lemonade entry.
if attempt.source_label == "lemonade":
approved_attempts.append(attempt)
continue
approved = approved_hash_for_attempt(attempt)
if approved is None:
missing_assets.append(attempt.name)
continue
attempt.expected_sha256 = approved.sha256
# Resolve the paired runtime archive's hash too. Drop the pair
# if the manifest does not list it -- never install an
# unverified archive.
if attempt.runtime_name and attempt.runtime_url:
runtime_approved = checksums.artifacts.get(attempt.runtime_name)
if runtime_approved is None:
attempt.runtime_name = None
attempt.runtime_url = None
attempt.runtime_sha256 = None
else:
attempt.runtime_sha256 = runtime_approved.sha256
approved_attempts.append(attempt)
if not approved_attempts:
missing_text = ", ".join(missing_assets) if missing_assets else "none"
raise PrebuiltFallback(
"approved checksum asset did not contain the selected prebuilt archive(s): "
f"{missing_text}"
)
return approved_attempts
def require_approved_source_hash(
checksums: ApprovedReleaseChecksums, llama_tag: str
) -> ApprovedArtifactHash:
source_asset_name = source_archive_logical_name(llama_tag)
approved_source = checksums.artifacts.get(source_asset_name)
if approved_source is None:
raise PrebuiltFallback(
f"approved checksum asset did not contain source archive {source_asset_name}"
)
return approved_source
def preferred_source_archive(
checksums: ApprovedReleaseChecksums, llama_tag: str
) -> tuple[str, str, ApprovedArtifactHash | None, bool]:
exact_source = exact_source_archive_hash(checksums)
exact_repo = repo_slug_from_source(checksums.source_repo) or repo_slug_from_source(
checksums.source_repo_url
)
if exact_source is not None and exact_repo and checksums.source_commit:
return (
exact_repo,
checksums.source_commit,
exact_source,
True,
)
legacy = checksums.artifacts.get(source_archive_logical_name(llama_tag))
return (
UPSTREAM_REPO,
llama_tag,
legacy,
False,
)
def selected_source_archive_metadata(
checksums: ApprovedReleaseChecksums, llama_tag: str
) -> tuple[str, str | None]:
_source_repo, _source_ref, source_archive, _exact_source = preferred_source_archive(
checksums, llama_tag
)
if source_archive is None:
return source_archive_logical_name(llama_tag), None
return source_archive.asset_name, source_archive.sha256
def resolve_install_attempts(
llama_tag: str, host: HostInfo, published_repo: str, published_release_tag: str
) -> tuple[str, str, list[AssetChoice], ApprovedReleaseChecksums]:
requested_tag, plans = resolve_install_release_plans(
llama_tag,
host,
published_repo,
published_release_tag,
)
if not plans:
raise PrebuiltFallback("no prebuilt release plans were available")
plan = plans[0]
return requested_tag, plan.llama_tag, plan.attempts, plan.approved_checksums
def resolve_install_release_plans(
llama_tag: str,
host: HostInfo,
published_repo: str,
published_release_tag: str,
*,
max_release_fallbacks: int = DEFAULT_MAX_PREBUILT_RELEASE_FALLBACKS,
) -> tuple[str, list[InstallReleasePlan]]:
requested_tag = normalized_requested_llama_tag(llama_tag)
allow_older_release_fallback = requested_tag == "latest" and not published_release_tag
release_limit = max(1, max_release_fallbacks)
# macOS may need to walk past a run of too-new prebuilts. Only when the host
# version is known; otherwise keep the default (cannot tell up front).
if host.is_macos and allow_older_release_fallback and host.macos_version is not None:
release_limit = max(release_limit, DEFAULT_MAX_MACOS_RELEASE_FALLBACKS)
plans: list[InstallReleasePlan] = []
last_error: PrebuiltFallback | None = None
for resolved_release in iter_resolved_published_releases(
llama_tag,
published_repo,
published_release_tag,
):
bundle = resolved_release.bundle
checksums = resolved_release.checksums
resolved_tag = bundle.upstream_tag
try:
if host.is_linux and host.is_x86_64 and host.has_usable_nvidia:
linux_cuda_selection = resolve_linux_cuda_choice(host, bundle)
attempts = apply_approved_hashes(linux_cuda_selection.attempts, checksums)
if not attempts:
raise PrebuiltFallback("no compatible Linux CUDA asset was found")
log_lines(linux_cuda_selection.selection_log)
else:
attempts = resolve_release_asset_choice(
host,
resolved_tag,
bundle,
checksums,
)
if not attempts:
raise PrebuiltFallback("no compatible prebuilt asset was found")
if attempts[0].selection_log:
log_lines(attempts[0].selection_log)
except PrebuiltFallback as exc:
last_error = exc
if not allow_older_release_fallback:
raise
log(
"published release skipped for install planning: "
f"{bundle.repo}@{bundle.release_tag} upstream_tag={resolved_tag} ({exc})"
)
continue
plans.append(
InstallReleasePlan(
requested_tag = requested_tag,
llama_tag = resolved_tag,
release_tag = bundle.release_tag,
attempts = attempts,
approved_checksums = checksums,
)
)
if not allow_older_release_fallback or len(plans) >= release_limit:
break
if plans:
return requested_tag, plans
if last_error is not None:
raise last_error
raise PrebuiltFallback("no installable published llama.cpp releases were found")
def write_prebuilt_metadata(
install_dir: Path,
*,
requested_tag: str,
llama_tag: str,
release_tag: str,
choice: AssetChoice,
approved_checksums: ApprovedReleaseChecksums,
prebuilt_fallback_used: bool,
) -> None:
source_asset_name, source_sha256 = selected_source_archive_metadata(
approved_checksums,
llama_tag,
)
# expected_install_fingerprint is the source of truth for what the
# fingerprint must contain. Calling it here -- instead of inlining a
# parallel payload -- prevents drift where new keys (e.g. the cudart
# pair fields added for #5106) are added to one side but not the
# other, which would cause every install to look stale.
fingerprint = expected_install_fingerprint(
llama_tag = llama_tag,
release_tag = release_tag,
choice = choice,
approved_checksums = approved_checksums,
)
if fingerprint is None:
raise PrebuiltFallback(f"cannot compute install fingerprint for {choice.name}")
metadata = {
"requested_tag": requested_tag,
"tag": llama_tag,
"release_tag": release_tag,
"published_repo": approved_checksums.repo,
"asset": choice.name,
"asset_sha256": choice.expected_sha256,
"source": choice.source_label,
# Binary-side repo/tag for non-upstream sources (e.g. lemonade).
# published_repo/release_tag always refer to the unsloth source tree;
# these capture where the actual binaries came from so the install
# summary can show both (e.g. "unslothai/llama.cpp@b9334 + lemonade@b1280").
"binary_repo": choice.repo,
"binary_release_tag": choice.tag,
"source_asset": source_asset_name,
"source_sha256": source_sha256,
"source_commit": approved_checksums.source_commit,
"source_commit_short": approved_checksums.source_commit_short,
"source_repo": approved_checksums.source_repo,
"source_repo_url": approved_checksums.source_repo_url,
"source_ref_kind": approved_checksums.source_ref_kind,
"requested_source_ref": approved_checksums.requested_source_ref,
"resolved_source_ref": approved_checksums.resolved_source_ref,
"bundle_profile": choice.bundle_profile,
"runtime_line": choice.runtime_line,
"coverage_class": choice.coverage_class,
"install_fingerprint": fingerprint,
"prebuilt_fallback_used": prebuilt_fallback_used,
"installed_at_utc": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
}
(install_dir / "UNSLOTH_PREBUILT_INFO.json").write_text(json.dumps(metadata, indent = 2) + "\n")
def expected_install_fingerprint(
*,
llama_tag: str,
release_tag: str,
choice: AssetChoice,
approved_checksums: ApprovedReleaseChecksums,
) -> str | None:
source_asset_name, source_sha256 = selected_source_archive_metadata(
approved_checksums,
llama_tag,
)
payload = {
"published_repo": approved_checksums.repo,
"release_tag": release_tag,
"upstream_tag": llama_tag,
"asset": choice.name,
"asset_sha256": choice.expected_sha256,
"source": choice.source_label,
"source_asset": source_asset_name,
"source_sha256": source_sha256,
"runtime_line": choice.runtime_line,
# Including the paired runtime archive (Windows cudart bundle)
# in the fingerprint is what forces existing #5106 installs to
# refresh: pre-PR installs hashed nothing in this slot, post-PR
# paired installs hash the cudart sha. Without these two keys
# an existing cudart-less install would keep matching the new
# choice and never re-overlay the cudart DLLs.
"runtime_asset": choice.runtime_name,
"runtime_sha256": choice.runtime_sha256,
"bundle_profile": choice.bundle_profile,
"coverage_class": choice.coverage_class,
}
return hashlib.sha256(
json.dumps(payload, sort_keys = True, separators = (",", ":")).encode("utf-8")
).hexdigest()
def load_prebuilt_metadata(install_dir: Path) -> dict[str, Any] | None:
metadata_path = install_dir / "UNSLOTH_PREBUILT_INFO.json"
if not metadata_path.is_file():
return None
try:
payload = json.loads(metadata_path.read_text(encoding = "utf-8"))
except Exception:
return None
if not isinstance(payload, dict):
return None
return payload
def runtime_payload_health_groups(choice: AssetChoice) -> list[list[str]]:
if choice.install_kind in {"linux-cpu", "linux-arm64"}:
return [
["libllama-common.so*"],
["libllama.so*"],
["libggml.so*"],
["libggml-base.so*"],
["libggml-cpu*.so*"],
["libmtmd.so*"],
]
if choice.install_kind == "linux-cuda":
return [
["libllama-common.so*"],
["libllama.so*"],
["libggml.so*"],
["libggml-base.so*"],
["libggml-cpu*.so*"],
["libmtmd.so*"],
["libggml-cuda.so*"],
]
if choice.install_kind in {"macos-arm64", "macos-x64"}:
return [
["libllama*.dylib"],
["libggml*.dylib"],
["libmtmd*.dylib"],
]
if choice.install_kind == "linux-rocm":
return [
["libllama-common.so*"],
["libllama.so*"],
["libggml.so*"],
["libggml-base.so*"],
["libggml-cpu*.so*"],
["libmtmd.so*"],
["libggml-hip.so*"],
]
if choice.install_kind in {"windows-cpu", "windows-arm64"}:
return [["llama.dll"]]
if choice.install_kind == "windows-cuda":
groups = [["llama.dll"], ["ggml-cuda.dll"]]
# When the cudart bundle was paired in (#5106) require all
# three of its DLLs alongside the main archive's payload.
# install_kind alone is not enough -- legacy installs without
# the cudart pair must still pass the health check on the
# no-pair fallback path, otherwise pair-less builds would loop
# on reinstall forever. The upstream cudart bundle ships
# cudart64_X.dll + cublas64_X.dll + cublasLt64_X.dll; missing
# any one of them still breaks GPU initialisation.
if choice.runtime_name:
groups.append(["cudart64_*.dll"])
groups.append(["cublas64_*.dll"])
groups.append(["cublasLt64_*.dll"])
return groups
if choice.install_kind == "windows-hip":
return [["llama.dll"], ["*hip*.dll"]]
return []
def install_runtime_dir(install_dir: Path, host: HostInfo) -> Path:
if host.is_windows:
return install_dir / "build" / "bin" / "Release"
return install_dir / "build" / "bin"
def runtime_payload_is_healthy(install_dir: Path, host: HostInfo, choice: AssetChoice) -> bool:
runtime_dir = install_runtime_dir(install_dir, host)
if not runtime_dir.exists():
return False
for pattern_group in runtime_payload_health_groups(choice):
matched = False
for pattern in pattern_group:
if any(runtime_dir.glob(pattern)):
matched = True
break
if not matched:
return False
return True
def existing_install_matches_choice(
install_dir: Path,
host: HostInfo,
*,
llama_tag: str,
release_tag: str,
choice: AssetChoice,
approved_checksums: ApprovedReleaseChecksums,
) -> bool:
if not install_dir.exists():
return False
metadata = load_prebuilt_metadata(install_dir)
if metadata is None:
return False
try:
confirm_install_tree(install_dir, host)
except Exception:
return False
if not runtime_payload_is_healthy(install_dir, host, choice):
return False
# Verify primary executables still exist (catches partial deletion)
runtime_dir = install_runtime_dir(install_dir, host)
ext = ".exe" if host.is_windows else ""
for binary in ("llama-server", "llama-quantize"):
if not (runtime_dir / f"{binary}{ext}").exists():
return False
if host.is_linux:
try:
preflight_linux_installed_binaries(
[runtime_dir / "llama-server", runtime_dir / "llama-quantize"],
install_dir,
host,
)
except Exception:
return False
expected_fingerprint = expected_install_fingerprint(
llama_tag = llama_tag,
release_tag = release_tag,
choice = choice,
approved_checksums = approved_checksums,
)
if not expected_fingerprint:
return False
recorded_fingerprint = metadata.get("install_fingerprint")
if not isinstance(recorded_fingerprint, str) or not recorded_fingerprint:
return False
if recorded_fingerprint != expected_fingerprint:
return False
expected_pairs = {
"release_tag": release_tag,
"published_repo": approved_checksums.repo,
"tag": llama_tag,
"asset": choice.name,
"asset_sha256": choice.expected_sha256,
"source": choice.source_label,
"runtime_line": choice.runtime_line,
"bundle_profile": choice.bundle_profile,
"coverage_class": choice.coverage_class,
}
for key, expected in expected_pairs.items():
if metadata.get(key) != expected:
return False
return True
def existing_install_matches_plan(
install_dir: Path, host: HostInfo, plan: InstallReleasePlan
) -> bool:
if not plan.attempts:
return False
return existing_install_matches_choice(
install_dir,
host,
llama_tag = plan.llama_tag,
release_tag = plan.release_tag,
choice = plan.attempts[0],
approved_checksums = plan.approved_checksums,
)
def validate_prebuilt_choice(
choice: AssetChoice,
host: HostInfo,
install_dir: Path,
work_dir: Path,
probe_path: Path,
*,
requested_tag: str,
llama_tag: str,
release_tag: str,
approved_checksums: ApprovedReleaseChecksums,
prebuilt_fallback_used: bool,
quantized_path: Path,
) -> tuple[Path, Path]:
source_repo, source_ref, source_archive, exact_source = preferred_source_archive(
approved_checksums, llama_tag
)
if exact_source:
log(f"hydrating exact llama.cpp source for {source_repo}@{source_ref} into {install_dir}")
else:
log(f"hydrating upstream llama.cpp source for {llama_tag} into {install_dir}")
hydrate_source_tree(
source_ref,
install_dir,
work_dir,
source_repo = source_repo,
expected_sha256 = source_archive.sha256 if source_archive is not None else None,
source_label = (
f"llama.cpp source tree for {source_repo}@{source_ref}"
if exact_source
else f"llama.cpp source tree for {llama_tag}"
),
exact_source = exact_source,
)
log(f"overlaying prebuilt bundle {choice.name} into {install_dir}")
server_path, quantize_path = install_from_archives(choice, host, install_dir, work_dir)
preflight_linux_installed_binaries((server_path, quantize_path), install_dir, host)
preflight_macos_installed_binaries((server_path, quantize_path), install_dir, host)
ensure_repo_shape(install_dir)
write_prebuilt_metadata(
install_dir,
requested_tag = requested_tag,
llama_tag = llama_tag,
release_tag = release_tag,
choice = choice,
approved_checksums = approved_checksums,
prebuilt_fallback_used = prebuilt_fallback_used,
)
validate_quantize(
quantize_path,
probe_path,
quantized_path,
install_dir,
host,
runtime_line = choice.runtime_line,
)
validate_server(
server_path,
probe_path,
host,
install_dir,
runtime_line = choice.runtime_line,
install_kind = choice.install_kind,
)
log(f"staged prebuilt validation succeeded for {choice.name}")
return server_path, quantize_path
def validate_prebuilt_attempts(
attempts: Iterable[AssetChoice],
host: HostInfo,
install_dir: Path,
work_dir: Path,
probe_path: Path,
*,
requested_tag: str,
llama_tag: str,
release_tag: str,
approved_checksums: ApprovedReleaseChecksums,
initial_fallback_used: bool = False,
existing_install_dir: Path | None = None,
) -> tuple[AssetChoice, Path, bool]:
attempt_list = list(attempts)
if not attempt_list:
raise PrebuiltFallback("no prebuilt bundle attempts were available")
tried_fallback = initial_fallback_used
for index, attempt in enumerate(attempt_list):
if index > 0:
tried_fallback = True
log(
"retrying CUDA prebuilt "
f"{attempt.name} install_kind={attempt.install_kind} "
f"runtime_line={attempt.runtime_line} coverage_class={attempt.coverage_class}"
)
if existing_install_dir is not None and existing_install_matches_choice(
existing_install_dir,
host,
llama_tag = llama_tag,
release_tag = release_tag,
choice = attempt,
approved_checksums = approved_checksums,
):
log(
"existing llama.cpp install already matches fallback candidate "
f"{attempt.name}; skipping reinstall"
)
raise ExistingInstallSatisfied(attempt, tried_fallback)
staging_dir = create_install_staging_dir(install_dir)
quantized_path = work_dir / f"stories260K-q4-{index}.gguf"
if quantized_path.exists():
quantized_path.unlink()
try:
validate_prebuilt_choice(
attempt,
host,
staging_dir,
work_dir,
probe_path,
requested_tag = requested_tag,
llama_tag = llama_tag,
release_tag = release_tag,
approved_checksums = approved_checksums,
prebuilt_fallback_used = tried_fallback,
quantized_path = quantized_path,
)
except Exception as exc:
remove_tree(staging_dir)
prune_install_staging_root(install_dir)
if isinstance(exc, PrebuiltFallback):
attempt_error = exc
else:
attempt_error = PrebuiltFallback(
f"candidate attempt failed before activation for {attempt.name}: {exc}"
)
if index == len(attempt_list) - 1:
raise attempt_error from exc
log(
"selected CUDA bundle failed before activation; trying next prebuilt fallback "
f"({textwrap.shorten(str(attempt_error), width = 200, placeholder = '...')})"
)
continue
return attempt, staging_dir, tried_fallback
raise PrebuiltFallback("no prebuilt bundle passed validation")
def install_prebuilt(
install_dir: Path,
llama_tag: str,
published_repo: str,
published_release_tag: str,
*,
simple_policy: bool = False,
override_has_rocm: bool = False,
override_rocm_gfx: str | None = None,
force_cpu: bool = False,
) -> None:
host = detect_host()
host = _apply_host_overrides(
host,
override_has_rocm = override_has_rocm,
override_rocm_gfx = override_rocm_gfx,
force_cpu = force_cpu,
)
choice: AssetChoice | None = None
try:
with install_lock(install_lock_path(install_dir)):
if install_dir.exists():
log(
f"existing llama.cpp install detected at {install_dir}; validating staged prebuilt update before replacement"
)
else:
log(
f"no existing llama.cpp install detected at {install_dir}; performing fresh prebuilt install"
)
if simple_policy:
requested_tag, release_plans = resolve_simple_install_release_plans(
llama_tag,
host,
published_repo,
published_release_tag,
)
else:
requested_tag, release_plans = resolve_install_release_plans(
llama_tag,
host,
published_repo,
published_release_tag,
)
if release_plans and existing_install_matches_plan(install_dir, host, release_plans[0]):
current = release_plans[0]
log(
"existing llama.cpp install already matches selected release "
f"{current.release_tag} upstream_tag={current.llama_tag}; skipping download and install"
)
return
with tempfile.TemporaryDirectory(prefix = "unsloth-llama-prebuilt-") as tmp:
work_dir = Path(tmp)
probe_path = work_dir / "stories260K.gguf"
download_validation_model(probe_path, validation_model_cache_path(install_dir))
release_count = len(release_plans)
for release_index, plan in enumerate(release_plans):
choice = plan.attempts[0]
if existing_install_matches_plan(install_dir, host, plan):
log(
"existing llama.cpp install already matches fallback release "
f"{plan.release_tag} upstream_tag={plan.llama_tag}; skipping reinstall"
)
return
log(
"selected "
f"{choice.name} ({choice.source_label}) from published release "
f"{plan.release_tag} for {host.system} {host.machine}"
)
try:
choice, selected_staging_dir, _ = validate_prebuilt_attempts(
plan.attempts,
host,
install_dir,
work_dir,
probe_path,
requested_tag = requested_tag,
llama_tag = plan.llama_tag,
release_tag = plan.release_tag,
approved_checksums = plan.approved_checksums,
initial_fallback_used = release_index > 0,
existing_install_dir = install_dir,
)
except ExistingInstallSatisfied:
return
except PrebuiltFallback as exc:
if release_index == release_count - 1:
raise
log(
"published release "
f"{plan.release_tag} upstream_tag={plan.llama_tag} failed; "
"trying the next older published prebuilt "
f"({textwrap.shorten(str(exc), width = 200, placeholder = '...')})"
)
continue
activate_install_tree(selected_staging_dir, install_dir, host)
try:
ensure_converter_scripts(install_dir, plan.llama_tag)
except Exception as exc:
log(
"converter script fetch failed after activation; install remains valid "
f"({textwrap.shorten(str(exc), width = 200, placeholder = '...')})"
)
return
except BusyInstallConflict as exc:
log("prebuilt install path is blocked by an in-use llama.cpp install")
log(f"prebuilt busy reason: {exc}")
raise SystemExit(EXIT_BUSY) from exc
except PrebuiltFallback as exc:
log("prebuilt install path failed; falling back to source build")
log(f"prebuilt fallback reason: {exc}")
report = collect_system_report(host, choice, install_dir)
print(report)
raise SystemExit(EXIT_FALLBACK) from exc
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description = "Install and validate a prebuilt llama.cpp bundle for Unsloth Studio."
)
parser.add_argument("--install-dir", help = "Target ~/.unsloth/llama.cpp directory")
parser.add_argument(
"--llama-tag",
default = DEFAULT_LLAMA_TAG,
help = (
"llama.cpp release tag. Defaults to the latest usable published Unsloth "
"release unless UNSLOTH_LLAMA_TAG overrides it."
),
)
parser.add_argument(
"--published-repo",
default = DEFAULT_PUBLISHED_REPO,
help = "Published bundle repository",
)
parser.add_argument(
"--published-release-tag",
default = DEFAULT_PUBLISHED_TAG,
help = (
"Published GitHub release tag to pin. By default, scan releases "
"until a usable published llama.cpp release bundle is found."
),
)
parser.add_argument(
"--simple-policy",
action = "store_true",
help = "Use the simplified platform-specific prebuilt selection policy.",
)
parser.add_argument(
"--has-rocm",
action = "store_true",
default = False,
help = (
"Assert that an AMD ROCm GPU is present. When set, skips the internal "
"hipinfo/amd-smi probe and forces has_rocm=True in the host profile. "
"Used by setup.ps1/setup.sh to forward their own ROCm detection result "
"so the HIP llama.cpp prebuilt is selected even when hipinfo is not on PATH."
),
)
parser.add_argument(
"--rocm-gfx",
default = os.environ.get("UNSLOTH_ROCM_GFX_ARCH"),
help = (
"Forward the AMD gfx target (e.g. gfx1151) that setup.ps1/setup.sh "
"resolved, so the lemonade HIP prebuilt is selected even when the "
"installer's own hipinfo/amd-smi probe cannot report it. Implies "
"--has-rocm. Defaults to the UNSLOTH_ROCM_GFX_ARCH environment variable."
),
)
parser.add_argument(
"--cpu-fallback",
action = "store_true",
default = False,
help = (
"Select the CPU prebuilt for this OS/arch even when a GPU is present. "
"setup.sh uses this as a last resort for arm64 Linux GPU hosts whose "
"source build failed (no arm64 CUDA prebuilt exists anywhere)."
),
)
resolve_group = parser.add_mutually_exclusive_group()
resolve_group.add_argument(
"--resolve-llama-tag",
nargs = "?",
const = "latest",
help = "Resolve a llama.cpp tag such as 'latest' to the logical upstream release tag.",
)
resolve_group.add_argument(
"--resolve-install-tag",
nargs = "?",
const = "latest",
help = (
"Resolve a llama.cpp tag such as 'latest' to the concrete upstream tag "
"selected by the current published-release policy."
),
)
resolve_group.add_argument(
"--resolve-source-build",
nargs = "?",
const = "latest",
help = ("Resolve the source-build fallback plan."),
)
parser.add_argument(
"--output-format",
choices = ("plain", "json"),
default = "plain",
help = "Resolver output format. Defaults to plain.",
)
return parser.parse_args()
def emit_resolver_output(payload: dict[str, Any], *, output_format: str) -> None:
if output_format == "json":
print(json.dumps(payload, sort_keys = True))
return
if "llama_tag" in payload:
print(payload["llama_tag"])
return
if {
"source_url",
"source_ref_kind",
"source_ref",
}.issubset(payload):
print(
"\t".join(
(
str(payload["source_url"]),
str(payload["source_ref_kind"]),
str(payload["source_ref"]),
)
)
)
return
print(json.dumps(payload, sort_keys = True))
def main() -> int:
args = parse_args()
if args.resolve_llama_tag is not None:
resolved = resolve_requested_llama_tag(
args.resolve_llama_tag,
args.published_repo,
args.published_release_tag or "",
)
emit_resolver_output(
{
"requested_tag": normalized_requested_llama_tag(args.resolve_llama_tag),
"llama_tag": resolved,
},
output_format = args.output_format,
)
return EXIT_SUCCESS
if args.resolve_install_tag is not None:
resolved = resolve_requested_install_tag(
args.resolve_install_tag,
args.published_release_tag or "",
args.published_repo,
)
emit_resolver_output(
{
"requested_tag": normalized_requested_llama_tag(args.resolve_install_tag),
"llama_tag": resolved,
},
output_format = args.output_format,
)
return EXIT_SUCCESS
if args.resolve_source_build is not None:
plan = resolve_source_build_plan(
args.resolve_source_build,
args.published_repo,
args.published_release_tag or "",
)
emit_resolver_output(
{
"requested_tag": normalized_requested_llama_tag(args.resolve_source_build),
"source_url": plan.source_url,
"source_ref_kind": plan.source_ref_kind,
"source_ref": plan.source_ref,
"compatibility_upstream_tag": plan.compatibility_upstream_tag,
},
output_format = args.output_format,
)
return EXIT_SUCCESS
if not args.install_dir:
raise SystemExit(
"install_llama_prebuilt.py: --install-dir is required unless --resolve-llama-tag, --resolve-install-tag, or --resolve-source-build is used"
)
install_prebuilt(
install_dir = Path(args.install_dir).expanduser().resolve(),
llama_tag = args.llama_tag,
published_repo = args.published_repo,
published_release_tag = args.published_release_tag or "",
simple_policy = args.simple_policy,
override_has_rocm = args.has_rocm,
override_rocm_gfx = args.rocm_gfx,
force_cpu = args.cpu_fallback,
)
return EXIT_SUCCESS
if __name__ == "__main__":
try:
raise SystemExit(main())
except SystemExit:
raise
except BusyInstallConflict as exc:
log(
f"fatal helper busy conflict: {textwrap.shorten(str(exc), width = 400, placeholder = '...')}"
)
raise SystemExit(EXIT_BUSY)
except PrebuiltFallback as exc:
# Expected when the published repo (e.g. ggml-org/llama.cpp) has no
# prebuilt manifest. Exit quietly with EXIT_FALLBACK so the caller
# falls back to source build without a noisy "fatal helper error".
log(textwrap.shorten(str(exc), width = 400, placeholder = "..."))
raise SystemExit(EXIT_FALLBACK)
except Exception as exc:
message = textwrap.shorten(str(exc), width = 400, placeholder = "...")
log(f"fatal helper error: {message}")
raise SystemExit(EXIT_ERROR)