* fix(studio): set HIP_VISIBLE_DEVICES in apply_gpu_ids for ROCm training workers Training workers are spawned via multiprocessing spawn before detect_hardware() runs, so IS_ROCM is still False. If the user never set HIP_VISIBLE_DEVICES in their shell, _inherits_rocm_visibility is also False, leaving the worker with only CUDA_VISIBLE_DEVICES set. On ROCm hosts the HIP runtime honors HIP_VISIBLE_DEVICES over CUDA_VISIBLE_DEVICES, so the worker saw the full device list and torch raised "no usable HIP accelerator" on some setups. Fall back to probing torch.version.hip (a build-time attribute, safe to read before GPU init) to detect ROCm when neither IS_ROCM nor inherited env vars are available. Mirrors the existing fix in llama_cpp.py for llama-server subprocess GPU pinning. Fixes https://github.com/unslothai/unsloth/issues/5180 * test: tighten apply_gpu_ids ROCm fallback assertions Replace loose OR chain with exact string matches, split into three focused tests, and add a guard check for the try/except wrapper. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: detect ROCm unified memory (Strix Halo / AMD iGPU) via torch fallback amd-smi on iGPUs with shared/unified memory (e.g. Radeon 8060S on Strix Halo) reports only the dedicated VRAM slice (~512 MB) in its metric output, so get_visible_gpu_utilization() was returning usable_gb ≈ 0.35 GB instead of the full GTT pool (~128 GB). torch.cuda.mem_get_info() already surfaces the correct unified-pool size. Add _reconcile_rocm_unified_memory(): after amd-smi returns a valid result on a ROCm device, cross-check each device's vram_total_gb against torch.cuda.mem_get_info(). When torch reports a larger total, replace the amd-smi VRAM fields in-place. No-op for discrete AMD GPUs where the two sources agree. Fixes: "Falling back to all visible GPUs -- model may not fit" on AMD iGPU machines even when 100+ GB of unified memory is available. * Apply unified-memory reconciliation in get_gpu_utilization too The visible-GPU path was already corrected for AMD iGPUs with unified memory (Strix Halo / Radeon 8060S), but get_gpu_utilization was still returning the raw 512 MB amd-smi VRAM slice. Studio's /api/train/hardware endpoint and the live GPU monitor read from this primary path, so users continued seeing the wrong total even after auto_select_gpu_ids picked the right device. Refactor to share the per-device correction: * _apply_unified_memory_correction(metrics, torch_info) -- the actual replacement logic, in-place on a single metrics dict. * _reconcile_rocm_unified_memory(...) -- multi-device, iterates utilization["devices"] (visible-GPU path). * _reconcile_primary_rocm_unified_memory(...) -- single flat metrics dict (primary-GPU path), uses parent_visible_spec to pick the primary index, falls back to ordinal 0 when no visibility env is set. get_gpu_utilization now calls the primary reconciler under IS_ROCM, so both endpoints surface the real unified-memory pool on iGPUs while leaving discrete AMD GPUs untouched (torch_total <= smi_total -> no replace). * Use 'is not None' and log debug on torch.version.hip probe failures Two small follow-ups to the apply_gpu_ids ROCm fallback: 1. Match detect_hardware()'s 'getattr(torch.version, "hip", None) is not None' form so the entire codebase has one canonical 'this torch was built with HIP' check. On every shipping torch wheel hip is either None or a non-empty version string, so the new form agrees with the old bool() form on every real install. 2. Log the probe failure at debug level instead of swallowing it silently. The broad 'except Exception' is intentional (we never want apply_gpu_ids to crash a worker over a probe), but the silent pass made it impossible to tell whether the fallback was firing or being skipped. * fix(studio): honour HIP_VISIBLE_DEVICES in _get_parent_visible_gpu_spec before IS_ROCM is set When a user has HIP_VISIBLE_DEVICES set in their shell (e.g. "1" to select GPU 1) but detect_hardware() has not yet run in the Studio parent process, IS_ROCM is still False. _get_parent_visible_gpu_spec() was gated on IS_ROCM so it fell through to CUDA_VISIBLE_DEVICES (unset), saw all physical GPUs, and auto-selected index 0. apply_gpu_ids then overwrote HIP_VISIBLE_DEVICES with "0", making the intended GPU invisible to ROCm torch in the worker, which triggered the "no usable HIP accelerator" error (issue #5180). Apply the same _inherits_rocm_visibility pattern already used in apply_gpu_ids: check for HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES in the environment regardless of IS_ROCM so the correct GPU index is preserved. * fix(install): harden AMD ROCm GPU detection for multi-GPU and env-filtered setups The previous rocminfo awk pattern could miss discrete GPUs on machines where HIP_VISIBLE_DEVICES/ROCR_VISIBLE_DEVICES is used to mask an integrated GPU — the env vars filter rocminfo output but may not propagate into the install script subprocess, causing detection to fail entirely. Two changes: - Tighten rocminfo pattern from /gfx[0-9]/ && !/gfx000/ to /gfx[1-9][0-9]/ — simpler and correctly excludes the CPU agent (gfx000) without a negative lookahead - Add sysfs KFD topology fallback: reads /sys/class/kfd/kfd/topology/nodes/*/gpu_id which is a kernel-level view unaffected by HIP_VISIBLE_DEVICES or ROCR_VISIBLE_DEVICES Fixes detection failure reported in Discord by Chains (gfx1201 + iGPU machine where env var exclusion of the iGPU caused rocminfo to return no usable device). * Fix KFD sysfs awk fallback to read properties file The fallback added by this PR reads /sys/class/kfd/kfd/topology/nodes/*/gpu_id files but matches the literal token 'gpu_id' against their content. Those files contain only a single decimal value (e.g. '0' for CPU agents, '50432' for GPU agents), so the regex never matches and 'found' stays 0, making the fallback a no-op on every host. The properties file in the same directory contains key/value lines like 'gpu_id 50432' which is what the existing awk pattern expects. Reproduced with a synthetic sysfs layout: against gpu_id files awk exits 1; against properties files awk exits 0 when any node reports gpu_id > 0. * fix(setup.ps1): detect AMD ROCm GPU on Windows, bring to parity with setup.sh setup.ps1 only checked nvidia-smi and fell straight to "gpu: none" on AMD machines. setup.sh already probed rocminfo/amd-smi/hipconfig/hipinfo. Add three-tier detection mirroring install_llama_prebuilt.py's detect_host(): 1. hipinfo: gcnArchName in output confirms a real HIP GPU (not just SDK) 2. amd-smi list: "GPU: <digit>" data rows as fallback 3. WMI Win32_VideoController: last resort -- detects AMD GPU even without HIP SDK, then guides user to install it rather than silently going CPU Also corrects the "none" message to mention AMD ROCm alongside NVIDIA so users with AMD hardware understand the requirement. Fixes: rohit-style install where Strix Halo (Radeon 8060S) showed "gpu: none" even with the HIP SDK present. * fix(install.ps1): detect AMD ROCm GPU on Windows, bring to parity with setup.ps1 install.ps1 had the same nvidia-smi-only GPU detection as setup.ps1 before the setup.ps1 fix. Applies the same three-tier AMD detection: 1. hipinfo: gcnArchName confirms real HIP GPU 2. amd-smi list: GPU data rows as fallback 3. WMI Win32_VideoController: detects AMD GPU without HIP SDK and guides user to install it Fixes: install.ps1 showing "gpu: none" while setup.ps1 correctly showed "AMD GPU detected" on the same machine (reported by rohit, RX 7600 XT). * fix(install.ps1): suppress 'No NVIDIA GPU detected' when AMD GPU is present * feat: add Windows AMD ROCm PyTorch wheel installation install_python_stack.py: - Add _ROCM_WINDOWS_WHEEL_BASE and _ROCM_WINDOWS_RELEASES constants pointing to AMD repo.radeon.com (ROCm 7.2 -> torch 2.9.1+rocm7.2.1) - Extend _ensure_rocm_torch() with a Windows branch: detects ROCm via _has_rocm_gpu() / _detect_rocm_version(), requires Python 3.12 (cp312 is the only ABI AMD publishes for Windows), installs the direct wheel URL from repo.radeon.com install.ps1: - Capture ROCmVersion during AMD detection via hipconfig --version / amd-smi version (needed for wheel URL selection) - After Get-TorchIndexUrl, add an AMD wheel override block: when HasROCm and Python 3.12 detected, set ROCmTorchWheelUrl to AMD wheel URL - Expand torch install branch to handle ROCmTorchWheelUrl with uv pip install --force-reinstall --no-cache-dir * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: also install torchvision and torchaudio from AMD Windows repo AMD publishes matching torchvision-0.24.1+rocm7.2.1 and torchaudio-2.9.1+rocm7.2.1 cp312 wheels at the same repo.radeon.com release folder. Install all three in both install.ps1 and install_python_stack.py Windows ROCm path. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * feat: add ROCm 7.1.1 Windows wheel mapping AMD uses a different version string for 7.1.1 wheels: 2.9.0+rocmsdk20251116 (date-tagged) instead of +rocm7.1.1. Adds the 7.1.1 release folder to both install.ps1 and install_python_stack.py so users with ROCm 7.1 get ROCm torch instead of falling back to CPU. * fix: install rocm_sdk_core and rocm_sdk_libraries_custom alongside torch The AMD Windows torch wheels declare rocm[libraries]==<ver> as a hard dependency. Without installing rocm_sdk_core and rocm_sdk_libraries_custom from the same AMD release folder, uv cannot resolve the dependency and fails with 'No solution found'. Include all 5 wheels in one install call. * fix: expand ROCm wheel array to scalars for Invoke-InstallCommand @array splatting inside a scriptblock only works when the native command is prefixed with '&'. Invoke-InstallCommand uses '& $Command' to run the block, so @ROCmAllWheelUrls was not being expanded. Extract to scalar variables $rw0-$rw4 which are captured correctly by the closure. * fix: use --no-deps for AMD Windows torch wheel install uv's resolver looks up rocm[libraries]==0.1.dev0 on PyPI during dependency resolution before downloading any wheels, and fails because the package doesn't exist on PyPI. --no-deps skips resolution entirely and installs all 5 AMD wheels directly. The GPU runtime dependency is satisfied by the HIP SDK, not a Python package. * fix: setup.ps1 and install_python_stack.py now install ROCm torch on Windows setup.ps1 was always setting CuTag='cpu' for non-NVIDIA hosts and installing cpu-only PyTorch, overwriting the ROCm torch installed by install.ps1. Adds the same AMD wheel selection logic (ROCm version detection, Python 3.12 check, 5-wheel install with --no-deps) to setup.ps1's torch install block. install_python_stack.py: remove IS_WINDOWS guard from _ensure_rocm_torch() call site so the Windows path in _ensure_rocm_torch() is reachable during 'unsloth studio update' as well. * fix: suppress manual-install warning when ROCm torch already present; fix progress counter - Gate the 'must be installed manually' warning on torch.version.hip being empty so it doesn't fire when our ROCm torch install succeeded - Update _TOTAL counter to include the 3 ROCm steps on Windows now that _ensure_rocm_torch() is called there (fixes 10/9 display) * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * feat: add rocm step display in setup.ps1; fix warning and progress counter - Add 'rocm' step after 'cuda' in setup.ps1 showing ROCm version or HIP SDK missing - Move ROCm version detection up to GPU detection block so it's available early - Suppress 'must be installed manually' warning when torch.version.hip is set - Fix _TOTAL counter to include ROCm steps on Windows (fixes 10/9 display) * fix: detect AMD SDK ROCm torch via __version__ when torch.version.hip is unset AMD's repo.radeon.com wheels (e.g. 2.9.0+rocmsdk20251116) do not set torch.version.hip, leaving it None. All three probes that relied solely on torch.version.hip now also check for 'rocm' in torch.__version__.lower(): - hardware.py detect_hardware(): IS_ROCM was never set, causing the studio to report 'Hardware detected: CPU' even after AMD wheels were installed and HIP DLLs were on PATH. - install_python_stack.py _ensure_rocm_torch(): skip-if-already-installed probe would always reinstall on subsequent runs. - install_python_stack.py Windows AMD warning: suppression check always failed, so the 'must be installed manually' note kept appearing after a successful AMD wheel install. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * perf: drop --no-cache-dir from AMD ROCm torch wheel installs uv caches downloaded wheels by default; passing --no-cache-dir forced a full redownload of the ~2 GB torch wheel on every install run. CUDA installs never had this flag -- AMD was the only path affected. * fix: use install-state flag instead of subprocess probe for AMD Windows warning Replace the subprocess torch probe in the post-install warning block with a module-level _rocm_windows_torch_installed flag set by _ensure_rocm_torch(). Subprocess re-import of torch is unnecessary and fragile -- the install function already knows whether it succeeded. * fix: hoist global declaration to top of _ensure_rocm_torch Python requires the global statement to appear before any assignment to the variable within a function. Moving it to the function top fixes the SyntaxError on line 354. * fix: pass AMD torch install status via env var to suppress false warning setup.ps1 now sets UNSLOTH_ROCM_TORCH_INSTALLED=1 after a successful AMD wheel install. install_python_stack.py reads this at the top of _ensure_rocm_torch() to skip both the subprocess probe and the warning -- no re-import of torch needed, and the warning message now correctly says 'could not be auto-installed' rather than 'must be installed manually'. * fix: register ROCm DLL directory before torch import on Windows Python 3.8+ ignores PATH for extension DLL loading on Windows; amdhip64.dll and other HIP runtime DLLs must be registered via os.add_dll_directory(). Without this, torch.cuda.is_available() always returns False on AMD ROCm Windows even when HIP_PATH is correctly set in system environment variables. Reads HIP_PATH / ROCM_PATH env vars first, then falls back to scanning common ROCm install roots (C:\Program Files\AMD\ROCm, F:\ROCm, C:\ROCm). * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: remove hardcoded non-standard ROCm paths from DLL directory scan Only use HIP_PATH/ROCM_PATH (set by AMD installer) and the standard C:\Program Files\AMD\ROCm\<version>\bin location. Custom drive paths like F:\ROCm are user-specific and should not be hardcoded. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: prevent torchao overrides step from overwriting AMD ROCm torch torchao==0.14.0 in overrides.txt declares torch as a dependency. Without --no-deps, uv resolves torch from PyPI and installs 2.11.0+cpu on top of the AMD ROCm wheels (2.9.0+rocmsdk20251116). This was the root cause of 'Hardware detected: CPU' -- the AMD wheels were installed but then immediately overwritten by the overrides step. When _rocm_windows_torch_installed is True, add --no-deps to the overrides pip_install call so torchao is installed without pulling in CPU torch. * fix: add rocm_sdk namespace tarball to Windows ROCm wheel installs torch/_rocm_init.py calls `import rocm_sdk` at startup, which requires the rocm namespace tarball (rocm-*.tar.gz) in addition to the SDK wheel packages. This tarball was missing from both install.ps1 and setup.ps1, causing ModuleNotFoundError on first torch import. - Add rocm-0.1.dev0.tar.gz to ROCm 7.1.1 install (provides rocm_sdk namespace) - Add rocm-7.2.1.tar.gz + rocm_sdk_devel to ROCm 7.2.1 install - Install tarball in a dedicated step before main SDK/torch wheels - Switch to @array splatting in install.ps1 scriptblock for dynamic wheel count - Remove --no-cache-dir from Python-side ROCm wheel install (prevents ~2GB redownload) * feat: enable ROCm 7.2 torch install + warn on gfx1151 with ROCm < 7.2 Chigoma333 (AMD Radeon 8060S / gfx1151, Strix Halo) confirmed that ROCm 7.1 segfaults when tensors are moved to GPU, but ROCm 7.2 + torch 2.11.0+rocm7.2 works fully including training. Changes: - Uncomment (7,2): "rocm7.2" in _ROCM_TORCH_INDEX (was blocked by <2.11.0) - Add _ROCM_TORCH_PKG_SPECS dict with per-tag version bounds: rocm7.2 → torch>=2.11.0,<2.12.0; all older tags → <2.11.0 - Add _detect_amd_gfx_codes() helper that parses rocminfo output - Warn on gfx1151/gfx1150 (Strix Halo) when ROCm < 7.2 is installed, pointing users at the known segfault and recommending upgrade - install.sh get_torch_index_url(): enable rocm7.2 case (previously capped to rocm7.1), cap unknown future tags to rocm7.2 - install.sh: override TORCH_CONSTRAINT to >=2.11.0,<2.12.0 when rocm7.2 index is selected, so pip can actually resolve torch 2.11.0 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: prefer Python 3.12 for AMD ROCm users when 3.13 is also installed After GPU detection, if ROCm HIP SDK is found and the selected Python is not 3.12, run a second pass to locate a 3.12 install via py.exe and PATH (catches uv-managed installs). Switch $DetectedPython to 3.12 so the venv is created with a compatible interpreter for the cp312-only AMD Windows torch wheels. NVIDIA and Intel GPU paths are unaffected -- the re-detection block only runs when $HasROCm is true. Fixes: #5301 * fix: also check uv-managed Python 3.12 for AMD ROCm #5301 * fix: hide amd-smi console popups on Windows, guard torch.distributed.is_initialized for ROCm #5301 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: suppress remaining console popups on Windows, patch torch.distributed.is_initialized for ROCm #5301 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: stub all missing torch.distributed attrs for ROCm Windows wheel #5301 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: inject torch.distributed stub when C backend missing in ROCm Windows wheel #5301 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix(rocm/windows): pre-stub torch._C._distributed_c10d + raise amd-smi timeout Two fixes for Windows ROCm regressions reported by electroglyph on #5301: 1. worker.py — torch.distributed stub now fires unconditionally on Windows The previous stub only injected sys.modules in the except branch, meaning it was silently skipped when `import torch.distributed` happened to succeed (the C backend is lazily resolved). The crash then hit later when transformers/trl triggered the lazy load. Fix: on win32 we pre-populate sys.modules['torch._C._distributed_c10d'] AND set the attribute on the torch._C extension module *before* attempting the import, covering both the early-ImportError and lazy-load failure modes. 2. amd.py — increase amd-smi timeout from 5 s to 30 s on Windows (10 s Linux) amd-smi on Windows must cold-init the ROCm runtime on first invocation; 5 s was consistently too short, producing repeated 'Command timed out' warnings in the server log. 30 s gives enough headroom without blocking indefinitely on broken installs. 3. install.ps1 — widen Python 3.12 enforcement to ROCmGpuLabel (WMI-only path) Users whose HIP SDK is not on PATH were detected via WMI but not switched to Python 3.12 before the install started, causing a second pass. Guard now fires on (HasROCm -or ROCmGpuLabel). * fix(rocm): guard c10d stub, fix TorchIndexFamily for 7.1, clean dead code + comments - worker.py: wrap c10d stub injection in `if _c10d_key not in sys.modules` so Windows NVIDIA users with a real torch.distributed are never affected - install.ps1: fix Get-TauriTorchIndexFamily receiving hardcoded "rocm7.2" even when ROCm 7.1 wheels are installed; now branches on $ROCmVersion - main.py: remove dead `import ctypes as _ctypes` (ctypes is never called) - hardware.py, install_python_stack.py, worker.py, install.ps1: shorten verbose multi-line comment blocks throughout - tests: update 4 stale assertions that expected rocm7.2 to be absent/capped * fix(tests): match windows AMD warning assertion to actual source string * chore: trim verbose comment blocks across all ROCm-related files * fix: guard reconcile call against None numeric_ids; add torchvision lower bounds * fix(install.ps1): recreate venv with Python 3.12 after ROCm switch Venv was created with 3.13 before GPU detection ran; switching $DetectedPython to 3.12 had no effect since $VenvPython still pointed to the 3.13 interpreter inside the already-created venv. * ux: detect AMD GPU before Python selection to avoid double venv creation - Early hipinfo + WMI probe runs before Find-CompatiblePython so Python 3.12 is selected upfront when AMD is detected; venv is now created exactly once instead of 3.13 then immediately 3.12. - Post-venv recreation block replaced with a simple warning for the rare case where AMD was missed by the early probe. - setup.ps1: show venv's actual Python version (e.g. 3.12) instead of the system Python found by the pre-activation search (was showing 3.13). * fix(rocm/win): auto-stub all _distributed_c10d symbols via PEP-562 __getattr__ The bare ModuleType stub caused ImportError when torch._dynamo was imported (triggered by trainer.py accessing torch._dynamo.config at load time). torch._dynamo pulls in torch.distributed.fsdp._flat_param which does: from torch._C._distributed_c10d import FakeProcessGroup and potentially other symbols. Adding module __getattr__ auto-creates a stub class for any missing symbol so all such imports succeed without enumerating every individual symbol. Applied to both the primary stub and the fallback stub in the except branch. * chore: trim c10d stub comment * fix(rocm/win): auto-stub missing torch.distributed attrs (Store, ProcessGroup, …) * fix(rocm/win): pre-stub fsdp submodules in sys.modules; fix __getattr__ subpackage clash * feat(rocm/win): arch-aware wheel selector always picks newest ROCm release Replace HIP-SDK-version-gated wheel selection with GPU arch-based logic. Select-ROCmWheelRelease (PS) and _select_windows_rocm_release (Python) map gcnArchName → minimum ROCm version, then pick the newest available release that satisfies it (currently always rocm-rel-7.2.1 for any supported GPU). Wheels bundle their own ROCm runtime so the installed HIP SDK 7.1 does not prevent using 7.2.1 wheels on gfx1200 (RX 9060 XT) and similar RDNA 4 GPUs. Also installs the bitsandbytes Windows ROCm continuous-release wheel and sets BNB_ROCM_VERSION=72 in worker.py before ML imports so bnb loads the libbitsandbytes_rocm72.dll that ships in that wheel. * fix(rocm/win): stub class metaclass for ProcessGroup.BackendType; amd-smi circuit breaker torchao.float8.inference accesses ProcessGroup.BackendType as a class-level attribute. Plain type() stubs have no __getattr__ on the metaclass so this raises AttributeError. Introduce _StubClassMeta whose __getattr__ returns child stub classes, fixing the torchao import chain. Add an amd-smi circuit breaker in amd.py: after 3 consecutive failures the module stops spawning the process, eliminating the repeated Windows UAC / DiskPart elevation prompts caused by polling a non-functional amd-smi. Also guard BNB_ROCM_VERSION=72 behind a DLL existence check so bitsandbytes fails with its own detection message rather than a harder "DLL not found" when the Windows ROCm bnb wheel is not yet installed. * fix: stub __members__ so torchao float8 enum check doesn't crash on ROCm Windows torchao.float8.inference accesses ProcessGroup.BackendType.__members__ expecting a Python Enum registry dict. _StubClassMeta.__getattr__ was blocking all dunder attributes, causing AttributeError. Return {} for __members__ specifically so the isinstance/iteration checks pass cleanly. * fix: stub distributed tensor/functional_collectives to prevent missing C++ op crash on ROCm Windows torch._dynamo.trace_rules eagerly loads torch.distributed.tensor at import time, which pulls in _functional_collectives.py. That file registers Meta kernels for _c10d_functional C++ ops, but those ops are only registered by torch._C._distributed_c10d — a C extension absent from ROCm Windows wheels. Pre-stubbing the affected modules in sys.modules prevents the real import chain from running and avoids the "operator does not exist" crash. * fix: give mod stubs __path__ and pre-stub _tensor to fix 'not a package' import error _make_mod_stub now sets __path__=[] so Python treats stub modules as packages. Without it, any import of a submodule raises "is not a package". Also pre-stub torch.distributed._tensor and its submodules so that _tensor/__init__.py (which re-exports from torch.distributed.tensor) never runs and torchao's `from torch.distributed._tensor import DTensor` gets a harmless stub instead of crashing. * fix: stub torch.ops._c10d_functional namespace with hashable op sentinels torchao.dtypes.nf4tensor uses _c10d_functional ops as dict keys at import time (all_gather_into_tensor.default, wait_tensor.default) and torch.ops.c10d.scatter_.default. None of these ops are registered on ROCm Windows because torch._C._distributed_c10d (the C extension) doesn't ship. Replace the whole _c10d_functional namespace with a custom stub whose ops return hashable .default objects, so dict-key construction doesn't crash. Also inject a scatter_ stub into torch.ops.c10d if it's missing. * fix: stub entire torchao package on ROCm Windows instead of individual ops torchao is not supported on ROCm Windows and its import chain transitively requires torch._C._distributed_c10d (absent from the ROCm Windows wheel). Rather than stub each missing op one by one, stub the whole torchao package upfront. Unsloth uses bitsandbytes for quantization, not torchao, so this has no functional impact. transformers gracefully handles an importable-but- empty torchao by disabling TorchAoHfQuantizer. * fix: set __spec__ on mod stubs so importlib.util.find_spec doesn't raise Manually-injected sys.modules entries have __spec__=None by default. importlib.util.find_spec() raises ValueError when it finds a module in sys.modules with __spec__=None (transformers.utils.import_utils hits this when checking if torchao is available). Give every stub a minimal ModuleSpec(name, loader=None, is_package=True) to satisfy find_spec. * fix: add meta path finder to auto-stub subpackages of stub modules `import torchao.prototype` goes through the import machinery, not __getattr__, so an empty __path__ means ModuleNotFoundError. Rather than list every submodule explicitly, register a MetaPathFinder that intercepts any import whose parent is one of our stubs (detected by loader=None in the parent's ModuleSpec). Real installed packages always have a SourceFileLoader so they are never intercepted. Also register child stubs in sys.modules from __getattr__ as a belt-and-suspenders measure. * fix: use _unsloth_stub sentinel instead of loader=None for stub detection The import machinery overwrites module.__spec__ with the spec returned by find_spec (which has loader=_StubSubpackageLoader, not None), so the loader=None check broke for second-level subpackages. Switch to a custom _unsloth_stub object identity sentinel set directly on each stub module -- it survives __spec__ being replaced and correctly identifies stubs at any depth (torchao.prototype.safetensors, etc.). * refactor(rocm/win): switch to repo.amd.com arch-aware index, remove stubs AMD recommends repo.amd.com/rocm/whl/{arch}/ as the Windows ROCm wheel source. These wheels bundle their own ROCm runtime, support all Python versions (not just cp312), and include the full torch._C extension set (including _distributed_c10d) that the old repo.radeon.com wheel omitted. Changes: - install.ps1: remove Select-ROCmWheelRelease + hardcoded cp312 wheel URLs; remove Python 3.12 forced-preference logic; install via --index-url repo.amd.com/rocm/whl/{arch-family}/ - studio/setup.ps1: same -- remove Select-ROCmWheelRelease, switch to repo.amd.com arch-aware index URL - studio/install_python_stack.py: replace _ROCM_WINDOWS_RELEASES / _select_windows_rocm_release with _windows_rocm_index_url() using the _GFX_TO_AMD_INDEX_ARCH map; drop Python 3.12 restriction - studio/backend/core/training/worker.py: remove all stub machinery (_make_mod_stub, _StubSubpackageFinder, _StubSubpackageLoader, _StubClassMeta, torchao/fsdp/dtensor stubs, _c10d_functional ops stubs, BNB DLL detection) -- no longer needed with new wheel source * fix(rocm/win): restore _distributed_c10d + torchao stubs; fix BNB install repo.amd.com torch wheels also omit torch._C._distributed_c10d on Windows (RCCL is not shipped on Windows). torch/distributed/__init__.py imports from it unconditionally at module level, so the stub must land in sys.modules before any torch.distributed import. torchao (pulled in by transformers.quantizers) walks torchao.float8.distributed_utils -> torch.distributed._functional_collectives -> distributed_c10d at import time. Stubbing torchao up-front short-circuits that chain. worker.py: - Restore _make_mod_stub / _StubSubpackageFinder / _StubSubpackageLoader - Restore _StubClassMeta for ProcessGroup.BackendType attribute access - Restore _distributed_c10d stub with __getattr__ (Windows only) - Restore torchao stubs (5 modules, Windows only) install_python_stack.py: - BNB AMD wheel install was inside the early-return branch that fires when torch is already a ROCm build (installed by install.ps1). Move BNB install outside that branch so it always runs on Windows ROCm — the PyPI bitsandbytes has only CUDA DLLs and fails to load on ROCm. * worker: remove _distributed_c10d stub; stub only torchao The installed torch/distributed/__init__.py from repo.amd.com (torch==2.10.0+rocm7.12.0) is now properly guarded with `if is_available():`, so `import torch.distributed` alone is safe. The crash only comes via torchao's import chain: torchao.float8.distributed_utils → torch.distributed._functional_collectives (unguarded import) → torch.distributed.distributed_c10d → torch._C._distributed_c10d ← absent on Windows ROCm Stubbing torchao short-circuits the chain entirely. No need to stub _distributed_c10d. Remove _StubClassMeta and the _c10d stub block; keep only _make_mod_stub + _StubSubpackageFinder + torchao seeds. * fix: BNB AMD wheel skipped + torch.compile segfault on Windows ROCm install_python_stack.py: the UNSLOTH_ROCM_TORCH_INSTALLED=1 early-return path (set by setup.ps1 when it installed torch itself) returned before ever reaching the AMD BNB prerelease wheel install. The PyPI bitsandbytes==0.49.x ships only CUDA DLLs, so loading it on ROCm fails with "libbitsandbytes_rocm72.dll not found". Now installs the AMD Windows BNB wheel before returning on that path too. worker.py: torch._grouped_mm crashes on gfx1200 (null HIP kernel pointer, 0xC0000005) when torch.compile's JitDecomp system dispatches it during the first forward pass. Detect Windows ROCm via torch.version.hip (already in sys.modules from section 1e) and set TORCHDYNAMO_DISABLE=1 to bypass the broken kernel dispatch. * fix: BNB AMD wheel install fails uv wheel filename check The bitsandbytes continuous-release wheel is intentionally mismatched: filename encodes 1.33.7.preview (= 1.33.7rc0 in PEP 440) but wheel metadata reports 0.50.0.dev0. uv rejects this by default. Introduce _install_bnb_windows_rocm() helper that sets UV_SKIP_WHEEL_FILENAME_CHECK=1 only for this specific install, then restores the previous env value. Both BNB install call sites (the UNSLOTH_ROCM_TORCH_INSTALLED early-return path and the normal Windows ROCm path) now use this helper. * worker: patch _grouped_mm CUDA dispatch on Windows ROCm (gfx1200 null kernel) TORCHDYNAMO_DISABLE=1 stopped the compiler frontend but not the autograd JitDecomp system, which also dispatches _grouped_mm and hits the same null HIP kernel crash (0xC0000005). Verified that torch.library.Library("aten","IMPL").impl("_grouped_mm", fn, "CUDA") successfully overrides the broken HIP kernel with a Python mm fallback on torch==2.10.0+rocm7.12.0. Schema: _grouped_mm(Tensor self, Tensor mat2, Tensor? offs=None, Tensor? bias=None, ScalarType? out_dtype=None) -> Tensor The fallback handles both the simple case (offs=None → torch.mm) and the grouped case (offs provided → split self by offsets, multiply each group against the corresponding slice of mat2, then cat results). Keep _WINDOWS_ROCM_GROUPED_MM_LIB alive at function scope to prevent the C++ dispatch registration from being freed by GC. * worker: fix torchao stub — return stub classes not modules for isinstance() peft/tuners/lora/torchao.py does: from torchao.dtypes import AffineQuantizedTensor, LinearActivationQuantizedTensor isinstance(weight, (AffineQuantizedTensor, LinearActivationQuantizedTensor)) The stub __getattr__ was returning stub modules, which isinstance() rejects with "arg 2 must be a type, a tuple of types, or a union". Add _StubTypeMeta metaclass whose __instancecheck__ always returns False, and _make_stub_type() to create stub classes via it. Change _make_mod_stub __getattr__ to return stub classes instead of stub modules for leaf attribute access, so isinstance() gets a valid type and returns False. _StubSubpackageFinder still handles import-style subpackage creation (those still need module objects in sys.modules); __getattr__ only fires for from-import or direct attribute access, which are the isinstance paths. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * tests: add coverage for Windows ROCm install paths and worker patches Add conftest.py to fix pre-existing sys.path issue that prevented test_rocm_support.py from running at all (install_python_stack.py imports from backend.utils.wheel_utils which needs studio/ on sys.path). New test classes cover everything added in this session: - TestWindowsRocmIndexUrl: arch → AMD pip index URL mapping (gfx120X-all, gfx1151, gfx1150, gfx110X-all, unknown → None, trailing slash) - TestDetectWindowsGfxArch: hipinfo output parsing, missing/timeout/bad returncode/no-gcnArchName paths - TestInstallBnbWindowsRocm: UV_SKIP_WHEEL_FILENAME_CHECK set+restored, env restored on exception, no-op when URL missing - TestRocmTorchInstalledEnvVar: UNSLOTH_ROCM_TORCH_INSTALLED=1 skips pip_install, calls _install_bnb_windows_rocm, sets flag - TestWorkerWindowsRocmPatches: _grouped_mm CUDA dispatch override, offs/grouped variant handling, GC-prevention sentinel, _StubTypeMeta __instancecheck__, _StubSubpackageFinder registration, torchao key submodule pre-stubbing, TORCHDYNAMO_DISABLE guard - TestRocmTorchPkgSpecs: rocm7.2 torch 2.11.x spec, default <2.11 cap, 3-tuple shape, _GFX_TO_AMD_INDEX_ARCH RDNA4/3.5/3 coverage * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * tests: fix encoding, IS_WINDOWS patching, and wrong assertion - Add encoding="utf-8" to all read_text() calls (54 occurrences) so tests pass on Windows where the default codec is cp1252 and source files contain UTF-8 emoji (e.g. ⚠️ in install_python_stack.py) - Add @patch.object(stack_mod, "IS_WINDOWS", False) to Linux-path TestEnsureRocmTorch tests so they reach the Linux code path when run on a Windows machine instead of short-circuiting into the Windows branch - Fix test_grouped_mm_patch_guarded_by_windows_and_hip_check: the source uses getattr(_torch_for_rocm, "version", None) not torch.version, so check for '"version"' and '"hip"' substrings instead 137 passed, 2 skipped * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: pin BNB_ROCM_VERSION=72 for torch==2.11.0+rocm7.13.0 compatibility AMD's pip index now ships torch==2.11.0+rocm7.13.0 (ROCm 7.13). bitsandbytes auto-detects HIP 7.13 from torch.version.hip and looks for libbitsandbytes_rocm713.dll, which the AMD Windows prerelease wheel does not ship (it only ships rocm72.dll), causing a load error at training start. Fix: - worker.py section 1f: set BNB_ROCM_VERSION=72 (via setdefault) before section 2 ML imports, so bitsandbytes always loads rocm72.dll on Windows ROCm - install_python_stack.py: set BNB_ROCM_VERSION=72 in _install_bnb_windows_rocm() for any post-install imports; update comment to document root cause - tests: 4 new assertions covering the fix (141 passed, 2 skipped) * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: detect BNB ROCm DLL suffix dynamically instead of hardcoding '72' BNB_ROCM_VERSION was pinned to '72' which works today (AMD wheel ships rocm72.dll) but would break again if AMD ships a future wheel with a different DLL suffix (e.g. rocm713.dll). Add _detect_bnb_rocm_dll_ver() to install_python_stack.py: scans the installed bitsandbytes package dir for libbitsandbytes_rocm{VER}.dll using importlib.util.find_spec (no BNB import needed) and returns the suffix. '72' remains the fallback when detection fails. Apply the same detection inline in worker.py section 1f. Both paths still respect a pre-set BNB_ROCM_VERSION (caller override wins). Tests: +8 cases covering detection logic and fallback (147 passed, 2 skipped). * fix: patch torch.distributed stubs in server process for Windows ROCm On Windows ROCm, torch.distributed ships without process-group helpers (is_initialized, is_available, get_rank, get_world_size). The worker subprocess already patches these in section 1e, but the main server process calls _determine_attention_impl_for_gpu_estimate() which calls unsloth's resolve_attention_implementation() → is_initialized(), causing: "Could not resolve attention implementation for '...': module 'torch.distributed' has no attribute 'is_initialized'" Fix: patch the missing attrs onto torch.distributed at the top of _determine_attention_impl_for_gpu_estimate, matching the same stubs already applied in worker.py section 1e. No-ops on Linux/CUDA where torch.distributed is fully populated. * fix: gate _grouped_mm dispatch patch on HIP < 7.13 AMD fixed the gfx1200 null HIP kernel in ROCm 7.13 (torch 2.11+). Users on the new wheel now get the real GPU _grouped_mm kernel for MoE workloads instead of the Python mm fallback. Changes: - worker.py: add _hip_ver_at_least() helper; wrap full _grouped_mm patch in `if not _hip_ver_at_least(7, 13):` with else branch that logs the skip reason; update section-1f comment to document the fix - test_rocm_support.py: add 5 tests covering the helper definition, the (7, 13) gate expression, the else branch, the skip log message, and the AMD-format version string parsing (.split(".")[:2]) Verified: torch==2.11.0+rocm7.13.0 — 3D batch and grouped (offs) variants both succeed; null crash only present on rocm7.12 and earlier. * fix: stub is_torchelastic_launched on torch.distributed for Windows ROCm resolve_attention_implementation calls is_torchelastic_launched() which does not exist in the incomplete torch.distributed shipped with the Windows ROCm wheel, causing a warning on every model config load in the server process. Add it to the stub table alongside the four helpers already patched in _determine_attention_impl_for_gpu_estimate. Also adds two tests: one confirming the new stub and one confirming all five core distributed helpers are covered. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: explicit warnings on AMD ROCm arch/version fallbacks + Fast-Install arg order setup.ps1: - Fix Fast-Install argument order: packages before flags, consistent with all other Fast-Install calls in the file (was: Fast-Install --force-reinstall --index-url $url torch ...) (now: Fast-Install torch torchvision torchaudio --force-reinstall --index-url $url) - Add explicit [WARN] substep when $HasROCm is true but arch mapping fails: - GPU arch detected but not in supported wheel list → names the arch and lists supported families so user knows exactly what to report - HIP SDK present (amd-smi path) but gcnArchName unreadable → instructs user to re-install the HIP SDK; previously fell back silently to CPU install.sh: - Add [WARN] to stderr before silent CPU fallback when AMD GPU is confirmed (rocminfo/amd-smi) but ROCm version cannot be read from any source (amd-smi, /opt/rocm/.info/version, hipconfig, dpkg, rpm) - Add [WARN] to stderr when ROCm version is too old (< 6.0) with upgrade link install.ps1 and setup.sh: no changes needed (already handle these paths correctly) * fix: robust gfx arch detection for Strix Halo / HIP-runtime-only installs Covers users who have the HIP runtime (amd-smi available) but not the full HIP SDK (no hipinfo), which is common on Strix Halo iGPU systems. Without this, $ROCmGfxArch stays null and the installer silently falls back to CPU-only PyTorch despite a working GPU. Detection waterfall (setup.ps1 + install.ps1): 1. hipinfo gcnArchName -- full HIP SDK (existing, unchanged) 2. amd-smi list gfx pattern -- newer amd-smi versions embed arch 3. amd-smi static --asic -- ROCm 6+ ASIC details with GFX target 4. UNSLOTH_ROCM_GFX_ARCH env -- manual override escape hatch 5. GPU name → arch table -- best-effort from marketing name: 890M / Strix Halo → gfx1151 (RDNA 3.5 iGPU, Strix Halo) 880M / Strix Point → gfx1150 (RDNA 3.5 iGPU, Strix Point) 780M / Phoenix → gfx1103 (RDNA 3 iGPU) RX 7900/7800/7700 → gfx1100 (RDNA 3 desktop) RX 9070 XT / 9080 → gfx1201 (RDNA 4) RX 9070 / 9060 XT → gfx1200 (RDNA 4) When arch is inferred from name, a Cyan substep tells the user to set UNSLOTH_ROCM_GFX_ARCH to skip inference on future installs. WMI block intentionally does not set $HasROCm (no runtime confirmation). Tests: 11 new tests in TestStrixHaloGfxArchDetection covering all five detection levels, WMI safety, and gfx regex in both ps1 files. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: resolve hipinfo/hipconfig via HIP_PATH/ROCM_PATH when not on PATH AMD HIP SDK sets HIP_PATH on Windows but does not always add the bin directory to PATH. Get-Command hipinfo therefore silently fails and detection falls through to WMI, which cannot provide a gfx arch, leaving the user with a CPU-only PyTorch install and no warning. Changes: - setup.ps1 / install.ps1: before falling through to amd-smi, attempt to locate hipinfo.exe and hipconfig.exe under $env:HIP_PATH\bin (then $env:ROCM_PATH\bin) when Get-Command returns nothing - Emit a [WARN] with the resolved path and a one-liner to permanently fix PATH via SetEnvironmentVariable - Emit a [WARN] when HIP_PATH/ROCM_PATH is set but the exe is still not found (incomplete SDK install) - Emit a [WARN] with the first hipinfo output line when hipinfo runs but returns a non-zero exit code (e.g. "no ROCm-capable device detected") - 18 new tests in TestHipSdkEnvPathResolution; total 183 passed, 2 skipped * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * feat: print HIP SDK path and full hipconfig version in terminal on AMD detection Both install.ps1 and setup.ps1 now emit substeps under the gpu step when AMD ROCm is detected: gpu AMD ROCm (gfx1200) HIP SDK: C:\Program Files\AMD\ROCm\7.1 hipconfig: 7.1.51803-d3a86bd04 Previously only the gpu label (e.g. "AMD ROCm (gfx1200)") was shown with no indication of where the SDK was found or which exact build was active. The full hipconfig build string (e.g. 7.1.51803-d3a86bd04 instead of just 7.1) is now stored in ROCmVersionFull and also used in setup.ps1's 'rocm' step label. 9 new tests in TestHipSdkDetectedSubstep; total 192 passed, 2 skipped * fix: Strix rocm7.1 segfault bypass + Ubuntu 24.04 HIP gcc-install-dir Issue 1 (install.sh): gfx1151/gfx1150 + ROCm 7.1 causes a segfault in torch._grouped_mm (moe_utils.py:167). The Radeon repo now ships cp313 wheels for rocm-rel-7.1, so _amd_gpu_radeon=true silently lands on the broken combo. When Strix Halo/Point is detected and TORCH_INDEX_URL is rocm7.1, override to rocm7.2 PyTorch index, update TORCH_CONSTRAINT, and set _amd_gpu_radeon=false to bypass the Radeon repo entirely. Emits a clear [WARN] explaining the segfault and linking to the ROCm upgrade docs. Issue 2 (setup.sh): ROCm 7.x ships clang-20 which on Ubuntu 24.04+ picks /usr/lib/gcc/x86_64-linux-gnu/14/ (runtime dir, no C++ headers), causing 'cstdlib file not found' and a failed llama.cpp HIP build. Iterate gcc versions 14→11 to find the first install dir that has both runtime and /usr/include/c++/<ver> headers, then pass --gcc-install-dir to clang via CMAKE_HIP_FLAGS. Fix confirmed by h34v3nzc0dex (llama.cpp 417/417 clean). 11 new tests across TestStrixRocm71Override and TestSetupShGccInstallDir; total 203 passed, 2 skipped * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: BNB_ROCM_VERSION in server process + torch._C._distributed_c10d stubs Two errors visible in training logs on Windows ROCm: 1. Server process bitsandbytes crash: "Configured ROCm binary not found at libbitsandbytes_rocm713.dll" The installed BNB wheel ships rocm72.dll (not rocm713.dll). The training worker already sets BNB_ROCM_VERSION=72 via DLL detection but the server process (main.py) imported bitsandbytes before that ran. Fix: add the same DLL-scan + BNB_ROCM_VERSION assignment to main.py inside the existing win32 guard, before any downstream import can pull in bitsandbytes. 2. torch.distributed import failure: "No module named 'torch._C._distributed_c10d'; torch._C is not a package" torch._C is a C extension on Windows ROCm — Python cannot do submodule imports from it, so torch.distributed fails to import before our attribute stubs could ever run. Fix: inject empty ModuleType stubs for _distributed_c10d, _distributed_autograd and _distributed_rpc into sys.modules inside the win32 guard in hardware.py BEFORE importing torch.distributed, so the import succeeds and our attribute stubs take effect. 9 new tests in TestServerStartupRocmFixes; total 212 passed, 2 skipped * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix(win32): populate distributed c10d stub with dummy symbols torch.distributed tries to `from torch._C._distributed_c10d import FakeProcessGroup` (and ProcessGroup, Work, Store, etc.). The previous empty ModuleType stub caused an AttributeError on those names. Populate every stub with a _Dummy class for each known symbol so the import chain completes silently on Windows ROCm where torch._C is a compiled extension and its _distributed_c10d submodule doesn't exist. Adds four new tests in TestServerStartupRocmFixes covering FakeProcessGroup, ProcessGroup, setattr population, and all three _distributed_* siblings. * fix(win32): distinguish HIP SDK installed vs GPU not ROCm-accessible Previously, when hipinfo was found but exited non-zero (e.g. "no ROCm-capable device detected"), both install.ps1 and setup.ps1 fell through to the WMI-label-only branch and printed "AMD GPU detected -- HIP SDK not found" -- factually wrong since the SDK binary is present. Add $HipSdkInstalled flag (set true when hipinfo binary is found, regardless of exit code). When HipSdkInstalled && !HasROCm: - Show "AMD GPU detected -- not ROCm-accessible (HIP <ver>)" instead - Explain this is a driver issue, not an SDK issue, with a link - Still run hipconfig version capture so version shows in output - CPU-only hint now says "GPU not ROCm-accessible" not "require HIP SDK" Also applies to setup.ps1 (same detection block, same branches). Adds TestHipSdkInstalledButDeviceInaccessible (11 tests). * fix(win32): scope ROCm workarounds to AMD hosts only Three Codex-flagged issues where Windows ROCm workarounds incorrectly applied to Windows CUDA (NVIDIA) machines: main.py (P1): BNB_ROCM_VERSION was set unconditionally on all win32 hosts. On NVIDIA, bitsandbytes sees BNB_ROCM_VERSION and looks for a ROCm DLL that doesn't exist, breaking bitsandbytes initialisation. Fix: gate the block on HIP_PATH/ROCM_PATH being present (ROCm hosts only). worker.py (P2): torchao stubs were seeded for all win32 runs, shadowing real torchao on Windows CUDA and silently disabling torchao quantization for NVIDIA users. Fix: gate on HIP_PATH/ROCM_PATH (win32 ROCm only). install_python_stack.py (P1): _detect_windows_gfx_arch() only checked shutil.which("hipinfo"), skipping the HIP_PATH/ROCM_PATH fallback that the PowerShell installers use. On installs where the HIP SDK bin dir is not on PATH, _ensure_rocm_torch() returned early without installing ROCm wheels or bitsandbytes. Fix: mirror the env-var fallback. * fix(linux): route Strix + ROCm 7.1 to AMD arch-specific index Instead of falling back to pytorch.org/rocm7.2, the Strix override now routes to repo.amd.com/rocm/whl/gfx1151/ (or gfx1150/) which serves torch 2.11.0+rocm7.13.0 -- AMD's build containing the actual _grouped_mm kernel fix, verified on real gfx1151 hardware by h34v3nzc0dex. This exercises the real GPU kernel path rather than the rocm7.2 workaround. UNSLOTH_AMD_ROCM_MIRROR can override the base URL for air-gapped installs. Also teaches _tauri_torch_index_family to recognise AMD arch-specific URLs (repo.amd.com/rocm/whl/gfx*) and return the rocm7.13 family label so _tauri_gpu_branch correctly classifies these installs as rocm. Suggested by h34v3nzc0dex based on hardware-verified probe results. * fix(studio/rocm): gate ROCm-only side-effects on active torch runtime Address five edge cases flagged during PR review: 1. studio/backend/main.py: BNB_ROCM_VERSION was set whenever HIP_PATH or ROCM_PATH was present in the environment. A Windows CUDA user who once installed the HIP SDK and reverted to a CUDA torch wheel still has those env vars set, so bitsandbytes would try to load libbitsandbytes_rocm72.dll against a CUDA torch and crash. Now probe torch.version.hip inside the env-var guard (worker.py already does this). 2. studio/backend/main.py: os.add_dll_directory returned handles were discarded. Per CPython docs, the directory leaves the DLL search list when the handle is garbage collected. Retain handles in module-level _ROCM_DLL_HANDLES list so they survive process lifetime. 3. studio/install_python_stack.py: _install_bnb_windows_rocm() returned None regardless of pip_install_try outcome, and the caller flipped _rocm_windows_torch_installed to True unconditionally. On a failed BNB install the post-install "manual install may be required" warning was suppressed and the user was misled. Helper now returns bool; caller gates on it. 4. studio/install_python_stack.py: _detect_windows_gfx_arch returned the raw capture group, so mixed-case hipinfo output ("Gfx1151") missed the lowercase keys in _GFX_TO_AMD_INDEX_ARCH and silently fell back to CPU torch. Lowercase the token. 5. studio/install_python_stack.py: UNSLOTH_ROCM_TORCH_INSTALLED=1 early- return trusted the env var even when the venv was wiped between runs. Subprocess-probe torch importability first; fall through to the full install path if the probe fails. Tests: 231 passed, 1 skipped in tests/studio/install/test_rocm_support.py (adds one new test for case 5 fall-through). * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix(studio/rocm): worker.py parity + don't roll back ROCm torch on bnb failure Addresses findings from a 10x reviewer pass on the prior fix commit: 1. studio/backend/core/training/worker.py (parity with main.py): - Gate the torchao stub block on torch.version.hip / 'rocm' in torch.__version__ instead of HIP_PATH / ROCM_PATH env-var presence. Same root cause as main.py: HIP SDK env vars stick around on CUDA hosts. - Add module-level Windows ROCm DLL registration block. Worker subprocesses inherit env vars but not the parent's add_dll_directory handles, so the first `import torch` in the worker could fail to find amdhip64.dll when HIP_PATH\bin is not on PATH. Mirrors main.py setup. Handles retained at module scope via _ROCM_DLL_HANDLES. - Promote _WINDOWS_ROCM_GROUPED_MM_LIB to module scope with `global` in run_training_process so the torch.library.Library registration survives past function return / mid-run garbage collection. - Harden _torch_has_hip() to also accept 'rocm' in torch.__version__ (AMD SDK / Radeon wheels may not set torch.version.hip). 2. studio/install_python_stack.py: - Don't roll back ROCm torch when bitsandbytes install fails. The prior commit gated _rocm_windows_torch_installed on _install_bnb_windows_rocm() returning True; if torch installed successfully but bnb failed, the flag stayed False and later install steps could overwrite ROCm torch with the generic CPU torch wheel. Set the flag after torch install; surface bnb failure as a separate warning instead. - _detect_windows_gfx_arch now probes in three tiers: UNSLOTH_ROCM_GFX_ARCH env-var override (matches the PowerShell installer), then hipinfo (PATH or HIP_PATH\bin), then amd-smi (`static --asic`, `list`). Without the amd-smi fallback, runtime-only Radeon installs without hipinfo on PATH made `studio update` return early and leave the venv on CPU torch. - Linux torch-already-rocm probe in _ensure_rocm_torch now matches the Windows probe shape: accepts torch.version.hip OR 'rocm' in torch.__version__ to cover AMD SDK / Radeon Linux wheels. 3. studio/backend/utils/hardware/hardware.py: - apply_gpu_ids() final-fallback torch probe accepts 'rocm' in torch.__version__ in addition to torch.version.hip, matching detect_hardware(). AMD SDK wheels could otherwise leak through with CUDA-only visibility masks on a spawned ROCm worker. Tests: 231 passed, 1 skipped in tests/studio/install/test_rocm_support.py (no test changes needed; the probe shape that prints the hip version (or 'rocm' sentinel) preserves the existing non-empty-string contract). Not addressed in this commit (deferred or out of scope): - Tag drift / lemonade checksum (PR 5303 surface, not this PR). - install.sh rocm7.2.1 URL: small fix, separate. - install.ps1 / setup.ps1 'Radeon 8060S' marketing-name fallback table. - Strix Halo + ROCm 7.1 routing asymmetry in Python update path. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix(studio/rocm): robustness pass - rocm tag normalisation, Strix routing parity, hardened detection Robustness pass on top of76137b2d. Four targeted fixes: 1. install.sh ROCm-tag routing normalisation. `rocm7.2.1` would route to https://download.pytorch.org/whl/rocm7.2.1 which does not exist (PyTorch publishes major.minor URLs only). Same for any future patch-level tag. Normalise every rocm{maj.min}* pattern to the bare {maj.min} index URL. 2. install.ps1 + studio/setup.ps1 marketing-name fallback. The gfx1151 row matched 890M / Strix Halo / HX 37x / HX 38x / AI 9 HX but not the actual retail name 'AMD Radeon 8060S Graphics' shipped by OEMs (Ryzen AI MAX+ 395). Add '8060S' to the regex. 3. install_python_stack.py Strix + ROCm 7.1 routing parity with install.sh. The shell installer reroutes Strix Halo / Point + ROCm 7.1 to repo.amd.com/rocm/whl/{gfx}/ (which serves torch 2.11.0+rocm7.13.0 with the upstream _grouped_mm fix). The Python `studio update` path only warned and still installed the broken generic rocm7.1 wheel. Mirror the override: detect gfx1151/gfx1150 on ROCm 7.1, route to the AMD per-gfx index, honour UNSLOTH_AMD_ROCM_MIRROR override. 4. _detect_windows_gfx_arch amd-smi parsing tightened. The amd-smi fallback added in the prior commit used a bare `\bgfx[1-9][0-9a-z]{2,3}\b` match against the lowercased stdout, which could pick up stray gfx references in warnings / device-name strings. Anchor on labelled lines first (Target_Graphics_Version, ASIC, Arch, gfx) and fall back to the bare match only when no labelled line is present. Tests: 231 passed, 1 skipped in tests/studio/install/test_rocm_support.py; sim_5301 23 cases pass (6 new sims for the Strix override + amd-smi parsing). * fix(studio/rocm): multi-GPU selection, Strix sibling handling, defensive cleanups Round 4 robustness pass based on 5 parallel Opus reviewers of head21773215. Seven items from across regression / edge-case / error-paths / architecture reviews: 1. studio/backend/main.py BNB gate: aligned with the broad ROCm check used everywhere else in this PR (torch.version.hip OR 'rocm' in __version__). AMD SDK / Radeon Linux wheels do not always populate torch.version.hip; without this, main.py would silently skip BNB_ROCM_VERSION while worker.py set it. 2. studio/install_python_stack.py _install_bnb_windows_rocm: init _ok = False before the try block. Without this, if pip_install_try itself raises (e.g. OSError on uv binary missing), the finally block restored env vars correctly but the subsequent `if not _ok:` raised UnboundLocalError, masking the original exception. 3. studio/install_python_stack.py _detect_windows_gfx_arch: - Rewrote to use re.findall (not re.search) on both hipinfo and amd-smi output, dedup tokens preserving order, and select via new _pick_visible_index() helper. - HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES (first comma entry, integer) now picks the right GPU on multi-AMD-GPU hosts. Out-of-range or non-int values fall back to the first GPU (matches detect_host behaviour in install_llama_prebuilt.py). 4. studio/install_python_stack.py Strix override now consults the runtime target before flipping: - Previous behaviour intersected gfx_codes with {gfx1151, gfx1150} and picked the first Strix arch, ignoring whether HIP_VISIBLE_DEVICES selected a non-Strix sibling (e.g. discrete RX 7900 in a mixed APU+dGPU box). Could install Strix-specific wheels onto a gfx1100 dGPU. - Now resolves the runtime gfx via _pick_visible_index() and only overrides when that runtime target is in the Strix set. 5. studio/backend/main.py + studio/backend/core/training/worker.py: ROCm version dir scan no longer sorts lexically. Previous sort placed "10.0" before "7.0" alphabetically, which would mis-prioritise ROCm 10.x bin dirs once AMD ships them. New _ver_key() splits on "." and sorts numerically with a string fallback. 6. install.sh Strix override URL: replaced ${var%/} (strips one trailing slash) with a while-loop that strips all trailing slashes, matching Python's .rstrip("/"). A user setting UNSLOTH_AMD_ROCM_MIRROR with "http://corp/whl///" no longer ends up with "http://corp/whl///gfx1151/" which strict pip proxies (artifactory, sonatype) 404 on. 7. studio/install_python_stack.py: bumped torch import probe timeout from 30s to 90s. PyTorch's lazy .so loading can take 60-90s on cold NFS or USB-backed venvs. The shorter timeout was producing a false "torch missing" classification and reinstalling a working ROCm torch. Tests: 231 passed, 1 skipped. sim_5301 30 cases pass (added 7 new sims for multi-GPU detection, Strix sibling handling, and _ok-init regression). * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix(studio/rocm): worker BNB/grouped_mm broad gate, install.sh Strix visibility, runtime-only ROCm detection Round-5 robustness pass based on 20 parallel reviewers of head96b9e465. 1. studio/backend/core/training/worker.py - BNB version pin / dynamo disable / _grouped_mm fallback block was still gated on torch.version.hip alone despite the torchao stub block above already using the broad check. AMD SDK / Radeon Windows wheels (torch.__version__ contains "rocm" but torch.version.hip is None) silently skipped the Windows ROCm runtime patches. Aligned to the same broad check (8/20 reviewers). 2. studio/backend/core/training/worker.py - _hip_ver_at_least() now also parses the ROCm version out of torch.__version__ (e.g. "2.11.0+rocm7.13.0") when torch.version.hip is missing, so the kernel-fix gate is correct for SDK / Radeon wheels too. 3. studio/backend/core/training/worker.py - _grouped_mm_safe_impl with offs=None now picks torch.bmm/matmul for 3-D inputs instead of always calling torch.mm. The real _grouped_mm accepts 3-D batched matmul; the prior fallback raised "self must be a matrix" on MoE workloads (2/20). 4. studio/backend/main.py - dropped the HIP_PATH / ROCM_PATH env-var gate from the BNB block; probe torch directly. Runtime-only Radeon / AMD SDK Windows installs do not set those SDK env vars but still ship ROCm torch (5/20 reviewers). 5. install.sh - Strix override now collects every gfx token from rocminfo / amd-smi (in enumeration order), then indexes by HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES so a mixed Strix iGPU + non- Strix dGPU host where the user selected the dGPU does NOT get rerouted to the Strix per-gfx index. Mirrors the Python update path (5/20 reviewers). 6. install.sh - Strix detection chain now also probes `amd-smi static --asic`, matching the PowerShell installer (1/20). Closes the gap on runtime-only Strix hosts where `amd-smi list` does not surface a gfx token. 7. studio/install_python_stack.py - _has_rocm_gpu() now has the sysfs KFD topology fallback (/sys/class/kfd/kfd/topology/nodes/*/gpu_id), matching install.sh. On minimal package-managed installs without rocminfo / amd-smi GUI tools, `studio update` can now detect the GPU and repair the venv instead of returning early (2/20). 8. studio/install_python_stack.py - _detect_amd_gfx_codes() now falls back to `amd-smi list` and `amd-smi static --asic` when rocminfo is missing (2/20). Strix routing on runtime-only Radeon hosts now matches what install.sh has done for a while. 9. studio/install_python_stack.py - Strix override now applies even when has_hip_torch is True. The whole point of the override is to repair an existing broken torch.version.hip == "7.1" install; skipping the reinstall left users on the known _grouped_mm segfaulting stack (3/20). Tests: 231 passed, 1 skipped. sim_5301 30 cases pass. sim_cross 12 pass. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix(studio/rocm): code review hardening pass - main.py: numeric DLL sort (string sort picked rocm72 over rocm713); add basename() to regex; log warning on detection failure; log info when BNB_ROCM_VERSION is set (mirrors worker.py) - worker.py: explicit len-guard in _hip_ver_at_least() with warning logs instead of silent IndexError/ValueError swallow - hardware.py: isinstance(result, dict) guard before result.get() in _smi_query() to prevent AttributeError on non-dict backend returns - amd.py: round() before int() on parsed GPU IDs; log warning when truncation occurs (defensive against malformed amd-smi output) - setup.sh: quote --gcc-install-dir value in CMAKE_HIP_FLAGS so paths with spaces do not break the CMake argument - install.ps1, setup.ps1: apply colon-split + ToLower() to hipinfo gcnArchName match (consistent with each other and with setup.sh) - install.sh: tighten ROCm tag case patterns to explicit rocmX.Y|rocmX.Y.* to avoid unintended prefix matches * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix(studio/training): GPU OOM guard to prevent system freeze on VRAM exhaustion On RDNA 4 (gfx1200/gfx1201) and other ROCm GPUs, exhausting VRAM can cause a HIP driver hang that freezes the entire system rather than raising a recoverable Python exception. Two-part fix: - set_per_process_memory_fraction(0.90) caps the HIP/CUDA allocator at 90% of VRAM so PyTorch raises OutOfMemoryError before hitting the hardware limit, keeping the driver alive and the system responsive - top-level exception handler detects OOM errors by type and message and surfaces a clear actionable message to the UI (reduce max_seq_length, enable gradient_checkpointing, lower batch size) instead of the raw CUDA/HIP error string * fix(studio/rocm): OOM guard ROCm-only + unified memory, multi-GPU arch selection OOM guard (worker.py): - Scope to _hw.IS_ROCM only -- NVIDIA CUDA has a graceful OOM path and does not need the allocator cap - Detect unified memory by comparing torch VRAM against psutil system RAM; use 0.80 on unified-memory APUs (gfx1151 Strix Halo) where the GPU pool is carved from host RAM, 0.90 on discrete cards Multi-GPU arch selection: - install.ps1 / setup.ps1: replace -match (first hit only) with [regex]::Matches() to collect all gcnArchName entries, then index by HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES - install_python_stack.py: index into full token list before dedup so HIP_VISIBLE_DEVICES=2 on [gfx1100, gfx1100, gfx1151] resolves gfx1151 - install.sh: remove awk dedup from gfx token collection for same reason GCC multiarch (setup.sh): - Only append -linux-gnu when gcc -print-multiarch does not already return the full triple, fixing double-suffix on Ubuntu 24.04 * fix(tests): update ROCm version cap expectations from rocm7.1 to rocm7.2 Daniel's normalisation commit updated the cap from rocm7.1 to rocm7.2 since PyTorch now publishes that index and rocm7.2 ships torch 2.11.0. Test expectations were stale. * fix(tests): correct MLX smoke test losses_per_step assertion logging_steps=1 with max_steps=30 produces 30 loss entries, not 7. The assertion was stale from a previous config. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix(studio/worker): detect unified-memory APU by GPU name not VRAM/RAM ratio The previous heuristic (VRAM > 50 % of system RAM) false-positived on discrete cards in low-RAM systems — e.g. RX 9060 XT 16 GB on a 16 GB or 24 GB machine would trip the unified-memory path and log "unified memory host" when it should say "discrete". AMD iGPUs (gfx1150/gfx1151 Strix Halo, Strix Point, etc.) expose names with a digit+M suffix ("AMD Radeon 890M"), while discrete cards use "RX NNNN [XT|XTX]" naming. Matching that suffix is reliable across all current ROCm-capable AMD consumer GPUs and does not require psutil. Also includes the device name in the log line to ease future debugging. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix(install/setup.ps1): force array on hipinfo gcnArchName parse to fix single-GPU arch truncation When [regex]::Matches() finds exactly one match, PowerShell's pipeline unwraps the result to a scalar string. Indexing a scalar string with [0] returns the first *character*, so a one-GPU system would parse gcnArchName "gfx1200" as "g", which is not in the supported arch map and triggers the CPU-only fallback. Wrapping with @() forces the result to remain an array regardless of match count. On a single-GPU machine the arch is now correctly read as "gfx1200" (or whatever the full name is) so the ROCm wheel index is selected. Reproducer: hipinfo exits 0 and outputs exactly one gcnArchName line. Without @(), $_hipAllArches = "gfx1200" (String); $_hipAllArches[0] = 'g'. With @(), $_hipAllArches = @("gfx1200") (Object[]); $_hipAllArches[0] = "gfx1200". * fix(studio/rocm): classify unified-memory APU via VRAM/RAM ratio, not arch list Replace the gcnArchName allowlist {gfx1150, gfx1151} with a psutil-based heuristic: unified APUs expose the entire system RAM as the HIP pool (ratio ≥ 0.90), discrete cards are well below that. No arch name required — future APUs classify correctly without code changes. Also removes the stale import re / \d[Mm]\b device-name regex that5d84704left behind, and logs vram/sys GiB for easier on-hardware verification. Addresses h34v3nzc0dex review: Radeon 8060S (gfx1151, 128 GiB unified) now correctly gets 0.80 cap instead of 0.90. * fix(studio/rocm): revert to gcnArchName for unified-memory APU classification VRAM/RAM ratio >= 0.90 false-positives on machines where discrete VRAM equals system RAM (e.g. RX 9060 XT 16 GB + 16 GB system RAM → ratio 1.0, incorrectly classified as unified → wrong 0.80 cap applied). gcnArchName is the correct signal: naming-independent, stable within a product family, and already parsed throughout this PR. Unified set is {gfx1150, gfx1151} (Strix Point + Strix Halo). * fix(studio/llama-prebuilt): resolve hipinfo via HIP_PATH/ROCM_PATH on Windows shutil.which("hipinfo") returns None when the HIP SDK bin dir is not on PATH -- the HIP SDK installer sets HIP_PATH/ROCM_PATH but does not always add the bin dir to PATH. This caused has_rocm=False in the prebuilt asset selector, so AMD ROCm machines got the CPU llama.cpp zip instead of the HIP one, silently running all chat inference on CPU. Add _resolve_exe() that falls back to %HIP_PATH%\bin and %ROCM_PATH%\bin when shutil.which() finds nothing, mirroring the same fallback already present in setup.ps1. * fix(studio/llama-prebuilt): pass --has-rocm from setup.ps1 to skip re-detection The Python prebuilt installer re-detects ROCm independently via shutil.which("hipinfo"), which fails when hipinfo is not on PATH (HIP SDK sets HIP_PATH but doesn't always add the bin dir to PATH). This caused has_rocm=False and downloaded the CPU llama.cpp zip even on confirmed AMD ROCm machines. setup.ps1 already performs reliable ROCm detection with its own HIP_PATH/ROCM_PATH fallback. Add --has-rocm flag to install_llama_prebuilt.py so setup.ps1 can forward its result directly, and pass it whenever $HasROCm is true. The Python script then overrides has_rocm=True in the HostInfo without re-probing. * fix(studio/llama-prebuilt): add HIP asset to simple-policy Windows path direct_upstream_release_plan (used by --simple-policy, which setup.ps1 always passes) only checked has_usable_nvidia on Windows and fell straight to CPU for AMD ROCm machines, ignoring has_rocm entirely. The --has-rocm override had no effect because the simple-policy code path never reached resolve_asset_choice where has_rocm was checked. Add an elif branch for has_rocm that tries the upstream HIP asset (llama-TAG-bin-win-hip-radeon-x64.zip) before falling through to the CPU fallback, consistent with the non-simple-policy path. * fix(studio/setup.ps1): auto-remove mismatched llama.cpp install kind When an existing llama.cpp install is the wrong kind for the current GPU (e.g. windows-cpu on an AMD ROCm machine that should have windows-hip), the prebuilt installer skips on tag match and never upgrades. Read install_kind from UNSLOTH_PREBUILT_INFO.json before invoking the installer and remove the directory if the kind doesn't match, forcing a fresh download of the correct variant. * fix(studio/setup.ps1): show live PyTorch install output in verbose mode for ROCm The ROCm torch reinstall (setup.ps1 phase) always silently captured output, so in --verbose mode the torch downgrade mid-install (2.11.0+rocm → 2.10.0 → 2.11.0+rocm) looked like the final state was 2.10.0. Match the CPU/CUDA blocks which show live uv output when $script:UnslothVerbose is set. * fix(rocm/windows): set ROCBLAS_TENSILE_LIBPATH for bundled rocblas.dll The llama.cpp ROCm prebuilt bundles rocblas.dll next to the binary but not the Tensile kernel library files it depends on at runtime (rocblas/library/TensileLibrary*.dat + *.hsaco). The bundled DLL searches for these files relative to its own location by default, i.e. <binary_dir>/rocblas/library/, which does not exist in the prebuilt install tree. This causes a silent crash on the very first GEMM (prefill) with no output from llama-server, seen by the caller as WinError 10054 / 10061. Model load and the single-token warmup pass because they use simpler code paths that do not trigger rocBLAS GEMM. Fix: set ROCBLAS_TENSILE_LIBPATH in the subprocess env to <HIP_PATH>/bin/rocblas/library so the bundled DLL finds the kernel files from the system ROCm installation. Uses setdefault so a user- supplied env var is never overwritten. No-ops on CUDA and CPU (no HIP_PATH) and on Linux (win32 branch only). Reproducer log: rocBLAS error: Cannot read .../Release/rocblas/library/TensileLibrary.dat rocBLAS error: Could not initialize Tensile host: directory_iterator: The system cannot find the path specified. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix(install.sh): restore gfx token dedup in Strix multi-GPU awk indexer536a54dfremoved the per-source `| awk '!seen[$0]++'` dedup from the _gfx_all collection step but left the indexer awk as bare NF, so on a mixed-arch host (e.g. dGPU gfx1100 + Strix iGPU gfx1151) where rocminfo emits each gfx token twice (Name: field + ISA triple), HIP_VISIBLE_DEVICES=1 indexed vals[1] = the second gfx1100 occurrence instead of gfx1151, triggering the Strix routing on the wrong GPU. Add !seen[$0]++ to the indexer awk so duplicate tokens from the same GPU collapse to one entry before the HIP_VISIBLE_DEVICES index is applied -- matching exactly what the Python side does with dict.fromkeys() in _detect_amd_gfx_codes(). The comment above the block ("skip duplicates") already documented this as the intended behaviour. * fix(studio/install): correct _TOTAL progress count on Windows base_total += 3 fired for all non-macOS platforms including Windows, but flash-attn (line 1620) and ROCm torch final (line 1705) are both guarded by 'not IS_WINDOWS and not IS_MACOS', so on Windows with torch enabled _TOTAL was 13 while only 11 _progress() calls actually execute. Split into +1 for the ROCm torch check (all non-macOS) and +2 for the two Linux-only steps, so Windows gets _TOTAL=11 and Linux gets 14. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix(install.ps1): enforce torch>=2.11.0 for gfx120X and Strix on Windows The AMD arch-specific index (repo.amd.com/rocm/whl/gfx120X-all/ and gfx1151/) publishes torch wheels from 2.7.1 through 2.11.0. Without a version floor pip can resolve to torch 2.10.0+rocm7.12 on RDNA 4 (gfx120X) or torch 2.10.0+rocm7.1 on Strix (gfx1151/gfx1150), both of which have a null-pointer crash in torch._C._grouped_mm (TheRock issues #5284 / #3284). torch 2.11.0+rocm7.13 contains the fix. Add $ROCmTorchFloor alongside $ROCmIndexUrl: set to torch>=2.11.0 for the two affected arch families, null for all others. Wire it into the uv pip install call so the broken wheels are never selected. * fix(rocm/windows): address Codex nits - deterministic DLL suffix, CUDA llama.cpp kind, HIP_VISIBLE_DEVICES arch indexing - install_python_stack.py / worker.py: _detect_bnb_rocm_dll_ver() and the inline worker probe now collect ALL libbitsandbytes_rocm*.dll suffixes and return max() by numeric value instead of stopping at the first glob hit. Filesystem glob order is not guaranteed; this ensures '713' always wins over '72' when both variants are present in the wheel. - setup.ps1 (expectedKind): add 'windows-cuda' branch so NVIDIA hosts are not treated as 'windows-cpu'. Previously an existing windows-cuda prebuilt was always considered a mismatch on non-ROCm machines, forcing an unnecessary re-download on every update. - setup.ps1 (amd-smi gfx arch): collect ALL gfx tokens from amd-smi list output in GPU order and honour HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES when selecting which arch to use. On mixed-arch AMD systems where the visible GPU is not the first enumerated one, this prevents installing an incompatible wheel index. Falls back to index 0 (same as before) when the visibility var is unset or is a comma-separated list. - test_rocm_support.py: add test_picks_highest_suffix_when_multiple_dlls to cover the multi-DLL case that was previously untested. * fix(rocm): misleading amd-smi log, BNB spec consistency, torch ceiling for AMD index amd.py: split 'returncode != 0 or not stdout' into two separate branches. Previously, exit-0 with empty output logged 'amd-smi returned code 0' (which reads as success, not a warning) and incorrectly incremented the circuit-breaker counter. Now: non-zero exit logs the code and counts toward the limit as before; empty stdout on exit 0 logs at DEBUG level and does not penalise the counter (amd-smi --json always emits at least [] on exit 0, so this branch is rare and is not a tool failure). main.py: replace spec.origin / os.path.dirname() with spec.submodule_search_locations to match install_python_stack.py and worker.py. For normal wheel installs both approaches reach the same directory, but using submodule_search_locations is the canonical way and handles editable bitsandbytes installs correctly. Also use max() by numeric suffix (same as the other two sites) instead of a sort-then-break loop. install.ps1: add <2.12.0 ceiling to the torch constraint for gfx120X (RDNA 4) and gfx1151/gfx1150 (Strix). AMD actively publishes new versions on their per-arch index; without a ceiling, a future 2.12.0+rocmX.Y wheel would be pulled in automatically before being validated on these architectures. The ceiling matches the existing Linux install_python_stack.py constraint for the same arches. Bump both when 2.12.x is confirmed working. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix(rocm): torch floor in setup.ps1, torchvision pin for Strix, rocmsdk in _hip_ver_at_least setup.ps1: add \ (mirrors install.ps1) and derive \ from it. Previously the AMD index install called 'Fast-Install torch torchvision torchaudio --force-reinstall --index-url \' with no version constraint, so pip could resolve torch 2.10.0+rocm7.12 for gfx1151/gfx1200 -- the exact broken wheel the PR is meant to avoid. Now gfx120X and Strix enforce 'torch>=2.11.0,<2.12.0', matching install.ps1 and the Linux constraint. install_python_stack.py: pin torchvision and torchaudio in _strix_override_pkgs. The Strix Linux override uses --index-url (exclusive, no PyPI fallback); bare unversioned 'torchvision' and 'torchaudio' could resolve a build from AMD's index targeting a different torch major, causing ABI/version mismatches at runtime. Now pinned to '>=0.26.0,<0.27.0' and '>=2.11.0,<2.12.0' respectively, matching _ROCM_TORCH_CONSTRAINT['rocm7.2']. worker.py: extend _hip_ver_at_least to handle AMD SDK wheel version strings. The fallback regex r'rocm(\d+)\.(\d+)' cannot match '2.9.0+rocmsdk20251116' (no rocmX.Y component), so the function always returned False on SDK/Radeon wheels -- installing the Python _grouped_mm workaround on wheels that already have the working HIP kernel. Added a second check: if the version string contains '+rocmsdk', assume >= 7.13 (the rocmsdk format post-dates the gfx120X null-kernel fix) and skip the fallback. * fix(rocm): warn on OOB HIP_VISIBLE_DEVICES, bail on empty numeric_ids mask - setup.ps1: when HIP/ROCR_VISIBLE_DEVICES names an index beyond the detected GPU count, emit a yellow warning and fall back to GPU 0 instead of silently reading allGfxArches[-1] (wrong arch) - hardware.py _reconcile_primary_rocm_unified_memory: distinguish numeric_ids=None (no env var, use torch ordinal 0) from numeric_ids=[] (empty mask / HIP_VISIBLE_DEVICES=-1, no GPU visible); bail out early in the empty case to avoid querying torch.device(0) incorrectly * fix(rocm): gate StubSubpackageFinder on win32 ROCm, add gcnArchName fallbacks - worker.py _StubSubpackageFinder: the meta_path append was running on every platform on every call to run_training_process; moved it inside the if _is_win32_rocm: block since stubs are only seeded there and the finder is a pure accumulation on Linux/Windows CUDA - worker.py OOM guard: AMD SDK / Radeon wheels may not populate gcnArchName, causing Strix Halo to be misclassified as discrete and get the 0.90 cap (12.8 GB OS headroom) instead of 0.80 (25.6 GB); now tries gcn_arch_name / arch_name / gfx_arch_name variants first, then falls back to device-name matching (890M -> Strix Halo, 880M -> Strix Point) with a debug log when the fallback fires * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix(rocm): pin torchvision/torchaudio in setup.ps1, remove -Unique from arch array - setup.ps1 ROCm torch install: torchvision and torchaudio were passed bare alongside pinned torch>=2.11.0,<2.12.0 for gfx1151/gfx1200 arches. AMD publishes packages independently so a future torchvision 0.27 (for torch 2.12) on the same arch index would cause pip ResolutionImpossible or an ABI-incompatible install. Added torchvisionFloorMap and torchaudioFloorMap mirroring install_python_stack.py's strix override (torchvision>=0.26.0,<0.27.0, torchaudio>=2.11.0,<2.12.0) and derived ROCmVisionSpec/ROCmAudioSpec used in all three Fast-Install call sites. - setup.ps1 amd-smi arch detection: Select-Object -Unique was collapsing same-arch multi-GPU arrays (e.g. two gfx1151 APUs -> 1-element array) causing HIP_VISIBLE_DEVICES=1 to trigger a false out-of-range warning and fall back to GPU 0 even though the correct GPU would have been at index 1. Removed -Unique; added comment noting the positional-index assumption and its non-contiguous-GPU limitation. * fix(rocm): add 8060s/8050s to OOM guard device-name fallback, extract classifier helper Path 3 of the OOM guard device-name fallback only checked for 890m/880m (gfx1150 Strix Point SKU names). Strix Halo (gfx1151) ships as Radeon 8060S (Ryzen AI MAX+ 395) and Radeon 8050S (cut-down SKU) -- neither matches, so the fallback returned is_unified=False and applied the 0.90 fraction instead of 0.80, leaving ~12.8 GiB OS headroom on a 128 GiB pool instead of ~25.6 GiB. Fix: add 8060s and 8050s to the name-match set. Also correct the comment that mislabelled 890M as a Strix Halo name (it is Strix Point). Refactor: extract the three-path classifier into _rocm_classify_unified_memory() so it can be unit-tested directly. Add 31 test cases in test_rocm_oom_guard.py covering all three paths and the regression case (Radeon 8060S Graphics). Reported-by: h34v3nzc0dex * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix(rocm): pass explicit dtype on bf16-unsupported hardware (RDNA2) dtype=None lets unsloth auto-detect the model dtype. On RDNA2 (gfx103x, e.g. RX 6600) is_bfloat16_supported() incorrectly returns True, so unsloth picks bf16 and the first bf16 kernel dispatch triggers: LLVM ERROR: Cannot select: intrinsic %llvm.amdgcn.fdot2.bf16.bf16 Replace every dtype=None in load_model() with _auto_dtype which resolves to None when bf16 is supported (all modern NVIDIA + RDNA3+) and torch.float16 otherwise. This gives RDNA2 users a working float16 training path without touching NVIDIA behaviour at all. Fixes: https://github.com/unslothai/unsloth/issues/5337 * fix: reduce log noise for expected non-issues on Windows ROCm Three log lines fired at warning/error level for conditions that are completely expected on a Windows HIP SDK-only setup: amd.py - amd-smi WinError 2 (FileNotFoundError): downgrade warning -> debug. amd-smi ships with Adrenalin, not the HIP SDK; absence is normal. - 'disabling' message: downgrade warning -> info with clearer text 'not available (not installed; expected on HIP SDK-only systems); GPU VRAM polling disabled' hardware.py - torch.distributed.Store missing: downgrade warning -> debug. The distributed stub added in this PR intentionally omits Store; the attention-impl fallback to eager is expected and non-actionable. worker.py - causal-conv1d: add early Windows exit (info) in both _ensure_causal_conv1d_fast_path and _causal_conv1d_install hook; no cp313/win_amd64 wheel exists, so the install always fails. - FLA: add early Windows exit (info) in _ensure_flash_linear_attention_unconditional; triton dependency has no cp313/win_amd64 wheel. - Defense-in-depth: _install_package_wheel_first non-HIP PyPI failure logs info+debug on Windows instead of error; FLA failure logs info+debug on Windows instead of warning. * [AMD] FIx installation of bitsandbytes when it's from .dev and skip rebuilding llama.cpp if we build it manually. * fix: use force_pip for Windows ROCm bitsandbytes prebuilt wheel install uv rejects the bnb continuous-release wheel due to filename/metadata version mismatch (1.33.7.preview vs 0.50.0.dev0). Switch to force_pip=True (pip bypass) instead of the UV_SKIP_WHEEL_FILENAME_CHECK env var workaround -- cleaner and consistent with how the Linux path handles it. BNB_ROCM_VERSION is still set post-install to the detected DLL suffix so the worker subprocess loads the correct libbitsandbytes_rocm{VER}.dll even when torch.version.hip reports a newer HIP version than the wheel ships. * fix: three small correctness fixes found in PR review - _install_bnb_windows_rocm: use UV_SKIP_WHEEL_FILENAME_CHECK=1 with try/finally instead of force_pip=True so the env var is always restored and the failing CI test passes - _determine_attention_impl_for_gpu_estimate: gate torch._C distributed stubs on IS_ROCM so Windows CUDA users keep the real extension - install.ps1 amd-smi fallback: collect all gfx tokens and index by HIP_VISIBLE_DEVICES, matching the hipinfo path on multi-GPU hosts * fix: stub torchao in export subprocess on Windows ROCm On Windows, the ROCm build of PyTorch ships without the distributed C extension (torch._C._distributed_c10d). torchao, which is pulled in transitively by transformers.quantizers at import time, walks into torch.distributed._functional_collectives -> distributed_c10d and crashes with: No module named 'torch._C._distributed_c10d'; 'torch._C' is not a package This only affected the export subprocess because the training subprocess already applied an identical torchao stub (introduced separately to fix the same root cause). The export subprocess had no such guard and died during 'Importing Unsloth...' before any model loading could happen. Fix: apply the same _StubSubpackageFinder / torchao stub pattern to the export subprocess entry point, gated on Windows ROCm detection, before any import of transformers or unsloth_zoo. Root cause tracked in ROCm/TheRock#3284 (libuv / torch.distributed missing on Windows ROCm builds). Ref: https://github.com/ROCm/TheRock/issues/3284 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * install.sh, setup.sh: add GPU arch step logging to match PS1 scripts Both shell scripts were missing the step "gpu" terminal log block that install.ps1 and setup.ps1 emit. This adds equivalent output: GPU label with gfx arch (e.g. "AMD ROCm (gfx1151)"), ROCm root path, hipconfig version, and marketing name substep. Includes the same gfx arch detection chain (rocminfo → amd-smi list → amd-smi static --asic), UNSLOTH_ROCM_GFX_ARCH env override, and name-based arch inference table (Strix Halo/Point, RDNA 3/4) as the PS1 versions. install.sh also replaces bare echo blocks for the AMD ROCm and CPU-only cases with formatted substep output. * Fix BNB_ROCM_VERSION gate, ROCm GPU mask preference, APU unified memory and Release build for PR #5301 - main.py: gate BNB_ROCM_VERSION on the rocm bnb DLL or HIP_PATH/ROCM_PATH instead of importing torch on every Windows host - hardware.py: prefer HIP/ROCR visible-device masks only on ROCm hosts so a stale mask cannot override CUDA_VISIBLE_DEVICES on NVIDIA - llama_cpp.py: set GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 only for unified-memory APUs (gfx1150/gfx1151) - setup.sh: pass -DCMAKE_BUILD_TYPE=Release for the HIP source build - add test_amd_apu_unified_memory.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: guard recompile_limit + fix AMD VRAM monitor fallback trainer.py: torch._dynamo.config.recompile_limit does not exist in some ROCm torch builds (e.g. pytorch.org/whl/rocm6.2 wheels). Guard the assignment so training doesn't crash on RDNA2/RDNA3. hardware.py: when amd-smi/nvidia-smi is unavailable or returns no usable data (HIP SDK-only Windows, Docker, unexpected JSON format), the existing fallback used torch.cuda.memory_allocated() which is process-specific and reads near-zero even with a fully loaded model. Switch to torch.cuda.mem_get_info() via _torch_get_per_device_info() which reports system-wide VRAM occupancy so the GPU monitor shows real usage on all AMD systems without requiring amd-smi. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: Windows VRAM monitor via Performance Counter API When amd-smi/nvidia-smi is unavailable on Windows, query dedicated GPU VRAM via Windows Performance Counters (same source as Task Manager). This gives system-wide cross-process usage, fixing the near-zero reading caused by torch.cuda.mem_get_info only seeing the Studio server process. Linux fallback path unchanged (mem_get_info is system-wide on ROCm). * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: rename to _rocm_windows_perf_counter_vram_gb, scope to IS_ROCM Function is AMD ROCm specific — amd-smi absent on Windows when only the HIP SDK is installed. Scoped to IS_ROCM so NVIDIA Windows path is untouched (nvidia-smi handles that case). * fix: AMD VRAM monitor — Linux DRM sysfs + Windows perf counter Linux: read /sys/class/drm/card*/device/mem_info_vram_used|total for system-wide GPU memory across all processes. No tools required, always present on Linux AMD systems. Windows: Windows Performance Counter API (already added). Both paths are gated on IS_ROCM and only fire when amd-smi is absent. torch mem_get_info remains as last resort (process-local). * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: AMD GPU monitor — utilization, temperature, and power for Windows and Linux fallback paths - Windows: GPU utilization via \GPU Engine(*engtype_3D*)\Utilization Percentage perf counter - Windows: temperature and power via ADL (atiadlxx.dll, ships with Adrenalin) - Linux: GPU utilization via DRM sysfs gpu_busy_percent - Linux: temperature via hwmon temp1_input (millidegrees C) - Linux: power via hwmon power1_average / power1_input (microwatts) All paths are no-op fallbacks (None) when the source is unavailable. Mirrors what nvidia-smi provides on the CUDA path. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: remove ADL ctypes — does not support AMD iGPU (Strix Halo) * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com> Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com> Co-authored-by: Erland366 <erland.pg366@gmail.com> Co-authored-by: danielhanchen <michaelhan2050@gmail.com>
5893 lines
214 KiB
Python
5893 lines
214 KiB
Python
#!/usr/bin/env python3
|
|
# SPDX-License-Identifier: AGPL-3.0-only
|
|
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
|
|
|
"""Cross platform llama.cpp prebuilt installer for Unsloth Studio"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import argparse
|
|
import errno
|
|
import fnmatch
|
|
import hashlib
|
|
import json
|
|
import os
|
|
import platform
|
|
import random
|
|
import re
|
|
import shutil
|
|
import site
|
|
import socket
|
|
import subprocess
|
|
import sys
|
|
import tarfile
|
|
import tempfile
|
|
import textwrap
|
|
import time
|
|
import urllib.error
|
|
import urllib.parse
|
|
import urllib.request
|
|
import zipfile
|
|
from contextlib import contextmanager
|
|
from dataclasses import dataclass, field, replace as dataclasses_replace
|
|
|
|
try:
|
|
from filelock import FileLock, Timeout as FileLockTimeout
|
|
except ImportError:
|
|
FileLock = None
|
|
FileLockTimeout = None
|
|
from pathlib import Path
|
|
from typing import Any, Iterable, Iterator
|
|
|
|
|
|
EXIT_SUCCESS = 0
|
|
EXIT_FALLBACK = 2
|
|
EXIT_ERROR = 1
|
|
EXIT_BUSY = 3
|
|
|
|
|
|
def windows_hidden_subprocess_kwargs() -> dict[str, object]:
|
|
"""Return Windows-only subprocess kwargs that suppress console windows."""
|
|
if sys.platform != "win32":
|
|
return {}
|
|
|
|
kwargs: dict[str, object] = {}
|
|
create_no_window = getattr(subprocess, "CREATE_NO_WINDOW", 0)
|
|
if create_no_window:
|
|
kwargs["creationflags"] = create_no_window
|
|
|
|
startupinfo_factory = getattr(subprocess, "STARTUPINFO", None)
|
|
startf_use_showwindow = getattr(subprocess, "STARTF_USESHOWWINDOW", 0)
|
|
sw_hide = getattr(subprocess, "SW_HIDE", 0)
|
|
if startupinfo_factory is not None and startf_use_showwindow:
|
|
startupinfo = startupinfo_factory()
|
|
startupinfo.dwFlags |= startf_use_showwindow
|
|
startupinfo.wShowWindow = sw_hide
|
|
kwargs["startupinfo"] = startupinfo
|
|
|
|
return kwargs
|
|
|
|
|
|
def env_int(name: str, default: int, *, minimum: int | None = None) -> int:
|
|
raw = os.environ.get(name)
|
|
if raw is None:
|
|
value = default
|
|
else:
|
|
try:
|
|
value = int(str(raw).strip())
|
|
except (TypeError, ValueError):
|
|
value = default
|
|
if minimum is not None:
|
|
value = max(minimum, value)
|
|
return value
|
|
|
|
|
|
# Prefer "latest" over "master" -- "master" bypasses the prebuilt resolver
|
|
# (no matching GitHub release), forces a source build, and causes HTTP 422
|
|
# errors. Only use "master" temporarily when the latest release is missing
|
|
# support for a new model architecture.
|
|
DEFAULT_LLAMA_TAG = os.environ.get("UNSLOTH_LLAMA_TAG", "latest")
|
|
# Default published repo for prebuilt release resolution. Linux uses
|
|
# Unsloth prebuilts; setup.sh/setup.ps1 pass --published-repo explicitly
|
|
# for macOS/Windows to override with ggml-org/llama.cpp when needed.
|
|
DEFAULT_PUBLISHED_REPO = "unslothai/llama.cpp"
|
|
DEFAULT_PUBLISHED_TAG = os.environ.get("UNSLOTH_LLAMA_RELEASE_TAG")
|
|
DEFAULT_PUBLISHED_MANIFEST_ASSET = os.environ.get(
|
|
"UNSLOTH_LLAMA_RELEASE_MANIFEST_ASSET", "llama-prebuilt-manifest.json"
|
|
)
|
|
DEFAULT_PUBLISHED_SHA256_ASSET = os.environ.get(
|
|
"UNSLOTH_LLAMA_RELEASE_SHA256_ASSET", "llama-prebuilt-sha256.json"
|
|
)
|
|
UPSTREAM_REPO = "ggml-org/llama.cpp"
|
|
UPSTREAM_RELEASES_API = f"https://api.github.com/repos/{UPSTREAM_REPO}/releases/latest"
|
|
TEST_MODEL_URL = (
|
|
"https://huggingface.co/ggml-org/models/resolve/main/tinyllamas/stories260K.gguf"
|
|
)
|
|
TEST_MODEL_SHA256 = "270cba1bd5109f42d03350f60406024560464db173c0e387d91f0426d3bd256d"
|
|
VALIDATION_MODEL_CACHE_DIRNAME = ".cache"
|
|
VALIDATION_MODEL_CACHE_FILENAME = "stories260K.gguf"
|
|
INSTALL_LOCK_TIMEOUT_SECONDS = 300
|
|
INSTALL_STAGING_ROOT_NAME = ".staging"
|
|
GITHUB_AUTH_HOSTS = {"api.github.com", "github.com"}
|
|
RETRYABLE_HTTP_STATUS = {408, 429, 500, 502, 503, 504}
|
|
HTTP_FETCH_ATTEMPTS = 4
|
|
HTTP_FETCH_BASE_DELAY_SECONDS = 0.75
|
|
JSON_FETCH_ATTEMPTS = 3
|
|
DEFAULT_GITHUB_RELEASE_SCAN_MAX_PAGES = env_int(
|
|
"UNSLOTH_LLAMA_GITHUB_RELEASE_SCAN_MAX_PAGES",
|
|
5,
|
|
minimum = 1,
|
|
)
|
|
SERVER_PORT_BIND_ATTEMPTS = 3
|
|
SERVER_BIND_RETRY_WINDOW_SECONDS = 5.0
|
|
TTY_PROGRESS_START_DELAY_SECONDS = 0.5
|
|
DEFAULT_MAX_PREBUILT_RELEASE_FALLBACKS = env_int(
|
|
"UNSLOTH_LLAMA_MAX_PREBUILT_RELEASE_FALLBACKS",
|
|
2,
|
|
minimum = 1,
|
|
)
|
|
FORCE_COMPILE_DEFAULT_REF = os.environ.get("UNSLOTH_LLAMA_FORCE_COMPILE_REF", "master")
|
|
|
|
DIRECT_LINUX_BUNDLE_PROFILES: dict[str, dict[str, Any]] = {
|
|
"cuda12-older": {
|
|
"runtime_line": "cuda12",
|
|
"coverage_class": "older",
|
|
"supported_sms": ["70", "75", "80", "86", "89"],
|
|
"min_sm": 70,
|
|
"max_sm": 89,
|
|
"rank": 10,
|
|
},
|
|
"cuda12-newer": {
|
|
"runtime_line": "cuda12",
|
|
"coverage_class": "newer",
|
|
"supported_sms": ["86", "89", "90", "100", "120"],
|
|
"min_sm": 86,
|
|
"max_sm": 120,
|
|
"rank": 20,
|
|
},
|
|
"cuda12-portable": {
|
|
"runtime_line": "cuda12",
|
|
"coverage_class": "portable",
|
|
"supported_sms": ["70", "75", "80", "86", "89", "90", "100", "120"],
|
|
"min_sm": 70,
|
|
"max_sm": 120,
|
|
"rank": 30,
|
|
},
|
|
"cuda13-older": {
|
|
"runtime_line": "cuda13",
|
|
"coverage_class": "older",
|
|
"supported_sms": ["75", "80", "86", "89"],
|
|
"min_sm": 75,
|
|
"max_sm": 89,
|
|
"rank": 40,
|
|
},
|
|
"cuda13-newer": {
|
|
"runtime_line": "cuda13",
|
|
"coverage_class": "newer",
|
|
"supported_sms": ["86", "89", "90", "100", "120"],
|
|
"min_sm": 86,
|
|
"max_sm": 120,
|
|
"rank": 50,
|
|
},
|
|
"cuda13-portable": {
|
|
"runtime_line": "cuda13",
|
|
"coverage_class": "portable",
|
|
"supported_sms": ["75", "80", "86", "89", "90", "100", "120"],
|
|
"min_sm": 75,
|
|
"max_sm": 120,
|
|
"rank": 60,
|
|
},
|
|
}
|
|
|
|
|
|
@dataclass
|
|
class HostInfo:
|
|
system: str
|
|
machine: str
|
|
is_windows: bool
|
|
is_linux: bool
|
|
is_macos: bool
|
|
is_x86_64: bool
|
|
is_arm64: bool
|
|
nvidia_smi: str | None
|
|
driver_cuda_version: tuple[int, int] | None
|
|
compute_caps: list[str]
|
|
visible_cuda_devices: str | None
|
|
has_physical_nvidia: bool
|
|
has_usable_nvidia: bool
|
|
has_rocm: bool = False
|
|
|
|
|
|
@dataclass
|
|
class AssetChoice:
|
|
repo: str
|
|
tag: str
|
|
name: str
|
|
url: str
|
|
source_label: str
|
|
# Paired runtime archive (Windows CUDA cudart bundle). When set,
|
|
# install_from_archives also downloads it and overlays its DLLs on
|
|
# top of the main install. See unslothai/unsloth#5106.
|
|
runtime_name: str | None = None
|
|
runtime_url: str | None = None
|
|
runtime_sha256: str | None = None
|
|
is_ready_bundle: bool = False
|
|
install_kind: str = ""
|
|
bundle_profile: str | None = None
|
|
runtime_line: str | None = None
|
|
coverage_class: str | None = None
|
|
supported_sms: list[str] | None = None
|
|
min_sm: int | None = None
|
|
max_sm: int | None = None
|
|
selection_log: list[str] | None = None
|
|
expected_sha256: str | None = None
|
|
|
|
|
|
@dataclass(frozen = True)
|
|
class PublishedLlamaArtifact:
|
|
asset_name: str
|
|
install_kind: str
|
|
runtime_line: str | None
|
|
coverage_class: str | None
|
|
supported_sms: list[str]
|
|
min_sm: int | None
|
|
max_sm: int | None
|
|
bundle_profile: str | None
|
|
rank: int
|
|
|
|
|
|
@dataclass
|
|
class PublishedReleaseBundle:
|
|
repo: str
|
|
release_tag: str
|
|
upstream_tag: str
|
|
manifest_sha256: str | None = None
|
|
source_repo: str | None = None
|
|
source_repo_url: str | None = None
|
|
source_ref_kind: str | None = None
|
|
requested_source_ref: str | None = None
|
|
resolved_source_ref: str | None = None
|
|
source_commit: str | None = None
|
|
source_commit_short: str | None = None
|
|
assets: dict[str, str] = field(default_factory = dict)
|
|
manifest_asset_name: str = DEFAULT_PUBLISHED_MANIFEST_ASSET
|
|
artifacts: list[PublishedLlamaArtifact] = field(default_factory = list)
|
|
selection_log: list[str] = field(default_factory = list)
|
|
|
|
|
|
@dataclass
|
|
class LinuxCudaSelection:
|
|
attempts: list[AssetChoice]
|
|
selection_log: list[str]
|
|
|
|
@property
|
|
def primary(self) -> AssetChoice:
|
|
if not self.attempts:
|
|
raise RuntimeError("linux CUDA selection unexpectedly had no attempts")
|
|
return self.attempts[0]
|
|
|
|
|
|
@dataclass
|
|
class CudaRuntimePreference:
|
|
runtime_line: str | None
|
|
selection_log: list[str]
|
|
|
|
|
|
@dataclass(frozen = True)
|
|
class ApprovedArtifactHash:
|
|
asset_name: str
|
|
sha256: str
|
|
repo: str | None
|
|
kind: str | None
|
|
|
|
|
|
@dataclass
|
|
class ApprovedReleaseChecksums:
|
|
repo: str
|
|
release_tag: str
|
|
upstream_tag: str
|
|
source_repo: str | None = None
|
|
source_repo_url: str | None = None
|
|
source_ref_kind: str | None = None
|
|
requested_source_ref: str | None = None
|
|
resolved_source_ref: str | None = None
|
|
source_commit: str | None = None
|
|
source_commit_short: str | None = None
|
|
artifacts: dict[str, ApprovedArtifactHash] = field(default_factory = dict)
|
|
|
|
|
|
@dataclass(frozen = True)
|
|
class ResolvedPublishedRelease:
|
|
bundle: PublishedReleaseBundle
|
|
checksums: ApprovedReleaseChecksums
|
|
|
|
|
|
@dataclass(frozen = True)
|
|
class SourceBuildPlan:
|
|
source_url: str
|
|
source_ref: str
|
|
source_ref_kind: str
|
|
compatibility_upstream_tag: str
|
|
source_repo: str | None = None
|
|
source_repo_url: str | None = None
|
|
requested_source_ref: str | None = None
|
|
resolved_source_ref: str | None = None
|
|
source_commit: str | None = None
|
|
|
|
|
|
@dataclass(frozen = True)
|
|
class InstallReleasePlan:
|
|
requested_tag: str
|
|
llama_tag: str
|
|
release_tag: str
|
|
attempts: list[AssetChoice]
|
|
approved_checksums: ApprovedReleaseChecksums
|
|
|
|
|
|
class PrebuiltFallback(RuntimeError):
|
|
pass
|
|
|
|
|
|
class BusyInstallConflict(RuntimeError):
|
|
pass
|
|
|
|
|
|
class ExistingInstallSatisfied(RuntimeError):
|
|
def __init__(self, choice: AssetChoice, used_fallback: bool):
|
|
super().__init__(f"existing install already matches candidate {choice.name}")
|
|
self.choice = choice
|
|
self.used_fallback = used_fallback
|
|
|
|
|
|
def _os_error_messages(exc: BaseException) -> list[str]:
|
|
messages: list[str] = []
|
|
if isinstance(exc, OSError):
|
|
for value in (
|
|
getattr(exc, "strerror", None),
|
|
getattr(exc, "filename", None),
|
|
getattr(exc, "filename2", None),
|
|
):
|
|
if isinstance(value, str) and value:
|
|
messages.append(value)
|
|
text = str(exc)
|
|
if text:
|
|
messages.append(text)
|
|
return [message.lower() for message in messages if message]
|
|
|
|
|
|
def is_busy_lock_error(exc: BaseException) -> bool:
|
|
if isinstance(exc, BusyInstallConflict):
|
|
return True
|
|
if isinstance(exc, OSError):
|
|
if exc.errno in {
|
|
errno.EACCES,
|
|
errno.EBUSY,
|
|
errno.EPERM,
|
|
errno.ETXTBSY,
|
|
}:
|
|
return True
|
|
if getattr(exc, "winerror", None) in {5, 32, 145}:
|
|
return True
|
|
for message in _os_error_messages(exc):
|
|
if any(
|
|
needle in message
|
|
for needle in (
|
|
"access is denied",
|
|
"being used by another process",
|
|
"device or resource busy",
|
|
"permission denied",
|
|
"text file busy",
|
|
"file is in use",
|
|
"process cannot access the file",
|
|
"cannot create a file when that file already exists",
|
|
)
|
|
):
|
|
return True
|
|
return False
|
|
|
|
|
|
def log(message: str) -> None:
|
|
print(f"[llama-prebuilt] {message}", file = sys.stderr)
|
|
|
|
|
|
def log_lines(lines: Iterable[str]) -> None:
|
|
for line in lines:
|
|
log(line)
|
|
|
|
|
|
def parsed_hostname(url: str | None) -> str | None:
|
|
if not url:
|
|
return None
|
|
try:
|
|
hostname = urllib.parse.urlparse(url).hostname
|
|
except Exception:
|
|
return None
|
|
if not hostname:
|
|
return None
|
|
return hostname.lower()
|
|
|
|
|
|
def should_send_github_auth(url: str | None) -> bool:
|
|
return parsed_hostname(url) in GITHUB_AUTH_HOSTS
|
|
|
|
|
|
def auth_headers(url: str | None = None) -> dict[str, str]:
|
|
headers = {
|
|
"User-Agent": "unsloth-studio-llama-prebuilt",
|
|
}
|
|
token = os.environ.get("GH_TOKEN") or os.environ.get("GITHUB_TOKEN")
|
|
if token and should_send_github_auth(url):
|
|
headers["Authorization"] = f"Bearer {token}"
|
|
return headers
|
|
|
|
|
|
def github_api_headers(url: str | None = None) -> dict[str, str]:
|
|
return {
|
|
"Accept": "application/vnd.github+json",
|
|
**auth_headers(url),
|
|
}
|
|
|
|
|
|
def is_github_api_url(url: str | None) -> bool:
|
|
return parsed_hostname(url) == "api.github.com"
|
|
|
|
|
|
def is_retryable_url_error(exc: Exception) -> bool:
|
|
if isinstance(exc, urllib.error.HTTPError):
|
|
# GitHub returns 403 (not the standard 429) when the API rate
|
|
# limit is hit. Anonymous calls share a 60-req/hour bucket per
|
|
# runner IP, which CI fleets can exhaust trivially. Treat 403
|
|
# against api.github.com as retryable so we get one or two
|
|
# backoff cycles before the source-build fallback fires; honour
|
|
# Retry-After / X-RateLimit-Reset in sleep_backoff for accurate
|
|
# waits. Real 403s on other hosts (private artefact downloads,
|
|
# auth failures) stay non-retryable.
|
|
if exc.code == 403:
|
|
return is_github_api_url(getattr(exc, "url", None))
|
|
return exc.code in RETRYABLE_HTTP_STATUS
|
|
if isinstance(exc, urllib.error.URLError):
|
|
return True
|
|
if isinstance(exc, TimeoutError):
|
|
return True
|
|
if isinstance(exc, socket.timeout):
|
|
return True
|
|
return False
|
|
|
|
|
|
_RATE_LIMIT_WAIT_CAP_SECONDS = 60.0
|
|
|
|
|
|
def _http_error_retry_delay(exc: Exception) -> float | None:
|
|
"""Extract a recommended wait from rate-limit headers on a 403/429.
|
|
|
|
Returns None when no header is present or the indicated wait is
|
|
longer than _RATE_LIMIT_WAIT_CAP_SECONDS (in which case the caller
|
|
should not block on it -- the source-build fallback is faster).
|
|
"""
|
|
if not isinstance(exc, urllib.error.HTTPError):
|
|
return None
|
|
headers = getattr(exc, "headers", None)
|
|
if headers is None:
|
|
return None
|
|
retry_after = headers.get("Retry-After")
|
|
if retry_after and retry_after.strip().isdigit():
|
|
wait = float(retry_after.strip())
|
|
return wait if wait <= _RATE_LIMIT_WAIT_CAP_SECONDS else None
|
|
rate_reset = headers.get("X-RateLimit-Reset")
|
|
if rate_reset and rate_reset.strip().isdigit():
|
|
wait = float(rate_reset.strip()) - time.time()
|
|
if 0.0 < wait <= _RATE_LIMIT_WAIT_CAP_SECONDS:
|
|
return wait + 1.0 # +1s of slack so the bucket is fresh
|
|
return None
|
|
|
|
|
|
def sleep_backoff(
|
|
attempt: int,
|
|
*,
|
|
base_delay: float = HTTP_FETCH_BASE_DELAY_SECONDS,
|
|
exc: Exception | None = None,
|
|
) -> None:
|
|
delay = base_delay * (2 ** max(attempt - 1, 0))
|
|
header_delay = _http_error_retry_delay(exc) if exc is not None else None
|
|
if header_delay is not None:
|
|
delay = max(delay, header_delay)
|
|
delay += random.uniform(0.0, 0.2)
|
|
time.sleep(delay)
|
|
|
|
|
|
def atomic_write_bytes(destination: Path, data: bytes) -> None:
|
|
destination.parent.mkdir(parents = True, exist_ok = True)
|
|
with tempfile.NamedTemporaryFile(
|
|
prefix = destination.name + ".tmp-",
|
|
dir = destination.parent,
|
|
delete = False,
|
|
) as handle:
|
|
tmp_path = Path(handle.name)
|
|
handle.write(data)
|
|
handle.flush()
|
|
os.fsync(handle.fileno())
|
|
os.replace(tmp_path, destination)
|
|
|
|
|
|
def atomic_replace_from_tempfile(tmp_path: Path, destination: Path) -> None:
|
|
destination.parent.mkdir(parents = True, exist_ok = True)
|
|
os.replace(tmp_path, destination)
|
|
|
|
|
|
def source_archive_logical_name(upstream_tag: str) -> str:
|
|
return f"llama.cpp-source-{upstream_tag}.tar.gz"
|
|
|
|
|
|
def exact_source_archive_logical_name(source_commit: str) -> str:
|
|
return f"llama.cpp-source-commit-{source_commit}.tar.gz"
|
|
|
|
|
|
def sha256_file(path: Path) -> str:
|
|
digest = hashlib.sha256()
|
|
with path.open("rb") as handle:
|
|
for chunk in iter(lambda: handle.read(1024 * 1024), b""):
|
|
digest.update(chunk)
|
|
return digest.hexdigest()
|
|
|
|
|
|
def sha256_bytes(data: bytes) -> str:
|
|
return hashlib.sha256(data).hexdigest()
|
|
|
|
|
|
def normalize_sha256_digest(value: str | None) -> str | None:
|
|
if not isinstance(value, str) or not value:
|
|
return None
|
|
lowered = value.lower()
|
|
if lowered.startswith("sha256:"):
|
|
lowered = lowered.split(":", 1)[1]
|
|
if len(lowered) != 64 or any(ch not in "0123456789abcdef" for ch in lowered):
|
|
return None
|
|
return lowered
|
|
|
|
|
|
def normalize_source_ref_kind(value: str | None) -> str | None:
|
|
if not isinstance(value, str):
|
|
return None
|
|
normalized = value.strip().lower()
|
|
if normalized in {"tag", "branch", "pull", "commit", "custom"}:
|
|
return normalized
|
|
return None
|
|
|
|
|
|
def normalize_source_commit(value: str | None) -> str | None:
|
|
if not isinstance(value, str):
|
|
return None
|
|
normalized = value.strip().lower()
|
|
if len(normalized) < 7 or len(normalized) > 40:
|
|
return None
|
|
if any(ch not in "0123456789abcdef" for ch in normalized):
|
|
return None
|
|
return normalized
|
|
|
|
|
|
def validate_schema_version(payload: dict[str, Any], *, label: str) -> None:
|
|
schema_version = payload.get("schema_version")
|
|
if schema_version is None:
|
|
return
|
|
try:
|
|
normalized = int(schema_version)
|
|
except (TypeError, ValueError) as exc:
|
|
raise RuntimeError(f"{label} schema_version was not an integer") from exc
|
|
if normalized != 1:
|
|
raise RuntimeError(f"{label} schema_version={normalized} is unsupported")
|
|
|
|
|
|
def repo_slug_from_source(value: str | None) -> str | None:
|
|
if not isinstance(value, str):
|
|
return None
|
|
normalized = value.strip()
|
|
if not normalized:
|
|
return None
|
|
normalized = normalized.removesuffix(".git")
|
|
if normalized.startswith("https://github.com/"):
|
|
slug = normalized[len("https://github.com/") :]
|
|
elif normalized.startswith("http://github.com/"):
|
|
slug = normalized[len("http://github.com/") :]
|
|
elif normalized.startswith("git@github.com:"):
|
|
slug = normalized[len("git@github.com:") :]
|
|
else:
|
|
slug = normalized
|
|
slug = slug.strip("/")
|
|
parts = slug.split("/")
|
|
if len(parts) != 2 or not all(parts):
|
|
return None
|
|
return f"{parts[0]}/{parts[1]}"
|
|
|
|
|
|
def source_url_from_repo_slug(repo_slug: str | None) -> str | None:
|
|
if not isinstance(repo_slug, str) or not repo_slug:
|
|
return None
|
|
return f"https://github.com/{repo_slug}"
|
|
|
|
|
|
def source_repo_clone_url(repo: str | None, repo_url: str | None) -> str | None:
|
|
if isinstance(repo_url, str) and repo_url.strip():
|
|
return repo_url.strip().removesuffix(".git")
|
|
return source_url_from_repo_slug(repo_slug_from_source(repo))
|
|
|
|
|
|
def infer_source_ref_kind(ref: str | None) -> str:
|
|
if not isinstance(ref, str):
|
|
return "tag"
|
|
normalized = ref.strip()
|
|
lowered = normalized.lower()
|
|
if not normalized:
|
|
return "tag"
|
|
if lowered.startswith("refs/pull/") or lowered.startswith("pull/"):
|
|
return "pull"
|
|
if (
|
|
lowered.startswith("refs/heads/")
|
|
or lowered in {"main", "master", "head"}
|
|
or lowered.startswith("origin/")
|
|
):
|
|
return "branch"
|
|
normalized_commit = normalize_source_commit(normalized)
|
|
if normalized_commit is not None:
|
|
return "commit"
|
|
return "tag"
|
|
|
|
|
|
def normalized_ref_aliases(ref: str | None) -> set[str]:
|
|
if not isinstance(ref, str):
|
|
return set()
|
|
normalized = ref.strip()
|
|
if not normalized:
|
|
return set()
|
|
aliases = {normalized}
|
|
lowered = normalized.lower()
|
|
commit = normalize_source_commit(normalized)
|
|
if commit is not None:
|
|
aliases.add(commit)
|
|
if lowered.startswith("refs/heads/"):
|
|
aliases.add(normalized.split("/", 2)[2])
|
|
elif "/" not in normalized and infer_source_ref_kind(normalized) == "branch":
|
|
aliases.add(f"refs/heads/{normalized}")
|
|
if lowered.startswith("refs/pull/"):
|
|
aliases.add(normalized.removeprefix("refs/"))
|
|
elif lowered.startswith("pull/"):
|
|
aliases.add(f"refs/{normalized}")
|
|
return aliases
|
|
|
|
|
|
def refs_match(candidate_ref: str | None, requested_ref: str | None) -> bool:
|
|
candidate_aliases = normalized_ref_aliases(candidate_ref)
|
|
requested_aliases = normalized_ref_aliases(requested_ref)
|
|
if not candidate_aliases or not requested_aliases:
|
|
return False
|
|
if candidate_aliases & requested_aliases:
|
|
return True
|
|
candidate_commit = normalize_source_commit(candidate_ref)
|
|
requested_commit = normalize_source_commit(requested_ref)
|
|
if candidate_commit and requested_commit:
|
|
return candidate_commit.startswith(
|
|
requested_commit
|
|
) or requested_commit.startswith(candidate_commit)
|
|
return False
|
|
|
|
|
|
def checkout_friendly_ref(ref_kind: str | None, ref: str | None) -> str | None:
|
|
"""Normalize a source ref to a form that ``git clone --branch`` accepts.
|
|
|
|
Fully qualified branch refs like ``refs/heads/main`` are stripped to
|
|
``main``; tag refs like ``refs/tags/b8508`` are stripped to ``b8508``.
|
|
Pull refs like ``refs/pull/123/head`` are left as-is since they are
|
|
always fetched explicitly rather than cloned with ``--branch``.
|
|
"""
|
|
if not isinstance(ref, str) or not ref:
|
|
return ref
|
|
lowered = ref.lower()
|
|
if ref_kind == "branch" and lowered.startswith("refs/heads/"):
|
|
return ref.split("/", 2)[2]
|
|
if ref_kind == "tag" and lowered.startswith("refs/tags/"):
|
|
return ref.split("/", 2)[2]
|
|
return ref
|
|
|
|
|
|
def windows_cuda_upstream_asset_names(llama_tag: str, runtime: str) -> list[str]:
|
|
return [
|
|
f"llama-{llama_tag}-bin-win-cuda-{runtime}-x64.zip",
|
|
f"cudart-llama-bin-win-cuda-{runtime}-x64.zip",
|
|
]
|
|
|
|
|
|
def windows_cuda_asset_aliases(
|
|
asset_name: str,
|
|
*,
|
|
compatibility_tag: str | None = None,
|
|
) -> list[str]:
|
|
aliases: list[str] = []
|
|
legacy_match = re.fullmatch(
|
|
r"llama-(?P<tag>[^/]+)-bin-win-cuda-(?P<runtime>\d+\.\d+)-x64\.zip",
|
|
asset_name,
|
|
)
|
|
if legacy_match:
|
|
runtime = legacy_match.group("runtime")
|
|
aliases.append(f"cudart-llama-bin-win-cuda-{runtime}-x64.zip")
|
|
if compatibility_tag:
|
|
aliases.append(f"llama-{compatibility_tag}-bin-win-cuda-{runtime}-x64.zip")
|
|
return aliases
|
|
|
|
current_match = re.fullmatch(
|
|
r"cudart-llama-bin-win-cuda-(?P<runtime>\d+\.\d+)-x64\.zip",
|
|
asset_name,
|
|
)
|
|
if current_match and compatibility_tag:
|
|
runtime = current_match.group("runtime")
|
|
aliases.append(f"llama-{compatibility_tag}-bin-win-cuda-{runtime}-x64.zip")
|
|
return aliases
|
|
|
|
|
|
def format_byte_count(num_bytes: float) -> str:
|
|
units = ["B", "KiB", "MiB", "GiB", "TiB"]
|
|
value = float(num_bytes)
|
|
for unit in units:
|
|
if abs(value) < 1024.0 or unit == units[-1]:
|
|
if unit == "B":
|
|
return f"{int(value)} {unit}"
|
|
return f"{value:.1f} {unit}"
|
|
value /= 1024.0
|
|
return f"{num_bytes:.1f} B"
|
|
|
|
|
|
class DownloadProgress:
|
|
def __init__(self, label: str, total_bytes: int | None) -> None:
|
|
self.label = label
|
|
self.total_bytes = total_bytes if total_bytes and total_bytes > 0 else None
|
|
self.start_time = time.monotonic()
|
|
self.last_emit = 0.0
|
|
term_ok = os.environ.get("TERM", "").lower() != "dumb"
|
|
self.stream = (
|
|
sys.stderr
|
|
if sys.stderr.isatty()
|
|
else sys.stdout
|
|
if sys.stdout.isatty()
|
|
else sys.stderr
|
|
)
|
|
self.is_tty = term_ok and self.stream.isatty()
|
|
self.completed = False
|
|
self.last_milestone_percent = -1
|
|
self.last_milestone_bytes = 0
|
|
self.has_rendered_tty_progress = False
|
|
|
|
def _render(self, downloaded_bytes: int, *, final: bool = False) -> str:
|
|
elapsed = max(time.monotonic() - self.start_time, 1e-6)
|
|
speed = downloaded_bytes / elapsed
|
|
speed_text = f"{format_byte_count(speed)}/s"
|
|
if self.total_bytes is not None:
|
|
percent = min(100.0, (downloaded_bytes / self.total_bytes) * 100.0)
|
|
return (
|
|
f"{self.label}: {percent:5.1f}% "
|
|
f"({format_byte_count(downloaded_bytes)}/{format_byte_count(self.total_bytes)}) "
|
|
f"at {speed_text}"
|
|
)
|
|
if final:
|
|
return f"{self.label}: {format_byte_count(downloaded_bytes)} downloaded at {speed_text}"
|
|
return f"{self.label}: {format_byte_count(downloaded_bytes)} downloaded at {speed_text}"
|
|
|
|
def update(self, downloaded_bytes: int) -> None:
|
|
now = time.monotonic()
|
|
if self.is_tty:
|
|
elapsed = now - self.start_time
|
|
if not self.has_rendered_tty_progress:
|
|
if (
|
|
self.total_bytes is not None
|
|
and downloaded_bytes >= self.total_bytes
|
|
):
|
|
return
|
|
if elapsed < TTY_PROGRESS_START_DELAY_SECONDS:
|
|
return
|
|
min_interval = 0.2
|
|
if (
|
|
self.has_rendered_tty_progress
|
|
and not self.completed
|
|
and (now - self.last_emit) < min_interval
|
|
):
|
|
return
|
|
self.last_emit = now
|
|
line = self._render(downloaded_bytes)
|
|
self.stream.write("\r\033[K" + line)
|
|
self.stream.flush()
|
|
self.has_rendered_tty_progress = True
|
|
return
|
|
|
|
should_emit = False
|
|
if self.total_bytes is not None:
|
|
percent = int((downloaded_bytes * 100) / max(self.total_bytes, 1))
|
|
milestone_percent = min((percent // 25) * 25, 100)
|
|
if (
|
|
milestone_percent > self.last_milestone_percent
|
|
and milestone_percent < 100
|
|
):
|
|
self.last_milestone_percent = milestone_percent
|
|
should_emit = True
|
|
else:
|
|
byte_step = 25 * 1024 * 1024
|
|
if (
|
|
downloaded_bytes - self.last_milestone_bytes >= byte_step
|
|
and (now - self.last_emit) >= 5.0
|
|
):
|
|
self.last_milestone_bytes = downloaded_bytes
|
|
should_emit = True
|
|
|
|
if not should_emit:
|
|
return
|
|
|
|
self.last_emit = now
|
|
self.stream.write(self._render(downloaded_bytes) + "\n")
|
|
self.stream.flush()
|
|
|
|
def finish(self, downloaded_bytes: int) -> None:
|
|
self.completed = True
|
|
line = self._render(downloaded_bytes, final = True)
|
|
if self.is_tty:
|
|
if not self.has_rendered_tty_progress:
|
|
return
|
|
self.stream.write("\r\033[K")
|
|
else:
|
|
self.stream.write(line + "\n")
|
|
self.stream.flush()
|
|
|
|
|
|
def download_label_from_url(url: str) -> str:
|
|
name = Path(urllib.parse.urlparse(url).path).name
|
|
return name or url
|
|
|
|
|
|
def download_bytes(
|
|
url: str,
|
|
*,
|
|
timeout: int = 120,
|
|
attempts: int = HTTP_FETCH_ATTEMPTS,
|
|
headers: dict[str, str] | None = None,
|
|
progress_label: str | None = None,
|
|
) -> bytes:
|
|
last_exc: Exception | None = None
|
|
for attempt in range(1, attempts + 1):
|
|
try:
|
|
request = urllib.request.Request(url, headers = headers or auth_headers(url))
|
|
with urllib.request.urlopen(request, timeout = timeout) as response:
|
|
total_bytes: int | None = None
|
|
content_length = response.headers.get("Content-Length")
|
|
if content_length and content_length.isdigit():
|
|
total_bytes = int(content_length)
|
|
progress = (
|
|
DownloadProgress(progress_label, total_bytes)
|
|
if progress_label
|
|
else None
|
|
)
|
|
data = bytearray()
|
|
while True:
|
|
chunk = response.read(1024 * 1024)
|
|
if not chunk:
|
|
break
|
|
data.extend(chunk)
|
|
if progress is not None:
|
|
progress.update(len(data))
|
|
if progress is not None:
|
|
progress.finish(len(data))
|
|
return bytes(data)
|
|
except Exception as exc:
|
|
last_exc = exc
|
|
if attempt >= attempts or not is_retryable_url_error(exc):
|
|
raise
|
|
log(f"fetch failed ({attempt}/{attempts}) for {url}: {exc}; retrying")
|
|
sleep_backoff(attempt, exc = exc)
|
|
assert last_exc is not None
|
|
raise last_exc
|
|
|
|
|
|
def fetch_json(url: str) -> Any:
|
|
attempts = JSON_FETCH_ATTEMPTS if is_github_api_url(url) else 1
|
|
last_decode_exc: Exception | None = None
|
|
for attempt in range(1, attempts + 1):
|
|
try:
|
|
data = download_bytes(
|
|
url,
|
|
timeout = 30,
|
|
headers = github_api_headers(url)
|
|
if is_github_api_url(url)
|
|
else auth_headers(url),
|
|
)
|
|
except urllib.error.HTTPError as exc:
|
|
if exc.code == 403 and is_github_api_url(url):
|
|
hint = ""
|
|
if not (os.environ.get("GH_TOKEN") or os.environ.get("GITHUB_TOKEN")):
|
|
hint = (
|
|
"; set GH_TOKEN or GITHUB_TOKEN to avoid GitHub API rate limits"
|
|
)
|
|
raise RuntimeError(f"GitHub API returned 403 for {url}{hint}") from exc
|
|
raise
|
|
if not data:
|
|
last_decode_exc = RuntimeError(f"downloaded empty JSON payload from {url}")
|
|
else:
|
|
try:
|
|
payload = json.loads(data.decode("utf-8"))
|
|
except (UnicodeDecodeError, json.JSONDecodeError) as exc:
|
|
last_decode_exc = RuntimeError(
|
|
f"downloaded invalid JSON from {url}: {exc}"
|
|
)
|
|
else:
|
|
if not isinstance(payload, dict) and not isinstance(payload, list):
|
|
raise RuntimeError(
|
|
f"downloaded unexpected JSON type from {url}: {type(payload).__name__}"
|
|
)
|
|
return payload
|
|
if attempt >= attempts:
|
|
assert last_decode_exc is not None
|
|
raise last_decode_exc
|
|
log(f"json fetch failed ({attempt}/{attempts}) for {url}; retrying")
|
|
sleep_backoff(attempt)
|
|
assert last_decode_exc is not None
|
|
raise last_decode_exc
|
|
|
|
|
|
def download_file(url: str, destination: Path) -> None:
|
|
destination.parent.mkdir(parents = True, exist_ok = True)
|
|
last_exc: Exception | None = None
|
|
for attempt in range(1, HTTP_FETCH_ATTEMPTS + 1):
|
|
tmp_path: Path | None = None
|
|
try:
|
|
request = urllib.request.Request(url, headers = auth_headers(url))
|
|
with tempfile.NamedTemporaryFile(
|
|
prefix = destination.name + ".tmp-",
|
|
dir = destination.parent,
|
|
delete = False,
|
|
) as handle:
|
|
tmp_path = Path(handle.name)
|
|
with urllib.request.urlopen(request, timeout = 120) as response:
|
|
total_bytes: int | None = None
|
|
content_length = response.headers.get("Content-Length")
|
|
if content_length and content_length.isdigit():
|
|
total_bytes = int(content_length)
|
|
progress = DownloadProgress(
|
|
f"Downloading {destination.name}", total_bytes
|
|
)
|
|
downloaded_bytes = 0
|
|
while True:
|
|
chunk = response.read(1024 * 1024)
|
|
if not chunk:
|
|
break
|
|
handle.write(chunk)
|
|
downloaded_bytes += len(chunk)
|
|
progress.update(downloaded_bytes)
|
|
progress.finish(downloaded_bytes)
|
|
handle.flush()
|
|
os.fsync(handle.fileno())
|
|
if not tmp_path.exists() or tmp_path.stat().st_size == 0:
|
|
raise RuntimeError(f"downloaded empty file from {url}")
|
|
atomic_replace_from_tempfile(tmp_path, destination)
|
|
return
|
|
except Exception as exc:
|
|
last_exc = exc
|
|
if tmp_path is not None:
|
|
try:
|
|
tmp_path.unlink(missing_ok = True)
|
|
except Exception:
|
|
pass
|
|
if attempt >= HTTP_FETCH_ATTEMPTS or not is_retryable_url_error(exc):
|
|
raise
|
|
log(
|
|
f"download failed ({attempt}/{HTTP_FETCH_ATTEMPTS}) for {url}: {exc}; retrying"
|
|
)
|
|
sleep_backoff(attempt, exc = exc)
|
|
assert last_exc is not None
|
|
raise last_exc
|
|
|
|
|
|
def download_file_verified(
|
|
url: str,
|
|
destination: Path,
|
|
*,
|
|
expected_sha256: str | None,
|
|
label: str,
|
|
) -> None:
|
|
normalized_expected = normalize_sha256_digest(expected_sha256)
|
|
if not normalized_expected:
|
|
download_file(url, destination)
|
|
log(
|
|
f"downloaded {label} without a published sha256; relying on install validation"
|
|
)
|
|
return
|
|
|
|
for attempt in range(1, 3):
|
|
download_file(url, destination)
|
|
actual_sha256 = sha256_file(destination)
|
|
if actual_sha256 == normalized_expected:
|
|
log(f"verified {label} sha256={actual_sha256}")
|
|
return
|
|
|
|
log(
|
|
f"{label} checksum mismatch on attempt {attempt}/2: "
|
|
f"expected={normalized_expected} actual={actual_sha256}"
|
|
)
|
|
destination.unlink(missing_ok = True)
|
|
if attempt == 2:
|
|
raise PrebuiltFallback(
|
|
f"{label} checksum mismatch after retry: expected={normalized_expected} actual={actual_sha256}"
|
|
)
|
|
log(f"retrying {label} download after checksum mismatch")
|
|
|
|
|
|
def upstream_source_archive_urls(tag: str) -> list[str]:
|
|
encoded_tag = urllib.parse.quote(tag, safe = "")
|
|
return [
|
|
f"https://codeload.github.com/{UPSTREAM_REPO}/tar.gz/refs/tags/{encoded_tag}",
|
|
f"https://github.com/{UPSTREAM_REPO}/archive/refs/tags/{encoded_tag}.tar.gz",
|
|
]
|
|
|
|
|
|
def commit_source_archive_urls(repo: str, source_commit: str) -> list[str]:
|
|
encoded_commit = urllib.parse.quote(source_commit, safe = "")
|
|
return [
|
|
f"https://codeload.github.com/{repo}/tar.gz/{encoded_commit}",
|
|
f"https://github.com/{repo}/archive/{encoded_commit}.tar.gz",
|
|
]
|
|
|
|
|
|
def github_release_assets(repo: str, tag: str) -> dict[str, str]:
|
|
payload = fetch_json(
|
|
f"https://api.github.com/repos/{repo}/releases/tags/{urllib.parse.quote(tag, safe = '')}"
|
|
)
|
|
if not isinstance(payload, dict):
|
|
raise RuntimeError(f"unexpected release payload for {repo}@{tag}")
|
|
return release_asset_map(payload)
|
|
|
|
|
|
def github_release(repo: str, tag: str) -> dict[str, Any]:
|
|
payload = fetch_json(
|
|
f"https://api.github.com/repos/{repo}/releases/tags/{urllib.parse.quote(tag, safe = '')}"
|
|
)
|
|
if not isinstance(payload, dict):
|
|
raise RuntimeError(f"unexpected release payload for {repo}@{tag}")
|
|
return payload
|
|
|
|
|
|
def github_releases(
|
|
repo: str,
|
|
*,
|
|
per_page: int = 100,
|
|
max_pages: int = 0,
|
|
) -> list[dict[str, Any]]:
|
|
releases: list[dict[str, Any]] = []
|
|
page = 1
|
|
while True:
|
|
payload = fetch_json(
|
|
f"https://api.github.com/repos/{repo}/releases?per_page={per_page}&page={page}"
|
|
)
|
|
if not isinstance(payload, list):
|
|
raise RuntimeError(f"unexpected releases payload for {repo}")
|
|
page_items = [item for item in payload if isinstance(item, dict)]
|
|
releases.extend(page_items)
|
|
if len(payload) < per_page:
|
|
break
|
|
page += 1
|
|
if max_pages > 0 and page > max_pages:
|
|
break
|
|
return releases
|
|
|
|
|
|
def latest_upstream_release_tag() -> str:
|
|
payload = fetch_json(UPSTREAM_RELEASES_API)
|
|
tag = payload.get("tag_name")
|
|
if not isinstance(tag, str) or not tag:
|
|
raise RuntimeError(
|
|
f"latest release tag was missing from {UPSTREAM_RELEASES_API}"
|
|
)
|
|
return tag
|
|
|
|
|
|
def is_release_tag_like(value: str | None) -> bool:
|
|
return isinstance(value, str) and bool(re.fullmatch(r"b\d+", value.strip()))
|
|
|
|
|
|
def release_time_sort_key(release: dict[str, Any]) -> tuple[str, int]:
|
|
published_at = release.get("published_at")
|
|
created_at = release.get("created_at")
|
|
release_id = release.get("id")
|
|
timestamp = (
|
|
published_at
|
|
if isinstance(published_at, str) and published_at
|
|
else created_at
|
|
if isinstance(created_at, str) and created_at
|
|
else ""
|
|
)
|
|
try:
|
|
normalized_id = int(release_id)
|
|
except (TypeError, ValueError):
|
|
normalized_id = 0
|
|
return (timestamp, normalized_id)
|
|
|
|
|
|
def iter_release_payloads_by_time(
|
|
repo: str,
|
|
published_release_tag: str = "",
|
|
requested_tag: str = "",
|
|
) -> Iterable[dict[str, Any]]:
|
|
if published_release_tag:
|
|
yield github_release(repo, published_release_tag)
|
|
return
|
|
|
|
if (
|
|
requested_tag
|
|
and requested_tag != "latest"
|
|
and is_release_tag_like(requested_tag)
|
|
):
|
|
try:
|
|
yield github_release(repo, requested_tag)
|
|
return
|
|
except urllib.error.HTTPError as exc:
|
|
if exc.code == 404:
|
|
log(
|
|
f"release tag {requested_tag} not found in {repo}; scanning recent releases"
|
|
)
|
|
else:
|
|
raise
|
|
except Exception:
|
|
raise
|
|
|
|
releases = [
|
|
release
|
|
for release in github_releases(
|
|
repo, max_pages = DEFAULT_GITHUB_RELEASE_SCAN_MAX_PAGES
|
|
)
|
|
if isinstance(release, dict)
|
|
and not release.get("draft")
|
|
and not release.get("prerelease")
|
|
]
|
|
releases.sort(key = release_time_sort_key, reverse = True)
|
|
for release in releases:
|
|
yield release
|
|
|
|
|
|
def direct_release_matches_request(
|
|
*, release_tag: str, llama_tag: str, requested_tag: str
|
|
) -> bool:
|
|
if requested_tag == "latest":
|
|
return True
|
|
for candidate in (release_tag, llama_tag):
|
|
if refs_match(candidate, requested_tag):
|
|
return True
|
|
return False
|
|
|
|
|
|
def synthetic_checksums_for_release(
|
|
repo: str, release_tag: str, upstream_tag: str
|
|
) -> ApprovedReleaseChecksums:
|
|
return ApprovedReleaseChecksums(
|
|
repo = repo,
|
|
release_tag = release_tag,
|
|
upstream_tag = upstream_tag,
|
|
artifacts = {},
|
|
)
|
|
|
|
|
|
def parse_direct_linux_release_bundle(
|
|
repo: str, release: dict[str, Any]
|
|
) -> PublishedReleaseBundle | None:
|
|
release_tag = release.get("tag_name")
|
|
if not isinstance(release_tag, str) or not release_tag:
|
|
return None
|
|
|
|
assets = release_asset_map(release)
|
|
artifacts: list[PublishedLlamaArtifact] = []
|
|
inferred_labels: list[str] = []
|
|
|
|
linux_asset_re = re.compile(
|
|
r"^app-(?P<label>.+)-(?P<target>linux-x64(?:-cpu)?|linux-x64-(?:cuda12|cuda13)-(?:older|newer|portable))\.tar\.gz$"
|
|
)
|
|
for asset_name in sorted(assets):
|
|
match = linux_asset_re.fullmatch(asset_name)
|
|
if not match:
|
|
continue
|
|
inferred_labels.append(match.group("label"))
|
|
target = match.group("target")
|
|
if target in {"linux-x64", "linux-x64-cpu"}:
|
|
artifacts.append(
|
|
PublishedLlamaArtifact(
|
|
asset_name = asset_name,
|
|
install_kind = "linux-cpu",
|
|
runtime_line = None,
|
|
coverage_class = None,
|
|
supported_sms = [],
|
|
min_sm = None,
|
|
max_sm = None,
|
|
bundle_profile = None,
|
|
rank = 1000,
|
|
)
|
|
)
|
|
continue
|
|
|
|
bundle_profile = target.removeprefix("linux-x64-")
|
|
profile = DIRECT_LINUX_BUNDLE_PROFILES.get(bundle_profile)
|
|
if profile is None:
|
|
continue
|
|
artifacts.append(
|
|
PublishedLlamaArtifact(
|
|
asset_name = asset_name,
|
|
install_kind = "linux-cuda",
|
|
runtime_line = str(profile["runtime_line"]),
|
|
coverage_class = str(profile["coverage_class"]),
|
|
supported_sms = [str(value) for value in profile["supported_sms"]],
|
|
min_sm = int(profile["min_sm"]),
|
|
max_sm = int(profile["max_sm"]),
|
|
bundle_profile = bundle_profile,
|
|
rank = int(profile["rank"]),
|
|
)
|
|
)
|
|
|
|
if not artifacts:
|
|
return None
|
|
|
|
upstream_tag = (
|
|
release_tag
|
|
if is_release_tag_like(release_tag)
|
|
else inferred_labels[0]
|
|
if len(set(inferred_labels)) == 1 and inferred_labels
|
|
else release_tag
|
|
)
|
|
selection_log = [
|
|
f"published_release: repo={repo}",
|
|
f"published_release: tag={release_tag}",
|
|
f"published_release: upstream_tag={upstream_tag}",
|
|
"published_release: direct_asset_scan=linux",
|
|
]
|
|
return PublishedReleaseBundle(
|
|
repo = repo,
|
|
release_tag = release_tag,
|
|
upstream_tag = upstream_tag,
|
|
assets = assets,
|
|
manifest_asset_name = DEFAULT_PUBLISHED_MANIFEST_ASSET,
|
|
artifacts = artifacts,
|
|
selection_log = selection_log,
|
|
)
|
|
|
|
|
|
def direct_linux_release_plan(
|
|
release: dict[str, Any],
|
|
host: HostInfo,
|
|
repo: str,
|
|
requested_tag: str,
|
|
) -> InstallReleasePlan | None:
|
|
bundle = parse_direct_linux_release_bundle(repo, release)
|
|
if bundle is None:
|
|
return None
|
|
if not direct_release_matches_request(
|
|
release_tag = bundle.release_tag,
|
|
llama_tag = bundle.upstream_tag,
|
|
requested_tag = requested_tag,
|
|
):
|
|
return None
|
|
|
|
attempts: list[AssetChoice] = []
|
|
if host.has_usable_nvidia:
|
|
selection = linux_cuda_choice_from_release(host, bundle)
|
|
if selection is not None:
|
|
attempts.extend(selection.attempts)
|
|
cpu_choice = published_asset_choice_for_kind(bundle, "linux-cpu")
|
|
if cpu_choice is not None:
|
|
attempts.append(cpu_choice)
|
|
if not attempts:
|
|
raise PrebuiltFallback("no compatible Linux prebuilt asset was found")
|
|
approved_checksums = synthetic_checksums_for_release(
|
|
repo,
|
|
bundle.release_tag,
|
|
bundle.upstream_tag,
|
|
)
|
|
resolved_upstream_tag = bundle.upstream_tag
|
|
if DEFAULT_PUBLISHED_SHA256_ASSET in bundle.assets and not is_release_tag_like(
|
|
bundle.upstream_tag
|
|
):
|
|
approved_checksums = load_approved_release_checksums(repo, bundle.release_tag)
|
|
# Require exact source provenance for branch/pull/commit releases.
|
|
# Mirrors validated_checksums_for_bundle so incomplete metadata
|
|
# fails closed instead of degrading to the legacy branch-as-tag
|
|
# source hydration path that this PR is meant to eliminate.
|
|
if (
|
|
not approved_checksums.source_commit
|
|
or exact_source_archive_hash(approved_checksums) is None
|
|
or source_clone_url_from_checksums(approved_checksums) is None
|
|
):
|
|
raise PrebuiltFallback(
|
|
f"approved checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET} for "
|
|
f"{repo}@{bundle.release_tag} did not contain exact source provenance"
|
|
)
|
|
attempts = apply_approved_hashes(attempts, approved_checksums)
|
|
return InstallReleasePlan(
|
|
requested_tag = requested_tag,
|
|
llama_tag = resolved_upstream_tag,
|
|
release_tag = bundle.release_tag,
|
|
attempts = attempts,
|
|
approved_checksums = approved_checksums,
|
|
)
|
|
|
|
|
|
def direct_upstream_release_plan(
|
|
release: dict[str, Any],
|
|
host: HostInfo,
|
|
repo: str,
|
|
requested_tag: str,
|
|
) -> InstallReleasePlan | None:
|
|
release_tag = release.get("tag_name")
|
|
if not isinstance(release_tag, str) or not release_tag:
|
|
return None
|
|
if not direct_release_matches_request(
|
|
release_tag = release_tag,
|
|
llama_tag = release_tag,
|
|
requested_tag = requested_tag,
|
|
):
|
|
return None
|
|
|
|
assets = release_asset_map(release)
|
|
attempts: list[AssetChoice] = []
|
|
if host.is_windows and host.is_x86_64:
|
|
if host.has_usable_nvidia:
|
|
torch_preference = detect_torch_cuda_runtime_preference(host)
|
|
attempts.extend(
|
|
windows_cuda_attempts(
|
|
host,
|
|
release_tag,
|
|
assets,
|
|
torch_preference.runtime_line,
|
|
torch_preference.selection_log,
|
|
)
|
|
)
|
|
elif host.has_rocm:
|
|
hip_asset = f"llama-{release_tag}-bin-win-hip-radeon-x64.zip"
|
|
hip_url = assets.get(hip_asset)
|
|
if hip_url:
|
|
attempts.append(
|
|
AssetChoice(
|
|
repo = repo,
|
|
tag = release_tag,
|
|
name = hip_asset,
|
|
url = hip_url,
|
|
source_label = "upstream",
|
|
install_kind = "windows-hip",
|
|
)
|
|
)
|
|
cpu_asset = f"llama-{release_tag}-bin-win-cpu-x64.zip"
|
|
cpu_url = assets.get(cpu_asset)
|
|
if cpu_url:
|
|
attempts.append(
|
|
AssetChoice(
|
|
repo = repo,
|
|
tag = release_tag,
|
|
name = cpu_asset,
|
|
url = cpu_url,
|
|
source_label = "upstream",
|
|
install_kind = "windows-cpu",
|
|
)
|
|
)
|
|
elif host.is_windows and host.is_arm64:
|
|
# Upstream ggml-org/llama.cpp ships llama-bNNNN-bin-win-cpu-arm64.zip
|
|
# (visible in the b9334 release manifest). Without this branch the
|
|
# selector returned 0 attempts and the installer fell back to a
|
|
# source build on every Windows ARM64 host.
|
|
cpu_asset = f"llama-{release_tag}-bin-win-cpu-arm64.zip"
|
|
cpu_url = assets.get(cpu_asset)
|
|
if cpu_url:
|
|
attempts.append(
|
|
AssetChoice(
|
|
repo = repo,
|
|
tag = release_tag,
|
|
name = cpu_asset,
|
|
url = cpu_url,
|
|
source_label = "upstream",
|
|
install_kind = "windows-arm64",
|
|
)
|
|
)
|
|
elif host.is_macos and host.is_arm64:
|
|
asset_name = f"llama-{release_tag}-bin-macos-arm64.tar.gz"
|
|
asset_url = assets.get(asset_name)
|
|
if asset_url:
|
|
attempts.append(
|
|
AssetChoice(
|
|
repo = repo,
|
|
tag = release_tag,
|
|
name = asset_name,
|
|
url = asset_url,
|
|
source_label = "upstream",
|
|
install_kind = "macos-arm64",
|
|
)
|
|
)
|
|
elif host.is_macos and host.is_x86_64:
|
|
asset_name = f"llama-{release_tag}-bin-macos-x64.tar.gz"
|
|
asset_url = assets.get(asset_name)
|
|
if asset_url:
|
|
attempts.append(
|
|
AssetChoice(
|
|
repo = repo,
|
|
tag = release_tag,
|
|
name = asset_name,
|
|
url = asset_url,
|
|
source_label = "upstream",
|
|
install_kind = "macos-x64",
|
|
)
|
|
)
|
|
elif host.is_linux and host.is_x86_64 and not host.has_usable_nvidia:
|
|
asset_name = f"llama-{release_tag}-bin-ubuntu-x64.tar.gz"
|
|
asset_url = assets.get(asset_name)
|
|
if asset_url:
|
|
attempts.append(
|
|
AssetChoice(
|
|
repo = repo,
|
|
tag = release_tag,
|
|
name = asset_name,
|
|
url = asset_url,
|
|
source_label = "upstream",
|
|
install_kind = "linux-cpu",
|
|
)
|
|
)
|
|
elif host.is_linux and host.is_arm64 and not host.has_usable_nvidia:
|
|
# Upstream ggml-org/llama.cpp ships llama-bNNNN-bin-ubuntu-arm64.tar.gz
|
|
# (visible in the b9334 release manifest). Without this branch the
|
|
# selector returned 0 attempts and the installer fell back to a
|
|
# source build on every Linux ARM64 host (DGX Spark, Ampere
|
|
# Altra, GitHub-hosted ubuntu-24.04-arm runners, etc.).
|
|
asset_name = f"llama-{release_tag}-bin-ubuntu-arm64.tar.gz"
|
|
asset_url = assets.get(asset_name)
|
|
if asset_url:
|
|
attempts.append(
|
|
AssetChoice(
|
|
repo = repo,
|
|
tag = release_tag,
|
|
name = asset_name,
|
|
url = asset_url,
|
|
source_label = "upstream",
|
|
install_kind = "linux-arm64",
|
|
)
|
|
)
|
|
if not attempts:
|
|
raise PrebuiltFallback("no compatible upstream prebuilt asset was found")
|
|
return InstallReleasePlan(
|
|
requested_tag = requested_tag,
|
|
llama_tag = release_tag,
|
|
release_tag = release_tag,
|
|
attempts = attempts,
|
|
approved_checksums = synthetic_checksums_for_release(
|
|
repo,
|
|
release_tag,
|
|
release_tag,
|
|
),
|
|
)
|
|
|
|
|
|
def resolve_simple_install_release_plans(
|
|
llama_tag: str,
|
|
host: HostInfo,
|
|
published_repo: str,
|
|
published_release_tag: str,
|
|
*,
|
|
max_release_fallbacks: int = DEFAULT_MAX_PREBUILT_RELEASE_FALLBACKS,
|
|
) -> tuple[str, list[InstallReleasePlan]]:
|
|
repo = published_repo or DEFAULT_PUBLISHED_REPO
|
|
requested_tag = normalized_requested_llama_tag(llama_tag)
|
|
allow_older_release_fallback = (
|
|
requested_tag == "latest" and not published_release_tag
|
|
)
|
|
release_limit = max(1, max_release_fallbacks)
|
|
plans: list[InstallReleasePlan] = []
|
|
last_error: PrebuiltFallback | None = None
|
|
|
|
try:
|
|
releases = iter_release_payloads_by_time(
|
|
repo, published_release_tag, requested_tag
|
|
)
|
|
for release in releases:
|
|
try:
|
|
if host.is_linux and repo == "unslothai/llama.cpp":
|
|
plan = direct_linux_release_plan(release, host, repo, requested_tag)
|
|
else:
|
|
plan = direct_upstream_release_plan(
|
|
release, host, repo, requested_tag
|
|
)
|
|
if plan is None:
|
|
continue
|
|
except PrebuiltFallback as exc:
|
|
last_error = exc
|
|
if not allow_older_release_fallback:
|
|
raise
|
|
release_tag = release.get("tag_name") or "unknown"
|
|
log(
|
|
"published release skipped for install planning: "
|
|
f"{repo}@{release_tag} ({exc})"
|
|
)
|
|
continue
|
|
|
|
plans.append(plan)
|
|
if not allow_older_release_fallback or len(plans) >= release_limit:
|
|
break
|
|
except PrebuiltFallback:
|
|
raise
|
|
except Exception as exc:
|
|
raise PrebuiltFallback(
|
|
f"failed to inspect published releases in {repo}: {exc}"
|
|
) from exc
|
|
|
|
if plans:
|
|
return requested_tag, plans
|
|
if last_error is not None:
|
|
raise last_error
|
|
raise PrebuiltFallback(
|
|
f"no installable published llama.cpp releases were found in {repo}"
|
|
)
|
|
|
|
|
|
def normalized_requested_llama_tag(requested_tag: str | None) -> str:
|
|
if isinstance(requested_tag, str):
|
|
normalized = requested_tag.strip()
|
|
if normalized:
|
|
return normalized
|
|
return "latest"
|
|
|
|
|
|
def normalize_compute_cap(value: Any) -> str | None:
|
|
raw = str(value).strip()
|
|
if not raw:
|
|
return None
|
|
if "." in raw:
|
|
parts = raw.split(".", 1)
|
|
if len(parts) != 2:
|
|
return None
|
|
major, minor = parts
|
|
if not major.isdigit() or not minor.isdigit():
|
|
return None
|
|
return f"{int(major)}{int(minor)}"
|
|
if raw.isdigit():
|
|
return str(int(raw))
|
|
return None
|
|
|
|
|
|
def normalize_compute_caps(compute_caps: Iterable[str]) -> list[str]:
|
|
normalized: list[str] = []
|
|
seen: set[str] = set()
|
|
for raw in compute_caps:
|
|
normalized_value = normalize_compute_cap(raw)
|
|
if normalized_value is None:
|
|
continue
|
|
if normalized_value in seen:
|
|
continue
|
|
seen.add(normalized_value)
|
|
normalized.append(normalized_value)
|
|
normalized.sort(key = int)
|
|
return normalized
|
|
|
|
|
|
def parse_cuda_visible_devices(value: str | None) -> list[str] | None:
|
|
if value is None:
|
|
return None
|
|
raw = value.strip()
|
|
if not raw or raw == "-1":
|
|
return []
|
|
return [token.strip() for token in raw.split(",") if token.strip()]
|
|
|
|
|
|
def supports_explicit_visible_device_matching(
|
|
visible_devices: list[str] | None,
|
|
) -> bool:
|
|
if not visible_devices:
|
|
return False
|
|
for token in visible_devices:
|
|
lowered = token.lower()
|
|
if token.isdigit() or lowered.startswith("gpu-"):
|
|
continue
|
|
return False
|
|
return True
|
|
|
|
|
|
def select_visible_gpu_rows(
|
|
gpu_rows: Iterable[tuple[str, str, str]],
|
|
visible_devices: list[str] | None,
|
|
) -> list[tuple[str, str, str]]:
|
|
rows = list(gpu_rows)
|
|
if visible_devices is None:
|
|
return rows
|
|
if not visible_devices:
|
|
return []
|
|
|
|
by_index = {index: (index, uuid, cap) for index, uuid, cap in rows}
|
|
by_uuid = {uuid.lower(): (index, uuid, cap) for index, uuid, cap in rows}
|
|
selected: list[tuple[str, str, str]] = []
|
|
seen_indices: set[str] = set()
|
|
for token in visible_devices:
|
|
row = by_index.get(token)
|
|
if row is None:
|
|
normalized_token = token.lower()
|
|
row = by_uuid.get(normalized_token)
|
|
if row is None and normalized_token.startswith("gpu-"):
|
|
row = by_uuid.get(normalized_token)
|
|
if row is None and not normalized_token.startswith("gpu-"):
|
|
row = by_uuid.get("gpu-" + normalized_token)
|
|
if row is None:
|
|
continue
|
|
index = row[0]
|
|
if index in seen_indices:
|
|
continue
|
|
seen_indices.add(index)
|
|
selected.append(row)
|
|
return selected
|
|
|
|
|
|
def dir_provides_exact_library(directory: str | Path, library: str) -> bool:
|
|
if not library:
|
|
return False
|
|
candidate = Path(directory) / library
|
|
return candidate.exists() and (candidate.is_file() or candidate.is_symlink())
|
|
|
|
|
|
def linux_runtime_dirs_for_required_libraries(
|
|
required_libraries: Iterable[str],
|
|
) -> list[str]:
|
|
required = [library for library in required_libraries if library]
|
|
candidates: list[str | Path] = []
|
|
|
|
env_dirs = os.environ.get("CUDA_RUNTIME_LIB_DIR", "")
|
|
if env_dirs:
|
|
candidates.extend(part for part in env_dirs.split(os.pathsep) if part)
|
|
ld_library_path = os.environ.get("LD_LIBRARY_PATH", "")
|
|
if ld_library_path:
|
|
candidates.extend(part for part in ld_library_path.split(os.pathsep) if part)
|
|
|
|
cuda_roots: list[Path] = []
|
|
for name in ("CUDA_HOME", "CUDA_PATH", "CUDA_ROOT"):
|
|
value = os.environ.get(name)
|
|
if value:
|
|
cuda_roots.append(Path(value))
|
|
cuda_roots.extend(
|
|
Path(path) for path in glob_paths("/usr/local/cuda", "/usr/local/cuda-*")
|
|
)
|
|
|
|
for root in cuda_roots:
|
|
candidates.extend(
|
|
[
|
|
root / "lib",
|
|
root / "lib64",
|
|
root / "targets" / "x86_64-linux" / "lib",
|
|
]
|
|
)
|
|
|
|
candidates.extend(
|
|
Path(path)
|
|
for path in glob_paths(
|
|
"/lib",
|
|
"/lib64",
|
|
"/usr/lib",
|
|
"/usr/lib64",
|
|
"/usr/local/lib",
|
|
"/usr/local/lib64",
|
|
"/lib/x86_64-linux-gnu",
|
|
"/usr/lib/x86_64-linux-gnu",
|
|
)
|
|
)
|
|
candidates.extend(
|
|
Path(path)
|
|
for path in glob_paths("/usr/local/lib/ollama/cuda_v*", "/usr/lib/wsl/lib")
|
|
)
|
|
candidates.extend(Path(path) for path in python_runtime_dirs())
|
|
candidates.extend(Path(path) for path in ldconfig_runtime_dirs(required))
|
|
|
|
resolved = dedupe_existing_dirs(candidates)
|
|
if not required:
|
|
return resolved
|
|
|
|
matched: list[tuple[int, str]] = []
|
|
for directory in resolved:
|
|
base = Path(directory)
|
|
provided = sum(
|
|
1 for library in required if dir_provides_exact_library(directory, library)
|
|
)
|
|
if provided:
|
|
matched.append((provided, directory))
|
|
|
|
matched.sort(key = lambda item: item[0], reverse = True)
|
|
return [directory for _, directory in matched]
|
|
|
|
|
|
def detected_linux_runtime_lines() -> tuple[list[str], dict[str, list[str]]]:
|
|
line_requirements = {
|
|
"cuda13": ["libcudart.so.13", "libcublas.so.13"],
|
|
"cuda12": ["libcudart.so.12", "libcublas.so.12"],
|
|
}
|
|
detected: list[str] = []
|
|
runtime_dirs: dict[str, list[str]] = {}
|
|
for line, required in line_requirements.items():
|
|
dirs = linux_runtime_dirs_for_required_libraries(required)
|
|
library_matches: dict[str, list[str]] = {}
|
|
matching_dirs: list[str] = []
|
|
for library in required:
|
|
matched_dirs = [
|
|
directory
|
|
for directory in dirs
|
|
if any(Path(directory).glob(f"{library}*"))
|
|
]
|
|
if not matched_dirs:
|
|
library_matches = {}
|
|
matching_dirs = []
|
|
break
|
|
library_matches[library] = matched_dirs
|
|
for directory in matched_dirs:
|
|
if directory not in matching_dirs:
|
|
matching_dirs.append(directory)
|
|
if library_matches:
|
|
detected.append(line)
|
|
runtime_dirs[line] = matching_dirs
|
|
return detected, runtime_dirs
|
|
|
|
|
|
def release_asset_map(release: dict[str, Any]) -> dict[str, str]:
|
|
assets = release.get("assets")
|
|
if not isinstance(assets, list):
|
|
return {}
|
|
return {
|
|
asset["name"]: asset.get("browser_download_url", "")
|
|
for asset in assets
|
|
if isinstance(asset, dict)
|
|
and isinstance(asset.get("name"), str)
|
|
and isinstance(asset.get("browser_download_url"), str)
|
|
}
|
|
|
|
|
|
def parse_published_artifact(raw: Any) -> PublishedLlamaArtifact | None:
|
|
if not isinstance(raw, dict):
|
|
raise ValueError("artifact entry was not an object")
|
|
asset_name = raw.get("asset_name")
|
|
install_kind = raw.get("install_kind")
|
|
if not isinstance(asset_name, str) or not asset_name:
|
|
raise ValueError("artifact.asset_name was missing or not a string")
|
|
if not isinstance(install_kind, str) or not install_kind:
|
|
raise ValueError(
|
|
f"artifact {asset_name} install_kind was missing or not a string"
|
|
)
|
|
|
|
supported_sms_raw = raw.get("supported_sms", [])
|
|
if not isinstance(supported_sms_raw, (list, tuple)):
|
|
raise ValueError(f"artifact {asset_name} supported_sms must be a list or tuple")
|
|
if any(not isinstance(value, (int, str)) for value in supported_sms_raw):
|
|
raise ValueError(
|
|
f"artifact {asset_name} supported_sms entries must be ints or strings"
|
|
)
|
|
supported_sms = normalize_compute_caps(supported_sms_raw)
|
|
|
|
min_sm_raw = raw.get("min_sm")
|
|
max_sm_raw = raw.get("max_sm")
|
|
try:
|
|
min_sm = int(min_sm_raw) if min_sm_raw is not None else None
|
|
max_sm = int(max_sm_raw) if max_sm_raw is not None else None
|
|
except (TypeError, ValueError) as exc:
|
|
raise ValueError(
|
|
f"artifact {asset_name} min_sm/max_sm were not integers"
|
|
) from exc
|
|
runtime_line = raw.get("runtime_line")
|
|
coverage_class = raw.get("coverage_class")
|
|
bundle_profile = raw.get("bundle_profile")
|
|
rank_raw = raw.get("rank", 1000)
|
|
if runtime_line is not None and not isinstance(runtime_line, str):
|
|
raise ValueError(f"artifact {asset_name} runtime_line was not a string")
|
|
if coverage_class is not None and not isinstance(coverage_class, str):
|
|
raise ValueError(f"artifact {asset_name} coverage_class was not a string")
|
|
if bundle_profile is not None and not isinstance(bundle_profile, str):
|
|
raise ValueError(f"artifact {asset_name} bundle_profile was not a string")
|
|
try:
|
|
rank = int(rank_raw)
|
|
except (TypeError, ValueError):
|
|
raise ValueError(f"artifact {asset_name} rank was not an integer")
|
|
return PublishedLlamaArtifact(
|
|
asset_name = asset_name,
|
|
install_kind = install_kind,
|
|
runtime_line = runtime_line
|
|
if isinstance(runtime_line, str) and runtime_line
|
|
else None,
|
|
coverage_class = coverage_class
|
|
if isinstance(coverage_class, str) and coverage_class
|
|
else None,
|
|
supported_sms = supported_sms,
|
|
min_sm = min_sm,
|
|
max_sm = max_sm,
|
|
bundle_profile = bundle_profile
|
|
if isinstance(bundle_profile, str) and bundle_profile
|
|
else None,
|
|
rank = rank,
|
|
)
|
|
|
|
|
|
def parse_published_release_bundle(
|
|
repo: str, release: dict[str, Any]
|
|
) -> PublishedReleaseBundle | None:
|
|
release_tag = release.get("tag_name")
|
|
if not isinstance(release_tag, str) or not release_tag:
|
|
return None
|
|
|
|
assets = release_asset_map(release)
|
|
manifest_url = assets.get(DEFAULT_PUBLISHED_MANIFEST_ASSET)
|
|
if not manifest_url:
|
|
return None
|
|
|
|
# Mixed repos are filtered by an explicit release-side manifest rather than
|
|
# by release tag or asset filename conventions.
|
|
manifest_bytes = download_bytes(
|
|
manifest_url,
|
|
timeout = 30,
|
|
headers = auth_headers(manifest_url),
|
|
)
|
|
manifest_sha256 = sha256_bytes(manifest_bytes)
|
|
try:
|
|
manifest_payload = json.loads(manifest_bytes.decode("utf-8"))
|
|
except (UnicodeDecodeError, json.JSONDecodeError) as exc:
|
|
raise RuntimeError(
|
|
f"published manifest {DEFAULT_PUBLISHED_MANIFEST_ASSET} was not valid JSON"
|
|
) from exc
|
|
if not isinstance(manifest_payload, dict):
|
|
raise RuntimeError(
|
|
f"published manifest {DEFAULT_PUBLISHED_MANIFEST_ASSET} was not a JSON object"
|
|
)
|
|
validate_schema_version(
|
|
manifest_payload,
|
|
label = f"published manifest {DEFAULT_PUBLISHED_MANIFEST_ASSET} in {repo}@{release_tag}",
|
|
)
|
|
component = manifest_payload.get("component")
|
|
upstream_tag = manifest_payload.get("upstream_tag")
|
|
source_repo = manifest_payload.get("source_repo")
|
|
source_repo_url = manifest_payload.get("source_repo_url")
|
|
source_ref_kind = normalize_source_ref_kind(manifest_payload.get("source_ref_kind"))
|
|
requested_source_ref = manifest_payload.get("requested_source_ref")
|
|
resolved_source_ref = manifest_payload.get("resolved_source_ref")
|
|
source_commit = normalize_source_commit(manifest_payload.get("source_commit"))
|
|
source_commit_short = manifest_payload.get("source_commit_short")
|
|
if component != "llama.cpp":
|
|
return None
|
|
if not isinstance(upstream_tag, str) or not upstream_tag:
|
|
raise RuntimeError(
|
|
f"published manifest {DEFAULT_PUBLISHED_MANIFEST_ASSET} in {repo}@{release_tag} omitted upstream_tag"
|
|
)
|
|
|
|
artifacts_payload = manifest_payload.get("artifacts")
|
|
if not isinstance(artifacts_payload, list):
|
|
raise RuntimeError(
|
|
f"published manifest {DEFAULT_PUBLISHED_MANIFEST_ASSET} in {repo}@{release_tag} omitted artifacts"
|
|
)
|
|
|
|
artifacts: list[PublishedLlamaArtifact] = []
|
|
for index, raw_artifact in enumerate(artifacts_payload):
|
|
try:
|
|
artifact = parse_published_artifact(raw_artifact)
|
|
except ValueError as exc:
|
|
log(
|
|
f"published artifact ignored for {repo}@{release_tag} artifact[{index}]: {exc}"
|
|
)
|
|
continue
|
|
if artifact is not None:
|
|
artifacts.append(artifact)
|
|
selection_log = [
|
|
f"published_release: repo={repo}",
|
|
f"published_release: tag={release_tag}",
|
|
f"published_release: manifest={DEFAULT_PUBLISHED_MANIFEST_ASSET}",
|
|
f"published_release: upstream_tag={upstream_tag}",
|
|
]
|
|
if isinstance(source_repo, str) and source_repo:
|
|
selection_log.append(f"published_release: source_repo={source_repo}")
|
|
if source_commit:
|
|
selection_log.append(f"published_release: source_commit={source_commit}")
|
|
return PublishedReleaseBundle(
|
|
repo = repo,
|
|
release_tag = release_tag,
|
|
upstream_tag = upstream_tag,
|
|
manifest_sha256 = manifest_sha256,
|
|
source_repo = source_repo
|
|
if isinstance(source_repo, str) and source_repo
|
|
else None,
|
|
source_repo_url = source_repo_url
|
|
if isinstance(source_repo_url, str) and source_repo_url
|
|
else None,
|
|
source_ref_kind = source_ref_kind,
|
|
requested_source_ref = requested_source_ref
|
|
if isinstance(requested_source_ref, str) and requested_source_ref
|
|
else None,
|
|
resolved_source_ref = resolved_source_ref
|
|
if isinstance(resolved_source_ref, str) and resolved_source_ref
|
|
else None,
|
|
source_commit = source_commit,
|
|
source_commit_short = source_commit_short
|
|
if isinstance(source_commit_short, str) and source_commit_short
|
|
else None,
|
|
assets = assets,
|
|
manifest_asset_name = DEFAULT_PUBLISHED_MANIFEST_ASSET,
|
|
artifacts = artifacts,
|
|
selection_log = selection_log,
|
|
)
|
|
|
|
|
|
def parse_approved_release_checksums(
|
|
repo: str,
|
|
release_tag: str,
|
|
payload: Any,
|
|
) -> ApprovedReleaseChecksums:
|
|
if not isinstance(payload, dict):
|
|
raise RuntimeError(
|
|
f"published checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET} was not a JSON object"
|
|
)
|
|
validate_schema_version(
|
|
payload,
|
|
label = f"published checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET}",
|
|
)
|
|
if payload.get("component") != "llama.cpp":
|
|
raise RuntimeError(
|
|
f"published checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET} did not describe llama.cpp"
|
|
)
|
|
payload_release_tag = payload.get("release_tag")
|
|
if not isinstance(payload_release_tag, str) or not payload_release_tag:
|
|
raise RuntimeError(
|
|
f"published checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET} omitted release_tag"
|
|
)
|
|
if payload_release_tag != release_tag:
|
|
raise RuntimeError(
|
|
f"published checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET} release_tag={payload_release_tag} "
|
|
f"did not match pinned release tag {release_tag}"
|
|
)
|
|
upstream_tag = payload.get("upstream_tag")
|
|
if not isinstance(upstream_tag, str) or not upstream_tag:
|
|
raise RuntimeError(
|
|
f"published checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET} omitted upstream_tag"
|
|
)
|
|
artifacts_payload = payload.get("artifacts")
|
|
if not isinstance(artifacts_payload, dict):
|
|
raise RuntimeError(
|
|
f"published checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET} omitted artifacts"
|
|
)
|
|
|
|
artifacts: dict[str, ApprovedArtifactHash] = {}
|
|
for asset_name, raw_entry in artifacts_payload.items():
|
|
if not isinstance(asset_name, str) or not asset_name:
|
|
raise RuntimeError(
|
|
"published checksum asset used a non-string artifact key"
|
|
)
|
|
if not isinstance(raw_entry, dict):
|
|
raise RuntimeError(
|
|
f"published checksum entry for {asset_name} was not an object"
|
|
)
|
|
digest = normalize_sha256_digest(raw_entry.get("sha256"))
|
|
if not digest:
|
|
raise RuntimeError(
|
|
f"published checksum entry for {asset_name} omitted a valid sha256"
|
|
)
|
|
repo_value = raw_entry.get("repo")
|
|
kind_value = raw_entry.get("kind")
|
|
artifacts[asset_name] = ApprovedArtifactHash(
|
|
asset_name = asset_name,
|
|
sha256 = digest,
|
|
repo = repo_value if isinstance(repo_value, str) and repo_value else None,
|
|
kind = kind_value if isinstance(kind_value, str) and kind_value else None,
|
|
)
|
|
|
|
source_commit = normalize_source_commit(payload.get("source_commit"))
|
|
source_commit_short = payload.get("source_commit_short")
|
|
source_repo = payload.get("source_repo")
|
|
source_repo_url = payload.get("source_repo_url")
|
|
source_ref_kind = normalize_source_ref_kind(payload.get("source_ref_kind"))
|
|
requested_source_ref = payload.get("requested_source_ref")
|
|
resolved_source_ref = payload.get("resolved_source_ref")
|
|
return ApprovedReleaseChecksums(
|
|
repo = repo,
|
|
release_tag = release_tag,
|
|
upstream_tag = upstream_tag,
|
|
source_repo = source_repo
|
|
if isinstance(source_repo, str) and source_repo
|
|
else None,
|
|
source_repo_url = source_repo_url
|
|
if isinstance(source_repo_url, str) and source_repo_url
|
|
else None,
|
|
source_ref_kind = source_ref_kind,
|
|
requested_source_ref = requested_source_ref
|
|
if isinstance(requested_source_ref, str) and requested_source_ref
|
|
else None,
|
|
resolved_source_ref = resolved_source_ref
|
|
if isinstance(resolved_source_ref, str) and resolved_source_ref
|
|
else None,
|
|
source_commit = source_commit,
|
|
source_commit_short = source_commit_short
|
|
if isinstance(source_commit_short, str) and source_commit_short
|
|
else None,
|
|
artifacts = artifacts,
|
|
)
|
|
|
|
|
|
def load_approved_release_checksums(
|
|
repo: str, release_tag: str
|
|
) -> ApprovedReleaseChecksums:
|
|
try:
|
|
release = github_release(repo, release_tag)
|
|
except Exception as exc:
|
|
raise PrebuiltFallback(
|
|
f"approved prebuilt release {repo}@{release_tag} was not available"
|
|
) from exc
|
|
assets = release_asset_map(release)
|
|
checksum_url = assets.get(DEFAULT_PUBLISHED_SHA256_ASSET)
|
|
if not checksum_url:
|
|
raise PrebuiltFallback(
|
|
f"approved prebuilt release {repo}@{release_tag} did not expose {DEFAULT_PUBLISHED_SHA256_ASSET}"
|
|
)
|
|
try:
|
|
payload = fetch_json(checksum_url)
|
|
checksums = parse_approved_release_checksums(repo, release_tag, payload)
|
|
except PrebuiltFallback:
|
|
raise
|
|
except Exception as exc:
|
|
raise PrebuiltFallback(
|
|
f"approved checksum asset {DEFAULT_PUBLISHED_SHA256_ASSET} in {repo}@{release_tag} was invalid"
|
|
) from exc
|
|
return checksums
|
|
|
|
|
|
def iter_published_release_bundles(
|
|
repo: str, published_release_tag: str = ""
|
|
) -> Iterable[PublishedReleaseBundle]:
|
|
releases = (
|
|
[github_release(repo, published_release_tag)]
|
|
if published_release_tag
|
|
else github_releases(repo, max_pages = DEFAULT_GITHUB_RELEASE_SCAN_MAX_PAGES)
|
|
)
|
|
for release in releases:
|
|
if not published_release_tag and (
|
|
release.get("draft") or release.get("prerelease")
|
|
):
|
|
continue
|
|
try:
|
|
bundle = parse_published_release_bundle(repo, release)
|
|
except Exception as exc:
|
|
release_tag = release.get("tag_name", "unknown")
|
|
log(f"published release metadata ignored for {repo}@{release_tag}: {exc}")
|
|
continue
|
|
if bundle is None:
|
|
continue
|
|
yield bundle
|
|
|
|
|
|
def linux_cuda_choice_from_release(
|
|
host: HostInfo,
|
|
release: PublishedReleaseBundle,
|
|
preferred_runtime_line: str | None = None,
|
|
selection_preamble: Iterable[str] = (),
|
|
) -> LinuxCudaSelection | None:
|
|
host_sms = normalize_compute_caps(host.compute_caps)
|
|
detected_runtime_lines, runtime_dirs = detected_linux_runtime_lines()
|
|
driver_runtime_lines = compatible_linux_runtime_lines(host)
|
|
runtime_lines = [
|
|
runtime_line
|
|
for runtime_line in detected_runtime_lines
|
|
if runtime_line in driver_runtime_lines
|
|
]
|
|
ordered_runtime_lines = list(runtime_lines)
|
|
selection_log = (
|
|
list(release.selection_log)
|
|
+ list(selection_preamble)
|
|
+ [
|
|
f"linux_cuda_selection: release={release.release_tag}",
|
|
f"linux_cuda_selection: detected_sms={','.join(host_sms) if host_sms else 'unknown'}",
|
|
"linux_cuda_selection: detected_runtime_lines="
|
|
+ (",".join(detected_runtime_lines) if detected_runtime_lines else "none"),
|
|
"linux_cuda_selection: driver_runtime_lines="
|
|
+ (",".join(driver_runtime_lines) if driver_runtime_lines else "none"),
|
|
"linux_cuda_selection: compatible_runtime_lines="
|
|
+ (",".join(runtime_lines) if runtime_lines else "none"),
|
|
]
|
|
)
|
|
for runtime_line in ("cuda13", "cuda12"):
|
|
selection_log.append(
|
|
"linux_cuda_selection: runtime_dirs "
|
|
f"{runtime_line}="
|
|
+ (
|
|
",".join(runtime_dirs.get(runtime_line, []))
|
|
if runtime_dirs.get(runtime_line)
|
|
else "none"
|
|
)
|
|
)
|
|
published_artifacts = [
|
|
artifact
|
|
for artifact in release.artifacts
|
|
if artifact.install_kind == "linux-cuda"
|
|
]
|
|
published_asset_names = sorted(
|
|
artifact.asset_name for artifact in published_artifacts
|
|
)
|
|
selection_log.append(
|
|
"linux_cuda_selection: published_assets="
|
|
+ (",".join(published_asset_names) if published_asset_names else "none")
|
|
)
|
|
|
|
if not host_sms:
|
|
selection_log.append(
|
|
"linux_cuda_selection: compute capability detection unavailable; prefer portable by runtime line"
|
|
)
|
|
if not runtime_lines:
|
|
selection_log.append(
|
|
"linux_cuda_selection: no Linux CUDA runtime line satisfied both runtime libraries and driver compatibility"
|
|
)
|
|
return None
|
|
|
|
if preferred_runtime_line:
|
|
if preferred_runtime_line in ordered_runtime_lines:
|
|
ordered_runtime_lines = [preferred_runtime_line] + [
|
|
runtime_line
|
|
for runtime_line in ordered_runtime_lines
|
|
if runtime_line != preferred_runtime_line
|
|
]
|
|
selection_log.append(
|
|
"linux_cuda_selection: torch_preferred_runtime_line="
|
|
f"{preferred_runtime_line} reordered_attempts={','.join(ordered_runtime_lines)}"
|
|
)
|
|
else:
|
|
selection_log.append(
|
|
"linux_cuda_selection: torch_preferred_runtime_line="
|
|
f"{preferred_runtime_line} unavailable_on_host"
|
|
)
|
|
|
|
attempts: list[AssetChoice] = []
|
|
seen_attempts: set[str] = set()
|
|
|
|
def add_attempt(
|
|
artifact: PublishedLlamaArtifact, asset_url: str, reason: str
|
|
) -> None:
|
|
asset_name = artifact.asset_name
|
|
if asset_name in seen_attempts:
|
|
return
|
|
seen_attempts.add(asset_name)
|
|
attempts.append(
|
|
AssetChoice(
|
|
repo = release.repo,
|
|
tag = release.release_tag,
|
|
name = asset_name,
|
|
url = asset_url,
|
|
source_label = "published",
|
|
is_ready_bundle = True,
|
|
install_kind = "linux-cuda",
|
|
bundle_profile = artifact.bundle_profile,
|
|
runtime_line = artifact.runtime_line,
|
|
coverage_class = artifact.coverage_class,
|
|
supported_sms = artifact.supported_sms,
|
|
min_sm = artifact.min_sm,
|
|
max_sm = artifact.max_sm,
|
|
selection_log = list(selection_log)
|
|
+ [
|
|
"linux_cuda_selection: selected "
|
|
f"{asset_name} runtime_line={artifact.runtime_line} coverage_class={artifact.coverage_class} reason={reason}"
|
|
],
|
|
)
|
|
)
|
|
|
|
for runtime_line in ordered_runtime_lines:
|
|
coverage_candidates: list[tuple[PublishedLlamaArtifact, str]] = []
|
|
portable_candidate: tuple[PublishedLlamaArtifact, str] | None = None
|
|
for artifact in published_artifacts:
|
|
if artifact.runtime_line != runtime_line:
|
|
continue
|
|
asset_name = artifact.asset_name
|
|
asset_url = release.assets.get(asset_name)
|
|
if not asset_url:
|
|
selection_log.append(
|
|
f"linux_cuda_selection: reject {asset_name} missing asset"
|
|
)
|
|
continue
|
|
if not host_sms and artifact.coverage_class != "portable":
|
|
selection_log.append(
|
|
"linux_cuda_selection: reject "
|
|
f"{asset_name} runtime_line={runtime_line} coverage_class={artifact.coverage_class} "
|
|
"reason=unknown_compute_caps_prefer_portable"
|
|
)
|
|
continue
|
|
|
|
if not artifact.supported_sms:
|
|
selection_log.append(
|
|
"linux_cuda_selection: reject "
|
|
f"{asset_name} runtime_line={runtime_line} coverage_class={artifact.coverage_class} "
|
|
"reason=artifact_missing_supported_sms"
|
|
)
|
|
continue
|
|
if artifact.min_sm is None or artifact.max_sm is None:
|
|
selection_log.append(
|
|
"linux_cuda_selection: reject "
|
|
f"{asset_name} runtime_line={runtime_line} coverage_class={artifact.coverage_class} "
|
|
"reason=artifact_missing_sm_bounds"
|
|
)
|
|
continue
|
|
|
|
supported_sms = {str(value) for value in artifact.supported_sms}
|
|
missing_sms = [sm for sm in host_sms if sm not in supported_sms]
|
|
out_of_range_sms = [
|
|
sm
|
|
for sm in host_sms
|
|
if not (artifact.min_sm <= int(sm) <= artifact.max_sm)
|
|
]
|
|
reasons: list[str] = []
|
|
if missing_sms:
|
|
reasons.append(f"missing_sms={','.join(missing_sms)}")
|
|
if out_of_range_sms:
|
|
reasons.append(f"out_of_range_sms={','.join(out_of_range_sms)}")
|
|
if reasons:
|
|
selection_log.append(
|
|
"linux_cuda_selection: reject "
|
|
f"{asset_name} runtime_line={runtime_line} coverage_class={artifact.coverage_class} "
|
|
f"coverage={artifact.min_sm}-{artifact.max_sm} supported={','.join(artifact.supported_sms)} "
|
|
f"reasons={' '.join(reasons)}"
|
|
)
|
|
continue
|
|
|
|
selection_log.append(
|
|
"linux_cuda_selection: accept "
|
|
f"{asset_name} runtime_line={runtime_line} coverage_class={artifact.coverage_class} "
|
|
f"coverage={artifact.min_sm}-{artifact.max_sm} supported={','.join(artifact.supported_sms)}"
|
|
)
|
|
if artifact.coverage_class == "portable":
|
|
portable_candidate = (artifact, asset_url)
|
|
else:
|
|
coverage_candidates.append((artifact, asset_url))
|
|
|
|
if coverage_candidates:
|
|
artifact, url = sorted(
|
|
coverage_candidates,
|
|
key = lambda item: (
|
|
(item[0].max_sm or 0) - (item[0].min_sm or 0),
|
|
item[0].rank,
|
|
item[0].max_sm or 0,
|
|
),
|
|
)[0]
|
|
add_attempt(artifact, url, "best coverage for runtime line")
|
|
if portable_candidate:
|
|
artifact, url = portable_candidate
|
|
add_attempt(artifact, url, "portable fallback for runtime line")
|
|
|
|
if not attempts:
|
|
return None
|
|
|
|
selection_log.append(
|
|
"linux_cuda_selection: attempt_order="
|
|
+ ",".join(choice.name for choice in attempts)
|
|
)
|
|
for attempt in attempts:
|
|
attempt.selection_log = list(selection_log) + [
|
|
"linux_cuda_selection: attempt "
|
|
f"{attempt.name} runtime_line={attempt.runtime_line} coverage_class={attempt.coverage_class}"
|
|
]
|
|
return LinuxCudaSelection(attempts = attempts, selection_log = selection_log)
|
|
|
|
|
|
def latest_published_linux_cuda_tag(host: HostInfo, published_repo: str) -> str | None:
|
|
for release in iter_published_release_bundles(published_repo):
|
|
if linux_cuda_choice_from_release(host, release):
|
|
return release.upstream_tag
|
|
return None
|
|
|
|
|
|
def iter_upstream_releases() -> Iterable[dict[str, Any]]:
|
|
for release in github_releases(
|
|
UPSTREAM_REPO, max_pages = DEFAULT_GITHUB_RELEASE_SCAN_MAX_PAGES
|
|
):
|
|
if release.get("draft") or release.get("prerelease"):
|
|
continue
|
|
yield release
|
|
|
|
|
|
def pinned_published_release_bundle(
|
|
repo: str, published_release_tag: str
|
|
) -> PublishedReleaseBundle:
|
|
bundle = next(iter_published_release_bundles(repo, published_release_tag), None)
|
|
if bundle is None:
|
|
raise PrebuiltFallback(
|
|
f"published release {repo}@{published_release_tag} did not expose a usable llama.cpp manifest"
|
|
)
|
|
return bundle
|
|
|
|
|
|
def validated_checksums_for_bundle(
|
|
repo: str, bundle: PublishedReleaseBundle
|
|
) -> ApprovedReleaseChecksums:
|
|
checksums = load_approved_release_checksums(repo, bundle.release_tag)
|
|
manifest_hash = checksums.artifacts.get(bundle.manifest_asset_name)
|
|
if manifest_hash is not None and bundle.manifest_sha256 is not None:
|
|
if manifest_hash.sha256 != bundle.manifest_sha256:
|
|
raise PrebuiltFallback(
|
|
"published manifest checksum did not match the approved checksum asset"
|
|
)
|
|
# Accept bundles that carry only an exact-commit source archive
|
|
# (e.g. llama.cpp-source-commit-<sha>.tar.gz) without requiring the
|
|
# legacy llama.cpp-source-<upstream_tag>.tar.gz entry.
|
|
if exact_source_archive_hash(checksums) is None:
|
|
require_approved_source_hash(checksums, bundle.upstream_tag)
|
|
return checksums
|
|
|
|
|
|
def published_release_matches_request(
|
|
bundle: PublishedReleaseBundle, requested_ref: str
|
|
) -> bool:
|
|
if requested_ref == "latest":
|
|
return True
|
|
for candidate in (
|
|
bundle.upstream_tag,
|
|
bundle.requested_source_ref,
|
|
bundle.resolved_source_ref,
|
|
bundle.source_commit,
|
|
):
|
|
if refs_match(candidate, requested_ref):
|
|
return True
|
|
return False
|
|
|
|
|
|
def resolve_published_release(
|
|
requested_tag: str | None,
|
|
published_repo: str,
|
|
published_release_tag: str = "",
|
|
) -> ResolvedPublishedRelease:
|
|
repo = published_repo or DEFAULT_PUBLISHED_REPO
|
|
normalized_requested = normalized_requested_llama_tag(requested_tag)
|
|
|
|
if published_release_tag:
|
|
bundle = pinned_published_release_bundle(repo, published_release_tag)
|
|
if not published_release_matches_request(bundle, normalized_requested):
|
|
raise PrebuiltFallback(
|
|
"published release "
|
|
f"{repo}@{published_release_tag} targeted upstream tag {bundle.upstream_tag}, "
|
|
f"but requested {normalized_requested}"
|
|
)
|
|
return ResolvedPublishedRelease(
|
|
bundle = bundle,
|
|
checksums = validated_checksums_for_bundle(repo, bundle),
|
|
)
|
|
|
|
skipped_invalid = 0
|
|
for bundle in iter_published_release_bundles(repo):
|
|
if not published_release_matches_request(bundle, normalized_requested):
|
|
continue
|
|
try:
|
|
checksums = validated_checksums_for_bundle(repo, bundle)
|
|
except PrebuiltFallback as exc:
|
|
skipped_invalid += 1
|
|
log(
|
|
"published release ignored for install resolution: "
|
|
f"{repo}@{bundle.release_tag} ({exc})"
|
|
)
|
|
continue
|
|
return ResolvedPublishedRelease(bundle = bundle, checksums = checksums)
|
|
|
|
if normalized_requested == "latest":
|
|
if skipped_invalid:
|
|
raise PrebuiltFallback(
|
|
f"no usable published llama.cpp releases were available in {repo}"
|
|
)
|
|
raise PrebuiltFallback(
|
|
f"no published llama.cpp releases were available in {repo}"
|
|
)
|
|
|
|
raise PrebuiltFallback(
|
|
f"no published prebuilt release in {repo} matched upstream tag {normalized_requested}"
|
|
)
|
|
|
|
|
|
def iter_resolved_published_releases(
|
|
requested_tag: str | None,
|
|
published_repo: str,
|
|
published_release_tag: str = "",
|
|
) -> Iterable[ResolvedPublishedRelease]:
|
|
repo = published_repo or DEFAULT_PUBLISHED_REPO
|
|
normalized_requested = normalized_requested_llama_tag(requested_tag)
|
|
|
|
if published_release_tag:
|
|
bundle = pinned_published_release_bundle(repo, published_release_tag)
|
|
if not published_release_matches_request(bundle, normalized_requested):
|
|
raise PrebuiltFallback(
|
|
"published release "
|
|
f"{repo}@{published_release_tag} targeted upstream tag {bundle.upstream_tag}, "
|
|
f"but requested {normalized_requested}"
|
|
)
|
|
yield ResolvedPublishedRelease(
|
|
bundle = bundle,
|
|
checksums = validated_checksums_for_bundle(repo, bundle),
|
|
)
|
|
return
|
|
|
|
matched_any = False
|
|
skipped_invalid = 0
|
|
yielded_valid = False
|
|
for bundle in iter_published_release_bundles(repo):
|
|
if not published_release_matches_request(bundle, normalized_requested):
|
|
continue
|
|
matched_any = True
|
|
try:
|
|
checksums = validated_checksums_for_bundle(repo, bundle)
|
|
except PrebuiltFallback as exc:
|
|
skipped_invalid += 1
|
|
log(
|
|
"published release ignored for install resolution: "
|
|
f"{repo}@{bundle.release_tag} ({exc})"
|
|
)
|
|
continue
|
|
yielded_valid = True
|
|
yield ResolvedPublishedRelease(bundle = bundle, checksums = checksums)
|
|
|
|
if yielded_valid:
|
|
return
|
|
|
|
if matched_any:
|
|
if skipped_invalid:
|
|
raise PrebuiltFallback(
|
|
f"no usable published llama.cpp releases were available in {repo}"
|
|
)
|
|
return
|
|
|
|
if normalized_requested == "latest":
|
|
raise PrebuiltFallback(
|
|
f"no published llama.cpp releases were available in {repo}"
|
|
)
|
|
|
|
raise PrebuiltFallback(
|
|
f"no published prebuilt release in {repo} matched upstream tag {normalized_requested}"
|
|
)
|
|
|
|
|
|
def resolve_requested_llama_tag(
|
|
requested_tag: str | None,
|
|
published_repo: str = "",
|
|
published_release_tag: str = "",
|
|
) -> str:
|
|
"""Resolve a llama.cpp tag for source-build fallback.
|
|
|
|
Resolution order:
|
|
1. Concrete tag (e.g. "b8508") -- returned as-is.
|
|
2. "latest" with published_repo -- resolve the latest usable Unsloth
|
|
published release bundle and return its upstream_tag. This is the
|
|
preferred version that matches the published prebuilt metadata.
|
|
3. "latest" without published_repo or if (2) fails -- query the upstream
|
|
ggml-org/llama.cpp repo. This may return a newer, untested tag.
|
|
|
|
The Unsloth repo is preferred because its releases are pinned to specific
|
|
upstream tags that have been validated with Unsloth Studio. Using the
|
|
upstream bleeding-edge tag risks API/ABI incompatibilities.
|
|
"""
|
|
normalized_requested = normalized_requested_llama_tag(requested_tag)
|
|
if normalized_requested != "latest":
|
|
return normalized_requested
|
|
# Prefer the Unsloth release repo tag (tested/approved) over bleeding-edge
|
|
# upstream. For example, unslothai/llama.cpp may publish b8508 while
|
|
# ggml-org/llama.cpp latest is b8514. The source-build fallback should
|
|
# compile the same version the prebuilt path would have installed.
|
|
if published_repo:
|
|
try:
|
|
return resolve_published_release(
|
|
"latest",
|
|
published_repo,
|
|
published_release_tag,
|
|
).bundle.upstream_tag
|
|
except Exception:
|
|
pass
|
|
# Fall back to upstream ggml-org latest release tag
|
|
return latest_upstream_release_tag()
|
|
|
|
|
|
def resolve_requested_install_tag(
|
|
requested_tag: str | None,
|
|
published_release_tag: str = "",
|
|
published_repo: str = DEFAULT_PUBLISHED_REPO,
|
|
) -> str:
|
|
return resolve_published_release(
|
|
requested_tag,
|
|
published_repo,
|
|
published_release_tag,
|
|
).bundle.upstream_tag
|
|
|
|
|
|
def exact_source_archive_hash(
|
|
checksums: ApprovedReleaseChecksums,
|
|
) -> ApprovedArtifactHash | None:
|
|
if not checksums.source_commit:
|
|
return None
|
|
return checksums.artifacts.get(
|
|
exact_source_archive_logical_name(checksums.source_commit)
|
|
)
|
|
|
|
|
|
def source_clone_url_from_checksums(checksums: ApprovedReleaseChecksums) -> str | None:
|
|
return source_repo_clone_url(checksums.source_repo, checksums.source_repo_url)
|
|
|
|
|
|
def source_build_plan_for_release(
|
|
release: ResolvedPublishedRelease,
|
|
) -> SourceBuildPlan:
|
|
checksums = release.checksums
|
|
exact_source = exact_source_archive_hash(checksums)
|
|
source_repo = checksums.source_repo or release.bundle.source_repo
|
|
source_repo_url = checksums.source_repo_url or release.bundle.source_repo_url
|
|
requested_source_ref = (
|
|
checksums.requested_source_ref or release.bundle.requested_source_ref
|
|
)
|
|
resolved_source_ref = (
|
|
checksums.resolved_source_ref or release.bundle.resolved_source_ref
|
|
)
|
|
source_commit = checksums.source_commit or release.bundle.source_commit
|
|
source_ref_kind = checksums.source_ref_kind or release.bundle.source_ref_kind
|
|
source_url = source_repo_clone_url(source_repo, source_repo_url)
|
|
if exact_source is not None and source_url and source_commit:
|
|
return SourceBuildPlan(
|
|
source_url = source_url,
|
|
source_ref = source_commit,
|
|
source_ref_kind = "commit",
|
|
compatibility_upstream_tag = release.bundle.upstream_tag,
|
|
source_repo = source_repo,
|
|
source_repo_url = source_repo_url,
|
|
requested_source_ref = requested_source_ref,
|
|
resolved_source_ref = resolved_source_ref,
|
|
source_commit = source_commit,
|
|
)
|
|
source_ref = checkout_friendly_ref(
|
|
source_ref_kind, resolved_source_ref or requested_source_ref
|
|
)
|
|
if (
|
|
source_url
|
|
and source_ref
|
|
and source_ref_kind in {"tag", "branch", "pull", "commit"}
|
|
):
|
|
return SourceBuildPlan(
|
|
source_url = source_url,
|
|
source_ref = source_ref,
|
|
source_ref_kind = source_ref_kind,
|
|
compatibility_upstream_tag = release.bundle.upstream_tag,
|
|
source_repo = source_repo,
|
|
source_repo_url = source_repo_url,
|
|
requested_source_ref = requested_source_ref,
|
|
resolved_source_ref = resolved_source_ref,
|
|
source_commit = source_commit,
|
|
)
|
|
return SourceBuildPlan(
|
|
source_url = source_url_from_repo_slug(UPSTREAM_REPO)
|
|
or "https://github.com/ggml-org/llama.cpp",
|
|
source_ref = release.bundle.upstream_tag,
|
|
source_ref_kind = "tag",
|
|
compatibility_upstream_tag = release.bundle.upstream_tag,
|
|
source_repo = source_repo,
|
|
source_repo_url = source_repo_url,
|
|
requested_source_ref = requested_source_ref,
|
|
resolved_source_ref = resolved_source_ref,
|
|
source_commit = source_commit,
|
|
)
|
|
|
|
|
|
def resolve_source_build_plan(
|
|
requested_tag: str | None,
|
|
published_repo: str,
|
|
published_release_tag: str = "",
|
|
) -> SourceBuildPlan:
|
|
normalized_requested = normalized_requested_llama_tag(requested_tag)
|
|
if normalized_requested != "latest":
|
|
try:
|
|
release = resolve_published_release(
|
|
normalized_requested,
|
|
published_repo,
|
|
published_release_tag,
|
|
)
|
|
return source_build_plan_for_release(release)
|
|
except Exception:
|
|
pass
|
|
inferred_kind = infer_source_ref_kind(normalized_requested)
|
|
return SourceBuildPlan(
|
|
source_url = "https://github.com/ggml-org/llama.cpp",
|
|
source_ref = checkout_friendly_ref(inferred_kind, normalized_requested)
|
|
or normalized_requested,
|
|
source_ref_kind = inferred_kind,
|
|
compatibility_upstream_tag = normalized_requested,
|
|
)
|
|
|
|
if published_repo:
|
|
try:
|
|
release = resolve_published_release(
|
|
"latest",
|
|
published_repo,
|
|
published_release_tag,
|
|
)
|
|
return source_build_plan_for_release(release)
|
|
except Exception:
|
|
pass
|
|
latest_tag = latest_upstream_release_tag()
|
|
return SourceBuildPlan(
|
|
source_url = "https://github.com/ggml-org/llama.cpp",
|
|
source_ref = latest_tag,
|
|
source_ref_kind = "tag",
|
|
compatibility_upstream_tag = latest_tag,
|
|
)
|
|
|
|
|
|
def run_capture(
|
|
command: list[str],
|
|
*,
|
|
timeout: int = 30,
|
|
check: bool = False,
|
|
env: dict[str, str] | None = None,
|
|
) -> subprocess.CompletedProcess[str]:
|
|
result = subprocess.run(
|
|
command,
|
|
capture_output = True,
|
|
text = True,
|
|
timeout = timeout,
|
|
env = env,
|
|
**windows_hidden_subprocess_kwargs(),
|
|
)
|
|
if check and result.returncode != 0:
|
|
raise subprocess.CalledProcessError(
|
|
result.returncode, command, result.stdout, result.stderr
|
|
)
|
|
return result
|
|
|
|
|
|
def detect_host() -> HostInfo:
|
|
system = platform.system()
|
|
machine = platform.machine().lower()
|
|
is_windows = system == "Windows"
|
|
is_linux = system == "Linux"
|
|
is_macos = system == "Darwin"
|
|
is_x86_64 = machine in {"x86_64", "amd64"}
|
|
is_arm64 = machine in {"arm64", "aarch64"}
|
|
|
|
nvidia_smi = shutil.which("nvidia-smi")
|
|
driver_cuda_version = None
|
|
compute_caps: list[str] = []
|
|
visible_cuda_devices = os.environ.get("CUDA_VISIBLE_DEVICES")
|
|
visible_device_tokens = parse_cuda_visible_devices(visible_cuda_devices)
|
|
has_physical_nvidia = False
|
|
has_usable_nvidia = False
|
|
if nvidia_smi:
|
|
# Require `nvidia-smi -L` to actually list a GPU before treating the
|
|
# host as NVIDIA. The banner text "NVIDIA-SMI ..." is printed even
|
|
# when the command fails to communicate with the driver (e.g. stale
|
|
# container leftovers), which would otherwise misclassify an AMD
|
|
# ROCm host as NVIDIA and short-circuit the ROCm path.
|
|
try:
|
|
listing = run_capture([nvidia_smi, "-L"], timeout = 20)
|
|
gpu_lines = [
|
|
line for line in listing.stdout.splitlines() if line.startswith("GPU ")
|
|
]
|
|
if gpu_lines:
|
|
has_physical_nvidia = True
|
|
has_usable_nvidia = visible_device_tokens != []
|
|
except Exception:
|
|
pass
|
|
|
|
try:
|
|
result = run_capture([nvidia_smi], timeout = 20)
|
|
merged = "\n".join(part for part in (result.stdout, result.stderr) if part)
|
|
# Newer NVIDIA drivers (e.g. 610.x on Windows) print
|
|
# "CUDA UMD Version: X.Y" instead of the legacy
|
|
# "CUDA Version: X.Y"; accept both spellings.
|
|
cuda_match = re.search(
|
|
r"CUDA(?: UMD)? Version:\s*(\d+)\.(\d+)",
|
|
merged,
|
|
)
|
|
if cuda_match is not None:
|
|
driver_cuda_version = (
|
|
int(cuda_match.group(1)),
|
|
int(cuda_match.group(2)),
|
|
)
|
|
except Exception:
|
|
pass
|
|
|
|
try:
|
|
caps = run_capture(
|
|
[
|
|
nvidia_smi,
|
|
"--query-gpu=index,uuid,compute_cap",
|
|
"--format=csv,noheader",
|
|
],
|
|
timeout = 20,
|
|
)
|
|
visible_gpu_rows: list[tuple[str, str, str]] = []
|
|
for raw in caps.stdout.splitlines():
|
|
parts = [part.strip() for part in raw.split(",")]
|
|
if len(parts) != 3:
|
|
continue
|
|
index, uuid, cap = parts
|
|
visible_gpu_row = select_visible_gpu_rows(
|
|
[(index, uuid, cap)],
|
|
visible_device_tokens,
|
|
)
|
|
if not visible_gpu_row:
|
|
continue
|
|
visible_gpu_rows.extend(visible_gpu_row)
|
|
normalized_cap = normalize_compute_cap(cap)
|
|
if normalized_cap is None:
|
|
continue
|
|
if normalized_cap not in compute_caps:
|
|
compute_caps.append(normalized_cap)
|
|
|
|
if visible_gpu_rows:
|
|
has_usable_nvidia = True
|
|
# Older nvidia-smi versions (pre -L support) hit the
|
|
# except in the first try block but still succeed here,
|
|
# leaving has_physical_nvidia unset. Mirror the -L path
|
|
# so downstream diagnostics on line ~4390 still run.
|
|
if not has_physical_nvidia:
|
|
has_physical_nvidia = True
|
|
elif visible_device_tokens == []:
|
|
has_usable_nvidia = False
|
|
elif supports_explicit_visible_device_matching(visible_device_tokens):
|
|
has_usable_nvidia = False
|
|
elif has_physical_nvidia:
|
|
has_usable_nvidia = True
|
|
except Exception:
|
|
pass
|
|
|
|
# Detect AMD ROCm (HIP) -- require actual GPU, not just tools installed
|
|
|
|
def _amd_smi_has_gpu(stdout: str) -> bool:
|
|
"""Check for 'GPU: <number>' data rows, not just a table header."""
|
|
return bool(re.search(r"(?im)^gpu\s*[:\[]\s*\d", stdout))
|
|
|
|
has_rocm = False
|
|
if is_linux:
|
|
for _cmd, _check in (
|
|
# rocminfo: look for a real gfx GPU id (3-4 chars, nonzero first digit).
|
|
# gfx000 is the CPU agent; ROCm 6.1+ also emits generic ISA lines like
|
|
# "gfx11-generic" or "gfx9-4-generic" which only have 1-2 digits before
|
|
# the dash and must not be treated as a real GPU.
|
|
(
|
|
["rocminfo"],
|
|
lambda out: bool(re.search(r"gfx[1-9][0-9a-z]{2,3}", out.lower())),
|
|
),
|
|
(["amd-smi", "list"], _amd_smi_has_gpu),
|
|
):
|
|
_exe = shutil.which(_cmd[0])
|
|
if not _exe:
|
|
continue
|
|
try:
|
|
_result = run_capture([_exe, *_cmd[1:]], timeout = 10)
|
|
except Exception:
|
|
continue
|
|
if _result.returncode == 0 and _result.stdout.strip():
|
|
if _check(_result.stdout):
|
|
has_rocm = True
|
|
break
|
|
elif is_windows:
|
|
# Windows: prefer active probes that validate GPU presence.
|
|
# hipinfo / amd-smi are often NOT on PATH -- the HIP SDK installer
|
|
# sets HIP_PATH / ROCM_PATH but does not always add the bin dir to
|
|
# the system PATH. Mirror setup.ps1's fallback: check the env-var
|
|
# bin dirs before giving up so that `has_rocm` is not silently False
|
|
# on machines where the PATH is not yet updated.
|
|
def _resolve_exe(name: str) -> str | None:
|
|
"""Return full path to `name`, checking PATH then HIP_PATH/ROCM_PATH bin."""
|
|
found = shutil.which(name)
|
|
if found:
|
|
return found
|
|
for _env in ("HIP_PATH", "ROCM_PATH"):
|
|
_root = os.environ.get(_env)
|
|
if _root:
|
|
_candidate = os.path.join(_root, "bin", f"{name}.exe")
|
|
if os.path.isfile(_candidate):
|
|
return _candidate
|
|
return None
|
|
|
|
for _cmd, _check in (
|
|
(["hipinfo"], lambda out: "gcnarchname" in out.lower()),
|
|
(["amd-smi", "list"], _amd_smi_has_gpu),
|
|
):
|
|
_exe = _resolve_exe(_cmd[0])
|
|
if not _exe:
|
|
continue
|
|
try:
|
|
_result = run_capture([_exe, *_cmd[1:]], timeout = 10)
|
|
except Exception:
|
|
continue
|
|
if _result.returncode == 0 and _result.stdout.strip():
|
|
if _check(_result.stdout):
|
|
has_rocm = True
|
|
break
|
|
# Note: amdhip64.dll presence alone is NOT treated as GPU evidence
|
|
# since the HIP SDK can be installed without an AMD GPU.
|
|
|
|
return HostInfo(
|
|
system = system,
|
|
machine = machine,
|
|
is_windows = is_windows,
|
|
is_linux = is_linux,
|
|
is_macos = is_macos,
|
|
is_x86_64 = is_x86_64,
|
|
is_arm64 = is_arm64,
|
|
nvidia_smi = nvidia_smi,
|
|
driver_cuda_version = driver_cuda_version,
|
|
compute_caps = compute_caps,
|
|
visible_cuda_devices = visible_cuda_devices,
|
|
has_physical_nvidia = has_physical_nvidia,
|
|
has_usable_nvidia = has_usable_nvidia,
|
|
has_rocm = has_rocm,
|
|
)
|
|
|
|
|
|
def pick_windows_cuda_runtime(host: HostInfo) -> str | None:
|
|
if not host.driver_cuda_version:
|
|
return None
|
|
major, minor = host.driver_cuda_version
|
|
if major > 13 or (major == 13): # and minor >= 1):
|
|
return "13.1"
|
|
if major > 12 or (major == 12 and minor >= 4):
|
|
return "12.4"
|
|
return None
|
|
|
|
|
|
def compatible_linux_runtime_lines(host: HostInfo) -> list[str]:
|
|
if not host.driver_cuda_version:
|
|
return []
|
|
major, _minor = host.driver_cuda_version
|
|
if major >= 13:
|
|
return ["cuda13", "cuda12"]
|
|
if major >= 12:
|
|
return ["cuda12"]
|
|
return []
|
|
|
|
|
|
def windows_runtime_line_info() -> dict[str, tuple[str, ...]]:
|
|
return {
|
|
"cuda13": ("cudart64_13*.dll", "cublas64_13*.dll", "cublasLt64_13*.dll"),
|
|
"cuda12": ("cudart64_12*.dll", "cublas64_12*.dll", "cublasLt64_12*.dll"),
|
|
}
|
|
|
|
|
|
def detected_windows_runtime_lines() -> tuple[list[str], dict[str, list[str]]]:
|
|
dirs = windows_runtime_dirs()
|
|
detected: list[str] = []
|
|
runtime_dirs: dict[str, list[str]] = {}
|
|
for runtime_line, required_patterns in windows_runtime_line_info().items():
|
|
matching_dirs = windows_runtime_dirs_for_patterns(required_patterns, dirs)
|
|
if matching_dirs:
|
|
detected.append(runtime_line)
|
|
runtime_dirs[runtime_line] = matching_dirs
|
|
return detected, runtime_dirs
|
|
|
|
|
|
def compatible_windows_runtime_lines(host: HostInfo) -> list[str]:
|
|
driver_runtime = pick_windows_cuda_runtime(host)
|
|
if driver_runtime == "13.1":
|
|
return ["cuda13", "cuda12"]
|
|
if driver_runtime == "12.4":
|
|
return ["cuda12"]
|
|
return []
|
|
|
|
|
|
def runtime_line_from_cuda_version(cuda_version: str | None) -> str | None:
|
|
if not cuda_version:
|
|
return None
|
|
raw = str(cuda_version).strip()
|
|
if not raw:
|
|
return None
|
|
major, _, _ = raw.partition(".")
|
|
if major == "12":
|
|
return "cuda12"
|
|
if major == "13":
|
|
return "cuda13"
|
|
return None
|
|
|
|
|
|
def detect_torch_cuda_runtime_preference(host: HostInfo) -> CudaRuntimePreference:
|
|
selection_log: list[str] = []
|
|
if host.is_macos:
|
|
selection_log.append("torch_cuda_preference: skipped on macOS")
|
|
return CudaRuntimePreference(runtime_line = None, selection_log = selection_log)
|
|
if not (host.has_usable_nvidia and (host.is_linux or host.is_windows)):
|
|
selection_log.append(
|
|
"torch_cuda_preference: skipped because CUDA host prerequisites were not met"
|
|
)
|
|
return CudaRuntimePreference(runtime_line = None, selection_log = selection_log)
|
|
|
|
try:
|
|
import torch
|
|
except Exception as exc:
|
|
selection_log.append(f"torch_cuda_preference: import failed: {exc}")
|
|
return CudaRuntimePreference(runtime_line = None, selection_log = selection_log)
|
|
|
|
cuda_version = getattr(getattr(torch, "version", None), "cuda", None)
|
|
if not isinstance(cuda_version, str) or not cuda_version.strip():
|
|
selection_log.append(
|
|
"torch_cuda_preference: torch.version.cuda missing; skipping Torch shortcut"
|
|
)
|
|
return CudaRuntimePreference(runtime_line = None, selection_log = selection_log)
|
|
|
|
try:
|
|
cuda_available = bool(torch.cuda.is_available())
|
|
except Exception as exc:
|
|
selection_log.append(
|
|
f"torch_cuda_preference: torch.cuda.is_available() failed: {exc}"
|
|
)
|
|
return CudaRuntimePreference(runtime_line = None, selection_log = selection_log)
|
|
|
|
if not cuda_available:
|
|
selection_log.append(
|
|
"torch_cuda_preference: torch.cuda.is_available() returned False; falling back to normal selection"
|
|
)
|
|
return CudaRuntimePreference(runtime_line = None, selection_log = selection_log)
|
|
|
|
runtime_line = runtime_line_from_cuda_version(cuda_version)
|
|
if runtime_line is None:
|
|
selection_log.append(
|
|
f"torch_cuda_preference: unsupported torch.version.cuda={cuda_version}; falling back to normal selection"
|
|
)
|
|
return CudaRuntimePreference(runtime_line = None, selection_log = selection_log)
|
|
|
|
selection_log.append(
|
|
"torch_cuda_preference: selected runtime_line="
|
|
f"{runtime_line} from torch.version.cuda={cuda_version}"
|
|
)
|
|
return CudaRuntimePreference(runtime_line = runtime_line, selection_log = selection_log)
|
|
|
|
|
|
def windows_cuda_attempts(
|
|
host: HostInfo,
|
|
llama_tag: str,
|
|
upstream_assets: dict[str, str],
|
|
preferred_runtime_line: str | None,
|
|
selection_preamble: Iterable[str] = (),
|
|
) -> list[AssetChoice]:
|
|
selection_log = list(selection_preamble)
|
|
runtime_by_line = {"cuda12": "12.4", "cuda13": "13.1"}
|
|
driver_runtime = pick_windows_cuda_runtime(host)
|
|
detected_runtime_lines, runtime_dirs = detected_windows_runtime_lines()
|
|
compatible_runtime_lines = compatible_windows_runtime_lines(host)
|
|
normal_runtime_lines: list[str]
|
|
if detected_runtime_lines:
|
|
normal_runtime_lines = [
|
|
line for line in compatible_runtime_lines if line in detected_runtime_lines
|
|
]
|
|
else:
|
|
normal_runtime_lines = compatible_runtime_lines
|
|
selection_log.append(
|
|
"windows_cuda_selection: driver_runtime="
|
|
+ (driver_runtime if driver_runtime else "unknown")
|
|
)
|
|
selection_log.append(
|
|
"windows_cuda_selection: detected_runtime_lines="
|
|
+ (",".join(detected_runtime_lines) if detected_runtime_lines else "none")
|
|
)
|
|
for runtime_line in ("cuda13", "cuda12"):
|
|
selection_log.append(
|
|
"windows_cuda_selection: runtime_dirs "
|
|
f"{runtime_line}="
|
|
+ (
|
|
",".join(runtime_dirs.get(runtime_line, []))
|
|
if runtime_dirs.get(runtime_line)
|
|
else "none"
|
|
)
|
|
)
|
|
if detected_runtime_lines:
|
|
selection_log.append(
|
|
"windows_cuda_selection: host_runtime_order="
|
|
+ (",".join(normal_runtime_lines) if normal_runtime_lines else "none")
|
|
)
|
|
else:
|
|
selection_log.append(
|
|
"windows_cuda_selection: no CUDA runtime DLL line detected; falling back to driver order"
|
|
)
|
|
if not normal_runtime_lines:
|
|
if detected_runtime_lines:
|
|
selection_log.append(
|
|
"windows_cuda_selection: detected CUDA runtime DLLs were incompatible with the reported driver"
|
|
)
|
|
fallback_runtime_lines = (
|
|
["cuda13", "cuda12"]
|
|
if driver_runtime == "13.1"
|
|
else (["cuda12"] if driver_runtime == "12.4" else [])
|
|
)
|
|
normal_runtime_lines = fallback_runtime_lines
|
|
|
|
runtime_order: list[str] = []
|
|
if preferred_runtime_line and preferred_runtime_line in normal_runtime_lines:
|
|
runtime_order.append(preferred_runtime_line)
|
|
selection_log.append(
|
|
"windows_cuda_selection: torch_preferred_runtime_line="
|
|
f"{preferred_runtime_line} reordered_attempts"
|
|
)
|
|
elif preferred_runtime_line:
|
|
selection_log.append(
|
|
"windows_cuda_selection: torch_preferred_runtime_line="
|
|
f"{preferred_runtime_line} unavailable_or_incompatible"
|
|
)
|
|
else:
|
|
selection_log.append(
|
|
"windows_cuda_selection: no Torch runtime preference available"
|
|
)
|
|
|
|
runtime_order.extend(
|
|
runtime_line
|
|
for runtime_line in normal_runtime_lines
|
|
if runtime_line not in runtime_order
|
|
)
|
|
selection_log.append(
|
|
"windows_cuda_selection: normal_runtime_order="
|
|
+ (",".join(normal_runtime_lines) if normal_runtime_lines else "none")
|
|
)
|
|
selection_log.append(
|
|
"windows_cuda_selection: attempt_runtime_order="
|
|
+ (",".join(runtime_order) if runtime_order else "none")
|
|
)
|
|
|
|
attempts: list[AssetChoice] = []
|
|
for runtime_line in runtime_order:
|
|
runtime = runtime_by_line[runtime_line]
|
|
selected_name = None
|
|
asset_url = None
|
|
for candidate_name in windows_cuda_upstream_asset_names(llama_tag, runtime):
|
|
asset_url = upstream_assets.get(candidate_name)
|
|
if asset_url:
|
|
selected_name = candidate_name
|
|
break
|
|
if not asset_url or not selected_name:
|
|
selection_log.append(
|
|
"windows_cuda_selection: skip missing assets "
|
|
+ ",".join(windows_cuda_upstream_asset_names(llama_tag, runtime))
|
|
)
|
|
continue
|
|
# Pair the cudart bundle when upstream ships it. Without this
|
|
# the binary needs a system CUDA toolkit on PATH at runtime
|
|
# (#5106). Only pair when the selected main archive is the
|
|
# binary archive, not the cudart archive itself.
|
|
runtime_archive_name: str | None = None
|
|
runtime_archive_url: str | None = None
|
|
if selected_name.startswith("llama-"):
|
|
cudart_name = f"cudart-llama-bin-win-cuda-{runtime}-x64.zip"
|
|
cudart_url = upstream_assets.get(cudart_name)
|
|
if cudart_url and cudart_url != asset_url:
|
|
runtime_archive_name = cudart_name
|
|
runtime_archive_url = cudart_url
|
|
attempt_log = list(selection_log) + [
|
|
f"windows_cuda_selection: selected {selected_name} runtime={runtime}"
|
|
]
|
|
if runtime_archive_name:
|
|
attempt_log.append(
|
|
f"windows_cuda_selection: paired runtime archive {runtime_archive_name}"
|
|
)
|
|
else:
|
|
attempt_log.append(
|
|
"windows_cuda_selection: no paired runtime archive found; "
|
|
"binary will rely on a system CUDA toolkit at runtime"
|
|
)
|
|
attempts.append(
|
|
AssetChoice(
|
|
repo = UPSTREAM_REPO,
|
|
tag = llama_tag,
|
|
name = selected_name,
|
|
url = asset_url,
|
|
source_label = "upstream",
|
|
install_kind = "windows-cuda",
|
|
runtime_line = runtime_line,
|
|
runtime_name = runtime_archive_name,
|
|
runtime_url = runtime_archive_url,
|
|
selection_log = attempt_log,
|
|
)
|
|
)
|
|
return attempts
|
|
|
|
|
|
def published_windows_cuda_attempts(
|
|
host: HostInfo,
|
|
release: PublishedReleaseBundle,
|
|
preferred_runtime_line: str | None,
|
|
selection_preamble: Iterable[str] = (),
|
|
) -> list[AssetChoice]:
|
|
selection_log = list(release.selection_log) + list(selection_preamble)
|
|
runtime_by_line = {"cuda12": "12.4", "cuda13": "13.1"}
|
|
runtime_order = windows_cuda_attempts(
|
|
host,
|
|
release.upstream_tag,
|
|
{
|
|
f"llama-{release.upstream_tag}-bin-win-cuda-{runtime}-x64.zip": "published"
|
|
for runtime in runtime_by_line.values()
|
|
},
|
|
preferred_runtime_line,
|
|
selection_log,
|
|
)
|
|
published_artifacts = [
|
|
artifact
|
|
for artifact in release.artifacts
|
|
if artifact.install_kind == "windows-cuda"
|
|
]
|
|
artifacts_by_runtime: dict[str, list[PublishedLlamaArtifact]] = {}
|
|
for artifact in published_artifacts:
|
|
if not artifact.runtime_line:
|
|
continue
|
|
artifacts_by_runtime.setdefault(artifact.runtime_line, []).append(artifact)
|
|
|
|
attempts: list[AssetChoice] = []
|
|
for ordered_attempt in runtime_order:
|
|
runtime_line = ordered_attempt.runtime_line
|
|
if not runtime_line:
|
|
continue
|
|
candidates = sorted(
|
|
artifacts_by_runtime.get(runtime_line, []),
|
|
key = lambda artifact: (artifact.rank, artifact.asset_name),
|
|
)
|
|
for artifact in candidates:
|
|
asset_url = release.assets.get(artifact.asset_name)
|
|
if not asset_url:
|
|
continue
|
|
# See windows_cuda_attempts: pair the cudart bundle.
|
|
runtime_archive_name: str | None = None
|
|
runtime_archive_url: str | None = None
|
|
if artifact.asset_name.startswith("llama-"):
|
|
runtime = runtime_by_line[runtime_line]
|
|
cudart_name = f"cudart-llama-bin-win-cuda-{runtime}-x64.zip"
|
|
cudart_url = release.assets.get(cudart_name)
|
|
if cudart_url and cudart_url != asset_url:
|
|
runtime_archive_name = cudart_name
|
|
runtime_archive_url = cudart_url
|
|
attempt_log = list(ordered_attempt.selection_log or []) + [
|
|
"windows_cuda_selection: selected published asset "
|
|
f"{artifact.asset_name} for runtime_line={runtime_line}"
|
|
]
|
|
if runtime_archive_name:
|
|
attempt_log.append(
|
|
f"windows_cuda_selection: paired published runtime archive {runtime_archive_name}"
|
|
)
|
|
attempts.append(
|
|
AssetChoice(
|
|
repo = release.repo,
|
|
tag = release.release_tag,
|
|
name = artifact.asset_name,
|
|
url = asset_url,
|
|
source_label = "published",
|
|
install_kind = "windows-cuda",
|
|
runtime_line = runtime_line,
|
|
runtime_name = runtime_archive_name,
|
|
runtime_url = runtime_archive_url,
|
|
selection_log = attempt_log,
|
|
)
|
|
)
|
|
break
|
|
return attempts
|
|
|
|
|
|
def resolve_windows_cuda_choices(
|
|
host: HostInfo, llama_tag: str, upstream_assets: dict[str, str]
|
|
) -> list[AssetChoice]:
|
|
torch_preference = detect_torch_cuda_runtime_preference(host)
|
|
attempts = windows_cuda_attempts(
|
|
host,
|
|
llama_tag,
|
|
upstream_assets,
|
|
torch_preference.runtime_line,
|
|
torch_preference.selection_log,
|
|
)
|
|
return attempts
|
|
|
|
|
|
def resolve_linux_cuda_choice(
|
|
host: HostInfo, release: PublishedReleaseBundle
|
|
) -> LinuxCudaSelection:
|
|
torch_preference = detect_torch_cuda_runtime_preference(host)
|
|
selection = linux_cuda_choice_from_release(
|
|
host,
|
|
release,
|
|
preferred_runtime_line = torch_preference.runtime_line,
|
|
selection_preamble = torch_preference.selection_log,
|
|
)
|
|
if selection is not None:
|
|
return selection
|
|
raise PrebuiltFallback("no compatible published Linux CUDA bundle was found")
|
|
|
|
|
|
def published_asset_choice_for_kind(
|
|
release: PublishedReleaseBundle,
|
|
install_kind: str,
|
|
) -> AssetChoice | None:
|
|
candidates = sorted(
|
|
(
|
|
artifact
|
|
for artifact in release.artifacts
|
|
if artifact.install_kind == install_kind
|
|
),
|
|
key = lambda artifact: (artifact.rank, artifact.asset_name),
|
|
)
|
|
for artifact in candidates:
|
|
asset_url = release.assets.get(artifact.asset_name)
|
|
if not asset_url:
|
|
continue
|
|
return AssetChoice(
|
|
repo = release.repo,
|
|
tag = release.release_tag,
|
|
name = artifact.asset_name,
|
|
url = asset_url,
|
|
source_label = "published",
|
|
install_kind = install_kind,
|
|
runtime_line = artifact.runtime_line,
|
|
selection_log = list(release.selection_log)
|
|
+ [
|
|
f"published_selection: selected {artifact.asset_name} install_kind={install_kind}"
|
|
],
|
|
)
|
|
return None
|
|
|
|
|
|
def _detect_host_rocm_version() -> tuple[int, int] | None:
|
|
"""Return (major, minor) of the installed ROCm runtime, or None.
|
|
|
|
Best-effort read from /opt/rocm/.info/version, amd-smi version, and
|
|
hipconfig --version. Used to pick a compatible upstream llama.cpp
|
|
ROCm prebuilt rather than always taking the numerically newest one
|
|
(which can be newer than the host runtime).
|
|
"""
|
|
rocm_root = os.environ.get("ROCM_PATH") or "/opt/rocm"
|
|
for path in (
|
|
os.path.join(rocm_root, ".info", "version"),
|
|
os.path.join(rocm_root, "lib", "rocm_version"),
|
|
):
|
|
try:
|
|
with open(path) as fh:
|
|
parts = fh.read().strip().split("-")[0].split(".")
|
|
# Explicit length guard avoids relying on the broad except
|
|
# below to swallow IndexError when the version file contains
|
|
# a single component (e.g. "6\n" on a partial install).
|
|
if len(parts) >= 2:
|
|
return int(parts[0]), int(parts[1])
|
|
except Exception:
|
|
pass
|
|
amd_smi = shutil.which("amd-smi")
|
|
if amd_smi:
|
|
try:
|
|
result = subprocess.run(
|
|
[amd_smi, "version"],
|
|
stdout = subprocess.PIPE,
|
|
stderr = subprocess.DEVNULL,
|
|
text = True,
|
|
timeout = 5,
|
|
)
|
|
if result.returncode == 0:
|
|
m = re.search(r"ROCm version:\s*(\d+)\.(\d+)", result.stdout)
|
|
if m:
|
|
return int(m.group(1)), int(m.group(2))
|
|
except Exception:
|
|
pass
|
|
hipconfig = shutil.which("hipconfig")
|
|
if hipconfig:
|
|
try:
|
|
result = subprocess.run(
|
|
[hipconfig, "--version"],
|
|
stdout = subprocess.PIPE,
|
|
stderr = subprocess.DEVNULL,
|
|
text = True,
|
|
timeout = 5,
|
|
)
|
|
if result.returncode == 0:
|
|
raw = (result.stdout or "").strip().split("\n")[0]
|
|
parts = raw.split(".")
|
|
if (
|
|
len(parts) >= 2
|
|
and parts[0].isdigit()
|
|
and parts[1].split("-")[0].isdigit()
|
|
):
|
|
return int(parts[0]), int(parts[1].split("-")[0])
|
|
except Exception:
|
|
pass
|
|
|
|
# Distro package-manager fallbacks. Mirrors install.sh::get_torch_index_url
|
|
# and _detect_rocm_version() in install_python_stack.py so package-managed
|
|
# ROCm hosts without /opt/rocm/.info/version still report a usable version
|
|
# and the <= host version filter in resolve_upstream_asset_choice picks
|
|
# the correct upstream prebuilt instead of the newest-regardless fallback.
|
|
for _cmd in (
|
|
["dpkg-query", "-W", "-f=${Version}\n", "rocm-core"],
|
|
["rpm", "-q", "--qf", "%{VERSION}\n", "rocm-core"],
|
|
):
|
|
_exe = shutil.which(_cmd[0])
|
|
if not _exe:
|
|
continue
|
|
try:
|
|
_result = subprocess.run(
|
|
[_exe, *_cmd[1:]],
|
|
stdout = subprocess.PIPE,
|
|
stderr = subprocess.DEVNULL,
|
|
text = True,
|
|
timeout = 5,
|
|
)
|
|
except Exception:
|
|
continue
|
|
if _result.returncode != 0 or not _result.stdout.strip():
|
|
continue
|
|
_raw = _result.stdout.strip()
|
|
# dpkg can prepend an epoch ("1:6.3.0-1"); strip it before parsing.
|
|
_raw = re.sub(r"^\d+:", "", _raw)
|
|
_m = re.match(r"(\d+)[.-](\d+)", _raw)
|
|
if _m:
|
|
return int(_m.group(1)), int(_m.group(2))
|
|
return None
|
|
|
|
|
|
def resolve_upstream_asset_choice(host: HostInfo, llama_tag: str) -> AssetChoice:
|
|
upstream_assets = github_release_assets(UPSTREAM_REPO, llama_tag)
|
|
if host.is_linux and host.is_x86_64:
|
|
# AMD ROCm: try upstream ROCm prebuilt first, then fall back to source build.
|
|
# Source build (via setup.sh) compiles with -DGGML_HIP=ON and auto-detects
|
|
# the exact GPU target via rocminfo, which is more reliable for consumer
|
|
# GPUs (e.g. gfx1151) that may not be in the prebuilt.
|
|
if host.has_rocm and not host.has_usable_nvidia:
|
|
# Scan upstream assets for any rocm-<version> prebuilt. When the
|
|
# host ROCm runtime version is known, pick the newest candidate
|
|
# whose major.minor is <= host version -- otherwise a ROCm 6.4
|
|
# host would download the rocm-7.2 tarball, fail preflight, and
|
|
# fall back to a source build even though a compatible 6.4
|
|
# prebuilt exists. If no compatible candidate matches (e.g. host
|
|
# runtime is older than every published prebuilt), fall back to
|
|
# the numerically newest so we at least try something.
|
|
_rocm_pattern = re.compile(
|
|
rf"llama-{re.escape(llama_tag)}-bin-ubuntu-rocm-([0-9]+(?:\.[0-9]+)*)-x64\.tar\.gz"
|
|
)
|
|
rocm_candidates: list[tuple[tuple[int, ...], str]] = []
|
|
for _name in upstream_assets:
|
|
_m = _rocm_pattern.match(_name)
|
|
if _m is None:
|
|
continue
|
|
_parts = tuple(int(p) for p in _m.group(1).split("."))
|
|
rocm_candidates.append((_parts, _name))
|
|
rocm_candidates.sort(reverse = True)
|
|
_host_rocm_version = _detect_host_rocm_version()
|
|
_compatible: list[tuple[tuple[int, ...], str]] = rocm_candidates
|
|
if _host_rocm_version is not None:
|
|
_compatible = [
|
|
item
|
|
for item in rocm_candidates
|
|
if item[0][:2] <= _host_rocm_version
|
|
]
|
|
if rocm_candidates and not _compatible:
|
|
# Fall back to the newest candidate so a source build is
|
|
# not forced when the host runtime is older than every
|
|
# published prebuilt: preflight will still catch a true
|
|
# incompatibility and trigger a fallback.
|
|
_compatible = rocm_candidates[:1]
|
|
if _compatible:
|
|
rocm_name = _compatible[0][1]
|
|
if _host_rocm_version is not None:
|
|
log(
|
|
f"AMD ROCm {_host_rocm_version[0]}.{_host_rocm_version[1]} "
|
|
f"detected -- trying upstream prebuilt {rocm_name}"
|
|
)
|
|
else:
|
|
log(f"AMD ROCm detected -- trying upstream prebuilt {rocm_name}")
|
|
log(
|
|
"Note: if your ROCm runtime version differs significantly, "
|
|
"this may fail preflight and fall back to a source build (safe)"
|
|
)
|
|
return AssetChoice(
|
|
repo = UPSTREAM_REPO,
|
|
tag = llama_tag,
|
|
name = rocm_name,
|
|
url = upstream_assets[rocm_name],
|
|
source_label = "upstream",
|
|
install_kind = "linux-rocm",
|
|
)
|
|
# No ROCm prebuilt available -- fall back to source build
|
|
raise PrebuiltFallback(
|
|
"AMD ROCm detected but no upstream ROCm prebuilt found; "
|
|
"falling back to source build with HIP support"
|
|
)
|
|
|
|
upstream_name = f"llama-{llama_tag}-bin-ubuntu-x64.tar.gz"
|
|
if upstream_name not in upstream_assets:
|
|
raise PrebuiltFallback("upstream Linux CPU asset was not found")
|
|
return AssetChoice(
|
|
repo = UPSTREAM_REPO,
|
|
tag = llama_tag,
|
|
name = upstream_name,
|
|
url = upstream_assets[upstream_name],
|
|
source_label = "upstream",
|
|
install_kind = "linux-cpu",
|
|
)
|
|
|
|
if host.is_windows and host.is_x86_64:
|
|
if host.has_usable_nvidia:
|
|
attempts = resolve_windows_cuda_choices(host, llama_tag, upstream_assets)
|
|
if attempts:
|
|
return attempts[0]
|
|
raise PrebuiltFallback("no compatible Windows CUDA asset was found")
|
|
|
|
# AMD ROCm on Windows: try HIP prebuilt
|
|
if host.has_rocm:
|
|
hip_name = f"llama-{llama_tag}-bin-win-hip-radeon-x64.zip"
|
|
if hip_name in upstream_assets:
|
|
log(
|
|
f"AMD ROCm detected on Windows -- trying upstream HIP prebuilt {hip_name}"
|
|
)
|
|
return AssetChoice(
|
|
repo = UPSTREAM_REPO,
|
|
tag = llama_tag,
|
|
name = hip_name,
|
|
url = upstream_assets[hip_name],
|
|
source_label = "upstream",
|
|
install_kind = "windows-hip",
|
|
)
|
|
log(
|
|
"AMD ROCm detected on Windows but no HIP prebuilt found -- falling back to CPU"
|
|
)
|
|
|
|
upstream_name = f"llama-{llama_tag}-bin-win-cpu-x64.zip"
|
|
if upstream_name not in upstream_assets:
|
|
raise PrebuiltFallback("upstream Windows CPU asset was not found")
|
|
return AssetChoice(
|
|
repo = UPSTREAM_REPO,
|
|
tag = llama_tag,
|
|
name = upstream_name,
|
|
url = upstream_assets[upstream_name],
|
|
source_label = "upstream",
|
|
install_kind = "windows-cpu",
|
|
)
|
|
|
|
if host.is_macos and host.is_arm64:
|
|
upstream_name = f"llama-{llama_tag}-bin-macos-arm64.tar.gz"
|
|
if upstream_name not in upstream_assets:
|
|
raise PrebuiltFallback("upstream macOS arm64 asset was not found")
|
|
return AssetChoice(
|
|
repo = UPSTREAM_REPO,
|
|
tag = llama_tag,
|
|
name = upstream_name,
|
|
url = upstream_assets[upstream_name],
|
|
source_label = "upstream",
|
|
install_kind = "macos-arm64",
|
|
)
|
|
|
|
if host.is_macos and host.is_x86_64:
|
|
upstream_name = f"llama-{llama_tag}-bin-macos-x64.tar.gz"
|
|
if upstream_name not in upstream_assets:
|
|
raise PrebuiltFallback("upstream macOS x64 asset was not found")
|
|
return AssetChoice(
|
|
repo = UPSTREAM_REPO,
|
|
tag = llama_tag,
|
|
name = upstream_name,
|
|
url = upstream_assets[upstream_name],
|
|
source_label = "upstream",
|
|
install_kind = "macos-x64",
|
|
)
|
|
|
|
raise PrebuiltFallback(
|
|
f"no prebuilt policy exists for {host.system} {host.machine}"
|
|
)
|
|
|
|
|
|
def resolve_asset_choice(host: HostInfo, llama_tag: str) -> AssetChoice:
|
|
if host.is_linux and host.is_x86_64 and host.has_usable_nvidia:
|
|
raise PrebuiltFallback(
|
|
"Linux CUDA installs require a compatible published bundle; upstream fallback is not available"
|
|
)
|
|
return resolve_upstream_asset_choice(host, llama_tag)
|
|
|
|
|
|
def resolve_release_asset_choice(
|
|
host: HostInfo,
|
|
llama_tag: str,
|
|
release: PublishedReleaseBundle,
|
|
checksums: ApprovedReleaseChecksums,
|
|
) -> list[AssetChoice]:
|
|
if host.is_windows and host.is_x86_64 and host.has_usable_nvidia:
|
|
torch_preference = detect_torch_cuda_runtime_preference(host)
|
|
published_attempts = published_windows_cuda_attempts(
|
|
host,
|
|
release,
|
|
torch_preference.runtime_line,
|
|
torch_preference.selection_log,
|
|
)
|
|
if published_attempts:
|
|
try:
|
|
return apply_approved_hashes(published_attempts, checksums)
|
|
except PrebuiltFallback as exc:
|
|
log(
|
|
"published Windows CUDA assets ignored for install planning: "
|
|
f"{release.repo}@{release.release_tag} ({exc})"
|
|
)
|
|
upstream_assets = github_release_assets(UPSTREAM_REPO, llama_tag)
|
|
return apply_approved_hashes(
|
|
resolve_windows_cuda_choices(host, llama_tag, upstream_assets),
|
|
checksums,
|
|
)
|
|
|
|
published_choice: AssetChoice | None = None
|
|
if host.is_windows and host.is_x86_64:
|
|
# AMD Windows hosts should prefer a hash-approved published
|
|
# Windows HIP bundle when one exists, but otherwise fall through
|
|
# to resolve_asset_choice() so the upstream HIP prebuilt is
|
|
# tried before the CPU fallback. Hard-pinning the published
|
|
# windows-cpu bundle here would make the new HIP path
|
|
# unreachable.
|
|
if host.has_rocm:
|
|
published_choice = published_asset_choice_for_kind(release, "windows-hip")
|
|
else:
|
|
published_choice = published_asset_choice_for_kind(release, "windows-cpu")
|
|
elif host.is_macos and host.is_arm64:
|
|
published_choice = published_asset_choice_for_kind(release, "macos-arm64")
|
|
elif host.is_macos and host.is_x86_64:
|
|
published_choice = published_asset_choice_for_kind(release, "macos-x64")
|
|
|
|
if published_choice is not None:
|
|
try:
|
|
return apply_approved_hashes([published_choice], checksums)
|
|
except PrebuiltFallback as exc:
|
|
log(
|
|
"published platform asset ignored for install planning: "
|
|
f"{release.repo}@{release.release_tag} {published_choice.name} ({exc})"
|
|
)
|
|
|
|
return apply_approved_hashes([resolve_asset_choice(host, llama_tag)], checksums)
|
|
|
|
|
|
def extract_archive(archive_path: Path, destination: Path) -> None:
|
|
def safe_extract_path(base: Path, member_name: str) -> Path:
|
|
normalized = member_name.replace("\\", "/")
|
|
member_path = Path(normalized)
|
|
if member_path.is_absolute():
|
|
raise PrebuiltFallback(
|
|
f"archive member used an absolute path: {member_name}"
|
|
)
|
|
|
|
target = (base / member_path).resolve()
|
|
base_resolved = base.resolve()
|
|
try:
|
|
target.relative_to(base_resolved)
|
|
except ValueError as exc:
|
|
raise PrebuiltFallback(
|
|
f"archive member escaped destination: {member_name}"
|
|
) from exc
|
|
return target
|
|
|
|
def _try_repair_missing_slash(
|
|
member_name: str, link_name: str, archive_names: set[str]
|
|
) -> str | None:
|
|
"""Some upstream llama.cpp Mac releases (e.g. b9165, b9169) ship
|
|
symlinks whose linkname is missing the directory separator AND
|
|
the leading character of the file basename between the
|
|
top-level dir and the rest of the path:
|
|
|
|
llama-b9165/libggml-rpc.0.dylib -> llama-b9165ibggml-rpc.0.11.1.dylib
|
|
|
|
That cannot be resolved as written. Detect the pattern
|
|
(linkname starts with the top-level dir name but no following
|
|
slash) and search archive entries under that dir for a real
|
|
file whose basename ends with the mangled suffix. Only accept
|
|
when the suffix uniquely identifies a real archive entry.
|
|
Returns the corrected linkname expressed relative to the
|
|
member's parent directory -- callers join it with
|
|
`target.parent`, so a full `top/file` path would double the
|
|
prefix into `top/top/file`."""
|
|
if "/" not in member_name or "/" in link_name:
|
|
return None
|
|
top, _, _ = member_name.partition("/")
|
|
if not link_name.startswith(top) or len(link_name) <= len(top):
|
|
return None
|
|
bad_suffix = link_name[len(top) :]
|
|
if not bad_suffix or bad_suffix.startswith("/"):
|
|
return None
|
|
prefix = f"{top}/"
|
|
candidates = [
|
|
name
|
|
for name in archive_names
|
|
if name.startswith(prefix)
|
|
and "/" not in name[len(prefix) :]
|
|
and name[len(prefix) :].endswith(bad_suffix)
|
|
]
|
|
if len(candidates) != 1:
|
|
return None
|
|
# Strip the top-level dir so the caller's `target.parent / Path(...)`
|
|
# composition resolves inside the staging dir, not into a duplicate
|
|
# `top/top/...` path.
|
|
return candidates[0][len(prefix) :]
|
|
|
|
def safe_link_target(
|
|
base: Path,
|
|
member_name: str,
|
|
link_name: str,
|
|
target: Path,
|
|
archive_names: set[str],
|
|
) -> tuple[str, Path]:
|
|
normalized = link_name.replace("\\", "/")
|
|
repaired = _try_repair_missing_slash(member_name, normalized, archive_names)
|
|
if repaired is not None:
|
|
normalized = repaired
|
|
link_path = Path(normalized)
|
|
if link_path.is_absolute():
|
|
raise PrebuiltFallback(
|
|
f"archive link used an absolute target: {member_name} -> {link_name}"
|
|
)
|
|
if not normalized:
|
|
raise PrebuiltFallback(f"archive link used an empty target: {member_name}")
|
|
|
|
resolved = (target.parent / link_path).resolve()
|
|
base_resolved = base.resolve()
|
|
try:
|
|
resolved.relative_to(base_resolved)
|
|
except ValueError as exc:
|
|
raise PrebuiltFallback(
|
|
f"archive link escaped destination: {member_name} -> {link_name}"
|
|
) from exc
|
|
return normalized, resolved
|
|
|
|
def extract_zip_safely(source: Path, base: Path) -> None:
|
|
with zipfile.ZipFile(source) as archive:
|
|
for member in archive.infolist():
|
|
target = safe_extract_path(base, member.filename)
|
|
mode = (member.external_attr >> 16) & 0o170000
|
|
if mode == 0o120000:
|
|
raise PrebuiltFallback(
|
|
f"zip archive contained a symlink entry: {member.filename}"
|
|
)
|
|
if member.is_dir():
|
|
target.mkdir(parents = True, exist_ok = True)
|
|
continue
|
|
target.parent.mkdir(parents = True, exist_ok = True)
|
|
with archive.open(member, "r") as src, target.open("wb") as dst:
|
|
shutil.copyfileobj(src, dst)
|
|
|
|
def extract_tar_safely(source: Path, base: Path) -> None:
|
|
pending_links: list[tuple[tarfile.TarInfo, Path]] = []
|
|
archive_names: set[str] = set()
|
|
with tarfile.open(source, "r:gz") as archive:
|
|
for member in archive.getmembers():
|
|
archive_names.add(member.name)
|
|
target = safe_extract_path(base, member.name)
|
|
if member.isdir():
|
|
target.mkdir(parents = True, exist_ok = True)
|
|
continue
|
|
if member.islnk() or member.issym():
|
|
pending_links.append((member, target))
|
|
continue
|
|
if not member.isfile():
|
|
raise PrebuiltFallback(
|
|
f"tar archive contained an unsupported entry: {member.name}"
|
|
)
|
|
target.parent.mkdir(parents = True, exist_ok = True)
|
|
extracted = archive.extractfile(member)
|
|
if extracted is None:
|
|
raise PrebuiltFallback(
|
|
f"tar archive entry could not be read: {member.name}"
|
|
)
|
|
with extracted, target.open("wb") as dst:
|
|
shutil.copyfileobj(extracted, dst)
|
|
|
|
unresolved = list(pending_links)
|
|
while unresolved:
|
|
next_round: list[tuple[tarfile.TarInfo, Path]] = []
|
|
progressed = False
|
|
for member, target in unresolved:
|
|
normalized_link, resolved_target = safe_link_target(
|
|
base, member.name, member.linkname, target, archive_names
|
|
)
|
|
if not resolved_target.exists() and not resolved_target.is_symlink():
|
|
next_round.append((member, target))
|
|
continue
|
|
if resolved_target.is_dir():
|
|
raise PrebuiltFallback(
|
|
f"archive link targeted a directory: {member.name} -> {member.linkname}"
|
|
)
|
|
|
|
target.parent.mkdir(parents = True, exist_ok = True)
|
|
if target.exists() or target.is_symlink():
|
|
target.unlink()
|
|
|
|
if member.issym():
|
|
target.symlink_to(normalized_link)
|
|
else:
|
|
shutil.copy2(resolved_target, target)
|
|
progressed = True
|
|
|
|
if not progressed:
|
|
details = ", ".join(
|
|
f"{member.name} -> {member.linkname}" for member, _ in next_round
|
|
)
|
|
raise PrebuiltFallback(
|
|
f"tar archive contained unresolved link entries: {details}"
|
|
)
|
|
unresolved = next_round
|
|
|
|
destination.mkdir(parents = True, exist_ok = True)
|
|
if archive_path.name.endswith(".zip"):
|
|
extract_zip_safely(archive_path, destination)
|
|
return
|
|
if archive_path.name.endswith(".tar.gz"):
|
|
extract_tar_safely(archive_path, destination)
|
|
return
|
|
raise PrebuiltFallback(f"unsupported archive format: {archive_path.name}")
|
|
|
|
|
|
def copy_globs(
|
|
source_dir: Path, destination: Path, patterns: list[str], *, required: bool = True
|
|
) -> None:
|
|
destination.mkdir(parents = True, exist_ok = True)
|
|
matched_sources: dict[str, Path] = {}
|
|
for path in sorted(
|
|
(candidate for candidate in source_dir.rglob("*") if candidate.is_file()),
|
|
key = lambda candidate: (
|
|
len(candidate.relative_to(source_dir).parts),
|
|
str(candidate),
|
|
),
|
|
):
|
|
for pattern in patterns:
|
|
if fnmatch.fnmatch(path.name, pattern):
|
|
previous = matched_sources.get(path.name)
|
|
if previous is not None and previous != path:
|
|
raise PrebuiltFallback(
|
|
f"ambiguous archive layout for {path.name}: "
|
|
f"{previous.relative_to(source_dir)} and {path.relative_to(source_dir)}"
|
|
)
|
|
matched_sources[path.name] = path
|
|
break
|
|
|
|
if required and not matched_sources:
|
|
raise PrebuiltFallback(f"required files missing from {source_dir}: {patterns}")
|
|
|
|
for name, path in matched_sources.items():
|
|
shutil.copy2(path, destination / name)
|
|
|
|
|
|
def ensure_converter_scripts(install_dir: Path, llama_tag: str) -> None:
|
|
canonical = install_dir / "convert_hf_to_gguf.py"
|
|
if not canonical.exists():
|
|
# Hydrated source tree should have placed this file already.
|
|
# Fall back to a network fetch so the install is not blocked.
|
|
raw_base = f"https://raw.githubusercontent.com/ggml-org/llama.cpp/{llama_tag}"
|
|
source_url = f"{raw_base}/convert_hf_to_gguf.py"
|
|
data = download_bytes(
|
|
source_url,
|
|
progress_label = f"Downloading {download_label_from_url(source_url)}",
|
|
)
|
|
if not data:
|
|
raise RuntimeError(f"downloaded empty converter script from {source_url}")
|
|
if b"import " not in data and b"def " not in data and b"#!/" not in data:
|
|
raise RuntimeError(
|
|
f"downloaded converter script did not look like Python source: {source_url}"
|
|
)
|
|
atomic_write_bytes(canonical, data)
|
|
legacy = install_dir / "convert-hf-to-gguf.py"
|
|
if legacy.exists() or legacy.is_symlink():
|
|
legacy.unlink()
|
|
try:
|
|
legacy.symlink_to("convert_hf_to_gguf.py")
|
|
except OSError:
|
|
shutil.copy2(canonical, legacy)
|
|
|
|
|
|
def extracted_archive_root(extract_dir: Path) -> Path:
|
|
children = [path for path in extract_dir.iterdir()]
|
|
if len(children) == 1 and children[0].is_dir():
|
|
return children[0]
|
|
return extract_dir
|
|
|
|
|
|
def copy_directory_contents(source_dir: Path, destination: Path) -> None:
|
|
destination.mkdir(parents = True, exist_ok = True)
|
|
for item in source_dir.iterdir():
|
|
target = destination / item.name
|
|
if item.is_dir():
|
|
shutil.copytree(item, target, dirs_exist_ok = True)
|
|
else:
|
|
shutil.copy2(item, target)
|
|
|
|
|
|
def hydrate_source_tree(
|
|
source_ref: str,
|
|
install_dir: Path,
|
|
work_dir: Path,
|
|
*,
|
|
source_repo: str = UPSTREAM_REPO,
|
|
expected_sha256: str | None,
|
|
source_label: str | None = None,
|
|
exact_source: bool = False,
|
|
) -> None:
|
|
archive_path = work_dir / f"llama.cpp-source-{source_ref}.tar.gz"
|
|
source_urls = (
|
|
commit_source_archive_urls(source_repo, source_ref)
|
|
if exact_source
|
|
else upstream_source_archive_urls(source_ref)
|
|
)
|
|
label = source_label or f"llama.cpp source tree for {source_ref}"
|
|
extract_dir = Path(tempfile.mkdtemp(prefix = "source-extract-", dir = work_dir))
|
|
|
|
try:
|
|
log(f"downloading {label}")
|
|
last_exc: Exception | None = None
|
|
downloaded = False
|
|
for index, source_url in enumerate(source_urls):
|
|
try:
|
|
if index > 0:
|
|
log(
|
|
f"retrying source tree download from fallback URL: {source_url}"
|
|
)
|
|
download_file_verified(
|
|
source_url,
|
|
archive_path,
|
|
expected_sha256 = expected_sha256,
|
|
label = label,
|
|
)
|
|
downloaded = True
|
|
break
|
|
except Exception as exc:
|
|
last_exc = exc
|
|
if index == len(source_urls) - 1:
|
|
raise
|
|
log(f"source tree download failed from {source_url}: {exc}")
|
|
if not downloaded:
|
|
assert last_exc is not None
|
|
raise last_exc
|
|
extract_archive(archive_path, extract_dir)
|
|
source_root = extracted_archive_root(extract_dir)
|
|
required_paths = [
|
|
source_root / "CMakeLists.txt",
|
|
source_root / "convert_hf_to_gguf.py",
|
|
source_root / "gguf-py",
|
|
]
|
|
missing = [
|
|
str(path.relative_to(source_root))
|
|
for path in required_paths
|
|
if not path.exists()
|
|
]
|
|
if missing:
|
|
raise PrebuiltFallback(
|
|
"upstream source archive was missing required repo files: "
|
|
+ ", ".join(missing)
|
|
)
|
|
copy_directory_contents(source_root, install_dir)
|
|
except PrebuiltFallback:
|
|
raise
|
|
except Exception as exc:
|
|
raise PrebuiltFallback(f"failed to hydrate {label}: {exc}") from exc
|
|
finally:
|
|
remove_tree(extract_dir)
|
|
|
|
|
|
def normalize_install_layout(install_dir: Path, host: HostInfo) -> tuple[Path, Path]:
|
|
build_bin = install_dir / "build" / "bin"
|
|
if host.is_windows:
|
|
exec_dir = build_bin / "Release"
|
|
exec_dir.mkdir(parents = True, exist_ok = True)
|
|
return exec_dir / "llama-server.exe", exec_dir / "llama-quantize.exe"
|
|
|
|
install_dir.mkdir(parents = True, exist_ok = True)
|
|
build_bin.mkdir(parents = True, exist_ok = True)
|
|
return install_dir / "llama-server", install_dir / "llama-quantize"
|
|
|
|
|
|
def discover_installed_executable(install_dir: Path, executable_name: str) -> Path:
|
|
direct = install_dir / executable_name
|
|
if direct.exists() and direct.is_file():
|
|
return direct
|
|
candidate = next(
|
|
(path for path in install_dir.rglob(executable_name) if path.is_file()), None
|
|
)
|
|
if candidate is None:
|
|
raise PrebuiltFallback(f"{executable_name} was not installed")
|
|
return candidate
|
|
|
|
|
|
def write_exec_wrapper(entrypoint: Path, target: Path) -> None:
|
|
relative_target = os.path.relpath(target, entrypoint.parent)
|
|
script = "\n".join(
|
|
[
|
|
"#!/bin/sh",
|
|
f'exec "$(dirname "$0")/{relative_target}" "$@"',
|
|
"",
|
|
]
|
|
)
|
|
atomic_write_bytes(entrypoint, script.encode("utf-8"))
|
|
os.chmod(entrypoint, 0o755)
|
|
|
|
|
|
def create_exec_entrypoint(entrypoint: Path, target: Path) -> None:
|
|
if entrypoint == target:
|
|
return
|
|
if entrypoint.exists() or entrypoint.is_symlink():
|
|
entrypoint.unlink()
|
|
try:
|
|
entrypoint.symlink_to(os.path.relpath(target, entrypoint.parent))
|
|
except Exception:
|
|
write_exec_wrapper(entrypoint, target)
|
|
|
|
|
|
def overlay_directory_for_choice(
|
|
install_dir: Path, choice: AssetChoice, host: HostInfo
|
|
) -> Path:
|
|
if host.is_windows or choice.install_kind.startswith("windows"):
|
|
path = install_dir / "build" / "bin" / "Release"
|
|
else:
|
|
path = install_dir / "build" / "bin"
|
|
path.mkdir(parents = True, exist_ok = True)
|
|
return path
|
|
|
|
|
|
def paired_runtime_dll_patterns(choice: AssetChoice) -> list[str]:
|
|
"""Filename patterns the paired runtime archive is allowed to drop
|
|
into the install. Used for the second copy_globs pass in
|
|
install_from_archives, narrower than runtime_patterns_for_choice so
|
|
the runtime archive cannot overwrite main-archive payload like
|
|
llama-server.exe. Only Windows CUDA has paired runtimes today."""
|
|
if choice.install_kind == "windows-cuda":
|
|
return ["cudart64_*.dll", "cublas64_*.dll", "cublasLt64_*.dll"]
|
|
return []
|
|
|
|
|
|
def runtime_patterns_for_choice(choice: AssetChoice) -> list[str]:
|
|
# Broad shared-library glob + explicit binary names. Lets upstream
|
|
# repackage the SO/DLL set (e.g. ggml-org/llama.cpp#23462 split the
|
|
# per-binary entry code into paired ``lib<binary>-impl.so`` shared
|
|
# libraries between b9279 and b9283) without us re-enumerating
|
|
# every new file. Studio only invokes llama-server and llama-quantize;
|
|
# other CLIs upstream ships (llama-cli, llama-bench, ...) are skipped.
|
|
if choice.install_kind in {"linux-cpu", "linux-cuda", "linux-rocm", "linux-arm64"}:
|
|
return ["llama-server", "llama-quantize", "lib*.so*"]
|
|
if choice.install_kind in {"macos-arm64", "macos-x64"}:
|
|
return ["llama-server", "llama-quantize", "lib*.dylib"]
|
|
if choice.install_kind in {
|
|
"windows-cpu",
|
|
"windows-cuda",
|
|
"windows-hip",
|
|
"windows-arm64",
|
|
}:
|
|
return ["llama-server.exe", "llama-quantize.exe", "*.dll"]
|
|
raise PrebuiltFallback(
|
|
f"unsupported install kind for runtime overlay: {choice.install_kind}"
|
|
)
|
|
|
|
|
|
def metadata_patterns_for_choice(choice: AssetChoice) -> list[str]:
|
|
patterns = ["BUILD_INFO.txt", "THIRD_PARTY_LICENSES.txt"]
|
|
if choice.install_kind.startswith("windows"):
|
|
patterns.append("LICENSE.txt")
|
|
else:
|
|
patterns.append("LICENSE")
|
|
return patterns
|
|
|
|
|
|
@contextmanager
|
|
def install_lock(lock_path: Path) -> Iterator[None]:
|
|
lock_path.parent.mkdir(parents = True, exist_ok = True)
|
|
|
|
if FileLock is None:
|
|
# Fallback: exclusive file creation as a simple lock.
|
|
# Write our PID so stale locks from crashed processes can be detected.
|
|
fd: int | None = None
|
|
deadline = time.monotonic() + INSTALL_LOCK_TIMEOUT_SECONDS
|
|
while True:
|
|
try:
|
|
fd = os.open(str(lock_path), os.O_CREAT | os.O_EXCL | os.O_RDWR)
|
|
try:
|
|
os.write(fd, f"{os.getpid()}\n".encode())
|
|
os.fsync(fd)
|
|
except Exception:
|
|
os.close(fd)
|
|
fd = None
|
|
lock_path.unlink(missing_ok = True)
|
|
raise
|
|
break
|
|
except FileExistsError:
|
|
# Check if the holder process is still alive
|
|
stale = False
|
|
try:
|
|
raw = lock_path.read_text().strip()
|
|
except FileNotFoundError:
|
|
# Lock vanished between our open attempt and read -- retry
|
|
continue
|
|
if not raw:
|
|
# File exists but PID not yet written -- another process
|
|
# just created it. Wait briefly for the write to land.
|
|
if time.monotonic() >= deadline:
|
|
raise BusyInstallConflict(
|
|
f"timed out after {INSTALL_LOCK_TIMEOUT_SECONDS}s waiting for concurrent install lock: {lock_path}"
|
|
)
|
|
time.sleep(0.1)
|
|
continue
|
|
try:
|
|
holder_pid = int(raw)
|
|
os.kill(holder_pid, 0) # signal 0 = existence check
|
|
except ValueError:
|
|
# PID unreadable (corrupted file)
|
|
stale = True
|
|
except ProcessLookupError:
|
|
# Process is dead
|
|
stale = True
|
|
except PermissionError:
|
|
# Process is alive but owned by another user -- not stale
|
|
pass
|
|
if stale:
|
|
lock_path.unlink(missing_ok = True)
|
|
continue
|
|
if time.monotonic() >= deadline:
|
|
raise BusyInstallConflict(
|
|
f"timed out after {INSTALL_LOCK_TIMEOUT_SECONDS}s waiting for concurrent install lock: {lock_path}"
|
|
)
|
|
time.sleep(0.5)
|
|
try:
|
|
yield
|
|
finally:
|
|
if fd is not None:
|
|
os.close(fd)
|
|
lock_path.unlink(missing_ok = True)
|
|
return
|
|
|
|
try:
|
|
with FileLock(lock_path, timeout = INSTALL_LOCK_TIMEOUT_SECONDS):
|
|
yield
|
|
except FileLockTimeout as exc:
|
|
raise BusyInstallConflict(
|
|
f"timed out after {INSTALL_LOCK_TIMEOUT_SECONDS}s waiting for concurrent install lock: {lock_path}"
|
|
) from exc
|
|
|
|
|
|
def install_lock_path(install_dir: Path) -> Path:
|
|
return install_dir.parent / f".{install_dir.name}.install.lock"
|
|
|
|
|
|
def install_staging_root(install_dir: Path) -> Path:
|
|
root = install_dir.parent / INSTALL_STAGING_ROOT_NAME
|
|
root.mkdir(parents = True, exist_ok = True)
|
|
return root
|
|
|
|
|
|
def prune_install_staging_root(install_dir: Path) -> None:
|
|
root = install_dir.parent / INSTALL_STAGING_ROOT_NAME
|
|
try:
|
|
root.rmdir()
|
|
except OSError:
|
|
pass
|
|
|
|
|
|
def create_install_staging_dir(install_dir: Path) -> Path:
|
|
staging_dir = Path(
|
|
tempfile.mkdtemp(
|
|
prefix = f"{install_dir.name}.staging-", dir = install_staging_root(install_dir)
|
|
)
|
|
)
|
|
log(f"created install staging dir {staging_dir}")
|
|
return staging_dir
|
|
|
|
|
|
def unique_install_side_path(install_dir: Path, label: str) -> Path:
|
|
root = install_staging_root(install_dir)
|
|
timestamp = time.strftime("%Y%m%d%H%M%S", time.gmtime())
|
|
prefix = f"{install_dir.name}.{label}-{timestamp}-{os.getpid()}"
|
|
candidate = root / prefix
|
|
counter = 0
|
|
while candidate.exists():
|
|
counter += 1
|
|
candidate = root / f"{prefix}-{counter}"
|
|
return candidate
|
|
|
|
|
|
def remove_tree(path: Path | None) -> None:
|
|
if path and path.exists():
|
|
shutil.rmtree(path, ignore_errors = True)
|
|
|
|
|
|
def remove_tree_logged(path: Path | None, label: str) -> None:
|
|
if not path:
|
|
return
|
|
if not path.exists():
|
|
log(f"{label} already absent at {path}")
|
|
return
|
|
log(f"removing {label} at {path}")
|
|
try:
|
|
shutil.rmtree(path)
|
|
except Exception as exc:
|
|
log(f"failed to remove {label} at {path}: {exc}")
|
|
raise
|
|
|
|
|
|
def cleanup_install_side_paths(
|
|
install_dir: Path,
|
|
*,
|
|
staging_dir: Path | None = None,
|
|
rollback_dir: Path | None = None,
|
|
failed_dir: Path | None = None,
|
|
active_dir: Path | None = None,
|
|
) -> None:
|
|
cleanup_failures: list[str] = []
|
|
for label, path in (
|
|
("failed install path", failed_dir),
|
|
("rollback path", rollback_dir),
|
|
("active install path", active_dir),
|
|
("staging dir", staging_dir),
|
|
):
|
|
if not path:
|
|
continue
|
|
try:
|
|
remove_tree_logged(path, label)
|
|
except Exception as exc:
|
|
cleanup_failures.append(f"{label} ({path}): {exc}")
|
|
prune_install_staging_root(install_dir)
|
|
if cleanup_failures:
|
|
raise RuntimeError("cleanup failed for " + "; ".join(cleanup_failures))
|
|
|
|
|
|
def confirm_install_tree(install_dir: Path, host: HostInfo) -> None:
|
|
if host.is_windows:
|
|
expected = [
|
|
install_dir / "build" / "bin" / "Release" / "llama-server.exe",
|
|
install_dir / "build" / "bin" / "Release" / "llama-quantize.exe",
|
|
install_dir / "convert_hf_to_gguf.py",
|
|
install_dir / "gguf-py",
|
|
]
|
|
else:
|
|
expected = [
|
|
install_dir / "llama-server",
|
|
install_dir / "llama-quantize",
|
|
install_dir / "build" / "bin" / "llama-server",
|
|
install_dir / "build" / "bin" / "llama-quantize",
|
|
install_dir / "convert_hf_to_gguf.py",
|
|
install_dir / "gguf-py",
|
|
]
|
|
|
|
expected.append(install_dir / "UNSLOTH_PREBUILT_INFO.json")
|
|
missing = [str(path) for path in expected if not path.exists()]
|
|
if missing:
|
|
raise RuntimeError(
|
|
"activated install was missing expected files: " + ", ".join(missing)
|
|
)
|
|
|
|
|
|
def activate_install_tree(staging_dir: Path, install_dir: Path, host: HostInfo) -> None:
|
|
rollback_dir: Path | None = None
|
|
failed_dir: Path | None = None
|
|
try:
|
|
if install_dir.exists():
|
|
rollback_dir = unique_install_side_path(install_dir, "rollback")
|
|
log(f"moving existing install to rollback path {rollback_dir}")
|
|
os.replace(install_dir, rollback_dir)
|
|
log(f"moved existing install to rollback path {rollback_dir.name}")
|
|
|
|
log(f"activating staged install {staging_dir} -> {install_dir}")
|
|
os.replace(staging_dir, install_dir)
|
|
log(f"activated staged install at {install_dir}")
|
|
log(f"confirming activated install tree at {install_dir}")
|
|
confirm_install_tree(install_dir, host)
|
|
log(f"activated install tree confirmed at {install_dir}")
|
|
except Exception as exc:
|
|
log(f"activation failed for staged install: {exc}")
|
|
try:
|
|
if install_dir.exists():
|
|
failed_dir = unique_install_side_path(install_dir, "failed")
|
|
log(f"moving failed active install to {failed_dir}")
|
|
os.replace(install_dir, failed_dir)
|
|
elif staging_dir.exists():
|
|
failed_dir = staging_dir
|
|
staging_dir = None
|
|
log(f"retaining failed staging tree at {failed_dir}")
|
|
|
|
if rollback_dir and rollback_dir.exists():
|
|
log(f"restoring rollback path {rollback_dir} -> {install_dir}")
|
|
os.replace(rollback_dir, install_dir)
|
|
log(f"restored previous install from rollback path {rollback_dir.name}")
|
|
if is_busy_lock_error(exc):
|
|
raise BusyInstallConflict(
|
|
"staged prebuilt validation passed but the existing install could not be replaced "
|
|
"because llama.cpp appears to still be in use; restored previous install "
|
|
f"({textwrap.shorten(str(exc), width = 200, placeholder = '...')})"
|
|
) from exc
|
|
raise PrebuiltFallback(
|
|
"staged prebuilt validation passed but activation failed; restored previous install "
|
|
f"({textwrap.shorten(str(exc), width = 200, placeholder = '...')})"
|
|
) from exc
|
|
except (BusyInstallConflict, PrebuiltFallback):
|
|
raise
|
|
except Exception as rollback_exc:
|
|
log(f"rollback after failed activation also failed: {rollback_exc}")
|
|
|
|
log(
|
|
"rollback restoration failed; cleaning staging, install, and rollback paths before source build fallback"
|
|
)
|
|
cleanup_error: Exception | None = None
|
|
try:
|
|
cleanup_install_side_paths(
|
|
install_dir,
|
|
staging_dir = staging_dir,
|
|
rollback_dir = rollback_dir,
|
|
failed_dir = failed_dir,
|
|
active_dir = install_dir,
|
|
)
|
|
except Exception as cleanup_exc:
|
|
cleanup_error = cleanup_exc
|
|
log(f"cleanup after rollback failure also failed: {cleanup_exc}")
|
|
details = textwrap.shorten(str(exc), width = 200, placeholder = "...")
|
|
if cleanup_error is not None:
|
|
raise PrebuiltFallback(
|
|
"staged prebuilt validation passed but activation and rollback failed; "
|
|
f"cleanup also reported errors ({details}; cleanup={cleanup_error})"
|
|
) from exc
|
|
raise PrebuiltFallback(
|
|
"staged prebuilt validation passed but activation and rollback failed; "
|
|
f"cleaned install state for fresh source build ({details})"
|
|
) from exc
|
|
else:
|
|
if rollback_dir:
|
|
try:
|
|
remove_tree_logged(rollback_dir, "rollback path")
|
|
except Exception as cleanup_exc:
|
|
log(
|
|
f"non-fatal: rollback cleanup failed after successful activation: {cleanup_exc}"
|
|
)
|
|
finally:
|
|
remove_tree(failed_dir)
|
|
remove_tree(staging_dir)
|
|
prune_install_staging_root(install_dir)
|
|
|
|
|
|
def install_from_archives(
|
|
choice: AssetChoice, host: HostInfo, install_dir: Path, work_dir: Path
|
|
) -> tuple[Path, Path]:
|
|
main_archive = work_dir / choice.name
|
|
log(f"downloading {choice.name} from {choice.source_label} release")
|
|
download_file_verified(
|
|
choice.url,
|
|
main_archive,
|
|
expected_sha256 = choice.expected_sha256,
|
|
label = f"prebuilt archive {choice.name}",
|
|
)
|
|
|
|
install_dir.mkdir(parents = True, exist_ok = True)
|
|
extract_dir = Path(tempfile.mkdtemp(prefix = "extract-", dir = work_dir))
|
|
runtime_extract_dir: Path | None = None
|
|
|
|
try:
|
|
extract_archive(main_archive, extract_dir)
|
|
# Download the paired runtime archive into its own temp dir to
|
|
# avoid copy_globs's ambiguous-layout guard on shared names
|
|
# like LICENSE.txt. Two passes of copy_globs land both archives
|
|
# in the same overlay dir. Fixes #5106.
|
|
if choice.runtime_url and choice.runtime_name:
|
|
runtime_archive = work_dir / choice.runtime_name
|
|
log(
|
|
f"downloading paired runtime archive {choice.runtime_name} "
|
|
f"from {choice.source_label} release"
|
|
)
|
|
download_file_verified(
|
|
choice.runtime_url,
|
|
runtime_archive,
|
|
expected_sha256 = choice.runtime_sha256,
|
|
label = f"prebuilt runtime archive {choice.runtime_name}",
|
|
)
|
|
runtime_extract_dir = Path(
|
|
tempfile.mkdtemp(prefix = "extract-runtime-", dir = work_dir)
|
|
)
|
|
extract_archive(runtime_archive, runtime_extract_dir)
|
|
source_dir = extract_dir
|
|
overlay_dir = overlay_directory_for_choice(install_dir, choice, host)
|
|
copy_globs(
|
|
source_dir, overlay_dir, runtime_patterns_for_choice(choice), required = True
|
|
)
|
|
if runtime_extract_dir is not None:
|
|
# The runtime archive only contributes the CUDA DLLs.
|
|
# Restrict the overlay to the cudart bundle's known
|
|
# filenames (cudart64_X.dll / cublas64_X.dll /
|
|
# cublasLt64_X.dll) rather than the broad ``*.exe`` /
|
|
# ``*.dll`` set from runtime_patterns_for_choice, so a
|
|
# malformed runtime archive can never overwrite
|
|
# llama-server.exe or other main-archive payload. The
|
|
# upstream cudart-llama-bin-win-cuda-X.Y-x64.zip currently
|
|
# ships exactly these three DLLs (verified against b9103
|
|
# cuda-12.4 and cuda-13.1 bundles).
|
|
copy_globs(
|
|
runtime_extract_dir,
|
|
overlay_dir,
|
|
paired_runtime_dll_patterns(choice),
|
|
required = False,
|
|
)
|
|
copy_globs(
|
|
source_dir,
|
|
install_dir,
|
|
metadata_patterns_for_choice(choice),
|
|
required = False,
|
|
)
|
|
finally:
|
|
remove_tree(extract_dir)
|
|
if runtime_extract_dir is not None:
|
|
remove_tree(runtime_extract_dir)
|
|
|
|
if host.is_windows:
|
|
exec_dir = install_dir / "build" / "bin" / "Release"
|
|
server_src = next(exec_dir.glob("llama-server.exe"), None)
|
|
quantize_src = next(exec_dir.glob("llama-quantize.exe"), None)
|
|
if server_src is None or quantize_src is None:
|
|
raise PrebuiltFallback("windows executables were not installed correctly")
|
|
return server_src, quantize_src
|
|
|
|
build_bin = install_dir / "build" / "bin"
|
|
source_server = build_bin / "llama-server"
|
|
source_quantize = build_bin / "llama-quantize"
|
|
if not source_server.exists() or not source_quantize.exists():
|
|
raise PrebuiltFallback(
|
|
"unix executables were not installed correctly into build/bin"
|
|
)
|
|
os.chmod(source_server, 0o755)
|
|
os.chmod(source_quantize, 0o755)
|
|
|
|
root_server = install_dir / "llama-server"
|
|
root_quantize = install_dir / "llama-quantize"
|
|
if source_server != root_server:
|
|
create_exec_entrypoint(root_server, source_server)
|
|
if source_quantize != root_quantize:
|
|
create_exec_entrypoint(root_quantize, source_quantize)
|
|
build_server = build_bin / "llama-server"
|
|
build_quantize = build_bin / "llama-quantize"
|
|
if source_server != build_server:
|
|
create_exec_entrypoint(build_server, source_server)
|
|
if source_quantize != build_quantize:
|
|
create_exec_entrypoint(build_quantize, source_quantize)
|
|
|
|
return source_server, source_quantize
|
|
|
|
|
|
def ensure_repo_shape(install_dir: Path) -> None:
|
|
required = [
|
|
install_dir / "CMakeLists.txt",
|
|
install_dir / "convert_hf_to_gguf.py",
|
|
install_dir / "gguf-py",
|
|
]
|
|
missing = [
|
|
str(path.relative_to(install_dir)) for path in required if not path.exists()
|
|
]
|
|
if missing:
|
|
raise PrebuiltFallback(
|
|
"hydrated llama.cpp source tree was missing: " + ", ".join(missing)
|
|
)
|
|
|
|
|
|
def validation_model_cache_path(install_dir: Path) -> Path:
|
|
cache_dir = install_dir.parent / VALIDATION_MODEL_CACHE_DIRNAME
|
|
cache_dir.mkdir(parents = True, exist_ok = True)
|
|
return cache_dir / VALIDATION_MODEL_CACHE_FILENAME
|
|
|
|
|
|
def validated_validation_model_bytes(data: bytes) -> bytes:
|
|
if not data:
|
|
raise RuntimeError(f"downloaded empty validation model from {TEST_MODEL_URL}")
|
|
digest = hashlib.sha256(data).hexdigest()
|
|
if digest != TEST_MODEL_SHA256:
|
|
raise RuntimeError(
|
|
"validation model checksum mismatch: "
|
|
f"expected={TEST_MODEL_SHA256} actual={digest}"
|
|
)
|
|
return data
|
|
|
|
|
|
def download_validation_model(path: Path, cache_path: Path | None = None) -> None:
|
|
try:
|
|
data: bytes | None = None
|
|
if cache_path and cache_path.exists():
|
|
try:
|
|
data = validated_validation_model_bytes(cache_path.read_bytes())
|
|
log(f"using cached tiny GGUF validation model from {cache_path}")
|
|
except Exception as exc:
|
|
log(
|
|
f"cached tiny GGUF validation model was invalid; refreshing cache ({exc})"
|
|
)
|
|
data = None
|
|
if data is None:
|
|
log("downloading tiny GGUF validation model")
|
|
data = validated_validation_model_bytes(
|
|
download_bytes(
|
|
TEST_MODEL_URL,
|
|
progress_label = f"Downloading {download_label_from_url(TEST_MODEL_URL)}",
|
|
)
|
|
)
|
|
if cache_path is not None:
|
|
atomic_write_bytes(cache_path, data)
|
|
atomic_write_bytes(path, data)
|
|
except Exception as exc:
|
|
raise PrebuiltFallback(f"validation model unavailable: {exc}") from exc
|
|
|
|
|
|
def free_local_port() -> int:
|
|
sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
|
|
sock.bind(("127.0.0.1", 0))
|
|
_, port = sock.getsockname()
|
|
sock.close()
|
|
return int(port)
|
|
|
|
|
|
def read_log_excerpt(log_path: Path, *, max_lines: int = 60) -> str:
|
|
try:
|
|
content = log_path.read_text(encoding = "utf-8", errors = "replace")
|
|
except FileNotFoundError:
|
|
return ""
|
|
return "\n".join(content.splitlines()[-max_lines:])
|
|
|
|
|
|
def is_retryable_server_bind_error(
|
|
exc: Exception | None,
|
|
output: str = "",
|
|
*,
|
|
exited_quickly: bool = False,
|
|
) -> bool:
|
|
haystack = output.lower()
|
|
bind_markers = (
|
|
"address already in use",
|
|
"only one usage of each socket address",
|
|
"failed to bind",
|
|
"bind failed",
|
|
"failed to listen",
|
|
"errno 98",
|
|
"errno 10048",
|
|
)
|
|
if any(marker in haystack for marker in bind_markers):
|
|
return True
|
|
|
|
if isinstance(exc, urllib.error.URLError):
|
|
reason = exc.reason
|
|
if exited_quickly and isinstance(reason, ConnectionRefusedError):
|
|
return True
|
|
if isinstance(reason, OSError) and reason.errno in {
|
|
98,
|
|
99,
|
|
111,
|
|
10048,
|
|
10049,
|
|
10061,
|
|
}:
|
|
return exited_quickly
|
|
if exited_quickly and isinstance(exc, ConnectionRefusedError):
|
|
return True
|
|
if isinstance(exc, OSError) and exc.errno in {98, 99, 111, 10048, 10049, 10061}:
|
|
return exited_quickly
|
|
return False
|
|
|
|
|
|
def dedupe_existing_dirs(paths: Iterable[str | Path]) -> list[str]:
|
|
unique: list[str] = []
|
|
seen: set[str] = set()
|
|
for raw in paths:
|
|
if not raw:
|
|
continue
|
|
path = Path(raw).expanduser()
|
|
if not path.is_dir():
|
|
continue
|
|
resolved = str(path.resolve())
|
|
if resolved in seen:
|
|
continue
|
|
seen.add(resolved)
|
|
unique.append(resolved)
|
|
return unique
|
|
|
|
|
|
def linux_missing_libraries(
|
|
binary_path: Path, *, env: dict[str, str] | None = None
|
|
) -> list[str]:
|
|
try:
|
|
result = run_capture(["ldd", str(binary_path)], timeout = 20, env = env)
|
|
except Exception:
|
|
return []
|
|
|
|
missing: list[str] = []
|
|
for line in (result.stdout + result.stderr).splitlines():
|
|
line = line.strip()
|
|
if "=> not found" not in line:
|
|
continue
|
|
library = line.split("=>", 1)[0].strip()
|
|
if library and library not in missing:
|
|
missing.append(library)
|
|
return missing
|
|
|
|
|
|
def python_runtime_dirs() -> list[str]:
|
|
candidates: list[Path] = []
|
|
search_roots = [Path(entry) for entry in sys.path if entry]
|
|
try:
|
|
search_roots.extend(Path(path) for path in site.getsitepackages())
|
|
except Exception:
|
|
pass
|
|
try:
|
|
user_site = site.getusersitepackages()
|
|
if user_site:
|
|
search_roots.append(Path(user_site))
|
|
except Exception:
|
|
pass
|
|
|
|
for root in search_roots:
|
|
if not root.is_dir():
|
|
continue
|
|
# ``nvidia/<pkg>/lib`` -- Linux convention; harmless on Windows
|
|
# where the directory simply does not exist on real wheels.
|
|
candidates.extend(root.glob("nvidia/*/lib"))
|
|
# ``nvidia/<pkg>/bin`` -- legacy modular Windows wheels
|
|
# (``nvidia-cuda-runtime-cu12``, ``nvidia-cublas-cu12``).
|
|
candidates.extend(root.glob("nvidia/*/bin"))
|
|
# ``nvidia/<pkg>/bin/x86_64`` and ``.../bin/x64`` -- current
|
|
# CUDA 13 Windows wheel layout (the unsuffixed
|
|
# ``nvidia-cuda-runtime`` 13.x and ``nvidia-cublas`` 13.x
|
|
# packages ship under ``nvidia/cu13/bin/x86_64/cudart64_13.dll``).
|
|
# Without these, Windows preflight CUDA detection misses cu13
|
|
# installs and falls back to the upstream cudart bundle path
|
|
# even when usable DLLs are already on disk (#5106). Kept in
|
|
# sync with the backend resolver
|
|
# ``llama_cpp.LlamaCppBackend._windows_pip_nvidia_dll_dirs``.
|
|
candidates.extend(root.glob("nvidia/*/bin/x86_64"))
|
|
candidates.extend(root.glob("nvidia/*/bin/x64"))
|
|
# ``nvidia/<pkg>/Library/bin`` -- conda-style wheel repacks.
|
|
candidates.extend(root.glob("nvidia/*/Library/bin"))
|
|
candidates.extend(root.glob("nvidia/*/Library/bin/x86_64"))
|
|
candidates.extend(root.glob("nvidia/*/Library/bin/x64"))
|
|
candidates.extend(root.glob("torch/lib"))
|
|
return dedupe_existing_dirs(candidates)
|
|
|
|
|
|
def ldconfig_runtime_dirs(required_libraries: Iterable[str]) -> list[str]:
|
|
try:
|
|
result = run_capture(["ldconfig", "-p"], timeout = 20)
|
|
except Exception:
|
|
return []
|
|
|
|
required = set(required_libraries)
|
|
candidates: list[str] = []
|
|
for line in result.stdout.splitlines():
|
|
if "=>" not in line:
|
|
continue
|
|
library, _, location = line.partition("=>")
|
|
library = library.strip().split()[0]
|
|
if required and library not in required:
|
|
continue
|
|
path = Path(location.strip()).parent
|
|
candidates.append(str(path))
|
|
return dedupe_existing_dirs(candidates)
|
|
|
|
|
|
def linux_runtime_dirs(binary_path: Path) -> list[str]:
|
|
missing = linux_missing_libraries(binary_path)
|
|
if not missing:
|
|
return []
|
|
return linux_runtime_dirs_for_required_libraries(missing)
|
|
|
|
|
|
def preflight_linux_installed_binaries(
|
|
binaries: Iterable[Path],
|
|
install_dir: Path,
|
|
host: HostInfo,
|
|
) -> None:
|
|
if not host.is_linux:
|
|
return
|
|
|
|
issues: list[str] = []
|
|
for binary_path in binaries:
|
|
env = binary_env(binary_path, install_dir, host)
|
|
missing = linux_missing_libraries(binary_path, env = env)
|
|
if not missing:
|
|
continue
|
|
runtime_dirs = [
|
|
part for part in env.get("LD_LIBRARY_PATH", "").split(os.pathsep) if part
|
|
]
|
|
issues.append(
|
|
f"{binary_path.name}: missing={','.join(missing)} "
|
|
f"ld_library_path={','.join(runtime_dirs) if runtime_dirs else 'none'}"
|
|
)
|
|
|
|
if issues:
|
|
raise PrebuiltFallback(
|
|
"linux extracted binary preflight failed:\n" + "\n".join(issues)
|
|
)
|
|
|
|
|
|
def glob_paths(*patterns: str) -> list[str]:
|
|
matches: list[str] = []
|
|
for pattern in patterns:
|
|
if any(char in pattern for char in "*?[]"):
|
|
matches.extend(str(path) for path in Path("/").glob(pattern.lstrip("/")))
|
|
else:
|
|
matches.append(pattern)
|
|
return matches
|
|
|
|
|
|
def windows_runtime_dirs() -> list[str]:
|
|
candidates: list[str | Path] = []
|
|
|
|
env_dirs = os.environ.get("CUDA_RUNTIME_DLL_DIR", "")
|
|
if env_dirs:
|
|
candidates.extend(part for part in env_dirs.split(os.pathsep) if part)
|
|
|
|
path_dirs = os.environ.get("PATH", "")
|
|
if path_dirs:
|
|
candidates.extend(part for part in path_dirs.split(os.pathsep) if part)
|
|
|
|
cuda_roots: list[Path] = []
|
|
for name in ("CUDA_PATH", "CUDA_HOME", "CUDA_ROOT"):
|
|
value = os.environ.get(name)
|
|
if value:
|
|
cuda_roots.append(Path(value))
|
|
|
|
for root in cuda_roots:
|
|
candidates.extend([root / "bin", root / "lib" / "x64"])
|
|
|
|
program_files = os.environ.get("ProgramFiles", r"C:\Program Files")
|
|
toolkit_base = Path(program_files) / "NVIDIA GPU Computing Toolkit" / "CUDA"
|
|
if toolkit_base.is_dir():
|
|
candidates.extend(toolkit_base.glob("v*/bin"))
|
|
candidates.extend(toolkit_base.glob("v*/lib/x64"))
|
|
|
|
candidates.extend(Path(path) for path in python_runtime_dirs())
|
|
return dedupe_existing_dirs(candidates)
|
|
|
|
|
|
def windows_runtime_dirs_for_patterns(
|
|
required_patterns: Iterable[str],
|
|
candidate_dirs: Iterable[str] | None = None,
|
|
) -> list[str]:
|
|
directories = (
|
|
list(candidate_dirs) if candidate_dirs is not None else windows_runtime_dirs()
|
|
)
|
|
matching_dirs: list[str] = []
|
|
for pattern in required_patterns:
|
|
matched_dirs = [
|
|
directory for directory in directories if any(Path(directory).glob(pattern))
|
|
]
|
|
if not matched_dirs:
|
|
return []
|
|
for directory in matched_dirs:
|
|
if directory not in matching_dirs:
|
|
matching_dirs.append(directory)
|
|
return matching_dirs
|
|
|
|
|
|
def windows_runtime_dirs_for_runtime_line(runtime_line: str | None) -> list[str]:
|
|
if not runtime_line:
|
|
return []
|
|
patterns = windows_runtime_line_info().get(runtime_line)
|
|
if not patterns:
|
|
return []
|
|
return windows_runtime_dirs_for_patterns(patterns)
|
|
|
|
|
|
def binary_env(
|
|
binary_path: Path,
|
|
install_dir: Path,
|
|
host: HostInfo,
|
|
*,
|
|
runtime_line: str | None = None,
|
|
) -> dict[str, str]:
|
|
env = os.environ.copy()
|
|
if host.is_windows:
|
|
path_dirs = [
|
|
str(binary_path.parent),
|
|
*windows_runtime_dirs_for_runtime_line(runtime_line),
|
|
]
|
|
existing = [part for part in env.get("PATH", "").split(os.pathsep) if part]
|
|
env["PATH"] = os.pathsep.join(dedupe_existing_dirs([*path_dirs, *existing]))
|
|
elif host.is_linux:
|
|
ld_dirs = [
|
|
str(binary_path.parent),
|
|
str(install_dir),
|
|
*linux_runtime_dirs(binary_path),
|
|
]
|
|
existing = [
|
|
part for part in env.get("LD_LIBRARY_PATH", "").split(os.pathsep) if part
|
|
]
|
|
env["LD_LIBRARY_PATH"] = os.pathsep.join(
|
|
dedupe_existing_dirs([*ld_dirs, *existing])
|
|
)
|
|
elif host.is_macos:
|
|
dyld_dirs = [str(binary_path.parent), str(install_dir)]
|
|
existing = [
|
|
part for part in env.get("DYLD_LIBRARY_PATH", "").split(os.pathsep) if part
|
|
]
|
|
env["DYLD_LIBRARY_PATH"] = os.pathsep.join(
|
|
dedupe_existing_dirs([*dyld_dirs, *existing])
|
|
)
|
|
return env
|
|
|
|
|
|
def validate_quantize(
|
|
quantize_path: Path,
|
|
probe_path: Path,
|
|
quantized_path: Path,
|
|
install_dir: Path,
|
|
host: HostInfo,
|
|
*,
|
|
runtime_line: str | None = None,
|
|
) -> None:
|
|
command = [str(quantize_path), str(probe_path), str(quantized_path), "Q6_K", "2"]
|
|
result = subprocess.run(
|
|
command,
|
|
capture_output = True,
|
|
text = True,
|
|
timeout = 120,
|
|
env = binary_env(quantize_path, install_dir, host, runtime_line = runtime_line),
|
|
**windows_hidden_subprocess_kwargs(),
|
|
)
|
|
if (
|
|
result.returncode != 0
|
|
or not quantized_path.exists()
|
|
or quantized_path.stat().st_size == 0
|
|
):
|
|
raise PrebuiltFallback(
|
|
"llama-quantize validation failed:\n"
|
|
+ result.stdout
|
|
+ ("\n" + result.stderr if result.stderr else "")
|
|
)
|
|
|
|
|
|
def validate_server(
|
|
server_path: Path,
|
|
probe_path: Path,
|
|
host: HostInfo,
|
|
install_dir: Path,
|
|
*,
|
|
runtime_line: str | None = None,
|
|
install_kind: str | None = None,
|
|
) -> None:
|
|
last_failure: PrebuiltFallback | None = None
|
|
for port_attempt in range(1, SERVER_PORT_BIND_ATTEMPTS + 1):
|
|
port = free_local_port()
|
|
command = [
|
|
str(server_path),
|
|
"-m",
|
|
str(probe_path),
|
|
"--host",
|
|
"127.0.0.1",
|
|
"--port",
|
|
str(port),
|
|
"-c",
|
|
"32",
|
|
"--parallel",
|
|
"1",
|
|
"--threads",
|
|
"1",
|
|
"--ubatch-size",
|
|
"32",
|
|
"--batch-size",
|
|
"32",
|
|
]
|
|
# Only enable GPU offload for assets that actually ship GPU code.
|
|
# Gating on `host.has_rocm` alone breaks the intentional CPU
|
|
# fallback on AMD Windows hosts without a HIP prebuilt: the CPU
|
|
# binary would be launched with `--n-gpu-layers 1` and fail
|
|
# validation. Use the resolved install_kind as the source of
|
|
# truth and fall back to host detection when the caller did not
|
|
# pass one (keeps backwards compatibility with older call sites).
|
|
_gpu_kinds = {
|
|
"linux-cuda",
|
|
"linux-rocm",
|
|
"windows-cuda",
|
|
"windows-hip",
|
|
"macos-arm64",
|
|
}
|
|
if install_kind is not None:
|
|
_enable_gpu_layers = install_kind in _gpu_kinds
|
|
else:
|
|
# Older call sites that don't pass install_kind: keep ROCm
|
|
# hosts in the GPU-validation path so an AMD-only Linux host
|
|
# is exercised against the actual hardware rather than the
|
|
# CPU fallback. NVIDIA and macOS-arm64 are already covered.
|
|
_enable_gpu_layers = (
|
|
host.has_usable_nvidia
|
|
or host.has_rocm
|
|
or (host.is_macos and host.is_arm64)
|
|
)
|
|
if _enable_gpu_layers:
|
|
command.extend(["--n-gpu-layers", "1"])
|
|
|
|
log_fd, log_name = tempfile.mkstemp(prefix = "llama-server-", suffix = ".log")
|
|
os.close(log_fd)
|
|
log_path = Path(log_name)
|
|
process: subprocess.Popen[str] | None = None
|
|
try:
|
|
with log_path.open("w", encoding = "utf-8", errors = "replace") as log_handle:
|
|
process = subprocess.Popen(
|
|
command,
|
|
stdout = log_handle,
|
|
stderr = subprocess.STDOUT,
|
|
text = True,
|
|
env = binary_env(
|
|
server_path, install_dir, host, runtime_line = runtime_line
|
|
),
|
|
**windows_hidden_subprocess_kwargs(),
|
|
)
|
|
deadline = time.time() + 60
|
|
startup_started = time.time()
|
|
response_body = ""
|
|
last_error: Exception | None = None
|
|
while time.time() < deadline:
|
|
if process.poll() is not None:
|
|
process.wait(timeout = 5)
|
|
log_handle.flush()
|
|
output = read_log_excerpt(log_path)
|
|
exited_quickly = (
|
|
time.time() - startup_started
|
|
) <= SERVER_BIND_RETRY_WINDOW_SECONDS
|
|
failure = PrebuiltFallback(
|
|
"llama-server exited during startup:\n" + output
|
|
)
|
|
if (
|
|
port_attempt < SERVER_PORT_BIND_ATTEMPTS
|
|
and is_retryable_server_bind_error(
|
|
last_error,
|
|
output,
|
|
exited_quickly = exited_quickly,
|
|
)
|
|
):
|
|
log(
|
|
f"llama-server startup hit a port race on {port}; retrying with a fresh port "
|
|
f"({port_attempt}/{SERVER_PORT_BIND_ATTEMPTS})"
|
|
)
|
|
last_failure = failure
|
|
break
|
|
raise failure
|
|
|
|
payload = json.dumps({"prompt": "a", "n_predict": 1}).encode(
|
|
"utf-8"
|
|
)
|
|
request = urllib.request.Request(
|
|
f"http://127.0.0.1:{port}/completion",
|
|
data = payload,
|
|
headers = {"Content-Type": "application/json"},
|
|
)
|
|
try:
|
|
with urllib.request.urlopen(request, timeout = 5) as response:
|
|
status_code = response.status
|
|
response_body = response.read().decode("utf-8", "replace")
|
|
if status_code == 200:
|
|
return
|
|
last_error = RuntimeError(
|
|
f"unexpected HTTP status {status_code}"
|
|
)
|
|
except urllib.error.HTTPError as exc:
|
|
response_body = exc.read().decode("utf-8", "replace")
|
|
last_error = exc
|
|
except Exception as exc:
|
|
last_error = exc
|
|
time.sleep(0.5)
|
|
else:
|
|
log_handle.flush()
|
|
output = read_log_excerpt(log_path)
|
|
raise PrebuiltFallback(
|
|
"llama-server completion validation timed out"
|
|
+ (f" ({last_error})" if last_error else "")
|
|
+ ":\n"
|
|
+ output
|
|
+ ("\n" + response_body if response_body else "")
|
|
)
|
|
finally:
|
|
if process is not None and process.poll() is None:
|
|
process.terminate()
|
|
try:
|
|
process.wait(timeout = 5)
|
|
except subprocess.TimeoutExpired:
|
|
process.kill()
|
|
process.wait(timeout = 5)
|
|
try:
|
|
log_path.unlink(missing_ok = True)
|
|
except Exception:
|
|
pass
|
|
if last_failure is not None:
|
|
raise last_failure
|
|
raise PrebuiltFallback("llama-server validation failed unexpectedly")
|
|
|
|
|
|
def collect_system_report(
|
|
host: HostInfo, choice: AssetChoice | None, install_dir: Path
|
|
) -> str:
|
|
lines = [
|
|
f"platform={host.system} machine={host.machine}",
|
|
f"driver_cuda_version={host.driver_cuda_version}",
|
|
f"compute_caps={','.join(host.compute_caps) if host.compute_caps else 'unknown'}",
|
|
f"cuda_visible_devices={host.visible_cuda_devices if host.visible_cuda_devices is not None else 'unset'}",
|
|
f"has_physical_nvidia={host.has_physical_nvidia}",
|
|
f"has_usable_nvidia={host.has_usable_nvidia}",
|
|
f"chosen_asset={(choice.name if choice else 'none')}",
|
|
f"asset_source={(choice.source_label if choice else 'none')}",
|
|
]
|
|
if host.is_linux and host.has_physical_nvidia:
|
|
runtime_lines, runtime_dirs = detected_linux_runtime_lines()
|
|
lines.append(
|
|
"linux_runtime_lines="
|
|
+ (",".join(runtime_lines) if runtime_lines else "none")
|
|
)
|
|
for runtime_line in ("cuda13", "cuda12"):
|
|
lines.append(
|
|
f"linux_runtime_dirs_{runtime_line}="
|
|
+ (
|
|
",".join(runtime_dirs.get(runtime_line, []))
|
|
if runtime_dirs.get(runtime_line)
|
|
else "none"
|
|
)
|
|
)
|
|
if choice and choice.selection_log:
|
|
lines.append("selection_log:")
|
|
lines.extend(choice.selection_log)
|
|
if host.nvidia_smi:
|
|
try:
|
|
smi = run_capture([host.nvidia_smi], timeout = 20)
|
|
excerpt = "\n".join((smi.stdout + smi.stderr).splitlines()[:20])
|
|
lines.append("nvidia-smi:")
|
|
lines.append(excerpt)
|
|
except Exception as exc:
|
|
lines.append(f"nvidia-smi error: {exc}")
|
|
|
|
if host.is_linux:
|
|
server_binary = install_dir / "llama-server"
|
|
if server_binary.exists():
|
|
server_env = binary_env(server_binary, install_dir, host)
|
|
lines.append(
|
|
"linux_missing_libs="
|
|
+ (
|
|
",".join(linux_missing_libraries(server_binary, env = server_env))
|
|
or "none"
|
|
)
|
|
)
|
|
lines.append(
|
|
"linux_runtime_dirs="
|
|
+ (
|
|
",".join(
|
|
[
|
|
part
|
|
for part in server_env.get("LD_LIBRARY_PATH", "").split(
|
|
os.pathsep
|
|
)
|
|
if part
|
|
]
|
|
)
|
|
or "none"
|
|
)
|
|
)
|
|
try:
|
|
ldd = run_capture(
|
|
["ldd", str(server_binary)], timeout = 20, env = server_env
|
|
)
|
|
lines.append("ldd llama-server:")
|
|
lines.append((ldd.stdout + ldd.stderr).strip())
|
|
except Exception as exc:
|
|
lines.append(f"ldd error: {exc}")
|
|
elif host.is_windows:
|
|
lines.append(
|
|
"windows_runtime_dirs=" + (",".join(windows_runtime_dirs()) or "none")
|
|
)
|
|
runtime_lines, runtime_dirs = detected_windows_runtime_lines()
|
|
lines.append(
|
|
"windows_runtime_lines="
|
|
+ (",".join(runtime_lines) if runtime_lines else "none")
|
|
)
|
|
for runtime_line in ("cuda13", "cuda12"):
|
|
lines.append(
|
|
f"windows_runtime_dirs_{runtime_line}="
|
|
+ (
|
|
",".join(runtime_dirs.get(runtime_line, []))
|
|
if runtime_dirs.get(runtime_line)
|
|
else "none"
|
|
)
|
|
)
|
|
elif host.is_macos:
|
|
server_binary = install_dir / "llama-server"
|
|
if server_binary.exists():
|
|
try:
|
|
otool = run_capture(["otool", "-L", str(server_binary)], timeout = 20)
|
|
lines.append("otool -L llama-server:")
|
|
lines.append((otool.stdout + otool.stderr).strip())
|
|
except Exception as exc:
|
|
lines.append(f"otool error: {exc}")
|
|
|
|
return "\n".join(lines)
|
|
|
|
|
|
def apply_approved_hashes(
|
|
attempts: Iterable[AssetChoice],
|
|
checksums: ApprovedReleaseChecksums,
|
|
) -> list[AssetChoice]:
|
|
def approved_hash_for_attempt(attempt: AssetChoice) -> ApprovedArtifactHash | None:
|
|
candidate_names = [attempt.name]
|
|
if (
|
|
isinstance(attempt.tag, str)
|
|
and attempt.tag
|
|
and attempt.tag != checksums.upstream_tag
|
|
and attempt.name.startswith("llama-")
|
|
):
|
|
legacy_prefix = f"llama-{attempt.tag}-"
|
|
compatibility_prefix = f"llama-{checksums.upstream_tag}-"
|
|
compatibility_name = (
|
|
attempt.name.replace(legacy_prefix, compatibility_prefix, 1)
|
|
if attempt.name.startswith(legacy_prefix)
|
|
else attempt.name
|
|
)
|
|
candidate_names.append(compatibility_name)
|
|
candidate_names.extend(
|
|
windows_cuda_asset_aliases(
|
|
attempt.name,
|
|
compatibility_tag = checksums.upstream_tag,
|
|
)
|
|
)
|
|
seen_names: set[str] = set()
|
|
for candidate_name in candidate_names:
|
|
if candidate_name in seen_names:
|
|
continue
|
|
seen_names.add(candidate_name)
|
|
approved = checksums.artifacts.get(candidate_name)
|
|
if approved is not None:
|
|
return approved
|
|
return None
|
|
|
|
approved_attempts: list[AssetChoice] = []
|
|
missing_assets: list[str] = []
|
|
for attempt in attempts:
|
|
approved = approved_hash_for_attempt(attempt)
|
|
if approved is None:
|
|
missing_assets.append(attempt.name)
|
|
continue
|
|
attempt.expected_sha256 = approved.sha256
|
|
# Resolve the paired runtime archive's hash too. Drop the pair
|
|
# if the manifest does not list it -- never install an
|
|
# unverified archive.
|
|
if attempt.runtime_name and attempt.runtime_url:
|
|
runtime_approved = checksums.artifacts.get(attempt.runtime_name)
|
|
if runtime_approved is None:
|
|
attempt.runtime_name = None
|
|
attempt.runtime_url = None
|
|
attempt.runtime_sha256 = None
|
|
else:
|
|
attempt.runtime_sha256 = runtime_approved.sha256
|
|
approved_attempts.append(attempt)
|
|
if not approved_attempts:
|
|
missing_text = ", ".join(missing_assets) if missing_assets else "none"
|
|
raise PrebuiltFallback(
|
|
"approved checksum asset did not contain the selected prebuilt archive(s): "
|
|
f"{missing_text}"
|
|
)
|
|
return approved_attempts
|
|
|
|
|
|
def require_approved_source_hash(
|
|
checksums: ApprovedReleaseChecksums, llama_tag: str
|
|
) -> ApprovedArtifactHash:
|
|
source_asset_name = source_archive_logical_name(llama_tag)
|
|
approved_source = checksums.artifacts.get(source_asset_name)
|
|
if approved_source is None:
|
|
raise PrebuiltFallback(
|
|
f"approved checksum asset did not contain source archive {source_asset_name}"
|
|
)
|
|
return approved_source
|
|
|
|
|
|
def preferred_source_archive(
|
|
checksums: ApprovedReleaseChecksums, llama_tag: str
|
|
) -> tuple[str, str, ApprovedArtifactHash | None, bool]:
|
|
exact_source = exact_source_archive_hash(checksums)
|
|
exact_repo = repo_slug_from_source(checksums.source_repo) or repo_slug_from_source(
|
|
checksums.source_repo_url
|
|
)
|
|
if exact_source is not None and exact_repo and checksums.source_commit:
|
|
return (
|
|
exact_repo,
|
|
checksums.source_commit,
|
|
exact_source,
|
|
True,
|
|
)
|
|
legacy = checksums.artifacts.get(source_archive_logical_name(llama_tag))
|
|
return (
|
|
UPSTREAM_REPO,
|
|
llama_tag,
|
|
legacy,
|
|
False,
|
|
)
|
|
|
|
|
|
def selected_source_archive_metadata(
|
|
checksums: ApprovedReleaseChecksums,
|
|
llama_tag: str,
|
|
) -> tuple[str, str | None]:
|
|
_source_repo, _source_ref, source_archive, _exact_source = preferred_source_archive(
|
|
checksums, llama_tag
|
|
)
|
|
if source_archive is None:
|
|
return source_archive_logical_name(llama_tag), None
|
|
return source_archive.asset_name, source_archive.sha256
|
|
|
|
|
|
def resolve_install_attempts(
|
|
llama_tag: str,
|
|
host: HostInfo,
|
|
published_repo: str,
|
|
published_release_tag: str,
|
|
) -> tuple[str, str, list[AssetChoice], ApprovedReleaseChecksums]:
|
|
requested_tag, plans = resolve_install_release_plans(
|
|
llama_tag,
|
|
host,
|
|
published_repo,
|
|
published_release_tag,
|
|
)
|
|
if not plans:
|
|
raise PrebuiltFallback("no prebuilt release plans were available")
|
|
plan = plans[0]
|
|
return requested_tag, plan.llama_tag, plan.attempts, plan.approved_checksums
|
|
|
|
|
|
def resolve_install_release_plans(
|
|
llama_tag: str,
|
|
host: HostInfo,
|
|
published_repo: str,
|
|
published_release_tag: str,
|
|
*,
|
|
max_release_fallbacks: int = DEFAULT_MAX_PREBUILT_RELEASE_FALLBACKS,
|
|
) -> tuple[str, list[InstallReleasePlan]]:
|
|
requested_tag = normalized_requested_llama_tag(llama_tag)
|
|
allow_older_release_fallback = (
|
|
requested_tag == "latest" and not published_release_tag
|
|
)
|
|
release_limit = max(1, max_release_fallbacks)
|
|
plans: list[InstallReleasePlan] = []
|
|
last_error: PrebuiltFallback | None = None
|
|
|
|
for resolved_release in iter_resolved_published_releases(
|
|
llama_tag,
|
|
published_repo,
|
|
published_release_tag,
|
|
):
|
|
bundle = resolved_release.bundle
|
|
checksums = resolved_release.checksums
|
|
resolved_tag = bundle.upstream_tag
|
|
try:
|
|
if host.is_linux and host.is_x86_64 and host.has_usable_nvidia:
|
|
linux_cuda_selection = resolve_linux_cuda_choice(host, bundle)
|
|
attempts = apply_approved_hashes(
|
|
linux_cuda_selection.attempts, checksums
|
|
)
|
|
if not attempts:
|
|
raise PrebuiltFallback("no compatible Linux CUDA asset was found")
|
|
log_lines(linux_cuda_selection.selection_log)
|
|
else:
|
|
attempts = resolve_release_asset_choice(
|
|
host,
|
|
resolved_tag,
|
|
bundle,
|
|
checksums,
|
|
)
|
|
if not attempts:
|
|
raise PrebuiltFallback("no compatible prebuilt asset was found")
|
|
if attempts[0].selection_log:
|
|
log_lines(attempts[0].selection_log)
|
|
except PrebuiltFallback as exc:
|
|
last_error = exc
|
|
if not allow_older_release_fallback:
|
|
raise
|
|
log(
|
|
"published release skipped for install planning: "
|
|
f"{bundle.repo}@{bundle.release_tag} upstream_tag={resolved_tag} ({exc})"
|
|
)
|
|
continue
|
|
|
|
plans.append(
|
|
InstallReleasePlan(
|
|
requested_tag = requested_tag,
|
|
llama_tag = resolved_tag,
|
|
release_tag = bundle.release_tag,
|
|
attempts = attempts,
|
|
approved_checksums = checksums,
|
|
)
|
|
)
|
|
|
|
if not allow_older_release_fallback or len(plans) >= release_limit:
|
|
break
|
|
|
|
if plans:
|
|
return requested_tag, plans
|
|
if last_error is not None:
|
|
raise last_error
|
|
raise PrebuiltFallback("no installable published llama.cpp releases were found")
|
|
|
|
|
|
def write_prebuilt_metadata(
|
|
install_dir: Path,
|
|
*,
|
|
requested_tag: str,
|
|
llama_tag: str,
|
|
release_tag: str,
|
|
choice: AssetChoice,
|
|
approved_checksums: ApprovedReleaseChecksums,
|
|
prebuilt_fallback_used: bool,
|
|
) -> None:
|
|
source_asset_name, source_sha256 = selected_source_archive_metadata(
|
|
approved_checksums,
|
|
llama_tag,
|
|
)
|
|
# expected_install_fingerprint is the source of truth for what the
|
|
# fingerprint must contain. Calling it here -- instead of inlining a
|
|
# parallel payload -- prevents drift where new keys (e.g. the cudart
|
|
# pair fields added for #5106) are added to one side but not the
|
|
# other, which would cause every install to look stale.
|
|
fingerprint = expected_install_fingerprint(
|
|
llama_tag = llama_tag,
|
|
release_tag = release_tag,
|
|
choice = choice,
|
|
approved_checksums = approved_checksums,
|
|
)
|
|
if fingerprint is None:
|
|
raise PrebuiltFallback(f"cannot compute install fingerprint for {choice.name}")
|
|
metadata = {
|
|
"requested_tag": requested_tag,
|
|
"tag": llama_tag,
|
|
"release_tag": release_tag,
|
|
"published_repo": approved_checksums.repo,
|
|
"asset": choice.name,
|
|
"asset_sha256": choice.expected_sha256,
|
|
"source": choice.source_label,
|
|
"source_asset": source_asset_name,
|
|
"source_sha256": source_sha256,
|
|
"source_commit": approved_checksums.source_commit,
|
|
"source_commit_short": approved_checksums.source_commit_short,
|
|
"source_repo": approved_checksums.source_repo,
|
|
"source_repo_url": approved_checksums.source_repo_url,
|
|
"source_ref_kind": approved_checksums.source_ref_kind,
|
|
"requested_source_ref": approved_checksums.requested_source_ref,
|
|
"resolved_source_ref": approved_checksums.resolved_source_ref,
|
|
"bundle_profile": choice.bundle_profile,
|
|
"runtime_line": choice.runtime_line,
|
|
"coverage_class": choice.coverage_class,
|
|
"install_fingerprint": fingerprint,
|
|
"prebuilt_fallback_used": prebuilt_fallback_used,
|
|
"installed_at_utc": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
|
|
}
|
|
(install_dir / "UNSLOTH_PREBUILT_INFO.json").write_text(
|
|
json.dumps(metadata, indent = 2) + "\n"
|
|
)
|
|
|
|
|
|
def expected_install_fingerprint(
|
|
*,
|
|
llama_tag: str,
|
|
release_tag: str,
|
|
choice: AssetChoice,
|
|
approved_checksums: ApprovedReleaseChecksums,
|
|
) -> str | None:
|
|
source_asset_name, source_sha256 = selected_source_archive_metadata(
|
|
approved_checksums,
|
|
llama_tag,
|
|
)
|
|
payload = {
|
|
"published_repo": approved_checksums.repo,
|
|
"release_tag": release_tag,
|
|
"upstream_tag": llama_tag,
|
|
"asset": choice.name,
|
|
"asset_sha256": choice.expected_sha256,
|
|
"source": choice.source_label,
|
|
"source_asset": source_asset_name,
|
|
"source_sha256": source_sha256,
|
|
"runtime_line": choice.runtime_line,
|
|
# Including the paired runtime archive (Windows cudart bundle)
|
|
# in the fingerprint is what forces existing #5106 installs to
|
|
# refresh: pre-PR installs hashed nothing in this slot, post-PR
|
|
# paired installs hash the cudart sha. Without these two keys
|
|
# an existing cudart-less install would keep matching the new
|
|
# choice and never re-overlay the cudart DLLs.
|
|
"runtime_asset": choice.runtime_name,
|
|
"runtime_sha256": choice.runtime_sha256,
|
|
"bundle_profile": choice.bundle_profile,
|
|
"coverage_class": choice.coverage_class,
|
|
}
|
|
return hashlib.sha256(
|
|
json.dumps(payload, sort_keys = True, separators = (",", ":")).encode("utf-8")
|
|
).hexdigest()
|
|
|
|
|
|
def load_prebuilt_metadata(install_dir: Path) -> dict[str, Any] | None:
|
|
metadata_path = install_dir / "UNSLOTH_PREBUILT_INFO.json"
|
|
if not metadata_path.is_file():
|
|
return None
|
|
try:
|
|
payload = json.loads(metadata_path.read_text(encoding = "utf-8"))
|
|
except Exception:
|
|
return None
|
|
if not isinstance(payload, dict):
|
|
return None
|
|
return payload
|
|
|
|
|
|
def runtime_payload_health_groups(choice: AssetChoice) -> list[list[str]]:
|
|
if choice.install_kind in {"linux-cpu", "linux-arm64"}:
|
|
return [
|
|
["libllama-common.so*"],
|
|
["libllama.so*"],
|
|
["libggml.so*"],
|
|
["libggml-base.so*"],
|
|
["libggml-cpu-*.so*"],
|
|
["libmtmd.so*"],
|
|
]
|
|
if choice.install_kind == "linux-cuda":
|
|
return [
|
|
["libllama-common.so*"],
|
|
["libllama.so*"],
|
|
["libggml.so*"],
|
|
["libggml-base.so*"],
|
|
["libggml-cpu-*.so*"],
|
|
["libmtmd.so*"],
|
|
["libggml-cuda.so*"],
|
|
]
|
|
if choice.install_kind in {"macos-arm64", "macos-x64"}:
|
|
return [
|
|
["libllama*.dylib"],
|
|
["libggml*.dylib"],
|
|
["libmtmd*.dylib"],
|
|
]
|
|
if choice.install_kind == "linux-rocm":
|
|
return [
|
|
["libllama-common.so*"],
|
|
["libllama.so*"],
|
|
["libggml.so*"],
|
|
["libggml-base.so*"],
|
|
["libggml-cpu-*.so*"],
|
|
["libmtmd.so*"],
|
|
["libggml-hip.so*"],
|
|
]
|
|
if choice.install_kind in {"windows-cpu", "windows-arm64"}:
|
|
return [["llama.dll"]]
|
|
if choice.install_kind == "windows-cuda":
|
|
groups = [["llama.dll"], ["ggml-cuda.dll"]]
|
|
# When the cudart bundle was paired in (#5106) require all
|
|
# three of its DLLs alongside the main archive's payload.
|
|
# install_kind alone is not enough -- legacy installs without
|
|
# the cudart pair must still pass the health check on the
|
|
# no-pair fallback path, otherwise pair-less builds would loop
|
|
# on reinstall forever. The upstream cudart bundle ships
|
|
# cudart64_X.dll + cublas64_X.dll + cublasLt64_X.dll; missing
|
|
# any one of them still breaks GPU initialisation.
|
|
if choice.runtime_name:
|
|
groups.append(["cudart64_*.dll"])
|
|
groups.append(["cublas64_*.dll"])
|
|
groups.append(["cublasLt64_*.dll"])
|
|
return groups
|
|
if choice.install_kind == "windows-hip":
|
|
return [["llama.dll"], ["*hip*.dll"]]
|
|
return []
|
|
|
|
|
|
def install_runtime_dir(install_dir: Path, host: HostInfo) -> Path:
|
|
if host.is_windows:
|
|
return install_dir / "build" / "bin" / "Release"
|
|
return install_dir / "build" / "bin"
|
|
|
|
|
|
def runtime_payload_is_healthy(
|
|
install_dir: Path, host: HostInfo, choice: AssetChoice
|
|
) -> bool:
|
|
runtime_dir = install_runtime_dir(install_dir, host)
|
|
if not runtime_dir.exists():
|
|
return False
|
|
for pattern_group in runtime_payload_health_groups(choice):
|
|
matched = False
|
|
for pattern in pattern_group:
|
|
if any(runtime_dir.glob(pattern)):
|
|
matched = True
|
|
break
|
|
if not matched:
|
|
return False
|
|
return True
|
|
|
|
|
|
def existing_install_matches_choice(
|
|
install_dir: Path,
|
|
host: HostInfo,
|
|
*,
|
|
llama_tag: str,
|
|
release_tag: str,
|
|
choice: AssetChoice,
|
|
approved_checksums: ApprovedReleaseChecksums,
|
|
) -> bool:
|
|
if not install_dir.exists():
|
|
return False
|
|
|
|
metadata = load_prebuilt_metadata(install_dir)
|
|
if metadata is None:
|
|
return False
|
|
|
|
try:
|
|
confirm_install_tree(install_dir, host)
|
|
except Exception:
|
|
return False
|
|
|
|
if not runtime_payload_is_healthy(install_dir, host, choice):
|
|
return False
|
|
|
|
# Verify primary executables still exist (catches partial deletion)
|
|
runtime_dir = install_runtime_dir(install_dir, host)
|
|
ext = ".exe" if host.is_windows else ""
|
|
for binary in ("llama-server", "llama-quantize"):
|
|
if not (runtime_dir / f"{binary}{ext}").exists():
|
|
return False
|
|
if host.is_linux:
|
|
try:
|
|
preflight_linux_installed_binaries(
|
|
[runtime_dir / "llama-server", runtime_dir / "llama-quantize"],
|
|
install_dir,
|
|
host,
|
|
)
|
|
except Exception:
|
|
return False
|
|
expected_fingerprint = expected_install_fingerprint(
|
|
llama_tag = llama_tag,
|
|
release_tag = release_tag,
|
|
choice = choice,
|
|
approved_checksums = approved_checksums,
|
|
)
|
|
if not expected_fingerprint:
|
|
return False
|
|
|
|
recorded_fingerprint = metadata.get("install_fingerprint")
|
|
if not isinstance(recorded_fingerprint, str) or not recorded_fingerprint:
|
|
return False
|
|
|
|
if recorded_fingerprint != expected_fingerprint:
|
|
return False
|
|
|
|
expected_pairs = {
|
|
"release_tag": release_tag,
|
|
"published_repo": approved_checksums.repo,
|
|
"tag": llama_tag,
|
|
"asset": choice.name,
|
|
"asset_sha256": choice.expected_sha256,
|
|
"source": choice.source_label,
|
|
"runtime_line": choice.runtime_line,
|
|
"bundle_profile": choice.bundle_profile,
|
|
"coverage_class": choice.coverage_class,
|
|
}
|
|
for key, expected in expected_pairs.items():
|
|
if metadata.get(key) != expected:
|
|
return False
|
|
return True
|
|
|
|
|
|
def existing_install_matches_plan(
|
|
install_dir: Path,
|
|
host: HostInfo,
|
|
plan: InstallReleasePlan,
|
|
) -> bool:
|
|
if not plan.attempts:
|
|
return False
|
|
return existing_install_matches_choice(
|
|
install_dir,
|
|
host,
|
|
llama_tag = plan.llama_tag,
|
|
release_tag = plan.release_tag,
|
|
choice = plan.attempts[0],
|
|
approved_checksums = plan.approved_checksums,
|
|
)
|
|
|
|
|
|
def validate_prebuilt_choice(
|
|
choice: AssetChoice,
|
|
host: HostInfo,
|
|
install_dir: Path,
|
|
work_dir: Path,
|
|
probe_path: Path,
|
|
*,
|
|
requested_tag: str,
|
|
llama_tag: str,
|
|
release_tag: str,
|
|
approved_checksums: ApprovedReleaseChecksums,
|
|
prebuilt_fallback_used: bool,
|
|
quantized_path: Path,
|
|
) -> tuple[Path, Path]:
|
|
source_repo, source_ref, source_archive, exact_source = preferred_source_archive(
|
|
approved_checksums, llama_tag
|
|
)
|
|
if exact_source:
|
|
log(
|
|
f"hydrating exact llama.cpp source for {source_repo}@{source_ref} into {install_dir}"
|
|
)
|
|
else:
|
|
log(f"hydrating upstream llama.cpp source for {llama_tag} into {install_dir}")
|
|
hydrate_source_tree(
|
|
source_ref,
|
|
install_dir,
|
|
work_dir,
|
|
source_repo = source_repo,
|
|
expected_sha256 = source_archive.sha256 if source_archive is not None else None,
|
|
source_label = (
|
|
f"llama.cpp source tree for {source_repo}@{source_ref}"
|
|
if exact_source
|
|
else f"llama.cpp source tree for {llama_tag}"
|
|
),
|
|
exact_source = exact_source,
|
|
)
|
|
log(f"overlaying prebuilt bundle {choice.name} into {install_dir}")
|
|
server_path, quantize_path = install_from_archives(
|
|
choice, host, install_dir, work_dir
|
|
)
|
|
preflight_linux_installed_binaries((server_path, quantize_path), install_dir, host)
|
|
ensure_repo_shape(install_dir)
|
|
write_prebuilt_metadata(
|
|
install_dir,
|
|
requested_tag = requested_tag,
|
|
llama_tag = llama_tag,
|
|
release_tag = release_tag,
|
|
choice = choice,
|
|
approved_checksums = approved_checksums,
|
|
prebuilt_fallback_used = prebuilt_fallback_used,
|
|
)
|
|
validate_quantize(
|
|
quantize_path,
|
|
probe_path,
|
|
quantized_path,
|
|
install_dir,
|
|
host,
|
|
runtime_line = choice.runtime_line,
|
|
)
|
|
validate_server(
|
|
server_path,
|
|
probe_path,
|
|
host,
|
|
install_dir,
|
|
runtime_line = choice.runtime_line,
|
|
install_kind = choice.install_kind,
|
|
)
|
|
log(f"staged prebuilt validation succeeded for {choice.name}")
|
|
return server_path, quantize_path
|
|
|
|
|
|
def validate_prebuilt_attempts(
|
|
attempts: Iterable[AssetChoice],
|
|
host: HostInfo,
|
|
install_dir: Path,
|
|
work_dir: Path,
|
|
probe_path: Path,
|
|
*,
|
|
requested_tag: str,
|
|
llama_tag: str,
|
|
release_tag: str,
|
|
approved_checksums: ApprovedReleaseChecksums,
|
|
initial_fallback_used: bool = False,
|
|
existing_install_dir: Path | None = None,
|
|
) -> tuple[AssetChoice, Path, bool]:
|
|
attempt_list = list(attempts)
|
|
if not attempt_list:
|
|
raise PrebuiltFallback("no prebuilt bundle attempts were available")
|
|
|
|
tried_fallback = initial_fallback_used
|
|
for index, attempt in enumerate(attempt_list):
|
|
if index > 0:
|
|
tried_fallback = True
|
|
log(
|
|
"retrying CUDA prebuilt "
|
|
f"{attempt.name} install_kind={attempt.install_kind} "
|
|
f"runtime_line={attempt.runtime_line} coverage_class={attempt.coverage_class}"
|
|
)
|
|
|
|
if existing_install_dir is not None and existing_install_matches_choice(
|
|
existing_install_dir,
|
|
host,
|
|
llama_tag = llama_tag,
|
|
release_tag = release_tag,
|
|
choice = attempt,
|
|
approved_checksums = approved_checksums,
|
|
):
|
|
log(
|
|
"existing llama.cpp install already matches fallback candidate "
|
|
f"{attempt.name}; skipping reinstall"
|
|
)
|
|
raise ExistingInstallSatisfied(attempt, tried_fallback)
|
|
|
|
staging_dir = create_install_staging_dir(install_dir)
|
|
quantized_path = work_dir / f"stories260K-q4-{index}.gguf"
|
|
if quantized_path.exists():
|
|
quantized_path.unlink()
|
|
try:
|
|
validate_prebuilt_choice(
|
|
attempt,
|
|
host,
|
|
staging_dir,
|
|
work_dir,
|
|
probe_path,
|
|
requested_tag = requested_tag,
|
|
llama_tag = llama_tag,
|
|
release_tag = release_tag,
|
|
approved_checksums = approved_checksums,
|
|
prebuilt_fallback_used = tried_fallback,
|
|
quantized_path = quantized_path,
|
|
)
|
|
except Exception as exc:
|
|
remove_tree(staging_dir)
|
|
prune_install_staging_root(install_dir)
|
|
if isinstance(exc, PrebuiltFallback):
|
|
attempt_error = exc
|
|
else:
|
|
attempt_error = PrebuiltFallback(
|
|
f"candidate attempt failed before activation for {attempt.name}: {exc}"
|
|
)
|
|
if index == len(attempt_list) - 1:
|
|
raise attempt_error from exc
|
|
log(
|
|
"selected CUDA bundle failed before activation; trying next prebuilt fallback "
|
|
f"({textwrap.shorten(str(attempt_error), width = 200, placeholder = '...')})"
|
|
)
|
|
continue
|
|
|
|
return attempt, staging_dir, tried_fallback
|
|
|
|
raise PrebuiltFallback("no prebuilt bundle passed validation")
|
|
|
|
|
|
def install_prebuilt(
|
|
install_dir: Path,
|
|
llama_tag: str,
|
|
published_repo: str,
|
|
published_release_tag: str,
|
|
*,
|
|
simple_policy: bool = False,
|
|
override_has_rocm: bool = False,
|
|
) -> None:
|
|
host = detect_host()
|
|
if override_has_rocm and not host.has_rocm:
|
|
host = dataclasses_replace(host, has_rocm = True)
|
|
choice: AssetChoice | None = None
|
|
try:
|
|
with install_lock(install_lock_path(install_dir)):
|
|
if install_dir.exists():
|
|
log(
|
|
f"existing llama.cpp install detected at {install_dir}; validating staged prebuilt update before replacement"
|
|
)
|
|
else:
|
|
log(
|
|
f"no existing llama.cpp install detected at {install_dir}; performing fresh prebuilt install"
|
|
)
|
|
if simple_policy:
|
|
requested_tag, release_plans = resolve_simple_install_release_plans(
|
|
llama_tag,
|
|
host,
|
|
published_repo,
|
|
published_release_tag,
|
|
)
|
|
else:
|
|
requested_tag, release_plans = resolve_install_release_plans(
|
|
llama_tag,
|
|
host,
|
|
published_repo,
|
|
published_release_tag,
|
|
)
|
|
if release_plans and existing_install_matches_plan(
|
|
install_dir, host, release_plans[0]
|
|
):
|
|
current = release_plans[0]
|
|
log(
|
|
"existing llama.cpp install already matches selected release "
|
|
f"{current.release_tag} upstream_tag={current.llama_tag}; skipping download and install"
|
|
)
|
|
return
|
|
with tempfile.TemporaryDirectory(prefix = "unsloth-llama-prebuilt-") as tmp:
|
|
work_dir = Path(tmp)
|
|
probe_path = work_dir / "stories260K.gguf"
|
|
download_validation_model(
|
|
probe_path, validation_model_cache_path(install_dir)
|
|
)
|
|
release_count = len(release_plans)
|
|
for release_index, plan in enumerate(release_plans):
|
|
choice = plan.attempts[0]
|
|
if existing_install_matches_plan(install_dir, host, plan):
|
|
log(
|
|
"existing llama.cpp install already matches fallback release "
|
|
f"{plan.release_tag} upstream_tag={plan.llama_tag}; skipping reinstall"
|
|
)
|
|
return
|
|
log(
|
|
"selected "
|
|
f"{choice.name} ({choice.source_label}) from published release "
|
|
f"{plan.release_tag} for {host.system} {host.machine}"
|
|
)
|
|
try:
|
|
choice, selected_staging_dir, _ = validate_prebuilt_attempts(
|
|
plan.attempts,
|
|
host,
|
|
install_dir,
|
|
work_dir,
|
|
probe_path,
|
|
requested_tag = requested_tag,
|
|
llama_tag = plan.llama_tag,
|
|
release_tag = plan.release_tag,
|
|
approved_checksums = plan.approved_checksums,
|
|
initial_fallback_used = release_index > 0,
|
|
existing_install_dir = install_dir,
|
|
)
|
|
except ExistingInstallSatisfied:
|
|
return
|
|
except PrebuiltFallback as exc:
|
|
if release_index == release_count - 1:
|
|
raise
|
|
log(
|
|
"published release "
|
|
f"{plan.release_tag} upstream_tag={plan.llama_tag} failed; "
|
|
"trying the next older published prebuilt "
|
|
f"({textwrap.shorten(str(exc), width = 200, placeholder = '...')})"
|
|
)
|
|
continue
|
|
|
|
activate_install_tree(selected_staging_dir, install_dir, host)
|
|
try:
|
|
ensure_converter_scripts(install_dir, plan.llama_tag)
|
|
except Exception as exc:
|
|
log(
|
|
"converter script fetch failed after activation; install remains valid "
|
|
f"({textwrap.shorten(str(exc), width = 200, placeholder = '...')})"
|
|
)
|
|
return
|
|
except BusyInstallConflict as exc:
|
|
log("prebuilt install path is blocked by an in-use llama.cpp install")
|
|
log(f"prebuilt busy reason: {exc}")
|
|
raise SystemExit(EXIT_BUSY) from exc
|
|
except PrebuiltFallback as exc:
|
|
log("prebuilt install path failed; falling back to source build")
|
|
log(f"prebuilt fallback reason: {exc}")
|
|
report = collect_system_report(host, choice, install_dir)
|
|
print(report)
|
|
raise SystemExit(EXIT_FALLBACK) from exc
|
|
|
|
|
|
def parse_args() -> argparse.Namespace:
|
|
parser = argparse.ArgumentParser(
|
|
description = "Install and validate a prebuilt llama.cpp bundle for Unsloth Studio."
|
|
)
|
|
parser.add_argument("--install-dir", help = "Target ~/.unsloth/llama.cpp directory")
|
|
parser.add_argument(
|
|
"--llama-tag",
|
|
default = DEFAULT_LLAMA_TAG,
|
|
help = (
|
|
"llama.cpp release tag. Defaults to the latest usable published Unsloth "
|
|
"release unless UNSLOTH_LLAMA_TAG overrides it."
|
|
),
|
|
)
|
|
parser.add_argument(
|
|
"--published-repo",
|
|
default = DEFAULT_PUBLISHED_REPO,
|
|
help = "Published bundle repository",
|
|
)
|
|
parser.add_argument(
|
|
"--published-release-tag",
|
|
default = DEFAULT_PUBLISHED_TAG,
|
|
help = (
|
|
"Published GitHub release tag to pin. By default, scan releases "
|
|
"until a usable published llama.cpp release bundle is found."
|
|
),
|
|
)
|
|
parser.add_argument(
|
|
"--simple-policy",
|
|
action = "store_true",
|
|
help = "Use the simplified platform-specific prebuilt selection policy.",
|
|
)
|
|
parser.add_argument(
|
|
"--has-rocm",
|
|
action = "store_true",
|
|
default = False,
|
|
help = (
|
|
"Assert that an AMD ROCm GPU is present. When set, skips the internal "
|
|
"hipinfo/amd-smi probe and forces has_rocm=True in the host profile. "
|
|
"Used by setup.ps1/setup.sh to forward their own ROCm detection result "
|
|
"so the HIP llama.cpp prebuilt is selected even when hipinfo is not on PATH."
|
|
),
|
|
)
|
|
resolve_group = parser.add_mutually_exclusive_group()
|
|
resolve_group.add_argument(
|
|
"--resolve-llama-tag",
|
|
nargs = "?",
|
|
const = "latest",
|
|
help = "Resolve a llama.cpp tag such as 'latest' to the logical upstream release tag.",
|
|
)
|
|
resolve_group.add_argument(
|
|
"--resolve-install-tag",
|
|
nargs = "?",
|
|
const = "latest",
|
|
help = (
|
|
"Resolve a llama.cpp tag such as 'latest' to the concrete upstream tag "
|
|
"selected by the current published-release policy."
|
|
),
|
|
)
|
|
resolve_group.add_argument(
|
|
"--resolve-source-build",
|
|
nargs = "?",
|
|
const = "latest",
|
|
help = ("Resolve the source-build fallback plan."),
|
|
)
|
|
parser.add_argument(
|
|
"--output-format",
|
|
choices = ("plain", "json"),
|
|
default = "plain",
|
|
help = "Resolver output format. Defaults to plain.",
|
|
)
|
|
return parser.parse_args()
|
|
|
|
|
|
def emit_resolver_output(payload: dict[str, Any], *, output_format: str) -> None:
|
|
if output_format == "json":
|
|
print(json.dumps(payload, sort_keys = True))
|
|
return
|
|
if "llama_tag" in payload:
|
|
print(payload["llama_tag"])
|
|
return
|
|
if {
|
|
"source_url",
|
|
"source_ref_kind",
|
|
"source_ref",
|
|
}.issubset(payload):
|
|
print(
|
|
"\t".join(
|
|
(
|
|
str(payload["source_url"]),
|
|
str(payload["source_ref_kind"]),
|
|
str(payload["source_ref"]),
|
|
)
|
|
)
|
|
)
|
|
return
|
|
print(json.dumps(payload, sort_keys = True))
|
|
|
|
|
|
def main() -> int:
|
|
args = parse_args()
|
|
if args.resolve_llama_tag is not None:
|
|
resolved = resolve_requested_llama_tag(
|
|
args.resolve_llama_tag,
|
|
args.published_repo,
|
|
args.published_release_tag or "",
|
|
)
|
|
emit_resolver_output(
|
|
{
|
|
"requested_tag": normalized_requested_llama_tag(args.resolve_llama_tag),
|
|
"llama_tag": resolved,
|
|
},
|
|
output_format = args.output_format,
|
|
)
|
|
return EXIT_SUCCESS
|
|
|
|
if args.resolve_install_tag is not None:
|
|
resolved = resolve_requested_install_tag(
|
|
args.resolve_install_tag,
|
|
args.published_release_tag or "",
|
|
args.published_repo,
|
|
)
|
|
emit_resolver_output(
|
|
{
|
|
"requested_tag": normalized_requested_llama_tag(
|
|
args.resolve_install_tag
|
|
),
|
|
"llama_tag": resolved,
|
|
},
|
|
output_format = args.output_format,
|
|
)
|
|
return EXIT_SUCCESS
|
|
|
|
if args.resolve_source_build is not None:
|
|
plan = resolve_source_build_plan(
|
|
args.resolve_source_build,
|
|
args.published_repo,
|
|
args.published_release_tag or "",
|
|
)
|
|
emit_resolver_output(
|
|
{
|
|
"requested_tag": normalized_requested_llama_tag(
|
|
args.resolve_source_build
|
|
),
|
|
"source_url": plan.source_url,
|
|
"source_ref_kind": plan.source_ref_kind,
|
|
"source_ref": plan.source_ref,
|
|
"compatibility_upstream_tag": plan.compatibility_upstream_tag,
|
|
},
|
|
output_format = args.output_format,
|
|
)
|
|
return EXIT_SUCCESS
|
|
|
|
if not args.install_dir:
|
|
raise SystemExit(
|
|
"install_llama_prebuilt.py: --install-dir is required unless --resolve-llama-tag, --resolve-install-tag, or --resolve-source-build is used"
|
|
)
|
|
install_prebuilt(
|
|
install_dir = Path(args.install_dir).expanduser().resolve(),
|
|
llama_tag = args.llama_tag,
|
|
published_repo = args.published_repo,
|
|
published_release_tag = args.published_release_tag or "",
|
|
simple_policy = args.simple_policy,
|
|
override_has_rocm = args.has_rocm,
|
|
)
|
|
return EXIT_SUCCESS
|
|
|
|
|
|
if __name__ == "__main__":
|
|
try:
|
|
raise SystemExit(main())
|
|
except SystemExit:
|
|
raise
|
|
except BusyInstallConflict as exc:
|
|
log(
|
|
f"fatal helper busy conflict: {textwrap.shorten(str(exc), width = 400, placeholder = '...')}"
|
|
)
|
|
raise SystemExit(EXIT_BUSY)
|
|
except PrebuiltFallback as exc:
|
|
# Expected when the published repo (e.g. ggml-org/llama.cpp) has no
|
|
# prebuilt manifest. Exit quietly with EXIT_FALLBACK so the caller
|
|
# falls back to source build without a noisy "fatal helper error".
|
|
log(textwrap.shorten(str(exc), width = 400, placeholder = "..."))
|
|
raise SystemExit(EXIT_FALLBACK)
|
|
except Exception as exc:
|
|
message = textwrap.shorten(str(exc), width = 400, placeholder = "...")
|
|
log(f"fatal helper error: {message}")
|
|
raise SystemExit(EXIT_ERROR)
|