unsloth/studio/setup.ps1
Daniel Han 3ab8dce97a
install: let UNSLOTH_TORCH_INDEX_FAMILY / _URL override CUDA wheel detection (#6692)
* install: let UNSLOTH_TORCH_INDEX_FAMILY / _URL override CUDA wheel detection

get_torch_index_url (and the studio-update mirror _detect_cuda_torch_index_url)
chose the torch wheel family solely by probing the host GPU, with no override.
In a headless / container / CI build the host driver is visible via the
/proc/driver/nvidia/gpus fallback but nvidia-smi cannot report a CUDA version,
so the function fell back to its cu126 default and installed the wrong wheels
(e.g. a cu128 image got cu126 torch).

Add an explicit override checked before any probing, in both the shell installer
and the Python studio-update path:
  - UNSLOTH_TORCH_INDEX_URL   full index URL, used verbatim (wins)
  - UNSLOTH_TORCH_INDEX_FAMILY family (cpu, cu128, rocm6.4, ...) appended to the
                               mirror base (UNSLOTH_PYTORCH_MIRROR still honoured)

This matches how the published GPU images select CUDA -- vLLM and SGLang take the
CUDA version from an explicit build ARG rather than detecting it, and the Unsloth
Docker base image already pins the cu128 index directly. Desktop installs are
unchanged: with no override set, detection runs exactly as before.

Adds test_get_torch_index_url.sh cases for the override (family, full URL,
precedence, mirror base, trailing-slash strip, empty-ignored).

* install: make the torch-index override authoritative across ROCm paths

Address review feedback on the override added in this PR so a pinned index is
honoured everywhere, not just in get_torch_index_url:

- Skip the WSL ROCm bootstrap (root privilege + large downloads, probes
  /dev/dxg) when UNSLOTH_TORCH_INDEX_URL / _FAMILY is set; it previously ran
  before the override was consulted.
- Skip the Radeon/Strix rerouting (which re-probes the GPU and overwrites the
  resolved URL with repo.radeon.com / repo.amd.com) when the index is pinned, so
  an explicit ROCm override (e.g. UNSLOTH_TORCH_INDEX_FAMILY=rocm6.4) is kept.
- install_python_stack.py: derive _TORCH_BACKEND from the override when
  UNSLOTH_TORCH_BACKEND is unset (standalone studio update), so _ensure_rocm_torch
  / _ensure_cuda_torch repair to the requested family instead of re-detecting.
- Strip ALL leading/trailing slashes in the shell override to match the Python
  side (avoids 404s on strict pip proxies).

Adds test cases for double-slash and leading/trailing-slash overrides.

* install: honor pinned torch index in CUDA/ROCm repair paths

Follow-up to the override work in this PR: the get_torch_index_url / install.sh
reroute already respect a pinned UNSLOTH_TORCH_INDEX_URL / _FAMILY, but the
Python repair helpers in install_python_stack.py still re-probed the GPU and
could overwrite the pinned family. Make the pin authoritative there too:

- _ensure_cuda_torch: an explicit cu* pin commits to CUDA wheels, so repair a
  ROCm-poisoned venv even when no NVIDIA GPU is visible here (headless /
  container / CI cross-install), instead of bailing on the GPU-presence gate.
- _ensure_rocm_torch: skip the AMD per-gfx (Strix) reroute when a ROCm index is
  pinned, and in the generic reinstall path install from the pinned URL verbatim
  rather than re-detecting the host ROCm version. gfx*/rocm7.2 indexes serve
  torch 2.11+, so select the 2.11 package specs for a gfx leaf.
- install.sh: raise the torch constraint to 2.11 for */gfx* indexes too, matching
  rocm7.2, so a pinned full-URL/family override that returns early keeps a valid
  constraint.

Add _explicit_torch_index_url / _explicit_rocm_torch_index_url helpers and tests
covering the no-GPU CUDA pin repair and the explicit gfx index honored verbatim.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: honor torch-index override on the Windows installers too

The pinned-index work landed for install.sh and install_python_stack.py, but the
Windows installers still picked the wheel index from GPU probing. Extend the same
UNSLOTH_TORCH_INDEX_URL / _FAMILY contract so a pinned index wins on every platform:

- install.ps1: Get-TorchIndexUrl returns the pinned URL/family before nvidia-smi
  probing; the AMD ROCm reroute is skipped when the index is pinned, so an explicit
  cpu/cu* pin on an AMD host is not overwritten.
- studio/setup.ps1: add shared Get-PinnedTorchIndexUrl / Get-TorchIndexLeaf helpers;
  the stale-venv check, the install selection and the AMD reroute all honor the pin,
  and the CPU/CUDA install pulls from the resolved index URL.
- tests: parity test that all four installers read both override vars and the two
  Windows installers gate the AMD reroute on the pinned flag.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: complete pinned-index handling for ROCm/Windows edge cases

Follow-ups to the override work flagged in review:

- install.ps1: a pinned gfx*/rocm>=7.2 index previously skipped the AMD reroute
  that sets the torch>=2.11 floor, so the generic install used torch>=2.4,<2.11
  and could resolve the known-bad _grouped_mm wheel. Route a pinned ROCm index
  through the ROCm install path with the 2.11 floor + companions, and guard the
  companion-spec lookup so a skipped reroute block cannot null-deref.
- studio/setup.ps1: the stale-venv check compared the installed flavor (cuXXX/cpu,
  with +rocm misread as cpu) against the raw pinned leaf (gfx1151 / rocm6.4), so a
  correct pinned ROCm venv was always marked stale. Classify +rocm wheels as the
  generic 'rocm' flavor and normalize a pinned rocm*/gfx* leaf to 'rocm' before
  comparing (cu* stays specific so cu126-vs-cu128 still rebuilds).
- install_python_stack.py: _ensure_cuda_torch now also reinstalls from a pinned
  CUDA index when the venv carries a CPU wheel (headless CPU-venv-to-CUDA
  cross-install via 'studio update'), not only when it finds a ROCm build.
- tests: parity assertions already cover all four installers honoring the override.

* install: finish pinned ROCm/CUDA edge cases on Windows + repair path

Follow-ups to the previous round:

- studio/setup.ps1: a pinned gfx*/rocm>=7.2 index now routes through the ROCm
  install path with the 2.11 floor + companions (it previously fell through to the
  CUDA branch with bare torch/torchvision/torchaudio against the ROCm index). The
  CPU/CUDA fallback index is forced to the CPU wheel index when a ROCm index is
  active, so a failed pinned-ROCm install does not retry the ROCm mirror.
- studio/setup.ps1: the stale-venv check no longer treats an unrecognized pinned
  URL leaf (e.g. a PEP 503 mirror ending in /simple) as a torch flavor tag, which
  was marking a correct venv stale; cu*/cpu/rocm/gfx leaves are still compared.
- install.ps1: the post-failure CPU fallback uses an explicit CPU index instead of
  , which for a pinned ROCm index was the ROCm mirror itself (so the
  'fallback' just retried the failing index and aborted the installer).
- install_python_stack.py: _ensure_cuda_torch now also reinstalls when the venv's
  CUDA family differs from a pinned one (installed cu126 vs pinned cu128), not only
  CPU->CUDA; the probe reports the installed cuXXX tag for the comparison.

* install: keep the ROCm to CPU fallback install inside the retry-helper window

The pinned-ROCm CPU fallback computes an explicit CPU index, but the comment
explaining why it cannot reuse $TorchIndexUrl pushed the actual
Invoke-InstallCommandRetry / --force-reinstall call more than 600 chars past the
"ROCm PyTorch install failed" message, so test_pr5940_followups's window check
no longer saw the retry helper. Move the CPU-index computation and its comment
above the failure substep so the retrying force-reinstall stays adjacent to the
message. No behavior change: same explicit CPU index, same retry, same
--force-reinstall.

* install: address #6692 review round 5 (ROCm/CPU pin edge cases)

setup.ps1:
- Stale-venv check: treat an AMD/ROCm host (HasROCm or a resolved gfx arch) with
  no explicit pin as expecting "rocm", not "cpu", so a healthy +rocm venv is not
  flagged stale (which made installer-managed setup exit and direct update rebuild).
- Pinned-ROCm install failure now routes into the force-reinstall CPU branch:
  CuTag stays the rocm/gfx leaf on failure, so the condition also checks
  ROCmCpuFallback; otherwise the CUDA branch installed from the CPU index without
  --force-reinstall and kept the partial ROCm torch.
- Explicit ROCm pin compare no longer collapses gfx*/rocm* to a generic "rocm":
  it compares the +rocmX.Y version (and the torch 2.11 line for gfx pins) so
  changing the pinned family (e.g. rocm6.4 -> gfx1151) rebuilds and applies it.

install_python_stack.py:
- _ensure_rocm_torch: an explicit ROCm wheel-index pin now bypasses the
  NVIDIA-present / no-AMD-GPU / unreadable-ROCm gates (headless/container/CI
  cross-install), mirroring the explicit-CUDA-pin bypass in _ensure_cuda_torch.
- Add _ensure_cpu_torch: an explicit CPU pin (FAMILY=cpu or /cpu URL) now has a
  repair path that reinstalls CPU torch over an existing CUDA/ROCm build on a
  standalone update (which skips install.sh's flavor enforcement).

install.sh:
- Pin torchvision/torchaudio companions alongside torch for the rocm7.2 / per-gfx
  index and the Strix reroute (those AMD indexes publish companions independently
  and a bare name can resolve a torch-2.12-built wheel, an ABI mismatch).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* torch-index override: classify CUDA pin by leaf; trim blank shell overrides

_ensure_cuda_torch only overrode the NVIDIA-presence gate for *any* pinned index,
so a non-CUDA mirror URL (or a ROCm/CPU pin) on a non-NVIDIA host with ROCm torch
could force a CUDA reinstall over a working ROCm venv. Add
_explicit_cuda_torch_index_url() (leaf cu*), matching the ROCm/CPU helpers, and
gate on it instead.

install.sh::get_torch_index_url treated a whitespace-only UNSLOTH_TORCH_INDEX_URL
/ _FAMILY as authoritative (yielding an invalid index), unlike the Python .strip()
and PowerShell IsNullOrWhiteSpace paths; trim leading/trailing whitespace first.

* install: honor pinned torch index over CVD/GPU gates and fix leaf-based ROCm classification

- install_python_stack.py: an explicit cu* pin now clears the CUDA_VISIBLE_DEVICES
  empty/-1 hide gate as well as the NVIDIA-presence gate, so
  CVD=-1 UNSLOTH_TORCH_INDEX_FAMILY=cu128 studio update repairs to CUDA wheels
  (parity with install.sh's get_torch_index_url override, which skips all GPU
  probing). Unpinned CVD=-1 still skips.
- install_python_stack.py: _ensure_cpu_torch installs the bounded _CPU_TORCH_PKG_SPEC
  instead of a bare torch/torchvision/torchaudio trio; the /cpu index now also
  serves torch 2.11+, which is outside the supported <2.11 range.
- install.sh: the torch>=2.11 constraint case matches the index leaf (rocm7.2|gfx*)
  instead of the whole URL, so a mirror base path containing a gfx/rocm7.2 segment
  with a cu*/cpu family is not false-matched onto the 2.11 line.
- setup.ps1: the stale-venv check expects rocm torch only for arches the install
  path maps to a repo.amd.com wheel index; an unmapped/unreadable arch installs
  CPU, so a correct CPU venv is no longer marked stale.
- Tests for each of the above.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: tighten pinned torch-index override edge cases

- install.sh: trim whitespace-only UNSLOTH_TORCH_INDEX_URL/_FAMILY before the
  _torch_index_pinned guard, matching get_torch_index_url, so a blank override no
  longer skips the WSL bootstrap and Radeon/Strix reroutes while detection still
  picks the normal index.
- install.sh / install.ps1 / setup.ps1 / install_python_stack.py: force the torch
  2.11 floor only for the gfx families with the <2.11 _grouped_mm bug (gfx120X-all,
  gfx1151, gfx1150). A pinned override to gfx110X-all/gfx90a/gfx908 stays on the
  default range, matching the automatic AMD path.
- install_python_stack.py _ensure_cuda_torch: treat an untagged CUDA build under a
  CUDA pin as a family mismatch (reinstall), and match cuXXX pins narrowly (cu +
  digits) so a custom/current mirror leaf no longer forces CUDA over a CPU/ROCm venv.
- install_python_stack.py _ensure_rocm_torch: reinstall when an explicit ROCm pin
  names a different ROCm family than the already-installed ROCm torch (the ROCm
  analogue of the CUDA cuXXX mismatch repair).

Adds tests for each case.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: fix second-order edge cases in pinned torch-index ROCm/CUDA handling

Parse the ROCm torch probe positionally so an empty HIP marker is kept:
CPU/CUDA torch no longer reads as HIP, so the ROCm reinstall is not skipped.
Emit one "<marker>|<version>" line (like the CUDA probe) for a robust parse.

Limit the gfx torch 2.11 expectation to the install allowlist
(gfx120X-all/gfx1151/gfx1150). A pinned gfx110X-all/gfx90a/gfx908 index stays
on the default <2.11 specs, so a correct 2.10+rocm wheel is no longer judged a
mismatch and force-reinstalled every update.

Distinguish an AMD per-arch wheel (three-part +rocmA.B.C) from a generic
pytorch.org wheel (two-part +rocmA.B): a gfx per-arch pin over a generic 2.11
wheel now reinstalls the per-arch wheel, while an already-installed per-arch
wheel is not re-flagged (no reinstall loop).

Mirror all of the above in setup.ps1 via new Test-RocmGfx211Leaf /
Test-CudaFamilyLeaf / Get-RocmPinStaleTags helpers, reused by both the
install-spec path and the stale-venv check so they cannot diverge again.
Require a digit after "cu" (^cu[0-9]) in setup.ps1, install.ps1 and install.sh
so a mirror leaf like /custom or /current is not branded CUDA and does not
rebuild the venv every run.

Add tests: CPU/CUDA probe -> has_hip_torch False; gfx110X-all pin + 2.10 wheel
not stale; gfx1151 pin + generic 2.11 wheel stale; gfx1151 pin + per-arch wheel
not stale; /custom and /current not CUDA; plus cross-language allowlist and
cu-digit parity guards, and a PowerShell unit test for the new setup.ps1 helpers.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix ROCm/gfx pin case normalization, ROCm-tag requirement, and CUDA-leaf classification

Normalize torch-index leaves to lowercase before the gfx*/rocm*/cu* allowlist
matches so the canonical gfx120X-all (capital X) gets the torch 2.11 floor in
install.sh (leaf, flavor and repairable helpers). Require an installed +rocm
local tag before a rocmX.Y or non-2.11 gfx pin is judged satisfied in
setup.ps1 Get-RocmPinStaleTags and the Python _rocm_pin_family_mismatch, so an
untagged CPU/CUDA wheel never leaves the pin unapplied. Classify a leaf as CUDA
only via ^cu[0-9]: the Python _TORCH_BACKEND derivation now uses
_is_cuda_family_leaf, and install.sh brands cuda only on cu[0-9]* (unset on an
unknown /current /custom mirror leaf) so the stack probes the GPU instead of
skipping ROCm repair. Add bash, Python and PowerShell tests for capital
gfx120X-all floor, current/custom not-cuda, and untagged-wheel ROCm pins.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: converge torch-index pin detection via a per-venv marker

Introduce a torch-index MARKER that records the exact wheel --index-url used
after each successful torch install, so `unsloth studio update` / repair makes
the "did the pinned index change?" decision by an EXACT string compare rather
than inferring it from the wheel +rocm/+cu version tag. The tag cannot encode
the AMD per-arch gfx family (two 2.11 gfx indexes both install +rocm7.13.0), so
the tag heuristic missed a gfx1151 -> gfx120X-all switch and a custom-URL swap.

Marker path is per-venv (.unsloth-torch-index), one line = the resolved index
URL, written atomically (temp + rename). Path, format and normalization are
shared across all four installers (install.sh, install_python_stack.py,
setup.ps1, install.ps1).

- Reapply gfx pins on a per-arch target change: the marker's exact compare
  reinstalls when the pinned index differs, even when both wheels share a tag.
- Honor custom ROCm URL pins during repair: an explicit index whose leaf is not
  rocm/gfx/cu/cpu (e.g. simple, current) now reinstalls torch VERBATIM from the
  pin when it differs from the marker ("URL wins verbatim").
- Align the KNOWN-2.11 rocm/gfx set to exactly rocm7.2 plus the gfx allowlist
  gfx120x-all/gfx1151/gfx1150 in every language; stop treating an unknown newer
  rocm (rocm7.3, which does not exist) as the 2.11 line speculatively.

Backward compatible: with no marker (old venvs, torch installed out-of-band) the
existing +rocm/version-tag heuristics still decide, and a matching marker never
reinstall-loops. A cu128 CUDA pin stays a CUDA pin; custom and current leaves are
not CUDA. Adds marker tests (py/sh/ps) plus cross-installer parity checks.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: keep the torch-index marker additive to flavor validation

Three narrow fixes in the marker-based stale-venv detection:

- setup.ps1: a matching marker no longer overwrites the detected installed
  flavor. The marker compare is now an additional rebuild trigger, so a stale
  wheel (torch swapped to a +cpu build while the marker still records a cuXXX
  pin) is still caught by the flavor check instead of being masked as up to date.

- setup.ps1: a supported AMD arch carrying CPU torch is no longer marked stale
  and wiped. The downstream AMD Windows ROCm override upgrades CPU torch to ROCm
  in place, so wiping first would delete the venv and abort with "Virtual
  environment not found". Only a genuinely wrong CUDA wheel still rebuilds.

- install.sh: the Radeon --find-links path records its repo.radeon.com base in
  the marker instead of the generic pytorch.org ROCm fallback index, so a later
  pin to that generic family correctly reinstalls rather than comparing equal.
  Mirrors install.ps1/setup.ps1, which already record the real AMD index.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: honor custom pins and repair pinned venvs in place

Four follow-ups to the torch-index marker work:

- install_python_stack.py: _ensure_cuda_torch/_ensure_rocm_torch now bail when an
  explicit custom-index pin names no known torch family, so a verbatim URL override
  (a private/simple mirror) is not clobbered by auto-detected CUDA/ROCm wheels
  before _ensure_verbatim_torch_index applies it.

- install_python_stack.py: the ROCm marker is additive, not a substitute -- a
  matching marker still runs the family/version check so a wheel swapped after the
  marker was written is caught. Mirrors setup.ps1.

- setup.ps1: a stale venv under an explicit pin, whose torch still imports, is
  repaired in place (force-reinstall torch from the pin in the dependency pass)
  instead of wiped. The wipe path only delegates to install.ps1, so on a direct
  update it stranded the user at "Virtual environment not found" instead of
  applying the new pin. A broken venv or unpinned drift still wipes/delegates.

- install.ps1: when a pinned ROCm install fails over to a CPU base, the marker now
  records the CPU index actually used instead of the ROCm pin, so the next managed
  setup does not see CPU torch under a ROCm pin and abort as stale.

* setup.ps1: keep the ROCm CPU-fallback force line the pr5940 test guards

5c93ffd4 folded the pin-change force-reinstall into the ROCm CPU-fallback
condition on one line, so the exact literal that test_pr5940_followups.py checks
(if ($ROCmCpuFallback) { $cpuForce = @("--force-reinstall") }) no longer appeared
and the test failed. Split the two conditions into separate if lines: the ROCm
fallback line is restored verbatim and the pin-change force is its own line. Both
still set $cpuForce to the array, so @splat passes one arg.

* install: honor exact CUDA/custom index URL pins in the torch-index marker

Address three Codex review findings on the torch-index marker mechanism:

- install.sh: after the ROCm CPU repair reinstalls torch from the generic
  $TORCH_INDEX_URL, record that as the marker source. A Radeon --find-links
  install set _TORCH_MARKER_INDEX_URL to its repo.radeon.com base earlier, so
  leaving it made the marker misreport Radeon wheels and a later Radeon pin would
  compare equal and skip a needed reinstall.

- install_python_stack.py: _ensure_cuda_torch now consults the exact-URL marker
  (_marker_pin_mismatch) when the installed +cuXXX tag matches the pinned leaf,
  so a same-leaf CUDA mirror change (official cu128 to an internal cu128 mirror)
  is reinstalled and re-recorded instead of skipped.

- _normalize_index_url / _normalize_family_leaf (install.sh, setup.ps1,
  install_python_stack.py): lowercase only KNOWN wheel-family leaves (rocm/gfx/
  cpu/cuXXX) so gfx120X-all still matches gfx120x-all, while a custom
  (unknown-family) leaf keeps its case so a verbatim URL pin like /Current does
  not compare equal to /current. Tests updated to assert the refined behavior.

* install: fix 3 torch-index marker edge cases (CPU mirror pin, Radeon leaf, migrated venv)

Addresses three review findings on the torch-index override path:

1. CPU index URL change on an already-CPU venv. _ensure_cpu_torch returned
   early whenever torch was already a CPU build, so a standalone update that
   moved the pin (official /cpu -> a private UNSLOTH_PYTORCH_MIRROR /cpu, same
   +cpu tag) never reinstalled. It now consults the exact-URL marker and
   reinstalls only when _marker_pin_mismatch reports a different index,
   mirroring the CUDA/ROCm same-family handling. A matching marker (or none)
   still leaves CPU torch untouched, so there is no reinstall loop.

2. Radeon find-links directory misclassified as a pip ROCm family. A
   repo.radeon.com/.../rocm-rel-7.2.1 leaf starts with "rocm" but is a
   find-links listing, not a pip --index-url. The old startswith(("rocm",
   "gfx")) test routed it into a --index-url reinstall that fails against
   find-links. New _is_pip_rocm_family_leaf gates on ^rocm\d / gfx (matching
   install.sh's rocm[0-9]* and setup.ps1's ^(rocm[0-9]|gfx)), so a Radeon URL
   routes to the verbatim/marker path instead.

3. Migrated venv rewriting its marker to a pin it did not install. install.sh
   and install.ps1 write the marker unconditionally, so a migration that
   preserves existing torch recorded the newly requested pin and a later
   update then found a matching marker and skipped the reinstall the pin
   needs (e.g. a per-arch gfx1151 -> gfx120X-all switch, identical +rocm tag).
   Both now track _TORCH_INSTALLED_THIS_RUN and write the marker only when
   torch was actually installed or repaired this run.

Also add Get-NormalizedFamilyLeaf to the setup.ps1 helper-extraction list in
test_torch_index_marker.ps1 (it was added to setup.ps1 and the shell test in an
earlier round but missed here) and add two unit tests covering findings 1 and 2.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: keep pinned torch repairs on the pinned index

Two fixes for explicit index pins (UNSLOTH_TORCH_INDEX_FAMILY / _URL):

1. install_python_stack.py's repair paths ran uv without clearing the
   inherited uv index env vars. uv resolves the default index (--index-url
   or --default-index) at the LOWEST priority, so a UV_INDEX or
   UV_EXTRA_INDEX_URL mirror in the environment won for any package it
   served: a cu128-pinned repair could install torch from the mirror and
   then record the cu128 marker it never used. Verified empirically: with
   UV_EXTRA_INDEX_URL=.../cu126 exported, uv pip install torch
   --index-url .../cu128 resolves torch 2.13.0+cu126. Strip the four uv
   index env vars for pinned-index commands only, mirroring the gate
   install.sh, install.ps1 and setup.ps1 already have; non-pinned installs
   keep the user's mirror.

2. install.ps1 routed any pinned leaf matching rocm* through the ROCm
   --default-index path, so a custom find-links leaf like rocm-rel-7.2.1
   was treated as a PEP 503 ROCm index and could silently fall back to CPU
   torch on resolution failure. Require a digit after rocm, matching
   install.sh's rocm[0-9]* and install_python_stack.py's ^rocm\d.

Adds parity + unit tests for both (11 new tests).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: keep pinned repairs off UV_TORCH_BACKEND and narrow setup.ps1's rocm pin match

Round 2 of the pinned-index hardening:

1. _build_uv_cmd converted UV_TORCH_BACKEND into --torch-backend before the
   new env isolation could act, and uv's torch backend redirects torch
   resolution to its own per-backend index even when --index-url is given
   (verified: a cu128-pinned dry run with UV_TORCH_BACKEND=cpu resolves
   torch 2.13.0+cpu). Pinned-index commands now never receive the flag and
   UV_TORCH_BACKEND joins the stripped env vars, so uv cannot re-read it.

2. setup.ps1's pinned reroute had the same bare rocm* glob install.ps1 had:
   a custom find-links leaf like rocm-rel-7.2.1 was routed through the ROCm
   --index-url path instead of the verbatim unknown-pin path. Now requires
   a digit after rocm, matching install.ps1, install.sh and
   _is_pip_rocm_family_leaf.

3. The marker test's case-normalization checks used -eq, which is
   case-insensitive in PowerShell, making them vacuous, and the unknown-leaf
   expectation was written lowercased while the implementation deliberately
   preserves custom-leaf case. Tightened to -ceq with the case-preserving
   expected value.

Adds unit + parity tests for 1 and 2 (5 new tests).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: extend the pinned-index guards to every remaining surface

Round 3 of the pinned-index hardening, closing the same holes on the
surfaces the earlier rounds missed:

1. install.sh's pinned-install env scrub now clears UV_TORCH_BACKEND (uv's
   torch backend redirects torch resolution to its own per-backend index
   even against --default-index), and both PowerShell wrappers clear it in
   their pinned-install scrubs, matching install_python_stack.py.

2. setup.ps1's marker stale check still classified any rocm* leaf as a
   PyTorch ROCm family while the install selection is digit-gated, so a
   custom rocm-current / rocm-rel-7.2.1 pin stale-compared as
   not-rocm vs rocm and force-reinstalled on every studio update. The
   stale check now uses the same ^rocm\d gate.

3. install_python_stack.py's pinned-command scrub also strips
   PIP_EXTRA_INDEX_URL for the pip fallback: pip adds the env extra index
   in addition to --index-url, so an inherited mirror could satisfy torch
   off the pin while the marker recorded the pinned URL. PIP_INDEX_URL
   needs no strip since the explicit --index-url flag overrides it.

Parity + unit tests extended (4 new tests).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: scrub find-links and carry the pinned scrub through pip fallbacks

Round 4 of the pinned-index hardening:

1. UV_FIND_LINKS joins every pinned-install scrub (install.sh, install.ps1,
   setup.ps1, install_python_stack.py): uv's --find-links locations can
   satisfy torch off the pinned index the same way an extra index does.

2. setup.ps1's Fast-Install restored the scrubbed vars in its finally
   BEFORE the pip fallback ran, and never touched the pip env vars at all,
   so a failed uv attempt fell back to python -m pip with an inherited
   PIP_EXTRA_INDEX_URL / PIP_FIND_LINKS able to win over the pinned
   --index-url. The scrub now wraps the whole function (uv attempt + pip
   fallback) and includes the pip vars; restore happens after both.

3. install_python_stack.py's scrub also strips PIP_FIND_LINKS for its own
   pip fallback, completing the PIP_EXTRA_INDEX_URL fix from round 3.

Parity tests extended (2 new tests).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: digit-gate rocm leaves in marker normalization and ROCm side effects

Round 5 of the pinned-index hardening (three custom-rocm-leaf edge cases):

1. _normalize_family_leaf lowercased every leaf starting with rocm, so a
   custom mirror leaf like rocm-Current compared equal to its lowercase form
   and a case-only pin change was skipped. URL paths can be case-sensitive.
   The rocm prefix is now digit-gated (rocm[0-9]*, matching
   _is_pip_rocm_family_leaf) in install.sh, setup.ps1 and
   install_python_stack.py, so only true family leaves (rocm7.2) are
   lowercased; a custom rocm-* leaf keeps its case.

2. setup.ps1 Test-MarkerPinMismatch compared normalized URLs with -ne, which
   is case-insensitive in PowerShell, so a case-only marker change (Simple
   vs simple) was treated as matching and the reinstall skipped. Now -cne.

3. install.sh gated the AMD bitsandbytes install and the "repair ROCm torch"
   --default-index reinstall on a bare whole-URL rocm glob, so a custom
   CPU/CUDA/private index whose leaf merely starts with rocm (rocm-current)
   was force-repaired from the wrong ROCm-only path whenever torch.version.hip
   was empty. Both now gate on _torch_index_is_rocm_family, computed once from
   the digit-gated leaf (rocm[0-9]*/gfx*).

Tests: 4 new parity assertions plus 2 case-sensitivity marker checks.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: apply an explicit custom torch-index pin on the first update

Round 6: an explicitly-set custom (unknown-family) UNSLOTH_TORCH_INDEX_URL
was silently ignored on the first `studio update` of a venv that predates
the marker feature, on both platforms, because the no-marker case was
treated as "do nothing" and the version-tag heuristics cannot judge an
unknown leaf.

1. install_python_stack.py _ensure_verbatim_torch_index now reinstalls
   verbatim when the marker is ABSENT (None), not only when it differs, and
   short-circuits only when the marker already records this exact pin. It
   then writes the marker, so every later update is a no-op. A user who did
   not set the override gets pin=None and is untouched, so an out-of-band
   torch install is never clobbered.

2. setup.ps1: for an unknown-family pin on a marker-less venv the stale-venv
   check now sets PinChangedForceReinstall so the torch block reinstalls in
   place from the pin. It deliberately does NOT set shouldRebuild, which
   would wipe the venv and strand a direct `studio update`.

3. setup.sh (the Linux `studio update` entry point) skipped
   install_python_stack.py entirely when unsloth was already current, so the
   marker-driven reinstall (both the verbatim custom pin and the cu/rocm
   flavor and family-change repair, e.g. gfx1151 to gfx120X-all) never ran.
   It now forces the dependency pass when a torch-index pin env var is set;
   the pass is idempotent and no-ops when the marker already matches. This
   mirrors setup.ps1's stale-venv pre-check.

Tests: 3 new parity assertions.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* test: expect first-update reinstall for a no-marker custom index pin

Follow-up to d671d8fb2: _ensure_verbatim_torch_index now applies an
explicit unknown-family URL pin verbatim on the first update when the
marker is absent (instead of no-op), so the old
test_verbatim_custom_url_no_marker_is_noop assertion was stale. Rewritten
as test_verbatim_custom_url_no_marker_reinstalls_once: asserts the one
verbatim reinstall from the pinned URL, that the marker is written, and
that a second call with the pin still set is idempotent (no reinstall
loop).

* install: gate the pinned update pass on the marker and record a pin baseline

Round 8, two follow-ups to the round-6 first-update pin fix:

1. setup.sh forced the full dependency pass on EVERY `studio update` while a
   torch-index pin stayed exported, even after the marker already recorded the
   same pin, turning quick updates into the expensive pass every time. It now
   probes install_python_stack.py --torch-pin-needs-apply (which reuses the
   exact marker normalization) and forces the pass only when the pin is not yet
   applied (marker absent or different); an already-applied persistent pin keeps
   the fast path. A probe error fails safe toward running the pass. setup.ps1
   gets the same probe in its fast path for parity.

2. A known-family full-URL pin on a venv predating the marker (e.g. an installed
   cu128 build and UNSLOTH_TORCH_INDEX_URL pointing at a same-family mirror) left
   the marker absent forever: the _ensure_* helpers deliberately do not force a
   multi-GB reinstall of identical-family wheels on an old venv, so nothing
   recorded the pin and every update re-entered the pass. _record_torch_index_pin_baseline
   now records the resolved pin as a baseline after the ensure sequence when the
   family already matches and no marker exists, so the pin is tracked (a later
   genuine change is detected and applied) and the update loop is broken, without
   the redundant reinstall.

Tests: 3 new baseline unit tests, 4 new parity assertions, and the CLI probe.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* setup.sh: keep the pin probe's exit 1 from killing the update under set -e

The --torch-pin-needs-apply probe deliberately exits 1 for the common
steady-state answer (pin already recorded, keep the fast path), but it ran
as a bare command under set -euo pipefail, so the whole studio update
aborted before the exit code was even captured. Absorb the status with
|| _PIN_NEEDS_APPLY=$? and pre-seed 0 so all three outcomes route as
documented: 0 runs the pass, 1 keeps the fast path, anything else fails
safe into the pass. Parity test asserts the guard.

* install: strip pin credentials, disable uv config discovery, bound verbatim installs

Four verified fix groups from a 12-reviewer audit of the torch-index
override feature, each reproduced before fixing:

1. Credential persistence: all four marker writers stored the raw pin URL,
   so an authenticated pin (https://user:token@mirror/simple) persisted its
   credentials in .unsloth-torch-index (mode 0644 under a default POSIX
   umask) and install_python_stack.py printed pin URLs verbatim in repair
   messages. Userinfo is now stripped before persisting and in every
   log/substep that interpolates a pin, via lockstep helpers
   (_strip_index_url_credentials in install.sh / install_python_stack.py,
   Remove-IndexUrlCredentials in install.ps1 / setup.ps1). The three
   normalizers strip too, so an OLD marker that already carries credentials
   still compares equal to the same pin: no reinstall loop on upgrade.
   Query strings deliberately stay in the marker; two indexes distinguished
   only by query must not compare equal.

2. uv configuration discovery beat the explicit pin: with a discovered
   uv.toml declaring torch-backend = "cpu" or a [[index]] entry, uv 0.10.12
   resolves torch 2.13.0+cpu against an explicit --index-url/.../cu126 pin;
   UV_NO_CONFIG=1 restores +cu126 (reproduced both ways). The pinned-install
   scrub in all four installers now sets UV_NO_CONFIG=1 and drops
   UV_CONFIG_FILE.

3. The verbatim custom-index update path installed a bare, unconstrained
   torch trio while fresh installs from the same unknown-leaf pin apply the
   supported range; _ensure_verbatim_torch_index now installs the bounded
   trio spec, closing the fresh-vs-update asymmetry.

4. Query-bearing pins (.../cu128?token=x) classified by raw leaf split and
   force-reinstalled on every update (the installed cu128 never equals
   cu128?token=x). Query/fragment are now stripped before leaf
   classification in all four implementations; the marker comparison keeps
   the query per (1).

Rejected after verification (no change): the pin-baseline record cannot
produce a wrong later decision (every pin change still mismatches and
reinstalls from the new pin); the venv temp-file symlink scenarios require
an attacker who already owns the environment; pathological inputs like
" / cu128 / " have no realistic caller and fail loudly.

Parity, stack, rocm-support, marker (sh + ps1), pin-stale, index-url and
flavor suites all pass (455 python + full shell/ps1 batteries).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: harden custom-pin repair against clobber, broken torch, and pip config

Four follow-ups to the pinned-index audit fixes:

1. setup.ps1 routed an unknown-leaf custom pin through the CUDA branch with
   a bare torch trio while install.ps1 (fresh) and the Python verbatim path
   bound the supported range; the pinned unknown-leaf route now applies the
   same torch>=2.4,<2.11.0 bound. Known cu* leaves and unpinned runs are
   unchanged.

2. The final torch safety pass could not repair a clobbered unknown-family
   pin: intermediate dependency steps can pull torch from PyPI (the pass
   exists for exactly that reason), but the verbatim helper short-circuited
   on marker==pin and no flavor tag exists to probe. The helper now keeps a
   per-run snapshot of the installed trio (taken after a verbatim reinstall
   or on the first matching-marker pass) and reinstalls from the pin when
   the final pass sees the trio drifted. Probe failure skips the
   comparison; a reinstall refreshes the snapshot, so no loop.

3. _record_torch_index_pin_baseline could freeze a known-family pin as
   applied on a venv whose torch is missing or broken (every family helper
   returns without reinstalling when its probe fails), making
   --torch-pin-needs-apply report done forever. The baseline now probes the
   installed flavor and records only on a match: a cuXXX pin requires the
   matching +cuXXX tag, cpu requires a cpu build, rocm/gfx requires hip;
   probe failure records nothing.

4. The pinned pip fallback stripped PIP_* env vars but user/site pip config
   files still applied (a configured global.extra-index-url can satisfy
   torch off the pin). PIP_CONFIG_FILE is now pointed at the null device
   for pinned commands (pip loads no config files then), in
   _install_env_for_cmd and setup.ps1's Fast-Install pinned scrub.
   install.sh / install.ps1 have no pip fallback (uv-only), verified.

Tests: 7 new rocm_support tests (snapshot reset fixture), 1 stack test,
2 parity tests. Full battery green (464 python, sh and ps1 suites).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: complete the pin-repair coverage across the fast path and platforms

Three cross-platform follow-ups to the round-2 pin-repair fixes:

1. The --torch-pin-needs-apply probe only compared marker==pin, so a torch
   trio clobbered to the wrong family (a cpu wheel replacing cu128 via a
   later pip install) with a still-matching marker reported "already
   applied" and the _ensure_{cuda,rocm,cpu} repair never ran on the Linux
   fast path. The probe is now a testable _torch_pin_needs_apply() that also
   checks the installed flavor against a known-family pin (via a shared
   _torch_flavor_matches_pin() helper, so the baseline and the probe cannot
   drift). An unknown-family pin has no flavor to validate and a failed
   probe cannot prove drift, so both keep the fast path.

2. macOS ARM (real CPU/MPS torch, not NO_TORCH) never applied an unknown-
   family custom pin on update: both the verbatim path and the baseline
   returned on IS_MACOS while fresh install.sh honors the pin, so the marker
   was never written and setup.sh forced the dependency pass on every update
   forever. The guards are now IS_MAC_INTEL (Intel mac is already NO_TORCH),
   and the final pass applies the pin on macOS ARM.

3. The round-2 final verbatim repair sat in the step-13 sequence guarded
   not IS_WINDOWS, so on Windows a dependency step that clobbered torch after
   the pin was applied was masked by the matching marker (setup.ps1 does not
   re-validate the main venv's torch after calling this script -- verified).
   Step 13 now runs the verbatim snapshot-drift repair on Windows and macOS
   ARM too; the Linux-oriented cuda/rocm/cpu family helpers stay Linux-only.

Tests: 13 new rocm_support cases (flavor drift, macOS ARM, Windows repair),
parity updates. Full battery green (475 python, sh and ps1 suites).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: strip query tokens from the marker and tighten the pin-drift probe

Four follow-ups to the round-3 pin-repair fixes:

1. The credential stripper feeding the torch-index marker and the logged repair
   messages dropped only user:pass@ userinfo, so a private feed that carries its
   auth token in the query string (.../simple?token=SECRET) persisted the token
   in the world-readable marker (mode 0644 under a default umask) and printed it
   in substep output. All four strippers (install.sh, install.ps1,
   studio/setup.ps1, install_python_stack.py) now drop the query and fragment
   before building the sanitized URL. A query is not part of a PEP 503 index's
   identity, so this also stops a rotated token from spuriously mismatching the
   marker and forcing a needless reinstall.

2. The --torch-pin-needs-apply fast-path probe accepted an untagged CUDA build
   (no +cuXXX local tag) under a specific cuXXX pin, but _ensure_cuda_torch
   reinstalls exactly that build to enforce the pin. The probe was more lenient
   than the repair, so the repair pass was skipped on the fast path.
   _torch_flavor_matches_pin now reports a mismatch for an untagged build under a
   cuXXX pin, forcing the pass.

3. The probe's ROCm branch accepted any HIP build for a rocm/gfx pin, while
   _ensure_rocm_torch decides a reinstall with the per-arch
   _rocm_pin_family_mismatch predicate (a generic +rocm7.2 wheel under a per-arch
   gfx pin, or a wrong ROCm version, is a mismatch). The probe now reuses that
   predicate, so it is as strict as the repair. This needs the installed torch
   version, so _probe_torch_flavor now returns (marker, cutag, version) and
   _torch_flavor_matches_pin takes the pin URL (extracting the leaf internally).

4. On Windows a known-family cu*/cpu pin is applied to the main venv by setup.ps1
   before install_python_stack.py runs; a later dependency step can clobber it,
   and the GPU-aware _ensure_{cuda,cpu}_torch self-skip on Windows while the
   verbatim helper handles only unknown-family pins, so nothing repaired the
   clobber (setup.ps1 does not re-validate the main venv's torch afterward,
   verified). New _ensure_pinned_known_family_torch reinstalls a drifted cu*/cpu
   pin in the step-13 Windows/macOS-ARM branch; rocm/gfx per-arch specs stay owned
   by setup.ps1, unknown-family by the verbatim helper.

A speculative ROCm 2.11 floor was also raised but is unreachable: the rocm7.2
index publishes no 2.x wheel below 2.11.0, and an unknown newer rocm is not
floored speculatively.

Tests: query/fragment strip cases in the sh + ps1 marker suites and the Python
strip/marker tests; the tri-state helper and the probe/baseline harnesses moved
to the (marker, cutag, version) flavor with matching versions; new probe cases
(untagged CUDA, generic-rocm-under-gfx) and 8 _ensure_pinned_known_family_torch
tests; a four-way query-strip parity assertion. Full battery green (1150 python,
sh 26/26 marker, ps1 marker/flavor/pin-stale).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: reinstall markerless gfx pins and cap custom-index updates at torch 2.11

Two follow-ups from the pin-marker audit:

1. A markerless venv with a gfx per-arch 2.11 pin trusted the wheel version
   tag, which is byte-identical (+rocm7.13.0) across gfx120X-all / gfx1151 /
   gfx1150. A pre-marker install holding one gfx arch's wheel that is now
   pinned to a DIFFERENT gfx index was therefore never switched:
   _rocm_pin_family_mismatch returns no-mismatch for any three-part +rocm
   2.11 wheel, and _ensure_rocm_torch's absent-marker branch fell through to
   that heuristic. _ensure_rocm_torch now forces a one-time reinstall when the
   marker is absent AND the pin leaf is a 2.11 gfx per-arch index; the reinstall
   writes the marker, so the next update compares exactly and does not loop
   (the correctly-pinned no-reinstall guarantee then comes from the exact marker
   compare, not the ambiguous tag). Non-gfx-2.11 pins (rocmX.Y, non-2.11 gfx)
   stay on the tag heuristic -- their tags are distinguishable.

2. The verbatim custom-index update path used _CUDA_TORCH_PKG_SPEC (torch
   <2.12.0) while a FRESH install of the same unknown leaf caps torch at
   <2.11.0 (install.sh's default TORCH_CONSTRAINT, and setup.ps1's custom-pin
   branch), so a private /simple mirror publishing torch 2.11 could upgrade a
   `studio update` to a state the fresh installer never produces. Added
   _CUSTOM_INDEX_TORCH_PKG_SPEC (torch>=2.4,<2.11.0), used only by the verbatim
   path; companions stay pinned for the same exclusive --index-url ABI reason
   as _CUDA_TORCH_PKG_SPEC (a bare name could pull a torch-2.12-built
   torchvision). _CUDA_TORCH_PKG_SPEC is unchanged (known-family cu/cpu repair
   correctly tracks install.sh's widened cu ceiling).

Tests: 2 new markerless-gfx cases (one-time reinstall + marker write + no-loop
second run, and the rocmX.Y absent-marker no-op), the pre-existing markerless
gfx no-reinstall test flipped to assert the one-time reinstall (it had encoded
the old tag-trusting behavior), and the custom-index bound assertions. 488
passed.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: a matching marker must not mask a broken, clobbered, or misclassified torch

Four round-6 follow-ups, all closing cases where a matching torch-index
marker wrongly vouched for a torch that is not actually the pinned one:

1. _is_cuda_family_leaf matched cu+digits by PREFIX (^cu[0-9]), so a custom
   mirror leaf like cu128-private classified as CUDA family; the flavor check
   then compared the installed cu128 tag to the whole leaf cu128-private and
   forced a reinstall on EVERY update (never converging). The cu family is
   now matched EXACTLY (re.fullmatch cu[0-9]+), so a cu-suffixed custom leaf
   routes through the verbatim/unknown path with a stable marker. Mirrored in
   install.sh (_normalize_family_leaf: strip cu, require an all-digit
   remainder) and setup.ps1 / install.ps1 (^cu[0-9]+$).

2. _torch_pin_needs_apply returned False on a failed torch probe (missing or
   unimportable) under a matching marker, so setup.sh kept the fast path and
   a broken torch was never repaired. A failed probe now forces the pass: the
   marker cannot vouch for a torch that does not import, forcing is idempotent,
   and once torch imports again the probe succeeds and the forcing stops
   (self-resolving). Reverses the round-4 conservative choice for this case.

3. _ensure_verbatim_torch_index snapshotted the installed trio on the first
   pass with a matching marker and treated an unimportable torch (snapshot
   None) as "no drift, skip", so a torch clobbered to a broken state before
   the run was masked. A None snapshot now reapplies the pin. A torch
   clobbered to a WORKING-but-wrong build under an unknown-family pin remains
   undetectable from metadata (no flavor tag; reinstalling every update would
   be the loop this avoids) and is documented as a known limitation.

4. The step-13 Windows final repair reran only the verbatim (unknown-family)
   and known-family cu*/cpu paths, so a clobbered explicit rocm/gfx pin (the
   wheel setup.ps1 installed from AMD's per-arch index) was left in place. The
   branch now also runs _ensure_rocm_torch on Windows for an explicit rocm/gfx
   pin; it has a Windows path and no-ops when torch already links HIP, so it
   only reinstalls a genuinely clobbered ROCm venv (loop-safe).

Tests: the round-4 failed-probe-trusts-marker test flipped to force the pass;
new cases for the cu-suffix no-loop, the broken-torch verbatim reinstall, and
the Windows rocm final-repair structure; item-2 exact-cu parity assertions.
490 passed. sh/ps1 marker + flavor + pin-stale suites all green.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: repair Windows ROCm pins from the pinned URL and honor NO_TORCH

Four round-7 review items, two of them regressions in the round-6 work:

1. _torch_pin_needs_apply ignored UNSLOTH_NO_TORCH. With a torch-index env
   var set and no marker, the failed-probe branch forced the dependency pass
   on every `studio update`, and the pass (which also honors NO_TORCH) never
   installs torch or writes a marker, so nothing could ever stop the forcing.
   It now returns False immediately under NO_TORCH: the pin only matters once
   torch is actually installed.

2. The step-13 Windows final repair (round-6) restored a clobbered explicit
   rocm/gfx pin by calling _ensure_rocm_torch, whose Windows path reinstalls
   from the arch AUTO-DETECTED via hipinfo, not from the pin. A user pinning a
   different gfx family or a private mirror was restored from the wrong source
   (and the wrong marker written), and a headless box was skipped entirely
   (the arch probe returns nothing). The repair now goes through
   _ensure_pinned_known_family_torch, which reinstalls from the PINNED url with
   the same per-arch floor setup.ps1 uses (2.11-line gfx leaves) or a bare trio
   (older arches, rocmN mirrors). It is gated on IS_WINDOWS since macOS ARM has
   no ROCm, and the existing flavor check keeps it loop-safe (a matching HIP
   wheel is left alone).

3. _ensure_verbatim_torch_index's broken-torch check (round-6) used
   "_installed_trio_snapshot() is None", but that helper reports a REMOVED torch
   as "torch==absent" (a non-None tuple) and a broken import as the stale
   on-disk version, so a missing or unimportable torch under a matching marker
   was read as "no drift" and skipped. The matching-marker path now confirms
   torch health with an import probe (_probe_torch_flavor): a torch that does
   not import reapplies the pin, while a healthy torch keeps the snapshot-based
   intra-run drift detection.

4. A unit test for _ensure_cpu_torch did not pin NO_TORCH False like its
   siblings, so a suite run with UNSLOTH_NO_TORCH=1 in the environment made the
   guard return early and the reinstall assertions fail spuriously.

Tests: the round-6 broken-torch verbatim test re-encodes the non-None
"torch==absent" snapshot case (the exact state the old "is None" check missed);
new Windows-ROCm pinned-repair cases (reinstall from the pin, per-arch floor vs
bare spec, matching-wheel no-op, off-Windows no-op); a NO_TORCH fast-path probe
case; the parity test now asserts the Windows final branch does not auto-detect
the ROCm index and that the helper reinstalls from the explicit pin. 494 passed.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: floor the rocm7.2 index in the Windows pin repair; isolate marker tests

Three round-8 review items, two of them downstream of the round-7 changes:

1. _ensure_pinned_known_family_torch gave a rocm<d> index leaf a bare
   torch/torchvision/torchaudio trio while flooring only gfx* leaves, so a
   Windows venv clobbered under an explicit rocm7.2 pin could reinstall an
   unbounded or ABI-mismatched trio from that exclusive --index-url. It now
   mirrors the spec the initial ROCm paths pin: the rocm7.2 floor for 2.11-line
   gfx leaves and rocm<d> leaves that serve torch 2.11, the <2.11 default for
   older rocm versions, and a bare trio only for older gfx per-arch leaves
   (which publish no floor), matching _ROCM_TORCH_PKG_SPECS / _ensure_rocm_torch.

2. test_verbatim_custom_url_no_marker_reinstalls_once called
   _ensure_verbatim_torch_index twice; the second call now hits the
   matching-marker health probe, and with pip_install mocked torch never becomes
   importable, so in a no-torch environment _probe_torch_flavor returned None and
   forced another reinstall, failing the idempotence assertion. The test now pins
   a healthy flavor so the idempotence check is about the marker, not ambient
   torch.

3. The TestEnsureRocmTorchMarker fixture patched os.environ per test but not
   _TORCH_BACKEND, which install_python_stack.py computes once at import from
   UNSLOTH_TORCH_BACKEND. A runner starting with a cuda/cpu backend made
   _ensure_rocm_torch early-return and skip the mocked repair these tests
   exercise. The fixture now neutralizes _TORCH_BACKEND so the marker tests are
   independent of the caller's installer-pin environment.

Tests: the Windows floor-spec test now asserts a rocm7.2 mirror pin uses the
rocm7.2 floor (not bare), plus a new rocm7.1 case that must fall back to the
<2.11 default; the marker suite passes under a hostile
UNSLOTH_TORCH_BACKEND=cuda / UNSLOTH_TORCH_INDEX_URL env. 495 passed.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: apply same-flavor pin repoints, keep ROCm fallback nonfatal, bound custom companions

Four round-9 review items, two of them regressions in the round-7 pin helper:

1. _ensure_pinned_known_family_torch returned as satisfied whenever the installed
   flavor matched the pin, so a same-flavor SOURCE change (one /cpu or /cu128
   mirror to another, or a gfx1151 -> gfx120x-all per-arch switch, both carrying
   the same wheel tag) was never applied, while _torch_pin_needs_apply kept forcing
   the pass on the marker mismatch forever. It now also reinstalls when the marker
   records a DIFFERENT index of the same flavor, rewriting the marker so the next
   update matches (no loop), exactly as the Linux _ensure_{cuda,cpu}_torch helpers
   do. An absent marker on an already-matching venv is still left to the baseline
   recorder (no forced reinstall of a correct pre-marker venv).

2. That helper reinstalled a Windows ROCm pin with the FATAL pip_install, so when
   setup.ps1 had taken its CPU fallback (the pinned AMD index unavailable), the
   final repair re-hit the same missing index and aborted the whole install. The
   ROCm reinstall is now nonfatal (pip_install_try): on failure it leaves the CPU
   base in place and writes no ROCm marker, so the install completes -- matching
   _ensure_rocm_torch's Windows path. cu*/cpu pins stay fatal (authoritative source).

3. install.sh left torchvision/torchaudio bare for a pinned custom/unknown-leaf
   index (a private /simple mirror), unlike the Python update path's
   _CUSTOM_INDEX_TORCH_PKG_SPEC, so a mirror also exposing newer companion wheels
   could resolve a torch-2.12-built torchvision against the capped <2.11 torch. It
   now bounds the companions (torchvision>=0.19,<0.26.0 / torchaudio>=2.4,<2.11.0)
   for a custom leaf, gated on an empty _expected_torch_flavor_tag so known families
   keep their curated bare/floored companions.

4. install.sh's _expected_torch_flavor_tag matched cu[0-9]* by prefix, so a custom
   leaf like cu128-private classified as the cu128 family and force-reinstalled a
   correct +cu128 wheel on every run. It now requires exact cu+digits (routing the
   suffixed leaf to the custom path), matching the Python re.fullmatch(cu[0-9]+) and
   PowerShell, and feeding item 3's custom-leaf detection.

Tests: new cases for the same-flavor marker-change reinstall, the nonfatal ROCm
fallback (no marker on failure), the rocm7.2/older-rocm floor selection now split
across the nonfatal path, cu-suffixed custom leaves in test_torch_flavor.sh, and the
custom-leaf companion bounds in test_torch_constraint.sh. 497 python + 143 shell
assertions pass; the marker suite still passes under a hostile
UNSLOTH_TORCH_BACKEND=cuda env.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: bound custom-pin companions on the Windows setup path; isolate pin-probe tests

Two round-10 review items:

1. setup.ps1's custom/unknown-leaf pin branch capped only torch ($cudaTorchSpec)
   and still asked the exclusive index for bare torchvision/torchaudio, so a
   private mirror that also serves newer companion wheels could install a
   torch<2.11 wheel alongside a torchvision>=0.26 / torchaudio>=2.11 built for a
   newer torch ABI, after which the marker records the pin as applied. It now
   bounds the whole trio (torch>=2.4,<2.11.0 / torchvision>=0.19,<0.26.0 /
   torchaudio>=2.4,<2.11.0) for a pinned non-cu-family leaf, matching install.sh,
   install.ps1's fresh pinned install, and install_python_stack.py's
   _CUSTOM_INDEX_TORCH_PKG_SPEC. This completes the companion-bounds fix across all
   three installers; known cu* leaves keep bare specs (the family index bounds them).

2. The _torch_pin_needs_apply probe tests did not pin NO_TORCH False, so a test
   process launched with UNSLOTH_NO_TORCH=1 short-circuited the probe (the round-7
   guard) and returned False for cases that expect the pass to run. The _needs_apply
   helper now patches NO_TORCH (default False) around the call, and the dedicated
   no-torch case passes no_torch=True explicitly.

Tests: the cross-platform parity test now asserts setup.ps1 bounds the full trio
(not just torch) for a custom leaf; the pin-probe suite passes under a hostile
UNSLOTH_NO_TORCH=1 environment. setup.ps1 parses clean; 497 python + shell suites
green.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: bound custom rocm-* pins, redact diag tokens, snapshot custom pins before base update

Three round-11 review items, all reproduced before fixing:

1. install.sh's custom-index companion bounds gated on _expected_torch_flavor_tag
   returning empty, but that helper returned "rocm" for ANY rocm* leaf, so a custom
   mirror whose leaf starts with rocm but is not a pip family (a private rocm-current
   mirror, a Radeon find-links rocm-rel-7.2.1) escaped the bounds and installed bare
   torchvision/torchaudio. It now digit-gates rocm to rocm[0-9]* (matching the Python
   _is_pip_rocm_family_leaf ^rocm\d), so those custom leaves return "" and the <2.11
   companion caps apply; real rocm7.2 / gfx per-arch indexes still classify as rocm.

2. _tauri_torch_index_family classified by the raw last path segment, so a pinned URL
   carrying auth in the query (.../rocm7.2?token=SECRET) had the token echoed verbatim
   into the emitted [TAURI:DIAG] line. It now strips query/fragment before classifying
   (mirroring the marker/log credential stripping), so no token reaches the diagnostic
   output; as a side effect .../cu128?token=x now classifies as cu128 instead of auto.

3. On studio update, the core package step (a newer unsloth can require a torch the
   custom pin does not satisfy, pulling a default PyPI trio) runs BEFORE the step-2b
   verbatim check, which then recorded the already-clobbered trio as the baseline for a
   matching marker and left the pin unapplied. A new _capture_verbatim_baseline() records
   the pre-clobber trio before the core step, so the verbatim pass detects the drift and
   reapplies the pin. Captures only for a matching custom pin with importable torch; a
   mismatched/absent marker or broken torch is left to _ensure_verbatim_torch_index.

Tests: _expected_torch_flavor_tag rocm-current / rocm-rel cases; _tauri_torch_index_family
token/fragment redaction with a no-leak regression guard; _capture_verbatim_baseline
record/skip cases plus an end-to-end clobber-detection scenario; a structural guard that
the capture runs before the core step. 501 python + shell suites pass; install.sh bash -n
clean, shellcheck unchanged from base.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: match rocm family leaves exactly, enforce the rocm7.2 torch line, repair a broken pinned torch

A pinned index is a pip ROCm --index-url family only when its leaf is an exact
rocm<digits> / rocm<digits>.<digits> (rocm7.2) or a gfx* per-arch leaf. The prior
^rocm[0-9] prefix match also caught suffixed private-mirror leaves (rocm7.2-private,
rocm7-current), routing them through the ROCm/companion-family path instead of the
verbatim pin: the companion bounds were skipped and, on a pre-marker venv with a
compatible +rocm wheel, the pin was never applied. Match the family exactly through one
shared helper at every site:
  - install_python_stack.py: _is_pip_rocm_family_leaf (re.fullmatch), plus the two other
    loose gates it feeds (_normalize_family_leaf, _torch_flavor_matches_pin).
  - install.sh: a new _is_pip_rocm_family_leaf routes _expected_torch_flavor_tag,
    _torch_index_repairable, _normalize_family_leaf and the ROCm side-effect gate.
  - setup.ps1: a new Test-PipRocmFamilyLeaf routes Get-NormalizedFamilyLeaf and both
    pinned reroutes; install.ps1 anchors its reroute regex.

_rocm_pin_family_mismatch (and its setup.ps1 mirror Get-RocmPinStaleTags) compared only
the ROCm version, so a +rocm7.2 wheel whose torch release drifted off the 2.11 line
(2.12/2.13 from an out-of-band upgrade or a custom rocm7.2 mirror) satisfied the family
check while violating _ROCM_TORCH_PKG_SPECS['rocm7.2'] (torch>=2.11,<2.12). Flag it stale
so the repair reinstalls to floor; >=2.11 alone is not enough, so the release is compared
exactly against the 2.11 line for a KNOWN-2.11 rocm pin.

_ensure_pinned_known_family_torch returned on a failed import probe, but
_torch_pin_needs_apply forces the dependency pass on that same failed probe: a broken
torch under a known-family pin was left in place and the pass was forced on every update.
Treat an unimportable torch as drift and reinstall the pinned trio (the spec and marker
derive from the pinned leaf, not the absent flavor); once it lands the probe succeeds and
the fast path returns.

Tests: exact-match cases across test_torch_flavor.sh, test_rocm_support.py,
test_cross_platform_parity.py and the two .ps1 helper suites; the rocm7.2 release-line
and broken-probe-reinstall cases; extraction lists updated for the new helpers.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: anchor the PS pinned-ROCm floor gate and bound install.ps1 custom-pin companions

Round 12 made every family CLASSIFIER exact, but the Windows install-flow floor gate reads
$_pinRocm211 directly from the raw pinned leaf with an unanchored -match '^rocm(\d+)\.(\d+)'
BEFORE any exact classification runs. A suffixed custom leaf (rocm7.2-private) matches that
rocm7.2 prefix, so it takes the 2.11-floor branch and is force-routed through the ROCm
install path before the exact-match elseif can send it to the verbatim install. Anchor the
match ($) in both install.ps1 and setup.ps1 so only an exact rocmX.Y leaf is floored; a
suffixed or newer-suffix leaf falls through to the verbatim path. The Python floor
selection is already exact (dict lookups gated on _is_pip_rocm_family_leaf), so only the two
PS scripts needed this.

install.ps1's custom (non-cu-family) pinned-torch install bounded torch>=2.4,<2.11.0 but
left torchvision/torchaudio bare, so a private mirror serving newer companions could pull a
wheel built for a newer torch ABI while the marker records the pin as applied. Bound both
companions (torchvision>=0.19,<0.26.0 / torchaudio>=2.4,<2.11.0) when the leaf is not a
cu<digits> family index (a cu index bounds its own resolution), matching setup.ps1's
Test-CudaFamilyLeaf gate and _CUSTOM_INDEX_TORCH_PKG_SPEC.

Tests: parity guards for the anchored floor gate in both PS scripts and for install.ps1's
bounded custom-pin companions.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: tighten comments in the torch-index-override paths

Collapse the verbose comment and docstring blocks added across the installer
scripts and their tests to fewer, clearer lines without changing behaviour.
Remove a duplicated CUDA-spec comment block. Comments/docstrings only; no code
changes (AST-verified).

* install: repair a broken pinned torch on Linux, strip trailing slash in tauri family, count the final step

_ensure_cuda_torch / _ensure_cpu_torch returned on a failed import probe (torch present but
unimportable). With an explicit CUDA/CPU pin, _torch_pin_needs_apply forces the dependency
pass on that same failed probe, and the base package update does not force-reinstall an
already-installed torch distribution, so the broken torch was left in place and the pass
reran every update without repairing it. Treat a failed probe under a pin as drift and
reinstall from the pinned index (the reinstall rewrites the marker and the next probe
imports, so no loop). This is the Linux counterpart of the known-family repair fix.

_tauri_torch_index_family stripped the query/fragment before classifying but not a trailing
slash, so a token-authenticated pin like .../cu128/?token=x collapsed to .../cu128/ and fell
through the exact-suffix */cu128 and */cpu arms to "auto". Strip a trailing slash too,
mirroring _torch_index_url_leaf.

The Windows / macOS-ARM final torch-repair step (_ensure_pinned_known_family_torch) runs a
progress step that base_total never counted (the final-step increment was gated to Linux),
so _STEP ran one past _TOTAL on those platforms. Add the missing increment.

Tests: broken-probe reinstall for the CUDA (family and URL pins) and CPU paths; trailing
slash / slash+token cases for _tauri_torch_index_family; a full-flow progress-count guard
asserting _STEP == _TOTAL on Windows and Linux.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: tighten comments in the torch-index-override paths

* install: harden the torch-index pin across all four installers

Redact index-URL credentials from captured install logs before they print on
failure. uv/pip failure text embeds the failing --index-url verbatim, so a
user:token@ or ?token= secret could leak into the console. Add a shared
redaction pass (_redact_install_output / Redact-InstallOutput) wired into the
error-output dump in install.sh, install.ps1, setup.ps1 and
install_python_stack.py. Verbose mode still streams live uncaptured output, so
it is intentionally left unredacted (developer opt-in).

Trim trailing slashes on the PATH only for a verbatim UNSLOTH_TORCH_INDEX_URL
override, preserving a ?query/#fragment token. A whole-URL rstrip corrupted a
base64 token ending in "/", and a single-slash strip left .../cu128//
classifying as an empty leaf. Add _trim_index_path_slashes /
Trim-IndexPathSlashes and route the override through it; strip ALL trailing
slashes in the backend-branding leaf classifier so a double slash still yields
the real leaf.

Reject a trailing-dot ROCm leaf (rocm7.) in the bash family validator so it
matches Python re.fullmatch(rocm\d+(?:\.\d+)?) and the PowerShell regex: both the
major and the minor must be non-empty digits, so rocm7. is a custom verbatim pin,
not a pip ROCm family.

Scrub PIP_NO_INDEX and PIP_INDEX_URL for a pinned install in the two installers
that have a plain-pip fallback (install_python_stack.py, setup.ps1):
PIP_NO_INDEX=1 makes the fallback ignore every index including the pinned
--index-url, and PIP_INDEX_URL replaces it. install.sh and install.ps1 install
via uv --default-index (which ignores pip config/env), so they are unaffected.

Add unit tests (bash, Python, PowerShell) and cross-platform parity tests
covering credential redaction, path-only slash trimming, the rocm7. validator,
the double-slash leaf, and the PIP_NO_INDEX/PIP_INDEX_URL scrub.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: redact captured torch-install output and warn on a failed pinned ROCm repair

Close a redaction gap the earlier pass missed: setup.ps1's direct
`Fast-Install ... | Out-String` branches (ROCm from $ROCmIndexUrl, CPU/CUDA from
$TorchInstallIndexUrl, plus the Triton and T5 sub-venv installs) printed the
captured $output verbatim on failure, bypassing Redact-InstallOutput. A private
index carrying userinfo or a ?token= in the pin could leak into Windows Studio
setup logs. Route every `Write-Host $output` through Redact-InstallOutput.

Warn on a failed pinned Windows ROCm reinstall in
_ensure_pinned_known_family_torch: the branch printed "reinstalling from it" then
called pip_install_try, but had no else, so a failure continued silently and left
the user believing the pin was applied while the old CPU/wrong torch survived.
Mirror the auto-ROCm Windows path and warn, telling the user to retry.

* install: redact captured output on the pip fallback and optional-install failure paths

The uv install path already redacted its captured output, but pip_install's pip
fallback runs through run(), which printed result.stdout verbatim on failure, and
_print_optional_install_failure did the same. A pinned --index-url carrying
userinfo or a ?token= could still leak there when uv is unavailable or the pip
fallback also fails. Route both through _redact_install_output. The verbose
pip_install_try path stays raw (developer opt-in), matching the other installers.

* install: split the survive-updates marker subsystem into a follow-up

The torch-index override PR grew a persisted per-venv marker plus repair
machinery (stale-pin detection, verbatim re-apply, update-time reinstall
triggers) that roughly doubled it. That subsystem is orthogonal to the core
feature and is being reworked in a follow-up (versioned/hashed marker,
full-URL pin baseline), so it moves there wholesale instead of shipping
twice.

What this PR still does: UNSLOTH_TORCH_INDEX_URL / UNSLOTH_TORCH_INDEX_FAMILY
pick the torch wheel index at install time in all four installers, with the
exact rocm/gfx/cpu/cu leaf classification, the torch 2.11 floor for the
per-arch AMD indexes, bounded companions for custom leaves, credential
redaction of captured installer output, path-only slash trimming, and the
uv/pip index env scrubs. Flavor-based repair keeps honoring the pin: a wrong
family under an explicit pin still reinstalls from the pinned URL, and
setup.ps1 repairs a pinned stale venv in place instead of wiping it.

What moves to the follow-up: the .unsloth-torch-index marker file and its
writers/readers/normalizers, exact-URL pin-change detection on update
(same-tag gfx switches, custom-mirror repoints), the verbatim trio snapshot
and clobber re-apply, the pin-baseline recorder, and the
--torch-pin-needs-apply fast-path probe in setup.sh / setup.ps1. Their tests
(the marker sh/ps1 suites, the stale-pin suite, and the marker classes in the
rocm/cuda/parity suites) move with them; the removed code is preserved on a
local archive branch to seed that PR.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: re-apply a ROCm pin over an existing HIP wheel via the version tag

The subsystem split left an explicit ROCm/gfx pin unenforced on `studio
update` whenever the venv already imported ANY ROCm torch: the pinned
reinstall lived inside the `elif not has_hip_torch` branch, so a rocm6.4 to
rocm7.2 switch, a gfx1151 pin over a generic +rocm7.2 wheel, or a broken
2.12+rocm7.2 drift never re-applied the pin.

Restore the markerless half of that detection: _rocm_pin_family_mismatch
compares the pinned leaf against the installed wheel tag (exact rocmX.Y
compare, the 2.11 gfx per-arch allowlist, the untagged-wheel rule), the HIP
probe emits "<hip_marker>|<version>" again so the installed tag is available,
and _ensure_rocm_torch reinstalls from the pinned URL when the tag mismatches
even though HIP torch is present. setup.ps1 mirrors it: the stale-venv check
routes a pinned rocm/gfx leaf through Get-RocmPinStaleTags instead of
collapsing it to a generic "rocm" flavor, and the existing pinned in-place
repair (no wipe) applies the change.

What still waits for the follow-up marker PR, by design: pin changes the
wheel tag cannot see -- a per-arch switch between two 2.11 gfx indexes
(identical +rocm7.13.0 tag), a custom-mirror URL repoint under the same
family leaf, and unknown-family verbatim pins. Those need the persisted
index record.

Tests restored with the code: the _rocm_pin_family_mismatch table, the five
update-path cases (older-rocm reinstall, gfx-over-pre-2.11 reinstall,
matching-pin no-reinstall, non-2.11 gfx no-reinstall, gfx-over-generic-2.11
reinstall), the "|" probe-format guards, and the AST-extracted
Get-RocmPinStaleTags suite for setup.ps1.

* install: compare major-only rocm pins, redact URL fragments, bound pinned CPU trio

Three review fixes on the restored pin-repair path.

The family classifier accepts a major-only rocm<d> leaf (rocm7), but the
mismatch comparators only parsed rocmX.Y, so a rocm7 pin fell through to the
2.11-line fallback and INVERTED both verdicts: an installed +rocm6.4 wheel
compared as satisfied (pin never re-applied) while a matching +rocm7.2 wheel
compared as stale (reinstall loop). Major-only pins now compare on the major
alone in _rocm_pin_family_mismatch and Get-RocmPinStaleTags: rocm6.x under a
rocm7 pin is a mismatch, any rocm7.x satisfies it, an untagged wheel never
does, and a bare +rocm tag with an unreadable version is accepted (matching
the existing lenient unreadable fallback).

The output redactors scrubbed userinfo and ?query= values but not #fragments,
so a pin like https://mirror/whl/cu128#token=secret leaked the secret in
captured uv/pip failure text -- inconsistent with the URL handling itself,
which already treats fragments as sensitive. All four redactors gain a
URL-anchored fragment rule (anchored so a bare "# comment" line in tool
output is never touched).

setup.ps1's CPU branch installed a bare torch/torchvision/torchaudio trio;
fine for the unpinned host default, but a PINNED cpu index routes through the
same branch and the /cpu index serves newer torch, so a fresh pinned CPU
install could land an unsupported trio that _ensure_cpu_torch then keeps
(it accepts any CPU build). Under a pin the branch now installs the bounded
trio mirroring _CPU_TORCH_PKG_SPEC (torch>=2.4,<2.12.0 and matching
companions); the unpinned path is unchanged.

Tests: major-only rows in the Python mismatch table and the AST-extracted
setup.ps1 suite; fragment + query-plus-fragment + bare-hash-comment cases in
all four redactor suites; a parity check that the pinned CPU trio bounds
exist, are gated on the pin, and mirror the Python repair spec.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install: tighten comments in the torch index override paths

* tests: track the moved pass-through inheritance in the gguf order check

Main moved the llama_extra_args pass-through inheritance out of the
GGUF branch into _resolve_inherited_extra_args, which runs before it,
so the source-order assertion's "if request.llama_extra_args is None"
anchor no longer exists inside the branch and the check failed after
the main merge. The test now asserts the same property in the current
shape: inheritance before the GGUF branch (a carried --no-mmproj still
shapes the hub guard's companion requirement), and marker, hub guard,
unload in order within the branch. Full file passes (32 tests).

* tests: anchor the inheritance order check on the call, not the definition

source.index("_resolve_inherited_extra_args(") matched the function
definition, which always precedes the endpoint, so the ordering
assertion was vacuously true. Anchoring on "= _resolve_inherited_
extra_args(" pins the first call site inside the load endpoint (line
4505), which is the statement whose position relative to the GGUF
branch the test is meant to guard. 32 tests pass.

* tests: align the gguf order test with main

Main fixed the stale ordering assertion in PR 7252; adopting its
version verbatim removes this file from the branch diff entirely and
avoids a conflict on the next main merge. 32 tests pass.

* install: bound the companion constraints to torch's window everywhere

A full platform x vendor validation matrix over this branch surfaced a
real trio mismatch on the cpu/mac paths: torch is capped <2.11 (installs
2.10.0+cpu) but the bare torchaudio companion resolves 2.11.0+cpu,
because torchaudio 2.11 dropped its exact torch pin. Reproduced in a
sandboxed end to end cpu install. torchvision still exact-pins torch and
self-corrected.

The default companion constraints are now bounded to torch's window
(<0.26 / <2.11) and widen together with the cu* torch window (<0.27 /
<2.12), so every leaf resolves a paired trio. Verified with uv dry-runs
on the cpu, cu130, and rocm6.4 leaves (2.10.0/0.25.0/2.10.0,
2.11.0/0.26.0/2.11.0, 2.9.1/0.24.1/2.9.1) and a rerun of the sandboxed
cpu install, which now lands torch 2.10.0+cpu with torchaudio
2.10.0+cpu.

The Strix WSL reroute now also forwards UNSLOTH_TORCH_INDEX_URL and
UNSLOTH_TORCH_INDEX_FAMILY into the rerouted 24.04 distro; dropping
them silently reverted the child install to auto-detection, defeating
the pin this branch introduces.

test_torch_constraint.sh updated: the bounded companions must appear at
the defaults and the custom-leaf block, no bare companion may remain,
and the cu* widen must carry the companions with it.

* install: harden the override path against reroute drift and credential leaks

Review sweep focused on default-path idempotency found no defects on the
unset path; these fixes cover the override path and failure reporting.

install.sh:
- The early WSL Strix Halo distro reroute now honors an explicit index
  pin (UNSLOTH_TORCH_INDEX_URL / _FAMILY): the pin is used in the current
  distro instead of probing the GPU and re-entering another distribution,
  matching the contract of the later Radeon and Strix guards. Whitespace
  only values do not gate, in parity with get_torch_index_url.
- Verbose mode now streams installer output through the credential
  redactor; it previously bypassed the redaction the quiet path applies.
  The exit code survives the pipe via an rc file since the script runs
  under plain sh with no pipefail.
- The kept-release fallback warning now strips credentials from the
  index URL before printing it.

install.ps1:
- Bounded torchvision and torchaudio next to every capped torch install
  (custom pin, ROCm CPU fallback, CUDA flavor repair). torchaudio 2.11
  dropped its exact torch pin from the wheel metadata, so a bare
  companion beside torch<2.11 can resolve a mismatched 2.11.0 build,
  cu family indexes included. Mirrors the install.sh companion bounds.

studio/install_python_stack.py:
- The verbose failure path now redacts index URLs in pip and uv output
  before printing, matching every other output site in the file.

All sh, ps1 and python installer test suites pass (the host-defaults
suite has a known pre-existing failure unrelated to this change).

* install: redact verbose Windows installer output and repair the parity tests

Follow-ups to the override-hardening commit, from review:

- install.ps1 Invoke-InstallCommand and setup.ps1 Invoke-SetupCommand now
  pipe verbose output through Redact-InstallOutput per record, and the
  three verbose Fast-Install torch call sites (ROCm, CPU, CUDA) do the
  same: uv and pip echo the pinned index URL, credentials included, in
  their errors, and verbose mode previously bypassed the redaction the
  quiet paths apply. ForEach-Object and Out-Host leave $LASTEXITCODE
  untouched, verified with a native command exiting 7 behind the pipe.

- test_cross_platform_parity.py: the install.ps1 companion-bounds
  assertion now matches the implemented behavior (bounds on every index,
  no cu-family exemption, since torchaudio 2.11 dropped its exact torch
  pin) instead of requiring the removed $_pinCuLeaf gate.

- test_rocm_support.py: the WSL reroute guard test slices the whole
  function body to its closing brace instead of a fixed 1200-character
  window, which the new pin-gate preamble had outgrown.

428 tests pass across the parity, install stack and rocm support suites;
the sh and ps1 installer suites pass unchanged.

* install: tighten comments in the torch-index and ROCm/CUDA repair paths

* install: digit-gate the gfx family leaf and honor ROCm pins in the Windows repair

Two review follow-ups on the override path:

- The pip ROCm family predicate accepted ANY gfx-prefixed leaf, so a
  custom verbatim pin like /gfx-private classified as a ROCm family and
  enabled the ROCm-only side effects (AMD bitsandbytes, ROCm torch
  repair) on a mirror that may serve CPU/CUDA wheels. gfx now requires a
  following digit (gfx90a, gfx1151, gfx120X-all), consistently in
  install.sh, install_python_stack.py, install.ps1 (family gate and
  expected-flavor classifier) and setup.ps1, matching the strictness the
  rocm side already had (rocm7.2-private stays verbatim). The broader
  backend BRANDING globs are unchanged on purpose: radeon repo leaves
  (rocm-rel-X.Y) must still brand the rocm backend without being
  force-repaired as a family.

- The Windows branch of the ROCm torch repair always installed from the
  public per-arch index, ignoring an explicit ROCm-family pin: after a
  pinned setup.ps1 install failed to a CPU base, the repair retried
  repo.amd.com instead of the pinned index. The branch now resolves
  _explicit_rocm_torch_index_url() first, uses it as the install index
  when set, and mirrors the Linux pin contract by skipping the NVIDIA
  and gfx-detection gates a pin is documented to override.

Source-assertion tests updated to the tightened predicate and the new
repair label. 1165 tests pass across the parity, install stack and
studio install suites; the sh and ps1 suites pass; both PowerShell
installers parse clean.

* Remove scratch archives accidentally committed with the comment pass

The temp/ archive copies of installer and test files were working
scratch, not PR content, and inflated the diff by about nine thousand
lines.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-07-20 00:58:52 -07:00

4435 lines
218 KiB
PowerShell

#Requires -Version 5.1
# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
<#
.SYNOPSIS
Full environment setup for Unsloth Studio on Windows (bundled version).
.DESCRIPTION
Uses an isolated, Unsloth-managed Node.js for the frontend build when the
system Node/npm do not meet requirements (never modifies the system Node).
When running from pip install: skips frontend build (already bundled). When
running from git repo: full setup including frontend build.
Supports NVIDIA GPU (full training + inference) and CPU-only (GGUF chat mode).
.NOTES
Default output is minimal (step/substep), aligned with studio/setup.sh.
FULL / LEGACY LOGGING (defensible audit trail, detailed multi-line output):
unsloth studio setup --verbose
Or: $env:UNSLOTH_VERBOSE='1'; powershell -File .\studio\setup.ps1
Or: .\setup.ps1 --verbose
#>
$ErrorActionPreference = "Stop"
$ScriptDir = Split-Path -Parent $MyInvocation.MyCommand.Path
$PackageDir = Split-Path -Parent $ScriptDir
# --------------------------------------------------------------------------
# Maintainer-editable defaults
# Change these in the GitHub-hosted script so users get updated defaults.
# User env vars always override these baked-in values.
# --------------------------------------------------------------------------
# Prefer "latest" over "master" -- "master" bypasses the prebuilt resolver
# (no matching GitHub release), forces a source build, and causes HTTP 422
# errors. Only use "master" temporarily when the latest release is missing
# support for a new model architecture.
#
# UNSLOTH_LLAMA_CPP_BACKEND : "auto" (default) or "cpu". When "cpu", forces
# the CPU-only prebuilt bundle on GPU hosts. Fixes Intel iGPU Vulkan
# crashes (#7213).
$DefaultLlamaPrForce = ""
$DefaultLlamaSource = "https://github.com/ggml-org/llama.cpp"
$DefaultLlamaTag = "latest"
$DefaultLlamaForceCompileRef = "master"
# Corporate-mirror / proxy escape hatch for the frontend npm/bun install (#6491).
# studio/frontend/.npmrc pins registry=https://registry.npmjs.org/ as a supply-chain
# lock, which overrides a corporate user's ~/.npmrc proxy and causes 403s behind a
# firewall. UNSLOTH_NPM_REGISTRY is a deliberate opt-in: when set we splat it as
# `--registry <url>` into every npm/bun install. `--registry` is the highest-precedence
# override for BOTH tools and leaves min-release-age / save-exact in force. Empty array
# (the default) splats to nothing, so normal installs are unchanged.
$NpmRegistryArgs = @()
if ($env:UNSLOTH_NPM_REGISTRY) {
$NpmRegistryArgs = @('--registry', $env:UNSLOTH_NPM_REGISTRY)
}
# Verbose can be enabled either by CLI flag or by UNSLOTH_VERBOSE=1.
$script:UnslothVerbose = ($env:UNSLOTH_VERBOSE -eq '1')
foreach ($a in $args) {
if ($a -eq '--verbose' -or $a -eq '-v') {
$script:UnslothVerbose = $true
break
}
}
# Propagate to child processes (e.g. install_python_stack.py) so they
# also respect verbose mode. Process-scoped -- does not persist.
if ($script:UnslothVerbose) {
$env:UNSLOTH_VERBOSE = '1'
}
$script:LlamaCppDegraded = $false
# CUDA toolkit state, published by Resolve-CudaToolkit. Only the Phase 4 source
# build consumes these; the prebuilt path leaves them at these defaults.
$script:CudaToolkitReady = $false
$script:NvccPath = $null
$script:CudaToolkitRoot = $null
$script:CudaArch = $null
# Detect if running from pip install (no frontend/ dir in studio)
$FrontendDir = Join-Path $ScriptDir "frontend"
$OxcValidatorDir = Join-Path $ScriptDir "backend\core\data_recipe\oxc-validator"
$IsPipInstall = -not (Test-Path $FrontendDir)
# ─────────────────────────────────────────────
# Helper functions
# ─────────────────────────────────────────────
# Reload ALL environment variables from registry.
# Picks up changes made by installers (winget, msi, etc.) including
# Path, CUDA_PATH, CUDA_PATH_V*, and any other vars they set.
function Refresh-Environment {
foreach ($level in @('Machine', 'User')) {
$vars = [System.Environment]::GetEnvironmentVariables($level)
foreach ($key in $vars.Keys) {
if ($key -eq 'Path') { continue }
Set-Item -Path "Env:$key" -Value $vars[$key] -ErrorAction SilentlyContinue
}
}
$machinePath = [System.Environment]::GetEnvironmentVariable('Path', 'Machine')
$userPath = [System.Environment]::GetEnvironmentVariable('Path', 'User')
# Merge: venv Scripts (if active) > Machine > User > current $env:Path. Dedup raw+expanded.
$venvScripts = if ($env:VIRTUAL_ENV) { Join-Path $env:VIRTUAL_ENV 'Scripts' } else { $null }
$sources = @()
if ($venvScripts) { $sources += $venvScripts }
$sources += @($machinePath, $userPath, $env:Path)
$merged = ($sources | Where-Object { $_ }) -join ';'
$seen = @{}
$unique = New-Object System.Collections.Generic.List[string]
foreach ($p in $merged -split ";") {
$rawKey = $p.Trim().Trim('"').TrimEnd("\").ToLowerInvariant()
$expKey = [Environment]::ExpandEnvironmentVariables($p).Trim().Trim('"').TrimEnd("\").ToLowerInvariant()
if ($rawKey -and -not $seen.ContainsKey($rawKey) -and -not $seen.ContainsKey($expKey)) {
$seen[$rawKey] = $true
if ($expKey -and $expKey -ne $rawKey) { $seen[$expKey] = $true }
$unique.Add($p)
}
}
$env:Path = $unique -join ";"
}
# ── Helper: safely add a directory to the persistent User PATH ──
# Direct registry access preserves REG_EXPAND_SZ (avoids dotnet/runtime#1442).
# Append (default) keeps existing tools first; Prepend for must-win entries.
function Add-ToUserPath {
param(
[Parameter(Mandatory = $true)][string]$Directory,
[ValidateSet('Append','Prepend')]
[string]$Position = 'Append'
)
try {
$regKey = [Microsoft.Win32.Registry]::CurrentUser.CreateSubKey('Environment')
try {
$rawPath = $regKey.GetValue('Path', '', [Microsoft.Win32.RegistryValueOptions]::DoNotExpandEnvironmentNames)
[string[]]$entries = if ($rawPath) { $rawPath -split ';' } else { @() } # string[] prevents scalar collapse
$normalDir = $Directory.Trim().Trim('"').TrimEnd('\').ToLowerInvariant()
$expNormalDir = [Environment]::ExpandEnvironmentVariables($Directory).Trim().Trim('"').TrimEnd('\').ToLowerInvariant()
$kept = New-Object System.Collections.Generic.List[string]
$matchIndices = New-Object System.Collections.Generic.List[int]
for ($i = 0; $i -lt $entries.Count; $i++) {
$stripped = $entries[$i].Trim().Trim('"')
$rawNorm = $stripped.TrimEnd('\').ToLowerInvariant()
$expNorm = [Environment]::ExpandEnvironmentVariables($stripped).TrimEnd('\').ToLowerInvariant()
$isMatch = ($rawNorm -and ($rawNorm -eq $normalDir -or $rawNorm -eq $expNormalDir)) -or
($expNorm -and ($expNorm -eq $normalDir -or $expNorm -eq $expNormalDir))
if ($isMatch) {
$matchIndices.Add($i)
continue
}
$kept.Add($entries[$i])
}
$alreadyPresent = $matchIndices.Count -gt 0
if ($alreadyPresent -and $Position -eq 'Append') { # Append: idempotent no-op
return $false
}
if ($alreadyPresent -and $Position -eq 'Prepend' -and # Prepend: no-op if already at front
$matchIndices.Count -eq 1 -and $matchIndices[0] -eq 0) {
return $false
}
# One-time backup under HKCU\Software\Unsloth\PathBackup
if ($rawPath) {
try {
$backupKey = [Microsoft.Win32.Registry]::CurrentUser.CreateSubKey('Software\Unsloth')
try {
$existingBackup = $backupKey.GetValue('PathBackup', $null)
if (-not $existingBackup) {
$backupKey.SetValue('PathBackup', $rawPath, [Microsoft.Win32.RegistryValueKind]::ExpandString)
}
} finally {
$backupKey.Close()
}
} catch { }
}
if (-not $rawPath) {
Write-Host "[WARN] User PATH is empty - initializing with $Directory" -ForegroundColor Yellow
}
$newPath = if ($rawPath) {
if ($Position -eq 'Prepend') {
(@($Directory) + $kept) -join ';'
} else {
($kept + @($Directory)) -join ';'
}
} else {
$Directory
}
if ($newPath -ceq $rawPath) { # no actual change
return $false
}
$regKey.SetValue('Path', $newPath, [Microsoft.Win32.RegistryValueKind]::ExpandString)
# Broadcast WM_SETTINGCHANGE via dummy env-var roundtrip.
# [NullString]::Value avoids PS 7.5+/.NET 9 $null-to-"" coercion.
try {
$d = "UnslothPathRefresh_$([guid]::NewGuid().ToString('N').Substring(0,8))"
[Environment]::SetEnvironmentVariable($d, '1', 'User')
[Environment]::SetEnvironmentVariable($d, [NullString]::Value, 'User')
} catch { }
return $true
} finally {
$regKey.Close()
}
} catch {
Write-Host "[WARN] Could not update User PATH: $($_.Exception.Message)" -ForegroundColor Yellow
return $false
}
}
# PowerShell 5.1 compatibility helper: avoid relying on New-TemporaryFile.
function New-UnslothTemporaryFile {
$tempPath = [System.IO.Path]::GetTempFileName()
return Get-Item -LiteralPath $tempPath
}
function Remove-AgentInstructionFiles {
param([string[]]$Roots)
foreach ($root in $Roots) {
if (-not $root) { continue }
$item = Get-Item -LiteralPath $root -Force -ErrorAction SilentlyContinue
if (-not $item -or -not $item.PSIsContainer) { continue }
if ($item.Attributes -band [System.IO.FileAttributes]::ReparsePoint) { continue }
$pending = New-Object System.Collections.Stack
$pending.Push($item)
while ($pending.Count -gt 0) {
$current = $pending.Pop()
foreach ($child in @(Get-ChildItem -LiteralPath $current.FullName -Force -ErrorAction SilentlyContinue)) {
if ($child.PSIsContainer) {
if (-not ($child.Attributes -band [System.IO.FileAttributes]::ReparsePoint)) {
$pending.Push($child)
}
} elseif ($child.Name -in @("AGENTS.md", "CLAUDE.md")) {
Remove-Item -LiteralPath $child.FullName -Force -ErrorAction SilentlyContinue
}
}
}
}
}
function Get-InstalledLlamaPrebuiltRelease {
param([string]$InstallDir)
$metadataPath = Join-Path $InstallDir "UNSLOTH_PREBUILT_INFO.json"
if (-not (Test-Path $metadataPath)) {
return $null
}
try {
$payload = Get-Content $metadataPath -Raw | ConvertFrom-Json
} catch {
return $null
}
if (-not $payload.published_repo -or -not $payload.release_tag) {
return $null
}
$message = "installed release: $($payload.published_repo)@$($payload.release_tag)"
if ($payload.tag -and $payload.tag -ne $payload.release_tag) {
$message += " (tag $($payload.tag))"
}
return $message
}
# Find nvcc on PATH, CUDA_PATH, or standard toolkit dirs.
# Returns the path to nvcc.exe, or $null if not found.
function Find-Nvcc {
param([string]$MaxVersion = "")
$toolkitBase = 'C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA'
if ($MaxVersion -and (Test-Path $toolkitBase)) {
$drMajor = [int]$MaxVersion.Split('.')[0]
# Get all installed CUDA dirs, sorted descending (highest first)
$cudaDirs = Get-ChildItem -Directory $toolkitBase | Where-Object {
$_.Name -match '^v(\d+)\.(\d+)'
} | Sort-Object { [version]($_.Name -replace '^v','') } -Descending
foreach ($dir in $cudaDirs) {
if ($dir.Name -match '^v(\d+)\.(\d+)') {
$tkMajor = [int]$Matches[1]
$compatible = ($tkMajor -le $drMajor)
if ($compatible) {
$nvcc = Join-Path $dir.FullName 'bin\nvcc.exe'
if (Test-Path $nvcc) {
return $nvcc
}
}
}
}
# No compatible side-by-side version found
return $null
}
# Fallback: no version constraint — pick latest or whatever is available
# 1. Check nvcc on PATH
$cmd = Get-Command nvcc -ErrorAction SilentlyContinue
if ($cmd) { return $cmd.Source }
# 2. Check CUDA_PATH env var
$cudaRoot = [Environment]::GetEnvironmentVariable('CUDA_PATH', 'Process')
if (-not $cudaRoot) { $cudaRoot = [Environment]::GetEnvironmentVariable('CUDA_PATH', 'Machine') }
if (-not $cudaRoot) { $cudaRoot = [Environment]::GetEnvironmentVariable('CUDA_PATH', 'User') }
if ($cudaRoot -and (Test-Path (Join-Path $cudaRoot 'bin\nvcc.exe'))) {
return (Join-Path $cudaRoot 'bin\nvcc.exe')
}
# 3. Scan standard toolkit directory
if (Test-Path $toolkitBase) {
$latest = Get-ChildItem -Directory $toolkitBase | Where-Object {
$_.Name -match '^v(\d+)\.(\d+)'
} | Sort-Object { [version]($_.Name -replace '^v','') } -Descending | Select-Object -First 1
if ($latest -and (Test-Path (Join-Path $latest.FullName 'bin\nvcc.exe'))) {
return (Join-Path $latest.FullName 'bin\nvcc.exe')
}
}
return $null
}
function Write-CudaDriverToolkitMismatch {
param(
[Parameter(Mandatory = $true)][string]$ToolkitVersion,
[Parameter(Mandatory = $true)][string]$DriverMaxCuda,
[string]$Color = "Yellow"
)
$toolkitMajor = $ToolkitVersion.Split('.')[0]
$driverMajor = $DriverMaxCuda.Split('.')[0]
substep "CUDA Toolkit $ToolkitVersion is a major-version mismatch: toolkit major $toolkitMajor exceeds driver CUDA major $driverMajor ($DriverMaxCuda)." $Color
substep "Update the NVIDIA GPU driver to run CUDA Toolkit $ToolkitVersion, or install a CUDA $driverMajor.x toolkit." $Color
substep "Or let Unsloth use the prebuilt CUDA bundle; it does not need the local toolkit." $Color
}
# Detect CUDA Compute Capability via nvidia-smi.
# Returns e.g. "80" for A100 (8.0), "89" for RTX 4090 (8.9), etc.
# Returns $null if detection fails.
function Get-CudaComputeCapability {
# Use the resolved absolute path ($NvidiaSmiExe) to survive Refresh-Environment
$smiExe = if ($script:NvidiaSmiExe) { $script:NvidiaSmiExe } else {
$cmd = Get-Command nvidia-smi -ErrorAction SilentlyContinue
if ($cmd) { $cmd.Source } else { $null }
}
if (-not $smiExe) { return $null }
try {
# Bounded: a wedged nvidia-smi must not hang setup after the initial
# -L probe succeeded (the helper merges stderr after stdout, so the
# first line is still the compute_cap value).
$raw = Invoke-NvidiaSmiBounded $smiExe @('--query-gpu=compute_cap', '--format=csv,noheader')
if ($LASTEXITCODE -ne 0 -or -not $raw) { return $null }
# nvidia-smi may return multiple GPUs; take the first one
$cap = ($raw -split "`n")[0].Trim()
if ($cap -match '^(\d+)\.(\d+)$') {
$major = $Matches[1]
$minor = $Matches[2]
return "$major$minor"
}
} catch { }
return $null
}
# Check if an nvcc binary supports a given sm_ architecture.
# Uses `nvcc --list-gpu-code` which outputs sm_* tokens (--list-gpu-arch
# outputs compute_* tokens instead). Available since CUDA 11.6.
# Returns $false if the flag isn't supported (old toolkit) — safer to reject
# and fall back to scanning/PTX than to assume support and fail later.
function Test-NvccArchSupport {
param([string]$NvccExe, [string]$Arch)
try {
$listCode = & $NvccExe --list-gpu-code 2>&1 | Out-String
if ($LASTEXITCODE -ne 0) { return $false }
return ($listCode -match "sm_$Arch")
} catch {
return $false
}
}
# Given an nvcc binary, return the highest sm_ architecture it supports.
# Returns e.g. "90" for CUDA 12.4. Returns $null if detection fails.
function Get-NvccMaxArch {
param([string]$NvccExe)
try {
$listCode = & $NvccExe --list-gpu-code 2>&1 | Out-String
if ($LASTEXITCODE -ne 0) { return $null }
$arches = @()
foreach ($line in $listCode -split "`n") {
if ($line.Trim() -match '^sm_(\d+)') {
$arches += [int]$Matches[1]
}
}
if ($arches.Count -gt 0) {
return ($arches | Sort-Object | Select-Object -Last 1).ToString()
}
} catch { }
return $null
}
# Detect driver's max CUDA version from nvidia-smi and return the highest
# compatible PyTorch CUDA index tag (e.g. "cu128").
# PyTorch on Windows ships CPU-only by default from PyPI; CUDA wheels live at
# https://download.pytorch.org/whl/<tag>. The tag must not exceed the driver's
# capability: e.g. driver "CUDA Version: 12.9" → cu128 (not cu130).
function Get-PytorchCudaTag {
$smiExe = if ($script:NvidiaSmiExe) { $script:NvidiaSmiExe } else {
$cmd = Get-Command nvidia-smi -ErrorAction SilentlyContinue
if ($cmd) { $cmd.Source } else { $null }
}
if (-not $smiExe) { return "cu126" }
try {
# Bounded: a wedged nvidia-smi must not hang setup. The helper merges
# stderr into the returned string, matching the old 2>&1 | Out-String
# shape (plain 2>$null leaks ErrorRecord objects in PS 5.1).
$output = Invoke-NvidiaSmiBounded $smiExe
# Newer NVIDIA drivers (e.g. 610.x on Windows) print
# "CUDA UMD Version: X.Y" instead of the legacy "CUDA Version: X.Y".
# Accept both spellings so we don't fall through to the cu126 default.
if ($output -match 'CUDA(?: UMD)? Version:\s+(\d+)\.(\d+)') {
$major = [int]$Matches[1]
$minor = [int]$Matches[2]
# PyTorch 2.10 offers: cu124, cu126, cu128, cu130
if ($major -ge 13) { return "cu130" }
if ($major -eq 12 -and $minor -ge 8) { return "cu128" }
if ($major -eq 12 -and $minor -ge 6) { return "cu126" }
if ($major -ge 12) { return "cu124" }
if ($major -ge 11) { return "cu118" }
return "cpu"
}
} catch { }
return "cu126"
}
# Trim trailing slashes from the URL PATH only, preserving ?query / #fragment: a whole-URL
# TrimEnd corrupts a token ending in "/", a single strip leaves .../cu128// empty. Shared.
function Trim-IndexPathSlashes {
param([string]$Url)
$value = $Url.Trim()
$idx = $value.IndexOfAny([char[]]@('?', '#'))
if ($idx -lt 0) {
return $value.TrimEnd('/')
}
return $value.Substring(0, $idx).TrimEnd('/') + $value.Substring($idx)
}
# Explicit torch-index pin (UNSLOTH_TORCH_INDEX_URL / _FAMILY), shared by the stale-venv check
# and install selection so a pinned index wins over GPU probing (parity with the other
# installers). URL is verbatim; _FAMILY is the leaf joined to the mirror base.
function Get-PinnedTorchIndexUrl {
if (-not [string]::IsNullOrWhiteSpace($env:UNSLOTH_TORCH_INDEX_URL)) {
return (Trim-IndexPathSlashes $env:UNSLOTH_TORCH_INDEX_URL)
}
if (-not [string]::IsNullOrWhiteSpace($env:UNSLOTH_TORCH_INDEX_FAMILY)) {
$base = if ($env:UNSLOTH_PYTORCH_MIRROR) { $env:UNSLOTH_PYTORCH_MIRROR.TrimEnd('/') } else { "https://download.pytorch.org/whl" }
return "$base/$($env:UNSLOTH_TORCH_INDEX_FAMILY.Trim().Trim('/'))"
}
return $null
}
# Last path segment of a wheel index URL, query/fragment dropped first so a token-authenticated
# pin (.../cu128?token=x) classifies as cu128 (else it reinstalls every update). Classification
# only. Shared with the py / install.sh leaf extractors.
function Get-TorchIndexLeaf {
param([string]$Url)
if ([string]::IsNullOrWhiteSpace($Url)) { return $null }
$path = ($Url -split '[?#]', 2)[0]
if ([string]::IsNullOrWhiteSpace($path)) { return $null }
return ($path.TrimEnd('/') -split '/')[-1].ToLowerInvariant()
}
# Redact index-URL credentials (userinfo + ?query= + #fragment) from captured installer
# output before printing on failure; uv/pip errors echo the failing --index-url verbatim.
# Mirrors the other installers. Verbose mode streams uncaptured, so it isn't redacted.
function Redact-InstallOutput {
param([string]$Text)
if (-not $Text) { return $Text }
$Text = $Text -replace '(https?://)[^/@\s`]+@', '$1<redacted>@'
$Text = $Text -replace '([?&][^=\s&`]+)=[^&#\s`]+', '$1=<redacted>'
# A #token=... fragment is as sensitive as a query; URL-anchored.
return $Text -replace '(https?://[^\s`#]+)#[^\s`]+', '$1#<redacted>'
}
# AMD per-arch leaves needing the torch 2.11 floor (the _grouped_mm <2.11 bug). MUST match
# the install-spec path below and the other installers; other leaves ship <2.11 and stay default.
function Test-RocmGfx211Leaf {
param([string]$Leaf)
return @('gfx120x-all', 'gfx1151', 'gfx1150') -contains $Leaf
}
# rocmX.Y versions KNOWN to ship torch 2.11: rocm7.2 only today. Do NOT floor an unknown newer
# rocm speculatively. MUST match _ROCM_KNOWN_TORCH211_VERSIONS and the rocm7.2 leaf elsewhere.
function Test-RocmKnown211Version {
param([int]$Major, [int]$Minor)
return ($Major -eq 7 -and $Minor -eq 2)
}
# True only for a real CUDA family leaf: "cu" + digits (cu118, cu128, ...). A bare -like 'cu*'
# would match "custom"/"current" and rebuild the venv every run. Mirrors _is_cuda_family_leaf.
function Test-CudaFamilyLeaf {
param([string]$Leaf)
if ([string]::IsNullOrWhiteSpace($Leaf)) { return $false }
# EXACT cu+digits: cu128-private routes through the unknown-leaf path instead.
return $Leaf -match '^cu[0-9]+$'
}
# True only for a real pip ROCm family leaf: EXACT rocm<digits>[.<digits>] or a gfx leaf. A leaf
# that merely STARTS with rocm (rocm-rel-7.2.1, rocm7.2-private) is a custom pin the verbatim
# path owns, so anchor the match. Mirrors _is_pip_rocm_family_leaf / install.sh.
function Test-PipRocmFamilyLeaf {
param([string]$Leaf)
if ([string]::IsNullOrWhiteSpace($Leaf)) { return $false }
# gfx must be followed by a digit (an architecture leaf); gfx-private is custom.
return ($Leaf -match '^gfx[0-9]') -or ($Leaf -match '^rocm[0-9]+(\.[0-9]+)?$')
}
# Stale-venv ROCm comparison for a pinned gfx*/rocm* index. Returns @{ Expected; Installed } so
# the caller rebuilds when they differ. Mirrors _rocm_pin_family_mismatch (same rocmX.Y / gfx
# cases). An untagged (no +rocm) wheel never satisfies a ROCm pin -> stale.
function Get-RocmPinStaleTags {
param([string]$PinLeaf, [string]$TorchVersion)
$_pinRocm = [regex]::Match($PinLeaf, '^rocm(\d+)\.(\d+)')
$_pinVer = if ($_pinRocm.Success) { "$($_pinRocm.Groups[1].Value).$($_pinRocm.Groups[2].Value)" } else { $null }
# The family classifier accepts a major-only rocm<d> leaf too (rocm7).
$_pinMajorOnly = [regex]::Match($PinLeaf, '^rocm(\d+)$')
# Installed rocm version and whether the wheel is a per-arch (three-part) build.
$_instRocm = [regex]::Match($TorchVersion, '\+rocm(\d+)\.(\d+)')
$_instVer = if ($_instRocm.Success) { "$($_instRocm.Groups[1].Value).$($_instRocm.Groups[2].Value)" } else { $null }
$_instPerArch = [regex]::IsMatch($TorchVersion, '\+rocm\d+\.\d+\.\d+')
# A ROCm build MUST carry a +rocm tag; an untagged wheel can't satisfy any ROCm pin.
$_instHasRocm = [regex]::IsMatch($TorchVersion, '\+rocm')
$_instRel = [regex]::Match($TorchVersion, '^(\d+)\.(\d+)')
$_instIs211 = $false
if ($_instRel.Success) {
$_instIs211 = ([int]$_instRel.Groups[1].Value -gt 2) -or ([int]$_instRel.Groups[1].Value -eq 2 -and [int]$_instRel.Groups[2].Value -ge 11)
}
if ($PinLeaf -like 'gfx*') {
if (Test-RocmGfx211Leaf $PinLeaf) {
# Expect the AMD per-arch (three-part) 2.11 wheel: satisfied only when BOTH
# a 2.11 release AND a three-part rocm tag are installed.
$installed = if ($_instIs211 -and $_instPerArch) { "rocm-perarch(torch>=2.11)" } else { "rocm-generic-or-old" }
return @{ Expected = "rocm-perarch(torch>=2.11)"; Installed = $installed }
}
# Non-2.11 gfx leaf (<2.11 spec): stale on an untagged wheel or a 2.11+ build.
$installed = if (-not $_instHasRocm) { "not-rocm" } elseif ($_instIs211) { "rocm(torch>=2.11)" } else { "rocm(torch<2.11)" }
return @{
Expected = "rocm(torch<2.11)"
Installed = $installed
}
}
# Major-only rocm pin (rocm7): compare majors only -- a +rocm6.4 wheel under a rocm7
# pin is stale, any +rocm7.x wheel satisfies it (no pinned minor to compare, and the
# 2.11-line fallback below would invert both verdicts). Mirrors _rocm_pin_family_mismatch.
if ($_pinMajorOnly.Success) {
$_pinMaj = [int]$_pinMajorOnly.Groups[1].Value
if ($_instVer) {
$_instMaj = [int]$_instRocm.Groups[1].Value
$expected = if ($_instMaj -eq $_pinMaj) { "rocm$_instVer" } else { "rocm$_pinMaj.x" }
return @{ Expected = $expected; Installed = "rocm$_instVer" }
}
# Untagged wheel never satisfies a ROCm pin; a +rocm tag with an unreadable
# version is accepted (matches the lenient unreadable fallback below).
$installed = if ($_instHasRocm) { "rocm" } else { "not-rocm" }
return @{ Expected = "rocm"; Installed = $installed }
}
# rocmX.Y pin.
if ($_pinVer -and $_instVer) {
# Both readable: exact compare. When they match AND the pin is KNOWN-2.11, the
# installed release must also be 2.11 (a +rocm7.2 wheel drifted to 2.12 shares the
# tag but violates the spec), so fold the release into the tag. Mirrors _rocm_pin_family_mismatch.
$_pinKnown211 = Test-RocmKnown211Version -Major ([int]$_pinRocm.Groups[1].Value) -Minor ([int]$_pinRocm.Groups[2].Value)
$_instOn211 = $_instRel.Success -and [int]$_instRel.Groups[1].Value -eq 2 -and [int]$_instRel.Groups[2].Value -eq 11
if ($_pinKnown211 -and -not $_instOn211) {
return @{ Expected = "rocm$_pinVer(torch2.11)"; Installed = "rocm$_instVer(torch-off-2.11)" }
}
return @{ Expected = "rocm$_pinVer"; Installed = "rocm$_instVer" }
}
$_pinNeeds211 = $false
if ($_pinRocm.Success) {
# Only KNOWN-2.11 rocm (rocm7.2) is on the 2.11 line (no speculative floor).
# Matches _ROCM_KNOWN_TORCH211_VERSIONS.
$_pinNeeds211 = Test-RocmKnown211Version -Major ([int]$_pinRocm.Groups[1].Value) -Minor ([int]$_pinRocm.Groups[2].Value)
}
# Fallback (installed rocm version unreadable): compare on the 2.11 line; an untagged
# wheel never satisfies a rocmX.Y pin -> stale.
$installed = if (-not $_instHasRocm) { "not-rocm" } elseif ($_instIs211) { "rocm(torch>=2.11)" } else { "rocm(torch<2.11)" }
return @{
Expected = if ($_pinNeeds211) { "rocm(torch>=2.11)" } else { "rocm(torch<2.11)" }
Installed = $installed
}
}
# VS generator -> MSBuild BuildCustomizations dir; toolset tracks the VS major
# (18->v180, 17->v170), defaulting to v170 when unparseable.
function Get-VcBuildCustomizationsDir {
param(
[Parameter(Mandatory)][string]$VsInstallPath,
[string]$Generator
)
$toolset = 'v170'
if ($Generator -and ($Generator -match 'Visual Studio (\d+)\b')) {
$toolset = "v$($Matches[1])0"
}
return (Join-Path $VsInstallPath "MSBuild\Microsoft\VC\$toolset\BuildCustomizations")
}
# Installed cmake version, or $null if absent/unparseable.
function Get-CmakeVersion {
$raw = & cmake --version 2>$null | Select-Object -First 1
if ($raw -and ($raw -match '(\d+)\.(\d+)(?:\.(\d+))?')) {
$patch = if ($Matches[3]) { $Matches[3] } else { '0' }
return [version]"$($Matches[1]).$($Matches[2]).$patch"
}
return $null
}
# VS 18 2026 generator needs cmake >= 4.2 (added there); true for older VS generators.
function Test-CmakeSupportsGenerator {
param(
[Parameter(Mandatory)][string]$CmakeVersion,
[Parameter(Mandatory)][string]$Generator
)
if ($Generator -match 'Visual Studio 18\b') {
$clean = ($CmakeVersion -replace '[^0-9.].*$', '').TrimEnd('.')
try { $v = [version]$clean } catch { return $false }
return ($v -ge [version]'4.2')
}
return $true
}
function Test-CmakeListsGenerator {
# Does `cmake --help` actually list the generator? A VS-bundled cmake can drive
# VS 2026 below the 4.2 floor, so probe rather than trust the version. (#6473)
param([Parameter(Mandatory)][string]$Generator)
$help = & cmake --help 2>$null | Out-String
if (-not $help) { return $false }
$haystack = ($help -replace '\s+', ' ')
$needle = ($Generator -replace '\s+', ' ')
return $haystack.Contains($needle)
}
function Test-CmakeCanDriveGenerator {
# cmake can drive $Generator if it lists it (VS-bundled below 4.2) or meets the floor.
param([Parameter(Mandatory)][string]$Generator)
if (Test-CmakeListsGenerator -Generator $Generator) { return $true }
$verObj = Get-CmakeVersion
$verStr = if ($verObj) { $verObj.ToString() } else { '0.0' }
return (Test-CmakeSupportsGenerator -CmakeVersion $verStr -Generator $Generator)
}
function Add-DefaultCmakeToPath {
# Prepend the default CMake dir so a freshly winget-installed cmake wins over an
# older one already on PATH. $true if found. (#6473)
$cmakeDefaults = @(
"$env:ProgramFiles\CMake\bin",
"${env:ProgramFiles(x86)}\CMake\bin",
"$env:LOCALAPPDATA\CMake\bin"
)
foreach ($d in $cmakeDefaults) {
if (Test-Path (Join-Path $d "cmake.exe")) {
$env:Path = "$d;$env:Path"
Add-ToUserPath -Directory $d -Position 'Prepend' | Out-Null
return $true
}
}
return $false
}
function Get-FallbackVsGenerator {
# Newest pre-2026 VS whose generator the current cmake can drive, for when the
# VS 2026 generator is unusable (old/offline cmake) but an older toolchain exists.
# vswhere first (catches non-default roots like D:\), then Program Files; matches
# Find-VsBuildTools. Returns @{ Generator; InstallPath } or $null. (#6473)
$knownEditions = @('BuildTools', 'Community', 'Professional', 'Enterprise', 'Preview')
# install path if it holds a usable cl.exe, else $null
$tryCandidate = {
param($gen, $installPath)
if (-not $installPath) { return $null }
$vcDir = Join-Path $installPath "VC\Tools\MSVC"
if (-not (Test-Path $vcDir)) { return $null }
$cl = Get-ChildItem -Path $vcDir -Filter "cl.exe" -Recurse -ErrorAction SilentlyContinue | Select-Object -First 1
if ($cl) { return @{ Generator = $gen; InstallPath = $installPath } }
return $null
}
# vswhere (non-default roots)
$vsw = "${env:ProgramFiles(x86)}\Microsoft Visual Studio\Installer\vswhere.exe"
if (Test-Path $vsw) {
$json = & $vsw -all -requires Microsoft.VisualStudio.Component.VC.Tools.x86.x64 -format json 2>$null | Out-String
if ($json) {
try { $instances = @($json | ConvertFrom-Json) } catch { $instances = @() }
$ranked = $instances | ForEach-Object {
$label = if ($_.catalog -and $_.catalog.productLineVersion) { [string]$_.catalog.productLineVersion } else { '' }
[pscustomobject]@{ Gen = (Resolve-VsGeneratorFromLabel $label); Path = [string]$_.installationPath }
} | Where-Object { $_.Gen -and ($_.Gen -notmatch 'Visual Studio 18\b') }
# newest first: 2022 > 2019 > 2017
$ranked = $ranked | Sort-Object { switch -regex ($_.Gen) { '17 2022' {0} '16 2019' {1} '15 2017' {2} default {9} } }
foreach ($cand in $ranked) {
if (-not (Test-CmakeListsGenerator -Generator $cand.Gen)) { continue }
$res = & $tryCandidate $cand.Gen $cand.Path
if ($res) { return $res }
}
}
}
# Program Files scan
$roots = @($env:ProgramFiles, ${env:ProgramFiles(x86)}) | Where-Object { $_ }
$older = @(
@{ Dir = '2022'; Generator = 'Visual Studio 17 2022' },
@{ Dir = '2019'; Generator = 'Visual Studio 16 2019' },
@{ Dir = '2017'; Generator = 'Visual Studio 15 2017' }
)
foreach ($entry in $older) {
if (-not (Test-CmakeListsGenerator -Generator $entry.Generator)) { continue }
foreach ($r in $roots) {
$vsBase = Join-Path $r "Microsoft Visual Studio\$($entry.Dir)"
if (-not (Test-Path $vsBase)) { continue }
foreach ($ed in $knownEditions) {
$candidate = Join-Path $vsBase $ed
if (-not (Test-Path $candidate)) { continue }
$res = & $tryCandidate $entry.Generator $candidate
if ($res) { return $res }
}
}
}
return $null
}
# VS version label -> cmake generator. vswhere's productLineVersion is the year for
# VS <= 2022 but the internal major "18" for VS 2026, and dir names use either form,
# so accept both. (VS 2026 detection adapted from @LeoBorcherding's #6038.)
function Resolve-VsGeneratorFromLabel {
param([string]$Label)
if (-not $Label) { return $null }
$map = @{
'2026' = 'Visual Studio 18 2026'; '18' = 'Visual Studio 18 2026'
'2022' = 'Visual Studio 17 2022'; '17' = 'Visual Studio 17 2022'
'2019' = 'Visual Studio 16 2019'; '16' = 'Visual Studio 16 2019'
'2017' = 'Visual Studio 15 2017'; '15' = 'Visual Studio 15 2017'
}
return $map[$Label.Trim()]
}
# Find VS Build Tools for cmake -G: vswhere, then a filesystem scan (handles broken
# vswhere registration). Returns @{ Generator; InstallPath; Source } or $null.
function Find-VsBuildTools {
# vswhere first (works when VS is properly registered)
$vsw = "${env:ProgramFiles(x86)}\Microsoft Visual Studio\Installer\vswhere.exe"
if (Test-Path $vsw) {
$info = & $vsw -latest -requires Microsoft.VisualStudio.Component.VC.Tools.x86.x64 -property catalog_productLineVersion 2>$null
$path = & $vsw -latest -requires Microsoft.VisualStudio.Component.VC.Tools.x86.x64 -property installationPath 2>$null
if ($info -and $path) {
$gen = Resolve-VsGeneratorFromLabel $info
if ($gen) {
return @{ Generator = $gen; InstallPath = $path.Trim(); Source = 'vswhere' }
}
}
}
# filesystem scan (handles broken vswhere registration); VS 2026+ dir is "18"
$roots = @($env:ProgramFiles, ${env:ProgramFiles(x86)}) | Where-Object { $_ }
$knownEditions = @('BuildTools', 'Community', 'Professional', 'Enterprise', 'Preview')
$dirs = @('18', '2026', '2022', '2019', '2017')
foreach ($d in $dirs) {
$gen = Resolve-VsGeneratorFromLabel $d
if (-not $gen) { continue }
foreach ($r in $roots) {
$vsBase = Join-Path $r "Microsoft Visual Studio\$d"
if (-not (Test-Path $vsBase)) { continue }
# VS 2026 (dir "18") may use non-standard edition names, so scan every subdir
if ($d -eq '18' -or $d -eq '2026') {
$editionCandidates = Get-ChildItem -Path $vsBase -Directory -ErrorAction SilentlyContinue | ForEach-Object { $_.FullName }
} else {
$editionCandidates = $knownEditions | ForEach-Object { Join-Path $vsBase $_ }
}
foreach ($candidate in $editionCandidates) {
if (-not (Test-Path $candidate)) { continue }
$vcDir = Join-Path $candidate "VC\Tools\MSVC"
if (Test-Path $vcDir) {
$cl = Get-ChildItem -Path $vcDir -Filter "cl.exe" -Recurse -ErrorAction SilentlyContinue | Select-Object -First 1
if ($cl) {
$ed = Split-Path $candidate -Leaf
return @{ Generator = $gen; InstallPath = $candidate; Source = "filesystem ($ed)"; ClExe = $cl.FullName }
}
}
}
}
}
return $null
}
# Install CMake + VS Build Tools, deferred here from Phase 1 so the prebuilt path
# never pays for a multi-GB install. Called only when a source build is committed.
# CMake is best-effort (build skips downstream if absent); VS Build Tools are
# required, so exit 1 with guidance if missing. No-ops for VS when already detected.
function Ensure-BuildToolsForLlamaSourceBuild {
# CMake
if ($null -eq (Get-Command cmake -ErrorAction SilentlyContinue)) {
Write-Host "CMake not found -- installing via winget (needed for the llama.cpp source build)..." -ForegroundColor Yellow
if ($null -ne (Get-Command winget -ErrorAction SilentlyContinue)) {
try {
Invoke-SetupCommand { winget install Kitware.CMake --source winget --accept-package-agreements --accept-source-agreements } | Out-Null
Refresh-Environment
} catch { }
}
# winget may install cmake but not put it on PATH yet; try the default dir
if ($null -eq (Get-Command cmake -ErrorAction SilentlyContinue)) {
$cmakeDefaults = @(
"$env:ProgramFiles\CMake\bin",
"${env:ProgramFiles(x86)}\CMake\bin",
"$env:LOCALAPPDATA\CMake\bin"
)
foreach ($d in $cmakeDefaults) {
if (Test-Path (Join-Path $d "cmake.exe")) {
$env:Path = "$d;$env:Path"
Add-ToUserPath -Directory $d -Position 'Prepend' | Out-Null
break
}
}
}
if ($null -ne (Get-Command cmake -ErrorAction SilentlyContinue)) { step "cmake" "installed" }
}
# VS Build Tools
if ($script:VsInstallPath) { return } # already detected by the early probe
$vsResult = Find-VsBuildTools
if (-not $vsResult) {
Write-Host "Visual Studio Build Tools not found -- installing via winget..." -ForegroundColor Yellow
Write-Host " (Needed only for the llama.cpp source build; may take several minutes)" -ForegroundColor Gray
if ($null -ne (Get-Command winget -ErrorAction SilentlyContinue)) {
$prevEAPTemp = $ErrorActionPreference
$ErrorActionPreference = "Continue"
winget install Microsoft.VisualStudio.2022.BuildTools --source winget --accept-package-agreements --accept-source-agreements --override "--add Microsoft.VisualStudio.Workload.VCTools --includeRecommended --passive --wait"
$ErrorActionPreference = $prevEAPTemp
# Re-scan after install (don't trust vswhere catalog)
$vsResult = Find-VsBuildTools
}
}
if ($vsResult) {
$script:CmakeGenerator = $vsResult.Generator
$script:VsInstallPath = $vsResult.InstallPath
step "vs" "$($vsResult.Generator) ($($vsResult.Source))"
if ($vsResult.ClExe) { substep "cl.exe: $($vsResult.ClExe)" }
} else {
Write-Host "[ERROR] Visual Studio Build Tools are required for the llama.cpp source build but could not be found or installed." -ForegroundColor Red
Write-Host " Manual install:" -ForegroundColor Red
Write-Host ' 1. winget install Microsoft.VisualStudio.2022.BuildTools --source winget' -ForegroundColor Yellow
Write-Host ' 2. Open Visual Studio Installer -> Modify -> check "Desktop development with C++"' -ForegroundColor Yellow
exit 1
}
}
# Detect the VC++ 2015-2022 Redistributable that the prebuilt llama-server and
# PyTorch need (they link VCRUNTIME140_1.dll etc., which the Universal CRT lacks).
# Signal is System32\vcruntime140_1.dll (VS 2019+), registry as fallback.
function Test-VCRedistInstalled {
$sys = $env:SystemRoot
if ($sys -and (Test-Path (Join-Path $sys 'System32\vcruntime140_1.dll'))) { return $true }
foreach ($k in @(
'HKLM:\SOFTWARE\Microsoft\VisualStudio\14.0\VC\Runtimes\x64',
'HKLM:\SOFTWARE\WOW6432Node\Microsoft\VisualStudio\14.0\VC\Runtimes\x64'
)) {
try {
$r = Get-ItemProperty -Path $k -ErrorAction Stop
if ($r.Installed -eq 1 -and [int]$r.Major -ge 14 -and [int]$r.Minor -ge 20) { return $true }
} catch { }
}
return $false
}
# Install the VC++ 2015-2022 runtime if missing (non-fatal; usually a no-op).
function Ensure-VCRedist {
if (Test-VCRedistInstalled) { step "vcredist" "present"; return }
Write-Host "Microsoft Visual C++ Redistributable (2015-2022) is missing; the prebuilt llama.cpp and PyTorch need it. Installing the runtime..." -ForegroundColor Yellow
if ($null -ne (Get-Command winget -ErrorAction SilentlyContinue)) {
try {
Invoke-SetupCommand { winget install --id Microsoft.VCRedist.2015+.x64 --source winget --accept-package-agreements --accept-source-agreements } | Out-Null
Refresh-Environment
} catch { substep "VCRedist install failed: $($_.Exception.Message)" "Yellow" }
}
if (Test-VCRedistInstalled) { step "vcredist" "installed" }
else {
substep "Could not install the VC++ Redistributable automatically." "Yellow"
substep "If llama-server or torch reports a missing VCRUNTIME140.dll, install:" "Yellow"
substep "https://aka.ms/vs/17/release/vc_redist.x64.exe" "Yellow"
}
}
# ─────────────────────────────────────────────
# Output style (aligned with studio/setup.sh: step / substep)
# ─────────────────────────────────────────────
$Rule = [string]::new([char]0x2500, 52)
function Enable-StudioVirtualTerminal {
if ($env:NO_COLOR) { return $false }
try {
Add-Type -Namespace StudioVT -Name Native -MemberDefinition @'
[DllImport("kernel32.dll")] public static extern IntPtr GetStdHandle(int nStdHandle);
[DllImport("kernel32.dll")] public static extern bool GetConsoleMode(IntPtr h, out uint m);
[DllImport("kernel32.dll")] public static extern bool SetConsoleMode(IntPtr h, uint m);
'@ -ErrorAction Stop
$h = [StudioVT.Native]::GetStdHandle(-11)
[uint32]$mode = 0
if (-not [StudioVT.Native]::GetConsoleMode($h, [ref]$mode)) { return $false }
$mode = $mode -bor 0x0004
return [StudioVT.Native]::SetConsoleMode($h, $mode)
} catch {
return $false
}
}
$script:StudioVtOk = Enable-StudioVirtualTerminal
function Get-StudioAnsi {
param(
[Parameter(Mandatory = $true)]
[ValidateSet('Title', 'Dim', 'Ok', 'Warn', 'Err', 'Reset')]
[string]$Kind
)
$e = [char]27
switch ($Kind) {
'Title' { return "${e}[38;5;150m" }
'Dim' { return "${e}[38;5;245m" }
'Ok' { return "${e}[38;5;108m" }
'Warn' { return "${e}[38;5;136m" }
'Err' { return "${e}[91m" }
'Reset' { return "${e}[0m" }
}
}
function Write-SetupVerboseDetail {
param(
[Parameter(Mandatory = $true)][string]$Message,
[string]$Color = "Gray"
)
if (-not $script:UnslothVerbose) { return }
if ($script:StudioVtOk -and -not $env:NO_COLOR) {
$ansi = switch ($Color) {
'Green' { (Get-StudioAnsi Ok) }
'Gray' { (Get-StudioAnsi Dim) }
'DarkGray' { (Get-StudioAnsi Dim) }
'Yellow' { (Get-StudioAnsi Warn) }
'Cyan' { (Get-StudioAnsi Title) }
'Red' { (Get-StudioAnsi Err) }
default { (Get-StudioAnsi Dim) }
}
Write-Host ($ansi + $Message + (Get-StudioAnsi Reset))
} else {
$fc = switch ($Color) {
'Green' { 'DarkGreen' }
'Gray' { 'DarkGray' }
'Cyan' { 'Green' }
default { $Color }
}
Write-Host $Message -ForegroundColor $fc
}
}
function Invoke-SetupCommand {
param(
[Parameter(Mandatory = $true)][scriptblock]$Command,
[switch]$AlwaysQuiet
)
$prevEap = $ErrorActionPreference
$ErrorActionPreference = "Continue"
try {
# Reset to avoid stale values from prior native commands.
$global:LASTEXITCODE = 0
if ($script:UnslothVerbose -and -not $AlwaysQuiet) {
# Merge stderr into stdout so progress/warning output stays visible
# without flipping $? on successful native commands (PS 5.1 treats
# stderr records as errors that set $? = $false even on exit code 0).
# Redact per record: uv/pip echo index URLs (credentials and all) in
# their errors, and verbose mode must not bypass the quiet path's
# redaction. ForEach-Object/Out-Host leave $LASTEXITCODE untouched.
& $Command 2>&1 | ForEach-Object { Redact-InstallOutput "$_" } | Out-Host
} else {
$output = & $Command 2>&1 | Out-String
if ($LASTEXITCODE -ne 0) {
Write-Host (Redact-InstallOutput $output) -ForegroundColor Red
}
}
return [int]$LASTEXITCODE
} finally {
$ErrorActionPreference = $prevEap
}
}
function Write-LlamaFailureLog {
param(
[string]$Output,
[int]$MaxLines = 120
)
if (-not $Output) { return }
$lines = @(
($Output -split "`r?`n") | Where-Object { -not [string]::IsNullOrWhiteSpace($_) }
)
if ($lines.Count -eq 0) { return }
if ($lines.Count -gt $MaxLines) {
Write-Host " Showing last $MaxLines lines:" -ForegroundColor DarkGray
$lines = $lines | Select-Object -Last $MaxLines
}
foreach ($line in $lines) {
Write-Host " | $line" -ForegroundColor DarkGray
}
}
# Mirror the plain (no ANSI) form of step/substep messages to the
# OS-level stdout handle when a parent is consuming our stdout via
# a pipe (CI `tee`, Python subprocess.PIPE, CREATE_NO_WINDOW grandchild).
# Write-Host on PS 5.1 routes through $Host.UI / the Information
# stream, neither of which propagates reliably across the
# install.ps1 -> unsloth.exe -> python -> powershell.exe ->
# setup.ps1 process chain. [Console]::Out always lands on the OS
# stdout file handle. Gated on IsOutputRedirected so the
# interactive-console path keeps the colorized Write-Host output
# only (no double-print).
function Write-StudioStdoutMirror {
param([Parameter(Mandatory = $true)][string]$Line)
try {
if ([Console]::IsOutputRedirected) {
[Console]::Out.WriteLine($Line)
[Console]::Out.Flush()
}
} catch {}
}
function step {
param(
[Parameter(Mandatory = $true)][string]$Label,
[Parameter(Mandatory = $true)][string]$Value,
[string]$Color = "Green"
)
$padded = if ($Label.Length -ge 15) { $Label.Substring(0, 15) } else { $Label.PadRight(15) }
if ($script:StudioVtOk -and -not $env:NO_COLOR) {
$dim = Get-StudioAnsi Dim
$rst = Get-StudioAnsi Reset
$val = switch ($Color) {
'Green' { Get-StudioAnsi Ok }
'Yellow' { Get-StudioAnsi Warn }
'Red' { Get-StudioAnsi Err }
'DarkGray' { Get-StudioAnsi Dim }
default { Get-StudioAnsi Ok }
}
Write-Host (" {0}{1}{2}{3}{4}{2}" -f $dim, $padded, $rst, $val, $Value)
} else {
Write-Host (" {0}" -f $padded) -NoNewline -ForegroundColor DarkGray
$fc = switch ($Color) {
'Green' { 'DarkGreen' }
'Yellow' { 'Yellow' }
'Red' { 'Red' }
'DarkGray' { 'DarkGray' }
default { 'DarkGreen' }
}
Write-Host $Value -ForegroundColor $fc
}
Write-StudioStdoutMirror (" {0}{1}" -f $padded, $Value)
}
function substep {
param(
[Parameter(Mandatory = $true)][string]$Message,
[string]$Color = "DarkGray"
)
if ($script:StudioVtOk -and -not $env:NO_COLOR) {
$msgCol = switch ($Color) {
'Yellow' { (Get-StudioAnsi Warn) }
default { (Get-StudioAnsi Dim) }
}
$pad = "".PadRight(15)
Write-Host (" {0}{1}{2}{3}" -f $msgCol, $pad, $Message, (Get-StudioAnsi Reset))
} else {
$fc = switch ($Color) {
'Yellow' { 'Yellow' }
default { 'DarkGray' }
}
Write-Host (" {0,-15}{1}" -f "", $Message) -ForegroundColor $fc
}
Write-StudioStdoutMirror (" {0,-15}{1}" -f "", $Message)
}
function Show-NpmRegistryHint {
# Print actionable guidance when a frontend/OXC npm/bun install fails and the
# registry lock is the likely cause (corporate firewall/proxy). No-op once the
# user has opted in via UNSLOTH_NPM_REGISTRY. We never switch registries
# automatically -- we only guide.
if ($env:UNSLOTH_NPM_REGISTRY) { return }
$mirror = $env:NPM_CONFIG_REGISTRY
if (-not $mirror) {
# Read npm config from a dir with no project .npmrc so the frontend's pinned
# registry= does not mask the user's ~/.npmrc / global mirror.
$pushed = $false
try {
Push-Location ([System.IO.Path]::GetTempPath()) -ErrorAction Stop
$pushed = $true
$mirror = (& npm config get registry 2>$null | Out-String).Trim()
} catch { $mirror = "" } finally { if ($pushed) { Pop-Location } }
}
if ($mirror -in @("", "undefined", "null", "https://registry.npmjs.org", "https://registry.npmjs.org/")) {
$mirror = ""
}
Write-Host ""
step "frontend" "registry.npmjs.org looks blocked (corporate firewall/proxy?)" "Yellow"
if ($mirror) {
substep "Unsloth pins the public npm registry; your mirror is being ignored."
substep "Detected a registry in your npm config:"
substep " $mirror"
substep "Re-run pointing Unsloth at it:"
substep " `$env:UNSLOTH_NPM_REGISTRY='$mirror'; .\install.ps1 --local"
} else {
substep "If you use a private mirror/proxy, point Unsloth at it and re-run:"
substep " `$env:UNSLOTH_NPM_REGISTRY='https://your-mirror.example/api/npm/'; .\install.ps1 --local"
}
substep "(min-release-age and save-exact stay enforced.)"
}
# ─────────────────────────────────────────────
# Banner
# ─────────────────────────────────────────────
Write-Host ""
if ($script:StudioVtOk -and -not $env:NO_COLOR) {
Write-Host (" " + (Get-StudioAnsi Title) + [char]::ConvertFromUtf32(0x1F9A5) + " Unsloth Studio Setup" + (Get-StudioAnsi Reset))
Write-Host (" {0}{1}{2}" -f (Get-StudioAnsi Dim), $Rule, (Get-StudioAnsi Reset))
} else {
Write-Host (" " + [char]::ConvertFromUtf32(0x1F9A5) + " Unsloth Studio Setup") -ForegroundColor Green
Write-Host " $Rule" -ForegroundColor DarkGray
}
# Back up User PATH under HKCU\Software\Unsloth before any modifications.
try {
$envKey = [Microsoft.Win32.Registry]::CurrentUser.OpenSubKey('Environment', $false)
if ($envKey) {
try {
$rawPath = $envKey.GetValue('Path', '', [Microsoft.Win32.RegistryValueOptions]::DoNotExpandEnvironmentNames)
} finally {
$envKey.Close()
}
if ($rawPath) {
$backupKey = [Microsoft.Win32.Registry]::CurrentUser.CreateSubKey('Software\Unsloth')
try {
$existingBackup = $backupKey.GetValue('PathBackup', $null)
if (-not $existingBackup) {
$backupKey.SetValue('PathBackup', $rawPath, [Microsoft.Win32.RegistryValueKind]::ExpandString)
}
} finally {
$backupKey.Close()
}
}
}
} catch {
Write-Host "[DEBUG] Could not back up User PATH: $($_.Exception.Message)" -ForegroundColor DarkGray
}
# ==========================================================================
# PHASE 1: System-level prerequisites (winget installs, env vars)
# All heavy system tool installs happen here BEFORE touching Python.
# ==========================================================================
# ============================================
# 1a. GPU detection
# ============================================
# ── Helper: run nvidia-smi under a timeout ──
# A wedged NVIDIA driver can make nvidia-smi block during init or after a reset;
# WaitForExit bounds it (mirrors Invoke-AmdSmiNoElevate below) so detection
# cannot hang setup. No RunAsInvoker compat layer: nvidia-smi does not
# auto-elevate. Returns combined stdout+stderr; "" on timeout/failure.
function Invoke-NvidiaSmiBounded {
param(
[Parameter(Mandatory = $true, Position = 0)][string]$Exe,
[Parameter(Position = 1)][string[]]$SmiArgs = @(),
[int]$TimeoutSec = 10
)
try {
$psi = New-Object System.Diagnostics.ProcessStartInfo
$psi.FileName = $Exe
$psi.Arguments = ($SmiArgs -join ' ')
$psi.UseShellExecute = $false
$psi.RedirectStandardOutput = $true
$psi.RedirectStandardError = $true
$psi.CreateNoWindow = $true
$proc = [System.Diagnostics.Process]::Start($psi)
$outTask = $proc.StandardOutput.ReadToEndAsync()
$errTask = $proc.StandardError.ReadToEndAsync()
if (-not $proc.WaitForExit($TimeoutSec * 1000)) {
try { $proc.Kill() } catch {}
$global:LASTEXITCODE = 124
return ""
}
$global:LASTEXITCODE = $proc.ExitCode
return ($outTask.Result + "`n" + $errTask.Result)
} catch {
$global:LASTEXITCODE = 1
return ""
}
}
# ── Helper: nvidia-smi -L lists at least one real GPU ──
# Exit code 0 alone is not enough: a stale/driverless nvidia-smi can exit 0
# while listing no GPU, which would mark an AMD host NVIDIA and suppress ROCm
# detection. Require a "GPU <n>:" data row.
function Test-NvidiaSmiHasGpu {
param([Parameter(Mandatory = $true)][string]$Exe)
$out = Invoke-NvidiaSmiBounded $Exe @('-L')
return ($LASTEXITCODE -eq 0 -and $out -match '(?m)^GPU\s+\d+:')
}
$HasNvidiaSmi = $false
$NvidiaSmiExe = $null # Absolute path -- survives Refresh-Environment
try {
$nvSmiCmd = Get-Command nvidia-smi -ErrorAction SilentlyContinue
if ($nvSmiCmd -and (Test-NvidiaSmiHasGpu $nvSmiCmd.Source)) {
$HasNvidiaSmi = $true
$NvidiaSmiExe = $nvSmiCmd.Source
}
} catch {}
# Fallback: nvidia-smi may not be on PATH even though a GPU + driver exist.
# Check the default install location and the Windows driver store.
if (-not $HasNvidiaSmi) {
$nvSmiDefaults = @(
"$env:ProgramFiles\NVIDIA Corporation\NVSMI\nvidia-smi.exe",
"$env:SystemRoot\System32\nvidia-smi.exe"
)
foreach ($p in $nvSmiDefaults) {
if (Test-Path $p) {
try {
if (Test-NvidiaSmiHasGpu $p) {
$HasNvidiaSmi = $true
$NvidiaSmiExe = $p
Write-Host " Found nvidia-smi at $(Split-Path $p -Parent)" -ForegroundColor Gray
break
}
} catch {}
}
}
}
# ── Helper: run amd-smi without triggering a UAC elevation prompt ──
# amd-smi on Windows auto-elevates to read GPU/APU memory, surfacing a confusing
# DiskPart UAC prompt mid-install (Unsloth backend amd.py hits the same). RunAsInvoker
# forces it (and helpers it spawns) to run un-elevated; on failure the WMI name ->
# gfx fallback still resolves the arch.
function Invoke-AmdSmiNoElevate {
param(
[Parameter(Mandatory = $true, Position = 0)][string]$Exe,
[Parameter(Position = 1)][string[]]$SmiArgs = @(),
[int]$TimeoutSec = 30
)
# RunAsInvoker blocks the auto-elevation/UAC prompt; the timeout bounds a flaky
# amd-smi that can otherwise spin for minutes (30s mirrors the backend amd.py).
$prevCompat = [Environment]::GetEnvironmentVariable('__COMPAT_LAYER', 'Process')
$env:__COMPAT_LAYER = 'RunAsInvoker'
try {
# [Process]::Start, NOT Start-Process -PassThru: the latter leaves .ExitCode
# $null after WaitForExit on PS 5.1, so $LASTEXITCODE (checked by callers)
# reads non-zero and kills detection. Async reads drain the pipes (no
# deadlock); amd-smi args have no spaces so a plain join is safe.
$psi = New-Object System.Diagnostics.ProcessStartInfo
$psi.FileName = $Exe
$psi.Arguments = ($SmiArgs -join ' ')
$psi.UseShellExecute = $false
$psi.RedirectStandardOutput = $true
$psi.RedirectStandardError = $true
$psi.CreateNoWindow = $true
$proc = [System.Diagnostics.Process]::Start($psi)
$outTask = $proc.StandardOutput.ReadToEndAsync()
$errTask = $proc.StandardError.ReadToEndAsync()
if (-not $proc.WaitForExit($TimeoutSec * 1000)) {
try { $proc.Kill() } catch {}
$global:LASTEXITCODE = 124
return ""
}
$global:LASTEXITCODE = $proc.ExitCode
return ($outTask.Result + "`n" + $errTask.Result)
} catch {
$global:LASTEXITCODE = 1
return ""
} finally {
if ($null -eq $prevCompat) {
Remove-Item Env:__COMPAT_LAYER -ErrorAction SilentlyContinue
} else {
$env:__COMPAT_LAYER = $prevCompat
}
}
}
# ── AMD ROCm detection (Windows): probe hipinfo/amd-smi for actual GPU ──
$HasROCm = $false
$HipSdkInstalled = $false # HIP SDK binary found (independent of device accessibility)
$ROCmGpuLabel = $null
$script:ROCmGfxArch = $null
if (-not $HasNvidiaSmi) {
# hipinfo: PATH first, then HIP_PATH/ROCM_PATH bin fallback (mirrors NVIDIA smi path resolution).
# AMD HIP SDK sets HIP_PATH but may not add the bin dir to PATH depending on install type.
# Ignore the venv hipInfo.exe (AMD wheel, on PATH): not a HIP SDK, so amd-smi
# would still auto-elevate. Cf. _path_inside_venv().
function Test-HipinfoIsVenvInternal {
param([AllowNull()][string]$HipinfoPath)
if ([string]::IsNullOrWhiteSpace($HipinfoPath)) { return $false }
# VenvDir/VIRTUAL_ENV can be unset this early (the update flow probes before
# VenvDir is set), so also derive the venv from the setup python + default
# Unsloth home, else the venv hipInfo isn't caught.
$venvRoots = @()
if ($env:VIRTUAL_ENV) { $venvRoots += $env:VIRTUAL_ENV }
$vd = Get-Variable -Name VenvDir -ValueOnly -ErrorAction SilentlyContinue
if ($vd) { $venvRoots += $vd }
if ($env:UNSLOTH_SETUP_PYTHON) {
try { $venvRoots += (Split-Path -Parent (Split-Path -Parent $env:UNSLOTH_SETUP_PYTHON)) } catch {}
}
if ($env:USERPROFILE) { $venvRoots += (Join-Path $env:USERPROFILE ".unsloth\studio\unsloth_studio") }
# A custom Unsloth home (UNSLOTH_STUDIO_HOME / STUDIO_HOME alias) moves the
# venv off the default path; seed it too or its hipInfo escapes the filter.
$studioHomeEnv = if (-not [string]::IsNullOrWhiteSpace($env:UNSLOTH_STUDIO_HOME)) { $env:UNSLOTH_STUDIO_HOME.Trim() } elseif (-not [string]::IsNullOrWhiteSpace($env:STUDIO_HOME)) { $env:STUDIO_HOME.Trim() } else { $null }
if ($studioHomeEnv) {
# Expand a leading ~ like the canonical resolver below; else GetFullPath
# keeps the literal ~ (cwd-relative) and the hipInfo escapes the filter.
if (($studioHomeEnv -eq "~" -or $studioHomeEnv -like "~/*" -or $studioHomeEnv -like "~\*") -and -not [string]::IsNullOrWhiteSpace($env:USERPROFILE)) {
# A bare "~" leaves an empty child path; Join-Path rejects that on
# PS 5.1, so use USERPROFILE directly and only join a real remainder.
$studioHomeRest = $studioHomeEnv.Substring(1).TrimStart('/', '\')
$studioHomeEnv = if ($studioHomeRest) { Join-Path $env:USERPROFILE $studioHomeRest } else { $env:USERPROFILE }
}
$venvRoots += (Join-Path $studioHomeEnv "unsloth_studio")
}
try { $hip = [System.IO.Path]::GetFullPath($HipinfoPath).TrimEnd('\', '/') } catch { return $false }
foreach ($root in $venvRoots) {
if ([string]::IsNullOrWhiteSpace($root)) { continue }
try { $r = [System.IO.Path]::GetFullPath($root).TrimEnd('\', '/') } catch { continue }
# Skip a bare drive root (e.g. a non-venv UNSLOTH_SETUP_PYTHON like
# C:\Python311\python.exe yields C:) -- it would match every path on that drive.
if ($r -match '^[a-zA-Z]:$') { continue }
if ($hip.Equals($r, [System.StringComparison]::OrdinalIgnoreCase) -or
$hip.StartsWith($r + [System.IO.Path]::DirectorySeparatorChar, [System.StringComparison]::OrdinalIgnoreCase)) {
return $true
}
}
return $false
}
# Scan all hipinfo and keep the first non-venv one (the venv copy from the
# bnb fix could shadow a real HIP SDK's). -CommandType Application matches
# only real executables, not a user alias/function named hipinfo.
$hipinfoExe = Get-Command hipinfo -CommandType Application -All -ErrorAction SilentlyContinue |
Where-Object { -not (Test-HipinfoIsVenvInternal $_.Source) } |
Select-Object -First 1
if (-not $hipinfoExe) {
# Iterate the env roots (mirrors the Python list) and take the first non-venv
# bin\hipinfo.exe, so a venv-internal HIP_PATH can't mask a real SDK in ROCM_PATH.
$hipMissingLabel = $null; $hipMissingRoot = $null; $hipMissingCandidate = $null
foreach ($hipEnvLabel in @("HIP_PATH", "HIP_PATH_57", "ROCM_PATH")) {
$hipRoot = [Environment]::GetEnvironmentVariable($hipEnvLabel)
if ([string]::IsNullOrWhiteSpace($hipRoot)) { continue }
$hipinfoCandidate = Join-Path $hipRoot "bin\hipinfo.exe"
if (-not (Test-Path $hipinfoCandidate)) {
if (-not $hipMissingLabel) { $hipMissingLabel = $hipEnvLabel; $hipMissingRoot = $hipRoot; $hipMissingCandidate = $hipinfoCandidate }
continue
}
if (Test-HipinfoIsVenvInternal $hipinfoCandidate) { continue } # venv copy (AMD wheel): not a HIP SDK
substep "[WARN] hipinfo not on PATH -- located via ${hipEnvLabel}: $hipinfoCandidate" "Yellow"
substep " Add '$(Join-Path $hipRoot 'bin')' to your PATH to suppress this warning" "Yellow"
substep " Quick fix: [Environment]::SetEnvironmentVariable('PATH',`$env:PATH+';$(Join-Path $hipRoot 'bin')','User')" "Yellow"
$hipinfoExe = [PSCustomObject]@{ Source = $hipinfoCandidate }
break
}
if ((-not $hipinfoExe) -and $hipMissingLabel) {
substep "[WARN] ${hipMissingLabel}=$hipMissingRoot is set but hipinfo.exe not found at $hipMissingCandidate" "Yellow"
substep " HIP SDK install may be incomplete -- re-install from:" "Yellow"
substep " https://rocm.docs.amd.com/en/latest/deploy/windows/index.html" "Yellow"
}
}
if ($hipinfoExe) {
$HipSdkInstalled = $true # binary found → SDK is installed regardless of device state
try {
$hipOut = & $hipinfoExe.Source 2>&1 | Out-String
if ($hipOut -match "(?i)gcnArchName") {
# hipinfo can crash after printing gcnArchName (#6043).
# Once the arch is printed, keep the ROCm wheel path.
$HasROCm = $true
$_hipAllArches = @([regex]::Matches($hipOut, "(?im)^\s*gcnArchName\s*:\s*(\S+)") | ForEach-Object { ($_.Groups[1].Value -split ':')[0].Trim().ToLower() })
$_hipVisIdx = if ($env:HIP_VISIBLE_DEVICES -match '^\d') { [int]($env:HIP_VISIBLE_DEVICES -split ',')[0] } elseif ($env:ROCR_VISIBLE_DEVICES -match '^\d') { [int]($env:ROCR_VISIBLE_DEVICES -split ',')[0] } else { 0 }
if ($_hipAllArches.Count -gt 0) {
$script:ROCmGfxArch = if ($_hipVisIdx -lt $_hipAllArches.Count) { $_hipAllArches[$_hipVisIdx] } else { $_hipAllArches[0] }
$ROCmGpuLabel = "AMD ROCm ($script:ROCmGfxArch)"
} else {
$ROCmGpuLabel = "AMD ROCm"
}
if ($LASTEXITCODE -ne 0) {
substep "[INFO] hipinfo exited with code $LASTEXITCODE but reported gcnArchName -- treating as ROCm-capable (see #6043)" "Cyan"
}
} elseif ($LASTEXITCODE -ne 0) {
# hipinfo ran but returned a HIP runtime error without any gcnArchName
# output (e.g. "no ROCm-capable device detected"), or crashed before
# printing device info.
$firstLine = ($hipOut -split '\r?\n' | Where-Object { $_.Trim() } | Select-Object -First 1)
substep "[WARN] hipinfo returned a HIP runtime error (exit $LASTEXITCODE)" "Yellow"
substep " $firstLine" "Yellow"
substep " Ensure ROCm drivers are installed: https://rocm.docs.amd.com/en/latest/deploy/windows/index.html" "Yellow"
}
} catch {}
}
# amd-smi fallback: HIP runtime present but hipinfo unavailable (no full HIP SDK).
# 'list' confirms GPU visibility, 'static --asic' extracts the gfx arch hipinfo
# would give. Critical for Strix Halo (gfx1151) and other HIP-runtime-only iGPUs.
#
# BUT on hosts without a working HIP runtime amd-smi elevates a child at runtime,
# popping a UAC/DiskPart prompt RunAsInvoker can't suppress (its manifest is
# asInvoker; even 'amd-smi version' hangs). So only probe when a HIP SDK is present
# (hipinfo found -> un-elevated) or the user opts in; else fall through to WMI name
# inference (enough to pick ROCm wheels + the ROCm llama.cpp prebuilt).
# An explicit opt-out (UNSLOTH_ENABLE_AMD_SMI=0/false/no/off) wins over the HIP-SDK
# heuristic: a HIP SDK binary with a broken runtime can still pop the prompt, so
# $HipSdkInstalled must NOT silently re-enable it.
$amdSmiOptOut = $env:UNSLOTH_ENABLE_AMD_SMI -match '^(?i)(0|false|no|off)$'
$amdSmiAllowed = (-not $amdSmiOptOut) -and ($HipSdkInstalled -or ($env:UNSLOTH_ENABLE_AMD_SMI -match '^(?i)(1|true|yes|on)$'))
if (-not $HasROCm -and $amdSmiAllowed) {
$amdSmiExe = Get-Command "amd-smi" -ErrorAction SilentlyContinue
if ($amdSmiExe) {
try {
$smiOut = Invoke-AmdSmiNoElevate $amdSmiExe.Source @('list')
if ($LASTEXITCODE -eq 0 -and $smiOut -match "(?im)^GPU\s*[:\[]\s*\d") {
$HasROCm = $true
# Attempt 1: newer amd-smi versions embed the gfx arch in list output.
# Collect ALL gfx tokens in output order so that on mixed-arch systems
# we can honour HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES and pick the
# arch for the *runtime-visible* GPU rather than always the first one.
# Do NOT deduplicate: a dual same-arch system (e.g. two gfx1151 APUs)
# must produce a 2-element array so HIP_VISIBLE_DEVICES=1 selects the
# second GPU rather than triggering a false out-of-range warning.
# Note: this mapping assumes amd-smi lists GPUs in the same order as
# HIP enumerates them (both follow PCI bus order in practice); it may
# give the wrong arch when GPU indices are non-contiguous (very rare).
$allGfxArches = @([regex]::Matches($smiOut, '(?i)\b(gfx\d+[a-z]?)\b') |
ForEach-Object { $_.Groups[1].Value.ToLower() })
if ($allGfxArches.Count -gt 0) {
# Resolve which GPU index is runtime-visible. When a single
# integer index is set, use it; fall back to index 0 otherwise
# (comma-separated lists or unset → first GPU, same as before).
$visGpu = if ($env:HIP_VISIBLE_DEVICES) { $env:HIP_VISIBLE_DEVICES }
elseif ($env:ROCR_VISIBLE_DEVICES) { $env:ROCR_VISIBLE_DEVICES }
else { $null }
$gpuIdx = 0
if ($visGpu -match '^\s*(\d+)\s*$') { $gpuIdx = [int]$Matches[1] }
if ($gpuIdx -ge $allGfxArches.Count) {
substep "[WARN] HIP/ROCR_VISIBLE_DEVICES index $gpuIdx is out of range ($($allGfxArches.Count) GPU(s) detected); defaulting to GPU 0 for arch selection" "Yellow"
$gpuIdx = 0
}
$script:ROCmGfxArch = $allGfxArches[$gpuIdx]
$ROCmGpuLabel = "AMD ROCm ($script:ROCmGfxArch)"
} else {
# Attempt 2: 'static --asic' exposes ASIC details on ROCm 6+,
# including the GFX target needed for wheel index selection.
$smiAsicOut = ""
try { $smiAsicOut = Invoke-AmdSmiNoElevate $amdSmiExe.Source @('static','--asic') } catch {}
if ($smiAsicOut -match "(?i)\b(gfx\d+[a-z]?)\b") {
$script:ROCmGfxArch = $Matches[1].ToLower()
$ROCmGpuLabel = "AMD ROCm ($script:ROCmGfxArch)"
} elseif ($smiAsicOut -match "(?im)Market.?Name\s*[:\|]\s*([^\r\n]+)") {
$ROCmGpuLabel = "AMD ROCm ($($Matches[1].Trim()))"
} else {
$ROCmGpuLabel = "AMD ROCm"
}
}
}
} catch {}
}
}
# WMI fallback: AMD GPU in device list but no HIP SDK → guide the user.
# WMI gives a marketing name (e.g. "AMD Radeon 890M") but never a gfx arch.
# $HasROCm is intentionally NOT set here — we cannot confirm ROCm runtime
# support without hipinfo or amd-smi. The name is saved to $ROCmGpuLabel
# so the name-based inference below can still attempt an arch lookup.
if (-not $HasROCm) {
try {
$wmiGpu = Get-WmiObject Win32_VideoController -ErrorAction SilentlyContinue |
Where-Object { $_.Name -match "AMD|Radeon" } |
Select-Object -First 1
if ($wmiGpu) { $ROCmGpuLabel = $wmiGpu.Name }
} catch {}
}
# ── Arch resolution: env-var override → name inference ──────────────────
# Runs after all probes, even when none confirmed a ROCm runtime ($HasROCm false):
# the Adrenalin driver alone runs the per-gfx ROCm llama.cpp prebuilt (bundles its
# own runtime), and all it needs is the gfx arch, inferable from the WMI GPU name.
# Resolving it here lets setup.ps1 forward --rocm-gfx so a GPU llama.cpp is pulled
# instead of CPU. (PyTorch ROCm wheels still require a HIP SDK -- gated on $HasROCm
# below -- so this only affects llama.cpp / inference.)
if (-not $script:ROCmGfxArch) {
# 1. Manual override: set UNSLOTH_ROCM_GFX_ARCH=gfx1151 before running.
if ($env:UNSLOTH_ROCM_GFX_ARCH) {
$script:ROCmGfxArch = $env:UNSLOTH_ROCM_GFX_ARCH.Trim().ToLower()
$ROCmGpuLabel = "AMD ROCm ($script:ROCmGfxArch)"
substep "gfx arch from UNSLOTH_ROCM_GFX_ARCH env override: $script:ROCmGfxArch" "Cyan"
}
# 2. Best-effort name → arch lookup (amd-smi / WMI). Most-specific first,
# first match wins. Covers only arches the ROCm prebuilts support
# (gfx120X/110X/1151/1150/103X); unknown names fall back cleanly to CPU.
elseif ($ROCmGpuLabel) {
$nameArchTable = @(
@{ P = "9070 XT|9080"; A = "gfx1201" } # RDNA 4 (Radeon RX 9070 XT / 9080)
@{ P = "9070|9060"; A = "gfx1200" } # RDNA 4 (Radeon RX 9070 / 9060)
@{ P = "8060S|8050S|8040S|Strix Halo|Ryzen AI Max|AI Max"; A = "gfx1151" } # RDNA 3.5 (Strix Halo: Radeon 8060S/8050S/8040S iGPU, Ryzen AI Max+)
@{ P = "890M|880M|860M|840M|Strix Point|Krackan|HX 37[05]|AI 9 HX|AI 9 36[05]|AI 7 35[05]|AI 5 34[05]|AI 7 PRO 35|AI 5 33"; A = "gfx1150" } # RDNA 3.5 (Strix/Krackan Point: Radeon 890M/880M iGPU, Ryzen AI 9 HX 370/375)
@{ P = "RX 7900|RX 7800|RX 7700(?!S)|PRO W7900|PRO W7800|PRO W7700"; A = "gfx1100" } # RDNA 3 desktop / workstation (Navi 31)
@{ P = "RX 7600|RX 7700S|RX 7650|PRO W7600|PRO W7500|PRO V710"; A = "gfx1102" } # RDNA 3 (Navi 33)
@{ P = "780M|760M|740M|Phoenix|Hawk Point|Z1 Extreme|Z2 Extreme"; A = "gfx1103" } # RDNA 3 iGPU (Phoenix / Hawk Point)
@{ P = "RX 6900|RX 6800|RX 6750|RX 6700|PRO W6800|PRO W6900"; A = "gfx1030" } # RDNA 2 (Navi 21) -- gfx103X family
@{ P = "RX 6650|RX 6600|PRO W6600|PRO W6650"; A = "gfx1032" } # RDNA 2 (Navi 23) -- gfx103X family
@{ P = "RX 6500|RX 6400|RX 6300|PRO W6400|PRO W6500"; A = "gfx1034" } # RDNA 2 (Navi 24) -- gfx103X family
)
foreach ($row in $nameArchTable) {
if ($ROCmGpuLabel -match $row.P) {
$script:ROCmGfxArch = $row.A
$ROCmGpuLabel = "AMD ROCm ($script:ROCmGfxArch)"
substep "gfx arch inferred from GPU name: $script:ROCmGfxArch" "Cyan"
substep "Tip: set UNSLOTH_ROCM_GFX_ARCH=$script:ROCmGfxArch to skip inference next time" "Cyan"
break
}
}
}
}
# Capture ROCm version early for display and wheel selection.
# Run whenever the HIP SDK binary is present, not just when the device is accessible --
# hipconfig --version works even when hipinfo reports no ROCm device (driver issue).
if ($HasROCm -or $HipSdkInstalled) {
$script:ROCmVersion = $null
$hipConfigExe = Get-Command hipconfig -ErrorAction SilentlyContinue
if (-not $hipConfigExe) {
$hipRoot = if ($env:HIP_PATH) { $env:HIP_PATH } elseif ($env:ROCM_PATH) { $env:ROCM_PATH } else { $null }
if ($hipRoot) {
$hipConfigCandidate = Join-Path $hipRoot "bin\hipconfig.exe"
if (Test-Path $hipConfigCandidate) {
$hipConfigEnvLabel = if ($env:HIP_PATH) { "HIP_PATH" } else { "ROCM_PATH" }
substep "[WARN] hipconfig not on PATH -- located via ${hipConfigEnvLabel}: $hipConfigCandidate" "Yellow"
$hipConfigExe = [PSCustomObject]@{ Source = $hipConfigCandidate }
}
}
}
if ($hipConfigExe) {
try {
$hipVerOut = & $hipConfigExe.Source --version 2>&1 | Out-String
if ($LASTEXITCODE -eq 0) {
$hipVerLine = ($hipVerOut -split '\r?\n' | Where-Object { $_.Trim() } | Select-Object -First 1).Trim()
if ($hipVerLine -match '(\d+\.\d+)') {
$script:ROCmVersion = $Matches[1]
$script:ROCmVersionFull = $hipVerLine
}
}
} catch {}
}
if (-not $script:ROCmVersion -and $amdSmiAllowed) {
$amdSmiVer = Get-Command "amd-smi" -ErrorAction SilentlyContinue
if ($amdSmiVer) {
try {
$smiVerOut = Invoke-AmdSmiNoElevate $amdSmiVer.Source @('version')
if ($LASTEXITCODE -eq 0 -and $smiVerOut -match 'ROCm version:\s*(\d+\.\d+)') { $script:ROCmVersion = $Matches[1] }
} catch {}
}
}
}
}
if ($HasNvidiaSmi) {
step "gpu" "NVIDIA GPU detected"
} elseif ($HasROCm) {
step "gpu" $ROCmGpuLabel
$hipSdkPath = if ($env:HIP_PATH) { $env:HIP_PATH } elseif ($env:ROCM_PATH) { $env:ROCM_PATH } else { "on system PATH" }
substep "HIP SDK: $hipSdkPath"
if ($script:ROCmVersionFull) { substep "hipconfig: $script:ROCmVersionFull" }
} elseif ($HipSdkInstalled -and $ROCmGpuLabel) {
# HIP SDK is installed but ROCm can't see the device (driver issue, not SDK issue)
$sdkVer = if ($script:ROCmVersionFull) { " (HIP $script:ROCmVersionFull)" } else { "" }
Write-Host ""
step "gpu" "AMD GPU detected -- not ROCm-accessible$sdkVer" "Yellow"
substep "Detected: $ROCmGpuLabel" "Yellow"
substep "[WARN] HIP SDK is installed but hipinfo reports no ROCm-capable device." "Yellow"
substep " This is a driver issue, not an SDK issue." "Yellow"
substep " Ensure the ROCm compute driver is installed alongside the display driver:" "Yellow"
substep " https://rocm.docs.amd.com/en/latest/deploy/windows/index.html" "Yellow"
} elseif ($script:ROCmGfxArch) {
# Known arch: PyTorch comes from AMD's bundled-runtime ROCm wheels (repo.amd.com),
# which ship their own runtime -- HIP SDK optional (only adds the system toolchain).
Write-Host ""
step "gpu" "AMD ROCm ($script:ROCmGfxArch)" "Cyan"
substep "Detected: $ROCmGpuLabel" "Cyan"
substep "GPU PyTorch uses AMD's bundled-runtime ROCm wheels -- HIP SDK not required (optional)." "Cyan"
Write-Host ""
} elseif ($ROCmGpuLabel) {
Write-Host ""
step "gpu" "AMD GPU detected -- arch unknown" "Yellow"
substep "Detected: $ROCmGpuLabel" "Yellow"
substep "Could not determine the GPU arch (gfx...). Install the HIP SDK or set" "Yellow"
substep "UNSLOTH_ROCM_GFX_ARCH to enable GPU ROCm PyTorch:" "Yellow"
substep "https://rocm.docs.amd.com/en/latest/deploy/windows/index.html" "Yellow"
Write-Host ""
} else {
Write-Host ""
step "gpu" "none (chat-only / GGUF)" "Yellow"
substep "Training and GPU inference require an NVIDIA or AMD ROCm GPU." "Yellow"
Write-Host ""
}
# ============================================
# 1a.5. Windows Long Paths (required for deep node_modules / Python paths)
# ============================================
$LongPathsEnabled = $false
try {
$regVal = Get-ItemProperty -Path "HKLM:\SYSTEM\CurrentControlSet\Control\FileSystem" -Name "LongPathsEnabled" -ErrorAction SilentlyContinue
if ($regVal -and $regVal.LongPathsEnabled -eq 1) {
$LongPathsEnabled = $true
}
} catch {}
if ($LongPathsEnabled) {
step "long paths" "enabled"
} else {
Write-Host "Windows Long Paths not enabled (required for Triton compilation and deep dependency paths)." -ForegroundColor Yellow
Write-Host " Requesting admin access to fix..." -ForegroundColor Yellow
try {
# Spawn an elevated process to set the registry key (triggers UAC prompt)
$proc = Start-Process -FilePath "reg.exe" `
-ArgumentList 'add "HKLM\SYSTEM\CurrentControlSet\Control\FileSystem" /v LongPathsEnabled /t REG_DWORD /d 1 /f' `
-Verb RunAs -Wait -PassThru -ErrorAction Stop
if ($proc.ExitCode -eq 0) {
$LongPathsEnabled = $true
step "long paths" "enabled (via UAC)"
} else {
step "long paths" "failed to enable (exit code: $($proc.ExitCode))" "Yellow"
}
} catch {
step "long paths" "could not enable (UAC declined/unavailable)" "Yellow"
Write-Host " Run this manually in an Admin terminal:" -ForegroundColor Yellow
Write-Host ' reg add "HKLM\SYSTEM\CurrentControlSet\Control\FileSystem" /v LongPathsEnabled /t REG_DWORD /d 1 /f' -ForegroundColor Cyan
}
}
# ============================================
# 1b. Git (required by pip for git+https:// deps and by npm)
# ============================================
$HasGit = $null -ne (Get-Command git -ErrorAction SilentlyContinue)
if (-not $HasGit) {
Write-Host "Git not found -- installing via winget..." -ForegroundColor Yellow
$HasWinget = $null -ne (Get-Command winget -ErrorAction SilentlyContinue)
if ($HasWinget) {
try {
Invoke-SetupCommand { winget install Git.Git --source winget --accept-package-agreements --accept-source-agreements } | Out-Null
Refresh-Environment
$HasGit = $null -ne (Get-Command git -ErrorAction SilentlyContinue)
} catch { }
}
if (-not $HasGit) {
Write-Host "[ERROR] Git is required but could not be installed automatically." -ForegroundColor Red
Write-Host " Install Git from https://git-scm.com/download/win and re-run." -ForegroundColor Red
exit 1
}
step "git" "$(git --version)"
} else {
step "git" "$(git --version)"
}
# ============================================
# 1b.5. Visual C++ Redistributable (runtime for the prebuilt llama.cpp + PyTorch)
# ============================================
# Runtime dep, not a build tool: the prebuilt llama-server and PyTorch load it.
Ensure-VCRedist
# ============================================
# 1c. CMake (only needed for a llama.cpp SOURCE build -- detection only)
# ============================================
# Detection only: the prebuilt path needs no compiler, so do not install or exit
# here. Ensure-BuildToolsForLlamaSourceBuild installs CMake if a source build runs.
$HasCmake = $null -ne (Get-Command cmake -ErrorAction SilentlyContinue)
if ($HasCmake) {
step "cmake" "$(cmake --version | Select-Object -First 1)"
} else {
step "cmake" "not detected (only needed if a llama.cpp source build is required)" "Yellow"
}
# ============================================
# 1d. Visual Studio Build Tools (only needed for a llama.cpp SOURCE build -- detection only)
# ============================================
# Detection only: detect VS for a possible source build, but never install or exit
# here. Install is deferred to Ensure-BuildToolsForLlamaSourceBuild.
$CmakeGenerator = $null
$VsInstallPath = $null
$vsResult = Find-VsBuildTools
if ($vsResult) {
$CmakeGenerator = $vsResult.Generator
$VsInstallPath = $vsResult.InstallPath
step "vs" "$CmakeGenerator ($($vsResult.Source)) (only used if a source build is needed)"
if ($vsResult.ClExe) { substep "cl.exe: $($vsResult.ClExe)" }
} else {
step "vs" "not detected (only needed if a llama.cpp source build is required)" "Yellow"
}
# ============================================
# 1e. CUDA Toolkit (nvcc for llama.cpp build + env vars)
# ============================================
# Defined here but invoked lazily right before a Phase 4 source build; the
# prebuilt llama.cpp path needs no local toolkit. With -RequireOrExit a source
# build is committed, so hard-fail if no driver-compatible toolkit can be found
# or installed. Without it, detection is best-effort and only sets the flag.
function Resolve-CudaToolkit {
param([switch]$RequireOrExit)
# Toolkit major must be <= the driver's max CUDA major (nvidia-smi "CUDA Version: X.Y");
# a newer-major toolkit fails at runtime ("ggml_cuda_init: failed to initialize CUDA").
$DriverMaxCuda = $null
try {
# Bounded: source-build toolkit resolution must not hang on a wedged smi.
# test_resolve_cuda_toolkit.ps1 extracts this function alone into a child
# pwsh (no Invoke-NvidiaSmiBounded in scope) and stubs nvidia-smi with a
# .ps1 script, so fall back to direct invocation when the bounded runner
# is unavailable; production setup.ps1 always has it defined.
$smiOut = if (Get-Command Invoke-NvidiaSmiBounded -ErrorAction SilentlyContinue) {
Invoke-NvidiaSmiBounded $NvidiaSmiExe
} else {
& $NvidiaSmiExe 2>&1 | Out-String
}
# Newer drivers report "CUDA UMD Version: X.Y" instead of "CUDA Version: X.Y"; accept both.
if ($smiOut -match "CUDA(?: UMD)? Version:\s+([\d]+)\.([\d]+)") {
$DriverMaxCuda = "$($Matches[1]).$($Matches[2])"
substep "driver supports up to CUDA $DriverMaxCuda"
}
} catch {}
# Detect compute capability early so we can validate toolkit support
$CudaArch = Get-CudaComputeCapability
if ($CudaArch) {
substep "GPU Compute Capability = $($CudaArch.Insert($CudaArch.Length-1, '.')) (sm_$CudaArch)"
}
# -- Find a toolkit that's compatible with the driver AND the GPU --
# Strategy: prefer the toolkit at CUDA_PATH (user's existing setup) if it's
# compatible with the driver AND supports the GPU architecture. Only fall back
# to scanning side-by-side installs if CUDA_PATH is missing, points to an
# incompatible version, or can't compile for the GPU. This avoids
# header/binary mismatches when multiple toolkits are installed.
$IncompatibleToolkit = $null
$NvccPath = $null
if ($DriverMaxCuda) {
$drMajorCuda = [int]$DriverMaxCuda.Split('.')[0]
# --- Step 1: Check existing CUDA_PATH first ---
$existingCudaPath = [Environment]::GetEnvironmentVariable('CUDA_PATH', 'Machine')
if (-not $existingCudaPath) {
$existingCudaPath = [Environment]::GetEnvironmentVariable('CUDA_PATH', 'User')
}
if ($existingCudaPath -and (Test-Path (Join-Path $existingCudaPath 'bin\nvcc.exe'))) {
$candidateNvcc = Join-Path $existingCudaPath 'bin\nvcc.exe'
$verOut = & $candidateNvcc --version 2>&1 | Out-String
if ($verOut -match 'release\s+(\d+)\.(\d+)') {
$tkMaj = [int]$Matches[1]; $tkMin = [int]$Matches[2]
$isCompat = ($tkMaj -le $drMajorCuda)
if ($isCompat) {
# Also verify the toolkit supports our GPU architecture
$archOk = $true
if ($CudaArch) {
$archOk = Test-NvccArchSupport -NvccExe $candidateNvcc -Arch $CudaArch
if (-not $archOk) {
substep "CUDA_PATH toolkit (CUDA $tkMaj.$tkMin) does not support GPU arch sm_$CudaArch" "Yellow"
substep "Looking for a newer toolkit..." "Yellow"
}
}
if ($archOk) {
$NvccPath = $candidateNvcc
substep "using existing CUDA Toolkit at CUDA_PATH (nvcc: $NvccPath)"
}
} else {
substep "CUDA_PATH ($existingCudaPath) has CUDA $tkMaj.$tkMin with major $tkMaj, which exceeds driver CUDA major $drMajorCuda ($DriverMaxCuda)" "Yellow"
}
}
}
# --- Step 2: Fall back to scanning side-by-side installs ---
if (-not $NvccPath) {
$NvccPath = Find-Nvcc -MaxVersion $DriverMaxCuda
if ($NvccPath) {
substep "found compatible CUDA Toolkit (nvcc: $NvccPath)"
if ($existingCudaPath) {
$selectedRoot = Split-Path (Split-Path $NvccPath -Parent) -Parent
if ($existingCudaPath.TrimEnd('\') -ne $selectedRoot.TrimEnd('\')) {
substep "overriding CUDA_PATH from $existingCudaPath to $selectedRoot" "Yellow"
}
}
} else {
# No side-by-side match: a major-compatible toolkit may still be on
# PATH/CUDA_PATH/a custom dir; use it, else record it as too-new.
$AnyNvcc = Find-Nvcc
if ($AnyNvcc) {
$NvccOut = & $AnyNvcc --version 2>&1 | Out-String
if ($NvccOut -match "release\s+(\d+)\.(\d+)") {
$tkMaj = [int]$Matches[1]; $tkMin = [int]$Matches[2]
if ($tkMaj -le $drMajorCuda) {
$NvccPath = $AnyNvcc
substep "found compatible CUDA Toolkit (nvcc: $NvccPath)"
} else {
$IncompatibleToolkit = "$tkMaj.$tkMin"
}
}
}
}
}
} else {
$NvccPath = Find-Nvcc
}
# A newer-major toolkit blocked by the driver: explain the mismatch.
if (-not $NvccPath -and $IncompatibleToolkit) {
Write-CudaDriverToolkitMismatch -ToolkitVersion $IncompatibleToolkit -DriverMaxCuda $DriverMaxCuda
if (-not $RequireOrExit) {
$script:CudaToolkitReady = $false
return
}
# Reached only by a source build (forced, or after a prebuilt-install failure);
# with no compatible toolkit it must fail (setup.sh degrades to CPU instead).
Write-Host "" -ForegroundColor Red
Write-Host "========================================================================" -ForegroundColor Red
Write-Host "[ERROR] CUDA source build cannot use the installed toolkit with this driver." -ForegroundColor Red
Write-Host "========================================================================" -ForegroundColor Red
exit 1
}
# -- No toolkit at all: install via winget (only when a source build needs it) --
if (-not $NvccPath -and $RequireOrExit) {
Write-Host "CUDA toolkit (nvcc) not found -- installing via winget..." -ForegroundColor Yellow
$HasWinget = $null -ne (Get-Command winget -ErrorAction SilentlyContinue)
if ($HasWinget) {
if ($DriverMaxCuda) {
# Query winget for available CUDA Toolkit versions
$drMajor = [int]$DriverMaxCuda.Split('.')[0]
$AvailableVersions = @()
try {
$rawOutput = winget show Nvidia.CUDA --versions --source winget --accept-source-agreements 2>&1 | Out-String
# Parse version lines (e.g. "12.6", "12.5", "11.8")
foreach ($line in $rawOutput -split "`n") {
$line = $line.Trim()
if ($line -match '^\d+\.\d+') {
$AvailableVersions += $line
}
}
} catch {}
# Filter to compatible major versions and pick the highest
$BestVersion = $null
foreach ($ver in $AvailableVersions) {
$parts = $ver.Split('.')
$vMajor = [int]$parts[0]
if ($vMajor -le $drMajor) {
$BestVersion = $ver
break # list is descending, first match is highest compatible
}
}
if ($BestVersion) {
substep "Installing CUDA Toolkit $BestVersion via winget..."
$prevEAPCuda = $ErrorActionPreference
$ErrorActionPreference = "Continue"
Invoke-SetupCommand { winget install --id=Nvidia.CUDA --version=$BestVersion -e --source winget --accept-package-agreements --accept-source-agreements } | Out-Null
$ErrorActionPreference = $prevEAPCuda
Refresh-Environment
$NvccPath = Find-Nvcc -MaxVersion $DriverMaxCuda
if ($NvccPath) {
substep "CUDA Toolkit $BestVersion installed (nvcc: $NvccPath)"
}
} else {
substep "no compatible CUDA Toolkit version found in winget (need CUDA major <= $drMajor)" "Yellow"
}
} else {
substep "Installing CUDA Toolkit (latest) via winget..."
winget install --id=Nvidia.CUDA -e --source winget --accept-package-agreements --accept-source-agreements
Refresh-Environment
$NvccPath = Find-Nvcc
if ($NvccPath) {
substep "CUDA Toolkit installed (nvcc: $NvccPath)"
}
}
}
}
if (-not $NvccPath) {
if (-not $RequireOrExit) {
substep "no driver-compatible CUDA Toolkit found -- skipping; prebuilt llama.cpp needs no local toolkit" "Yellow"
$script:CudaToolkitReady = $false
return
}
Write-Host "[ERROR] CUDA Toolkit (nvcc) is required but could not be found or installed." -ForegroundColor Red
if ($DriverMaxCuda) {
Write-Host " Install a CUDA Toolkit with major version $($DriverMaxCuda.Split('.')[0]) from https://developer.nvidia.com/cuda-toolkit-archive" -ForegroundColor Yellow
} else {
Write-Host " Install CUDA Toolkit from https://developer.nvidia.com/cuda-downloads" -ForegroundColor Yellow
}
exit 1
}
# -- Set CUDA env vars so cmake AND MSBuild can find the toolkit --
$CudaToolkitRoot = Split-Path (Split-Path $NvccPath -Parent) -Parent
# CUDA_PATH: used by cmake's find_package(CUDAToolkit)
[Environment]::SetEnvironmentVariable('CUDA_PATH', $CudaToolkitRoot, 'Process')
# CudaToolkitDir: the MSBuild property that CUDA .targets checks directly
# Trailing backslash required -- the .targets file appends subpaths to it
[Environment]::SetEnvironmentVariable('CudaToolkitDir', "$CudaToolkitRoot\", 'Process')
# Always persist CUDA_PATH to User registry so the compatible toolkit is used
# in future sessions (overwrites any existing value pointing to a newer, incompatible version)
[Environment]::SetEnvironmentVariable('CUDA_PATH', $CudaToolkitRoot, 'User')
substep "Persisted CUDA_PATH=$CudaToolkitRoot to user environment"
# Clear all versioned CUDA_PATH_V* env vars in this process to prevent
# cmake/MSBuild from discovering a conflicting CUDA installation.
$cudaPathVars = @([Environment]::GetEnvironmentVariables('Process').Keys | Where-Object { $_ -match '^CUDA_PATH_V' })
foreach ($v in $cudaPathVars) {
[Environment]::SetEnvironmentVariable($v, $null, 'Process')
}
# Set only the versioned var matching the selected toolkit (e.g. CUDA_PATH_V13_0)
$tkDirName = Split-Path $CudaToolkitRoot -Leaf
if ($tkDirName -match '^v(\d+)\.(\d+)') {
$cudaPathVerVar = "CUDA_PATH_V$($Matches[1])_$($Matches[2])"
[Environment]::SetEnvironmentVariable($cudaPathVerVar, $CudaToolkitRoot, 'Process')
substep "Set $cudaPathVerVar (cleared other CUDA_PATH_V* vars)"
}
# Ensure nvcc's bin dir is on PATH for this process
$nvccBinDir = Split-Path $NvccPath -Parent
if ($env:PATH -notlike "*$nvccBinDir*") {
[Environment]::SetEnvironmentVariable('PATH', "$nvccBinDir;$env:PATH", 'Process')
}
# Persist nvcc bin dir (Prepend so the driver-compatible toolkit wins).
if (Add-ToUserPath -Directory $nvccBinDir -Position 'Prepend') {
substep "Persisted CUDA bin dir to user PATH"
}
# -- Ensure CUDA ↔ Visual Studio integration files exist --
# When CUDA is installed before VS Build Tools (or VS is reinstalled after CUDA),
# the MSBuild .targets/.props files that let VS compile .cu files are missing.
# cmake fails with "No CUDA toolset found". Fix: copy from CUDA extras dir.
if ($VsInstallPath -and $CudaToolkitRoot) {
$vsCustomizations = Get-VcBuildCustomizationsDir -VsInstallPath $VsInstallPath -Generator $CmakeGenerator
$cudaExtras = Join-Path $CudaToolkitRoot "extras\visual_studio_integration\MSBuildExtensions"
if ((Test-Path $cudaExtras) -and (Test-Path $vsCustomizations)) {
$hasTargets = Get-ChildItem $vsCustomizations -Filter "CUDA *.targets" -ErrorAction SilentlyContinue
if (-not $hasTargets) {
substep "CUDA VS integration missing -- copying .targets files..." "Yellow"
try {
Copy-Item "$cudaExtras\*" $vsCustomizations -Force -ErrorAction Stop
substep "CUDA VS integration files installed"
} catch {
# Direct copy failed (needs admin). Try elevated copy via Start-Process.
try {
$copyCmd = "Copy-Item '$cudaExtras\*' '$vsCustomizations' -Force"
Start-Process powershell -ArgumentList "-NoProfile -Command $copyCmd" -Verb RunAs -Wait -ErrorAction Stop
$hasTargetsRetry = Get-ChildItem $vsCustomizations -Filter "CUDA *.targets" -ErrorAction SilentlyContinue
if ($hasTargetsRetry) {
substep "CUDA VS integration files installed (elevated)"
} else {
throw "Copy did not produce .targets files"
}
} catch {
substep "could not copy CUDA VS integration files" "Yellow"
substep "The llama.cpp build may fail with 'No CUDA toolset found'." "Yellow"
substep "Manual fix: copy contents of" "Yellow"
substep "$cudaExtras"
substep "into:" "Yellow"
substep "$vsCustomizations"
}
}
}
}
}
step "cuda" $NvccPath
substep "CUDA_PATH = $CudaToolkitRoot"
substep "CudaToolkitDir = $CudaToolkitRoot\"
# $CudaArch was detected earlier (before toolkit selection) so it could
# influence which toolkit we picked. Just log the final state here.
if (-not $CudaArch) {
substep "could not detect compute capability -- cmake will use defaults" "Yellow"
}
# Publish the resolved toolkit to script scope for the Phase 4 build.
$script:NvccPath = $NvccPath
$script:CudaToolkitRoot = $CudaToolkitRoot
$script:CudaArch = $CudaArch
$script:CudaToolkitReady = $true
}
if ($HasROCm) {
$rocmVerLabel = if ($script:ROCmVersionFull) { "ROCm $script:ROCmVersionFull" } elseif ($script:ROCmVersion) { "ROCm $script:ROCmVersion" } else { "ROCm (version unknown)" }
step "rocm" $rocmVerLabel
} elseif ($script:ROCmGfxArch) {
# GPU training/inference works via AMD's bundled-runtime ROCm PyTorch wheels;
# the HIP SDK is optional (only the system ROCm toolchain).
step "rocm" "GPU via bundled ROCm wheels ($script:ROCmGfxArch) -- HIP SDK optional" "Cyan"
} elseif ($ROCmGpuLabel) {
step "rocm" "AMD GPU detected -- arch unknown; HIP SDK not found" "Yellow"
}
# ============================================
# 1f. Node.js / npm (skip if pip-installed or Tauri -- only needed for frontend build)
# ============================================
# Frontend and OXC share this Node floor. The helper returns:
# system | bundled | skip.
function Get-NodeDecision {
param(
[string]$NodeVersion, # `node -v` output, e.g. v22.17.1 (or empty)
[string]$NpmVersion, # `npm -v` output, e.g. 10.9.2 (or empty)
[string]$SkipInstall # "1" => never auto-install
)
$node = ($NodeVersion -replace '^v', '').Trim()
$npm = "$NpmVersion".Trim()
if ($node -match '^\d+\.\d+' -and $npm -match '^\d+') {
$nodeMajor = [int]($node.Split('.')[0])
$nodeMinor = [int]($node.Split('.')[1])
$npmMajor = [int]($npm.Split('.')[0])
$nodeOk = ($nodeMajor -eq 20 -and $nodeMinor -ge 19) -or
($nodeMajor -eq 22 -and $nodeMinor -ge 12) -or
($nodeMajor -ge 23)
if ($nodeOk -and $npmMajor -ge 11) { return "system" }
}
if ($SkipInstall -eq "1") { return "skip" }
return "bundled"
}
$SkipFrontend = ($env:SKIP_STUDIO_FRONTEND -eq "1")
$NodeOverride = $null
$NodeParent = $null
$NodeDir = $null
$SysNodeVersion = ""
$SysNpmVersion = ""
$NodeSource = $null
if (-not $IsPipInstall) {
# Put Node beside the Unsloth root. OXC can still need npm when the
# frontend build is skipped.
if (-not [string]::IsNullOrWhiteSpace($env:UNSLOTH_STUDIO_HOME)) { $NodeOverride = $env:UNSLOTH_STUDIO_HOME.Trim() }
elseif (-not [string]::IsNullOrWhiteSpace($env:STUDIO_HOME)) { $NodeOverride = $env:STUDIO_HOME.Trim() }
if ($NodeOverride) {
if ($NodeOverride -eq "~") {
$NodeOverride = $env:USERPROFILE
} elseif ($NodeOverride -like "~/*" -or $NodeOverride -like "~\*") {
$NodeOverride = (Join-Path $env:USERPROFILE $NodeOverride.Substring(1).TrimStart('/', '\'))
}
if (-not (Test-Path -LiteralPath $NodeOverride -PathType Container)) {
Write-Host "ERROR: UNSLOTH_STUDIO_HOME/STUDIO_HOME=$NodeOverride does not exist." -ForegroundColor Red
Write-Host " Run install.ps1 to create the install root before 'unsloth studio update'." -ForegroundColor Red
exit 1
}
$NodeParent = (Resolve-Path -LiteralPath $NodeOverride).Path
# An override pointing at the legacy default maps to the legacy sibling
# ~/.unsloth/node (what the runtime resolver and setup.sh use), not <root>/node.
$_legacyStudio = Join-Path $env:USERPROFILE ".unsloth\studio"
if (Test-Path -LiteralPath $_legacyStudio -PathType Container) {
$_legacyStudio = (Resolve-Path -LiteralPath $_legacyStudio).Path
}
if ($NodeParent -eq $_legacyStudio) {
$NodeParent = Join-Path $env:USERPROFILE ".unsloth"
$NodeOverride = $null
}
} else {
$NodeParent = Join-Path $env:USERPROFILE ".unsloth"
}
$NodeDir = Join-Path $NodeParent "node"
# Probe system node/npm without letting a missing/broken command abort setup.
# Under $ErrorActionPreference = "Stop" a bare `node -v` for an absent node
# throws a terminating error `2>$null` cannot swallow, and a present-but-broken
# shim throws too. Guard with Get-Command (node/npm independently) + try/catch;
# empty version => Get-NodeDecision returns "bundled".
$SysNodeVersion = try { if (Get-Command node -ErrorAction SilentlyContinue) { (node -v 2>$null) } else { "" } } catch { "" }
$SysNpmVersion = try { if (Get-Command npm -ErrorAction SilentlyContinue) { (npm -v 2>$null) } else { "" } } catch { "" }
$NodeSource = Get-NodeDecision -NodeVersion "$SysNodeVersion" -NpmVersion "$SysNpmVersion" -SkipInstall "$($env:UNSLOTH_SKIP_NODE_INSTALL)"
}
if ($IsPipInstall) {
step "frontend" "bundled (pip install)"
} elseif ($SkipFrontend) {
step "frontend" "bundled (Tauri)"
} else {
# Stale npm used to trigger system Node changes. Keep this process-local
# and provision only when the build or OXC needs Node.
if ($NodeSource -eq "system") {
substep "Node $SysNodeVersion and npm $SysNpmVersion already meet requirements (system)."
} elseif ($NodeSource -eq "bundled") {
substep "Node='$SysNodeVersion' npm='$SysNpmVersion' unsuitable; will use an isolated Node (system left untouched)."
} else {
substep "Node='$SysNodeVersion' npm='$SysNpmVersion' unsuitable and UNSLOTH_SKIP_NODE_INSTALL set; frontend build will be skipped." "Yellow"
}
}
# Conda CPython ships modified DLL search paths that break torch's c10.dll
# loading on Windows; a venv made from conda Python inherits its base_prefix,
# so check the executable path AND sys.base_prefix.
$CondaSkipPattern = '(?i)(conda|miniconda|anaconda|miniforge|mambaforge)'
function Test-IsConda {
param([string]$Exe)
if ($Exe -match $CondaSkipPattern) { return $true }
try {
$basePrefix = (& $Exe -c "import sys; print(sys.base_prefix)" 2>$null | Out-String).Trim()
if ($basePrefix -match $CondaSkipPattern) { return $true }
} catch { }
return $false
}
# 1g. Python (>= 3.11 and < 3.14). Prefer the interpreter install.ps1 already
# resolved and built the venv with (UNSLOTH_SETUP_PYTHON), or the existing
# venv python, before re-probing a system where a 3.14 or a WindowsApps stub
# ahead on PATH would trip the gate. setup.ps1 only updates packages in that
# venv, so the handoff is safe to reuse once validated.
function Resolve-ReusedSetupPython {
if (-not [string]::IsNullOrWhiteSpace($env:UNSLOTH_SETUP_PYTHON) -and
(Test-Path -LiteralPath $env:UNSLOTH_SETUP_PYTHON)) {
return $env:UNSLOTH_SETUP_PYTHON
}
# Standalone `unsloth studio setup/update` (install.ps1 did not run): derive
# the venv python from the studio root, mirroring the resolver below.
$root = if (-not [string]::IsNullOrWhiteSpace($env:UNSLOTH_STUDIO_HOME)) { $env:UNSLOTH_STUDIO_HOME.Trim() }
elseif (-not [string]::IsNullOrWhiteSpace($env:STUDIO_HOME)) { $env:STUDIO_HOME.Trim() }
else { Join-Path $env:USERPROFILE ".unsloth\studio" }
if ($root -eq "~") {
# Join-Path with an empty child throws on Windows PowerShell 5.1.
$root = $env:USERPROFILE
} elseif ($root -like "~/*" -or $root -like "~\*") {
$root = Join-Path $env:USERPROFILE $root.Substring(1).TrimStart('/', '\')
}
$venvPy = Join-Path $root "unsloth_studio\Scripts\python.exe"
if (Test-Path -LiteralPath $venvPy) { return $venvPy }
return $null
}
$ReusedSetupPython = Resolve-ReusedSetupPython
$HasPython = $null -ne (Get-Command python -ErrorAction SilentlyContinue)
$PythonOk = $false
$DetectedPyVer = $null
function Get-CompatiblePythonVersion {
param([string]$PythonExe)
try {
$out = & $PythonExe --version 2>&1 | Out-String
if ($out -match 'Python (3\.(11|12|13)(\.\d+)?)') {
return $Matches[1]
}
} catch { }
return $null
}
function Add-PythonDirToProcessPath {
param([string]$PythonExe)
try {
if ($PythonExe -and (Test-Path -LiteralPath $PythonExe)) {
$resolvedDir = Split-Path -Parent $PythonExe
$alreadyOnPath = ($env:PATH -split ';' | Where-Object { $_.TrimEnd('\') -ieq $resolvedDir.TrimEnd('\') }).Count -gt 0
if (-not $alreadyOnPath) {
$env:PATH = "$resolvedDir;$env:PATH"
}
$script:HasPython = $true
}
} catch { }
}
# Reuse the install.ps1 / venv interpreter before any system probe.
$ValidatedSetupPython = $null
if ($ReusedSetupPython) {
$_reusedVer = Get-CompatiblePythonVersion $ReusedSetupPython
if ($_reusedVer -and -not (Test-IsConda $ReusedSetupPython)) {
$DetectedPyVer = $_reusedVer
Add-PythonDirToProcessPath $ReusedSetupPython
$PythonOk = $true
$ValidatedSetupPython = $ReusedSetupPython
}
}
# Fall back to every py.exe on PATH (all-users and per-user launchers can both
# register). -All is required: Windows PowerShell 5.1 returns only the first
# launcher without it, and the PowerShell 7 multi-match array breaks the call
# operator if used directly.
$PyLaunchers = if ($PythonOk) { @() } else { @(Get-Command py -All -CommandType Application -ErrorAction SilentlyContinue) }
foreach ($PyLauncher in $PyLaunchers) {
if ($PyLauncher.Source -match $CondaSkipPattern) { continue }
foreach ($minor in @("3.13", "3.12", "3.11")) {
try {
$out = & $PyLauncher.Source "-$minor" --version 2>&1 | Out-String
if ($out -match 'Python (3\.\d+\.\d+)') {
$DetectedPyVer = $Matches[1]
# Make `python` resolvable for the rest of setup. Without this,
# py-launcher-only installs (no python.exe on PATH) pass the gate
# and then crash on the first bare `python` call below.
try {
$resolvedExe = (& $PyLauncher.Source "-$minor" -c "import sys; print(sys.executable)" 2>$null | Select-Object -First 1)
if ($resolvedExe -and (Test-Path $resolvedExe)) {
Add-PythonDirToProcessPath $resolvedExe
}
} catch { }
$PythonOk = $true
break
}
} catch { }
}
if ($PythonOk) { break }
}
if (-not $PythonOk -and $HasPython) {
$PyVer = python --version 2>&1
if ($PyVer -match "(\d+)\.(\d+)") {
$PyMajor = [int]$Matches[1]; $PyMinor = [int]$Matches[2]
if ($PyMajor -eq 3 -and $PyMinor -ge 11 -and $PyMinor -lt 14) {
$DetectedPyVer = "$PyMajor.$PyMinor"
$PythonOk = $true
}
}
}
if ($PythonOk) {
substep "Python $DetectedPyVer"
} elseif (-not $HasPython) {
# No `python` on PATH (and py.exe either absent or only had unsupported
# minors). Try winget as before -- gating on $HasPython alone, not also
# on $PyLauncher, so a launcher-only install with just 3.14 still gets
# an automatic 3.12 install instead of a hard error.
Write-Host "Python 3.11-3.13 not found -- installing Python 3.12 via winget..." -ForegroundColor Yellow
$HasWinget = $null -ne (Get-Command winget -ErrorAction SilentlyContinue)
if ($HasWinget) {
winget install -e --id Python.Python.3.12 --source winget --accept-package-agreements --accept-source-agreements
Refresh-Environment
}
$HasPython = $null -ne (Get-Command python -ErrorAction SilentlyContinue)
if (-not $HasPython) {
Write-Host "[ERROR] Python could not be installed automatically." -ForegroundColor Red
Write-Host " Install Python 3.12 from https://python.org/downloads/" -ForegroundColor Yellow
exit 1
}
step "python" "$(python --version 2>&1)"
$PythonOk = $true
} else {
# python.exe is on PATH but its version is unsupported, and py.exe (if
# present) had no supported minor either.
Write-Host "[ERROR] No supported Python (3.11-3.13) found on this system." -ForegroundColor Red
Write-Host " py.exe could not locate -3.11/-3.12/-3.13 and `python` on PATH is unsupported." -ForegroundColor Yellow
Write-Host " Install Python 3.12 from https://python.org/downloads/" -ForegroundColor Yellow
exit 1
}
# Add user-scheme Python Scripts dir to PATH (nt_user only, no venv fallback).
$ScriptsDir = python -c "import os, sysconfig; p = sysconfig.get_path('scripts', 'nt_user'); print(p if os.path.exists(p) else '')"
if ($LASTEXITCODE -eq 0 -and $ScriptsDir -and (Test-Path $ScriptsDir)) {
# Append (not Prepend) -- this dir has other pip scripts; shim handles unsloth.
if (Add-ToUserPath -Directory $ScriptsDir) {
# Also add to current process so it's available immediately
$ProcessPathEntries = $env:PATH.Split(';')
if (-not ($ProcessPathEntries | Where-Object { $_.TrimEnd('\') -eq $ScriptsDir })) {
$env:PATH = "$ScriptsDir;$env:PATH"
}
substep "Persisted Python Scripts dir to user PATH: $ScriptsDir"
}
}
Write-Host ""
step "system" "prerequisites ready"
Write-Host ""
# ==========================================================================
# PHASE 2: Frontend build (skip if pip-installed -- already bundled)
# ==========================================================================
$DistDir = Join-Path $FrontendDir "dist"
# Skip build if dist/ exists and no tracked input is newer than dist/.
# Checks src/, public/, package.json, config files -- not just src/.
$NeedFrontendBuild = $true
if ($IsPipInstall) {
$NeedFrontendBuild = $false
step "frontend" "bundled (pip install)"
} elseif ($SkipFrontend) {
$NeedFrontendBuild = $false
step "frontend" "bundled (Tauri)"
} elseif (Test-Path $DistDir) {
$DistTime = (Get-Item $DistDir).LastWriteTime
$NewerFile = $null
# Check src/ and public/ recursively (probe paths directly, not via -Include)
foreach ($subDir in @("src", "public")) {
$subPath = Join-Path $FrontendDir $subDir
if (Test-Path $subPath) {
$NewerFile = Get-ChildItem -Path $subPath -Recurse -File -ErrorAction SilentlyContinue |
Where-Object { $_.LastWriteTime -gt $DistTime } | Select-Object -First 1
if ($NewerFile) { break }
}
}
# Also check all top-level files (package.json, vite.config.ts, index.html, etc.)
if (-not $NewerFile) {
$NewerFile = Get-ChildItem -Path $FrontendDir -File -ErrorAction SilentlyContinue |
Where-Object { $_.Name -ne "bun.lock" -and $_.LastWriteTime -gt $DistTime } |
Select-Object -First 1
}
if (-not $NewerFile) {
$NeedFrontendBuild = $false
step "frontend" "up to date"
} else {
substep "Frontend source changed since last build -- rebuilding..." "Yellow"
}
}
# Provision Node when the frontend build OR the OXC runtime install needs it (the
# OXC `npm install` runs whenever its dir exists, regardless of dist staleness);
# never eagerly. System Node is used read-only; the isolated one is ours.
$NeedNodeForSetup = (-not $IsPipInstall) -and ($NeedFrontendBuild -or (Test-Path $OxcValidatorDir))
if ($NeedNodeForSetup) {
if ($NodeSource -eq "skip") {
if ($NeedFrontendBuild) {
step "frontend" "skipped (no suitable Node; system left untouched)" "Yellow"
}
$NeedFrontendBuild = $false
substep "found Node='$SysNodeVersion' npm='$SysNpmVersion'; Unsloth needs Node >=20.19/22.12/23 and npm >= 11" "Yellow"
substep "install a suitable Node + npm, or unset UNSLOTH_SKIP_NODE_INSTALL to let Unsloth manage an isolated Node" "Yellow"
} elseif ($NodeSource -eq "bundled") {
New-Item -ItemType Directory -Force -Path $NodeParent -ErrorAction SilentlyContinue | Out-Null
# Minimal ownership guard for a custom-home dir (the full Unsloth-owned
# helpers are defined later); never os.replace over a user-owned dir.
if ($NodeOverride -and (Test-Path -LiteralPath $NodeDir -PathType Container)) {
$nodeOwnedMarker = Join-Path $NodeDir ".unsloth-studio-owned"
$nodeMeta = Join-Path $NodeDir "UNSLOTH_NODE_PREBUILT_INFO.json"
if (-not (Test-Path -LiteralPath $nodeOwnedMarker) -and -not (Test-Path -LiteralPath $nodeMeta)) {
Write-Host "[ERROR] $NodeDir already exists and is not an Unsloth-owned Node install." -ForegroundColor Red
Write-Host " Move it aside or choose an empty UNSLOTH_STUDIO_HOME before re-running." -ForegroundColor Yellow
exit 1
}
}
substep "installing isolated Node (system Node/npm left untouched)..."
# The main Python resolver runs later; bare `python` may be a Store stub or
# absent this early, so prefer the validated handed-off/venv Python.
$NodeInstallPython = if ($ValidatedSetupPython) { $ValidatedSetupPython } else { "python" }
$nodeOut = & $NodeInstallPython "$PSScriptRoot\install_node_prebuilt.py" --install-dir $NodeDir 2>&1 | Out-String
$nodeExit = $LASTEXITCODE
if ($nodeExit -eq 3) {
Write-Host $nodeOut -ForegroundColor DarkGray
step "node" "install blocked by another active Unsloth install" "Red"
exit 3
} elseif ($nodeExit -ne 0) {
Write-Host $nodeOut -ForegroundColor DarkGray
Write-Host "[ERROR] Could not install an isolated Node automatically." -ForegroundColor Red
Write-Host " Install Node >= 20.19 (with npm >= 11) from https://nodejs.org/ and re-run, or check your network." -ForegroundColor Yellow
exit 1
}
if ($NodeOverride -and (Test-Path -LiteralPath $NodeDir -PathType Container)) {
New-Item -ItemType File -Force -Path (Join-Path $NodeDir ".unsloth-studio-owned") -ErrorAction SilentlyContinue | Out-Null
}
# Windows Node zip ships node.exe + npm.cmd at the root; prepend it (this
# process only) so node/npm/bun resolve here for the build.
$env:PATH = "$NodeDir;" + $env:PATH
# Keep npm and module resolution inside the isolated Node.
$env:NPM_CONFIG_PREFIX = $NodeDir
$env:npm_config_prefix = $NodeDir
Remove-Item Env:NODE_PATH -ErrorAction SilentlyContinue
step "node" "$(node -v) | npm $(npm -v) (isolated)"
# bun (optional, faster installs); npm -g stays in the isolated prefix.
if (-not (Get-Command bun -ErrorAction SilentlyContinue)) {
substep "installing bun (faster frontend package installs)..."
$prevEAP_bun = $ErrorActionPreference
$ErrorActionPreference = "Continue"
Invoke-SetupCommand { npm install -g bun --allow-scripts=bun @NpmRegistryArgs } | Out-Null
$ErrorActionPreference = $prevEAP_bun
Refresh-Environment
# Refresh-Environment rebuilds PATH (Machine;User;current), demoting the
# isolated-Node prepend; re-prepend so it wins for the build and OXC step.
$env:PATH = "$NodeDir;" + $env:PATH
$env:NPM_CONFIG_PREFIX = $NodeDir
$env:npm_config_prefix = $NodeDir
Remove-Item Env:NODE_PATH -ErrorAction SilentlyContinue
if (Get-Command bun -ErrorAction SilentlyContinue) {
substep "bun installed ($(bun --version))"
} else {
substep "bun install skipped (npm will be used instead)"
}
}
} else {
# system Node already satisfies requirements; use it as-is. We do NOT
# install global packages (bun) here -- the build falls back to npm.
step "node" "$SysNodeVersion | npm $SysNpmVersion (system)"
}
}
if ($NeedFrontendBuild -and -not $IsPipInstall) {
Write-Host ""
substep "building frontend..."
# ── Tailwind v4 .gitignore workaround ──
# Tailwind v4's oxide scanner respects .gitignore in parent directories.
# Python venvs create a .gitignore with "*" (ignore everything), which
# prevents Tailwind from scanning .tsx source files for class names.
# Temporarily hide any such .gitignore during the build, then restore it.
$HiddenGitignores = @()
$WalkDir = (Get-Item $FrontendDir).Parent.FullName
while ($WalkDir -and $WalkDir -ne [System.IO.Path]::GetPathRoot($WalkDir)) {
$gi = Join-Path $WalkDir ".gitignore"
if (Test-Path $gi) {
$content = Get-Content $gi -Raw -ErrorAction SilentlyContinue
if ($content -and ($content.Trim() -match '^\*$')) {
$hidden = "$gi._twbuild"
Rename-Item -Path $gi -NewName (Split-Path $hidden -Leaf) -Force
$HiddenGitignores += $gi
substep "Temporarily hiding $gi (venv .gitignore blocks Tailwind scanner)"
}
}
$WalkDir = Split-Path $WalkDir -Parent
}
# Use bun if available (faster install), fall back to npm.
# Bun is used only as package manager; Node runs the actual build (Vite 8).
$prevEAP_npm = $ErrorActionPreference
$ErrorActionPreference = "Continue"
Push-Location $FrontendDir
$UseBun = $null -ne (Get-Command bun -ErrorAction SilentlyContinue)
# bun's package cache can become corrupt -- packages get stored with only
# metadata but no actual content (bin/, lib/). When this happens bun install
# exits 0 but leaves binaries missing. We validate after install and clear
# the cache + retry once before falling back to npm.
if ($UseBun) {
Write-Host " Using bun for package install (faster)" -ForegroundColor DarkGray
$bunExit = Invoke-SetupCommand { bun install @NpmRegistryArgs }
# On Windows, .bin/ entries vary by package manager:
# npm → tsc, tsc.cmd, tsc.ps1
# bun → tsc.exe, tsc.bunx
$hasTsc = (Test-Path "node_modules\.bin\tsc") -or (Test-Path "node_modules\.bin\tsc.cmd") -or (Test-Path "node_modules\.bin\tsc.exe") -or (Test-Path "node_modules\.bin\tsc.bunx")
$hasVite = (Test-Path "node_modules\.bin\vite") -or (Test-Path "node_modules\.bin\vite.cmd") -or (Test-Path "node_modules\.bin\vite.exe") -or (Test-Path "node_modules\.bin\vite.bunx")
if ($bunExit -eq 0 -and $hasTsc -and $hasVite) {
# bun install succeeded and critical binaries are present
} elseif ($bunExit -eq 0) {
Write-Host " bun install exited 0 but critical binaries are missing, clearing cache and retrying..." -ForegroundColor Yellow
if (Test-Path "node_modules") {
Remove-Item "node_modules" -Recurse -Force -ErrorAction SilentlyContinue
}
Invoke-SetupCommand { bun pm cache rm } | Out-Null
$bunExit = Invoke-SetupCommand { bun install @NpmRegistryArgs }
$hasTsc = (Test-Path "node_modules\.bin\tsc") -or (Test-Path "node_modules\.bin\tsc.cmd") -or (Test-Path "node_modules\.bin\tsc.exe") -or (Test-Path "node_modules\.bin\tsc.bunx")
$hasVite = (Test-Path "node_modules\.bin\vite") -or (Test-Path "node_modules\.bin\vite.cmd") -or (Test-Path "node_modules\.bin\vite.exe") -or (Test-Path "node_modules\.bin\vite.bunx")
if ($bunExit -ne 0 -or -not $hasTsc -or -not $hasVite) {
Write-Host " bun retry failed, falling back to npm" -ForegroundColor Yellow
if (Test-Path "node_modules") {
Remove-Item "node_modules" -Recurse -Force -ErrorAction SilentlyContinue
}
$UseBun = $false
}
} else {
substep "bun install failed (exit $bunExit), falling back to npm" "Yellow"
if (Test-Path "node_modules") {
Remove-Item "node_modules" -Recurse -Force -ErrorAction SilentlyContinue
}
$UseBun = $false
}
}
if (-not $UseBun) {
$npmExit = Invoke-SetupCommand { npm install @NpmRegistryArgs }
if ($npmExit -ne 0) {
Pop-Location
$ErrorActionPreference = $prevEAP_npm
foreach ($gi in $HiddenGitignores) { Rename-Item -Path "$gi._twbuild" -NewName (Split-Path $gi -Leaf) -Force -ErrorAction SilentlyContinue }
Write-Host "[ERROR] npm install failed (exit code $npmExit)" -ForegroundColor Red
Write-Host " Try running 'npm install' manually in frontend/ to see errors" -ForegroundColor Yellow
Show-NpmRegistryHint
exit 1
}
}
# Always use npm to run the build (Node runtime — avoids bun Windows runtime issues)
$buildExit = Invoke-SetupCommand { npm run build }
if ($buildExit -ne 0) {
Pop-Location
$ErrorActionPreference = $prevEAP_npm
foreach ($gi in $HiddenGitignores) { Rename-Item -Path "$gi._twbuild" -NewName (Split-Path $gi -Leaf) -Force -ErrorAction SilentlyContinue }
Write-Host "[ERROR] npm run build failed (exit code $buildExit)" -ForegroundColor Red
exit 1
}
Pop-Location
$ErrorActionPreference = $prevEAP_npm
# ── Restore hidden .gitignore files ──
foreach ($gi in $HiddenGitignores) {
Rename-Item -Path "$gi._twbuild" -NewName (Split-Path $gi -Leaf) -Force -ErrorAction SilentlyContinue
}
# ── Validate CSS output ──
$CssFiles = Get-ChildItem (Join-Path $DistDir "assets") -Filter "*.css" -ErrorAction SilentlyContinue
$MaxCssSize = ($CssFiles | Measure-Object -Property Length -Maximum).Maximum
if ($MaxCssSize -lt 100000) {
step "frontend" "built (warning: CSS may be truncated)" "Yellow"
} else {
step "frontend" "built"
}
}
if ((Test-Path $OxcValidatorDir) -and $NodeSource -ne "skip" -and (Get-Command npm -ErrorAction SilentlyContinue)) {
substep "installing OXC validator runtime..."
$prevEAP_oxc = $ErrorActionPreference
$ErrorActionPreference = "Continue"
Push-Location $OxcValidatorDir
$oxcInstallExit = Invoke-SetupCommand { npm install @NpmRegistryArgs }
if ($oxcInstallExit -ne 0) {
Pop-Location
$ErrorActionPreference = $prevEAP_oxc
Write-Host "[ERROR] OXC validator npm install failed (exit code $oxcInstallExit)" -ForegroundColor Red
Show-NpmRegistryHint
exit 1
}
Pop-Location
$ErrorActionPreference = $prevEAP_oxc
step "oxc runtime" "installed"
} elseif ((Test-Path $OxcValidatorDir) -and $NodeSource -ne "skip") {
# No npm on PATH (e.g. a pip install with no system Node and no isolated Node
# provisioned). Skip rather than abort; the runtime resolver degrades. Mirrors setup.sh.
substep "OXC validator runtime skipped (no npm found); code validation degrades until Node is available" "Yellow"
}
Remove-AgentInstructionFiles -Roots @(
(Join-Path $FrontendDir "node_modules"),
(Join-Path $OxcValidatorDir "node_modules")
)
# ==========================================================================
# PHASE 3: Python environment + dependencies
# ==========================================================================
Write-Host ""
substep "setting up Python environment..."
# Find Python -- skip Anaconda/Miniconda distributions ($CondaSkipPattern and
# Test-IsConda are defined above the 1g gate). Standalone CPython (python.org,
# winget, uv) does not have conda's torch c10.dll loading issue.
$PythonCmd = $null
# 0. Reuse the interpreter install.ps1 already resolved and built the venv with
# (UNSLOTH_SETUP_PYTHON, or the existing venv python) before probing the
# system -- it is already validated as supported and non-conda.
if ($ReusedSetupPython) {
try {
$out = & $ReusedSetupPython --version 2>&1 | Out-String
if ($out -match 'Python 3\.(\d+)') {
$pyMinor = [int]$Matches[1]
if ($pyMinor -ge 11 -and $pyMinor -le 13 -and -not (Test-IsConda $ReusedSetupPython)) {
$PythonCmd = $ReusedSetupPython
}
}
} catch { }
}
# 1. Try the Python Launcher (py.exe) first -- most reliable on Windows.
# Enumerate every launcher with -All (Windows PowerShell 5.1 returns only
# the first match without it) and search each for a supported, non-conda
# interpreter.
$PyLaunchersResolve = if ($PythonCmd) { @() } else { @(Get-Command py -All -CommandType Application -ErrorAction SilentlyContinue) }
foreach ($pyLauncher in $PyLaunchersResolve) {
if ($pyLauncher.Source -match $CondaSkipPattern) { continue }
foreach ($minor in @("3.13", "3.12", "3.11")) {
try {
$out = & $pyLauncher.Source "-$minor" --version 2>&1 | Out-String
if ($out -match 'Python 3\.(\d+)') {
$pyMinor = [int]$Matches[1]
if ($pyMinor -ge 11 -and $pyMinor -le 13) {
# Resolve the actual executable path so venv creation
# does not re-resolve back to a conda interpreter.
$resolvedExe = (& $pyLauncher.Source "-$minor" -c "import sys; print(sys.executable)" 2>$null | Out-String).Trim()
if ($resolvedExe -and (Test-Path $resolvedExe) -and -not (Test-IsConda $resolvedExe)) {
$PythonCmd = $resolvedExe
break
}
}
}
} catch { }
}
if ($PythonCmd) { break }
}
# 2. Fall back to scanning python3.x / python3 / python on PATH.
# Use Get-Command -All to look past conda entries.
if (-not $PythonCmd) {
foreach ($candidate in @("python3.13", "python3.12", "python3.11", "python3", "python")) {
foreach ($cmdInfo in @(Get-Command $candidate -All -ErrorAction SilentlyContinue)) {
try {
if (-not $cmdInfo.Source) { continue }
if ($cmdInfo.Source -like "*\WindowsApps\*") { continue }
if (Test-IsConda $cmdInfo.Source) {
substep "skipping $($cmdInfo.Source) (conda Python breaks torch DLL loading)" "Yellow"
continue
}
$ver = & $cmdInfo.Source --version 2>&1
if ($ver -match 'Python 3\.(\d+)') {
$minor = [int]$Matches[1]
if ($minor -ge 11 -and $minor -le 13) {
$PythonCmd = $cmdInfo.Source
break
}
}
} catch { }
}
if ($PythonCmd) { break }
}
}
if (-not $PythonCmd) {
Write-Host "[ERROR] No standalone Python 3.11-3.13 found (conda Python is not supported)." -ForegroundColor Red
Write-Host " Install Python from https://python.org/downloads/ or via:" -ForegroundColor Yellow
Write-Host " winget install -e --id Python.Python.3.12" -ForegroundColor Yellow
exit 1
}
substep "Python found: $PythonCmd"
# The venv must already exist (created by install.ps1); this script only
# updates packages. UNSLOTH_STUDIO_HOME (or STUDIO_HOME alias) overrides the
# root. UNSLOTH_STUDIO_HOME wins when both are set. Whitespace-only values
# are treated as unset to match Python .strip() semantics.
$_studioOverrideVar = $null
$_studioOverride = $null
if (-not [string]::IsNullOrWhiteSpace($env:UNSLOTH_STUDIO_HOME)) {
$_studioOverrideVar = "UNSLOTH_STUDIO_HOME"
$_studioOverride = $env:UNSLOTH_STUDIO_HOME.Trim()
} elseif (-not [string]::IsNullOrWhiteSpace($env:STUDIO_HOME)) {
$_studioOverrideVar = "STUDIO_HOME"
$_studioOverride = $env:STUDIO_HOME.Trim()
}
if ($_studioOverride) {
if ($_studioOverride -eq "~" -or $_studioOverride -like "~/*" -or $_studioOverride -like "~\*") {
$_studioOverride = (Join-Path $env:USERPROFILE $_studioOverride.Substring(1).TrimStart('/','\'))
}
if (Test-Path -LiteralPath $_studioOverride -PathType Container) {
$StudioHome = (Resolve-Path -LiteralPath $_studioOverride).Path
# why: mirror setup.sh:417 and install.ps1:130 -- fail fast when the
# custom root is read-only instead of erroring later while creating
# sidecar venvs / installing packages.
$_setupWriteProbe = Join-Path $StudioHome (".unsloth-write-probe-" + [guid]::NewGuid())
try {
[System.IO.File]::WriteAllText($_setupWriteProbe, "")
Remove-Item -LiteralPath $_setupWriteProbe -Force -ErrorAction SilentlyContinue
} catch {
Write-Host "ERROR: $_studioOverrideVar=$StudioHome is not writable." -ForegroundColor Red
exit 1
}
} else {
Write-Host "ERROR: $_studioOverrideVar=$_studioOverride does not exist." -ForegroundColor Red
Write-Host " Run install.ps1 to create the install root before 'unsloth studio update'." -ForegroundColor Red
exit 1
}
} else {
$StudioHome = Join-Path $env:USERPROFILE ".unsloth\studio"
}
$VenvDir = Join-Path $StudioHome "unsloth_studio"
# why: in env-override mode $StudioHome is user-chosen; require the
# ownership marker before Remove-Item so unrelated dirs survive. Gated on
# the canonical comparison so an override pointing at the legacy default
# still behaves like a default install.
$StudioOwnedMarker = ".unsloth-studio-owned"
$LegacyStudioHome = Join-Path $env:USERPROFILE ".unsloth\studio"
$_studioHomeCanon = $StudioHome
if (Test-Path -LiteralPath $_studioHomeCanon -PathType Container) {
$_studioHomeCanon = (Resolve-Path -LiteralPath $_studioHomeCanon).Path
}
if (Test-Path -LiteralPath $LegacyStudioHome -PathType Container) {
$LegacyStudioHome = (Resolve-Path -LiteralPath $LegacyStudioHome).Path
}
$StudioHomeIsCustom = ($_studioHomeCanon -ne $LegacyStudioHome)
# Directory-local evidence that Unsloth created $Path, used to adopt a custom-home
# llama.cpp predating the .unsloth-studio-owned marker (see setup.sh). Only the
# prebuilt UNSLOTH_PREBUILT_INFO.json counts; source builds are indistinguishable
# from a user clone on Windows and stay under the strict guard.
function Test-StudioOwnedAdoptable {
param([Parameter(Mandatory = $true)][string]$Path)
if (Test-Path -LiteralPath (Join-Path $Path "UNSLOTH_PREBUILT_INFO.json") -PathType Leaf) { return $true }
return $false
}
function Assert-StudioOwnedOrAbsent {
param(
[Parameter(Mandatory = $true)][string]$Path,
[Parameter(Mandatory = $true)][string]$Label
)
if (-not (Test-Path -LiteralPath $Path -PathType Container)) { return }
if ($StudioHomeIsCustom -and -not (Test-Path -LiteralPath (Join-Path $Path $StudioOwnedMarker) -PathType Leaf)) {
if (Test-StudioOwnedAdoptable $Path) {
Mark-StudioOwned $Path
return
}
Write-Host "[ERROR] $Path already exists and is not marked as an Unsloth-owned $Label." -ForegroundColor Red
Write-Host " Move it aside or choose an empty UNSLOTH_STUDIO_HOME before re-running." -ForegroundColor Yellow
exit 1
}
}
function Mark-StudioOwned {
param([Parameter(Mandatory = $true)][string]$Path)
if (-not (Test-Path -LiteralPath $Path -PathType Container)) { return }
try {
[System.IO.File]::WriteAllText((Join-Path $Path $StudioOwnedMarker), "")
} catch {}
}
# Stale-venv detection: if the venv exists but its torch flavor no longer
# matches the current machine, repair according to invocation context.
# - install.ps1 sets UNSLOTH_INSTALL_ROLLBACK_MANAGED=1 so setup can delegate
# to the installer-level rollback that restores the previous environment.
# - direct `unsloth studio update` keeps the pre-existing self-repair behavior.
# In no-torch mode, a missing torch package is expected.
$NoTorchMode = $env:UNSLOTH_NO_TORCH -match '^(?i:true|1|yes)$'
$InstallerManagedSetup = $env:UNSLOTH_INSTALL_ROLLBACK_MANAGED -match '^(?i:true|1|yes)$'
if ((Test-Path -LiteralPath $VenvDir -PathType Container) -and -not $NoTorchMode) {
$VenvPyExe = Join-Path $VenvDir "Scripts\python.exe"
$installedTorchTag = $null
$shouldRebuild = $false
# Set when a stale venv under a pin is repaired in place (force-reinstall) not wiped.
$script:PinChangedForceReinstall = $false
if (Test-Path -LiteralPath $VenvPyExe) {
try {
$psi = New-Object System.Diagnostics.ProcessStartInfo
$psi.FileName = $VenvPyExe
$psi.Arguments = '-c "import torch; print(torch.__version__)"'
$psi.RedirectStandardOutput = $true
$psi.RedirectStandardError = $true
$psi.UseShellExecute = $false
$psi.CreateNoWindow = $true
$proc = [System.Diagnostics.Process]::Start($psi)
$torchVer = $proc.StandardOutput.ReadToEnd().Trim()
$finished = $proc.WaitForExit(30000)
if ($finished -and $proc.ExitCode -eq 0 -and $torchVer) {
if ($torchVer -match '\+(cu\d+)') {
$installedTorchTag = $Matches[1]
} elseif ($torchVer -match '\+rocm') {
# Any +rocm / gfx wheel -> generic "rocm" flavor (the exact version is
# repaired later by install_python_stack.py; here we only need the flavor).
$installedTorchTag = "rocm"
} elseif ($torchVer -match '\+cpu') {
$installedTorchTag = "cpu"
} else {
# Untagged wheel (plain "2.x.y" from PyPI) -> cpu.
$installedTorchTag = "cpu"
}
} else {
if (-not $finished) { try { $proc.Kill() } catch {} }
$shouldRebuild = $true
}
} catch {
$shouldRebuild = $true
}
} else {
# Missing python.exe means the venv is incomplete -- rebuild it.
$shouldRebuild = $true
}
if (-not $shouldRebuild) {
$_pinnedIdx = Get-PinnedTorchIndexUrl
$_expectedKnown = $true
if ($_pinnedIdx) {
$_pinLeaf = Get-TorchIndexLeaf $_pinnedIdx
# Digit-gated like the install selection: a custom rocm-* leaf (rocm-current /
# rocm-rel-7.2.1) is NOT a ROCm family and must not be stale-compared.
if (Test-PipRocmFamilyLeaf $_pinLeaf) {
# Don't collapse a pinned ROCm/gfx leaf to a generic "rocm" (would mask a family
# change, rocm6.4 -> gfx1151). Get-RocmPinStaleTags uses the SAME 2.11 allowlist
# as the install path, so a gfx110X-all/gfx90a/gfx908 pin on a <2.11 wheel is NOT stale.
$_rocmTags = Get-RocmPinStaleTags -PinLeaf $_pinLeaf -TorchVersion $torchVer
$expectedTorchTag = $_rocmTags.Expected
$installedTorchTag = $_rocmTags.Installed
} elseif ((Test-CudaFamilyLeaf $_pinLeaf) -or $_pinLeaf -eq 'cpu') {
# cu*/cpu leaves stay specific so a cu126-vs-cu128 mismatch rebuilds;
# /custom and /current fall through to the unknown-index branch below.
$expectedTorchTag = $_pinLeaf
} else {
# Custom index whose leaf is not a torch flavor (a /simple mirror): the
# flavor can't be inferred, so never treat the venv as stale over it.
$_expectedKnown = $false
$expectedTorchTag = $installedTorchTag
}
} elseif ($HasNvidiaSmi) {
$expectedTorchTag = Get-PytorchCudaTag
} elseif ($HasROCm -or $script:ROCmGfxArch) {
# AMD/ROCm host with no explicit pin: an existing +rocm wheel is correct (gfx arch
# counts even when $HasROCm is false). But only the arches the install path maps to a
# repo.amd.com index get ROCm torch; an unmapped arch installs CPU, so expect "cpu"
# for those or a correct CPU venv rebuilds every update.
$_rocmWheelArches = @(
"gfx1201", "gfx1200", # RDNA 4
"gfx1151", "gfx1150", # RDNA 3.5 (Strix Halo/Point)
"gfx1103", "gfx1102", "gfx1101", "gfx1100", # RDNA 3
"gfx90a", "gfx908" # MI200 / MI100
)
if ($script:ROCmGfxArch -and ($_rocmWheelArches -contains $script:ROCmGfxArch)) {
# A correct +rocm wheel is not stale. A CPU wheel on a supported AMD arch is
# NOT wiped either (the AMD Windows ROCm override below upgrades it in place);
# expect "cpu" for that case. A wrong CUDA wheel still rebuilds.
if ($installedTorchTag -eq "cpu") {
$expectedTorchTag = "cpu"
} else {
$expectedTorchTag = "rocm"
}
} else {
$expectedTorchTag = "cpu"
}
} else {
$expectedTorchTag = "cpu"
}
if ($_expectedKnown -and $installedTorchTag -and $installedTorchTag -ne $expectedTorchTag) {
$shouldRebuild = $true
}
}
# A stale venv under a pin whose torch still imports is repaired IN PLACE (the dependency
# pass force-reinstalls from the pin). The rebuild path wipes the venv and would strand a
# direct `studio update`; only a broken venv or an unpinned drift wipes.
if ($shouldRebuild -and $_pinnedIdx -and $installedTorchTag) {
substep "Torch-index pin changed ($installedTorchTag) -- reinstalling torch from the pin in place." "Cyan"
$script:PinChangedForceReinstall = $true
$shouldRebuild = $false
}
if ($shouldRebuild) {
$reason = if ($installedTorchTag) { "torch $installedTorchTag != required $expectedTorchTag" } else { "torch could not be imported" }
if ($InstallerManagedSetup) {
substep "Stale venv detected ($reason)." "Yellow"
Write-Host " [ERROR] The existing Unsloth environment needs repair." -ForegroundColor Red
Write-Host " Re-run install.ps1 so it can replace the environment safely with rollback." -ForegroundColor Yellow
exit 1
}
substep "Stale venv detected ($reason) -- rebuilding..." "Yellow"
# why: mirror install.ps1 env-mode guard so an update against a custom
# UNSLOTH_STUDIO_HOME never wipes an unrelated unsloth_studio venv;
# -PathType Leaf rejects a directory masquerading as the sentinel.
if (
$StudioHomeIsCustom -and
-not (Test-Path -LiteralPath (Join-Path $VenvDir $StudioOwnedMarker) -PathType Leaf) -and
-not (Test-Path -LiteralPath (Join-Path $StudioHome "share\studio.conf") -PathType Leaf) -and
-not (Test-Path -LiteralPath (Join-Path $StudioHome "bin\unsloth.exe") -PathType Leaf)
) {
Write-Host "[ERROR] $VenvDir already exists but does not look like an Unsloth Studio install." -ForegroundColor Red
Write-Host " Move it aside or choose an empty UNSLOTH_STUDIO_HOME before re-running." -ForegroundColor Yellow
exit 1
}
try {
Remove-Item -LiteralPath $VenvDir -Recurse -Force -ErrorAction Stop
} catch {
Write-Host " [ERROR] Could not remove stale venv: $($_.Exception.Message)" -ForegroundColor Red
Write-Host " Close any running Unsloth/Python processes and re-run setup." -ForegroundColor Red
exit 1
}
}
}
if (-not (Test-Path -LiteralPath $VenvDir)) {
Write-Host "[ERROR] Virtual environment not found at $VenvDir" -ForegroundColor Red
Write-Host " Run install.ps1 first to create the environment:" -ForegroundColor Yellow
Write-Host " irm https://unsloth.ai/install.ps1 | iex" -ForegroundColor Yellow
exit 1
} else {
substep "reusing existing virtual environment at $VenvDir"
$_venvPyExe = Join-Path $VenvDir "Scripts\python.exe"
if (Test-Path -LiteralPath $_venvPyExe) {
try {
$_venvPyVer = (& $_venvPyExe --version 2>&1 | Out-String).Trim()
if ($_venvPyVer) { substep $_venvPyVer }
} catch {}
}
}
# pip and python write to stderr even on success (progress bars, warnings).
# With $ErrorActionPreference = "Stop" (set at top of script), PS 5.1
# converts stderr lines into terminating ErrorRecords, breaking output.
# Lower to "Continue" for the pip/python section.
$prevEAP = $ErrorActionPreference
$ErrorActionPreference = "Continue"
$ActivateScript = Join-Path $VenvDir "Scripts\Activate.ps1"
. $ActivateScript
# Try to use uv (much faster than pip), fall back to pip if unavailable
$UseUv = $false
if (Get-Command uv -ErrorAction SilentlyContinue) {
$UseUv = $true
} else {
substep "installing uv package manager..."
try {
Invoke-SetupCommand { Invoke-Expression (Invoke-RestMethod -Uri "https://astral.sh/uv/install.ps1") } | Out-Null
Refresh-Environment
# Re-activate venv since Refresh-Environment rebuilds PATH from
# registry and drops the venv's Scripts directory
. $ActivateScript
if (Get-Command uv -ErrorAction SilentlyContinue) { $UseUv = $true }
} catch { }
}
# Helper: install a package, preferring uv with pip fallback
function Fast-Install {
param([Parameter(ValueFromRemainingArguments=$true)]$Args_)
# An explicit --index-url must win: inherited uv index vars otherwise pull CPU torch over
# the CUDA/ROCm build (#6898), so drop them for pinned installs (scrub covers the whole
# function since the pip fallback honours PIP_* too). UV_TORCH_BACKEND / UV_FIND_LINKS also
# reroute; UV_NO_CONFIG=1 (+ dropping UV_CONFIG_FILE) stops a uv.toml index outranking the
# pin (uv 0.10); PIP_NO_INDEX / PIP_INDEX_URL would defeat the pinned --index-url in pip.
$saved = @{}
$pinned = @($Args_) -contains '--index-url'
if ($pinned) {
foreach ($n in 'UV_DEFAULT_INDEX', 'UV_INDEX_URL', 'UV_INDEX', 'UV_EXTRA_INDEX_URL',
'UV_TORCH_BACKEND', 'UV_FIND_LINKS', 'PIP_EXTRA_INDEX_URL', 'PIP_FIND_LINKS',
'PIP_NO_INDEX', 'PIP_INDEX_URL',
'UV_CONFIG_FILE', 'UV_NO_CONFIG', 'PIP_CONFIG_FILE') {
$saved[$n] = [Environment]::GetEnvironmentVariable($n)
Remove-Item "Env:$n" -ErrorAction SilentlyContinue
}
$env:UV_NO_CONFIG = '1'
# A `pip config` global.extra-index-url still adds indexes to the pip FALLBACK;
# PIP_CONFIG_FILE = 'nul' (Windows devnull) loads NO config (uv ignores pip config).
$env:PIP_CONFIG_FILE = 'nul'
}
try {
if ($UseUv) {
$VenvPy = (Get-Command python).Source
$result = & uv pip install --python $VenvPy @Args_ 2>&1
if ($LASTEXITCODE -eq 0) { return }
}
& python -m pip install @Args_ 2>&1
}
finally {
if ($pinned) {
Remove-Item "Env:UV_NO_CONFIG" -ErrorAction SilentlyContinue
Remove-Item "Env:PIP_CONFIG_FILE" -ErrorAction SilentlyContinue
}
foreach ($n in $saved.Keys) { if ($null -ne $saved[$n]) { Set-Item "Env:$n" $saved[$n] } }
}
}
# ── Check if Python deps need updating ──
# Compare installed package version against PyPI latest.
# Skip all Python dependency work if versions match (fast update path).
$_PkgName = if ($env:STUDIO_PACKAGE_NAME) { $env:STUDIO_PACKAGE_NAME } else { "unsloth" }
$SkipPythonDeps = $false
if ($env:SKIP_STUDIO_BASE -ne "1" -and $env:STUDIO_LOCAL_INSTALL -ne "1") {
# Only check when NOT called from install.ps1 (which just installed the package)
$InstalledVer = try { (& python -c "from importlib.metadata import version; print(version('$_PkgName'))" 2>$null | Out-String).Trim() } catch { "" }
$LatestVer = ""
try {
$pypiJson = Invoke-RestMethod -Uri "https://pypi.org/pypi/$_PkgName/json" -TimeoutSec 5 -ErrorAction Stop
$LatestVer = "$($pypiJson.info.version)".Trim()
} catch { }
if ($InstalledVer -and $LatestVer -and ($InstalledVer -eq $LatestVer)) {
step "python" "$_PkgName $InstalledVer is up to date"
$SkipPythonDeps = $true
# A pre-#6483-fix install can be stuck on anyio>=4.14 even though
# $_PkgName itself is current; the fast path above would otherwise
# never reach install_python_stack's anyio repair (#6797).
$_anyioBad = $false
try {
& python -c "
import re, sys
from importlib.metadata import version, PackageNotFoundError
try:
parts = version('anyio').split('.')
major = int(parts[0])
minor = int(re.sub(r'[^0-9].*', '', parts[1])) if len(parts) > 1 else 0
except (PackageNotFoundError, ValueError, IndexError):
sys.exit(1)
sys.exit(0 if (major, minor) >= (4, 14) else 1)
" 2>$null
if ($LASTEXITCODE -eq 0) { $_anyioBad = $true }
} catch {}
if ($_anyioBad) {
substep "anyio >=4.14 found (#6483) -- forcing dependency pass to repair..." "Cyan"
$SkipPythonDeps = $false
}
# ...but not if an AMD GPU is present and installed PyTorch is CPU-only
# (host predates ROCm-wheel support, or GPU added later): the fast "up to
# date" path would leave the user on CPU torch with Train/Export disabled.
# Force the dependency pass so the ROCm wheels get installed.
if ($script:ROCmGfxArch) {
$_torchIsCpu = $true
try {
& python -c "import torch, sys; sys.exit(0 if torch.cuda.is_available() else 1)" 2>$null
if ($LASTEXITCODE -eq 0) { $_torchIsCpu = $false }
} catch {}
if ($_torchIsCpu) {
substep "AMD GPU ($script:ROCmGfxArch) detected but installed PyTorch is CPU-only -- reinstalling ROCm PyTorch" "Cyan"
$SkipPythonDeps = $false
}
}
} elseif ($InstalledVer -and $LatestVer) {
substep "$_PkgName $InstalledVer -> $LatestVer available, updating..."
} elseif (-not $LatestVer) {
substep "could not reach PyPI, updating to be safe..."
}
}
# if (-not $IsPipInstall) {
# # Running from repo: copy requirements and do editable install
# $RepoRoot = (Resolve-Path (Join-Path $ScriptDir "..\..")).Path
# $ReqsSrc = Join-Path $RepoRoot "backend\requirements"
# $ReqsDst = Join-Path $PackageDir "requirements"
# if (-not (Test-Path $ReqsDst)) { New-Item -ItemType Directory -Path $ReqsDst | Out-Null }
# Copy-Item (Join-Path $ReqsSrc "*.txt") $ReqsDst -Force
# Write-Host " Installing CLI entry point..." -ForegroundColor Cyan
# pip install -e $RepoRoot 2>&1 | Out-Null
# } else {
# # Running from pip install: the package is in system Python but not in
# # the fresh .venv. Install it so run_install() can find its modules
# # and bundled requirements files.
# Write-Host " Installing package into venv..." -ForegroundColor Cyan
# pip install unsloth 2>&1 | Out-Null
# }
# A torch-index pin change repairs in place: force the dependency pass so the torch install
# below force-reinstalls from the new pin (else the fast path keeps the old wheel).
if ($script:PinChangedForceReinstall) { $SkipPythonDeps = $false }
if (-not $SkipPythonDeps) {
if ($script:UnslothVerbose) {
Fast-Install --upgrade pip
} else {
Fast-Install --upgrade pip | Out-Null
}
# Pre-install PyTorch with CUDA support.
# On Windows, the default PyPI torch wheel is CPU-only.
# We need PyTorch's CUDA index to get GPU-enabled wheels.
# PyTorch bundles its own CUDA runtime, so this works regardless
# of whether the CUDA Toolkit is installed yet.
# The CUDA tag is chosen based on the driver's max supported CUDA version.
# Triton/inductor filenames are long and can hit Windows MAX_PATH (260). With long
# paths on, cache under Unsloth home; else use a short drive-root dir for headroom.
if ($LongPathsEnabled) {
$TorchCacheDir = Join-Path $StudioHome "TORCHINDUCTOR_CACHE_DIR"
} else {
$TorchCacheDir = "C:\tc"
}
if (-not (Test-Path -LiteralPath $TorchCacheDir)) { [System.IO.Directory]::CreateDirectory($TorchCacheDir) | Out-Null }
$env:TORCHINDUCTOR_CACHE_DIR = $TorchCacheDir
[Environment]::SetEnvironmentVariable('TORCHINDUCTOR_CACHE_DIR', $TorchCacheDir, 'User')
substep "TORCHINDUCTOR_CACHE_DIR set to $TorchCacheDir (avoids MAX_PATH issues)"
# Explicit pin (URL or family) wins over GPU probing and suppresses the AMD reroute below;
# matches install.sh / install.ps1 / install_python_stack.py.
$PinnedTorchIndexUrl = Get-PinnedTorchIndexUrl
$TorchIndexPinned = [bool]$PinnedTorchIndexUrl
if ($PinnedTorchIndexUrl) {
$CuTag = Get-TorchIndexLeaf $PinnedTorchIndexUrl
} elseif ($HasNvidiaSmi) {
$CuTag = Get-PytorchCudaTag
} else {
$CuTag = "cpu"
}
# ── GPU arch → newest compatible Windows ROCm wheel release ──
# Wheels bundle their own ROCm runtime; the installed HIP SDK version does
# not constrain which release to use. Always picks the newest release that
# supports the GPU architecture.
# ── AMD Windows ROCm torch override ──────────────────────────────────────────
# Uses AMD's arch-specific pip index (repo.amd.com/rocm/whl/{arch}/).
# Wheels bundle their own ROCm runtime; HIP SDK version is irrelevant.
$ROCmGfxArch = $script:ROCmGfxArch
$ROCmIndexUrl = $null
# Install AMD ROCm PyTorch wheels when ROCm is confirmed OR a gfx arch is known
# (name-inferred on Adrenalin-only hosts). The per-arch wheels bundle the runtime
# (rocm-sdk-libraries-<gfx>), so torch.cuda.is_available() is True without a HIP
# SDK -- which flips Unsloth out of chat-only (CHAT_ONLY) and enables Train/Export.
# Gating on $HasROCm alone left Strix Halo / Radeon 8060S on CPU torch; a failed
# ROCm install still falls back to CPU below, so this is safe.
if (-not $TorchIndexPinned -and ($HasROCm -or $ROCmGfxArch) -and $CuTag -eq "cpu") {
$amdIndexBase = if ($env:UNSLOTH_ROCM_WINDOWS_MIRROR) { $env:UNSLOTH_ROCM_WINDOWS_MIRROR.TrimEnd('/') } else { "https://repo.amd.com/rocm/whl" }
$archFamilyMap = @{
"gfx1201" = "gfx120X-all"; "gfx1200" = "gfx120X-all" # RDNA 4
"gfx1151" = "gfx1151"; "gfx1150" = "gfx1150" # RDNA 3.5 (Strix Halo/Point)
"gfx1103" = "gfx110X-all"; "gfx1102" = "gfx110X-all" # RDNA 3
"gfx1101" = "gfx110X-all"; "gfx1100" = "gfx110X-all"
"gfx90a" = "gfx90a"; "gfx908" = "gfx908" # MI200/MI100
}
# gfx120X and Strix have a null _grouped_mm kernel on torch <2.11.0.
# Mirrors the $torchFloorMap in install.ps1 so both installers enforce
# the same floor and ceiling when pulling from AMD's per-arch index.
$torchFloorMap = @{
"gfx1201" = "torch>=2.11.0,<2.12.0"; "gfx1200" = "torch>=2.11.0,<2.12.0"
"gfx1151" = "torch>=2.11.0,<2.12.0"; "gfx1150" = "torch>=2.11.0,<2.12.0"
}
# Companion ranges for torchvision/torchaudio -- must stay in sync with the
# torch ceiling so pip can always find a consistent trio on AMD's per-arch
# index. AMD publishes each package independently and may add a newer
# torchvision (e.g. 0.27 for torch 2.12) before removing 0.26, which would
# cause pip to resolve an ABI-incompatible set if these are left bare.
# Matches _ROCM_TORCH_PKG_SPECS["rocm7.2"] in install_python_stack.py.
# Bump all three ceilings together when torch 2.12.x is validated.
$torchvisionFloorMap = @{
"gfx1201" = "torchvision>=0.26.0,<0.27.0"; "gfx1200" = "torchvision>=0.26.0,<0.27.0"
"gfx1151" = "torchvision>=0.26.0,<0.27.0"; "gfx1150" = "torchvision>=0.26.0,<0.27.0"
}
$torchaudioFloorMap = @{
"gfx1201" = "torchaudio>=2.11.0,<2.12.0"; "gfx1200" = "torchaudio>=2.11.0,<2.12.0"
"gfx1151" = "torchaudio>=2.11.0,<2.12.0"; "gfx1150" = "torchaudio>=2.11.0,<2.12.0"
}
$archFamily = if ($ROCmGfxArch -and $archFamilyMap.ContainsKey($ROCmGfxArch)) { $archFamilyMap[$ROCmGfxArch] } else { $null }
$ROCmTorchSpec = if ($ROCmGfxArch -and $torchFloorMap.ContainsKey($ROCmGfxArch)) { $torchFloorMap[$ROCmGfxArch] } else { "torch" }
$ROCmVisionSpec = if ($ROCmGfxArch -and $torchvisionFloorMap.ContainsKey($ROCmGfxArch)) { $torchvisionFloorMap[$ROCmGfxArch] } else { "torchvision" }
$ROCmAudioSpec = if ($ROCmGfxArch -and $torchaudioFloorMap.ContainsKey($ROCmGfxArch)) { $torchaudioFloorMap[$ROCmGfxArch] } else { "torchaudio" }
if ($archFamily) {
$ROCmIndexUrl = "$amdIndexBase/$archFamily/"
} elseif ($ROCmGfxArch) {
# GPU arch detected but not in the supported wheel map — warn explicitly
# so the user knows why they are getting CPU PyTorch instead of ROCm.
substep "[WARN] AMD GPU ($ROCmGfxArch) not in supported arch list -- falling back to CPU-only PyTorch" "Yellow"
substep " Supported: gfx1200/1201 (RDNA 4), gfx1150/1151 (RDNA 3.5), gfx1100-1103 (RDNA 3), gfx90a, gfx908" "Yellow"
} else {
# HIP SDK present ($HasROCm=true via amd-smi) but gcnArchName was not
# readable — warn rather than silently falling back to CPU PyTorch.
substep "[WARN] AMD GPU detected (HIP SDK present) but GPU arch could not be read -- falling back to CPU-only PyTorch" "Yellow"
substep " Arch detection requires hipinfo to report gcnArchName. Re-install the HIP SDK if this is unexpected." "Yellow"
}
}
# A pinned gfx*/rocm index skips the auto-reroute above; route it through the ROCm install path
# with the same floor/companions the unpinned AMD path uses (mirrors install.ps1), else the CUDA
# branch installs bare torch and resolves a known-bad wheel for gfx115x/gfx120x/rocm>=7.2.
if ($TorchIndexPinned -and -not $ROCmIndexUrl -and $PinnedTorchIndexUrl) {
$_pinLeaf = Get-TorchIndexLeaf $PinnedTorchIndexUrl
$_pinRocm211 = $false
# Anchor the match ($) so a suffixed custom leaf (rocm7.2-private) falls through to the
# verbatim install instead of being floored by its rocm7.2 prefix.
if ($_pinLeaf -match '^rocm(\d+)\.(\d+)$') {
# Only KNOWN-2.11 rocm (rocm7.2) gets the floor (no speculative floor). Matches
# Test-RocmKnown211Version / _ROCM_KNOWN_TORCH211_VERSIONS.
$_pinRocm211 = Test-RocmKnown211Version -Major ([int]$Matches[1]) -Minor ([int]$Matches[2])
}
# Only the 2.11 gfx arches need the floor; others publish <2.11 and stay bare. Reuse
# Test-RocmGfx211Leaf so this allowlist and the stale-venv check never diverge.
$_pinGfx211 = Test-RocmGfx211Leaf $_pinLeaf
if ($_pinGfx211 -or $_pinRocm211) {
$ROCmIndexUrl = $PinnedTorchIndexUrl
$ROCmTorchSpec = "torch>=2.11.0,<2.12.0"
$ROCmVisionSpec = "torchvision>=0.26.0,<0.27.0"
$ROCmAudioSpec = "torchaudio>=2.11.0,<2.12.0"
substep "pinned ROCm index ($_pinLeaf) -- enforcing $ROCmTorchSpec" "Cyan"
} elseif (Test-PipRocmFamilyLeaf $_pinLeaf) {
# Other gfx / older rocm (<=7.1) ship torch <2.11; route via the ROCm path with
# bare specs. Only EXACT rocm<digits> and gfx* are --index-url families; a suffixed
# leaf stays on the verbatim path. Mirrors install.ps1 / _is_pip_rocm_family_leaf.
$ROCmIndexUrl = $PinnedTorchIndexUrl
$ROCmTorchSpec = "torch"
$ROCmVisionSpec = "torchvision"
$ROCmAudioSpec = "torchaudio"
}
}
$PyTorchWhlBase = if ($env:UNSLOTH_PYTORCH_MIRROR) { $env:UNSLOTH_PYTORCH_MIRROR.TrimEnd('/') } else { "https://download.pytorch.org/whl" }
# A full URL pin is used verbatim; a family pin already set $CuTag. A pinned ROCm install
# goes through $ROCmIndexUrl; on failure the fallback uses the CPU index, not the ROCm pin.
$TorchInstallIndexUrl = if ($ROCmIndexUrl) { "$PyTorchWhlBase/cpu" } elseif ($PinnedTorchIndexUrl) { $PinnedTorchIndexUrl } else { "$PyTorchWhlBase/$CuTag" }
$ROCmCpuFallback = $false
if ($ROCmIndexUrl) {
substep "installing PyTorch (AMD ROCm, $ROCmGfxArch)..."
if ($ROCmTorchSpec -ne "torch") {
substep " enforcing $ROCmTorchSpec $ROCmVisionSpec $ROCmAudioSpec (known _grouped_mm bug in older wheels)" "Cyan"
}
if ($script:UnslothVerbose) {
Fast-Install $ROCmTorchSpec $ROCmVisionSpec $ROCmAudioSpec --force-reinstall --index-url $ROCmIndexUrl | ForEach-Object { Redact-InstallOutput "$_" } | Out-Host
$torchInstallExit = $LASTEXITCODE
$output = ""
} else {
$output = Fast-Install $ROCmTorchSpec $ROCmVisionSpec $ROCmAudioSpec --force-reinstall --index-url $ROCmIndexUrl | Out-String
$torchInstallExit = $LASTEXITCODE
}
if ($torchInstallExit -ne 0) {
Write-Host "[WARN] AMD ROCm PyTorch install failed -- falling back to CPU" -ForegroundColor Yellow
Write-Host (Redact-InstallOutput $output) -ForegroundColor Yellow
$ROCmIndexUrl = $null
$ROCmCpuFallback = $true
} else {
# Tell install_python_stack.py to skip probe + suppress manual-install warning.
$env:UNSLOTH_ROCM_TORCH_INSTALLED = "1"
substep "GPU ROCm PyTorch installed ($ROCmGfxArch) -- training and GPU inference will use the GPU" "Cyan"
}
}
if (-not $ROCmIndexUrl -and ($CuTag -eq "cpu" -or $ROCmCpuFallback)) {
substep "installing PyTorch (CPU-only)..."
# After an AMD ROCm fallback, force-reinstall so a partial ROCm torch (which satisfies the
# CPU torch>= range) is replaced by the CPU build; skip on a genuine CPU host to stay fast.
# $ROCmCpuFallback matters when a PINNED ROCm index failed ($CuTag is still the rocm leaf).
# Build the array directly: an if-expression collapses @("x") to a scalar @splat would
# enumerate char-by-char.
$cpuForce = @()
if ($ROCmCpuFallback) { $cpuForce = @("--force-reinstall") }
# --force-reinstall on a pin change: a stale +cu / +rocm wheel still satisfies the CPU
# torch>= range, so uv would keep it and only swap companions.
if ($script:PinChangedForceReinstall) { $cpuForce = @("--force-reinstall") }
# A PINNED cpu index installs the bounded trio (parity with _CPU_TORCH_PKG_SPEC): the /cpu
# index serves newer torch, and _ensure_cpu_torch keeps any CPU build, so a bare trio could
# land an unsupported version. Unpinned CPU hosts keep the bare trio (pre-pin behavior).
$cpuTorchSpec = "torch"; $cpuVisionSpec = "torchvision"; $cpuAudioSpec = "torchaudio"
if ($TorchIndexPinned) {
$cpuTorchSpec = "torch>=2.4,<2.12.0"
$cpuVisionSpec = "torchvision>=0.19,<0.27.0"
$cpuAudioSpec = "torchaudio>=2.4,<2.12.0"
}
if ($script:UnslothVerbose) {
Fast-Install $cpuTorchSpec $cpuVisionSpec $cpuAudioSpec @cpuForce --index-url $TorchInstallIndexUrl | ForEach-Object { Redact-InstallOutput "$_" } | Out-Host
$torchInstallExit = $LASTEXITCODE
$output = ""
} else {
$output = Fast-Install $cpuTorchSpec $cpuVisionSpec $cpuAudioSpec @cpuForce --index-url $TorchInstallIndexUrl | Out-String
$torchInstallExit = $LASTEXITCODE
}
if ($torchInstallExit -ne 0) {
Write-Host "[FAILED] PyTorch install failed (exit code $torchInstallExit)" -ForegroundColor Red
Write-Host (Redact-InstallOutput $output) -ForegroundColor Red
exit 1
}
} elseif (-not $ROCmIndexUrl) {
substep "installing PyTorch with CUDA support ($CuTag)..."
substep "(This download is ~2.8 GB -- may take a few minutes)"
# --force-reinstall on a pin change: an installed cuXXX wheel satisfies the bare torch
# requirement (PEP 440 ignores the +cuXXX tag), so without it a changed CUDA pin (cu126
# -> cu128) never applies.
$cudaForce = @()
if ($script:PinChangedForceReinstall) { $cudaForce = @("--force-reinstall") }
# An unknown-leaf custom pin (/simple, /current) routes here with $CuTag as that leaf. Bound
# the trio like the fresh custom-pin paths so a mirror can't pull an ABI-newer companion
# against the capped torch. Known cu* leaves keep bare specs.
$cudaTorchSpec = "torch"
$cudaVisionSpec = "torchvision"
$cudaAudioSpec = "torchaudio"
if ($TorchIndexPinned -and -not (Test-CudaFamilyLeaf $CuTag)) {
$cudaTorchSpec = "torch>=2.4,<2.11.0"
$cudaVisionSpec = "torchvision>=0.19,<0.26.0"
$cudaAudioSpec = "torchaudio>=2.4,<2.11.0"
}
if ($script:UnslothVerbose) {
Fast-Install $cudaTorchSpec $cudaVisionSpec $cudaAudioSpec @cudaForce --index-url $TorchInstallIndexUrl | ForEach-Object { Redact-InstallOutput "$_" } | Out-Host
$torchInstallExit = $LASTEXITCODE
$output = ""
} else {
$output = Fast-Install $cudaTorchSpec $cudaVisionSpec $cudaAudioSpec @cudaForce --index-url $TorchInstallIndexUrl | Out-String
$torchInstallExit = $LASTEXITCODE
}
if ($torchInstallExit -ne 0) {
Write-Host "[FAILED] PyTorch CUDA install failed (exit code $torchInstallExit)" -ForegroundColor Red
Write-Host (Redact-InstallOutput $output) -ForegroundColor Red
exit 1
}
# Install Triton for Windows (enables torch.compile -- without it training can hang)
substep "installing Triton for Windows..."
if ($script:UnslothVerbose) {
Fast-Install "triton-windows<3.7"
$tritonInstallExit = $LASTEXITCODE
$output = ""
} else {
$output = Fast-Install "triton-windows<3.7" | Out-String
$tritonInstallExit = $LASTEXITCODE
}
if ($tritonInstallExit -ne 0) {
substep "Triton install failed -- torch.compile may not work" "Yellow"
Write-Host (Redact-InstallOutput $output) -ForegroundColor Yellow
} else {
substep "Triton for Windows installed (enables torch.compile)"
}
}
# No unsloth.exe rename needed. setup.ps1 runs *via* unsloth.exe, so renaming the
# running launcher only ever failed (WinError 32) and printed a scary warning. It's
# also unnecessary: install.ps1 sets SKIP_STUDIO_BASE=1 (base never reinstalled) and
# 'studio update' goes through uv (--upgrade-package), whose pip fallback no-ops on
# the already-satisfied bare unsloth/unsloth-zoo. Either way unsloth.exe stays.
# Ordered heavy dependency installation -- shared cross-platform script
substep "running ordered dependency installation..."
python "$PSScriptRoot\install_python_stack.py"
$stackExit = $LASTEXITCODE
# Restore ErrorActionPreference after pip/python work
$ErrorActionPreference = $prevEAP
if ($stackExit -ne 0) {
Write-Host "[FAILED] Python dependency installation failed (exit code $stackExit)" -ForegroundColor Red
Write-Host " Re-run the installer or check the error above for details." -ForegroundColor Red
exit 1
}
} else {
step "python" "dependencies up to date"
# Restore ErrorActionPreference (was lowered for pip/python section)
$ErrorActionPreference = $prevEAP
}
# ── Pre-install transformers 5.x into .venv_t5_530/, .venv_t5_550/, and .venv_t5_510/ ──
# Runs outside the deps fast-path gate so that upgrades from the legacy
# single .venv_t5 are always migrated to the tiered layout.
# T5 sidecar venvs live under the resolved $StudioHome so custom installs are self-contained.
$VenvT5_530Dir = Join-Path $StudioHome ".venv_t5_530"
$VenvT5_550Dir = Join-Path $StudioHome ".venv_t5_550"
$VenvT5_510Dir = Join-Path $StudioHome ".venv_t5_510"
$VenvT5Legacy = Join-Path $StudioHome ".venv_t5"
function Test-TargetPackageVersion {
param(
[Parameter(Mandatory = $true)][string]$TargetDir,
[Parameter(Mandatory = $true)][string]$PackageName,
[Parameter(Mandatory = $true)][string]$ExpectedVersion
)
if (-not (Test-Path -LiteralPath $TargetDir -PathType Container)) { return $false }
$packageNorm = $PackageName.Replace("-", "_")
foreach ($pattern in @("$packageNorm-*.dist-info", "$PackageName-*.dist-info")) {
foreach ($distInfo in @(Get-ChildItem -LiteralPath $TargetDir -Directory -Filter $pattern -ErrorAction SilentlyContinue)) {
$metadata = Join-Path $distInfo.FullName "METADATA"
if (-not (Test-Path -LiteralPath $metadata -PathType Leaf)) { continue }
foreach ($line in (Get-Content -LiteralPath $metadata -ErrorAction SilentlyContinue)) {
if ($line -eq "Version: $ExpectedVersion") { return $true }
}
}
}
return $false
}
$_NeedT5Install = $false
if (Test-Path -LiteralPath $VenvT5Legacy) {
Assert-StudioOwnedOrAbsent -Path $VenvT5Legacy -Label "legacy transformers sidecar venv"
Remove-Item -LiteralPath $VenvT5Legacy -Recurse -Force
$_NeedT5Install = $true
}
if (-not (Test-Path -LiteralPath $VenvT5_530Dir)) { $_NeedT5Install = $true }
if (-not (Test-Path -LiteralPath $VenvT5_550Dir)) { $_NeedT5Install = $true }
if (-not (Test-Path -LiteralPath $VenvT5_510Dir)) { $_NeedT5Install = $true }
if (-not (Test-TargetPackageVersion -TargetDir $VenvT5_530Dir -PackageName "transformers" -ExpectedVersion "5.3.0")) { $_NeedT5Install = $true }
if (-not (Test-TargetPackageVersion -TargetDir $VenvT5_550Dir -PackageName "transformers" -ExpectedVersion "5.5.0")) { $_NeedT5Install = $true }
if (-not (Test-TargetPackageVersion -TargetDir $VenvT5_510Dir -PackageName "transformers" -ExpectedVersion "5.10.2")) { $_NeedT5Install = $true }
# Also reinstall when python deps were updated
if (-not $SkipPythonDeps) { $_NeedT5Install = $true }
if ($_NeedT5Install) {
Write-Host ""
$prevEAP_t5 = $ErrorActionPreference
$ErrorActionPreference = "Continue"
# --- .venv_t5_530 (transformers 5.3.0) ---
substep "pre-installing transformers 5.3.0 for newer model support..."
Assert-StudioOwnedOrAbsent -Path $VenvT5_530Dir -Label "transformers 5.3 sidecar venv"
if (Test-Path -LiteralPath $VenvT5_530Dir) { Remove-Item -LiteralPath $VenvT5_530Dir -Recurse -Force }
[System.IO.Directory]::CreateDirectory($VenvT5_530Dir) | Out-Null
Mark-StudioOwned -Path $VenvT5_530Dir
foreach ($pkg in @("transformers==5.3.0", "huggingface_hub==1.8.0", "hf_xet==1.4.2")) {
if ($script:UnslothVerbose) {
Fast-Install --target $VenvT5_530Dir --no-deps $pkg
$t5PkgExit = $LASTEXITCODE
$output = ""
} else {
$output = Fast-Install --target $VenvT5_530Dir --no-deps $pkg | Out-String
$t5PkgExit = $LASTEXITCODE
}
if ($t5PkgExit -ne 0) {
Write-Host "[FAIL] Could not install $pkg into .venv_t5_530/" -ForegroundColor Red
Write-Host (Redact-InstallOutput $output) -ForegroundColor Red
$ErrorActionPreference = $prevEAP_t5
exit 1
}
}
if ($script:UnslothVerbose) {
Fast-Install --target $VenvT5_530Dir tiktoken
$tiktokenInstallExit = $LASTEXITCODE
$output = ""
} else {
$output = Fast-Install --target $VenvT5_530Dir tiktoken | Out-String
$tiktokenInstallExit = $LASTEXITCODE
}
if ($tiktokenInstallExit -ne 0) {
substep "Could not install tiktoken into .venv_t5_530/ -- Qwen tokenizers may fail" "Yellow"
}
step "transformers" "5.3.0 pre-installed"
# --- .venv_t5_550 (transformers 5.5.0) ---
substep "pre-installing transformers 5.5.0 for Gemma 4 support..."
Assert-StudioOwnedOrAbsent -Path $VenvT5_550Dir -Label "transformers 5.5 sidecar venv"
if (Test-Path -LiteralPath $VenvT5_550Dir) { Remove-Item -LiteralPath $VenvT5_550Dir -Recurse -Force }
[System.IO.Directory]::CreateDirectory($VenvT5_550Dir) | Out-Null
Mark-StudioOwned -Path $VenvT5_550Dir
foreach ($pkg in @("transformers==5.5.0", "huggingface_hub==1.8.0", "hf_xet==1.4.2")) {
if ($script:UnslothVerbose) {
Fast-Install --target $VenvT5_550Dir --no-deps $pkg
$t5PkgExit = $LASTEXITCODE
$output = ""
} else {
$output = Fast-Install --target $VenvT5_550Dir --no-deps $pkg | Out-String
$t5PkgExit = $LASTEXITCODE
}
if ($t5PkgExit -ne 0) {
Write-Host "[FAIL] Could not install $pkg into .venv_t5_550/" -ForegroundColor Red
Write-Host (Redact-InstallOutput $output) -ForegroundColor Red
$ErrorActionPreference = $prevEAP_t5
exit 1
}
}
if ($script:UnslothVerbose) {
Fast-Install --target $VenvT5_550Dir tiktoken
$tiktokenInstallExit = $LASTEXITCODE
$output = ""
} else {
$output = Fast-Install --target $VenvT5_550Dir tiktoken | Out-String
$tiktokenInstallExit = $LASTEXITCODE
}
if ($tiktokenInstallExit -ne 0) {
substep "Could not install tiktoken into .venv_t5_550/ -- Qwen tokenizers may fail" "Yellow"
}
step "transformers" "5.5.0 pre-installed"
# --- .venv_t5_510 (transformers 5.10.2) ---
substep "pre-installing transformers 5.10.2 for Gemma 4 Unified support..."
Assert-StudioOwnedOrAbsent -Path $VenvT5_510Dir -Label "transformers 5.10 sidecar venv"
if (Test-Path -LiteralPath $VenvT5_510Dir) { Remove-Item -LiteralPath $VenvT5_510Dir -Recurse -Force }
[System.IO.Directory]::CreateDirectory($VenvT5_510Dir) | Out-Null
Mark-StudioOwned -Path $VenvT5_510Dir
foreach ($pkg in @("transformers==5.10.2", "huggingface_hub==1.8.0", "hf_xet==1.4.2")) {
if ($script:UnslothVerbose) {
Fast-Install --target $VenvT5_510Dir --no-deps $pkg
$t5PkgExit = $LASTEXITCODE
$output = ""
} else {
$output = Fast-Install --target $VenvT5_510Dir --no-deps $pkg | Out-String
$t5PkgExit = $LASTEXITCODE
}
if ($t5PkgExit -ne 0) {
Write-Host "[FAIL] Could not install $pkg into .venv_t5_510/" -ForegroundColor Red
Write-Host (Redact-InstallOutput $output) -ForegroundColor Red
$ErrorActionPreference = $prevEAP_t5
exit 1
}
}
if ($script:UnslothVerbose) {
Fast-Install --target $VenvT5_510Dir tiktoken
$tiktokenInstallExit = $LASTEXITCODE
$output = ""
} else {
$output = Fast-Install --target $VenvT5_510Dir tiktoken | Out-String
$tiktokenInstallExit = $LASTEXITCODE
}
if ($tiktokenInstallExit -ne 0) {
substep "Could not install tiktoken into .venv_t5_510/ -- Qwen tokenizers may fail" "Yellow"
}
$ErrorActionPreference = $prevEAP_t5
step "transformers" "5.10.2 pre-installed"
} # end $_NeedT5Install
# ==========================================================================
# PHASE 3.4: Prefer prebuilt llama.cpp bundles before source build
# ==========================================================================
# Nest llama.cpp under $StudioHome only for real env-overrides, never the
# legacy default. Reuses $StudioHomeIsCustom from the canonical comparison
# computed above so the llama.cpp nest matches ownership-guard semantics.
if ($StudioHomeIsCustom) {
$UnslothHome = $StudioHome
} else {
$UnslothHome = Join-Path $env:USERPROFILE ".unsloth"
}
if (-not (Test-Path -LiteralPath $UnslothHome)) { [System.IO.Directory]::CreateDirectory($UnslothHome) | Out-Null }
$LlamaCppDir = Join-Path $UnslothHome "llama.cpp"
$NeedLlamaSourceBuild = $false
$SkipPrebuiltInstall = $false
$RequestedLlamaTag = if ($env:UNSLOTH_LLAMA_TAG) { $env:UNSLOTH_LLAMA_TAG } else { $DefaultLlamaTag }
# Every host installs the fork's app-* prebuilts now: GPU Windows (CUDA / ROCm)
# already did, and the fork now also ships the CPU bundles for Windows x64 and
# arm64 (windows-cpu / windows-arm64). ggml-org artifacts are no longer used by
# default. Mirrors setup.sh's routing.
$HelperReleaseRepo = "unslothai/llama.cpp"
$LlamaPr = if ($env:UNSLOTH_LLAMA_PR) { $env:UNSLOTH_LLAMA_PR.Trim() } else { "" }
$LlamaPrForce = if ($env:UNSLOTH_LLAMA_PR_FORCE) { $env:UNSLOTH_LLAMA_PR_FORCE.Trim() } else { $DefaultLlamaPrForce }
$LlamaSource = $DefaultLlamaSource
if ($LlamaSource.EndsWith('.git')) { $LlamaSource = $LlamaSource.Substring(0, $LlamaSource.Length - 4) }
$ResolvedSourceUrl = $LlamaSource
$ResolvedSourceRef = $RequestedLlamaTag
$ResolvedSourceRefKind = "tag"
$ResolvedLlamaTag = $RequestedLlamaTag
if ($env:UNSLOTH_LLAMA_FORCE_COMPILE -eq "1") {
$NeedLlamaSourceBuild = $true
$SkipPrebuiltInstall = $true
}
function Invoke-LlamaHelper {
param(
[string[]]$Arguments,
[string]$StderrPath = $null
)
$previousErrorActionPreference = $ErrorActionPreference
$previousNativeErrorPreference = $null
$restoreNativeErrorPreference = $false
$ErrorActionPreference = "Continue"
if ($PSVersionTable.PSVersion.Major -ge 7) {
$previousNativeErrorPreference = $PSNativeCommandUseErrorActionPreference
$PSNativeCommandUseErrorActionPreference = $false
$restoreNativeErrorPreference = $true
}
try {
# Capture all output (stdout + stderr) so that PowerShell does not
# convert stderr lines into visible ErrorRecord objects. Separate
# stdout from stderr afterwards.
$allOutput = & python "$PSScriptRoot\install_llama_prebuilt.py" @Arguments 2>&1
$exitCode = $LASTEXITCODE
$stdoutLines = @()
$stderrLines = @()
foreach ($line in $allOutput) {
if ($line -is [System.Management.Automation.ErrorRecord]) {
$stderrLines += $line.ToString()
} else {
$stdoutLines += $line
}
}
if ($StderrPath -and $stderrLines.Count -gt 0) {
$stderrLines | Out-File -FilePath $StderrPath -Encoding utf8
}
return [pscustomobject]@{
Output = $stdoutLines
ExitCode = $exitCode
}
} finally {
$ErrorActionPreference = $previousErrorActionPreference
if ($restoreNativeErrorPreference) {
$PSNativeCommandUseErrorActionPreference = $previousNativeErrorPreference
}
}
}
if ($LlamaSource -ne "https://github.com/ggml-org/llama.cpp") {
step "llama.cpp" "custom source: $LlamaSource -- forcing source build" "Yellow"
$NeedLlamaSourceBuild = $true
$SkipPrebuiltInstall = $true
}
if (-not $LlamaPr -and $LlamaPrForce -and $LlamaPrForce -match '^\d+$' -and [int]$LlamaPrForce -gt 0) {
$LlamaPr = $LlamaPrForce
step "llama.cpp" "baked-in PR_FORCE=$LlamaPrForce" "Yellow"
}
if ($LlamaPr) {
if ($LlamaPr -notmatch '^\d+$' -or [int]$LlamaPr -le 0) {
Write-Host "[ERROR] UNSLOTH_LLAMA_PR=$LlamaPr is not a valid PR number" -ForegroundColor Red
exit 1
}
step "llama.cpp" "UNSLOTH_LLAMA_PR=$LlamaPr -- will build from PR head" "Yellow"
$ResolvedLlamaTag = "pr-$LlamaPr"
$ResolvedSourceUrl = $LlamaSource
$ResolvedSourceRef = "pr-$LlamaPr"
$ResolvedSourceRefKind = "pull"
$NeedLlamaSourceBuild = $true
$SkipPrebuiltInstall = $true
}
$LocalLlamaCppLinked = $false
$LocalLlamaCppSrc = $env:UNSLOTH_LOCAL_LLAMA_CPP_DIR
if ($LocalLlamaCppSrc) {
if (-not (Test-Path -LiteralPath $LocalLlamaCppSrc -PathType Container)) {
step "llama.cpp" "UNSLOTH_LOCAL_LLAMA_CPP_DIR does not exist: $LocalLlamaCppSrc" "Red"
exit 1
}
$ResolvedLocal = (Resolve-Path -LiteralPath $LocalLlamaCppSrc).Path
# Reusing a local dir disables both the prebuilt download and the source
# build, so a runnable llama-server.exe must already be present. Accept any
# layout LlamaCppBackend._layout_candidates() resolves (root-level, build\bin,
# or build\bin\Release) so the flag never rejects a tree Unsloth could run.
$LocalLlamaServerFound = $false
foreach ($_cand in @(
(Join-Path $ResolvedLocal "llama-server.exe"),
(Join-Path $ResolvedLocal "build\bin\llama-server.exe"),
(Join-Path $ResolvedLocal "build\bin\Release\llama-server.exe"))) {
if (Test-Path -LiteralPath $_cand) { $LocalLlamaServerFound = $true; break }
}
if ($ResolvedLocal -eq $LlamaCppDir) {
# Points at the canonical install location itself: never delete-then-link
# onto itself. Reuse an existing build here (skip prebuilt + source) so the
# staged prebuilt installer can't replace a build the user asked to reuse;
# if nothing is built yet, fall through to the normal install.
if ($LocalLlamaServerFound) {
substep "UNSLOTH_LOCAL_LLAMA_CPP_DIR is the canonical install location and already holds a build; reusing it" "Yellow"
$LocalLlamaCppLinked = $true
$NeedLlamaSourceBuild = $false
} else {
substep "UNSLOTH_LOCAL_LLAMA_CPP_DIR points to the canonical install location with nothing built there yet; running the normal install" "Yellow"
}
} else {
# Fail clearly rather than junction an unbuilt or wrong-platform checkout
# and leave Unsloth with no usable binary.
if (-not $LocalLlamaServerFound) {
step "llama.cpp" "no llama-server.exe under $ResolvedLocal (looked for .\llama-server.exe, .\build\bin and .\build\bin\Release) -- build llama.cpp there first, or drop --with-llama-cpp-dir" "Red"
exit 1
}
# If the target is already a junction/symlink (e.g. a previous
# --with-llama-cpp-dir run), delete only the link via DirectoryInfo.Delete().
# Remove-Item -Recurse -Force on a reparse point can traverse the link and
# wipe the user's real llama.cpp directory on PowerShell 5.1. Dropping the
# stale link here also keeps the custom-home ownership check below idempotent.
# Use Get-Item -Force (not Test-Path): a *broken* junction whose target was
# moved/deleted makes Test-Path return false, which would leave the dangling
# link in place and make mklink below fail; Get-Item still resolves it so we
# can remove it and relink to a new valid directory.
$existing = Get-Item -LiteralPath $LlamaCppDir -Force -ErrorAction SilentlyContinue
if ($existing -and ($existing.Attributes -band [System.IO.FileAttributes]::ReparsePoint)) {
$existing.Delete()
}
if ($StudioHomeIsCustom) {
Assert-StudioOwnedOrAbsent -Path $LlamaCppDir -Label "llama.cpp install"
}
if (Test-Path -LiteralPath $LlamaCppDir) {
Remove-Item -Recurse -Force -LiteralPath $LlamaCppDir -ErrorAction SilentlyContinue
# A locked/in-use tree can silently survive removal (SilentlyContinue
# masks it). Don't then junction/copy over a half-present dir; mirror the
# prebuilt path's active-process handling and stop with a clear message.
if (Test-Path -LiteralPath $LlamaCppDir) {
step "llama.cpp" "install blocked by active llama.cpp process" "Yellow"
substep "Close Unsloth or other llama.cpp users and retry" "Yellow"
exit 3
}
}
cmd /c "mklink /J `"$LlamaCppDir`" `"$ResolvedLocal`"" 2>&1 | Out-Null
if ($LASTEXITCODE -ne 0) {
substep "Could not create directory junction; copying instead..." "Yellow"
Copy-Item -Recurse -LiteralPath $ResolvedLocal -Destination $LlamaCppDir
Remove-AgentInstructionFiles -Roots @($LlamaCppDir)
}
Write-Host ""
step "llama.cpp" "linked local directory: $ResolvedLocal"
$LocalLlamaCppLinked = $true
$NeedLlamaSourceBuild = $false
}
}
if ($LocalLlamaCppLinked) {
# local directory linked above; skip prebuilt install
} elseif ($env:UNSLOTH_LLAMA_FORCE_COMPILE -eq "1") {
Write-Host ""
substep "UNSLOTH_LLAMA_FORCE_COMPILE=1 -- skipping prebuilt llama.cpp install" "Yellow"
$NeedLlamaSourceBuild = $true
} elseif ($SkipPrebuiltInstall) {
Write-Host ""
substep "Skipping prebuilt install -- falling back to source build" "Yellow"
} else {
Write-Host ""
if (Test-Path -LiteralPath $LlamaCppDir) {
substep "Existing llama.cpp install detected -- validating staged prebuilt update before replacement"
# If the existing install is the wrong kind (e.g. windows-cpu on a ROCm
# machine that should have windows-rocm), remove it so the installer is
# forced to download the correct variant rather than skipping on tag match.
$existingMetaPath = Join-Path $LlamaCppDir "UNSLOTH_PREBUILT_INFO.json"
if (Test-Path $existingMetaPath) {
try {
$existingMeta = Get-Content $existingMetaPath -Raw | ConvertFrom-Json
$existingKind = $existingMeta.install_kind
# A ROCm host may legitimately carry the fork's windows-rocm bundle
# or the upstream windows-hip fallback, so accept either and never
# treat a valid ROCm install as mismatched. A name-inferred gfx
# arch (Adrenalin-only, no confirmed runtime) still counts as
# ROCm-capable -- the ROCm prebuilt bundles its own runtime,
# mirroring the --rocm-gfx forward below. The CPU branch covers both
# the x64 windows-cpu and arm64 windows-arm64 bundles (Windows arm64
# has no GPU prebuilt). NOTE: this block is currently inert --
# write_prebuilt_metadata does not persist an install_kind key, so
# $existingKind is always null; keep $expectedKinds in sync with the
# kinds install_llama_prebuilt.py installs before relying on it.
$expectedKinds = if ($HasROCm -or $script:ROCmGfxArch) { @("windows-rocm", "windows-hip") } elseif ($HasNvidiaSmi) { @("windows-cuda") } else { @("windows-cpu", "windows-arm64") }
if ($existingKind -and ($existingKind -notin $expectedKinds)) {
substep "Removing mismatched llama.cpp install (found '$existingKind', need one of: $($expectedKinds -join ', '))..."
Remove-Item -Recurse -Force -LiteralPath $LlamaCppDir -ErrorAction SilentlyContinue
}
} catch {
# unreadable metadata -- let the installer handle it
}
}
}
substep "installing prebuilt llama.cpp bundle (preferred path)..."
# why: install_llama_prebuilt.py uses os.replace(), which would displace
# an unrelated $env:UNSLOTH_STUDIO_HOME\llama.cpp before the source-build
# ownership check below ever runs.
if ($StudioHomeIsCustom) {
Assert-StudioOwnedOrAbsent -Path $LlamaCppDir -Label "llama.cpp install"
}
$prebuiltArgs = @(
"$PSScriptRoot\install_llama_prebuilt.py",
"--install-dir", $LlamaCppDir,
"--llama-tag", $RequestedLlamaTag,
"--published-repo", $HelperReleaseRepo
)
if ($HasROCm) {
$prebuiltArgs += "--has-rocm"
}
# Forward the resolved gfx arch so the per-gfx ROCm prebuilt is picked even
# when the installer's probe can't confirm the runtime (amd-smi-only /
# Adrenalin-only, name-inferred arch). --rocm-gfx is authoritative and
# implies ROCm in install_llama_prebuilt.py, so the GPU prebuilt is selected
# even with $HasROCm false. Gating on $HasROCm gave Strix Halo / 8060S CPU.
if ($script:ROCmGfxArch) {
$prebuiltArgs += @("--rocm-gfx", $script:ROCmGfxArch)
}
if ($env:UNSLOTH_LLAMA_RELEASE_TAG) {
$prebuiltArgs += @("--published-release-tag", $env:UNSLOTH_LLAMA_RELEASE_TAG)
}
# UNSLOTH_LLAMA_CPP_BACKEND=cpu (case-insensitive, whitespace-trimmed) forces the
# CPU-only prebuilt via --force-cpu (persisted so updates keep it). Fixes Intel
# iGPU Vulkan crash (#7213).
$llamaBackend = "$($env:UNSLOTH_LLAMA_CPP_BACKEND)".Trim().ToLowerInvariant()
if ($llamaBackend -eq "cpu") {
$prebuiltArgs += "--force-cpu"
} elseif ($llamaBackend -and $llamaBackend -ne "auto") {
Write-Host "[WARN] Ignoring UNSLOTH_LLAMA_CPP_BACKEND='$($env:UNSLOTH_LLAMA_CPP_BACKEND)' (expected 'auto' or 'cpu')" -ForegroundColor Yellow
}
$prevEAPPrebuilt = $ErrorActionPreference
$ErrorActionPreference = "Continue"
$previousNativeErrorPreference = $null
$restoreNativeErrorPreference = $false
if ($PSVersionTable.PSVersion.Major -ge 7) {
$previousNativeErrorPreference = $PSNativeCommandUseErrorActionPreference
$PSNativeCommandUseErrorActionPreference = $false
$restoreNativeErrorPreference = $true
}
try {
if ($script:UnslothVerbose) {
# Show live output in verbose mode while still capturing for error log
$prebuiltLog = Join-Path $env:TEMP "unsloth-prebuilt-$PID.log"
& python @prebuiltArgs 2>&1 | Tee-Object -FilePath $prebuiltLog | Out-Host
$prebuiltExit = $LASTEXITCODE
$prebuiltOutput = if (Test-Path $prebuiltLog) { Get-Content $prebuiltLog -Raw } else { "" }
Remove-Item $prebuiltLog -ErrorAction SilentlyContinue
} else {
$prebuiltOutput = & python @prebuiltArgs 2>&1 | Out-String
$prebuiltExit = $LASTEXITCODE
}
} finally {
if ($restoreNativeErrorPreference) {
$PSNativeCommandUseErrorActionPreference = $previousNativeErrorPreference
}
}
$ErrorActionPreference = $prevEAPPrebuilt
if ($prebuiltExit -eq 0) {
if ($prebuiltOutput -match "already matches") {
step "llama.cpp" "prebuilt up to date and validated"
} else {
step "llama.cpp" "prebuilt installed and validated"
}
if ($StudioHomeIsCustom -and (Test-Path -LiteralPath $LlamaCppDir -PathType Container)) {
Mark-StudioOwned -Path $LlamaCppDir
}
$installedRelease = Get-InstalledLlamaPrebuiltRelease -InstallDir $LlamaCppDir
if ($installedRelease) {
substep $installedRelease
}
} elseif ($prebuiltExit -eq 3) {
step "llama.cpp" "install blocked by active llama.cpp process" "Yellow"
Write-LlamaFailureLog -Output $prebuiltOutput
if (Test-Path -LiteralPath $LlamaCppDir) {
substep "Existing install was restored" "Yellow"
}
substep "Close Unsloth or other llama.cpp users and retry" "Yellow"
exit 3
} else {
step "llama.cpp" "prebuilt install failed (continuing)" "Yellow"
Write-LlamaFailureLog -Output $prebuiltOutput
if (Test-Path -LiteralPath $LlamaCppDir) {
substep "Prebuilt update failed; existing install was restored or cleaned before source build fallback" "Yellow"
}
substep "Prebuilt llama.cpp path unavailable or failed validation -- falling back to source build" "Yellow"
$NeedLlamaSourceBuild = $true
}
}
# ==========================================================================
# PHASE 3.5: Install OpenSSL dev (for HTTPS support in llama-server)
# ==========================================================================
# llama-server needs OpenSSL to download models from HuggingFace via -hf.
# ShiningLight.OpenSSL.Dev includes headers + libs that cmake can find.
$OpenSslAvailable = $false
if ($NeedLlamaSourceBuild) {
# Check if OpenSSL dev is already installed (look for include dir)
$OpenSslRoots = @(
'C:\Program Files\OpenSSL-Win64',
'C:\Program Files\OpenSSL',
'C:\OpenSSL-Win64'
)
$OpenSslRoot = $null
foreach ($root in $OpenSslRoots) {
if (Test-Path (Join-Path $root 'include\openssl\ssl.h')) {
$OpenSslRoot = $root
break
}
}
if ($OpenSslRoot) {
$OpenSslAvailable = $true
substep "OpenSSL dev found at $OpenSslRoot"
} else {
Write-Host ""
substep "installing OpenSSL dev (for HTTPS in llama-server)..."
$HasWinget = $null -ne (Get-Command winget -ErrorAction SilentlyContinue)
if ($HasWinget) {
winget install -e --id ShiningLight.OpenSSL.Dev --source winget --accept-package-agreements --accept-source-agreements
# Re-check after install
foreach ($root in $OpenSslRoots) {
if (Test-Path (Join-Path $root 'include\openssl\ssl.h')) {
$OpenSslRoot = $root
$OpenSslAvailable = $true
substep "OpenSSL dev installed at $OpenSslRoot"
break
}
}
}
if (-not $OpenSslAvailable) {
substep "OpenSSL dev not available -- llama-server will be built without HTTPS" "Yellow"
}
}
} else {
substep "OpenSSL dev install skipped -- prebuilt llama.cpp already validated" "Yellow"
}
# ==========================================================================
# PHASE 4: Build llama.cpp with CUDA for GGUF inference + export
# ==========================================================================
# Builds at ~/.unsloth/llama.cpp — a single shared location under the user's
# home directory. This is used by both the inference server and the GGUF
# export pipeline (unsloth-zoo).
# We build:
# - llama-server: for GGUF model inference (with HTTPS if OpenSSL available)
# - llama-quantize: for GGUF export quantization
# Prerequisites git, cmake, VS Build Tools were installed in Phase 1; the CUDA
# Toolkit is resolved lazily just below via Resolve-CudaToolkit (source build only).
$OriginalLlamaCppDir = $LlamaCppDir
$BuildDir = Join-Path $LlamaCppDir "build"
$LlamaServerBin = Join-Path $BuildDir "bin\Release\llama-server.exe"
$HasCmakeForBuild = $null -ne (Get-Command cmake -ErrorAction SilentlyContinue)
# Check if existing llama-server matches current GPU mode. A CUDA-built binary
# on a now-CPU-only machine (or vice versa) needs to be rebuilt.
$NeedRebuild = $false
if (Test-Path -LiteralPath $LlamaServerBin) {
$CmakeCacheFile = Join-Path $BuildDir "CMakeCache.txt"
if (Test-Path -LiteralPath $CmakeCacheFile) {
$cachedCuda = Select-String -LiteralPath $CmakeCacheFile -Pattern 'GGML_CUDA:BOOL=ON' -Quiet
if ($HasNvidiaSmi -and -not $cachedCuda) {
Write-Host " Existing llama-server is CPU-only but GPU is available -- rebuilding" -ForegroundColor Yellow
$NeedRebuild = $true
} elseif (-not $HasNvidiaSmi -and $cachedCuda) {
Write-Host " Existing llama-server was built with CUDA but no GPU detected -- rebuilding" -ForegroundColor Yellow
$NeedRebuild = $true
}
}
}
# Install build tools now (last resort) rather than eagerly in Phase 1, so the
# prebuilt path stays fast. Same condition as the if/elseif chain below: a source
# build runs only when needed and no usable binary is already present. A linked
# local dir sets $NeedLlamaSourceBuild = $false, so this no-ops for that path.
$WillBuildLlamaFromSource = $NeedLlamaSourceBuild -and `
-not ((Test-Path -LiteralPath $LlamaServerBin) -and -not $NeedRebuild -and $RequestedLlamaTag -ne "master")
if ($WillBuildLlamaFromSource) {
Ensure-BuildToolsForLlamaSourceBuild
# refresh so the chain below sees a newly installed cmake
$HasCmakeForBuild = $null -ne (Get-Command cmake -ErrorAction SilentlyContinue)
}
if ($LocalLlamaCppLinked) {
# Local dir linked above -- honor the flag's contract: skip BOTH the prebuilt
# download and the source build. Falling through here would run CMake inside
# the user's checkout (via the junction) when it lacks build\bin\Release\llama-server.exe.
Write-Host ""
step "llama.cpp" "linked (skipping build)"
} elseif (-not $NeedLlamaSourceBuild) {
Write-Host ""
step "llama.cpp" "prebuilt (validated)"
} elseif ((Test-Path -LiteralPath $LlamaServerBin) -and -not $NeedRebuild -and $RequestedLlamaTag -ne "master") {
# Skip rebuild only for pinned tags (e.g. b8635). When the requested
# tag is "master" (a moving target), always rebuild so the binary picks
# up new model architecture support (e.g. Gemma 4).
Write-Host ""
step "llama.cpp" "already built"
} elseif (-not $HasCmakeForBuild) {
Write-Host ""
if (-not $HasNvidiaSmi) {
# CPU-only machines depend entirely on llama-server for GGUF chat -- cmake is required
substep "CMake is required to build llama-server for GGUF chat mode." "Yellow"
substep "Continuing setup without llama.cpp build." "Yellow"
substep "Install CMake from https://cmake.org/download/ and re-run setup." "Yellow"
}
step "llama.cpp" "build skipped (cmake not available)" "Yellow"
substep "GGUF inference and export will not be available." "Yellow"
substep "Install CMake from https://cmake.org/download/ and re-run setup." "Yellow"
$script:LlamaCppDegraded = $true
} else {
# Finalize the VS generator (gate/fallback below) BEFORE Resolve-CudaToolkit,
# which copies the CUDA .targets into the current generator's dir; a later swap
# would strand them. The CMake 4.2 gate for VS 2026 is checked only here, in the
# source-build path, so a VS 2026 + cmake < 4.2 host can still use the prebuilt. (#6473)
if ($CmakeGenerator -match 'Visual Studio 18\b') {
if (-not (Test-CmakeCanDriveGenerator -Generator $CmakeGenerator)) {
$cmakeVerObj = Get-CmakeVersion
$cmakeVerStr = if ($cmakeVerObj) { $cmakeVerObj.ToString() } else { '0.0' }
substep "CMake $cmakeVerStr cannot drive the Visual Studio 2026 generator (need 4.2+ or a VS-bundled cmake) -- updating via winget..." "Yellow"
if ($null -ne (Get-Command winget -ErrorAction SilentlyContinue)) {
# upgrade first (fast if Kitware.CMake is already a winget app), then
# prepend the default dir so the new cmake wins over an older one on PATH
try {
Invoke-SetupCommand { winget upgrade Kitware.CMake --source winget --accept-package-agreements --accept-source-agreements } | Out-Null
Refresh-Environment
} catch { substep "CMake winget upgrade failed: $($_.Exception.Message)" "Yellow" }
Add-DefaultCmakeToPath | Out-Null
# upgrade no-ops if the cmake came from Scoop/Chocolatey/VS, not the
# Kitware winget package; install it so a 4.2+ cmake is available
if (-not (Test-CmakeCanDriveGenerator -Generator $CmakeGenerator)) {
try {
Invoke-SetupCommand { winget install Kitware.CMake --source winget --accept-package-agreements --accept-source-agreements } | Out-Null
Refresh-Environment
} catch { substep "CMake winget install failed: $($_.Exception.Message)" "Yellow" }
Add-DefaultCmakeToPath | Out-Null
}
}
if (-not (Test-CmakeCanDriveGenerator -Generator $CmakeGenerator)) {
# cmake still cannot drive VS 2026; before failing, fall back to an
# older installed VS whose generator it can drive (e.g. VS 2022 + old
# cmake on an offline box keeps building)
$fallback = Get-FallbackVsGenerator
if ($fallback) {
substep "CMake cannot drive $CmakeGenerator; falling back to $($fallback.Generator)" "Yellow"
$CmakeGenerator = $fallback.Generator
$VsInstallPath = $fallback.InstallPath
} else {
Write-Host "[ERROR] CMake 4.2+ is required to build llama.cpp with the Visual Studio 2026 generator, and no older Visual Studio toolchain was found to fall back to." -ForegroundColor Red
Write-Host " Upgrade CMake from https://cmake.org/download/ and re-run, or use a prebuilt llama.cpp bundle." -ForegroundColor Red
exit 1
}
}
}
substep "CMake can drive the $CmakeGenerator generator"
}
# CUDA resolved here (fail fast if none), after the final VS generator so its
# .targets land in the toolset cmake actually uses.
if ($HasNvidiaSmi) { Resolve-CudaToolkit -RequireOrExit }
Write-Host ""
if ($HasNvidiaSmi) {
substep "building llama.cpp with CUDA support..."
} elseif ($HasROCm -or $script:ROCmGfxArch) {
# AMD GPU present but in the CPU-only source-build fallback: a HIP source
# build needs the full HIP SDK + ROCm clang toolchain. AMD GPU acceleration
# comes from the per-gfx ROCm prebuilt (bundles the runtime, no SDK) -- reaching
# here means it couldn't be installed. Warn loudly, don't ship a slow CPU build.
$_amdArch = if ($script:ROCmGfxArch) { $script:ROCmGfxArch } else { "ROCm" }
substep "[WARN] AMD GPU ($_amdArch) detected, but the GPU-accelerated ROCm" "Yellow"
substep " llama.cpp prebuilt could not be installed -- falling back to a CPU build." "Yellow"
substep " The prebuilt is the AMD GPU path (no HIP SDK required). To restore GPU" "Yellow"
substep " acceleration: re-run the installer (check your network / proxy), or set" "Yellow"
substep " UNSLOTH_LLAMA_RELEASE_TAG to a tag with a gfx prebuilt for your GPU." "Yellow"
substep "building llama.cpp (CPU-only fallback)..." "Yellow"
} else {
substep "building llama.cpp (CPU-only, no NVIDIA GPU detected)..."
}
substep "This typically takes 5-10 minutes on first build."
Write-Host ""
# Start total build timer
$totalSw = [System.Diagnostics.Stopwatch]::StartNew()
# Native commands (git, cmake) write to stderr even on success.
# With $ErrorActionPreference = "Stop" (set at top of script), PS 5.1
# converts stderr lines into terminating ErrorRecords, breaking output.
# Lower to "Continue" for the build section.
$prevEAP = $ErrorActionPreference
$ErrorActionPreference = "Continue"
$BuildOk = $true
$FailedStep = ""
# Re-sanitize CUDA_PATH_V* vars — Refresh-Environment (called during
# Node/Python installs above) may have repopulated conflicting versioned
# vars from the Machine registry.
if ($HasNvidiaSmi -and $CudaToolkitRoot) {
$cudaPathVars2 = @([Environment]::GetEnvironmentVariables('Process').Keys | Where-Object { $_ -match '^CUDA_PATH_V' })
foreach ($v2 in $cudaPathVars2) {
[Environment]::SetEnvironmentVariable($v2, $null, 'Process')
}
$tkDirName2 = Split-Path $CudaToolkitRoot -Leaf
if ($tkDirName2 -match '^v(\d+)\.(\d+)') {
[Environment]::SetEnvironmentVariable("CUDA_PATH_V$($Matches[1])_$($Matches[2])", $CudaToolkitRoot, 'Process')
}
# Also re-assert CUDA_PATH and CudaToolkitDir in case they were overwritten
[Environment]::SetEnvironmentVariable('CUDA_PATH', $CudaToolkitRoot, 'Process')
[Environment]::SetEnvironmentVariable('CudaToolkitDir', "$CudaToolkitRoot\", 'Process')
}
if (-not $LlamaPr) {
$ResolvedSourceUrl = $LlamaSource
if ($env:UNSLOTH_LLAMA_FORCE_COMPILE -eq "1") {
if ($RequestedLlamaTag -eq "latest") {
$ResolvedSourceRef = if ($env:UNSLOTH_LLAMA_FORCE_COMPILE_REF) {
$env:UNSLOTH_LLAMA_FORCE_COMPILE_REF
} else {
$DefaultLlamaForceCompileRef
}
$ResolvedSourceRefKind = "branch"
} else {
$ResolvedSourceRef = $RequestedLlamaTag
$ResolvedSourceRefKind = "tag"
}
} elseif ($RequestedLlamaTag -eq "latest") {
$resolveTagArgs = @("--resolve-llama-tag", "latest", "--published-repo", "ggml-org/llama.cpp", "--output-format", "json")
$resolveTagResult = Invoke-LlamaHelper -Arguments $resolveTagArgs
$resolveTagOutput = $resolveTagResult.Output
$resolveTagExit = $resolveTagResult.ExitCode
if ($resolveTagExit -eq 0 -and $resolveTagOutput) {
try {
$ResolvedSourceRef = (($resolveTagOutput | Out-String) | ConvertFrom-Json).llama_tag
} catch {
$ResolvedSourceRef = ""
}
} else {
$ResolvedSourceRef = ""
}
if ([string]::IsNullOrWhiteSpace($ResolvedSourceRef)) {
$ResolvedSourceRef = "latest"
}
$ResolvedSourceRefKind = "tag"
} else {
$ResolvedSourceRef = $RequestedLlamaTag
$ResolvedSourceRefKind = "tag"
}
if ([string]::IsNullOrWhiteSpace($ResolvedSourceUrl)) { $ResolvedSourceUrl = $LlamaSource }
if ([string]::IsNullOrWhiteSpace($ResolvedSourceRef)) { $ResolvedSourceRef = $RequestedLlamaTag }
}
# -- Step A: Clone or pull llama.cpp --
$UseConcreteRef = ($ResolvedSourceRef -ne "latest" -and -not [string]::IsNullOrWhiteSpace($ResolvedSourceRef))
if (Test-Path -LiteralPath (Join-Path $LlamaCppDir ".git")) {
# why: in-place git mutation (remote set-url, checkout -B, clean -fdx)
# rewrites $LlamaCppDir; mirror the prebuilt and temp-dir-swap guards
# so an unrelated workspace .git tree is never silently overwritten.
if ($StudioHomeIsCustom) {
Assert-StudioOwnedOrAbsent -Path $LlamaCppDir -Label "llama.cpp install"
}
Write-Host " Syncing llama.cpp to $ResolvedSourceRef..." -ForegroundColor Gray
# Always sync the remote URL so switching between default/fork sources works
Invoke-SetupCommand -AlwaysQuiet { git -C $LlamaCppDir remote set-url origin "$ResolvedSourceUrl.git" } | Out-Null
if ($LlamaPr) {
$gitFetchExit = Invoke-SetupCommand -AlwaysQuiet { git -C $LlamaCppDir fetch --depth 1 origin "pull/$LlamaPr/head" }
if ($gitFetchExit -ne 0) {
$BuildOk = $false
$FailedStep = "git fetch PR #$LlamaPr"
} else {
$gitCheckoutExit = Invoke-SetupCommand -AlwaysQuiet { git -C $LlamaCppDir checkout -B "pr-$LlamaPr" FETCH_HEAD }
if ($gitCheckoutExit -ne 0) {
$BuildOk = $false
$FailedStep = "git checkout PR #$LlamaPr"
} else {
Invoke-SetupCommand -AlwaysQuiet { git -C $LlamaCppDir clean -fdx } | Out-Null
}
}
} elseif ($ResolvedSourceRefKind -eq "pull") {
$gitFetchExit = Invoke-SetupCommand -AlwaysQuiet { git -C $LlamaCppDir fetch --depth 1 origin $ResolvedSourceRef }
if ($gitFetchExit -ne 0) {
substep "git fetch failed -- using existing source" "Yellow"
} else {
$gitCheckoutExit = Invoke-SetupCommand -AlwaysQuiet { git -C $LlamaCppDir checkout -B unsloth-llama-build FETCH_HEAD }
if ($gitCheckoutExit -ne 0) {
$BuildOk = $false
$FailedStep = "git checkout"
} else {
Invoke-SetupCommand -AlwaysQuiet { git -C $LlamaCppDir clean -fdx } | Out-Null
}
}
} elseif ($ResolvedSourceRefKind -eq "commit") {
$gitFetchExit = Invoke-SetupCommand -AlwaysQuiet { git -C $LlamaCppDir fetch --depth 1 origin $ResolvedSourceRef }
if ($gitFetchExit -ne 0) {
substep "git fetch failed -- using existing source" "Yellow"
} else {
$gitCheckoutExit = Invoke-SetupCommand -AlwaysQuiet { git -C $LlamaCppDir checkout -B unsloth-llama-build FETCH_HEAD }
if ($gitCheckoutExit -ne 0) {
$BuildOk = $false
$FailedStep = "git checkout"
} else {
Invoke-SetupCommand -AlwaysQuiet { git -C $LlamaCppDir clean -fdx } | Out-Null
}
}
} elseif ($UseConcreteRef) {
$gitFetchExit = Invoke-SetupCommand -AlwaysQuiet { git -C $LlamaCppDir fetch --depth 1 origin $ResolvedSourceRef }
if ($gitFetchExit -ne 0) {
substep "git fetch failed -- using existing source" "Yellow"
} else {
$gitCheckoutExit = Invoke-SetupCommand -AlwaysQuiet { git -C $LlamaCppDir checkout -B unsloth-llama-build FETCH_HEAD }
if ($gitCheckoutExit -ne 0) {
$BuildOk = $false
$FailedStep = "git checkout"
} else {
Invoke-SetupCommand -AlwaysQuiet { git -C $LlamaCppDir clean -fdx } | Out-Null
}
}
} else {
$gitFetchExit = Invoke-SetupCommand -AlwaysQuiet { git -C $LlamaCppDir fetch --depth 1 origin }
if ($gitFetchExit -ne 0) {
substep "git fetch failed -- using existing source" "Yellow"
} else {
$gitCheckoutExit = Invoke-SetupCommand -AlwaysQuiet { git -C $LlamaCppDir checkout -B unsloth-llama-build FETCH_HEAD }
if ($gitCheckoutExit -ne 0) {
$BuildOk = $false
$FailedStep = "git checkout"
} else {
Invoke-SetupCommand -AlwaysQuiet { git -C $LlamaCppDir clean -fdx } | Out-Null
}
}
}
# why: in-place git-sync (the temp-dir clone path calls Mark-StudioOwned
# at swap-time) must mark the existing tree so a subsequent prebuilt
# update path's Assert-StudioOwnedOrAbsent does not exit on the same root.
if ($BuildOk -and $StudioHomeIsCustom) {
Mark-StudioOwned -Path $LlamaCppDir
}
} else {
Write-Host " Cloning llama.cpp @ $ResolvedSourceRef..." -ForegroundColor Gray
$buildTmp = "$LlamaCppDir.build.$PID"
$null = [System.IO.Directory]::CreateDirectory((Split-Path -LiteralPath $LlamaCppDir))
if (Test-Path -LiteralPath $buildTmp) { Remove-Item -LiteralPath $buildTmp -Recurse -Force }
if ($LlamaPr) {
$cloneExit = Invoke-SetupCommand -AlwaysQuiet { git clone --depth 1 "$LlamaSource.git" $buildTmp }
if ($cloneExit -ne 0) {
$BuildOk = $false
$FailedStep = "git clone"
if (Test-Path -LiteralPath $buildTmp) { Remove-Item -LiteralPath $buildTmp -Recurse -Force }
}
if ($BuildOk) {
$fetchExit = Invoke-SetupCommand -AlwaysQuiet { git -C $buildTmp fetch --depth 1 origin "pull/$LlamaPr/head:pr-$LlamaPr" }
if ($fetchExit -ne 0) {
$BuildOk = $false
$FailedStep = "git fetch PR #$LlamaPr"
if (Test-Path -LiteralPath $buildTmp) { Remove-Item -LiteralPath $buildTmp -Recurse -Force }
}
}
if ($BuildOk) {
$checkoutExit = Invoke-SetupCommand -AlwaysQuiet { git -C $buildTmp checkout "pr-$LlamaPr" }
if ($checkoutExit -ne 0) {
$BuildOk = $false
$FailedStep = "git checkout PR #$LlamaPr"
if (Test-Path -LiteralPath $buildTmp) { Remove-Item -LiteralPath $buildTmp -Recurse -Force }
}
}
} elseif ($ResolvedSourceRefKind -eq "pull") {
$cloneExit = Invoke-SetupCommand -AlwaysQuiet { git clone --depth 1 "$ResolvedSourceUrl.git" $buildTmp }
if ($cloneExit -ne 0) {
$BuildOk = $false
$FailedStep = "git clone"
if (Test-Path -LiteralPath $buildTmp) { Remove-Item -LiteralPath $buildTmp -Recurse -Force }
}
if ($BuildOk) {
$fetchExit = Invoke-SetupCommand -AlwaysQuiet { git -C $buildTmp fetch --depth 1 origin $ResolvedSourceRef }
if ($fetchExit -ne 0) {
$BuildOk = $false
$FailedStep = "git fetch source PR ref"
if (Test-Path -LiteralPath $buildTmp) { Remove-Item -LiteralPath $buildTmp -Recurse -Force }
}
}
if ($BuildOk) {
$checkoutExit = Invoke-SetupCommand -AlwaysQuiet { git -C $buildTmp checkout -B unsloth-llama-build FETCH_HEAD }
if ($checkoutExit -ne 0) {
$BuildOk = $false
$FailedStep = "git checkout source PR ref"
if (Test-Path -LiteralPath $buildTmp) { Remove-Item -LiteralPath $buildTmp -Recurse -Force }
}
}
} elseif ($ResolvedSourceRefKind -eq "commit") {
$cloneExit = Invoke-SetupCommand -AlwaysQuiet { git clone --depth 1 "$ResolvedSourceUrl.git" $buildTmp }
if ($cloneExit -ne 0) {
$BuildOk = $false
$FailedStep = "git clone"
if (Test-Path -LiteralPath $buildTmp) { Remove-Item -LiteralPath $buildTmp -Recurse -Force }
}
if ($BuildOk) {
$fetchExit = Invoke-SetupCommand -AlwaysQuiet { git -C $buildTmp fetch --depth 1 origin $ResolvedSourceRef }
if ($fetchExit -ne 0) {
$BuildOk = $false
$FailedStep = "git fetch source commit"
if (Test-Path -LiteralPath $buildTmp) { Remove-Item -LiteralPath $buildTmp -Recurse -Force }
}
}
if ($BuildOk) {
$checkoutExit = Invoke-SetupCommand -AlwaysQuiet { git -C $buildTmp checkout -B unsloth-llama-build FETCH_HEAD }
if ($checkoutExit -ne 0) {
$BuildOk = $false
$FailedStep = "git checkout source commit"
if (Test-Path -LiteralPath $buildTmp) { Remove-Item -LiteralPath $buildTmp -Recurse -Force }
}
}
} else {
$cloneArgs = @("clone", "--depth", "1")
if ($UseConcreteRef) {
$cloneArgs += @("--branch", $ResolvedSourceRef)
}
$cloneArgs += @("$ResolvedSourceUrl.git", $buildTmp)
$cloneExit = Invoke-SetupCommand -AlwaysQuiet { git @cloneArgs }
if ($cloneExit -ne 0) {
$BuildOk = $false
$FailedStep = "git clone"
if (Test-Path -LiteralPath $buildTmp) { Remove-Item -LiteralPath $buildTmp -Recurse -Force }
}
}
# Use temp dir for build; swap into $LlamaCppDir only after build succeeds
if ($BuildOk) {
$LlamaCppDir = $buildTmp
$BuildDir = Join-Path $LlamaCppDir "build"
}
}
# -- Step B: cmake configure --
if ($BuildOk) {
Write-Host ""
Write-Host "--- cmake configure ---" -ForegroundColor Cyan
$CmakeArgs = @(
'-S', $LlamaCppDir,
'-B', $BuildDir,
'-G', $CmakeGenerator,
'-Wno-dev'
)
# Tell cmake exactly where VS is (bypasses registry lookup)
if ($VsInstallPath) {
$CmakeArgs += "-DCMAKE_GENERATOR_INSTANCE=$VsInstallPath"
}
# Common flags
$CmakeArgs += '-DBUILD_SHARED_LIBS=OFF'
$CmakeArgs += '-DLLAMA_BUILD_TESTS=OFF'
$CmakeArgs += '-DLLAMA_BUILD_EXAMPLES=OFF'
$CmakeArgs += '-DLLAMA_BUILD_SERVER=ON'
$CmakeArgs += '-DGGML_NATIVE=ON'
# HTTPS support via OpenSSL
if ($OpenSslAvailable -and $OpenSslRoot) {
$CmakeArgs += "-DOPENSSL_ROOT_DIR=$OpenSslRoot"
$CmakeArgs += '-DLLAMA_OPENSSL=ON'
} else {
$CmakeArgs += '-DLLAMA_CURL=OFF'
}
$CmakeArgs += '-DCMAKE_EXE_LINKER_FLAGS=/NODEFAULTLIB:LIBCMT'
# CUDA flags -- only if GPU available, otherwise explicitly disable
if ($HasNvidiaSmi -and $NvccPath) {
# UNSLOTH_LLAMA_CUDA_ARCHS (e.g. "120" or "89;86") forces the build
# arch and wins over detection, matching setup.sh.
$CudaArchOverride = if ($env:UNSLOTH_LLAMA_CUDA_ARCHS) { ($env:UNSLOTH_LLAMA_CUDA_ARCHS -replace '\s', '') } else { '' }
if ((-not $CudaArch) -and (-not $CudaArchOverride)) {
# No detectable compute capability (#5854): -DGGML_CUDA=ON with no
# arch builds a PTX-only binary, so build CPU instead. Mirrors the
# Linux fix; set UNSLOTH_LLAMA_CUDA_ARCHS=120 to force a CUDA build.
substep "could not detect a CUDA compute capability; building CPU llama.cpp instead of a PTX-only binary (set UNSLOTH_LLAMA_CUDA_ARCHS=120 to force a CUDA build)." "Yellow"
$CmakeArgs += '-DGGML_CUDA=OFF'
} else {
$CmakeArgs += '-DGGML_CUDA=ON'
# Accept a host MSVC newer than nvcc's whitelist; a fresh toolkit
# (e.g. CUDA 13.3) otherwise aborts with "#error -- unsupported
# Microsoft Visual Studio version!". Mirrors the Linux fix. Via env
# (covers the configure probe + build), after Refresh-Environment, idempotent.
$nvccAllowFlag = '-allow-unsupported-compiler'
if ([string]::IsNullOrEmpty($env:NVCC_PREPEND_FLAGS)) {
$env:NVCC_PREPEND_FLAGS = $nvccAllowFlag
} elseif ($env:NVCC_PREPEND_FLAGS -notlike "*$nvccAllowFlag*") {
$env:NVCC_PREPEND_FLAGS = "$($env:NVCC_PREPEND_FLAGS) $nvccAllowFlag"
}
substep "NVCC_PREPEND_FLAGS = $env:NVCC_PREPEND_FLAGS"
$CmakeArgs += "-DCUDAToolkit_ROOT=$CudaToolkitRoot"
$CmakeArgs += "-DCUDA_TOOLKIT_ROOT_DIR=$CudaToolkitRoot"
$CmakeArgs += "-DCMAKE_CUDA_COMPILER=$NvccPath"
if ($CudaArchOverride) {
# Forced arch wins verbatim (no nvcc validation), matching setup.sh.
$CmakeArgs += "-DCMAKE_CUDA_ARCHITECTURES=$CudaArchOverride"
} elseif ($CudaArch) {
# Validate nvcc actually supports this architecture
if (Test-NvccArchSupport -NvccExe $NvccPath -Arch $CudaArch) {
$CmakeArgs += "-DCMAKE_CUDA_ARCHITECTURES=$CudaArch"
} else {
# GPU arch too new for this toolkit -- fall back to highest supported.
# PTX forward-compatibility will JIT-compile for the actual GPU at runtime.
$maxArch = Get-NvccMaxArch -NvccExe $NvccPath
if ($maxArch) {
$CmakeArgs += "-DCMAKE_CUDA_ARCHITECTURES=$maxArch"
substep "GPU is sm_$CudaArch but nvcc only supports up to sm_$maxArch" "Yellow"
substep "Building with sm_$maxArch (PTX will JIT for your GPU at runtime)" "Yellow"
}
# else: omit flag entirely, let cmake pick defaults
}
}
}
} else {
$CmakeArgs += '-DGGML_CUDA=OFF'
}
$cmakeOutput = cmake @CmakeArgs 2>&1 | Out-String
$cmakeConfigureExit = $LASTEXITCODE
if ($cmakeConfigureExit -ne 0) {
$BuildOk = $false
$FailedStep = "cmake configure"
Write-LlamaFailureLog -Output $cmakeOutput
if ($cmakeOutput -match 'No CUDA toolset found|CUDA_TOOLKIT_ROOT_DIR|nvcc') {
Write-Host ""
Write-Host " Hint: CUDA VS integration may be missing. Try running as admin:" -ForegroundColor Yellow
Write-Host " Copy contents of:" -ForegroundColor Yellow
Write-Host " <CUDA_PATH>\extras\visual_studio_integration\MSBuildExtensions" -ForegroundColor Yellow
Write-Host " into:" -ForegroundColor Yellow
$hintCustomizations = if ($VsInstallPath) { Get-VcBuildCustomizationsDir -VsInstallPath $VsInstallPath -Generator $CmakeGenerator } else { "<VS_PATH>\MSBuild\Microsoft\VC\v170\BuildCustomizations" }
Write-Host " $hintCustomizations" -ForegroundColor Yellow
}
}
}
# -- Step C: Build llama-server --
$NumCpu = [Environment]::ProcessorCount
if ($NumCpu -lt 1) { $NumCpu = 4 }
if ($BuildOk) {
Write-Host ""
Write-Host "--- cmake build (llama-server) ---" -ForegroundColor Cyan
Write-Host " Parallel jobs: $NumCpu" -ForegroundColor Gray
Write-Host ""
$output = cmake --build $BuildDir --config Release --target llama-server -j $NumCpu 2>&1 | Out-String
$cmakeBuildServerExit = $LASTEXITCODE
if ($cmakeBuildServerExit -ne 0) {
$BuildOk = $false
$FailedStep = "cmake build (llama-server)"
Write-LlamaFailureLog -Output $output
}
}
# -- Step D: Build llama-quantize (optional, best-effort) --
if ($BuildOk) {
Write-Host ""
Write-Host "--- cmake build (llama-quantize) ---" -ForegroundColor Cyan
$output = cmake --build $BuildDir --config Release --target llama-quantize -j $NumCpu 2>&1 | Out-String
$cmakeBuildQuantizeExit = $LASTEXITCODE
if ($cmakeBuildQuantizeExit -ne 0) {
substep "llama-quantize build failed (GGUF export may be unavailable)" "Yellow"
Write-LlamaFailureLog -Output $output
}
}
# -- Step E: Build the DiffusionGemma visual server (optional, best-effort) --
# An example target present on llama.cpp PR #24423; lets Unsloth serve
# DiffusionGemma GGUFs without DG_VISUAL_BIN. No-op when not configured.
if ($BuildOk) {
$null = cmake --build $BuildDir --config Release --target llama-diffusion-gemma-visual-server -j $NumCpu 2>&1 | Out-String
}
# Swap temp build dir into final location (only if we built in a temp dir)
if ($BuildOk -and $LlamaCppDir -ne $OriginalLlamaCppDir) {
Assert-StudioOwnedOrAbsent -Path $OriginalLlamaCppDir -Label "llama.cpp install"
if (Test-Path -LiteralPath $OriginalLlamaCppDir) { Remove-Item -LiteralPath $OriginalLlamaCppDir -Recurse -Force }
Move-Item -LiteralPath $LlamaCppDir -Destination $OriginalLlamaCppDir
$LlamaCppDir = $OriginalLlamaCppDir
$BuildDir = Join-Path $LlamaCppDir "build"
$LlamaServerBin = Join-Path $BuildDir "bin\Release\llama-server.exe"
Mark-StudioOwned -Path $LlamaCppDir
} elseif (-not $BuildOk -and $LlamaCppDir -ne $OriginalLlamaCppDir) {
# Build failed -- clean up temp dir, preserve existing install
if (Test-Path -LiteralPath $LlamaCppDir) { Remove-Item -LiteralPath $LlamaCppDir -Recurse -Force }
$LlamaCppDir = $OriginalLlamaCppDir
$BuildDir = Join-Path $LlamaCppDir "build"
$LlamaServerBin = Join-Path $BuildDir "bin\Release\llama-server.exe"
}
# Restore ErrorActionPreference
$ErrorActionPreference = $prevEAP
# Stop timer
$totalSw.Stop()
$totalMin = [math]::Floor($totalSw.Elapsed.TotalMinutes)
$totalSec = [math]::Round($totalSw.Elapsed.TotalSeconds % 60, 1)
# -- Summary --
if ($BuildOk -and (Test-Path -LiteralPath $LlamaServerBin)) {
step "llama.cpp" "built"
$QuantizeBin = Join-Path $BuildDir "bin\Release\llama-quantize.exe"
if (Test-Path -LiteralPath $QuantizeBin) {
step "llama-quantize" "built"
}
step "build time" "${totalMin}m ${totalSec}s" "DarkGray"
} else {
$altBin = Join-Path $BuildDir "bin\llama-server.exe"
if ($BuildOk -and (Test-Path -LiteralPath $altBin)) {
step "llama.cpp" "built"
step "build time" "${totalMin}m ${totalSec}s" "DarkGray"
} else {
step "llama.cpp" "build failed at: $FailedStep (${totalMin}m ${totalSec}s); continuing" "Yellow"
substep "To retry: delete $LlamaCppDir and re-run setup." "Yellow"
$script:LlamaCppDegraded = $true
}
}
}
$llamaCppItem = Get-Item -LiteralPath $LlamaCppDir -Force -ErrorAction SilentlyContinue
$llamaCppIsLink = $llamaCppItem -and ($llamaCppItem.Attributes -band [System.IO.FileAttributes]::ReparsePoint)
if (-not $llamaCppIsLink -and (
-not $StudioHomeIsCustom -or
(Test-Path -LiteralPath (Join-Path $LlamaCppDir $StudioOwnedMarker) -PathType Leaf) -or
(Test-StudioOwnedAdoptable $LlamaCppDir)
)) {
Remove-AgentInstructionFiles -Roots @($LlamaCppDir)
}
# ─────────────────────────────────────────────
# Footer
# ─────────────────────────────────────────────
$DoneLabel = if ($env:SKIP_STUDIO_BASE -eq "1") { "Unsloth Studio Setup Complete" } else { "Unsloth Studio Updated" }
if ($script:StudioVtOk -and -not $env:NO_COLOR) {
Write-Host (" {0}{1}{2}" -f (Get-StudioAnsi Dim), $Rule, (Get-StudioAnsi Reset))
if ($script:LlamaCppDegraded) {
Write-Host (" " + (Get-StudioAnsi Warn) + "$DoneLabel (limited: llama.cpp unavailable)" + (Get-StudioAnsi Reset))
} else {
Write-Host (" " + (Get-StudioAnsi Title) + $DoneLabel + (Get-StudioAnsi Reset))
}
Write-Host (" {0}{1}{2}" -f (Get-StudioAnsi Dim), $Rule, (Get-StudioAnsi Reset))
} else {
Write-Host " $Rule" -ForegroundColor DarkGray
if ($script:LlamaCppDegraded) {
Write-Host " $DoneLabel (limited: llama.cpp unavailable)" -ForegroundColor Yellow
} else {
Write-Host " $DoneLabel" -ForegroundColor Green
}
Write-Host " $Rule" -ForegroundColor DarkGray
}
step "launch" "unsloth studio -p 8888"
substep "(add -H 0.0.0.0 for LAN / cloud access; exposes the raw port only, not a public URL)"
substep "(add -H 0.0.0.0 --cloudflare for a public Cloudflare HTTPS link, or --secure to keep the raw port private; anyone with the API key can run code)"
Write-Host ""
# Match studio/setup.sh: exit non-zero for degraded llama.cpp when called
# from install.ps1 (SKIP_STUDIO_BASE=1) so the installer can detect the
# failure. Direct 'unsloth studio update' does not set SKIP_STUDIO_BASE,
# so it keeps degraded installs successful.
if ($script:LlamaCppDegraded -and $env:SKIP_STUDIO_BASE -eq "1") {
exit 1
}