Commit graph

5,228 commits

Author SHA1 Message Date
LeoBorcherding
643a797969 fix(install.ps1): recreate venv with Python 3.12 after ROCm switch
Venv was created with 3.13 before GPU detection ran; switching
$DetectedPython to 3.12 had no effect since $VenvPython still
pointed to the 3.13 interpreter inside the already-created venv.
2026-05-11 01:59:14 -05:00
LeoBorcherding
265a09a341 fix: guard reconcile call against None numeric_ids; add torchvision lower bounds 2026-05-11 00:27:10 -05:00
LeoBorcherding
42c5d984b7 chore: trim verbose comment blocks across all ROCm-related files 2026-05-11 00:20:28 -05:00
LeoBorcherding
f892b65a18 fix(tests): match windows AMD warning assertion to actual source string 2026-05-11 00:13:42 -05:00
LeoBorcherding
85841b5cad fix(rocm): guard c10d stub, fix TorchIndexFamily for 7.1, clean dead code + comments
- worker.py: wrap c10d stub injection in `if _c10d_key not in sys.modules` so
  Windows NVIDIA users with a real torch.distributed are never affected
- install.ps1: fix Get-TauriTorchIndexFamily receiving hardcoded "rocm7.2"
  even when ROCm 7.1 wheels are installed; now branches on $ROCmVersion
- main.py: remove dead `import ctypes as _ctypes` (ctypes is never called)
- hardware.py, install_python_stack.py, worker.py, install.ps1: shorten
  verbose multi-line comment blocks throughout
- tests: update 4 stale assertions that expected rocm7.2 to be absent/capped
2026-05-11 00:08:14 -05:00
LeoBorcherding
fe5546d837 fix(rocm/windows): pre-stub torch._C._distributed_c10d + raise amd-smi timeout
Two fixes for Windows ROCm regressions reported by electroglyph on #5301:

1. worker.py — torch.distributed stub now fires unconditionally on Windows
   The previous stub only injected sys.modules in the except branch, meaning
   it was silently skipped when `import torch.distributed` happened to succeed
   (the C backend is lazily resolved).  The crash then hit later when
   transformers/trl triggered the lazy load.  Fix: on win32 we pre-populate
   sys.modules['torch._C._distributed_c10d'] AND set the attribute on the
   torch._C extension module *before* attempting the import, covering both
   the early-ImportError and lazy-load failure modes.

2. amd.py — increase amd-smi timeout from 5 s to 30 s on Windows (10 s Linux)
   amd-smi on Windows must cold-init the ROCm runtime on first invocation;
   5 s was consistently too short, producing repeated 'Command timed out'
   warnings in the server log.  30 s gives enough headroom without blocking
   indefinitely on broken installs.

3. install.ps1 — widen Python 3.12 enforcement to ROCmGpuLabel (WMI-only path)
   Users whose HIP SDK is not on PATH were detected via WMI but not switched
   to Python 3.12 before the install started, causing a second pass.  Guard
   now fires on (HasROCm -or ROCmGpuLabel).
2026-05-10 23:06:13 -05:00
Leo Borcherding
4d401e7645
Merge branch 'unslothai:main' into fix/rocm-strix-halo-unified-memory 2026-05-10 13:56:26 -05:00
Tai An
b364080225
fix(gh_client): fail fast on 401/403 auth errors instead of retrying forever (#5325) (#5329)
* fix(gh_client): fail fast on 401/403 auth errors instead of retrying forever (#5325)

Fixes #5325. The Studio data-recipe GitHub Crawler swallows 401 Unauthorized
(and 403 Forbidden without rate-limit headers) into the generic
"network error" retry path, so a job with a stale or wrong-scoped GitHub
token spins indefinitely emitting "Retry." lines until the user cancels.

Changes:

- Add GitHubAuthError. Raised on 401, and on 403 unless the response carries
  a clear rate-limit signal (Retry-After header for secondary limits, or
  X-RateLimit-Remaining: 0 for primary limits).
- Track which token source resolved at construction time: explicit argument
  (recipe-level field), GH_TOKEN, or GITHUB_TOKEN. Surfaced in the error
  message so the user knows which credential to rotate.
- Insert the auth-failure check before the existing 403/429 rate-limit branch
  in both .graphql() and .rest() so auth failures bypass the sleep-and-retry
  loop and abort the recipe immediately.

Genuine rate limiting still retries via the existing path. requests.RequestException
handling is unchanged because GitHubAuthError does not inherit from it.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* style: apply black formatting per pre-commit.ci

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix GitHub auth failure handling

Preserve GitHub token source through the repo seed scraper and fail fast on non-rate-limit auth errors while keeping genuine rate-limit retries.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Wasim Yousef Said <wasimysdev@gmail.com>
2026-05-08 21:57:41 +04:00
pre-commit-ci[bot]
a288ff208c [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-08 09:58:27 +00:00
LeoBorcherding
bfac7c1015 fix: inject torch.distributed stub when C backend missing in ROCm Windows wheel #5301 2026-05-08 04:57:27 -05:00
pre-commit-ci[bot]
18690c6b67 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-08 09:47:58 +00:00
LeoBorcherding
739e5d555d fix: stub all missing torch.distributed attrs for ROCm Windows wheel #5301 2026-05-08 04:47:42 -05:00
pre-commit-ci[bot]
a7047229ff [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-08 09:22:21 +00:00
LeoBorcherding
1707169680 fix: suppress remaining console popups on Windows, patch torch.distributed.is_initialized for ROCm #5301 2026-05-08 04:22:04 -05:00
Roland Tannous
c57a97958a
Studio: stop truncating long log lines as suspected base64 (#5335)
* Studio: stop truncating long log lines as suspected base64

filter_sensitive_data carried a heuristic from the original Studio
import that truncated any string >100 chars containing ',' or '/'
to value[:20] + '...'. The block was dormant until #5246 wired
filter_sensitive_data into the structlog processor chain to redact
native-path leases. Once active, the heuristic ate normal log lines
- llama_cpp_backend's GGUF size summary, mmproj selection, the full
llama-server command line, and any traceback containing a path -
all rendered as a 20-char prefix, defeating debugging of llama-server
exceptions and GPU selection.

Drop the base64 truncation. No call site in the codebase logs raw
base64; if one ever does, it should truncate at the source rather
than in a global filter. Native-path lease redaction added by #5246
is preserved.

* Studio: regression test for filter_sensitive_data truncation

Pins two properties in studio/backend/loggers/handlers.py:

1. Long log messages with ',' or '/' (the GGUF size summary, mmproj
   selection, full llama-server command, exception tracebacks) flow
   through filter_sensitive_data unchanged. Exercises the exact call
   sites that regressed when #5246 wired the processor in.

2. Native-path lease redaction still fires for both the inline
   native_path_lease=... regex form and the nativePathLease dict-key
   form, so a future cleanup of the truncation logic can't quietly
   strip #5246's redaction along with it.
2026-05-08 13:07:18 +04:00
pre-commit-ci[bot]
ba6b279e61 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-08 08:51:06 +00:00
LeoBorcherding
2de2c29d3a fix: hide amd-smi console popups on Windows, guard torch.distributed.is_initialized for ROCm #5301 2026-05-08 03:50:46 -05:00
LeoBorcherding
bafb3f54df fix: also check uv-managed Python 3.12 for AMD ROCm #5301 2026-05-08 03:20:06 -05:00
LeoBorcherding
5deb230cfb fix: prefer Python 3.12 for AMD ROCm users when 3.13 is also installed
After GPU detection, if ROCm HIP SDK is found and the selected Python
is not 3.12, run a second pass to locate a 3.12 install via py.exe and
PATH (catches uv-managed installs). Switch $DetectedPython to 3.12 so
the venv is created with a compatible interpreter for the cp312-only AMD
Windows torch wheels.

NVIDIA and Intel GPU paths are unaffected -- the re-detection block only
runs when $HasROCm is true.

Fixes: #5301
2026-05-08 03:05:12 -05:00
Leo Borcherding
550317d738
Merge branch 'unslothai:main' into fix/rocm-strix-halo-unified-memory 2026-05-07 08:01:40 -07:00
Etherll
d1f9ab659f
fix: harden Studio IME composer sends (#5327)
* fix: harden Studio IME composer sends

* fix: address IME composer review feedback
2026-05-07 18:29:10 +04:00
Lee Jackson
b65a7450ca
Studio: Dark theme refactor, right sidebar redesign, and chat UI polish (#5150)
* Dark theme refactor, right sidebar redesign, and chat UI polish

- Dark theme refactor
- Redesign right sidebar
- Further left sidebar adjustments
- Wider chat and content area; layout tweaks for chat content
- Rounded corners across elements for consistency
- Show chat message menu icons on menu-area hover, not only on message hover
- Assistant message menu icons now always visible; user messages keep on-hover
- Redesigned copy icon used consistently across chat blocks and messages
- Redesigned trash icon, applied consistently
- Unified icon sizing and style with the sidebar
- Adjusted icon colors across chat
- Fix on-hover background design for chat icons
- Fix tooltip from 'more' button staying visible after clicking elsewhere
- Adjust position and design of generation speed info text below messages
- Adjust design of token speed info popup
- Adjust sidebar scrollbar to cover recent chats only

* Recents sidebar rename, UI/theme refactor, layout and chat polish

UI & Theme:
- Dark theme refactor
- Consistent rounded corners across elements
- CSS polish and cleanup
- Remove unused logo image assets

Recents sidebar:
- Add 'more' button for options menu
- Support renaming conversations and training runs
- Confirmation dialog before deleting chats
- Add optional display_name column to training_runs (idempotent ALTER TABLE) so renaming doesn't lose model_name/dataset_name from the run config
- New PATCH /api/train/runs/{run_id} endpoint accepts { display_name: string | null }; empty/whitespace clears the override
- Sidebar shows display_name ?? model_name and exposes Rename in the row's More menu, mirroring the chat rename flow
- Cache last list response in localStorage and hydrate from it on mount, so recents paint instantly on F5 / route revisit; cached items are shape-validated and dropped if malformed
- Optimistic updates on rename and delete (apply locally + cache before background refresh)
- Visible toast on rename/delete failure instead of swallowed errors

Layout:
- Redesigned right sidebar
- Further left sidebar adjustments
- Updated chat content layout; chat and content area slightly widened
- Sidebar scrollbar covers recent chats only

Icons:
- Redesigned copy icon, unified across chat blocks and messages
- Redesigned trash icon to match
- Consistent icon sizing and style across chat and sidebar
- Adjusted icon colors across chat
- Fix icon on-hover background design

Chat messages:
- Menu icons now appear on hover over the menu area, not just the message
- Assistant message menu icons always visible; user messages keep on-hover (next/previous response stays visible for edited prompts)
- Repositioned and restyled generation speed info text below messages
- Restyled token generation speed popup

Tooltips:
- Removed tooltip on hover for previous/next assistant response icons
- Unified tooltip design across sidebars and chat
- Removed tooltip animations (also fixes related lag)

Model & Chat Template config:
- Merged Chat Template config into Model Configuration section
- Added revert-to-original for chat template
- Fix Chat Template config disappearing on page refresh until model reload

Performance & scroll:
- Removed chatbox movement animations across pages/navigation (fixes related UI lag)
- Fix scroll flicker at end of streaming when a code block is the final element
- Additional chat scroll improvements

Bug fixes:
- Fix 'more' button tooltip remaining visible after clicking elsewhere

* Remove sidebar localStorage cache and optimistic updates

Drops the localStorage hydration and optimistic rename/delete logic from the recents sidebar; reverts to fetching fresh on mount.

* Fix missing cn import in shared-composer (regression from merge)

* chore(sidebar): import sidebar deps from feature indexes

Re-export deleteChatItem / renameChatItem / useChatSidebarItems / SidebarItem / useChatSearchStore / ChatSearchDialog from @/features/chat, and removeTrainingUnloadGuard from @/features/training. Switch app-sidebar.tsx to consume them via the public feature indexes instead of deep paths, clearing the no-restricted-imports eslint errors. No behavior or UX change.

* fix(studio/frontend): reload training Recents sidebar after F5 refresh

The Recents sidebar showed empty after a hard refresh. The hook's inFlightRef dedup guard collided with React StrictMode's double-mount in dev: the second mount's fetch returned silently with no error, no retry, and no toast — leaving the sidebar empty until navigation.

Replace skip-if-busy dedup with abort-previous via a hook-level AbortController. This also fixes a latent race where a slow poll could resurrect a just-deleted row by clobbering the optimistic update.

Changes (all in use-training-history-sidebar.ts):
- fetchRuns aborts any in-flight request before starting a new one; post-await signal.aborted check drops stale responses.
- Optimistic helpers (applyRunUpdate, removeRun) abort in-flight fetches so they don't depend on caller discipline to invalidate stale data.
- Initial load gets bounded retry-with-backoff (500ms / 1.5s / 3.5s) and surfaces a sonner toast with a Retry action on final failure.
- Failure toast auto-dismisses on any successful load (initial retry, Retry click, or polling recovery).
- Polling pauses while the tab is hidden and catches up on visible, avoiding wasted requests during long training runs.
- Both effects own their teardown explicitly (abort + clear timer).

* Apply unified tooltip design and behavior across remaining pages for consistency

* UI polish: spacing, tooltip on source icons, letter spacing, smaller icons, consistent edit icon

- Adjust tiny spacing between elements around the UI for subtle polish
- Redesign tooltip on source icons for web search / tool use, consistent with the new design
- Adjust chat text letter spacing
- Smaller icon sizes
- Replace 'edit message' icon in chat with the new Rename icon used in Recents for consistency

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Adjust CSS for right sidebar

* Fix scrollbar UI compatibility across browsers

* fix: preserve chat preset settings on model load

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(studio): remove duplicate chat template status field

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* chore: remove creative preset assumption

* fix(studio): align speculative decoding default

* fix(studio/chat): snap numeric param inputs to step grid

- Type a value in any param input (Temperature, Top K, Max Tokens, etc.)
  now clamps to [min, max] and snaps to the slider's step grid, killing
  off-grid values like 1.051234 and FP residue from slider drags.
- Branch picker chevrons share the action bar's 32px height + 10px radius
  via a new .aui-branch-chevron-btn utility; hover area aligns visually
  while staying narrower than the sibling icon buttons.

* fix(studio/chat): keep training-run polls converging and drop dead preset code

- Keep training-run polls converging when responses outrun the 5s interval
  (don't unconditionally abort prior in-flight; skip if one is still pending,
  mutation race still guarded).
- Drop dead Creative/Precise preset code paths (remove 'builtin-fixed' source
  variant + unreachable branches).

* fix(studio): training-run cards show custom name + model + dataset

- Training-run cards now display custom display_name + model + dataset,
  with cross-view sync on rename/delete.
- Enhance clarity of borders and colors in dark theme on export etc.

* fix(studio): match active state green to unsloth brand color

* fix(studio): preserve can_resume on training rename

* fix(studio): keep GGUF chat template override distinct

* fix(studio): treat audio input models as multimodal

* fix(studio): cancel numeric draft on Escape

* fix(studio): use default speculative mode on toggle

* fix(studio): detect GGUF audio VLM input models

* fix(studio): address final PR review findings

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(studio): refresh sidebar/history when a new training run starts so it appears without a manual reload

* fix: API and svg

* fix(studio/sidebar): align run rename dirty check with displayed baseline

* fix(studio/sidebar): use leading-tight on account block to prevent descender clipping with truncate

---------

Co-authored-by: sneakr <hauzin@hotmail.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Roland Tannous <115670425+rolandtannous@users.noreply.github.com>
Co-authored-by: shine1i <wasimysdev@gmail.com>
2026-05-07 14:33:31 +04:00
Daniel Han
7af8cac014
tests/studio/install: parallel UNSLOTH_STUDIO_HOME smoke test (#5306)
* tests/studio/install: parallel UNSLOTH_STUDIO_HOME smoke test

Adds tests/studio/install/smoke_test_parallel_studio_home.py to lock in
the install-time and runtime isolation guarantees added by #5190.

The runner spawns N concurrent install.sh --local --no-torch jobs, each
with its own UNSLOTH_STUDIO_HOME and a redirected HOME, then launches N
backends on dynamically allocated ports and cross-checks every install
against its running process. Asserts:

  install-time
    - all N installs exit 0
    - per-install bin / share / llama.cpp / unsloth_studio venv tree
    - shim symlink resolves into its own venv, no cross-resolution
    - share/studio_install_id is unique across the N installs
    - share/studio.conf exports UNSLOTH_EXE / UNSLOTH_STUDIO_HOME /
      UNSLOTH_LLAMA_CPP_PATH all pointing inside the install
    - share/launch-studio.sh has @@DATA_DIR@@ substituted to its own
      share/ at install time
    - the redirected HOME stays clean: no rc-file append, no
      .desktop file, no Studio.app stub, no shared marker

  runtime
    - /api/health returns 200 with status healthy and chat_only true
    - /api/health.studio_root_id matches share/studio_install_id
      (runtime resolver agrees with install-time write)
    - studio_root_id values are pairwise distinct
    - GET / and GET /api/chat return 200 on each backend
    - /proc/PID/exe is the install's own venv python

Standalone smoke runner, not pytest collected. Default --n 4 finishes
in about 60 seconds on a warm uv cache; artifacts are removed on PASS
unless --keep is passed and kept on FAIL or ERROR for inspection.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* tests/studio/install: portability + log fd cleanup in parallel smoke

Two cleanups on the parallel UNSLOTH_STUDIO_HOME smoke runner:

- Skip the /proc/PID/exe runtime cross-resolution check on platforms
  without /proc (macOS, BSD, Windows). install.sh supports macOS, so
  the smoke should not hard-error there. The install-time symlink,
  studio.conf and launch-studio.sh assertions already pin the venv
  python target statically; the proc check stays as a Linux-only
  redundant cross-resolution catch and now returns None cleanly on
  other platforms instead of raising.

- Wrap the per-backend log file in a with-statement so its parent fd
  is released deterministically at function return. The child still
  holds its own dup'd fd via Popen, so logging continues unchanged.
  The prior code relied on local-scope GC and was fine in CPython,
  but the with form makes the intent explicit.

Smoke still passes locally: 4 parallel installs in 42s, 4 backends
healthy in 5s, all install + runtime invariants hold.

* tests/studio/install: pin UNSLOTH_STUDIO_HOME on backend launch

The launch step copied os.environ unchanged except for HOME. If the
parent shell already exports UNSLOTH_STUDIO_HOME or STUDIO_HOME (for
example, when the developer is sourcing studio.conf from an existing
install), every backend inherits it and the Studio resolver prioritises
those env vars over the per-label sys.prefix inference. The runtime
invariant block then reports the caller's install_id on every port
instead of the per-label one, and the test fails spuriously rather
than testing the right roots.

Pin UNSLOTH_STUDIO_HOME to the per-label studio_home and pop the
STUDIO_HOME alias for each launch, mirroring what _run_one_install
already does for the install step.

Verified by running the smoke with UNSLOTH_STUDIO_HOME=/nonexistent
and STUDIO_HOME=/also-bogus exported in the parent env: PASS, all four
backends report their own install_id rather than the parent value.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-07 02:26:58 -07:00
Lee Jackson
4ab096970d
Studio: API settings overflow with long Colab URLs (#5286)
* fix: API settings overflow with long Colab URLs

* fix: gentle wrapping for API usage snippets

---------

Co-authored-by: Roland Tannous <115670425+rolandtannous@users.noreply.github.com>
2026-05-07 13:17:23 +04:00
Daniel Han
948ce43584
fix: 3 patch_* helpers — fast_lora import, sft_trainer Union, openenv OSError (#5319)
* fix: import fast_lora_forward inside patch_fast_lora

patch_fast_lora has referenced an unbound `fast_lora_forward` since
ddf118a8f (2024-11-21). The function is defined at
unsloth/kernels/fast_lora.py:652 and re-exported through
unsloth/kernels/__init__.py:45, but it was never imported into
unsloth/models/_utils.py, so calling patch_fast_lora() raises
NameError: name 'fast_lora_forward' is not defined.

The bug went unnoticed because no production code path calls
patch_fast_lora() unconditionally. Surfaced by a new CPU-CI check that
invokes every zero-arg patch_* helper across unsloth + unsloth_zoo
(consolidated-tests-ci.yml on PR #5312).

Importing inside the function (rather than at module top) keeps the
import surface narrow and avoids a circular-import risk if
unsloth.kernels.fast_lora ever needs to import from
unsloth.models._utils.

* fix: inject typing imports into patch_sft_trainer_tokenizer's exec namespace

patch_sft_trainer_tokenizer rewrites the source of TRL's SFTTrainer
methods (_prepare_non_packed_dataloader, _prepare_dataset) and re-execs
them. With TRL 1.x, those methods carry `Union[...]` type hints in
their signatures. The current rewrite only injects identifiers found by
`dir(trl.trainer.sft_trainer)` into the exec namespace, which does not
include `Union`, so exec(function, ...) raises NameError: name 'Union'
is not defined.

Fix: import Union, Optional, List, Any, Callable, Tuple, Dict, Iterator
inside the function. exec receives `locals()` as its globals dict, so
those names are visible to the executed source body. Same pattern as
unsloth/models/_utils.py:patch_linear_scaling, which already injects
`from typing import Union, Optional, List, Any, Callable, Tuple` into
its own exec_code.

Surfaced by the consolidated CPU-CI runtime patch_* check on PR #5312
in the matrix cell `transformers>=5,<6 + trl>=1,<2`.

* fix: guard openenv_vllm_reload_weights against OSError from inspect.getsource

TRL 0.29.1 and the 1.x line ship some openenv helpers as compiled
bytecode without accessible source on disk. inspect.getsource(patch_target)
raises OSError("could not get source code") in that case, which surfaces
as a hard failure in patch_trl_openenv() and aborts the rest of the
RL_ADDITIONAL_FUNCTIONS["openenv"] iteration.

Wrap the getsource call in a try/except OSError and log a warning
instead. The wake_up(tags=...) rewrite is the only thing skipped; the
core weight-reload patch path stays functional.

Surfaced by the consolidated CPU-CI runtime patch_* check on PR #5312
in matrix cells running TRL 0.29.1 (latest <1.0.0) and TRL 1.3.0
(latest 1.x). The pyproject pin (TRL 0.18.2-0.24.0) still gets source
for this function so the original code path runs unchanged there.
2026-05-07 00:12:09 -07:00
pre-commit-ci[bot]
1680dacfee [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-07 02:45:15 +00:00
LeoBorcherding
6fe91e7749 feat: enable ROCm 7.2 torch install + warn on gfx1151 with ROCm < 7.2
Chigoma333 (AMD Radeon 8060S / gfx1151, Strix Halo) confirmed that ROCm
7.1 segfaults when tensors are moved to GPU, but ROCm 7.2 + torch
2.11.0+rocm7.2 works fully including training.

Changes:
- Uncomment (7,2): "rocm7.2" in _ROCM_TORCH_INDEX (was blocked by <2.11.0)
- Add _ROCM_TORCH_PKG_SPECS dict with per-tag version bounds:
  rocm7.2 → torch>=2.11.0,<2.12.0; all older tags → <2.11.0
- Add _detect_amd_gfx_codes() helper that parses rocminfo output
- Warn on gfx1151/gfx1150 (Strix Halo) when ROCm < 7.2 is installed,
  pointing users at the known segfault and recommending upgrade
- install.sh get_torch_index_url(): enable rocm7.2 case (previously capped
  to rocm7.1), cap unknown future tags to rocm7.2
- install.sh: override TORCH_CONSTRAINT to >=2.11.0,<2.12.0 when rocm7.2
  index is selected, so pip can actually resolve torch 2.11.0
2026-05-06 21:44:50 -05:00
LeoBorcherding
301d6c0aa7 fix: add rocm_sdk namespace tarball to Windows ROCm wheel installs
torch/_rocm_init.py calls `import rocm_sdk` at startup, which requires
the rocm namespace tarball (rocm-*.tar.gz) in addition to the SDK wheel
packages. This tarball was missing from both install.ps1 and setup.ps1,
causing ModuleNotFoundError on first torch import.

- Add rocm-0.1.dev0.tar.gz to ROCm 7.1.1 install (provides rocm_sdk namespace)
- Add rocm-7.2.1.tar.gz + rocm_sdk_devel to ROCm 7.2.1 install
- Install tarball in a dedicated step before main SDK/torch wheels
- Switch to @array splatting in install.ps1 scriptblock for dynamic wheel count
- Remove --no-cache-dir from Python-side ROCm wheel install (prevents ~2GB redownload)
2026-05-06 21:31:44 -05:00
LeoBorcherding
efcaccbbcf fix: prevent torchao overrides step from overwriting AMD ROCm torch
torchao==0.14.0 in overrides.txt declares torch as a dependency. Without
--no-deps, uv resolves torch from PyPI and installs 2.11.0+cpu on top of
the AMD ROCm wheels (2.9.0+rocmsdk20251116). This was the root cause of
'Hardware detected: CPU' -- the AMD wheels were installed but then
immediately overwritten by the overrides step.

When _rocm_windows_torch_installed is True, add --no-deps to the overrides
pip_install call so torchao is installed without pulling in CPU torch.
2026-05-06 20:54:08 -05:00
pre-commit-ci[bot]
cc77737b78 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-07 01:46:35 +00:00
LeoBorcherding
0facfbc4bc fix: remove hardcoded non-standard ROCm paths from DLL directory scan
Only use HIP_PATH/ROCM_PATH (set by AMD installer) and the standard
C:\Program Files\AMD\ROCm\<version>\bin location. Custom drive paths
like F:\ROCm are user-specific and should not be hardcoded.
2026-05-06 20:46:16 -05:00
pre-commit-ci[bot]
05f5cda295 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-07 01:45:11 +00:00
LeoBorcherding
f09a4240a1 fix: register ROCm DLL directory before torch import on Windows
Python 3.8+ ignores PATH for extension DLL loading on Windows; amdhip64.dll
and other HIP runtime DLLs must be registered via os.add_dll_directory().
Without this, torch.cuda.is_available() always returns False on AMD ROCm
Windows even when HIP_PATH is correctly set in system environment variables.

Reads HIP_PATH / ROCM_PATH env vars first, then falls back to scanning
common ROCm install roots (C:\Program Files\AMD\ROCm, F:\ROCm, C:\ROCm).
2026-05-06 20:43:18 -05:00
LeoBorcherding
7fbdce171c fix: pass AMD torch install status via env var to suppress false warning
setup.ps1 now sets UNSLOTH_ROCM_TORCH_INSTALLED=1 after a successful AMD
wheel install. install_python_stack.py reads this at the top of
_ensure_rocm_torch() to skip both the subprocess probe and the warning --
no re-import of torch needed, and the warning message now correctly says
'could not be auto-installed' rather than 'must be installed manually'.
2026-05-06 20:37:44 -05:00
LeoBorcherding
4b2f7fb8e2 fix: hoist global declaration to top of _ensure_rocm_torch
Python requires the global statement to appear before any assignment
to the variable within a function. Moving it to the function top fixes
the SyntaxError on line 354.
2026-05-06 20:31:55 -05:00
LeoBorcherding
9439657ee8 fix: use install-state flag instead of subprocess probe for AMD Windows warning
Replace the subprocess torch probe in the post-install warning block with a
module-level _rocm_windows_torch_installed flag set by _ensure_rocm_torch().
Subprocess re-import of torch is unnecessary and fragile -- the install
function already knows whether it succeeded.
2026-05-06 20:29:06 -05:00
LeoBorcherding
77f7adef62 perf: drop --no-cache-dir from AMD ROCm torch wheel installs
uv caches downloaded wheels by default; passing --no-cache-dir forced a
full redownload of the ~2 GB torch wheel on every install run. CUDA installs
never had this flag -- AMD was the only path affected.
2026-05-06 20:16:12 -05:00
pre-commit-ci[bot]
7d3de8b317 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-07 01:14:35 +00:00
LeoBorcherding
f036ee0022 fix: detect AMD SDK ROCm torch via __version__ when torch.version.hip is unset
AMD's repo.radeon.com wheels (e.g. 2.9.0+rocmsdk20251116) do not set
torch.version.hip, leaving it None. All three probes that relied solely on
torch.version.hip now also check for 'rocm' in torch.__version__.lower():

- hardware.py detect_hardware(): IS_ROCM was never set, causing the studio
  to report 'Hardware detected: CPU' even after AMD wheels were installed
  and HIP DLLs were on PATH.
- install_python_stack.py _ensure_rocm_torch(): skip-if-already-installed
  probe would always reinstall on subsequent runs.
- install_python_stack.py Windows AMD warning: suppression check always
  failed, so the 'must be installed manually' note kept appearing after
  a successful AMD wheel install.
2026-05-06 20:14:14 -05:00
LeoBorcherding
67d8b7481a feat: add rocm step display in setup.ps1; fix warning and progress counter
- Add 'rocm' step after 'cuda' in setup.ps1 showing ROCm version or HIP SDK missing
- Move ROCm version detection up to GPU detection block so it's available early
- Suppress 'must be installed manually' warning when torch.version.hip is set
- Fix _TOTAL counter to include ROCm steps on Windows (fixes 10/9 display)
2026-05-06 19:51:15 -05:00
pre-commit-ci[bot]
b9c6882df4 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-07 00:27:27 +00:00
LeoBorcherding
6f792137bb fix: suppress manual-install warning when ROCm torch already present; fix progress counter
- Gate the 'must be installed manually' warning on torch.version.hip being empty
  so it doesn't fire when our ROCm torch install succeeded
- Update _TOTAL counter to include the 3 ROCm steps on Windows now that
  _ensure_rocm_torch() is called there (fixes 10/9 display)
2026-05-06 19:24:46 -05:00
LeoBorcherding
b550948b39 fix: setup.ps1 and install_python_stack.py now install ROCm torch on Windows
setup.ps1 was always setting CuTag='cpu' for non-NVIDIA hosts and installing
cpu-only PyTorch, overwriting the ROCm torch installed by install.ps1.
Adds the same AMD wheel selection logic (ROCm version detection, Python 3.12
check, 5-wheel install with --no-deps) to setup.ps1's torch install block.

install_python_stack.py: remove IS_WINDOWS guard from _ensure_rocm_torch()
call site so the Windows path in _ensure_rocm_torch() is reachable during
'unsloth studio update' as well.
2026-05-06 19:07:17 -05:00
LeoBorcherding
79670ab4a9 fix: use --no-deps for AMD Windows torch wheel install
uv's resolver looks up rocm[libraries]==0.1.dev0 on PyPI during
dependency resolution before downloading any wheels, and fails because
the package doesn't exist on PyPI. --no-deps skips resolution entirely
and installs all 5 AMD wheels directly. The GPU runtime dependency is
satisfied by the HIP SDK, not a Python package.
2026-05-06 18:48:57 -05:00
LeoBorcherding
a74cb8b94e fix: expand ROCm wheel array to scalars for Invoke-InstallCommand
@array splatting inside a scriptblock only works when the native command
is prefixed with '&'. Invoke-InstallCommand uses '& $Command' to run the
block, so @ROCmAllWheelUrls was not being expanded. Extract to scalar
variables $rw0-$rw4 which are captured correctly by the closure.
2026-05-06 18:12:42 -05:00
LeoBorcherding
e53b543ef0 fix: install rocm_sdk_core and rocm_sdk_libraries_custom alongside torch
The AMD Windows torch wheels declare rocm[libraries]==<ver> as a hard
dependency. Without installing rocm_sdk_core and rocm_sdk_libraries_custom
from the same AMD release folder, uv cannot resolve the dependency and
fails with 'No solution found'. Include all 5 wheels in one install call.
2026-05-06 18:08:01 -05:00
LeoBorcherding
7e2943597d feat: add ROCm 7.1.1 Windows wheel mapping
AMD uses a different version string for 7.1.1 wheels:
2.9.0+rocmsdk20251116 (date-tagged) instead of +rocm7.1.1.
Adds the 7.1.1 release folder to both install.ps1 and
install_python_stack.py so users with ROCm 7.1 get ROCm
torch instead of falling back to CPU.
2026-05-06 17:55:58 -05:00
pre-commit-ci[bot]
ec40e9f207 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-06 21:34:36 +00:00
LeoBorcherding
8299842271 fix: also install torchvision and torchaudio from AMD Windows repo
AMD publishes matching torchvision-0.24.1+rocm7.2.1 and
torchaudio-2.9.1+rocm7.2.1 cp312 wheels at the same repo.radeon.com
release folder. Install all three in both install.ps1 and
install_python_stack.py Windows ROCm path.
2026-05-06 16:33:29 -05:00
pre-commit-ci[bot]
14f655917c [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-06 21:32:24 +00:00