AMD: enable ROCm torch on gfx906 (MI50 / Radeon VII) on Linux (#7354)
* Add community-maintained legacy support path for gfx906 (MI50 / Radeon VII) rocm6.4+/7.x torch wheels bundle ROCm libraries whose Tensile kernels dropped gfx906 (rocBLAS 'TensileLibrary.dat ... not read for gfx906', ROCm/TheRock#1844), so on MI50/Vega 20 hosts with newer ROCm the installer picked wheels that fail at the first BLAS call. The rocm6.3 index is the last one whose wheels run on gfx906 (torch 2.7.0 verified on MI50 32GB, up to 2.9 in community use). Dynamo/Inductor codegen is also broken on this arch, crashing compiled graphs that train fine in eager mode. - install.sh: when the runtime GPU is gfx906 and the picked index is newer than rocm6.3, reroute torch to the rocm6.3 index and reset the constraint trio to the default <2.11 window (a rocm7.2 pick raises the floor to 2.11, which rocm6.3 cannot satisfy), with a legacy-path warning. - install_python_stack.py: mirror the reroute in _ensure_rocm_torch using the _default pkg specs, including repairing an existing +rocm7.x torch and leaving a working rocm6.3 install alone. - device_type.py: default TORCHDYNAMO_DISABLE / TORCH_COMPILE_DISABLE / UNSLOTH_COMPILE_DISABLE on gfx906 (setdefault, user override wins). Windows allowlists are untouched: repo.amd.com publishes no gfx906 wheel family (verified in the RDNA2 enablement PR). 16-bit LoRA and full finetuning work out of the box; 4-bit QLoRA needs a source-built bitsandbytes for gfx906. Based on the verified MI50 32GB setup in namnguyen0503/mi50-gfx906-unsloth-bnb4bit-lab. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * gfx906: second Codex pass (bnb skip under pin, override beats Strix) - Compute the gfx906 runtime-target flag independently of any torch-index pin or Strix override, so the bitsandbytes skip still applies when a user pins the ROCm index and sets UNSLOTH_ROCM_GFX_ARCH=gfx906 (the pin suppresses the torch reroute, not the bnb skip). Probe only when no pin is set (an explicit pin means don't second-guess it, matching the Strix path's asserted no-probe invariant); an explicit gfx906 override needs no probe. - Let UNSLOTH_ROCM_GFX_ARCH=gfx906 suppress the Strix reroute (both install.sh and install_python_stack.py) so a mixed Strix + MI50 host routes to rocm6.3 instead of the gfx1151 wheels probe order would pick. - Fix test_hardcoded_torch_constraint: the default <2.11 window literal now legitimately appears on two TORCH_CONSTRAINT= assignments (default + the gfx906 reroute reset after the rocm7.2 floor bump); assert it only ever appears on assignment lines, never on a pip install line (its real intent). New tests: bnb skipped under an explicit pin, gfx906 override wins over Strix, install.sh suppresses Strix on the override. rocm_support + selection + cross-platform parity: 667 passed; structural constraint 9/9. * gfx906: collapse single-line asserts to match pre-commit formatting * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * gfx906: keep bnb skip + rocm6.3 routing correct under pins and suffixed overrides Address the four Codex P2 findings on #7354: - bnb skip under a pinned index (install.sh + install_python_stack.py): a real gfx906 host that pins UNSLOTH_TORCH_INDEX_URL to rocm6.3 without also setting UNSLOTH_ROCM_GFX_ARCH no longer reinstalls the generic bitsandbytes wheel over a source-built gfx906 bnb. A pin now suppresses only the torch reroute, not the gfx906 detection used for the bnb skip (Python drops the pin gate on _runtime_is_gfx906; bash _is_gfx906_bnb_skip probes via _probe_amd_gfx_arch when the index is pinned). - clear the Radeon marketing-name flag for every gfx906 target, not only when the >=6.4 reroute fires, so a Radeon VII already on rocm6.3 does not divert to the repo.radeon.com branch (whose wheels lack gfx906 kernels). - normalize a copied HIP gcnArchName (gfx906:sramecc-:xnack- -> gfx906) before the exact comparisons in install.sh and install_python_stack.py, mirroring device_type.py. Tests: relax the three Strix-pin tests (the gfx probe may now run for the bnb flag but must not reroute the pinned index) and add coverage for the pinned bnb skip, the suffixed override, and the bash Radeon-clear / pinned-probe paths. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * gfx906: log skipped vLLM aimv2 fix + robust source-scan test bounds Follow-up review polish: - import_fixes: log at info level when the vLLM aimv2 fix is skipped because the dist metadata is unreadable, so the skip is diagnosable instead of silent. - test_rocm_support: bound the gfx906 install.sh source-scan on the ';;' that closes its case arm via a shared _gfx906_reroute_block helper, replacing the brittle fixed-length (3200/3800) slices that shift when the block grows. * gfx906: trim whitespace on UNSLOTH_ROCM_GFX_ARCH in install.sh (py parity) The bash gfx906 comparisons lowercased and stripped the gfx906:… feature suffix but not surrounding whitespace, while the Python paths do .strip(). A stray newline (e.g. export UNSLOTH_ROCM_GFX_ARCH=$(cmd)) would make bash miss gfx906 while Python catches it. Trim with `tr -d '[:space:]'` at both comparison sites so the reroute target and bnb-skip agree across bash/Python. * gfx906: remove generic bitsandbytes pulled in transitively after the skip --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: danielhanchen <unslothai@gmail.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com>
This commit is contained in:
parent
74295d93d8
commit
f03e669442
6 changed files with 659 additions and 26 deletions
128
install.sh
128
install.sh
|
|
@ -257,6 +257,51 @@ run_install_cmd_retry() {
|
|||
done
|
||||
}
|
||||
|
||||
# True when the runtime target is gfx906 (MI50/Radeon VII): the prebuilt AMD
|
||||
# bitsandbytes wheel carries no gfx906 kernels, and force-reinstalling it would
|
||||
# clobber a user's source-built bnb (the only 4-bit path on this arch) on every
|
||||
# `studio update`. So skip the auto-install and leave whatever bnb is present.
|
||||
# _gfx906_target is set during torch-index resolution; also honor an explicit
|
||||
# UNSLOTH_ROCM_GFX_ARCH so a pinned-index install still skips. The override is
|
||||
# normalized (gfx906:sramecc-:xnack- -> gfx906) so a copied HIP gcnArchName counts.
|
||||
_is_gfx906_bnb_skip() {
|
||||
[ "${_gfx906_target:-false}" = true ] && return 0
|
||||
_bnb_gfx_env=$(printf '%s' "${UNSLOTH_ROCM_GFX_ARCH:-}" | tr '[:upper:]' '[:lower:]' | tr -d '[:space:]')
|
||||
_bnb_gfx_env=${_bnb_gfx_env%%:*}
|
||||
[ "$_bnb_gfx_env" = "gfx906" ] && return 0
|
||||
# A pinned index (UNSLOTH_TORCH_INDEX_URL/_FAMILY) skips the reroute block that
|
||||
# sets _gfx906_target, so a real gfx906 host with a pinned rocm6.3 index and no
|
||||
# UNSLOTH_ROCM_GFX_ARCH would otherwise clobber a source-built bnb. Probe here
|
||||
# in that gap; skip only when gfx906 is the SOLE distinct arch (mixed hosts
|
||||
# opt in via the env var, mirroring the reroute block's de-dup rule).
|
||||
if [ -z "$_bnb_gfx_env" ] && [ "${_torch_index_pinned:-false}" = true ]; then
|
||||
_bnb_gfx_probe=$(_probe_amd_gfx_arch | awk 'NF && !seen[$0]++')
|
||||
[ "$_bnb_gfx_probe" = "gfx906" ] && return 0
|
||||
fi
|
||||
return 1
|
||||
}
|
||||
|
||||
# `pip install unsloth` resolves its unconditional bitsandbytes dep to a generic
|
||||
# CUDA wheel (no gfx906 kernels) once we skip the prebuilt one. Snapshot bnb before
|
||||
# the unsloth install, then drop a freshly pulled wheel afterwards while leaving a
|
||||
# pre-existing source build in place.
|
||||
_gfx906_bnb_installed() {
|
||||
"$_VENV_PY" -c "import importlib.util as u, sys; sys.exit(0 if u.find_spec('bitsandbytes') else 1)" >/dev/null 2>&1
|
||||
}
|
||||
_gfx906_bnb_snapshot() {
|
||||
_gfx906_bnb_absent_before=false
|
||||
_is_gfx906_bnb_skip || return 0
|
||||
_gfx906_bnb_installed || _gfx906_bnb_absent_before=true
|
||||
}
|
||||
_gfx906_bnb_prune() {
|
||||
_is_gfx906_bnb_skip || return 0
|
||||
[ "${_gfx906_bnb_absent_before:-false}" = true ] || return 0
|
||||
_gfx906_bnb_installed || return 0
|
||||
substep "gfx906: removing generic bitsandbytes pulled in as a dependency (no gfx906 kernels; build from source for 4-bit QLoRA)" "$C_WARN"
|
||||
uv pip uninstall --python "$_VENV_PY" bitsandbytes >/dev/null 2>&1 \
|
||||
|| "$_VENV_PY" -m pip uninstall -y bitsandbytes >/dev/null 2>&1 || true
|
||||
}
|
||||
|
||||
# Install bitsandbytes on AMD ROCm hosts. Uses the continuous-release_main
|
||||
# wheel for the ROCm 4-bit GEMV fix (bnb PR #1887, post-0.49.2); bnb <= 0.49.2
|
||||
# NaNs at decode shape on every AMD GPU. Falls back to PyPI >=0.49.1 if the
|
||||
|
|
@ -3296,10 +3341,20 @@ case "$_torch_index_leaf" in
|
|||
if (n > 0) print vals[idx]
|
||||
}')
|
||||
fi
|
||||
# An explicit UNSLOTH_ROCM_GFX_ARCH=gfx906 pins the runtime target to the
|
||||
# MI50 / Radeon VII path and must win over Strix probe-order detection on a
|
||||
# mixed Strix + MI50 host, so the Strix reroute is suppressed when it is set.
|
||||
# Normalize a copied HIP gcnArchName (gfx906:sramecc-:xnack- -> gfx906) and
|
||||
# trim whitespace (mirrors the Python .strip()) so the feature-flag suffix or
|
||||
# a stray newline does not defeat the exact gfx906 comparisons below.
|
||||
_gfx906_env=$(printf '%s' "${UNSLOTH_ROCM_GFX_ARCH:-}" | tr '[:upper:]' '[:lower:]' | tr -d '[:space:]')
|
||||
_gfx906_env=${_gfx906_env%%:*}
|
||||
_strix_gfx=""
|
||||
case "$_runtime_gfx" in
|
||||
gfx1151|gfx1150|gfx1152) _strix_gfx="$_runtime_gfx" ;;
|
||||
esac
|
||||
if [ "$_gfx906_env" != "gfx906" ]; then
|
||||
case "$_runtime_gfx" in
|
||||
gfx1151|gfx1150|gfx1152) _strix_gfx="$_runtime_gfx" ;;
|
||||
esac
|
||||
fi
|
||||
# Skip rocm7.13+ generic indexes: they already ship the fixes, so the
|
||||
# arch build (rocm7.13) would be a downgrade rather than a rescue.
|
||||
if [ -n "$_strix_gfx" ] && _rocm_leaf_below "$_torch_index_leaf" 7 13; then
|
||||
|
|
@ -3327,6 +3382,57 @@ case "$_torch_index_leaf" in
|
|||
TORCHAUDIO_CONSTRAINT="torchaudio>=2.11.0,<2.12.0"
|
||||
_amd_gpu_radeon=false
|
||||
fi
|
||||
# ── MI50 / Radeon VII (gfx906, Vega 20): legacy community-supported path ──
|
||||
# Newer rocm wheel families bundle ROCm libraries whose Tensile kernels
|
||||
# dropped gfx906 (rocBLAS "TensileLibrary.dat ... not read for gfx906",
|
||||
# ROCm/TheRock#1844), so a rocm6.4+/7.x index installs a torch that fails
|
||||
# at the first BLAS call. The rocm6.3 index is the last one whose wheels
|
||||
# run on gfx906 (torch 2.7.0 verified on MI50 32GB; up to 2.9 in community
|
||||
# use). Reroute any newer picked index; leave rocm6.0-6.3 alone.
|
||||
#
|
||||
# Target resolution: an explicit UNSLOTH_ROCM_GFX_ARCH wins (lets a host
|
||||
# whose rocminfo/amd-smi emit no gfx token still opt in; _gfx906_env was
|
||||
# lowercased above, before the Strix block it suppresses). Otherwise only
|
||||
# treat gfx906 as the target when it is the SOLE distinct arch present:
|
||||
# _gfx_all is de-duplicated by visible index, which loses per-device
|
||||
# ordinals on a mixed host, so a non-gfx906 selection must never be
|
||||
# downgraded to rocm6.3 -- such hosts set UNSLOTH_ROCM_GFX_ARCH to opt in.
|
||||
_gfx906_target=false
|
||||
if [ -n "$_gfx906_env" ]; then
|
||||
[ "$_gfx906_env" = "gfx906" ] && _gfx906_target=true
|
||||
elif [ -n "$_gfx_all" ]; then
|
||||
_gfx906_uniq=$(printf '%s\n' "$_gfx_all" | awk 'NF && !seen[$0]++')
|
||||
[ "$_gfx906_uniq" = "gfx906" ] && _gfx906_target=true
|
||||
fi
|
||||
# gfx906 always trains from the PyTorch rocm6.3 wheels, never the Radeon repo
|
||||
# (repo.radeon.com wheels carry no gfx906 BLAS kernels). Clear the Radeon
|
||||
# marketing-name flag as soon as gfx906 is the target -- even when the host
|
||||
# already picks rocm6.0-6.3 and the reroute below is a no-op -- so a Radeon VII
|
||||
# does not divert to the radeon branch on those versions.
|
||||
if [ "$_gfx906_target" = true ]; then
|
||||
_amd_gpu_radeon=false
|
||||
fi
|
||||
if [ "$_gfx906_target" = true ] && ! _rocm_leaf_below "$_torch_index_leaf" 6 4; then
|
||||
echo "" >&2
|
||||
echo " [WARN] gfx906 (MI50 / Radeon VII / Vega 20) detected -- routing torch to the" >&2
|
||||
echo " [WARN] rocm6.3 index: it is the last wheel family that runs on gfx906 (newer" >&2
|
||||
echo " [WARN] rocm wheels ship without gfx906 BLAS kernels and fail at first use)." >&2
|
||||
echo " [WARN] gfx906 is a community-maintained legacy path: 16-bit LoRA and full" >&2
|
||||
echo " [WARN] finetuning work out of the box; bitsandbytes 4-bit QLoRA requires a" >&2
|
||||
echo " [WARN] source build of bitsandbytes for gfx906 (see docs.unsloth.ai/amd)." >&2
|
||||
echo "" >&2
|
||||
_amd_gfx906_base="${UNSLOTH_PYTORCH_MIRROR:-https://download.pytorch.org/whl}"
|
||||
while [ "${_amd_gfx906_base%/}" != "$_amd_gfx906_base" ]; do
|
||||
_amd_gfx906_base="${_amd_gfx906_base%/}"
|
||||
done
|
||||
TORCH_INDEX_URL="${_amd_gfx906_base}/rocm6.3"
|
||||
# Reset to the default (<2.11) window: a rocm7.2 pick raised the floor
|
||||
# to 2.11 above, which the rocm6.3 index (torch <= 2.9.x) cannot satisfy.
|
||||
TORCH_CONSTRAINT="torch>=2.4,<2.11.0"
|
||||
TORCHVISION_CONSTRAINT="torchvision>=0.19,<0.26.0"
|
||||
TORCHAUDIO_CONSTRAINT="torchaudio>=2.4,<2.11.0"
|
||||
# (_amd_gpu_radeon already cleared above for every gfx906 target.)
|
||||
fi
|
||||
;;
|
||||
esac
|
||||
fi # _torch_index_pinned guard (Radeon + Strix reroute)
|
||||
|
|
@ -3553,6 +3659,7 @@ for _p in ('torch', 'torchvision', 'torchaudio'):
|
|||
if [ "$_MIGRATED" = true ]; then
|
||||
# Migrated env: force-reinstall unsloth+unsloth-zoo for a clean state, preserving
|
||||
# existing torch/CUDA unless the ROCm repair below fires.
|
||||
_gfx906_bnb_snapshot
|
||||
substep "upgrading unsloth in migrated environment..."
|
||||
if [ "$SKIP_TORCH" = true ]; then
|
||||
# No-torch: install unsloth + unsloth-zoo with --no-deps (current
|
||||
|
|
@ -3594,13 +3701,18 @@ if [ "$_MIGRATED" = true ]; then
|
|||
# existing ROCm installs gain the AMD bitsandbytes build without a
|
||||
# fresh reinstall.
|
||||
if [ "$SKIP_TORCH" = false ] && [ "$_torch_index_is_rocm_family" = true ]; then
|
||||
_install_bnb_rocm "install bitsandbytes (AMD)" "$_VENV_PY"
|
||||
if _is_gfx906_bnb_skip; then
|
||||
substep "gfx906: skipping prebuilt bitsandbytes (no gfx906 kernels); build from source for 4-bit QLoRA -- https://docs.unsloth.ai/get-started/install-and-update/amd" "$C_WARN"
|
||||
else
|
||||
_install_bnb_rocm "install bitsandbytes (AMD)" "$_VENV_PY"
|
||||
fi
|
||||
# Repair ROCm torch if overwritten during migrated install
|
||||
_has_hip=$("$_VENV_PY" -c "import torch; print(getattr(torch.version,'hip','') or '')" 2>/dev/null || true)
|
||||
if [ -z "$_has_hip" ]; then
|
||||
substep "repairing ROCm torch (overwritten by dependency resolution)..."
|
||||
_install_torch_default_index --force-reinstall
|
||||
fi
|
||||
_gfx906_bnb_prune
|
||||
fi
|
||||
elif [ -n "$TORCH_INDEX_URL" ]; then
|
||||
# Fresh: Step 1 - install torch from explicit index (skip when --no-torch or Intel Mac)
|
||||
|
|
@ -3791,8 +3903,13 @@ elif [ -n "$TORCH_INDEX_URL" ]; then
|
|||
# host stays in GGUF-only mode rather than pulling in bitsandbytes,
|
||||
# which is only useful once torch is present for training.
|
||||
if [ "$SKIP_TORCH" = false ] && [ "$_torch_index_is_rocm_family" = true ]; then
|
||||
_install_bnb_rocm "install bitsandbytes (AMD)" "$_VENV_PY"
|
||||
if _is_gfx906_bnb_skip; then
|
||||
substep "gfx906: skipping prebuilt bitsandbytes (no gfx906 kernels); build from source for 4-bit QLoRA -- https://docs.unsloth.ai/get-started/install-and-update/amd" "$C_WARN"
|
||||
else
|
||||
_install_bnb_rocm "install bitsandbytes (AMD)" "$_VENV_PY"
|
||||
fi
|
||||
fi
|
||||
_gfx906_bnb_snapshot
|
||||
# Fresh: Step 2 - install unsloth, preserving the torch Step 1 installed
|
||||
tauri_log "STEP" "Installing Unsloth"
|
||||
substep "installing unsloth (this may take a few minutes)..."
|
||||
|
|
@ -3843,6 +3960,7 @@ elif [ -n "$TORCH_INDEX_URL" ]; then
|
|||
substep "repairing ROCm torch (overwritten by dependency resolution)..."
|
||||
_install_torch_default_index --force-reinstall
|
||||
fi
|
||||
_gfx906_bnb_prune
|
||||
fi
|
||||
else
|
||||
# Fallback: GPU detection failed to produce a URL -- let uv resolve torch
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue