unsloth/studio/backend/core/training/trainer.py
Leo Borcherding b6d5636cc0
fix/strix halo and windows AMD ROCm support (#5301)
* fix(studio): set HIP_VISIBLE_DEVICES in apply_gpu_ids for ROCm training workers

Training workers are spawned via multiprocessing spawn before detect_hardware()
runs, so IS_ROCM is still False. If the user never set HIP_VISIBLE_DEVICES in
their shell, _inherits_rocm_visibility is also False, leaving the worker with
only CUDA_VISIBLE_DEVICES set. On ROCm hosts the HIP runtime honors
HIP_VISIBLE_DEVICES over CUDA_VISIBLE_DEVICES, so the worker saw the full
device list and torch raised "no usable HIP accelerator" on some setups.

Fall back to probing torch.version.hip (a build-time attribute, safe to read
before GPU init) to detect ROCm when neither IS_ROCM nor inherited env vars
are available. Mirrors the existing fix in llama_cpp.py for llama-server
subprocess GPU pinning.

Fixes https://github.com/unslothai/unsloth/issues/5180

* test: tighten apply_gpu_ids ROCm fallback assertions

Replace loose OR chain with exact string matches, split into three
focused tests, and add a guard check for the try/except wrapper.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: detect ROCm unified memory (Strix Halo / AMD iGPU) via torch fallback

amd-smi on iGPUs with shared/unified memory (e.g. Radeon 8060S on Strix
Halo) reports only the dedicated VRAM slice (~512 MB) in its metric output,
so get_visible_gpu_utilization() was returning usable_gb ≈ 0.35 GB instead
of the full GTT pool (~128 GB).  torch.cuda.mem_get_info() already surfaces
the correct unified-pool size.

Add _reconcile_rocm_unified_memory(): after amd-smi returns a valid result
on a ROCm device, cross-check each device's vram_total_gb against
torch.cuda.mem_get_info().  When torch reports a larger total, replace the
amd-smi VRAM fields in-place.  No-op for discrete AMD GPUs where the two
sources agree.

Fixes: "Falling back to all visible GPUs -- model may not fit" on AMD iGPU
machines even when 100+ GB of unified memory is available.

* Apply unified-memory reconciliation in get_gpu_utilization too

The visible-GPU path was already corrected for AMD iGPUs with unified memory
(Strix Halo / Radeon 8060S), but get_gpu_utilization was still returning the
raw 512 MB amd-smi VRAM slice. Studio's /api/train/hardware endpoint and the
live GPU monitor read from this primary path, so users continued seeing the
wrong total even after auto_select_gpu_ids picked the right device.

Refactor to share the per-device correction:
  * _apply_unified_memory_correction(metrics, torch_info) -- the actual
    replacement logic, in-place on a single metrics dict.
  * _reconcile_rocm_unified_memory(...)                   -- multi-device,
    iterates utilization["devices"] (visible-GPU path).
  * _reconcile_primary_rocm_unified_memory(...)           -- single flat
    metrics dict (primary-GPU path), uses parent_visible_spec to pick the
    primary index, falls back to ordinal 0 when no visibility env is set.

get_gpu_utilization now calls the primary reconciler under IS_ROCM, so both
endpoints surface the real unified-memory pool on iGPUs while leaving
discrete AMD GPUs untouched (torch_total <= smi_total -> no replace).

* Use 'is not None' and log debug on torch.version.hip probe failures

Two small follow-ups to the apply_gpu_ids ROCm fallback:

1. Match detect_hardware()'s 'getattr(torch.version, "hip", None) is not None'
   form so the entire codebase has one canonical 'this torch was built with
   HIP' check. On every shipping torch wheel hip is either None or a non-empty
   version string, so the new form agrees with the old bool() form on every
   real install.

2. Log the probe failure at debug level instead of swallowing it silently.
   The broad 'except Exception' is intentional (we never want apply_gpu_ids
   to crash a worker over a probe), but the silent pass made it impossible
   to tell whether the fallback was firing or being skipped.

* fix(studio): honour HIP_VISIBLE_DEVICES in _get_parent_visible_gpu_spec before IS_ROCM is set

When a user has HIP_VISIBLE_DEVICES set in their shell (e.g. "1" to select
GPU 1) but detect_hardware() has not yet run in the Studio parent process,
IS_ROCM is still False.  _get_parent_visible_gpu_spec() was gated on IS_ROCM
so it fell through to CUDA_VISIBLE_DEVICES (unset), saw all physical GPUs,
and auto-selected index 0.  apply_gpu_ids then overwrote HIP_VISIBLE_DEVICES
with "0", making the intended GPU invisible to ROCm torch in the worker,
which triggered the "no usable HIP accelerator" error (issue #5180).

Apply the same _inherits_rocm_visibility pattern already used in
apply_gpu_ids: check for HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES in the
environment regardless of IS_ROCM so the correct GPU index is preserved.

* fix(install): harden AMD ROCm GPU detection for multi-GPU and env-filtered setups

The previous rocminfo awk pattern could miss discrete GPUs on machines
where HIP_VISIBLE_DEVICES/ROCR_VISIBLE_DEVICES is used to mask an
integrated GPU — the env vars filter rocminfo output but may not
propagate into the install script subprocess, causing detection to
fail entirely.

Two changes:
- Tighten rocminfo pattern from /gfx[0-9]/ && !/gfx000/ to
  /gfx[1-9][0-9]/ — simpler and correctly excludes the CPU agent
  (gfx000) without a negative lookahead
- Add sysfs KFD topology fallback: reads
  /sys/class/kfd/kfd/topology/nodes/*/gpu_id which is a kernel-level
  view unaffected by HIP_VISIBLE_DEVICES or ROCR_VISIBLE_DEVICES

Fixes detection failure reported in Discord by Chains (gfx1201 + iGPU
machine where env var exclusion of the iGPU caused rocminfo to return
no usable device).

* Fix KFD sysfs awk fallback to read properties file

The fallback added by this PR reads /sys/class/kfd/kfd/topology/nodes/*/gpu_id
files but matches the literal token 'gpu_id' against their content. Those
files contain only a single decimal value (e.g. '0' for CPU agents, '50432'
for GPU agents), so the regex never matches and 'found' stays 0, making the
fallback a no-op on every host. The properties file in the same directory
contains key/value lines like 'gpu_id 50432' which is what the existing awk
pattern expects.

Reproduced with a synthetic sysfs layout: against gpu_id files awk exits 1;
against properties files awk exits 0 when any node reports gpu_id > 0.

* fix(setup.ps1): detect AMD ROCm GPU on Windows, bring to parity with setup.sh

setup.ps1 only checked nvidia-smi and fell straight to "gpu: none" on AMD
machines. setup.sh already probed rocminfo/amd-smi/hipconfig/hipinfo.

Add three-tier detection mirroring install_llama_prebuilt.py's detect_host():
1. hipinfo: gcnArchName in output confirms a real HIP GPU (not just SDK)
2. amd-smi list: "GPU: <digit>" data rows as fallback
3. WMI Win32_VideoController: last resort -- detects AMD GPU even without
   HIP SDK, then guides user to install it rather than silently going CPU

Also corrects the "none" message to mention AMD ROCm alongside NVIDIA so
users with AMD hardware understand the requirement.

Fixes: rohit-style install where Strix Halo (Radeon 8060S) showed
"gpu: none" even with the HIP SDK present.

* fix(install.ps1): detect AMD ROCm GPU on Windows, bring to parity with setup.ps1

install.ps1 had the same nvidia-smi-only GPU detection as setup.ps1 before
the setup.ps1 fix. Applies the same three-tier AMD detection:
1. hipinfo: gcnArchName confirms real HIP GPU
2. amd-smi list: GPU data rows as fallback
3. WMI Win32_VideoController: detects AMD GPU without HIP SDK and guides
   user to install it

Fixes: install.ps1 showing "gpu: none" while setup.ps1 correctly showed
"AMD GPU detected" on the same machine (reported by rohit, RX 7600 XT).

* fix(install.ps1): suppress 'No NVIDIA GPU detected' when AMD GPU is present

* feat: add Windows AMD ROCm PyTorch wheel installation

install_python_stack.py:
- Add _ROCM_WINDOWS_WHEEL_BASE and _ROCM_WINDOWS_RELEASES constants
  pointing to AMD repo.radeon.com (ROCm 7.2 -> torch 2.9.1+rocm7.2.1)
- Extend _ensure_rocm_torch() with a Windows branch: detects ROCm via
  _has_rocm_gpu() / _detect_rocm_version(), requires Python 3.12 (cp312
  is the only ABI AMD publishes for Windows), installs the direct wheel
  URL from repo.radeon.com

install.ps1:
- Capture ROCmVersion during AMD detection via hipconfig --version /
  amd-smi version (needed for wheel URL selection)
- After Get-TorchIndexUrl, add an AMD wheel override block: when HasROCm
  and Python 3.12 detected, set ROCmTorchWheelUrl to AMD wheel URL
- Expand torch install branch to handle ROCmTorchWheelUrl with
  uv pip install --force-reinstall --no-cache-dir

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: also install torchvision and torchaudio from AMD Windows repo

AMD publishes matching torchvision-0.24.1+rocm7.2.1 and
torchaudio-2.9.1+rocm7.2.1 cp312 wheels at the same repo.radeon.com
release folder. Install all three in both install.ps1 and
install_python_stack.py Windows ROCm path.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* feat: add ROCm 7.1.1 Windows wheel mapping

AMD uses a different version string for 7.1.1 wheels:
2.9.0+rocmsdk20251116 (date-tagged) instead of +rocm7.1.1.
Adds the 7.1.1 release folder to both install.ps1 and
install_python_stack.py so users with ROCm 7.1 get ROCm
torch instead of falling back to CPU.

* fix: install rocm_sdk_core and rocm_sdk_libraries_custom alongside torch

The AMD Windows torch wheels declare rocm[libraries]==<ver> as a hard
dependency. Without installing rocm_sdk_core and rocm_sdk_libraries_custom
from the same AMD release folder, uv cannot resolve the dependency and
fails with 'No solution found'. Include all 5 wheels in one install call.

* fix: expand ROCm wheel array to scalars for Invoke-InstallCommand

@array splatting inside a scriptblock only works when the native command
is prefixed with '&'. Invoke-InstallCommand uses '& $Command' to run the
block, so @ROCmAllWheelUrls was not being expanded. Extract to scalar
variables $rw0-$rw4 which are captured correctly by the closure.

* fix: use --no-deps for AMD Windows torch wheel install

uv's resolver looks up rocm[libraries]==0.1.dev0 on PyPI during
dependency resolution before downloading any wheels, and fails because
the package doesn't exist on PyPI. --no-deps skips resolution entirely
and installs all 5 AMD wheels directly. The GPU runtime dependency is
satisfied by the HIP SDK, not a Python package.

* fix: setup.ps1 and install_python_stack.py now install ROCm torch on Windows

setup.ps1 was always setting CuTag='cpu' for non-NVIDIA hosts and installing
cpu-only PyTorch, overwriting the ROCm torch installed by install.ps1.
Adds the same AMD wheel selection logic (ROCm version detection, Python 3.12
check, 5-wheel install with --no-deps) to setup.ps1's torch install block.

install_python_stack.py: remove IS_WINDOWS guard from _ensure_rocm_torch()
call site so the Windows path in _ensure_rocm_torch() is reachable during
'unsloth studio update' as well.

* fix: suppress manual-install warning when ROCm torch already present; fix progress counter

- Gate the 'must be installed manually' warning on torch.version.hip being empty
  so it doesn't fire when our ROCm torch install succeeded
- Update _TOTAL counter to include the 3 ROCm steps on Windows now that
  _ensure_rocm_torch() is called there (fixes 10/9 display)

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* feat: add rocm step display in setup.ps1; fix warning and progress counter

- Add 'rocm' step after 'cuda' in setup.ps1 showing ROCm version or HIP SDK missing
- Move ROCm version detection up to GPU detection block so it's available early
- Suppress 'must be installed manually' warning when torch.version.hip is set
- Fix _TOTAL counter to include ROCm steps on Windows (fixes 10/9 display)

* fix: detect AMD SDK ROCm torch via __version__ when torch.version.hip is unset

AMD's repo.radeon.com wheels (e.g. 2.9.0+rocmsdk20251116) do not set
torch.version.hip, leaving it None. All three probes that relied solely on
torch.version.hip now also check for 'rocm' in torch.__version__.lower():

- hardware.py detect_hardware(): IS_ROCM was never set, causing the studio
  to report 'Hardware detected: CPU' even after AMD wheels were installed
  and HIP DLLs were on PATH.
- install_python_stack.py _ensure_rocm_torch(): skip-if-already-installed
  probe would always reinstall on subsequent runs.
- install_python_stack.py Windows AMD warning: suppression check always
  failed, so the 'must be installed manually' note kept appearing after
  a successful AMD wheel install.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* perf: drop --no-cache-dir from AMD ROCm torch wheel installs

uv caches downloaded wheels by default; passing --no-cache-dir forced a
full redownload of the ~2 GB torch wheel on every install run. CUDA installs
never had this flag -- AMD was the only path affected.

* fix: use install-state flag instead of subprocess probe for AMD Windows warning

Replace the subprocess torch probe in the post-install warning block with a
module-level _rocm_windows_torch_installed flag set by _ensure_rocm_torch().
Subprocess re-import of torch is unnecessary and fragile -- the install
function already knows whether it succeeded.

* fix: hoist global declaration to top of _ensure_rocm_torch

Python requires the global statement to appear before any assignment
to the variable within a function. Moving it to the function top fixes
the SyntaxError on line 354.

* fix: pass AMD torch install status via env var to suppress false warning

setup.ps1 now sets UNSLOTH_ROCM_TORCH_INSTALLED=1 after a successful AMD
wheel install. install_python_stack.py reads this at the top of
_ensure_rocm_torch() to skip both the subprocess probe and the warning --
no re-import of torch needed, and the warning message now correctly says
'could not be auto-installed' rather than 'must be installed manually'.

* fix: register ROCm DLL directory before torch import on Windows

Python 3.8+ ignores PATH for extension DLL loading on Windows; amdhip64.dll
and other HIP runtime DLLs must be registered via os.add_dll_directory().
Without this, torch.cuda.is_available() always returns False on AMD ROCm
Windows even when HIP_PATH is correctly set in system environment variables.

Reads HIP_PATH / ROCM_PATH env vars first, then falls back to scanning
common ROCm install roots (C:\Program Files\AMD\ROCm, F:\ROCm, C:\ROCm).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: remove hardcoded non-standard ROCm paths from DLL directory scan

Only use HIP_PATH/ROCM_PATH (set by AMD installer) and the standard
C:\Program Files\AMD\ROCm\<version>\bin location. Custom drive paths
like F:\ROCm are user-specific and should not be hardcoded.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: prevent torchao overrides step from overwriting AMD ROCm torch

torchao==0.14.0 in overrides.txt declares torch as a dependency. Without
--no-deps, uv resolves torch from PyPI and installs 2.11.0+cpu on top of
the AMD ROCm wheels (2.9.0+rocmsdk20251116). This was the root cause of
'Hardware detected: CPU' -- the AMD wheels were installed but then
immediately overwritten by the overrides step.

When _rocm_windows_torch_installed is True, add --no-deps to the overrides
pip_install call so torchao is installed without pulling in CPU torch.

* fix: add rocm_sdk namespace tarball to Windows ROCm wheel installs

torch/_rocm_init.py calls `import rocm_sdk` at startup, which requires
the rocm namespace tarball (rocm-*.tar.gz) in addition to the SDK wheel
packages. This tarball was missing from both install.ps1 and setup.ps1,
causing ModuleNotFoundError on first torch import.

- Add rocm-0.1.dev0.tar.gz to ROCm 7.1.1 install (provides rocm_sdk namespace)
- Add rocm-7.2.1.tar.gz + rocm_sdk_devel to ROCm 7.2.1 install
- Install tarball in a dedicated step before main SDK/torch wheels
- Switch to @array splatting in install.ps1 scriptblock for dynamic wheel count
- Remove --no-cache-dir from Python-side ROCm wheel install (prevents ~2GB redownload)

* feat: enable ROCm 7.2 torch install + warn on gfx1151 with ROCm < 7.2

Chigoma333 (AMD Radeon 8060S / gfx1151, Strix Halo) confirmed that ROCm
7.1 segfaults when tensors are moved to GPU, but ROCm 7.2 + torch
2.11.0+rocm7.2 works fully including training.

Changes:
- Uncomment (7,2): "rocm7.2" in _ROCM_TORCH_INDEX (was blocked by <2.11.0)
- Add _ROCM_TORCH_PKG_SPECS dict with per-tag version bounds:
  rocm7.2 → torch>=2.11.0,<2.12.0; all older tags → <2.11.0
- Add _detect_amd_gfx_codes() helper that parses rocminfo output
- Warn on gfx1151/gfx1150 (Strix Halo) when ROCm < 7.2 is installed,
  pointing users at the known segfault and recommending upgrade
- install.sh get_torch_index_url(): enable rocm7.2 case (previously capped
  to rocm7.1), cap unknown future tags to rocm7.2
- install.sh: override TORCH_CONSTRAINT to >=2.11.0,<2.12.0 when rocm7.2
  index is selected, so pip can actually resolve torch 2.11.0

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: prefer Python 3.12 for AMD ROCm users when 3.13 is also installed

After GPU detection, if ROCm HIP SDK is found and the selected Python
is not 3.12, run a second pass to locate a 3.12 install via py.exe and
PATH (catches uv-managed installs). Switch $DetectedPython to 3.12 so
the venv is created with a compatible interpreter for the cp312-only AMD
Windows torch wheels.

NVIDIA and Intel GPU paths are unaffected -- the re-detection block only
runs when $HasROCm is true.

Fixes: #5301

* fix: also check uv-managed Python 3.12 for AMD ROCm #5301

* fix: hide amd-smi console popups on Windows, guard torch.distributed.is_initialized for ROCm #5301

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: suppress remaining console popups on Windows, patch torch.distributed.is_initialized for ROCm #5301

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: stub all missing torch.distributed attrs for ROCm Windows wheel #5301

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: inject torch.distributed stub when C backend missing in ROCm Windows wheel #5301

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(rocm/windows): pre-stub torch._C._distributed_c10d + raise amd-smi timeout

Two fixes for Windows ROCm regressions reported by electroglyph on #5301:

1. worker.py — torch.distributed stub now fires unconditionally on Windows
   The previous stub only injected sys.modules in the except branch, meaning
   it was silently skipped when `import torch.distributed` happened to succeed
   (the C backend is lazily resolved).  The crash then hit later when
   transformers/trl triggered the lazy load.  Fix: on win32 we pre-populate
   sys.modules['torch._C._distributed_c10d'] AND set the attribute on the
   torch._C extension module *before* attempting the import, covering both
   the early-ImportError and lazy-load failure modes.

2. amd.py — increase amd-smi timeout from 5 s to 30 s on Windows (10 s Linux)
   amd-smi on Windows must cold-init the ROCm runtime on first invocation;
   5 s was consistently too short, producing repeated 'Command timed out'
   warnings in the server log.  30 s gives enough headroom without blocking
   indefinitely on broken installs.

3. install.ps1 — widen Python 3.12 enforcement to ROCmGpuLabel (WMI-only path)
   Users whose HIP SDK is not on PATH were detected via WMI but not switched
   to Python 3.12 before the install started, causing a second pass.  Guard
   now fires on (HasROCm -or ROCmGpuLabel).

* fix(rocm): guard c10d stub, fix TorchIndexFamily for 7.1, clean dead code + comments

- worker.py: wrap c10d stub injection in `if _c10d_key not in sys.modules` so
  Windows NVIDIA users with a real torch.distributed are never affected
- install.ps1: fix Get-TauriTorchIndexFamily receiving hardcoded "rocm7.2"
  even when ROCm 7.1 wheels are installed; now branches on $ROCmVersion
- main.py: remove dead `import ctypes as _ctypes` (ctypes is never called)
- hardware.py, install_python_stack.py, worker.py, install.ps1: shorten
  verbose multi-line comment blocks throughout
- tests: update 4 stale assertions that expected rocm7.2 to be absent/capped

* fix(tests): match windows AMD warning assertion to actual source string

* chore: trim verbose comment blocks across all ROCm-related files

* fix: guard reconcile call against None numeric_ids; add torchvision lower bounds

* fix(install.ps1): recreate venv with Python 3.12 after ROCm switch

Venv was created with 3.13 before GPU detection ran; switching
$DetectedPython to 3.12 had no effect since $VenvPython still
pointed to the 3.13 interpreter inside the already-created venv.

* ux: detect AMD GPU before Python selection to avoid double venv creation

- Early hipinfo + WMI probe runs before Find-CompatiblePython so Python
  3.12 is selected upfront when AMD is detected; venv is now created
  exactly once instead of 3.13 then immediately 3.12.
- Post-venv recreation block replaced with a simple warning for the rare
  case where AMD was missed by the early probe.
- setup.ps1: show venv's actual Python version (e.g. 3.12) instead of
  the system Python found by the pre-activation search (was showing 3.13).

* fix(rocm/win): auto-stub all _distributed_c10d symbols via PEP-562 __getattr__

The bare ModuleType stub caused ImportError when torch._dynamo was imported
(triggered by trainer.py accessing torch._dynamo.config at load time).
torch._dynamo pulls in torch.distributed.fsdp._flat_param which does:
  from torch._C._distributed_c10d import FakeProcessGroup
and potentially other symbols. Adding module __getattr__ auto-creates a
stub class for any missing symbol so all such imports succeed without
enumerating every individual symbol. Applied to both the primary stub
and the fallback stub in the except branch.

* chore: trim c10d stub comment

* fix(rocm/win): auto-stub missing torch.distributed attrs (Store, ProcessGroup, …)

* fix(rocm/win): pre-stub fsdp submodules in sys.modules; fix __getattr__ subpackage clash

* feat(rocm/win): arch-aware wheel selector always picks newest ROCm release

Replace HIP-SDK-version-gated wheel selection with GPU arch-based logic.
Select-ROCmWheelRelease (PS) and _select_windows_rocm_release (Python) map
gcnArchName → minimum ROCm version, then pick the newest available release
that satisfies it (currently always rocm-rel-7.2.1 for any supported GPU).
Wheels bundle their own ROCm runtime so the installed HIP SDK 7.1 does not
prevent using 7.2.1 wheels on gfx1200 (RX 9060 XT) and similar RDNA 4 GPUs.

Also installs the bitsandbytes Windows ROCm continuous-release wheel and sets
BNB_ROCM_VERSION=72 in worker.py before ML imports so bnb loads the
libbitsandbytes_rocm72.dll that ships in that wheel.

* fix(rocm/win): stub class metaclass for ProcessGroup.BackendType; amd-smi circuit breaker

torchao.float8.inference accesses ProcessGroup.BackendType as a class-level
attribute.  Plain type() stubs have no __getattr__ on the metaclass so this
raises AttributeError.  Introduce _StubClassMeta whose __getattr__ returns
child stub classes, fixing the torchao import chain.

Add an amd-smi circuit breaker in amd.py: after 3 consecutive failures the
module stops spawning the process, eliminating the repeated Windows UAC /
DiskPart elevation prompts caused by polling a non-functional amd-smi.

Also guard BNB_ROCM_VERSION=72 behind a DLL existence check so bitsandbytes
fails with its own detection message rather than a harder "DLL not found" when
the Windows ROCm bnb wheel is not yet installed.

* fix: stub __members__ so torchao float8 enum check doesn't crash on ROCm Windows

torchao.float8.inference accesses ProcessGroup.BackendType.__members__
expecting a Python Enum registry dict. _StubClassMeta.__getattr__ was
blocking all dunder attributes, causing AttributeError. Return {} for
__members__ specifically so the isinstance/iteration checks pass cleanly.

* fix: stub distributed tensor/functional_collectives to prevent missing C++ op crash on ROCm Windows

torch._dynamo.trace_rules eagerly loads torch.distributed.tensor at import
time, which pulls in _functional_collectives.py. That file registers Meta
kernels for _c10d_functional C++ ops, but those ops are only registered
by torch._C._distributed_c10d — a C extension absent from ROCm Windows
wheels. Pre-stubbing the affected modules in sys.modules prevents the real
import chain from running and avoids the "operator does not exist" crash.

* fix: give mod stubs __path__ and pre-stub _tensor to fix 'not a package' import error

_make_mod_stub now sets __path__=[] so Python treats stub modules as
packages. Without it, any import of a submodule raises "is not a package".
Also pre-stub torch.distributed._tensor and its submodules so that
_tensor/__init__.py (which re-exports from torch.distributed.tensor) never
runs and torchao's `from torch.distributed._tensor import DTensor` gets a
harmless stub instead of crashing.

* fix: stub torch.ops._c10d_functional namespace with hashable op sentinels

torchao.dtypes.nf4tensor uses _c10d_functional ops as dict keys at import
time (all_gather_into_tensor.default, wait_tensor.default) and
torch.ops.c10d.scatter_.default. None of these ops are registered on ROCm
Windows because torch._C._distributed_c10d (the C extension) doesn't ship.
Replace the whole _c10d_functional namespace with a custom stub whose ops
return hashable .default objects, so dict-key construction doesn't crash.
Also inject a scatter_ stub into torch.ops.c10d if it's missing.

* fix: stub entire torchao package on ROCm Windows instead of individual ops

torchao is not supported on ROCm Windows and its import chain transitively
requires torch._C._distributed_c10d (absent from the ROCm Windows wheel).
Rather than stub each missing op one by one, stub the whole torchao package
upfront. Unsloth uses bitsandbytes for quantization, not torchao, so this
has no functional impact. transformers gracefully handles an importable-but-
empty torchao by disabling TorchAoHfQuantizer.

* fix: set __spec__ on mod stubs so importlib.util.find_spec doesn't raise

Manually-injected sys.modules entries have __spec__=None by default.
importlib.util.find_spec() raises ValueError when it finds a module in
sys.modules with __spec__=None (transformers.utils.import_utils hits this
when checking if torchao is available). Give every stub a minimal
ModuleSpec(name, loader=None, is_package=True) to satisfy find_spec.

* fix: add meta path finder to auto-stub subpackages of stub modules

`import torchao.prototype` goes through the import machinery, not
__getattr__, so an empty __path__ means ModuleNotFoundError. Rather than
list every submodule explicitly, register a MetaPathFinder that intercepts
any import whose parent is one of our stubs (detected by loader=None in the
parent's ModuleSpec). Real installed packages always have a SourceFileLoader
so they are never intercepted. Also register child stubs in sys.modules
from __getattr__ as a belt-and-suspenders measure.

* fix: use _unsloth_stub sentinel instead of loader=None for stub detection

The import machinery overwrites module.__spec__ with the spec returned by
find_spec (which has loader=_StubSubpackageLoader, not None), so the
loader=None check broke for second-level subpackages. Switch to a custom
_unsloth_stub object identity sentinel set directly on each stub module --
it survives __spec__ being replaced and correctly identifies stubs at any
depth (torchao.prototype.safetensors, etc.).

* refactor(rocm/win): switch to repo.amd.com arch-aware index, remove stubs

AMD recommends repo.amd.com/rocm/whl/{arch}/ as the Windows ROCm wheel
source. These wheels bundle their own ROCm runtime, support all Python
versions (not just cp312), and include the full torch._C extension set
(including _distributed_c10d) that the old repo.radeon.com wheel omitted.

Changes:
- install.ps1: remove Select-ROCmWheelRelease + hardcoded cp312 wheel
  URLs; remove Python 3.12 forced-preference logic; install via
  --index-url repo.amd.com/rocm/whl/{arch-family}/
- studio/setup.ps1: same -- remove Select-ROCmWheelRelease, switch to
  repo.amd.com arch-aware index URL
- studio/install_python_stack.py: replace _ROCM_WINDOWS_RELEASES /
  _select_windows_rocm_release with _windows_rocm_index_url() using the
  _GFX_TO_AMD_INDEX_ARCH map; drop Python 3.12 restriction
- studio/backend/core/training/worker.py: remove all stub machinery
  (_make_mod_stub, _StubSubpackageFinder, _StubSubpackageLoader,
  _StubClassMeta, torchao/fsdp/dtensor stubs, _c10d_functional ops
  stubs, BNB DLL detection) -- no longer needed with new wheel source

* fix(rocm/win): restore _distributed_c10d + torchao stubs; fix BNB install

repo.amd.com torch wheels also omit torch._C._distributed_c10d on Windows
(RCCL is not shipped on Windows). torch/distributed/__init__.py imports
from it unconditionally at module level, so the stub must land in
sys.modules before any torch.distributed import.

torchao (pulled in by transformers.quantizers) walks
torchao.float8.distributed_utils -> torch.distributed._functional_collectives
-> distributed_c10d at import time. Stubbing torchao up-front short-circuits
that chain.

worker.py:
- Restore _make_mod_stub / _StubSubpackageFinder / _StubSubpackageLoader
- Restore _StubClassMeta for ProcessGroup.BackendType attribute access
- Restore _distributed_c10d stub with __getattr__ (Windows only)
- Restore torchao stubs (5 modules, Windows only)

install_python_stack.py:
- BNB AMD wheel install was inside the early-return branch that fires when
  torch is already a ROCm build (installed by install.ps1). Move BNB install
  outside that branch so it always runs on Windows ROCm — the PyPI
  bitsandbytes has only CUDA DLLs and fails to load on ROCm.

* worker: remove _distributed_c10d stub; stub only torchao

The installed torch/distributed/__init__.py from repo.amd.com
(torch==2.10.0+rocm7.12.0) is now properly guarded with
`if is_available():`, so `import torch.distributed` alone is safe.

The crash only comes via torchao's import chain:
  torchao.float8.distributed_utils
    → torch.distributed._functional_collectives (unguarded import)
    → torch.distributed.distributed_c10d
    → torch._C._distributed_c10d  ← absent on Windows ROCm

Stubbing torchao short-circuits the chain entirely. No need to stub
_distributed_c10d. Remove _StubClassMeta and the _c10d stub block;
keep only _make_mod_stub + _StubSubpackageFinder + torchao seeds.

* fix: BNB AMD wheel skipped + torch.compile segfault on Windows ROCm

install_python_stack.py: the UNSLOTH_ROCM_TORCH_INSTALLED=1 early-return
path (set by setup.ps1 when it installed torch itself) returned before
ever reaching the AMD BNB prerelease wheel install.  The PyPI
bitsandbytes==0.49.x ships only CUDA DLLs, so loading it on ROCm fails
with "libbitsandbytes_rocm72.dll not found".  Now installs the AMD
Windows BNB wheel before returning on that path too.

worker.py: torch._grouped_mm crashes on gfx1200 (null HIP kernel pointer,
0xC0000005) when torch.compile's JitDecomp system dispatches it during
the first forward pass.  Detect Windows ROCm via torch.version.hip
(already in sys.modules from section 1e) and set TORCHDYNAMO_DISABLE=1
to bypass the broken kernel dispatch.

* fix: BNB AMD wheel install fails uv wheel filename check

The bitsandbytes continuous-release wheel is intentionally mismatched:
filename encodes 1.33.7.preview (= 1.33.7rc0 in PEP 440) but wheel
metadata reports 0.50.0.dev0.  uv rejects this by default.

Introduce _install_bnb_windows_rocm() helper that sets
UV_SKIP_WHEEL_FILENAME_CHECK=1 only for this specific install, then
restores the previous env value.  Both BNB install call sites (the
UNSLOTH_ROCM_TORCH_INSTALLED early-return path and the normal Windows
ROCm path) now use this helper.

* worker: patch _grouped_mm CUDA dispatch on Windows ROCm (gfx1200 null kernel)

TORCHDYNAMO_DISABLE=1 stopped the compiler frontend but not the autograd
JitDecomp system, which also dispatches _grouped_mm and hits the same
null HIP kernel crash (0xC0000005).

Verified that torch.library.Library("aten","IMPL").impl("_grouped_mm", fn,
"CUDA") successfully overrides the broken HIP kernel with a Python mm
fallback on torch==2.10.0+rocm7.12.0.

Schema: _grouped_mm(Tensor self, Tensor mat2, Tensor? offs=None,
                    Tensor? bias=None, ScalarType? out_dtype=None) -> Tensor

The fallback handles both the simple case (offs=None → torch.mm) and the
grouped case (offs provided → split self by offsets, multiply each group
against the corresponding slice of mat2, then cat results).

Keep _WINDOWS_ROCM_GROUPED_MM_LIB alive at function scope to prevent the
C++ dispatch registration from being freed by GC.

* worker: fix torchao stub — return stub classes not modules for isinstance()

peft/tuners/lora/torchao.py does:
  from torchao.dtypes import AffineQuantizedTensor, LinearActivationQuantizedTensor
  isinstance(weight, (AffineQuantizedTensor, LinearActivationQuantizedTensor))

The stub __getattr__ was returning stub modules, which isinstance() rejects
with "arg 2 must be a type, a tuple of types, or a union".

Add _StubTypeMeta metaclass whose __instancecheck__ always returns False,
and _make_stub_type() to create stub classes via it. Change _make_mod_stub
__getattr__ to return stub classes instead of stub modules for leaf
attribute access, so isinstance() gets a valid type and returns False.

_StubSubpackageFinder still handles import-style subpackage creation
(those still need module objects in sys.modules); __getattr__ only fires
for from-import or direct attribute access, which are the isinstance paths.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* tests: add coverage for Windows ROCm install paths and worker patches

Add conftest.py to fix pre-existing sys.path issue that prevented
test_rocm_support.py from running at all (install_python_stack.py
imports from backend.utils.wheel_utils which needs studio/ on sys.path).

New test classes cover everything added in this session:
- TestWindowsRocmIndexUrl: arch → AMD pip index URL mapping (gfx120X-all,
  gfx1151, gfx1150, gfx110X-all, unknown → None, trailing slash)
- TestDetectWindowsGfxArch: hipinfo output parsing, missing/timeout/bad
  returncode/no-gcnArchName paths
- TestInstallBnbWindowsRocm: UV_SKIP_WHEEL_FILENAME_CHECK set+restored,
  env restored on exception, no-op when URL missing
- TestRocmTorchInstalledEnvVar: UNSLOTH_ROCM_TORCH_INSTALLED=1 skips
  pip_install, calls _install_bnb_windows_rocm, sets flag
- TestWorkerWindowsRocmPatches: _grouped_mm CUDA dispatch override,
  offs/grouped variant handling, GC-prevention sentinel,
  _StubTypeMeta __instancecheck__, _StubSubpackageFinder registration,
  torchao key submodule pre-stubbing, TORCHDYNAMO_DISABLE guard
- TestRocmTorchPkgSpecs: rocm7.2 torch 2.11.x spec, default <2.11 cap,
  3-tuple shape, _GFX_TO_AMD_INDEX_ARCH RDNA4/3.5/3 coverage

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* tests: fix encoding, IS_WINDOWS patching, and wrong assertion

- Add encoding="utf-8" to all read_text() calls (54 occurrences) so
  tests pass on Windows where the default codec is cp1252 and source
  files contain UTF-8 emoji (e.g. ⚠️ in install_python_stack.py)
- Add @patch.object(stack_mod, "IS_WINDOWS", False) to Linux-path
  TestEnsureRocmTorch tests so they reach the Linux code path when run
  on a Windows machine instead of short-circuiting into the Windows branch
- Fix test_grouped_mm_patch_guarded_by_windows_and_hip_check: the source
  uses getattr(_torch_for_rocm, "version", None) not torch.version, so
  check for '"version"' and '"hip"' substrings instead

137 passed, 2 skipped

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: pin BNB_ROCM_VERSION=72 for torch==2.11.0+rocm7.13.0 compatibility

AMD's pip index now ships torch==2.11.0+rocm7.13.0 (ROCm 7.13).
bitsandbytes auto-detects HIP 7.13 from torch.version.hip and looks for
libbitsandbytes_rocm713.dll, which the AMD Windows prerelease wheel does
not ship (it only ships rocm72.dll), causing a load error at training start.

Fix:
- worker.py section 1f: set BNB_ROCM_VERSION=72 (via setdefault) before
  section 2 ML imports, so bitsandbytes always loads rocm72.dll on Windows ROCm
- install_python_stack.py: set BNB_ROCM_VERSION=72 in _install_bnb_windows_rocm()
  for any post-install imports; update comment to document root cause
- tests: 4 new assertions covering the fix (141 passed, 2 skipped)

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: detect BNB ROCm DLL suffix dynamically instead of hardcoding '72'

BNB_ROCM_VERSION was pinned to '72' which works today (AMD wheel ships
rocm72.dll) but would break again if AMD ships a future wheel with a
different DLL suffix (e.g. rocm713.dll).

Add _detect_bnb_rocm_dll_ver() to install_python_stack.py: scans the
installed bitsandbytes package dir for libbitsandbytes_rocm{VER}.dll
using importlib.util.find_spec (no BNB import needed) and returns the
suffix.  '72' remains the fallback when detection fails.

Apply the same detection inline in worker.py section 1f.  Both paths
still respect a pre-set BNB_ROCM_VERSION (caller override wins).

Tests: +8 cases covering detection logic and fallback (147 passed, 2 skipped).

* fix: patch torch.distributed stubs in server process for Windows ROCm

On Windows ROCm, torch.distributed ships without process-group helpers
(is_initialized, is_available, get_rank, get_world_size).  The worker
subprocess already patches these in section 1e, but the main server
process calls _determine_attention_impl_for_gpu_estimate() which calls
unsloth's resolve_attention_implementation() → is_initialized(), causing:

  "Could not resolve attention implementation for '...':
   module 'torch.distributed' has no attribute 'is_initialized'"

Fix: patch the missing attrs onto torch.distributed at the top of
_determine_attention_impl_for_gpu_estimate, matching the same stubs
already applied in worker.py section 1e.  No-ops on Linux/CUDA where
torch.distributed is fully populated.

* fix: gate _grouped_mm dispatch patch on HIP < 7.13

AMD fixed the gfx1200 null HIP kernel in ROCm 7.13 (torch 2.11+).
Users on the new wheel now get the real GPU _grouped_mm kernel for
MoE workloads instead of the Python mm fallback.

Changes:
- worker.py: add _hip_ver_at_least() helper; wrap full _grouped_mm
  patch in `if not _hip_ver_at_least(7, 13):` with else branch that
  logs the skip reason; update section-1f comment to document the fix
- test_rocm_support.py: add 5 tests covering the helper definition,
  the (7, 13) gate expression, the else branch, the skip log message,
  and the AMD-format version string parsing (.split(".")[:2])

Verified: torch==2.11.0+rocm7.13.0 — 3D batch and grouped (offs)
variants both succeed; null crash only present on rocm7.12 and earlier.

* fix: stub is_torchelastic_launched on torch.distributed for Windows ROCm

resolve_attention_implementation calls is_torchelastic_launched() which
does not exist in the incomplete torch.distributed shipped with the
Windows ROCm wheel, causing a warning on every model config load in the
server process. Add it to the stub table alongside the four helpers
already patched in _determine_attention_impl_for_gpu_estimate.

Also adds two tests: one confirming the new stub and one confirming all
five core distributed helpers are covered.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: explicit warnings on AMD ROCm arch/version fallbacks + Fast-Install arg order

setup.ps1:
- Fix Fast-Install argument order: packages before flags, consistent with
  all other Fast-Install calls in the file
  (was: Fast-Install --force-reinstall --index-url $url torch ...)
  (now: Fast-Install torch torchvision torchaudio --force-reinstall --index-url $url)
- Add explicit [WARN] substep when $HasROCm is true but arch mapping fails:
  - GPU arch detected but not in supported wheel list → names the arch and
    lists supported families so user knows exactly what to report
  - HIP SDK present (amd-smi path) but gcnArchName unreadable → instructs
    user to re-install the HIP SDK; previously fell back silently to CPU

install.sh:
- Add [WARN] to stderr before silent CPU fallback when AMD GPU is confirmed
  (rocminfo/amd-smi) but ROCm version cannot be read from any source
  (amd-smi, /opt/rocm/.info/version, hipconfig, dpkg, rpm)
- Add [WARN] to stderr when ROCm version is too old (< 6.0) with upgrade link

install.ps1 and setup.sh: no changes needed (already handle these paths correctly)

* fix: robust gfx arch detection for Strix Halo / HIP-runtime-only installs

Covers users who have the HIP runtime (amd-smi available) but not the
full HIP SDK (no hipinfo), which is common on Strix Halo iGPU systems.
Without this, $ROCmGfxArch stays null and the installer silently falls
back to CPU-only PyTorch despite a working GPU.

Detection waterfall (setup.ps1 + install.ps1):
  1. hipinfo gcnArchName          -- full HIP SDK (existing, unchanged)
  2. amd-smi list gfx pattern     -- newer amd-smi versions embed arch
  3. amd-smi static --asic        -- ROCm 6+ ASIC details with GFX target
  4. UNSLOTH_ROCM_GFX_ARCH env    -- manual override escape hatch
  5. GPU name → arch table        -- best-effort from marketing name:
       890M / Strix Halo  → gfx1151 (RDNA 3.5 iGPU, Strix Halo)
       880M / Strix Point → gfx1150 (RDNA 3.5 iGPU, Strix Point)
       780M / Phoenix     → gfx1103 (RDNA 3 iGPU)
       RX 7900/7800/7700  → gfx1100 (RDNA 3 desktop)
       RX 9070 XT / 9080  → gfx1201 (RDNA 4)
       RX 9070 / 9060 XT  → gfx1200 (RDNA 4)

When arch is inferred from name, a Cyan substep tells the user to set
UNSLOTH_ROCM_GFX_ARCH to skip inference on future installs.
WMI block intentionally does not set $HasROCm (no runtime confirmation).

Tests: 11 new tests in TestStrixHaloGfxArchDetection covering all five
detection levels, WMI safety, and gfx regex in both ps1 files.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: resolve hipinfo/hipconfig via HIP_PATH/ROCM_PATH when not on PATH

AMD HIP SDK sets HIP_PATH on Windows but does not always add the bin
directory to PATH.  Get-Command hipinfo therefore silently fails and
detection falls through to WMI, which cannot provide a gfx arch, leaving
the user with a CPU-only PyTorch install and no warning.

Changes:
- setup.ps1 / install.ps1: before falling through to amd-smi, attempt to
  locate hipinfo.exe and hipconfig.exe under $env:HIP_PATH\bin (then
  $env:ROCM_PATH\bin) when Get-Command returns nothing
- Emit a [WARN] with the resolved path and a one-liner to permanently fix
  PATH via SetEnvironmentVariable
- Emit a [WARN] when HIP_PATH/ROCM_PATH is set but the exe is still not
  found (incomplete SDK install)
- Emit a [WARN] with the first hipinfo output line when hipinfo runs but
  returns a non-zero exit code (e.g. "no ROCm-capable device detected")
- 18 new tests in TestHipSdkEnvPathResolution; total 183 passed, 2 skipped

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* feat: print HIP SDK path and full hipconfig version in terminal on AMD detection

Both install.ps1 and setup.ps1 now emit substeps under the gpu step when
AMD ROCm is detected:

  gpu  AMD ROCm (gfx1200)
       HIP SDK: C:\Program Files\AMD\ROCm\7.1
       hipconfig: 7.1.51803-d3a86bd04

Previously only the gpu label (e.g. "AMD ROCm (gfx1200)") was shown with
no indication of where the SDK was found or which exact build was active.
The full hipconfig build string (e.g. 7.1.51803-d3a86bd04 instead of just
7.1) is now stored in ROCmVersionFull and also used in setup.ps1's
'rocm' step label.

9 new tests in TestHipSdkDetectedSubstep; total 192 passed, 2 skipped

* fix: Strix rocm7.1 segfault bypass + Ubuntu 24.04 HIP gcc-install-dir

Issue 1 (install.sh): gfx1151/gfx1150 + ROCm 7.1 causes a segfault in
torch._grouped_mm (moe_utils.py:167). The Radeon repo now ships cp313
wheels for rocm-rel-7.1, so _amd_gpu_radeon=true silently lands on the
broken combo. When Strix Halo/Point is detected and TORCH_INDEX_URL is
rocm7.1, override to rocm7.2 PyTorch index, update TORCH_CONSTRAINT, and
set _amd_gpu_radeon=false to bypass the Radeon repo entirely. Emits a
clear [WARN] explaining the segfault and linking to the ROCm upgrade docs.

Issue 2 (setup.sh): ROCm 7.x ships clang-20 which on Ubuntu 24.04+ picks
/usr/lib/gcc/x86_64-linux-gnu/14/ (runtime dir, no C++ headers), causing
'cstdlib file not found' and a failed llama.cpp HIP build. Iterate gcc
versions 14→11 to find the first install dir that has both runtime and
/usr/include/c++/<ver> headers, then pass --gcc-install-dir to clang via
CMAKE_HIP_FLAGS. Fix confirmed by h34v3nzc0dex (llama.cpp 417/417 clean).

11 new tests across TestStrixRocm71Override and TestSetupShGccInstallDir;
total 203 passed, 2 skipped

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: BNB_ROCM_VERSION in server process + torch._C._distributed_c10d stubs

Two errors visible in training logs on Windows ROCm:

1. Server process bitsandbytes crash:
   "Configured ROCm binary not found at libbitsandbytes_rocm713.dll"
   The installed BNB wheel ships rocm72.dll (not rocm713.dll). The
   training worker already sets BNB_ROCM_VERSION=72 via DLL detection
   but the server process (main.py) imported bitsandbytes before that
   ran. Fix: add the same DLL-scan + BNB_ROCM_VERSION assignment to
   main.py inside the existing win32 guard, before any downstream
   import can pull in bitsandbytes.

2. torch.distributed import failure:
   "No module named 'torch._C._distributed_c10d'; torch._C is not a package"
   torch._C is a C extension on Windows ROCm — Python cannot do
   submodule imports from it, so torch.distributed fails to import
   before our attribute stubs could ever run. Fix: inject empty
   ModuleType stubs for _distributed_c10d, _distributed_autograd and
   _distributed_rpc into sys.modules inside the win32 guard in
   hardware.py BEFORE importing torch.distributed, so the import
   succeeds and our attribute stubs take effect.

9 new tests in TestServerStartupRocmFixes; total 212 passed, 2 skipped

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(win32): populate distributed c10d stub with dummy symbols

torch.distributed tries to `from torch._C._distributed_c10d import
FakeProcessGroup` (and ProcessGroup, Work, Store, etc.).  The previous
empty ModuleType stub caused an AttributeError on those names.

Populate every stub with a _Dummy class for each known symbol so the
import chain completes silently on Windows ROCm where torch._C is a
compiled extension and its _distributed_c10d submodule doesn't exist.

Adds four new tests in TestServerStartupRocmFixes covering FakeProcessGroup,
ProcessGroup, setattr population, and all three _distributed_* siblings.

* fix(win32): distinguish HIP SDK installed vs GPU not ROCm-accessible

Previously, when hipinfo was found but exited non-zero (e.g. "no
ROCm-capable device detected"), both install.ps1 and setup.ps1 fell
through to the WMI-label-only branch and printed "AMD GPU detected --
HIP SDK not found" -- factually wrong since the SDK binary is present.

Add $HipSdkInstalled flag (set true when hipinfo binary is found,
regardless of exit code). When HipSdkInstalled && !HasROCm:
- Show "AMD GPU detected -- not ROCm-accessible (HIP <ver>)" instead
- Explain this is a driver issue, not an SDK issue, with a link
- Still run hipconfig version capture so version shows in output
- CPU-only hint now says "GPU not ROCm-accessible" not "require HIP SDK"

Also applies to setup.ps1 (same detection block, same branches).

Adds TestHipSdkInstalledButDeviceInaccessible (11 tests).

* fix(win32): scope ROCm workarounds to AMD hosts only

Three Codex-flagged issues where Windows ROCm workarounds incorrectly
applied to Windows CUDA (NVIDIA) machines:

main.py (P1): BNB_ROCM_VERSION was set unconditionally on all win32
hosts. On NVIDIA, bitsandbytes sees BNB_ROCM_VERSION and looks for a
ROCm DLL that doesn't exist, breaking bitsandbytes initialisation.
Fix: gate the block on HIP_PATH/ROCM_PATH being present (ROCm hosts only).

worker.py (P2): torchao stubs were seeded for all win32 runs, shadowing
real torchao on Windows CUDA and silently disabling torchao quantization
for NVIDIA users. Fix: gate on HIP_PATH/ROCM_PATH (win32 ROCm only).

install_python_stack.py (P1): _detect_windows_gfx_arch() only checked
shutil.which("hipinfo"), skipping the HIP_PATH/ROCM_PATH fallback that
the PowerShell installers use. On installs where the HIP SDK bin dir is
not on PATH, _ensure_rocm_torch() returned early without installing
ROCm wheels or bitsandbytes. Fix: mirror the env-var fallback.

* fix(linux): route Strix + ROCm 7.1 to AMD arch-specific index

Instead of falling back to pytorch.org/rocm7.2, the Strix override now
routes to repo.amd.com/rocm/whl/gfx1151/ (or gfx1150/) which serves
torch 2.11.0+rocm7.13.0 -- AMD's build containing the actual _grouped_mm
kernel fix, verified on real gfx1151 hardware by h34v3nzc0dex.

This exercises the real GPU kernel path rather than the rocm7.2 workaround.
UNSLOTH_AMD_ROCM_MIRROR can override the base URL for air-gapped installs.

Also teaches _tauri_torch_index_family to recognise AMD arch-specific URLs
(repo.amd.com/rocm/whl/gfx*) and return the rocm7.13 family label so
_tauri_gpu_branch correctly classifies these installs as rocm.

Suggested by h34v3nzc0dex based on hardware-verified probe results.

* fix(studio/rocm): gate ROCm-only side-effects on active torch runtime

Address five edge cases flagged during PR review:

1. studio/backend/main.py: BNB_ROCM_VERSION was set whenever HIP_PATH or
   ROCM_PATH was present in the environment. A Windows CUDA user who once
   installed the HIP SDK and reverted to a CUDA torch wheel still has those
   env vars set, so bitsandbytes would try to load libbitsandbytes_rocm72.dll
   against a CUDA torch and crash. Now probe torch.version.hip inside the
   env-var guard (worker.py already does this).

2. studio/backend/main.py: os.add_dll_directory returned handles were
   discarded. Per CPython docs, the directory leaves the DLL search list when
   the handle is garbage collected. Retain handles in module-level
   _ROCM_DLL_HANDLES list so they survive process lifetime.

3. studio/install_python_stack.py: _install_bnb_windows_rocm() returned None
   regardless of pip_install_try outcome, and the caller flipped
   _rocm_windows_torch_installed to True unconditionally. On a failed BNB
   install the post-install "manual install may be required" warning was
   suppressed and the user was misled. Helper now returns bool; caller gates
   on it.

4. studio/install_python_stack.py: _detect_windows_gfx_arch returned the raw
   capture group, so mixed-case hipinfo output ("Gfx1151") missed the
   lowercase keys in _GFX_TO_AMD_INDEX_ARCH and silently fell back to CPU
   torch. Lowercase the token.

5. studio/install_python_stack.py: UNSLOTH_ROCM_TORCH_INSTALLED=1 early-
   return trusted the env var even when the venv was wiped between runs.
   Subprocess-probe torch importability first; fall through to the full
   install path if the probe fails.

Tests: 231 passed, 1 skipped in tests/studio/install/test_rocm_support.py
(adds one new test for case 5 fall-through).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(studio/rocm): worker.py parity + don't roll back ROCm torch on bnb failure

Addresses findings from a 10x reviewer pass on the prior fix commit:

1. studio/backend/core/training/worker.py (parity with main.py):
   - Gate the torchao stub block on torch.version.hip / 'rocm' in
     torch.__version__ instead of HIP_PATH / ROCM_PATH env-var presence.
     Same root cause as main.py: HIP SDK env vars stick around on CUDA hosts.
   - Add module-level Windows ROCm DLL registration block. Worker subprocesses
     inherit env vars but not the parent's add_dll_directory handles, so the
     first `import torch` in the worker could fail to find amdhip64.dll when
     HIP_PATH\bin is not on PATH. Mirrors main.py setup. Handles retained at
     module scope via _ROCM_DLL_HANDLES.
   - Promote _WINDOWS_ROCM_GROUPED_MM_LIB to module scope with `global` in
     run_training_process so the torch.library.Library registration survives
     past function return / mid-run garbage collection.
   - Harden _torch_has_hip() to also accept 'rocm' in torch.__version__
     (AMD SDK / Radeon wheels may not set torch.version.hip).

2. studio/install_python_stack.py:
   - Don't roll back ROCm torch when bitsandbytes install fails. The prior
     commit gated _rocm_windows_torch_installed on _install_bnb_windows_rocm()
     returning True; if torch installed successfully but bnb failed, the flag
     stayed False and later install steps could overwrite ROCm torch with the
     generic CPU torch wheel. Set the flag after torch install; surface bnb
     failure as a separate warning instead.
   - _detect_windows_gfx_arch now probes in three tiers: UNSLOTH_ROCM_GFX_ARCH
     env-var override (matches the PowerShell installer), then hipinfo (PATH
     or HIP_PATH\bin), then amd-smi (`static --asic`, `list`). Without the
     amd-smi fallback, runtime-only Radeon installs without hipinfo on PATH
     made `studio update` return early and leave the venv on CPU torch.
   - Linux torch-already-rocm probe in _ensure_rocm_torch now matches the
     Windows probe shape: accepts torch.version.hip OR 'rocm' in
     torch.__version__ to cover AMD SDK / Radeon Linux wheels.

3. studio/backend/utils/hardware/hardware.py:
   - apply_gpu_ids() final-fallback torch probe accepts 'rocm' in
     torch.__version__ in addition to torch.version.hip, matching
     detect_hardware(). AMD SDK wheels could otherwise leak through with
     CUDA-only visibility masks on a spawned ROCm worker.

Tests: 231 passed, 1 skipped in tests/studio/install/test_rocm_support.py
(no test changes needed; the probe shape that prints the hip version (or
'rocm' sentinel) preserves the existing non-empty-string contract).

Not addressed in this commit (deferred or out of scope):
- Tag drift / lemonade checksum (PR 5303 surface, not this PR).
- install.sh rocm7.2.1 URL: small fix, separate.
- install.ps1 / setup.ps1 'Radeon 8060S' marketing-name fallback table.
- Strix Halo + ROCm 7.1 routing asymmetry in Python update path.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(studio/rocm): robustness pass - rocm tag normalisation, Strix routing parity, hardened detection

Robustness pass on top of 76137b2d. Four targeted fixes:

1. install.sh ROCm-tag routing normalisation.
   `rocm7.2.1` would route to https://download.pytorch.org/whl/rocm7.2.1
   which does not exist (PyTorch publishes major.minor URLs only). Same
   for any future patch-level tag. Normalise every rocm{maj.min}* pattern
   to the bare {maj.min} index URL.

2. install.ps1 + studio/setup.ps1 marketing-name fallback.
   The gfx1151 row matched 890M / Strix Halo / HX 37x / HX 38x / AI 9 HX
   but not the actual retail name 'AMD Radeon 8060S Graphics' shipped by
   OEMs (Ryzen AI MAX+ 395). Add '8060S' to the regex.

3. install_python_stack.py Strix + ROCm 7.1 routing parity with install.sh.
   The shell installer reroutes Strix Halo / Point + ROCm 7.1 to
   repo.amd.com/rocm/whl/{gfx}/ (which serves torch 2.11.0+rocm7.13.0
   with the upstream _grouped_mm fix). The Python `studio update` path
   only warned and still installed the broken generic rocm7.1 wheel.
   Mirror the override: detect gfx1151/gfx1150 on ROCm 7.1, route to
   the AMD per-gfx index, honour UNSLOTH_AMD_ROCM_MIRROR override.

4. _detect_windows_gfx_arch amd-smi parsing tightened.
   The amd-smi fallback added in the prior commit used a bare
   `\bgfx[1-9][0-9a-z]{2,3}\b` match against the lowercased stdout,
   which could pick up stray gfx references in warnings / device-name
   strings. Anchor on labelled lines first (Target_Graphics_Version,
   ASIC, Arch, gfx) and fall back to the bare match only when no
   labelled line is present.

Tests: 231 passed, 1 skipped in tests/studio/install/test_rocm_support.py;
sim_5301 23 cases pass (6 new sims for the Strix override + amd-smi parsing).

* fix(studio/rocm): multi-GPU selection, Strix sibling handling, defensive cleanups

Round 4 robustness pass based on 5 parallel Opus reviewers of head 21773215.
Seven items from across regression / edge-case / error-paths / architecture
reviews:

1. studio/backend/main.py BNB gate: aligned with the broad ROCm check used
   everywhere else in this PR (torch.version.hip OR 'rocm' in __version__).
   AMD SDK / Radeon Linux wheels do not always populate torch.version.hip;
   without this, main.py would silently skip BNB_ROCM_VERSION while worker.py
   set it.

2. studio/install_python_stack.py _install_bnb_windows_rocm: init _ok = False
   before the try block. Without this, if pip_install_try itself raises
   (e.g. OSError on uv binary missing), the finally block restored env vars
   correctly but the subsequent `if not _ok:` raised UnboundLocalError,
   masking the original exception.

3. studio/install_python_stack.py _detect_windows_gfx_arch:
   - Rewrote to use re.findall (not re.search) on both hipinfo and amd-smi
     output, dedup tokens preserving order, and select via new
     _pick_visible_index() helper.
   - HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES (first comma entry, integer)
     now picks the right GPU on multi-AMD-GPU hosts. Out-of-range or non-int
     values fall back to the first GPU (matches detect_host behaviour in
     install_llama_prebuilt.py).

4. studio/install_python_stack.py Strix override now consults the runtime
   target before flipping:
   - Previous behaviour intersected gfx_codes with {gfx1151, gfx1150} and
     picked the first Strix arch, ignoring whether HIP_VISIBLE_DEVICES
     selected a non-Strix sibling (e.g. discrete RX 7900 in a mixed APU+dGPU
     box). Could install Strix-specific wheels onto a gfx1100 dGPU.
   - Now resolves the runtime gfx via _pick_visible_index() and only
     overrides when that runtime target is in the Strix set.

5. studio/backend/main.py + studio/backend/core/training/worker.py: ROCm
   version dir scan no longer sorts lexically. Previous sort placed "10.0"
   before "7.0" alphabetically, which would mis-prioritise ROCm 10.x bin
   dirs once AMD ships them. New _ver_key() splits on "." and sorts
   numerically with a string fallback.

6. install.sh Strix override URL: replaced ${var%/} (strips one trailing
   slash) with a while-loop that strips all trailing slashes, matching
   Python's .rstrip("/"). A user setting UNSLOTH_AMD_ROCM_MIRROR with
   "http://corp/whl///" no longer ends up with "http://corp/whl///gfx1151/"
   which strict pip proxies (artifactory, sonatype) 404 on.

7. studio/install_python_stack.py: bumped torch import probe timeout from
   30s to 90s. PyTorch's lazy .so loading can take 60-90s on cold NFS or
   USB-backed venvs. The shorter timeout was producing a false "torch
   missing" classification and reinstalling a working ROCm torch.

Tests: 231 passed, 1 skipped. sim_5301 30 cases pass (added 7 new sims for
multi-GPU detection, Strix sibling handling, and _ok-init regression).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(studio/rocm): worker BNB/grouped_mm broad gate, install.sh Strix visibility, runtime-only ROCm detection

Round-5 robustness pass based on 20 parallel reviewers of head 96b9e465.

1. studio/backend/core/training/worker.py - BNB version pin / dynamo disable
   / _grouped_mm fallback block was still gated on torch.version.hip alone
   despite the torchao stub block above already using the broad check. AMD
   SDK / Radeon Windows wheels (torch.__version__ contains "rocm" but
   torch.version.hip is None) silently skipped the Windows ROCm runtime
   patches. Aligned to the same broad check (8/20 reviewers).

2. studio/backend/core/training/worker.py - _hip_ver_at_least() now also
   parses the ROCm version out of torch.__version__ (e.g. "2.11.0+rocm7.13.0")
   when torch.version.hip is missing, so the kernel-fix gate is correct for
   SDK / Radeon wheels too.

3. studio/backend/core/training/worker.py - _grouped_mm_safe_impl with
   offs=None now picks torch.bmm/matmul for 3-D inputs instead of always
   calling torch.mm. The real _grouped_mm accepts 3-D batched matmul; the
   prior fallback raised "self must be a matrix" on MoE workloads (2/20).

4. studio/backend/main.py - dropped the HIP_PATH / ROCM_PATH env-var gate
   from the BNB block; probe torch directly. Runtime-only Radeon / AMD SDK
   Windows installs do not set those SDK env vars but still ship ROCm torch
   (5/20 reviewers).

5. install.sh - Strix override now collects every gfx token from
   rocminfo / amd-smi (in enumeration order), then indexes by
   HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES so a mixed Strix iGPU + non-
   Strix dGPU host where the user selected the dGPU does NOT get rerouted
   to the Strix per-gfx index. Mirrors the Python update path (5/20 reviewers).

6. install.sh - Strix detection chain now also probes `amd-smi static --asic`,
   matching the PowerShell installer (1/20). Closes the gap on runtime-only
   Strix hosts where `amd-smi list` does not surface a gfx token.

7. studio/install_python_stack.py - _has_rocm_gpu() now has the sysfs KFD
   topology fallback (/sys/class/kfd/kfd/topology/nodes/*/gpu_id), matching
   install.sh. On minimal package-managed installs without rocminfo /
   amd-smi GUI tools, `studio update` can now detect the GPU and repair the
   venv instead of returning early (2/20).

8. studio/install_python_stack.py - _detect_amd_gfx_codes() now falls back
   to `amd-smi list` and `amd-smi static --asic` when rocminfo is missing
   (2/20). Strix routing on runtime-only Radeon hosts now matches what
   install.sh has done for a while.

9. studio/install_python_stack.py - Strix override now applies even when
   has_hip_torch is True. The whole point of the override is to repair an
   existing broken torch.version.hip == "7.1" install; skipping the
   reinstall left users on the known _grouped_mm segfaulting stack (3/20).

Tests: 231 passed, 1 skipped. sim_5301 30 cases pass. sim_cross 12 pass.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(studio/rocm): code review hardening pass

- main.py: numeric DLL sort (string sort picked rocm72 over rocm713);
  add basename() to regex; log warning on detection failure; log info
  when BNB_ROCM_VERSION is set (mirrors worker.py)
- worker.py: explicit len-guard in _hip_ver_at_least() with warning
  logs instead of silent IndexError/ValueError swallow
- hardware.py: isinstance(result, dict) guard before result.get() in
  _smi_query() to prevent AttributeError on non-dict backend returns
- amd.py: round() before int() on parsed GPU IDs; log warning when
  truncation occurs (defensive against malformed amd-smi output)
- setup.sh: quote --gcc-install-dir value in CMAKE_HIP_FLAGS so paths
  with spaces do not break the CMake argument
- install.ps1, setup.ps1: apply colon-split + ToLower() to hipinfo
  gcnArchName match (consistent with each other and with setup.sh)
- install.sh: tighten ROCm tag case patterns to explicit
  rocmX.Y|rocmX.Y.* to avoid unintended prefix matches

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(studio/training): GPU OOM guard to prevent system freeze on VRAM exhaustion

On RDNA 4 (gfx1200/gfx1201) and other ROCm GPUs, exhausting VRAM can
cause a HIP driver hang that freezes the entire system rather than
raising a recoverable Python exception.

Two-part fix:
- set_per_process_memory_fraction(0.90) caps the HIP/CUDA allocator at
  90% of VRAM so PyTorch raises OutOfMemoryError before hitting the
  hardware limit, keeping the driver alive and the system responsive
- top-level exception handler detects OOM errors by type and message
  and surfaces a clear actionable message to the UI (reduce
  max_seq_length, enable gradient_checkpointing, lower batch size)
  instead of the raw CUDA/HIP error string

* fix(studio/rocm): OOM guard ROCm-only + unified memory, multi-GPU arch selection

OOM guard (worker.py):
- Scope to _hw.IS_ROCM only -- NVIDIA CUDA has a graceful OOM path and
  does not need the allocator cap
- Detect unified memory by comparing torch VRAM against psutil system RAM;
  use 0.80 on unified-memory APUs (gfx1151 Strix Halo) where the GPU pool
  is carved from host RAM, 0.90 on discrete cards

Multi-GPU arch selection:
- install.ps1 / setup.ps1: replace -match (first hit only) with
  [regex]::Matches() to collect all gcnArchName entries, then index by
  HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES
- install_python_stack.py: index into full token list before dedup so
  HIP_VISIBLE_DEVICES=2 on [gfx1100, gfx1100, gfx1151] resolves gfx1151
- install.sh: remove awk dedup from gfx token collection for same reason

GCC multiarch (setup.sh):
- Only append -linux-gnu when gcc -print-multiarch does not already return
  the full triple, fixing double-suffix on Ubuntu 24.04

* fix(tests): update ROCm version cap expectations from rocm7.1 to rocm7.2

Daniel's normalisation commit updated the cap from rocm7.1 to rocm7.2
since PyTorch now publishes that index and rocm7.2 ships torch 2.11.0.
Test expectations were stale.

* fix(tests): correct MLX smoke test losses_per_step assertion

logging_steps=1 with max_steps=30 produces 30 loss entries, not 7.
The assertion was stale from a previous config.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(studio/worker): detect unified-memory APU by GPU name not VRAM/RAM ratio

The previous heuristic (VRAM > 50 % of system RAM) false-positived on discrete
cards in low-RAM systems — e.g. RX 9060 XT 16 GB on a 16 GB or 24 GB machine
would trip the unified-memory path and log "unified memory host" when it should
say "discrete".

AMD iGPUs (gfx1150/gfx1151 Strix Halo, Strix Point, etc.) expose names with a
digit+M suffix ("AMD Radeon 890M"), while discrete cards use "RX NNNN [XT|XTX]"
naming.  Matching that suffix is reliable across all current ROCm-capable AMD
consumer GPUs and does not require psutil.

Also includes the device name in the log line to ease future debugging.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(install/setup.ps1): force array on hipinfo gcnArchName parse to fix single-GPU arch truncation

When [regex]::Matches() finds exactly one match, PowerShell's pipeline
unwraps the result to a scalar string.  Indexing a scalar string with [0]
returns the first *character*, so a one-GPU system would parse
gcnArchName "gfx1200" as "g", which is not in the supported arch map
and triggers the CPU-only fallback.

Wrapping with @() forces the result to remain an array regardless of
match count.  On a single-GPU machine the arch is now correctly read as
"gfx1200" (or whatever the full name is) so the ROCm wheel index is
selected.

Reproducer: hipinfo exits 0 and outputs exactly one gcnArchName line.
Without @(), $_hipAllArches = "gfx1200" (String); $_hipAllArches[0] = 'g'.
With @(), $_hipAllArches = @("gfx1200") (Object[]); $_hipAllArches[0] = "gfx1200".

* fix(studio/rocm): classify unified-memory APU via VRAM/RAM ratio, not arch list

Replace the gcnArchName allowlist {gfx1150, gfx1151} with a
psutil-based heuristic: unified APUs expose the entire system RAM
as the HIP pool (ratio ≥ 0.90), discrete cards are well below that.
No arch name required — future APUs classify correctly without code changes.

Also removes the stale import re / \d[Mm]\b device-name regex that
5d84704 left behind, and logs vram/sys GiB for easier on-hardware
verification.

Addresses h34v3nzc0dex review: Radeon 8060S (gfx1151, 128 GiB
unified) now correctly gets 0.80 cap instead of 0.90.

* fix(studio/rocm): revert to gcnArchName for unified-memory APU classification

VRAM/RAM ratio >= 0.90 false-positives on machines where discrete VRAM
equals system RAM (e.g. RX 9060 XT 16 GB + 16 GB system RAM → ratio 1.0,
incorrectly classified as unified → wrong 0.80 cap applied).

gcnArchName is the correct signal: naming-independent, stable within a
product family, and already parsed throughout this PR. Unified set is
{gfx1150, gfx1151} (Strix Point + Strix Halo).

* fix(studio/llama-prebuilt): resolve hipinfo via HIP_PATH/ROCM_PATH on Windows

shutil.which("hipinfo") returns None when the HIP SDK bin dir is not on
PATH -- the HIP SDK installer sets HIP_PATH/ROCM_PATH but does not always
add the bin dir to PATH. This caused has_rocm=False in the prebuilt asset
selector, so AMD ROCm machines got the CPU llama.cpp zip instead of the
HIP one, silently running all chat inference on CPU.

Add _resolve_exe() that falls back to %HIP_PATH%\bin and %ROCM_PATH%\bin
when shutil.which() finds nothing, mirroring the same fallback already
present in setup.ps1.

* fix(studio/llama-prebuilt): pass --has-rocm from setup.ps1 to skip re-detection

The Python prebuilt installer re-detects ROCm independently via
shutil.which("hipinfo"), which fails when hipinfo is not on PATH
(HIP SDK sets HIP_PATH but doesn't always add the bin dir to PATH).
This caused has_rocm=False and downloaded the CPU llama.cpp zip even
on confirmed AMD ROCm machines.

setup.ps1 already performs reliable ROCm detection with its own
HIP_PATH/ROCM_PATH fallback. Add --has-rocm flag to
install_llama_prebuilt.py so setup.ps1 can forward its result directly,
and pass it whenever $HasROCm is true. The Python script then overrides
has_rocm=True in the HostInfo without re-probing.

* fix(studio/llama-prebuilt): add HIP asset to simple-policy Windows path

direct_upstream_release_plan (used by --simple-policy, which setup.ps1
always passes) only checked has_usable_nvidia on Windows and fell
straight to CPU for AMD ROCm machines, ignoring has_rocm entirely.
The --has-rocm override had no effect because the simple-policy code
path never reached resolve_asset_choice where has_rocm was checked.

Add an elif branch for has_rocm that tries the upstream HIP asset
(llama-TAG-bin-win-hip-radeon-x64.zip) before falling through to the
CPU fallback, consistent with the non-simple-policy path.

* fix(studio/setup.ps1): auto-remove mismatched llama.cpp install kind

When an existing llama.cpp install is the wrong kind for the current
GPU (e.g. windows-cpu on an AMD ROCm machine that should have
windows-hip), the prebuilt installer skips on tag match and never
upgrades. Read install_kind from UNSLOTH_PREBUILT_INFO.json before
invoking the installer and remove the directory if the kind doesn't
match, forcing a fresh download of the correct variant.

* fix(studio/setup.ps1): show live PyTorch install output in verbose mode for ROCm

The ROCm torch reinstall (setup.ps1 phase) always silently captured
output, so in --verbose mode the torch downgrade mid-install
(2.11.0+rocm → 2.10.0 → 2.11.0+rocm) looked like the final state was
2.10.0. Match the CPU/CUDA blocks which show live uv output when
$script:UnslothVerbose is set.

* fix(rocm/windows): set ROCBLAS_TENSILE_LIBPATH for bundled rocblas.dll

The llama.cpp ROCm prebuilt bundles rocblas.dll next to the binary but
not the Tensile kernel library files it depends on at runtime
(rocblas/library/TensileLibrary*.dat + *.hsaco).  The bundled DLL
searches for these files relative to its own location by default, i.e.
<binary_dir>/rocblas/library/, which does not exist in the prebuilt
install tree.  This causes a silent crash on the very first GEMM
(prefill) with no output from llama-server, seen by the caller as
WinError 10054 / 10061.  Model load and the single-token warmup pass
because they use simpler code paths that do not trigger rocBLAS GEMM.

Fix: set ROCBLAS_TENSILE_LIBPATH in the subprocess env to
<HIP_PATH>/bin/rocblas/library so the bundled DLL finds the kernel
files from the system ROCm installation.  Uses setdefault so a user-
supplied env var is never overwritten.  No-ops on CUDA and CPU (no
HIP_PATH) and on Linux (win32 branch only).

Reproducer log:
  rocBLAS error: Cannot read .../Release/rocblas/library/TensileLibrary.dat
  rocBLAS error: Could not initialize Tensile host:
  directory_iterator: The system cannot find the path specified.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(install.sh): restore gfx token dedup in Strix multi-GPU awk indexer

536a54df removed the per-source `| awk '!seen[$0]++'` dedup from the
_gfx_all collection step but left the indexer awk as bare NF, so on a
mixed-arch host (e.g. dGPU gfx1100 + Strix iGPU gfx1151) where
rocminfo emits each gfx token twice (Name: field + ISA triple),
HIP_VISIBLE_DEVICES=1 indexed vals[1] = the second gfx1100 occurrence
instead of gfx1151, triggering the Strix routing on the wrong GPU.

Add !seen[$0]++ to the indexer awk so duplicate tokens from the same
GPU collapse to one entry before the HIP_VISIBLE_DEVICES index is
applied -- matching exactly what the Python side does with dict.fromkeys()
in _detect_amd_gfx_codes(). The comment above the block ("skip
duplicates") already documented this as the intended behaviour.

* fix(studio/install): correct _TOTAL progress count on Windows

base_total += 3 fired for all non-macOS platforms including Windows,
but flash-attn (line 1620) and ROCm torch final (line 1705) are both
guarded by 'not IS_WINDOWS and not IS_MACOS', so on Windows with torch
enabled _TOTAL was 13 while only 11 _progress() calls actually execute.

Split into +1 for the ROCm torch check (all non-macOS) and +2 for the
two Linux-only steps, so Windows gets _TOTAL=11 and Linux gets 14.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(install.ps1): enforce torch>=2.11.0 for gfx120X and Strix on Windows

The AMD arch-specific index (repo.amd.com/rocm/whl/gfx120X-all/ and
gfx1151/) publishes torch wheels from 2.7.1 through 2.11.0. Without a
version floor pip can resolve to torch 2.10.0+rocm7.12 on RDNA 4
(gfx120X) or torch 2.10.0+rocm7.1 on Strix (gfx1151/gfx1150), both of
which have a null-pointer crash in torch._C._grouped_mm (TheRock
issues #5284 / #3284). torch 2.11.0+rocm7.13 contains the fix.

Add $ROCmTorchFloor alongside $ROCmIndexUrl: set to torch>=2.11.0 for
the two affected arch families, null for all others. Wire it into the
uv pip install call so the broken wheels are never selected.

* fix(rocm/windows): address Codex nits - deterministic DLL suffix, CUDA llama.cpp kind, HIP_VISIBLE_DEVICES arch indexing

- install_python_stack.py / worker.py: _detect_bnb_rocm_dll_ver() and the
  inline worker probe now collect ALL libbitsandbytes_rocm*.dll suffixes and
  return max() by numeric value instead of stopping at the first glob hit.
  Filesystem glob order is not guaranteed; this ensures '713' always wins
  over '72' when both variants are present in the wheel.

- setup.ps1 (expectedKind): add 'windows-cuda' branch so NVIDIA hosts are
  not treated as 'windows-cpu'. Previously an existing windows-cuda prebuilt
  was always considered a mismatch on non-ROCm machines, forcing an
  unnecessary re-download on every update.

- setup.ps1 (amd-smi gfx arch): collect ALL gfx tokens from amd-smi list
  output in GPU order and honour HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES
  when selecting which arch to use. On mixed-arch AMD systems where the
  visible GPU is not the first enumerated one, this prevents installing an
  incompatible wheel index. Falls back to index 0 (same as before) when the
  visibility var is unset or is a comma-separated list.

- test_rocm_support.py: add test_picks_highest_suffix_when_multiple_dlls to
  cover the multi-DLL case that was previously untested.

* fix(rocm): misleading amd-smi log, BNB spec consistency, torch ceiling for AMD index

amd.py: split 'returncode != 0 or not stdout' into two separate branches.
Previously, exit-0 with empty output logged 'amd-smi returned code 0' (which
reads as success, not a warning) and incorrectly incremented the circuit-breaker
counter. Now: non-zero exit logs the code and counts toward the limit as before;
empty stdout on exit 0 logs at DEBUG level and does not penalise the counter
(amd-smi --json always emits at least [] on exit 0, so this branch is rare and
is not a tool failure).

main.py: replace spec.origin / os.path.dirname() with
spec.submodule_search_locations to match install_python_stack.py and worker.py.
For normal wheel installs both approaches reach the same directory, but using
submodule_search_locations is the canonical way and handles editable bitsandbytes
installs correctly. Also use max() by numeric suffix (same as the other two sites)
instead of a sort-then-break loop.

install.ps1: add <2.12.0 ceiling to the torch constraint for gfx120X (RDNA 4)
and gfx1151/gfx1150 (Strix). AMD actively publishes new versions on their
per-arch index; without a ceiling, a future 2.12.0+rocmX.Y wheel would be
pulled in automatically before being validated on these architectures. The
ceiling matches the existing Linux install_python_stack.py constraint for the
same arches. Bump both when 2.12.x is confirmed working.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(rocm): torch floor in setup.ps1, torchvision pin for Strix, rocmsdk in _hip_ver_at_least

setup.ps1: add \ (mirrors install.ps1) and derive \
from it. Previously the AMD index install called 'Fast-Install torch torchvision
torchaudio --force-reinstall --index-url \' with no version
constraint, so pip could resolve torch 2.10.0+rocm7.12 for gfx1151/gfx1200 --
the exact broken wheel the PR is meant to avoid. Now gfx120X and Strix enforce
'torch>=2.11.0,<2.12.0', matching install.ps1 and the Linux constraint.

install_python_stack.py: pin torchvision and torchaudio in _strix_override_pkgs.
The Strix Linux override uses --index-url (exclusive, no PyPI fallback); bare
unversioned 'torchvision' and 'torchaudio' could resolve a build from AMD's
index targeting a different torch major, causing ABI/version mismatches at
runtime. Now pinned to '>=0.26.0,<0.27.0' and '>=2.11.0,<2.12.0' respectively,
matching _ROCM_TORCH_CONSTRAINT['rocm7.2'].

worker.py: extend _hip_ver_at_least to handle AMD SDK wheel version strings.
The fallback regex r'rocm(\d+)\.(\d+)' cannot match '2.9.0+rocmsdk20251116'
(no rocmX.Y component), so the function always returned False on SDK/Radeon
wheels -- installing the Python _grouped_mm workaround on wheels that already
have the working HIP kernel. Added a second check: if the version string
contains '+rocmsdk', assume >= 7.13 (the rocmsdk format post-dates the
gfx120X null-kernel fix) and skip the fallback.

* fix(rocm): warn on OOB HIP_VISIBLE_DEVICES, bail on empty numeric_ids mask

- setup.ps1: when HIP/ROCR_VISIBLE_DEVICES names an index beyond the
  detected GPU count, emit a yellow warning and fall back to GPU 0
  instead of silently reading allGfxArches[-1] (wrong arch)
- hardware.py _reconcile_primary_rocm_unified_memory: distinguish
  numeric_ids=None (no env var, use torch ordinal 0) from numeric_ids=[]
  (empty mask / HIP_VISIBLE_DEVICES=-1, no GPU visible); bail out early
  in the empty case to avoid querying torch.device(0) incorrectly

* fix(rocm): gate StubSubpackageFinder on win32 ROCm, add gcnArchName fallbacks

- worker.py _StubSubpackageFinder: the meta_path append was running on
  every platform on every call to run_training_process; moved it inside
  the if _is_win32_rocm: block since stubs are only seeded there and the
  finder is a pure accumulation on Linux/Windows CUDA
- worker.py OOM guard: AMD SDK / Radeon wheels may not populate
  gcnArchName, causing Strix Halo to be misclassified as discrete and
  get the 0.90 cap (12.8 GB OS headroom) instead of 0.80 (25.6 GB);
  now tries gcn_arch_name / arch_name / gfx_arch_name variants first,
  then falls back to device-name matching (890M -> Strix Halo,
  880M -> Strix Point) with a debug log when the fallback fires

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(rocm): pin torchvision/torchaudio in setup.ps1, remove -Unique from arch array

- setup.ps1 ROCm torch install: torchvision and torchaudio were passed
  bare alongside pinned torch>=2.11.0,<2.12.0 for gfx1151/gfx1200 arches.
  AMD publishes packages independently so a future torchvision 0.27 (for
  torch 2.12) on the same arch index would cause pip ResolutionImpossible
  or an ABI-incompatible install. Added torchvisionFloorMap and
  torchaudioFloorMap mirroring install_python_stack.py's strix override
  (torchvision>=0.26.0,<0.27.0, torchaudio>=2.11.0,<2.12.0) and derived
  ROCmVisionSpec/ROCmAudioSpec used in all three Fast-Install call sites.

- setup.ps1 amd-smi arch detection: Select-Object -Unique was collapsing
  same-arch multi-GPU arrays (e.g. two gfx1151 APUs -> 1-element array)
  causing HIP_VISIBLE_DEVICES=1 to trigger a false out-of-range warning
  and fall back to GPU 0 even though the correct GPU would have been at
  index 1. Removed -Unique; added comment noting the positional-index
  assumption and its non-contiguous-GPU limitation.

* fix(rocm): add 8060s/8050s to OOM guard device-name fallback, extract classifier helper

Path 3 of the OOM guard device-name fallback only checked for 890m/880m
(gfx1150 Strix Point SKU names). Strix Halo (gfx1151) ships as Radeon 8060S
(Ryzen AI MAX+ 395) and Radeon 8050S (cut-down SKU) -- neither matches, so
the fallback returned is_unified=False and applied the 0.90 fraction instead
of 0.80, leaving ~12.8 GiB OS headroom on a 128 GiB pool instead of ~25.6 GiB.

Fix: add 8060s and 8050s to the name-match set. Also correct the comment that
mislabelled 890M as a Strix Halo name (it is Strix Point).

Refactor: extract the three-path classifier into _rocm_classify_unified_memory()
so it can be unit-tested directly. Add 31 test cases in test_rocm_oom_guard.py
covering all three paths and the regression case (Radeon 8060S Graphics).

Reported-by: h34v3nzc0dex

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(rocm): pass explicit dtype on bf16-unsupported hardware (RDNA2)

dtype=None lets unsloth auto-detect the model dtype. On RDNA2 (gfx103x,
e.g. RX 6600) is_bfloat16_supported() incorrectly returns True, so unsloth
picks bf16 and the first bf16 kernel dispatch triggers:

  LLVM ERROR: Cannot select: intrinsic %llvm.amdgcn.fdot2.bf16.bf16

Replace every dtype=None in load_model() with _auto_dtype which resolves
to None when bf16 is supported (all modern NVIDIA + RDNA3+) and
torch.float16 otherwise. This gives RDNA2 users a working float16
training path without touching NVIDIA behaviour at all.

Fixes: https://github.com/unslothai/unsloth/issues/5337

* fix: reduce log noise for expected non-issues on Windows ROCm

Three log lines fired at warning/error level for conditions that are
completely expected on a Windows HIP SDK-only setup:

amd.py
- amd-smi WinError 2 (FileNotFoundError): downgrade warning -> debug.
  amd-smi ships with Adrenalin, not the HIP SDK; absence is normal.
- 'disabling' message: downgrade warning -> info with clearer text
  'not available (not installed; expected on HIP SDK-only systems);
  GPU VRAM polling disabled'

hardware.py
- torch.distributed.Store missing: downgrade warning -> debug.
  The distributed stub added in this PR intentionally omits Store; the
  attention-impl fallback to eager is expected and non-actionable.

worker.py
- causal-conv1d: add early Windows exit (info) in both
  _ensure_causal_conv1d_fast_path and _causal_conv1d_install hook;
  no cp313/win_amd64 wheel exists, so the install always fails.
- FLA: add early Windows exit (info) in
  _ensure_flash_linear_attention_unconditional; triton dependency has
  no cp313/win_amd64 wheel.
- Defense-in-depth: _install_package_wheel_first non-HIP PyPI failure
  logs info+debug on Windows instead of error; FLA failure logs
  info+debug on Windows instead of warning.

* [AMD] FIx installation of bitsandbytes when it's from .dev and skip rebuilding llama.cpp if we build it manually.

* fix: use force_pip for Windows ROCm bitsandbytes prebuilt wheel install

uv rejects the bnb continuous-release wheel due to filename/metadata
version mismatch (1.33.7.preview vs 0.50.0.dev0). Switch to force_pip=True
(pip bypass) instead of the UV_SKIP_WHEEL_FILENAME_CHECK env var workaround
-- cleaner and consistent with how the Linux path handles it.

BNB_ROCM_VERSION is still set post-install to the detected DLL suffix so
the worker subprocess loads the correct libbitsandbytes_rocm{VER}.dll even
when torch.version.hip reports a newer HIP version than the wheel ships.

* fix: three small correctness fixes found in PR review

- _install_bnb_windows_rocm: use UV_SKIP_WHEEL_FILENAME_CHECK=1 with
  try/finally instead of force_pip=True so the env var is always
  restored and the failing CI test passes
- _determine_attention_impl_for_gpu_estimate: gate torch._C distributed
  stubs on IS_ROCM so Windows CUDA users keep the real extension
- install.ps1 amd-smi fallback: collect all gfx tokens and index by
  HIP_VISIBLE_DEVICES, matching the hipinfo path on multi-GPU hosts

* fix: stub torchao in export subprocess on Windows ROCm

On Windows, the ROCm build of PyTorch ships without the distributed
C extension (torch._C._distributed_c10d). torchao, which is pulled in
transitively by transformers.quantizers at import time, walks into
torch.distributed._functional_collectives -> distributed_c10d and
crashes with:

  No module named 'torch._C._distributed_c10d'; 'torch._C' is not a package

This only affected the export subprocess because the training subprocess
already applied an identical torchao stub (introduced separately to fix
the same root cause). The export subprocess had no such guard and died
during 'Importing Unsloth...' before any model loading could happen.

Fix: apply the same _StubSubpackageFinder / torchao stub pattern to the
export subprocess entry point, gated on Windows ROCm detection, before
any import of transformers or unsloth_zoo.

Root cause tracked in ROCm/TheRock#3284 (libuv / torch.distributed
missing on Windows ROCm builds).

Ref: https://github.com/ROCm/TheRock/issues/3284

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* install.sh, setup.sh: add GPU arch step logging to match PS1 scripts

Both shell scripts were missing the step "gpu" terminal log block that
install.ps1 and setup.ps1 emit. This adds equivalent output: GPU label
with gfx arch (e.g. "AMD ROCm (gfx1151)"), ROCm root path, hipconfig
version, and marketing name substep. Includes the same gfx arch detection
chain (rocminfo → amd-smi list → amd-smi static --asic), UNSLOTH_ROCM_GFX_ARCH
env override, and name-based arch inference table (Strix Halo/Point, RDNA 3/4)
as the PS1 versions. install.sh also replaces bare echo blocks for the AMD
ROCm and CPU-only cases with formatted substep output.

* Fix BNB_ROCM_VERSION gate, ROCm GPU mask preference, APU unified memory and Release build for PR #5301

- main.py: gate BNB_ROCM_VERSION on the rocm bnb DLL or HIP_PATH/ROCM_PATH instead of importing torch on every Windows host
- hardware.py: prefer HIP/ROCR visible-device masks only on ROCm hosts so a stale mask cannot override CUDA_VISIBLE_DEVICES on NVIDIA
- llama_cpp.py: set GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 only for unified-memory APUs (gfx1150/gfx1151)
- setup.sh: pass -DCMAKE_BUILD_TYPE=Release for the HIP source build
- add test_amd_apu_unified_memory.py

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: guard recompile_limit + fix AMD VRAM monitor fallback

trainer.py: torch._dynamo.config.recompile_limit does not exist in
some ROCm torch builds (e.g. pytorch.org/whl/rocm6.2 wheels). Guard
the assignment so training doesn't crash on RDNA2/RDNA3.

hardware.py: when amd-smi/nvidia-smi is unavailable or returns no
usable data (HIP SDK-only Windows, Docker, unexpected JSON format),
the existing fallback used torch.cuda.memory_allocated() which is
process-specific and reads near-zero even with a fully loaded model.
Switch to torch.cuda.mem_get_info() via _torch_get_per_device_info()
which reports system-wide VRAM occupancy so the GPU monitor shows
real usage on all AMD systems without requiring amd-smi.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: Windows VRAM monitor via Performance Counter API

When amd-smi/nvidia-smi is unavailable on Windows, query dedicated GPU
VRAM via Windows Performance Counters (same source as Task Manager).
This gives system-wide cross-process usage, fixing the near-zero reading
caused by torch.cuda.mem_get_info only seeing the Studio server process.

Linux fallback path unchanged (mem_get_info is system-wide on ROCm).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: rename to _rocm_windows_perf_counter_vram_gb, scope to IS_ROCM

Function is AMD ROCm specific — amd-smi absent on Windows when only the
HIP SDK is installed. Scoped to IS_ROCM so NVIDIA Windows path is
untouched (nvidia-smi handles that case).

* fix: AMD VRAM monitor — Linux DRM sysfs + Windows perf counter

Linux: read /sys/class/drm/card*/device/mem_info_vram_used|total for
system-wide GPU memory across all processes. No tools required, always
present on Linux AMD systems.

Windows: Windows Performance Counter API (already added).

Both paths are gated on IS_ROCM and only fire when amd-smi is absent.
torch mem_get_info remains as last resort (process-local).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: AMD GPU monitor — utilization, temperature, and power for Windows and Linux fallback paths

- Windows: GPU utilization via \GPU Engine(*engtype_3D*)\Utilization Percentage perf counter
- Windows: temperature and power via ADL (atiadlxx.dll, ships with Adrenalin)
- Linux: GPU utilization via DRM sysfs gpu_busy_percent
- Linux: temperature via hwmon temp1_input (millidegrees C)
- Linux: power via hwmon power1_average / power1_input (microwatts)

All paths are no-op fallbacks (None) when the source is unavailable.
Mirrors what nvidia-smi provides on the CUDA path.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: remove ADL ctypes — does not support AMD iGPU (Strix Halo)

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com>
Co-authored-by: Erland366 <erland.pg366@gmail.com>
Co-authored-by: danielhanchen <michaelhan2050@gmail.com>
2026-05-29 22:29:56 -07:00

3757 lines
156 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""
Unsloth Training Backend
Integrates Unsloth training capabilities with the FastAPI backend
"""
import os
import sys
# Prevent tokenizer parallelism deadlocks when datasets uses multiprocessing fork
os.environ["TOKENIZERS_PARALLELISM"] = "false"
# Ensure compiled cache modules are importable by any subprocess.
# On spawn-based platforms (Windows, macOS), spawned dataset.map() workers must
# re-import all top-level modules. The compiled cache's trainer files import
# torch and unsloth_zoo (which initializes CUDA), making spawn impractical.
# Propagating UNSLOTH_COMPILE_LOCATION via PYTHONPATH ensures any subprocess
# (not just Pool workers) can find compiled modules.
# NOTE: Do NOT import unsloth_zoo.compiler here -- it triggers heavy torch/triton imports.
if sys.platform in ("win32", "darwin"):
_compile_cache = os.environ.get(
"UNSLOTH_COMPILE_LOCATION", "unsloth_compiled_cache"
)
if not os.path.isabs(_compile_cache):
_compile_cache = os.path.abspath(_compile_cache)
os.environ["UNSLOTH_COMPILE_LOCATION"] = _compile_cache
_pp = os.environ.get("PYTHONPATH", "")
if _compile_cache not in _pp.split(os.pathsep):
os.environ["PYTHONPATH"] = _compile_cache + (os.pathsep + _pp if _pp else "")
if _compile_cache not in sys.path:
sys.path.insert(0, _compile_cache)
import torch
from utils.hardware import (
clear_gpu_cache,
safe_num_proc,
dataset_map_num_proc,
get_device_map,
raise_if_offloaded,
get_visible_gpu_count,
)
# recompile_limit was removed in some ROCm torch builds (e.g. pytorch.org/whl/rocm6.2).
# Guard so training doesn't crash on RDNA2/RDNA3 with older ROCm torch wheels.
if hasattr(torch._dynamo.config, "recompile_limit"):
torch._dynamo.config.recompile_limit = 64
from unsloth import FastLanguageModel, FastVisionModel, is_bfloat16_supported
from unsloth.chat_templates import get_chat_template
import json
import threading
import math
import subprocess
import structlog
from loggers import get_logger
import time
from pathlib import Path
from typing import Optional, Callable
from dataclasses import dataclass
import pandas as pd
from datasets import Dataset, load_dataset
from core.inference.llama_cpp import _hf_offline_if_dns_dead
from utils.models import is_vision_model, detect_audio_type
from utils.models.model_config import _env_offline
from utils.datasets import format_and_template_dataset
from utils.datasets import MODEL_TO_TEMPLATE_MAPPER, TEMPLATE_TO_RESPONSES_MAPPER
from utils.datasets.raw_text import prepare_raw_text_dataset
from utils.paths import (
ensure_dir,
resolve_dataset_path,
resolve_output_dir,
resolve_tensorboard_dir,
)
from trl import SFTTrainer, SFTConfig
from utils.native_path_leases import child_env_without_native_path_secret
from utils.subprocess_compat import (
windows_hidden_subprocess_kwargs as _windows_hidden_subprocess_kwargs,
)
logger = get_logger(__name__)
def _build_report_targets(training_args) -> list[str] | str:
report_to: list[str] = []
if training_args.get("enable_wandb", False):
report_to.append("wandb")
if training_args.get("enable_tensorboard", False):
report_to.append("tensorboard")
return report_to or "none"
@dataclass
class TrainingProgress:
"""Training progress tracking"""
epoch: float = 0
step: int = 0
total_steps: int = 0
loss: Optional[float] = None
learning_rate: Optional[float] = None
is_training: bool = False
is_completed: bool = False
error: Optional[str] = None
status_message: str = "Ready to train" # Current stage message
elapsed_seconds: Optional[float] = None
eta_seconds: Optional[float] = None
grad_norm: Optional[float] = None
num_tokens: Optional[int] = None
eval_loss: Optional[float] = None
class UnslothTrainer:
"""
Unsloth Training Backend
"""
def __init__(self):
self.model = None
self.tokenizer = None
self.trainer = None
self.training_thread = None
self.training_progress = TrainingProgress()
self.progress_callbacks = []
self.is_training = False
self.should_stop = False
self.save_on_stop = True
self.load_in_4bit = True # Track quantization mode for metadata
# Model state tracking
self.is_cpt = False # Set to True for Continued Pretraining
self.is_vlm = False
self.is_audio = False
self.is_audio_vlm = (
False # Multimodal model (e.g. Gemma 3N) trained on audio data
)
self._audio_type = None # 'csm', 'whisper', 'snac', 'bicodec', 'dac'
self._cuda_audio_used = (
False # Set once after audio CUDA preprocessing; never cleared
)
self._spark_tts_repo_dir = (
None # Path to downloaded Spark-TTS repo (for BiCodecTokenizer)
)
self.model_name = None
# Training metrics tracking
self.training_start_time: Optional[float] = None
self.batch_size: Optional[int] = None
self.max_seq_length: Optional[int] = None
self.gradient_accumulation_steps: Optional[int] = None
# Thread safety
self._lock = threading.Lock()
# Store training context for later transfer
self.training_context = {
"base_model_name": None,
"output_dir": None,
"is_lora": True, # Default to LoRA
}
def pre_detect_and_load_tokenizer(
self,
model_name: str,
max_seq_length: int = 2048,
hf_token: Optional[str] = None,
is_dataset_image: bool = False,
is_dataset_audio: bool = False,
trust_remote_code: bool = False,
) -> None:
"""Lightweight detection and tokenizer load — no model weights, no VRAM.
Sets is_vlm, _audio_type, is_audio_vlm, model_name and loads a
lightweight tokenizer for dataset formatting. Call this before
load_and_format_dataset() when you want to process the dataset
BEFORE loading the training model (avoids VRAM contention with
the LLM-assisted detection helper).
load_model() may be called afterwards — it will re-detect and load
the full model + tokenizer, overwriting the lightweight one set here.
"""
self.model_name = model_name
self.max_seq_length = max_seq_length
self.trust_remote_code = trust_remote_code
if hf_token:
os.environ["HF_TOKEN"] = hf_token
# --- Detect audio type (reads config.json only, no VRAM) ---
self._audio_type = detect_audio_type(model_name, hf_token)
if self._audio_type == "audio_vlm":
self.is_audio = False
self.is_audio_vlm = is_dataset_audio
self._audio_type = None
else:
self.is_audio = self._audio_type is not None
self.is_audio_vlm = False
if not self.is_audio and not self.is_audio_vlm:
self._cuda_audio_used = False
# --- Detect VLM ---
vision = (
is_vision_model(model_name, hf_token = hf_token)
if not self.is_audio
else False
)
self.is_vlm = not self.is_audio_vlm and vision and is_dataset_image
logger.info(
"pre_detect: audio_type=%s, is_audio=%s, is_audio_vlm=%s, is_vlm=%s",
self._audio_type,
self.is_audio,
self.is_audio_vlm,
self.is_vlm,
)
# --- Load lightweight tokenizer/processor (CPU only, no VRAM) ---
# Whisper needs AutoProcessor (has feature_extractor + tokenizer).
# All others work with AutoTokenizer (CSM loads its own processor inline).
if self._audio_type == "whisper":
from transformers import AutoProcessor
self.tokenizer = AutoProcessor.from_pretrained(
model_name,
trust_remote_code = trust_remote_code,
token = hf_token,
)
else:
from transformers import AutoTokenizer
self.tokenizer = AutoTokenizer.from_pretrained(
model_name,
trust_remote_code = trust_remote_code,
token = hf_token,
)
logger.info("Pre-loaded tokenizer for %s", model_name)
def add_progress_callback(self, callback: Callable[[TrainingProgress], None]):
"""Add callback for training progress updates"""
self.progress_callbacks.append(callback)
def _update_progress(self, **kwargs):
"""Update training progress and notify callbacks"""
with self._lock:
for key, value in kwargs.items():
if hasattr(self.training_progress, key):
setattr(self.training_progress, key, value)
# Notify all callbacks
for callback in self.progress_callbacks:
try:
callback(self.training_progress)
except Exception as e:
logger.error(f"Error in progress callback: {e}")
def _create_progress_callback(self):
"""Create a TrainerCallback for progress tracking. Reused by all training branches."""
from transformers import TrainerCallback
trainer_ref = self
class _ProgressCallback(TrainerCallback):
def on_log(self, args, state, control, logs = None, **kwargs):
if not logs:
return
loss_value = logs.get("loss", logs.get("train_loss", None))
current_step = state.global_step
grad_norm = logs.get("grad_norm", None)
elapsed_seconds = None
if trainer_ref.training_start_time is not None:
elapsed_seconds = time.time() - trainer_ref.training_start_time
eta_seconds = None
if elapsed_seconds is not None and current_step > 0:
total_steps = trainer_ref.training_progress.total_steps
if total_steps > 0:
steps_remaining = total_steps - current_step
if steps_remaining > 0:
eta_seconds = (
elapsed_seconds / current_step
) * steps_remaining
num_tokens = getattr(state, "num_input_tokens_seen", None)
trainer_ref._update_progress(
step = current_step,
epoch = round(state.epoch, 2) if state.epoch else 0,
loss = loss_value,
learning_rate = logs.get("learning_rate", None),
elapsed_seconds = elapsed_seconds,
eta_seconds = eta_seconds,
grad_norm = grad_norm,
num_tokens = num_tokens,
eval_loss = logs.get("eval_loss", None),
status_message = "",
)
def on_epoch_end(self, args, state, control, **kwargs):
trainer_ref._update_progress(epoch = state.epoch, step = state.global_step)
def on_step_end(self, args, state, control, **kwargs):
if trainer_ref.should_stop:
logger.info(f"Stop detected at step {state.global_step}\n")
control.should_training_stop = True
return control
return _ProgressCallback()
def _calculate_total_steps(
self, num_samples, batch_size, grad_accum, num_epochs, max_steps
):
"""Calculate total training steps from dataset size and training params."""
if max_steps and max_steps > 0:
return max_steps
len_dataloader = math.ceil(num_samples / batch_size)
steps_per_epoch = max(
len_dataloader // grad_accum + int(len_dataloader % grad_accum > 0), 1
)
return steps_per_epoch * num_epochs
def _build_audio_training_args(self, training_args, output_dir, *, extra_args = None):
"""Build training args dict for audio branches.
Constructs the common config (batch size, lr, warmup, fp16/bf16, etc.)
and applies per-branch overrides via extra_args.
"""
batch_size = training_args.get("batch_size", 2)
gradient_accumulation_steps = training_args.get(
"gradient_accumulation_steps", 4
)
warmup_steps_val = training_args.get("warmup_steps", 5)
max_steps_val = training_args.get("max_steps", 0)
learning_rate = training_args.get("learning_rate", 2e-4)
weight_decay = training_args.get("weight_decay", 0.001)
lr_scheduler_type = training_args.get("lr_scheduler_type", "linear")
random_seed = training_args.get("random_seed", 3407)
optim_value = training_args.get("optim", "adamw_8bit")
config = {
"per_device_train_batch_size": batch_size,
"gradient_accumulation_steps": gradient_accumulation_steps,
"warmup_steps": warmup_steps_val if warmup_steps_val is not None else 5,
"learning_rate": learning_rate,
"fp16": not is_bfloat16_supported(),
"bf16": is_bfloat16_supported(),
"logging_steps": 1,
"optim": optim_value,
"weight_decay": weight_decay,
"lr_scheduler_type": lr_scheduler_type,
"seed": random_seed,
"output_dir": output_dir,
"report_to": _build_report_targets(training_args),
}
if training_args.get("enable_tensorboard", False):
config["logging_dir"] = str(
resolve_tensorboard_dir(training_args.get("tensorboard_dir"))
)
# max_steps vs epochs
if max_steps_val and max_steps_val > 0:
config["max_steps"] = max_steps_val
else:
config["num_train_epochs"] = training_args.get("num_epochs", 3)
# save_steps
save_steps_val = training_args.get("save_steps", 0)
if save_steps_val and save_steps_val > 0:
config["save_steps"] = save_steps_val
config["save_strategy"] = "steps"
# Apply per-branch overrides
if extra_args:
config.update(extra_args)
return config
def _finalize_training(self, output_dir, label = ""):
"""Save model after training and update progress. Used by all training branches."""
if self.should_stop and self.save_on_stop:
self.trainer._save_checkpoint(self.trainer.model, trial = None)
self.trainer.save_model()
self.tokenizer.save_pretrained(output_dir)
self._patch_adapter_config(output_dir)
msg = f"{label} training stopped" if label else "Training stopped"
logger.info(f"\n{msg}. Model saved to {output_dir}\n")
self._update_progress(
is_training = False,
status_message = f"Training stopped. Model saved to {output_dir}",
)
elif self.should_stop:
msg = f"{label} training cancelled" if label else "Training cancelled"
logger.info(f"\n{msg}.\n")
self._update_progress(
is_training = False, status_message = "Training cancelled."
)
else:
self.trainer.save_model()
self.tokenizer.save_pretrained(output_dir)
self._patch_adapter_config(output_dir)
msg = f"{label} training completed" if label else "Training completed"
logger.info(f"\n{msg}! Model saved to {output_dir}\n")
self._update_progress(
is_training = False,
is_completed = True,
status_message = f"Training completed! Model saved to {output_dir}",
)
def _cleanup_audio_artifacts(self):
"""Remove sys.path entries and sys.modules from previous audio preprocessing.
After audio training, cloned repo dirs (OuteTTS, Spark-TTS) remain on
sys.path and heavy audio modules (snac, whisper, sparktts, outetts) stay
in sys.modules. When the next training run calls dataset.map(num_proc=N),
forked child processes inherit this stale state and deadlock.
"""
import sys as _sys
# Remove cloned audio repo paths from sys.path
base_dir = os.path.dirname(os.path.abspath(__file__))
audio_paths = [
os.path.join(base_dir, "inference", "OuteTTS"), # DAC/OuteTTS
]
# Spark-TTS path is relative to the downloaded repo
if self._spark_tts_repo_dir:
spark_code_dir = os.path.join(
os.path.dirname(self._spark_tts_repo_dir), "Spark-TTS"
)
audio_paths.append(spark_code_dir)
removed_paths = []
for path in audio_paths:
if path in _sys.path:
_sys.path.remove(path)
removed_paths.append(path)
# Remove stale audio modules from sys.modules
prefixes = ("snac", "whisper", "sparktts", "outetts")
removed_modules = [key for key in _sys.modules if key.startswith(prefixes)]
for key in removed_modules:
del _sys.modules[key]
if removed_paths or removed_modules:
logger.info(
f"Cleaned up audio artifacts: {len(removed_paths)} paths, "
f"{len(removed_modules)} modules\n"
)
def _resolve_audio_columns(self, dataset, custom_format_mapping: dict = None):
"""Resolve audio, text, and speaker columns from user mapping or hardcoded fallback.
Returns:
dict with keys: audio_col, text_col, speaker_col (speaker_col may be None)
"""
cols = dataset.column_names
if custom_format_mapping:
audio_col = None
text_col = None
speaker_col = None
for col, role in custom_format_mapping.items():
if role == "audio":
audio_col = col
elif role == "text":
text_col = col
elif role == "speaker_id":
speaker_col = col
# Use mapping if both required columns exist in the dataset
if audio_col and audio_col in cols and text_col and text_col in cols:
return {
"audio_col": audio_col,
"text_col": text_col,
"speaker_col": speaker_col,
}
# Hardcoded fallback (existing behavior)
audio_col = next((c for c in cols if c.lower() in ("audio", "speech")), None)
text_col = next(
(
c
for c in cols
if c.lower() in ("text", "sentence", "transcript", "transcription")
),
None,
)
speaker_col = None
if "source" in cols:
speaker_col = "source"
elif "speaker_id" in cols:
speaker_col = "speaker_id"
return {
"audio_col": audio_col,
"text_col": text_col,
"speaker_col": speaker_col,
}
def load_model(
self,
model_name: str,
max_seq_length: int = 2048,
load_in_4bit: bool = True,
hf_token: Optional[str] = None,
is_dataset_image: bool = False,
is_dataset_audio: bool = False,
trust_remote_code: bool = False,
full_finetuning: bool = False,
gpu_ids: Optional[list[int]] = None,
) -> bool:
"""Load model for training (supports both text and vision models)"""
self.load_in_4bit = load_in_4bit # Store for training_meta.json
self.trust_remote_code = (
trust_remote_code # For AutoProcessor etc. used during training
)
try:
if self.model is not None:
del self.model
if self.tokenizer is not None:
del self.tokenizer
if self.trainer is not None:
del self.trainer
logger.info("\nClearing GPU memory before training...")
clear_gpu_cache()
# Clean up sys.path and sys.modules from previous audio preprocessing
# to prevent deadlocks when forking worker processes in dataset.map()
self._cleanup_audio_artifacts()
# Reload Unsloth-patched transformers modeling modules before clearing
# the compiled cache. unsloth_compile_transformers() sets __UNSLOTH_PATCHED__
# on each modeling module and replaces methods with exec'd code.
# clear_unsloth_compiled_cache() deletes the disk cache, but the flag
# prevents re-compilation — leaving missing cache files. Reloading
# restores original class definitions so Unsloth can re-compile cleanly.
import sys as _sys
import importlib
for _key, _mod in list(_sys.modules.items()):
if "transformers.models." in _key and ".modeling_" in _key:
if hasattr(_mod, "__UNSLOTH_PATCHED__"):
try:
importlib.reload(_mod)
except Exception:
pass # Non-critical — Unsloth will handle stale modules
# Remove stale compiled cache so the new model gets a fresh one
from utils.cache_cleanup import clear_unsloth_compiled_cache
_preserve = (
["Unsloth*Trainer.py"] if sys.platform in ("win32", "darwin") else None
)
clear_unsloth_compiled_cache(preserve_patterns = _preserve)
# Detect audio model type dynamically (config.json + tokenizer)
self._audio_type = detect_audio_type(model_name, hf_token)
# audio_vlm is detected as an audio_type now, handle it separately
if self._audio_type == "audio_vlm":
self.is_audio = False
self.is_audio_vlm = (
is_dataset_audio # Only use audio VLM path if dataset has audio
)
self._audio_type = None
else:
self.is_audio = self._audio_type is not None
self.is_audio_vlm = False
if not self.is_audio and not self.is_audio_vlm:
self._cuda_audio_used = False
# VLM: vision model with image dataset (mutually exclusive with audio paths)
vision = (
is_vision_model(model_name, hf_token = hf_token)
if not self.is_audio
else False
)
self.is_vlm = not self.is_audio_vlm and vision and is_dataset_image
self.model_name = model_name
self.max_seq_length = max_seq_length
logger.info(
f"Audio type: {self._audio_type}, is_audio: {self.is_audio}, is_audio_vlm: {self.is_audio_vlm}"
)
logger.info(
f"Dataset has images: {is_dataset_image}, audio: {is_dataset_audio}"
)
logger.info(f"Using VLM path: {self.is_vlm}")
# Reset training state for new run
self._update_progress(
is_training = True,
is_completed = False,
error = None,
step = 0,
loss = 0.0,
epoch = 0,
)
# Update UI immediately with loading message
model_display = (
model_name.split("/")[-1] if "/" in model_name else model_name
)
model_type_label = (
"audio" if self.is_audio else ("vision" if self.is_vlm else "text")
)
self._update_progress(
status_message = f"Loading {model_type_label} model... {model_display}"
)
logger.info(f"\nLoading {model_type_label} model: {model_name}")
# Set HF token if provided
if hf_token:
os.environ["HF_TOKEN"] = hf_token
# Proactive gated-model check: verify access BEFORE from_pretrained.
# Catches ALL gated/private models (text, vision, audio) globally.
# Skip when offline -- from_pretrained will use the cache.
if "/" in model_name and not _env_offline():
try:
from huggingface_hub import model_info as hf_model_info
info = hf_model_info(model_name, token = hf_token or None)
# model_info succeeds even for gated repos (metadata is public),
# but info.gated tells us if files require acceptance/token.
if info.gated and not hf_token:
friendly = (
f"Access denied for '{model_name}'. This model is gated. "
f"Please add a Hugging Face token with access and try again."
)
logger.error(
f"Model '{model_name}' is gated (gated={info.gated}) and no HF token provided"
)
self._update_progress(error = friendly, is_training = False)
return False
except Exception as gate_err:
from huggingface_hub.utils import (
GatedRepoError,
RepositoryNotFoundError,
)
if isinstance(gate_err, (GatedRepoError, RepositoryNotFoundError)):
friendly = (
f"Access denied for '{model_name}'. This model is gated or private. "
f"Please add a Hugging Face token with access and try again."
)
logger.error(f"Gated model check failed: {gate_err}")
self._update_progress(error = friendly, is_training = False)
return False
device_map = get_device_map(gpu_ids)
logger.info(
f"Using device_map='{device_map}' ({get_visible_gpu_count()} GPU(s) visible)"
)
# On hardware without native bfloat16 support (e.g. RDNA2 / gfx103x),
# passing dtype=None lets unsloth auto-detect and incorrectly choose
# bf16, triggering an LLVM error at the first bf16 kernel dispatch.
# Explicitly pass float16 as the fallback so unsloth never reaches
# that path. Modern NVIDIA (Ampere+) and RDNA3+ return True here so
# they are unaffected — dtype stays None and unsloth picks bf16 as
# before.
_auto_dtype = None if is_bfloat16_supported() else torch.float16
# Branch based on model type
if self._audio_type == "csm":
# CSM: FastModel + auto_model=CsmForConditionalGeneration + load_in_4bit=False
from unsloth import FastModel
from transformers import CsmForConditionalGeneration
self.model, self.tokenizer = FastModel.from_pretrained(
model_name = model_name,
max_seq_length = max_seq_length,
dtype = _auto_dtype,
auto_model = CsmForConditionalGeneration,
load_in_4bit = False,
device_map = device_map,
full_finetuning = full_finetuning,
token = hf_token,
trust_remote_code = trust_remote_code,
)
logger.info("Loaded CSM audio model")
elif self._audio_type == "whisper":
# Whisper: FastModel + auto_model=WhisperForConditionalGeneration + load_in_4bit=False
from unsloth import FastModel
from transformers import WhisperForConditionalGeneration
self.model, self.tokenizer = FastModel.from_pretrained(
model_name = model_name,
dtype = _auto_dtype,
load_in_4bit = False,
device_map = device_map,
full_finetuning = full_finetuning,
auto_model = WhisperForConditionalGeneration,
whisper_language = "English",
whisper_task = "transcribe",
token = hf_token,
trust_remote_code = trust_remote_code,
)
# Configure generation settings (notebook lines 100-105)
self.model.generation_config.language = "<|en|>"
self.model.generation_config.task = "transcribe"
self.model.config.suppress_tokens = []
self.model.generation_config.forced_decoder_ids = None
logger.info("Loaded Whisper audio model (FastModel)")
elif self._audio_type == "snac":
# Orpheus: language model with audio codec tokens
self.model, self.tokenizer = FastLanguageModel.from_pretrained(
model_name = model_name,
max_seq_length = max_seq_length,
dtype = _auto_dtype,
load_in_4bit = load_in_4bit,
device_map = device_map,
full_finetuning = full_finetuning,
token = hf_token,
trust_remote_code = trust_remote_code,
)
logger.info(
f"Loaded {self._audio_type} audio model (FastLanguageModel)"
)
elif self._audio_type == "bicodec":
# Spark-TTS: download full repo (contains sparktts package + BiCodec weights),
# then load only the LLM subfolder with FastModel.
# model_name may be:
# "Spark-TTS-0.5B/LLM" (local-style, from YAML mapping)
# "unsloth/Spark-TTS-0.5B" (HF repo ID)
from unsloth import FastModel
from huggingface_hub import snapshot_download
if model_name.endswith("/LLM"):
# "Spark-TTS-0.5B/LLM" → parent="Spark-TTS-0.5B"
local_dir = model_name.rsplit("/", 1)[0]
hf_repo = f"unsloth/{local_dir}"
llm_path = model_name
else:
# "unsloth/Spark-TTS-0.5B" → local_dir="Spark-TTS-0.5B"
hf_repo = model_name
local_dir = model_name.split("/")[-1]
llm_path = f"{local_dir}/LLM"
repo_path = snapshot_download(hf_repo, local_dir = local_dir)
self._spark_tts_repo_dir = os.path.abspath(
repo_path
) # Absolute path for sys.path
llm_path = os.path.join(self._spark_tts_repo_dir, "LLM")
self.model, self.tokenizer = FastModel.from_pretrained(
model_name = llm_path,
max_seq_length = max_seq_length,
dtype = torch.float32, # Spark-TTS requires float32
load_in_4bit = False,
device_map = device_map,
full_finetuning = full_finetuning,
token = hf_token,
trust_remote_code = trust_remote_code,
)
logger.info("Loaded Spark-TTS (bicodec) model")
elif self._audio_type == "dac":
# OuteTTS: uses FastModel (not FastLanguageModel) with load_in_4bit=False
from unsloth import FastModel
self.model, self.tokenizer = FastModel.from_pretrained(
model_name,
max_seq_length = max_seq_length,
load_in_4bit = False,
device_map = device_map,
full_finetuning = full_finetuning,
token = hf_token,
trust_remote_code = trust_remote_code,
)
logger.info("Loaded OuteTTS (dac) model (FastModel)")
elif self.is_audio_vlm:
# Audio VLM: multimodal model trained on audio (e.g. Gemma 3N)
# Uses FastModel (general loader) — returns (model, processor)
from unsloth import FastModel
self.model, self.tokenizer = FastModel.from_pretrained(
model_name = model_name,
max_seq_length = max_seq_length,
dtype = _auto_dtype,
load_in_4bit = load_in_4bit,
device_map = device_map,
full_finetuning = full_finetuning,
token = hf_token,
trust_remote_code = trust_remote_code,
)
logger.info("Loaded audio VLM model (FastModel)")
elif self.is_vlm:
# Load vision model - returns (model, tokenizer)
self.model, self.tokenizer = FastVisionModel.from_pretrained(
model_name = model_name,
max_seq_length = max_seq_length,
dtype = _auto_dtype,
load_in_4bit = load_in_4bit,
device_map = device_map,
full_finetuning = full_finetuning,
token = hf_token,
trust_remote_code = trust_remote_code,
)
logger.info("Loaded vision model")
# Diagnostic: check if FastVisionModel returned a real Processor or a raw tokenizer
from transformers import ProcessorMixin
tok = self.tokenizer
has_image_proc = isinstance(tok, ProcessorMixin) or hasattr(
tok, "image_processor"
)
logger.info(
f"\n[VLM Diagnostic] FastVisionModel returned: {type(tok).__name__}"
)
logger.info(
f"[VLM Diagnostic] Is ProcessorMixin: {isinstance(tok, ProcessorMixin)}"
)
logger.info(
f"[VLM Diagnostic] Has image_processor: {hasattr(tok, 'image_processor')}"
)
logger.info(
f"[VLM Diagnostic] Usable as vision processor: {has_image_proc}\n"
)
else:
# Load text model - returns (model, tokenizer)
self.model, self.tokenizer = FastLanguageModel.from_pretrained(
model_name = model_name,
max_seq_length = max_seq_length,
dtype = _auto_dtype,
load_in_4bit = load_in_4bit,
device_map = device_map,
full_finetuning = full_finetuning,
token = hf_token,
trust_remote_code = trust_remote_code,
)
logger.info("Loaded text model")
raise_if_offloaded(self.model, device_map, "Studio training")
if self.should_stop:
return False
if full_finetuning:
# Enable training mode for full fine-tuning
# This ensures all model parameters are trainable; otherwise, they might be frozen.
self.model.for_training()
self._update_progress(status_message = "Model loaded successfully")
logger.info("Model loaded successfully")
return True
except OSError as e:
if "could not get source code" in str(e) and not getattr(
self, "_source_code_retried", False
):
# Unsloth's patching can leave stale state that makes
# inspect.getsource() fail when switching model families
# (e.g. gemma3 → gemma3n). The load always succeeds on a
# second attempt because the failed first call's partial
# imports clean up the stale state as a side effect.
self._source_code_retried = True
logger.info(f"\n'could not get source code' — retrying once...\n")
return self.load_model(
model_name = model_name,
max_seq_length = max_seq_length,
load_in_4bit = load_in_4bit,
hf_token = hf_token,
is_dataset_image = is_dataset_image,
is_dataset_audio = is_dataset_audio,
trust_remote_code = trust_remote_code,
full_finetuning = full_finetuning,
gpu_ids = gpu_ids,
)
error_msg = str(e)
error_lower = error_msg.lower()
if any(
k in error_lower
for k in (
"gated repo",
"access to it at",
"401",
"403",
"unauthorized",
"forbidden",
)
):
error_msg = (
f"Access denied for '{model_name}'. This model is gated or private. "
f"Please add a Hugging Face token with access and try again."
)
logger.error(f"Error loading model: {e}")
self._update_progress(error = error_msg, is_training = False)
return False
except Exception as e:
error_msg = str(e)
# Catch gated/auth errors and surface a friendly message
error_lower = error_msg.lower()
if any(
k in error_lower
for k in (
"gated repo",
"access to it at",
"401",
"403",
"unauthorized",
"forbidden",
)
):
error_msg = (
f"Access denied for '{model_name}'. This model is gated or private. "
f"Please add a Hugging Face token with access and try again."
)
logger.error(f"Error loading model: {e}")
self._update_progress(error = error_msg, is_training = False)
return False
finally:
self._source_code_retried = False
def prepare_model_for_training(
self,
use_lora: bool = True,
# Vision-specific LoRA parameters (only used if is_vlm=True)
finetune_vision_layers: bool = True,
finetune_language_layers: bool = True,
finetune_attention_modules: bool = True,
finetune_mlp_modules: bool = True,
# Standard LoRA parameters
target_modules: list = None,
lora_r: int = 16,
lora_alpha: int = 16,
lora_dropout: float = 0.0,
use_gradient_checkpointing: str = "unsloth",
use_rslora: bool = False,
use_loftq: bool = False,
modules_to_save: list = None,
) -> bool:
"""
Prepare model for training (with optional LoRA).
"""
try:
if self.model is None:
raise ValueError("Model not loaded. Call load_model() first.")
# Full finetuning mode - skip PEFT entirely
if not use_lora:
self._update_progress(
status_message = "Full finetuning mode - no LoRA adapters"
)
logger.info("Full finetuning mode - training all parameters\n")
return True
# LoRA/QLoRA mode - apply PEFT
# "all-linear" is a PEFT keyword that targets every linear layer
if isinstance(target_modules, list) and "all-linear" in target_modules:
if len(target_modules) == 1:
target_modules = "all-linear"
else:
target_modules = [m for m in target_modules if m != "all-linear"]
elif target_modules is None or (
isinstance(target_modules, list) and len(target_modules) == 0
):
target_modules = [
"q_proj",
"k_proj",
"v_proj",
"o_proj",
"gate_proj",
"up_proj",
"down_proj",
]
# Validate and normalize gradient_checkpointing
# Must be one of: True, False, or "unsloth"
if isinstance(use_gradient_checkpointing, str):
use_gradient_checkpointing = use_gradient_checkpointing.strip().lower()
if (
use_gradient_checkpointing == ""
or use_gradient_checkpointing == "unsloth"
):
use_gradient_checkpointing = "unsloth"
elif use_gradient_checkpointing in ("true", "1", "yes"):
use_gradient_checkpointing = True
elif use_gradient_checkpointing in ("false", "0", "no"):
use_gradient_checkpointing = False
else:
# Invalid value, default to "unsloth"
logger.warning(
f"Invalid gradient_checkpointing value: {use_gradient_checkpointing}, defaulting to 'unsloth'"
)
use_gradient_checkpointing = "unsloth"
elif use_gradient_checkpointing not in (True, False, "unsloth"):
# Invalid type or value, default to "unsloth"
logger.warning(
f"Invalid gradient_checkpointing type/value: {use_gradient_checkpointing}, defaulting to 'unsloth'"
)
use_gradient_checkpointing = "unsloth"
# Verify model is loaded
if self.model is None:
error_msg = "Model is None - model was not loaded properly"
logger.error(error_msg)
self._update_progress(error = error_msg)
return False
# Check if model has the expected attributes
if not hasattr(self.model, "config"):
error_msg = "Model does not have config attribute - model may not be loaded correctly"
logger.error(error_msg)
self._update_progress(error = error_msg)
return False
logger.info(
f"Configuring LoRA adapters (r={lora_r}, alpha={lora_alpha})...\n"
)
logger.info(
f"Gradient checkpointing: {use_gradient_checkpointing} (type: {type(use_gradient_checkpointing).__name__})\n"
)
# Branch based on model type: audio, audio_vlm, vision, or text
if self._audio_type in ("csm", "bicodec", "dac") or self.is_audio_vlm:
# Models using FastModel.get_peft_model (codec audio + audio VLM)
from unsloth import FastModel
label = self._audio_type or "audio_vlm"
logger.info(f"{label} LoRA configuration:")
logger.info(f" - Target modules: {target_modules}")
if self.is_audio_vlm:
logger.info(f" - Finetune vision layers: {finetune_vision_layers}")
logger.info(
f" - Finetune language layers: {finetune_language_layers}"
)
logger.info(
f" - Finetune attention modules: {finetune_attention_modules}"
)
logger.info(f" - Finetune MLP modules: {finetune_mlp_modules}")
logger.info()
peft_kwargs = dict(
r = lora_r,
target_modules = target_modules,
lora_alpha = lora_alpha,
lora_dropout = lora_dropout,
bias = "none",
use_gradient_checkpointing = use_gradient_checkpointing,
random_state = 3407,
use_rslora = use_rslora,
loftq_config = {"loftq_bits": 4, "loftq_iter": 1}
if use_loftq
else None,
)
# Audio VLM models support VLM-style layer selection
if self.is_audio_vlm:
peft_kwargs.update(
finetune_vision_layers = finetune_vision_layers,
finetune_language_layers = finetune_language_layers,
finetune_attention_modules = finetune_attention_modules,
finetune_mlp_modules = finetune_mlp_modules,
)
self.model = FastModel.get_peft_model(self.model, **peft_kwargs)
elif self._audio_type == "whisper":
# Phase 2: Whisper uses FastModel.get_peft_model with task_type=None
from unsloth import FastModel
logger.info(f"Audio model (whisper) LoRA configuration:")
logger.info(f" - Target modules: {target_modules}\n")
self.model = FastModel.get_peft_model(
self.model,
r = lora_r,
target_modules = target_modules,
lora_alpha = lora_alpha,
lora_dropout = lora_dropout,
bias = "none",
use_gradient_checkpointing = use_gradient_checkpointing,
random_state = 3407,
use_rslora = use_rslora,
loftq_config = {"loftq_bits": 4, "loftq_iter": 1}
if use_loftq
else None,
task_type = None,
)
elif self._audio_type == "snac":
# Orpheus uses FastLanguageModel.get_peft_model
logger.info(f"Audio model ({self._audio_type}) LoRA configuration:")
logger.info(f" - Target modules: {target_modules}\n")
self.model = FastLanguageModel.get_peft_model(
self.model,
r = lora_r,
target_modules = target_modules,
lora_alpha = lora_alpha,
lora_dropout = lora_dropout,
bias = "none",
use_gradient_checkpointing = use_gradient_checkpointing,
random_state = 3407,
use_rslora = use_rslora,
loftq_config = {"loftq_bits": 4, "loftq_iter": 1}
if use_loftq
else None,
)
elif self.is_vlm:
# Vision model LoRA
logger.info(f"Vision model LoRA configuration:")
logger.info(f" - Finetune vision layers: {finetune_vision_layers}")
logger.info(f" - Finetune language layers: {finetune_language_layers}")
logger.info(
f" - Finetune attention modules: {finetune_attention_modules}"
)
logger.info(f" - Finetune MLP modules: {finetune_mlp_modules}\n")
self.model = FastVisionModel.get_peft_model(
self.model,
finetune_vision_layers = finetune_vision_layers,
finetune_language_layers = finetune_language_layers,
finetune_attention_modules = finetune_attention_modules,
finetune_mlp_modules = finetune_mlp_modules,
r = lora_r,
target_modules = target_modules,
lora_alpha = lora_alpha,
lora_dropout = lora_dropout,
bias = "none",
use_gradient_checkpointing = use_gradient_checkpointing,
random_state = 3407,
use_rslora = use_rslora,
loftq_config = {"loftq_bits": 4, "loftq_iter": 1}
if use_loftq
else None,
modules_to_save = modules_to_save,
)
else:
# Text model LoRA
logger.info(f"Text model LoRA configuration:")
logger.info(f" - Target modules: {target_modules}\n")
if modules_to_save:
logger.info(f" - Modules to save: {modules_to_save}\n")
self.model = FastLanguageModel.get_peft_model(
self.model,
r = lora_r,
target_modules = target_modules,
lora_alpha = lora_alpha,
lora_dropout = lora_dropout,
bias = "none",
use_gradient_checkpointing = use_gradient_checkpointing,
random_state = 3407,
use_rslora = use_rslora,
loftq_config = {"loftq_bits": 4, "loftq_iter": 1}
if use_loftq
else None,
modules_to_save = modules_to_save,
)
# Check if stopped during LoRA preparation
if self.should_stop:
logger.info("Stopped during LoRA configuration\n")
return False
self._update_progress(status_message = "LoRA adapters configured")
logger.info("LoRA adapters configured successfully\n")
return True
except Exception as e:
import traceback
import sys
error_details = (
f"{type(e).__name__}: {str(e)}"
if str(e)
else f"{type(e).__name__} (no message)"
)
full_traceback = traceback.format_exc()
logger.error(f"Error preparing model: {error_details}")
logger.error(f"Full traceback:\n{full_traceback}")
logger.info(f"\n[ERROR] Error preparing model: {error_details}")
logger.info(f"[ERROR] Full traceback:\n{full_traceback}")
self._update_progress(error = error_details)
return False
def _apply_csm_forward_fix(self):
"""Monkey-patch CsmForConditionalGeneration.forward to fix depth decoder kwargs.
The original transformers forward passes raw **kwargs (num_items_in_batch,
causal_mask, etc.) from the Trainer/PEFT through to the depth decoder,
causing depth_decoder_loss=None and 'Tensor + NoneType' crash.
We patch at both instance AND class level for maximum reliability,
and strip non-TransformersKwargs params that Unsloth/PEFT inject.
"""
import types
import torch
import torch.nn as nn
from transformers.models.csm.modeling_csm import (
CsmForConditionalGeneration,
CsmOutputWithPast,
)
base_csm = self.model.base_model.model # CsmForConditionalGeneration
# Save original forward (the @can_return_tuple wrapped version)
_original_forward = CsmForConditionalGeneration.forward
# Keys that the depth decoder and its sub-layers actually understand
_TRANSFORMERS_KWARGS = {
"num_items_in_batch",
"output_hidden_states",
"output_attentions",
"output_router_logits",
"cu_seq_lens_q",
"cu_seq_lens_k",
"max_length_q",
"max_length_k",
}
def _fixed_csm_forward(
self,
input_ids = None,
input_values = None,
attention_mask = None,
input_values_cutoffs = None,
position_ids = None,
past_key_values = None,
inputs_embeds = None,
labels = None,
use_cache = None,
cache_position = None,
logits_to_keep = 0,
**kwargs,
):
# Strip non-standard kwargs injected by Unsloth/PEFT (causal_mask,
# num_logits_to_keep, task_ids, return_dict, etc.)
output_attentions = kwargs.pop("output_attentions", None)
output_hidden_states = kwargs.pop("output_hidden_states", None)
kwargs.pop("return_dict", None)
kwargs.pop("causal_mask", None)
kwargs.pop("num_logits_to_keep", None)
kwargs.pop("task_ids", None)
# Only keep recognized TransformersKwargs
clean_kwargs = {
k: v for k, v in kwargs.items() if k in _TRANSFORMERS_KWARGS
}
if input_ids is not None and input_ids.ndim == 2:
merged = self._merge_input_ids_with_input_values(
input_ids, input_values, input_values_cutoffs, labels
)
inputs_embeds = merged["inputs_embeds"]
labels = merged["labels"]
input_ids = None
backbone_outputs = self.backbone_model(
input_ids = input_ids,
attention_mask = attention_mask,
position_ids = position_ids,
past_key_values = past_key_values,
inputs_embeds = inputs_embeds,
use_cache = use_cache,
cache_position = cache_position,
output_attentions = output_attentions,
output_hidden_states = output_hidden_states,
**clean_kwargs,
)
backbone_hidden_states = backbone_outputs[0]
slice_indices = (
slice(-logits_to_keep, None)
if isinstance(logits_to_keep, int)
else logits_to_keep
)
backbone_logits = self.lm_head(backbone_hidden_states[:, slice_indices, :])
loss = None
backbone_loss = None
depth_decoder_loss = None
depth_decoder_outputs = None
if labels is not None:
backbone_labels = labels[:, :, 0]
backbone_loss = self.loss_function(
logits = backbone_logits,
labels = backbone_labels,
vocab_size = self.config.vocab_size,
**clean_kwargs,
)
train_mask = ~(labels[:, :, 1:] == -100).all(dim = -1)
depth_decoder_input_ids = labels[train_mask][
..., : self.config.num_codebooks - 1
]
depth_decoder_input_ids = nn.functional.pad(
depth_decoder_input_ids, (1, 0), value = 0
)
train_idxs = train_mask.nonzero(as_tuple = True)
backbone_last_hidden_states = backbone_hidden_states[
train_idxs[0], train_idxs[1] - 1, :
]
depth_decoder_labels = labels[train_mask]
# Build clean kwargs for depth decoder
dd_kwargs = clean_kwargs.copy()
# Scale num_items_in_batch for depth decoder (31 codebooks)
if "num_items_in_batch" in dd_kwargs:
dd_kwargs["num_items_in_batch"] = dd_kwargs[
"num_items_in_batch"
] * (self.config.num_codebooks - 1)
depth_decoder_outputs = self.depth_decoder(
input_ids = depth_decoder_input_ids,
backbone_last_hidden_state = backbone_last_hidden_states,
use_cache = False,
return_dict = True,
labels = depth_decoder_labels,
output_attentions = output_attentions,
output_hidden_states = output_hidden_states,
**dd_kwargs,
)
depth_decoder_loss = depth_decoder_outputs.loss
if depth_decoder_loss is None:
logger.warning(
"CSM depth_decoder_loss is None! "
f"labels shape={depth_decoder_labels.shape}, "
f"train_mask sum={train_mask.sum().item()}"
)
# Fallback: use only backbone loss instead of crashing
loss = backbone_loss
else:
loss = backbone_loss + depth_decoder_loss
return CsmOutputWithPast(
loss = loss,
backbone_loss = backbone_loss,
depth_decoder_loss = depth_decoder_loss,
logits = backbone_logits,
past_key_values = backbone_outputs.past_key_values,
hidden_states = backbone_outputs.hidden_states,
attentions = backbone_outputs.attentions,
depth_decoder_logits = (
depth_decoder_outputs.logits if depth_decoder_outputs else None
),
depth_decoder_past_key_values = (
depth_decoder_outputs.past_key_values
if depth_decoder_outputs
else None
),
depth_decoder_hidden_states = (
depth_decoder_outputs.hidden_states
if depth_decoder_outputs
else None
),
depth_decoder_attentions = (
depth_decoder_outputs.attentions if depth_decoder_outputs else None
),
)
# Patch at BOTH instance and class level for maximum reliability.
# Instance-level: catches calls via BaseTuner.forward -> self.model.forward()
base_csm.forward = types.MethodType(_fixed_csm_forward, base_csm)
# Class-level: catches any path that resolves through the class dict
CsmForConditionalGeneration.forward = _fixed_csm_forward
logger.info("Applied CSM forward fix (class + instance level)\n")
def _preprocess_csm_dataset(self, dataset, custom_format_mapping = None):
"""Preprocess dataset for CSM TTS training (exact notebook copy)."""
from transformers import AutoProcessor
from datasets import Audio
import torch
processor = AutoProcessor.from_pretrained(
self.model_name,
trust_remote_code = getattr(self, "trust_remote_code", False),
)
# Strip pad_to_multiple_of from tokenizer init_kwargs — fine-tuned models
# (e.g. keanteng/sesame-csm-elise) save it in tokenizer_config.json, and
# _merge_kwargs leaks it into audio_kwargs where EncodecFeatureExtractor rejects it.
processor.tokenizer.init_kwargs.pop("pad_to_multiple_of", None)
# Resolve columns from user mapping or hardcoded fallback
resolved = self._resolve_audio_columns(dataset, custom_format_mapping)
audio_col = resolved["audio_col"]
text_col = resolved["text_col"]
speaker_key = resolved["speaker_col"]
if audio_col is None:
raise ValueError(
f"No audio column found in dataset. Columns: {dataset.column_names}"
)
if text_col is None:
raise ValueError(
f"No text column found in dataset. Columns: {dataset.column_names}"
)
if speaker_key is None:
logger.info(
"No speaker found, adding default 'source' of 0 for all examples\n"
)
dataset = dataset.add_column("source", ["0"] * len(dataset))
speaker_key = "source"
logger.info(
f"CSM preprocessing: audio_col='{audio_col}', text_col='{text_col}', speaker_key='{speaker_key}'\n"
)
dataset = dataset.cast_column(audio_col, Audio(sampling_rate = 24000))
required_keys = [
"input_ids",
"attention_mask",
"labels",
"input_values",
"input_values_cutoffs",
]
self._update_progress(status_message = "Preprocessing CSM dataset...")
processed_examples = []
skipped = 0
for idx in range(len(dataset)):
if self.should_stop:
logger.info("Stopped during CSM preprocessing\n")
break
example = dataset[idx]
try:
conversation = [
{
"role": str(example[speaker_key]),
"content": [
{"type": "text", "text": example.get(text_col, "")},
{"type": "audio", "path": example[audio_col]["array"]},
],
}
]
# NOTE: pad_to_multiple_of intentionally omitted from text_kwargs —
# CsmProcessor._merge_kwargs leaks it to EncodecFeatureExtractor which rejects it.
model_inputs = processor.apply_chat_template(
conversation,
tokenize = True,
return_dict = True,
output_labels = True,
text_kwargs = {
"padding": "max_length",
"max_length": 256,
"padding_side": "right",
},
audio_kwargs = {
"sampling_rate": 24_000,
"max_length": 240001,
"padding": "max_length",
},
common_kwargs = {"return_tensors": "pt"},
)
out = {}
for k in required_keys:
if k not in model_inputs:
raise KeyError(f"Missing required key '{k}' in model outputs")
out[k] = model_inputs[k][0]
if not all(isinstance(out[k], torch.Tensor) for k in out):
skipped += 1
continue
processed_examples.append(out)
except Exception as e:
logger.warning(f"Error processing CSM example {idx}: {e}")
skipped += 1
continue
if (idx + 1) % 100 == 0:
self._update_progress(
status_message = f"Preprocessing CSM... {idx + 1}/{len(dataset)}"
)
if not processed_examples:
raise ValueError(
f"No valid examples after CSM preprocessing (skipped {skipped})"
)
result_dataset = Dataset.from_list(processed_examples)
logger.info(
f"CSM preprocessing complete: {len(result_dataset)} examples "
f"({skipped} skipped)\n"
)
return result_dataset
def _format_audio_vlm_dataset(self, dataset, custom_format_mapping = None):
"""Format dataset as audio chat messages for multimodal models (e.g. Gemma 3N).
Expects columns: audio (Audio), text (str).
Produces: messages column with system/user/assistant chat format.
"""
from datasets import Audio
resolved = self._resolve_audio_columns(dataset, custom_format_mapping)
audio_col = resolved["audio_col"]
text_col = resolved["text_col"]
if not audio_col or not text_col:
raise ValueError(
f"Audio VLM dataset needs 'audio' and 'text' columns, got: {dataset.column_names}"
)
# Store resolved audio column name for the collator closure
self._audio_vlm_audio_col = audio_col
# Cast audio to 16kHz (standard for speech models)
dataset = dataset.cast_column(audio_col, Audio(sampling_rate = 16000))
def format_messages(samples):
formatted = {"messages": []}
for idx in range(len(samples[audio_col])):
audio = samples[audio_col][idx]["array"]
label = str(samples[text_col][idx])
message = [
{
"role": "system",
"content": [
{
"type": "text",
"text": "You are an assistant that transcribes speech accurately.",
}
],
},
{
"role": "user",
"content": [
{"type": "audio", "audio": audio},
{"type": "text", "text": "Please transcribe this audio."},
],
},
{"role": "assistant", "content": [{"type": "text", "text": label}]},
]
formatted["messages"].append(message)
return formatted
self._update_progress(status_message = "Formatting audio VLM dataset...")
dataset = dataset.map(
format_messages,
batched = True,
batch_size = 4,
num_proc = dataset_map_num_proc(4),
)
logger.info(f"Audio VLM dataset formatted: {len(dataset)} examples\n")
return dataset
def _preprocess_snac_dataset(self, dataset, custom_format_mapping = None):
"""Preprocess dataset for Orpheus TTS training with SNAC codec.
Mirrors Orpheus_(3B)-TTS.ipynb: encode audio with SNAC (24kHz, 3 hierarchical
layers), interleave 7 codes per frame, wrap with Orpheus special tokens,
train on full sequence (no label masking).
"""
import torch
import torchaudio.transforms as T
SNAC_MODEL_NAME = "hubertsiuzdak/snac_24khz"
SNAC_SAMPLE_RATE = 24000
device = "cuda" if torch.cuda.is_available() else "cpu"
max_length = self.max_seq_length or 2048
tokenizer = self.tokenizer
# Orpheus special token IDs (hardcoded in tokenizer vocabulary)
START_OF_HUMAN = 128259
END_OF_HUMAN = 128260
START_OF_AI = 128261
END_OF_AI = 128262
START_OF_SPEECH = 128257
END_OF_SPEECH = 128258
END_OF_TEXT = 128009
AUDIO_OFFSET = 128266
resolved = self._resolve_audio_columns(dataset, custom_format_mapping)
audio_col = resolved["audio_col"]
text_col = resolved["text_col"]
speaker_col = resolved["speaker_col"]
has_source = speaker_col is not None
if not audio_col or not text_col:
raise ValueError(
f"SNAC dataset needs 'audio' and 'text' columns, got: {dataset.column_names}"
)
# Cast audio column so datasets 4.x AudioDecoder objects are decoded to dicts
from datasets import Audio
dataset = dataset.cast_column(audio_col, Audio(sampling_rate = SNAC_SAMPLE_RATE))
# Get dataset sample rate from first example (after cast, always SNAC_SAMPLE_RATE)
first_audio = dataset[0][audio_col]
ds_sample_rate = (
first_audio.get("sampling_rate", SNAC_SAMPLE_RATE)
if isinstance(first_audio, dict)
else SNAC_SAMPLE_RATE
)
# Load SNAC codec model
self._update_progress(status_message = "Loading SNAC codec model...")
logger.info("Loading SNAC codec model...\n")
from snac import SNAC
snac_model = SNAC.from_pretrained(SNAC_MODEL_NAME)
snac_model = snac_model.to(device).eval()
# Resample transform (created once)
resample_transform = (
T.Resample(orig_freq = ds_sample_rate, new_freq = SNAC_SAMPLE_RATE)
if ds_sample_rate != SNAC_SAMPLE_RATE
else None
)
self._update_progress(status_message = "Encoding audio with SNAC...")
logger.info(
f"SNAC preprocessing: audio_col='{audio_col}', text_col='{text_col}', "
f"has_source={has_source}, ds_sample_rate={ds_sample_rate}\n"
)
processed_examples = []
skipped = 0
for idx in range(len(dataset)):
if self.should_stop:
logger.info("Stopped during SNAC preprocessing\n")
break
example = dataset[idx]
try:
text = example.get(text_col)
if not text:
skipped += 1
continue
audio_data = example.get(audio_col)
if audio_data is None or audio_data.get("array") is None:
skipped += 1
continue
# --- Encode audio with SNAC (notebook lines 122-142) ---
waveform = (
torch.from_numpy(audio_data["array"])
.unsqueeze(0)
.to(dtype = torch.float32)
)
if resample_transform is not None:
waveform = resample_transform(waveform)
waveform = waveform.unsqueeze(0).to(device)
with torch.inference_mode():
codes = snac_model.encode(waveform)
# Interleave 7 codes per frame with layer offsets (notebook lines 134-142)
all_codes = []
for i in range(codes[0].shape[1]):
all_codes.append(codes[0][0][i].item() + AUDIO_OFFSET)
all_codes.append(codes[1][0][2 * i].item() + AUDIO_OFFSET + 4096)
all_codes.append(
codes[2][0][4 * i].item() + AUDIO_OFFSET + (2 * 4096)
)
all_codes.append(
codes[2][0][(4 * i) + 1].item() + AUDIO_OFFSET + (3 * 4096)
)
all_codes.append(
codes[1][0][(2 * i) + 1].item() + AUDIO_OFFSET + (4 * 4096)
)
all_codes.append(
codes[2][0][(4 * i) + 2].item() + AUDIO_OFFSET + (5 * 4096)
)
all_codes.append(
codes[2][0][(4 * i) + 3].item() + AUDIO_OFFSET + (6 * 4096)
)
if len(all_codes) == 0:
skipped += 1
continue
# Deduplicate consecutive frames with same first code (notebook lines 185-207)
deduped = all_codes[:7]
for i in range(7, len(all_codes), 7):
if all_codes[i] != deduped[-7]:
deduped.extend(all_codes[i : i + 7])
all_codes = deduped
# --- Build text tokens (notebook lines 217-224) ---
text_prompt = (
f"{example[speaker_col]}: {text}"
if has_source and example.get(speaker_col)
else text
)
text_ids = tokenizer.encode(text_prompt, add_special_tokens = True)
text_ids.append(END_OF_TEXT)
# --- Build full input_ids (notebook lines 225-234) ---
input_ids = (
[START_OF_HUMAN]
+ text_ids
+ [END_OF_HUMAN]
+ [START_OF_AI]
+ [START_OF_SPEECH]
+ all_codes
+ [END_OF_SPEECH]
+ [END_OF_AI]
)
# Truncate to max_length
input_ids = input_ids[:max_length]
# Labels = input_ids (no masking — Orpheus trains on full sequence)
labels = list(input_ids)
attention_mask = [1] * len(input_ids)
processed_examples.append(
{
"input_ids": input_ids,
"labels": labels,
"attention_mask": attention_mask,
}
)
except Exception as e:
logger.warning(f"Error processing SNAC example {idx}: {e}")
skipped += 1
continue
# Progress update every 100 examples
if (idx + 1) % 100 == 0:
self._update_progress(
status_message = f"Encoding audio... {idx + 1}/{len(dataset)}"
)
# Free SNAC model from GPU
logger.info("Freeing SNAC codec model from GPU...\n")
snac_model.to("cpu")
del snac_model
import gc
gc.collect()
torch.cuda.empty_cache()
self._cuda_audio_used = True
if not processed_examples:
raise ValueError(
f"No valid examples after SNAC preprocessing (skipped {skipped})"
)
result_dataset = Dataset.from_list(processed_examples)
logger.info(
f"SNAC preprocessing complete: {len(result_dataset)} examples "
f"({skipped} skipped)\n"
)
return result_dataset
def _preprocess_bicodec_dataset(self, dataset, custom_format_mapping = None):
"""Preprocess dataset for Spark-TTS training with BiCodec tokenizer.
Mirrors Spark_TTS_(0_5B).ipynb: encode audio with BiCodec (semantic + global tokens),
format as special-token text strings for SFTTrainer with dataset_text_field="text".
"""
import sys
import torch
import numpy as np
import torchaudio.transforms as T
import subprocess
device = "cuda" if torch.cuda.is_available() else "cpu"
# The sparktts Python package lives in the SparkAudio/Spark-TTS GitHub repo,
# NOT in the unsloth/Spark-TTS-0.5B HF model repo. Clone it if needed.
spark_code_dir = os.path.join(
os.path.dirname(self._spark_tts_repo_dir), "Spark-TTS"
)
sparktts_pkg = os.path.join(spark_code_dir, "sparktts")
if not os.path.isdir(sparktts_pkg):
self._update_progress(status_message = "Cloning Spark-TTS code repo...")
logger.info(f"Cloning SparkAudio/Spark-TTS to {spark_code_dir}...\n")
subprocess.run(
[
"git",
"clone",
"--depth",
"1",
"https://github.com/SparkAudio/Spark-TTS",
spark_code_dir,
],
check = True,
env = child_env_without_native_path_secret(),
**_windows_hidden_subprocess_kwargs(),
)
if spark_code_dir not in sys.path:
sys.path.insert(0, spark_code_dir)
from sparktts.models.audio_tokenizer import BiCodecTokenizer
from sparktts.utils.audio import audio_volume_normalize
# Resolve audio and text columns
resolved = self._resolve_audio_columns(dataset, custom_format_mapping)
audio_col = resolved["audio_col"]
text_col = resolved["text_col"]
speaker_col = resolved["speaker_col"]
has_source = speaker_col is not None
if not audio_col or not text_col:
raise ValueError(
f"BiCodec dataset needs 'audio' and 'text' columns, got: {dataset.column_names}"
)
# Cast audio column so datasets 4.x AudioDecoder objects are decoded to dicts.
# Don't resample here — BiCodec's target_sr may differ; the loop handles resampling.
from datasets import Audio
dataset = dataset.cast_column(audio_col, Audio())
# Load BiCodec tokenizer
self._update_progress(status_message = "Loading BiCodec tokenizer...")
logger.info("Loading BiCodec tokenizer...\n")
audio_tokenizer = BiCodecTokenizer(self._spark_tts_repo_dir, device)
target_sr = audio_tokenizer.config["sample_rate"]
self._update_progress(status_message = "Encoding audio with BiCodec...")
logger.info(
f"BiCodec preprocessing: audio_col='{audio_col}', text_col='{text_col}', "
f"has_source={has_source}, target_sr={target_sr}\n"
)
def extract_wav2vec2_features(wavs: torch.Tensor) -> torch.Tensor:
"""Extract wav2vec2 features (average of layers 11, 14, 16)."""
if wavs.shape[0] != 1:
raise ValueError(f"Expected batch size 1, but got shape {wavs.shape}")
wav_np = wavs.squeeze(0).cpu().numpy()
processed = audio_tokenizer.processor(
wav_np,
sampling_rate = 16000,
return_tensors = "pt",
padding = True,
)
input_values = processed.input_values.to(
audio_tokenizer.feature_extractor.device
)
model_output = audio_tokenizer.feature_extractor(input_values)
if model_output.hidden_states is None:
raise ValueError("Wav2Vec2Model did not return hidden states.")
feats_mix = (
model_output.hidden_states[11]
+ model_output.hidden_states[14]
+ model_output.hidden_states[16]
) / 3
return feats_mix
processed_examples = []
skipped = 0
for idx in range(len(dataset)):
if self.should_stop:
logger.info("Stopped during BiCodec preprocessing\n")
break
example = dataset[idx]
try:
text = example.get(text_col)
if not text:
skipped += 1
continue
audio_data = example.get(audio_col)
if audio_data is None or audio_data.get("array") is None:
skipped += 1
continue
audio_array = audio_data["array"]
sampling_rate = audio_data.get("sampling_rate", target_sr)
# Resample if needed
if sampling_rate != target_sr:
resampler = T.Resample(orig_freq = sampling_rate, new_freq = target_sr)
audio_tensor_temp = torch.from_numpy(audio_array).float()
audio_array = resampler(audio_tensor_temp).numpy()
# Volume normalize if configured
if audio_tokenizer.config.get("volume_normalize", False):
audio_array = audio_volume_normalize(audio_array)
# Get reference clip
ref_wav_np = audio_tokenizer.get_ref_clip(audio_array)
# Prepare tensors
audio_tensor = (
torch.from_numpy(audio_array).unsqueeze(0).float().to(device)
)
ref_wav_tensor = (
torch.from_numpy(ref_wav_np).unsqueeze(0).float().to(device)
)
# Extract wav2vec2 features
feat = extract_wav2vec2_features(audio_tensor)
batch = {
"wav": audio_tensor,
"ref_wav": ref_wav_tensor,
"feat": feat.to(device),
}
# BiCodec tokenize
semantic_token_ids, global_token_ids = audio_tokenizer.model.tokenize(
batch
)
global_tokens = "".join(
[
f"<|bicodec_global_{i}|>"
for i in global_token_ids.squeeze().cpu().numpy()
]
)
semantic_tokens = "".join(
[
f"<|bicodec_semantic_{i}|>"
for i in semantic_token_ids.squeeze().cpu().numpy()
]
)
# Format text with source prefix if available
text_content = (
f"{example[speaker_col]}: {text}"
if has_source and example.get(speaker_col)
else text
)
formatted = "".join(
[
"<|task_tts|>",
"<|start_content|>",
text_content,
"<|end_content|>",
"<|start_global_token|>",
global_tokens,
"<|end_global_token|>",
"<|start_semantic_token|>",
semantic_tokens,
"<|end_semantic_token|>",
"<|im_end|>",
]
)
processed_examples.append({"text": formatted})
except Exception as e:
logger.warning(f"Error processing BiCodec example {idx}: {e}")
skipped += 1
continue
# Progress update every 100 examples
if (idx + 1) % 100 == 0:
self._update_progress(
status_message = f"Encoding audio with BiCodec... {idx + 1}/{len(dataset)}"
)
# Free BiCodec model from GPU
logger.info("Freeing BiCodec tokenizer from GPU...\n")
audio_tokenizer.model.cpu()
audio_tokenizer.feature_extractor.cpu()
del audio_tokenizer
import gc
gc.collect()
torch.cuda.empty_cache()
self._cuda_audio_used = True
if not processed_examples:
raise ValueError(
f"No valid examples after BiCodec preprocessing (skipped {skipped})"
)
result_dataset = Dataset.from_list(processed_examples)
logger.info(
f"BiCodec preprocessing complete: {len(result_dataset)} examples "
f"({skipped} skipped)\n"
)
# Debug: show first example text (truncated)
sample = result_dataset[0]["text"]
logger.info(f"Sample text (first 200 chars): {sample[:200]}...\n")
logger.info(f"Sample text length: {len(sample)} chars\n")
return result_dataset
def _preprocess_dac_dataset(self, dataset, custom_format_mapping = None):
"""Preprocess dataset for OuteTTS training with DAC codec.
Mirrors Oute_TTS_(1B).ipynb DataCreationV3: uses Whisper for word timings,
OuteTTS AudioProcessor for speaker representations, PromptProcessor for
training prompts. Outputs text strings for SFTTrainer with dataset_text_field="text".
"""
import sys
import io
import tempfile
import torch
import numpy as np
import soundfile as sf
from datasets import Dataset as HFDataset
from utils.paths import ensure_dir, tmp_root
device = "cuda" if torch.cuda.is_available() else "cpu"
# Clone OuteTTS repo (same as audio_codecs._load_dac)
base_dir = os.path.dirname(os.path.abspath(__file__))
outetts_code_dir = os.path.join(base_dir, "inference", "OuteTTS")
outetts_pkg = os.path.join(outetts_code_dir, "outetts")
if not os.path.isdir(outetts_pkg):
self._update_progress(status_message = "Cloning OuteTTS code repo...")
logger.info(f"Cloning edwko/OuteTTS to {outetts_code_dir}...\n")
subprocess.run(
[
"git",
"clone",
"--depth",
"1",
"https://github.com/edwko/OuteTTS",
outetts_code_dir,
],
check = True,
env = child_env_without_native_path_secret(),
**_windows_hidden_subprocess_kwargs(),
)
for fpath in [
os.path.join(outetts_pkg, "models", "gguf_model.py"),
os.path.join(outetts_pkg, "interface.py"),
os.path.join(outetts_pkg, "__init__.py"),
]:
if os.path.exists(fpath):
os.remove(fpath)
logger.info(f"Removed {fpath}\n")
if outetts_code_dir not in sys.path:
sys.path.insert(0, outetts_code_dir)
from outetts.version.v3.audio_processor import AudioProcessor
from outetts.version.v3.prompt_processor import PromptProcessor
from outetts.models.config import ModelConfig as OuteTTSModelConfig
from outetts.utils.preprocessing import text_normalizations
# Resolve audio and text columns
resolved = self._resolve_audio_columns(dataset, custom_format_mapping)
audio_col = resolved["audio_col"]
text_col = resolved["text_col"]
if not audio_col or not text_col:
raise ValueError(
f"DAC dataset needs 'audio' and 'text' columns, got: {dataset.column_names}"
)
# Cast audio to 24kHz (notebook: dataset.cast_column("audio", Audio(sampling_rate=24000)))
from datasets import Audio
dataset = dataset.cast_column(audio_col, Audio(sampling_rate = 24000))
logger.info("Cast audio column to 24kHz\n")
# Load Whisper for word timings
self._update_progress(
status_message = "Loading Whisper model for word timings..."
)
logger.info("Loading Whisper model for word timings...\n")
import whisper
whisper_model = whisper.load_model("turbo", device = device)
# Load OuteTTS AudioProcessor + PromptProcessor
self._update_progress(status_message = "Loading OuteTTS AudioProcessor...")
logger.info("Loading OuteTTS AudioProcessor...\n")
model_tokenizer_path = "OuteAI/Llama-OuteTTS-1.0-1B"
dummy_config = OuteTTSModelConfig(
tokenizer_path = model_tokenizer_path,
device = device,
audio_codec_path = None,
)
audio_processor = AudioProcessor(config = dummy_config)
prompt_processor = PromptProcessor(model_tokenizer_path)
self._update_progress(status_message = "Preprocessing audio with OuteTTS...")
logger.info(
f"DAC preprocessing: audio_col='{audio_col}', text_col='{text_col}'\n"
)
processed_examples = []
skipped = 0
for idx in range(len(dataset)):
if self.should_stop:
logger.info("Stopped during DAC preprocessing\n")
break
example = dataset[idx]
try:
text = example.get(text_col)
if not text or not isinstance(text, str):
skipped += 1
continue
audio_data = example.get(audio_col)
if audio_data is None or audio_data.get("array") is None:
skipped += 1
continue
audio_array = np.array(audio_data["array"], dtype = np.float32)
sampling_rate = audio_data.get("sampling_rate", 24000)
# Convert to WAV bytes (Whisper needs a file path)
buf = io.BytesIO()
sf.write(buf, audio_array, sampling_rate, format = "WAV", subtype = "FLOAT")
buf.seek(0)
audio_bytes = buf.getvalue()
# 1. Get word timings from Whisper
with tempfile.NamedTemporaryFile(
suffix = ".wav",
delete = False,
dir = str(ensure_dir(tmp_root())),
) as tmp:
tmp.write(audio_bytes)
tmp.flush()
tmp_path = tmp.name
try:
whisper_result = whisper_model.transcribe(
tmp_path, word_timestamps = True
)
finally:
Path(tmp_path).unlink(missing_ok = True)
normalized_transcript = text_normalizations(text)
words_with_timings = []
if whisper_result and "segments" in whisper_result:
for segment in whisper_result["segments"]:
for word_info in segment.get("words", []):
cleaned = word_info["word"].strip()
if cleaned:
words_with_timings.append(
{
"word": cleaned,
"start": float(word_info["start"]),
"end": float(word_info["end"]),
}
)
if not words_with_timings:
skipped += 1
continue
# 2. Create speaker representation with AudioProcessor
speaker_data_dict = {
"audio": {"bytes": audio_bytes},
"text": normalized_transcript,
"words": words_with_timings,
}
speaker = audio_processor.create_speaker_from_dict(speaker_data_dict)
if speaker is None:
skipped += 1
continue
# 3. Get training prompt from PromptProcessor
prompt = prompt_processor.get_training_prompt(speaker)
if prompt:
processed_examples.append({"text": prompt})
except Exception as e:
logger.warning(f"Error processing DAC example {idx}: {e}")
skipped += 1
continue
if (idx + 1) % 100 == 0:
self._update_progress(
status_message = f"Preprocessing audio with OuteTTS... {idx + 1}/{len(dataset)}"
)
# Free Whisper from GPU (notebook: data_processor.whisper_model.to('cpu'))
logger.info("Moving Whisper model to CPU...\n")
whisper_model.to("cpu")
del whisper_model
del audio_processor
del prompt_processor
import gc
gc.collect()
torch.cuda.empty_cache()
self._cuda_audio_used = True
if not processed_examples:
raise ValueError(
f"No valid examples after DAC preprocessing (skipped {skipped})"
)
result_dataset = HFDataset.from_list(processed_examples)
logger.info(
f"DAC preprocessing complete: {len(result_dataset)} examples "
f"({skipped} skipped)\n"
)
sample = result_dataset[0]["text"]
logger.info(f"Sample text (first 200 chars): {sample[:200]}...\n")
return result_dataset
def _preprocess_whisper_dataset(
self, dataset, eval_split = None, custom_format_mapping = None
):
"""Preprocess dataset for Whisper speech-to-text training.
Mirrors Whisper.ipynb: extract audio features with Whisper's feature
extractor, tokenize text labels. Returns (train_data, eval_data) where
each is a list of dicts with 'input_features' and 'labels'.
"""
from datasets import Audio
WHISPER_SAMPLE_RATE = 16000
resolved = self._resolve_audio_columns(dataset, custom_format_mapping)
audio_col = resolved["audio_col"]
text_col = resolved["text_col"]
if not audio_col or not text_col:
raise ValueError(
f"Whisper dataset needs 'audio' and 'text' columns, got: {dataset.column_names}"
)
# Cast audio to 16kHz (Whisper's expected sample rate)
dataset = dataset.cast_column(
audio_col, Audio(sampling_rate = WHISPER_SAMPLE_RATE)
)
# Train/eval split (notebook does dataset.train_test_split)
eval_dataset_raw = None
if eval_split:
splits = dataset.train_test_split(test_size = 0.06, seed = 42)
dataset = splits["train"]
eval_dataset_raw = splits["test"]
self._update_progress(status_message = "Processing audio for Whisper...")
logger.info(
f"Whisper preprocessing: audio_col='{audio_col}', text_col='{text_col}', "
f"samples={len(dataset)}\n"
)
def process_split(ds, split_name = "train"):
processed = []
skipped = 0
for idx in range(len(ds)):
if self.should_stop:
logger.info(f"Stopped during Whisper {split_name} preprocessing\n")
break
example = ds[idx]
try:
audio_data = example.get(audio_col)
text = example.get(text_col)
if (
audio_data is None
or audio_data.get("array") is None
or not text
):
skipped += 1
continue
# Extract audio features (notebook line 112-115)
features = self.tokenizer.feature_extractor(
audio_data["array"], sampling_rate = audio_data["sampling_rate"]
)
# Tokenize text (notebook line 116)
tokenized_text = self.tokenizer.tokenizer(text)
processed.append(
{
"input_features": features.input_features[0],
"labels": tokenized_text.input_ids,
}
)
except Exception as e:
logger.warning(
f"Error processing Whisper {split_name} example {idx}: {e}"
)
skipped += 1
continue
if (idx + 1) % 100 == 0:
self._update_progress(
status_message = f"Processing {split_name} audio... {idx + 1}/{len(ds)}"
)
logger.info(
f"Whisper {split_name} preprocessing: {len(processed)} examples ({skipped} skipped)\n"
)
return processed
train_data = process_split(dataset, "train")
eval_data = (
process_split(eval_dataset_raw, "eval") if eval_dataset_raw else None
)
if not train_data:
raise ValueError("No valid examples after Whisper preprocessing")
return (train_data, eval_data)
@staticmethod
def _resolve_local_files(file_paths: list) -> list[str]:
"""Resolve a list of local dataset paths to concrete file paths."""
all_files: list[str] = []
for dataset_file in file_paths:
if os.path.isabs(dataset_file):
file_path = dataset_file
else:
file_path = str(resolve_dataset_path(dataset_file))
file_path_obj = Path(file_path)
if file_path_obj.is_dir():
parquet_dir = (
file_path_obj / "parquet-files"
if (file_path_obj / "parquet-files").exists()
else file_path_obj
)
parquet_files = sorted(parquet_dir.glob("*.parquet"))
if parquet_files:
all_files.extend(str(p) for p in parquet_files)
continue
candidates: list[Path] = []
for ext in (".json", ".jsonl", ".csv", ".parquet"):
candidates.extend(sorted(file_path_obj.glob(f"*{ext}")))
if candidates:
all_files.extend(str(c) for c in candidates)
continue
raise ValueError(
f"No supported data files in directory: {file_path_obj}"
)
else:
all_files.append(str(file_path_obj))
return all_files
@staticmethod
def _loader_for_files(files: list[str]) -> str:
"""Determine the HF datasets loader type from file extensions."""
first_ext = Path(files[0]).suffix.lower()
if first_ext in (".json", ".jsonl"):
return "json"
elif first_ext == ".csv":
return "csv"
elif first_ext == ".parquet":
return "parquet"
raise ValueError(f"Unsupported dataset format: {files[0]}")
def load_and_format_dataset(
self,
dataset_source: str,
format_type: str = "auto",
local_datasets: list = None,
local_eval_datasets: list = None,
custom_format_mapping: dict = None,
subset: str = None,
train_split: str = "train",
eval_split: str = None,
eval_steps: float = 0.00,
dataset_slice_start: int = None,
dataset_slice_end: int = None,
is_cpt: bool = False,
) -> Optional[tuple]:
"""
Load and prepare dataset for training.
Strategy: format first, then split — ensures both train and eval
portions are properly formatted and templated.
Returns:
Tuple of (dataset_info, eval_dataset) or None on error.
eval_dataset may be None if no eval split is available.
"""
try:
dataset = None
eval_dataset = None
has_separate_eval_source = (
False # True if eval comes from a separate HF split
)
eval_enabled = eval_steps is not None and eval_steps > 0
raw_text_mode = is_cpt or format_type == "raw"
def _raw_mode_label() -> str:
return "CPT" if is_cpt else "raw text"
def _apply_raw_text_prep(ds: Dataset, split_name: str) -> Dataset:
try:
result = prepare_raw_text_dataset(
ds,
mode_label = _raw_mode_label(),
split_name = split_name,
eos_token = getattr(self.tokenizer, "eos_token", None),
append_eos = True,
)
except ValueError as exc:
error_msg = str(exc)
logger.error(error_msg)
self._update_progress(error = error_msg)
raise
for notice in result.notices:
if notice.level == "warning":
logger.warning(notice.message)
if notice.update_status:
self._update_progress(status_message = notice.message)
else:
logger.info(f"{notice.message}\n")
return result.dataset
if local_datasets:
# Load local datasets using load_dataset() so the result is
# Arrow-backed (has cache files). Dataset.from_list() creates
# an in-memory dataset with no cache, which forces num_proc=1
# during tokenization/map because sharding requires Arrow files.
all_files = self._resolve_local_files(local_datasets)
if all_files:
loader = self._loader_for_files(all_files)
dataset = load_dataset(loader, data_files = all_files, split = "train")
# Check if stopped during dataset loading
if self.should_stop:
logger.info("Stopped during dataset loading\n")
return None
self._update_progress(
status_message = f"Loaded {len(dataset)} samples from local files"
)
logger.info(f"Loaded {len(dataset)} samples from local files\n")
logger.info(f"[DEBUG] Dataset cache_files: {dataset.cache_files}\n")
# Load local eval datasets if provided
if local_eval_datasets and eval_enabled:
eval_all_files = self._resolve_local_files(local_eval_datasets)
if eval_all_files:
eval_loader = self._loader_for_files(eval_all_files)
eval_dataset = load_dataset(
eval_loader, data_files = eval_all_files, split = "train"
)
has_separate_eval_source = True
logger.info(
f"Loaded {len(eval_dataset)} eval samples from local eval files\n"
)
elif dataset_source:
# Load from Hugging Face
split_name = train_split or "train"
load_kwargs = {"path": dataset_source, "split": split_name}
if subset:
load_kwargs["name"] = subset
_slice_start = dataset_slice_start or 0
if (
dataset_slice_end is not None
and dataset_slice_end >= 0
and dataset_slice_end >= _slice_start
):
# Manual slice — stream only the rows we need instead of
# downloading the entire dataset.
rows_to_stream = dataset_slice_end + 1
logger.info(
f"[dataset-slice] Manual slice specified "
f"(start={dataset_slice_start}, end={dataset_slice_end}), "
f"streaming {rows_to_stream} rows\n"
)
stream = load_dataset(**load_kwargs, streaming = True)
dataset = Dataset.from_list(list(stream.take(rows_to_stream)))
logger.info(
f"[dataset-slice] Downloaded {len(dataset)} rows "
f"(requested {rows_to_stream})\n"
)
self._update_progress(
status_message = f"Streamed {len(dataset)} rows from HuggingFace"
)
else:
self._update_progress(
status_message = f"Downloading dataset: {dataset_source}..."
)
dataset = load_dataset(**load_kwargs)
# Check if stopped during dataset loading
if self.should_stop:
logger.info("Stopped during dataset loading\n")
return None
n_rows = len(dataset) if hasattr(dataset, "__len__") else 0
self._update_progress(
status_message = f"Downloaded {dataset_source} ({n_rows:,} rows)"
)
logger.info(
f"Loaded dataset from Hugging Face: {dataset_source} ({n_rows:,} rows)\n"
)
# Resolve eval split from a separate HF split (explicit or auto-detected)
if eval_enabled:
effective_train = train_split or "train"
if eval_split and eval_split != effective_train:
# Explicit eval split provided - load it directly
logger.info(f"Loading explicit eval split: '{eval_split}'\n")
eval_load_kwargs = {"path": dataset_source, "split": eval_split}
if subset:
eval_load_kwargs["name"] = subset
eval_dataset = load_dataset(**eval_load_kwargs)
has_separate_eval_source = True
logger.info(
f"Loaded eval split '{eval_split}' with {len(eval_dataset)} rows\n"
)
elif eval_split and eval_split == effective_train:
# Same split as training — will do 80/20 split after formatting
logger.info(
f"Eval split '{eval_split}' is the same as train split — will split 80/20\n"
)
else:
# Auto-detect eval split from HF (returns a separate dataset, or None)
eval_dataset = self._auto_detect_eval_split_from_hf(
dataset_source = dataset_source,
subset = subset,
)
if eval_dataset is not None:
has_separate_eval_source = True
else:
logger.info(
"Eval disabled (eval_steps <= 0), skipping eval split detection\n"
)
if dataset is None:
raise ValueError("No dataset provided")
# Apply index range slicing if requested (inclusive on both ends)
if dataset_slice_start is not None or dataset_slice_end is not None:
total_rows = len(dataset)
start = dataset_slice_start if dataset_slice_start is not None else 0
end = (
dataset_slice_end
if dataset_slice_end is not None
else total_rows - 1
)
# Clamp to valid range
start = max(0, min(start, total_rows - 1))
end = max(start, min(end, total_rows - 1))
dataset = dataset.select(range(start, end + 1))
logger.info(
f"Sliced dataset to rows [{start}, {end}]: {len(dataset)} of {total_rows} rows\n"
)
self._update_progress(
status_message = f"Sliced dataset to {len(dataset)} rows (indices {start}-{end})"
)
# Check if stopped before applying template
if self.should_stop:
logger.info("Stopped before applying chat template\n")
return None
# ========== AUDIO MODELS: custom preprocessing ==========
if self._audio_type == "csm":
processed = self._preprocess_csm_dataset(dataset, custom_format_mapping)
return (processed, None)
elif self._audio_type == "whisper":
train_data, eval_data = self._preprocess_whisper_dataset(
dataset,
eval_split = eval_split,
custom_format_mapping = custom_format_mapping,
)
return (train_data, eval_data)
elif self._audio_type == "snac":
processed = self._preprocess_snac_dataset(
dataset, custom_format_mapping
)
return (processed, None)
elif self._audio_type == "bicodec":
processed = self._preprocess_bicodec_dataset(
dataset, custom_format_mapping
)
return ({"dataset": processed, "final_format": "audio_bicodec"}, None)
elif self._audio_type == "dac":
processed = self._preprocess_dac_dataset(dataset, custom_format_mapping)
return ({"dataset": processed, "final_format": "audio_dac"}, None)
# ========== RAW TEXT BYPASS ==========
if raw_text_mode:
logger.info(
f"{_raw_mode_label().capitalize()} mode: bypassing chat template, "
"using raw text\n"
)
dataset = _apply_raw_text_prep(dataset, "train")
if has_separate_eval_source and eval_dataset is not None:
eval_dataset = _apply_raw_text_prep(eval_dataset, "eval")
dataset_info = {
"dataset": dataset,
"detected_format": "raw_text",
"final_format": "raw_text",
"success": True,
}
if has_separate_eval_source and eval_dataset is not None:
logger.info(
f"{_raw_mode_label().capitalize()}: eval dataset "
f"({len(eval_dataset)} rows) kept as raw text\n"
)
elif eval_enabled and not has_separate_eval_source:
split_result = self._resolve_eval_split_from_dataset(dataset)
if split_result is not None:
train_portion, eval_dataset = split_result
dataset_info["dataset"] = train_portion
train_dataset = dataset_info["dataset"]
n = len(train_dataset) if hasattr(train_dataset, "__len__") else None
n_display = f"{n:,}" if isinstance(n, int) else "streaming"
self._update_progress(
status_message = f"Dataset ready ({n_display} samples, raw text)"
)
logger.info(f"Raw-text dataset ready ({n_display} samples)\n")
if "text" not in train_dataset.column_names:
raise ValueError(
f"Raw-text dataset missing 'text' column: {train_dataset.column_names}"
)
return (dataset_info, eval_dataset)
elif self.is_audio_vlm:
formatted = self._format_audio_vlm_dataset(
dataset, custom_format_mapping
)
return (formatted, None)
# ========== FORMAT FIRST ==========
logger.info(f"Formatting dataset with format_type='{format_type}'...\n")
dataset_info = format_and_template_dataset(
dataset,
model_name = self.model_name,
tokenizer = self.tokenizer,
is_vlm = self.is_vlm,
format_type = format_type,
dataset_name = dataset_source,
custom_format_mapping = custom_format_mapping,
progress_callback = self._update_progress,
)
# Check if stopped during formatting
if self.should_stop:
logger.info("Stopped during dataset formatting\n")
return None
# Abort if dataset formatting/conversion failed
if not dataset_info.get("success", True):
errors = dataset_info.get("errors", [])
error_msg = "; ".join(errors) if errors else "Dataset formatting failed"
logger.error(f"Dataset conversion failed: {error_msg}")
self._update_progress(error = error_msg)
return None
detected = dataset_info.get("detected_format", "unknown")
final_ds = dataset_info.get("dataset")
final_n = len(final_ds) if hasattr(final_ds, "__len__") else "?"
self._update_progress(
status_message = f"Dataset ready ({final_n:,} samples, {detected} format)"
)
logger.info(
f"Dataset formatted successfully ({final_n} samples, {detected})\n"
)
# ========== THEN SPLIT ==========
if has_separate_eval_source and eval_dataset is not None:
# Eval came from a separate HF split — format it too
logger.info(f"Formatting eval dataset ({len(eval_dataset)} rows)...\n")
eval_info = format_and_template_dataset(
eval_dataset,
model_name = self.model_name,
tokenizer = self.tokenizer,
is_vlm = self.is_vlm,
format_type = format_type,
dataset_name = dataset_source,
custom_format_mapping = custom_format_mapping,
)
eval_dataset = eval_info["dataset"]
logger.info(f"Eval dataset formatted successfully\n")
elif eval_enabled and not has_separate_eval_source:
# No separate eval source — split the already-formatted dataset
formatted_dataset = dataset_info["dataset"]
split_result = self._resolve_eval_split_from_dataset(formatted_dataset)
if split_result is not None:
train_portion, eval_dataset = split_result
dataset_info["dataset"] = train_portion
return (dataset_info, eval_dataset)
except Exception as e:
logger.error(f"Error loading dataset: {e}")
self._update_progress(error = str(e))
return None
def _auto_detect_eval_split_from_hf(
self, dataset_source: str, subset: str
) -> Optional[Dataset]:
"""Auto-detect an eval split from HF dataset (separate named split only)."""
try:
from datasets import get_dataset_split_names
load_kwargs = {"path": dataset_source}
if subset:
load_kwargs["config_name"] = subset
available_splits = get_dataset_split_names(**load_kwargs)
logger.info(f"Available splits: {available_splits}\n")
# Check for common eval split names
for candidate in ["eval", "validation", "valid", "val", "test"]:
if candidate in available_splits:
eval_load_kwargs = {"path": dataset_source, "split": candidate}
if subset:
eval_load_kwargs["name"] = subset
candidate_ds = load_dataset(**eval_load_kwargs)
if len(candidate_ds) >= 16:
logger.info(
f"Auto-detected eval split '{candidate}' with {len(candidate_ds)} rows\n"
)
return candidate_ds
else:
logger.info(
f"Found eval split '{candidate}' but only {len(candidate_ds)} rows (< 16), skipping\n"
)
except Exception as e:
logger.warning(f"Could not check dataset splits: {e}")
# No separate HF eval split found — caller will handle programmatic splitting
return None
def _resolve_eval_split_from_dataset(self, dataset) -> Optional[tuple]:
"""Split a dataset into train and eval portions.
Returns:
Tuple of (train_dataset, eval_dataset), or None if dataset too small.
"""
MIN_EVAL_ROWS = 16
MIN_TOTAL_ROWS = 32 # Need at least 16 train + 16 eval
n = len(dataset)
if n < MIN_TOTAL_ROWS:
logger.info(f"Dataset too small ({n} rows) for eval split, skipping eval\n")
return None
eval_size = max(MIN_EVAL_ROWS, min(128, int(0.05 * n)))
# Ensure we don't take more than half the dataset
eval_size = min(eval_size, n // 2)
logger.info(f"Auto-splitting: {eval_size} rows for eval from {n} total\n")
split_result = dataset.train_test_split(test_size = eval_size, seed = 3407)
logger.info(
f"Split complete: {len(split_result['train'])} train, {len(split_result['test'])} eval\n"
)
return (split_result["train"], split_result["test"])
def start_training(
self,
dataset: Dataset,
eval_dataset: Dataset = None,
eval_steps: float = 0.00,
output_dir: str | None = None,
num_epochs: int = 3,
learning_rate: float = 2e-4,
embedding_learning_rate: float | None = None,
batch_size: int = 2,
gradient_accumulation_steps: int = 4,
warmup_steps: int = None,
warmup_ratio: float = None,
max_steps: int = 0,
save_steps: int = 0,
weight_decay: float = 0.001,
random_seed: int = 3407,
packing: bool = False,
train_on_completions: bool = False,
enable_wandb: bool = False,
wandb_project: str = "unsloth-training",
wandb_token: str = None,
enable_tensorboard: bool = False,
tensorboard_dir: str | None = None,
**kwargs,
) -> bool:
"""Start training in a separate thread"""
if self.is_training:
logger.warning("Training already in progress")
return False
if self.model is None or self.tokenizer is None:
self._update_progress(error = "Model not loaded")
return False
# Pre-import heavy transformers modules on the main thread.
# Unsloth's patched_import hook (deepseek_v3_moe.py) is not thread-safe
# with Python's importlib cache, causing KeyError: 'size' if these are
# first imported inside the worker thread.
import transformers # noqa: F401 ensures submodules are cached
from transformers import ( # noqa: F401
Trainer as _HFTrainer,
TrainingArguments as _TrainingArguments,
TrainerCallback as _TrainerCallback,
)
if self._audio_type == "whisper":
from transformers import ( # noqa: F401
Seq2SeqTrainer as _Seq2SeqTrainer,
Seq2SeqTrainingArguments as _Seq2SeqTrainingArguments,
)
# Start training in separate thread
self.training_thread = threading.Thread(
target = self._train_worker,
args = (dataset,),
kwargs = {
"output_dir": output_dir,
"num_epochs": num_epochs,
"learning_rate": learning_rate,
"embedding_learning_rate": embedding_learning_rate,
"batch_size": batch_size,
"gradient_accumulation_steps": gradient_accumulation_steps,
"warmup_steps": warmup_steps,
"warmup_ratio": warmup_ratio,
"max_steps": max_steps,
"save_steps": save_steps,
"weight_decay": weight_decay,
"random_seed": random_seed,
"packing": packing,
"train_on_completions": train_on_completions,
"enable_wandb": enable_wandb,
"wandb_project": wandb_project,
"wandb_token": wandb_token,
"enable_tensorboard": enable_tensorboard,
"tensorboard_dir": tensorboard_dir,
"eval_dataset": eval_dataset,
"eval_steps": eval_steps,
**kwargs,
},
)
self.should_stop = False
self.is_training = True
try:
self.training_thread.start()
return True
except Exception as e:
self.is_training = False
logger.error(f"Failed to start training thread: {e}")
return False
def _train_worker(self, dataset: Dataset, **training_args):
"""Worker function for training (runs in separate thread)"""
try:
# On spawn-based platforms (Windows, macOS), register all known
# compiled-cache directories on sys.path and PYTHONPATH before any
# dataset.map() call so spawned workers can import dynamically
# compiled modules such as UnslothSFTTrainer.
if sys.platform in ("win32", "darwin"):
from utils.cache_cleanup import register_compiled_cache_on_path
register_compiled_cache_on_path()
# Store training parameters for metrics calculation
self.batch_size = training_args.get("batch_size", 2)
self.max_seq_length = training_args.get("max_seq_length", 2048)
self.gradient_accumulation_steps = training_args.get(
"gradient_accumulation_steps", 4
)
# Set training start time
self.training_start_time = time.time()
self._update_progress(is_training = True, error = None)
# Setup logging
if training_args.get("enable_wandb", False) and training_args.get(
"wandb_token"
):
os.environ["WANDB_API_KEY"] = training_args["wandb_token"]
import wandb
wandb.init(
project = training_args.get("wandb_project", "unsloth-training")
)
# Create output directory
output_dir = str(resolve_output_dir(training_args.get("output_dir")))
ensure_dir(Path(output_dir))
# ========== AUDIO TRAINER BRANCH ==========
if self._audio_type == "csm":
# CSM uses plain HF Trainer (NOT SFTTrainer)
# Needs remove_unused_columns=False for depth decoder (input_values + cutoffs)
from transformers import Trainer as HFTrainer, TrainingArguments
self._apply_csm_forward_fix()
config = self._build_audio_training_args(
training_args,
output_dir,
extra_args = {
"remove_unused_columns": False,
},
)
self.trainer = HFTrainer(
model = self.model,
train_dataset = dataset,
args = TrainingArguments(**config),
)
self.trainer.add_callback(self._create_progress_callback())
batch_size = training_args.get("batch_size", 2)
total = self._calculate_total_steps(
len(dataset),
batch_size,
training_args.get("gradient_accumulation_steps", 4),
training_args.get("num_epochs", 3),
training_args.get("max_steps", 0),
)
self._update_progress(
total_steps = total, status_message = "Starting CSM training..."
)
logger.info(f"CSM training config: {config}\n")
self.trainer.train(
resume_from_checkpoint = training_args.get("resume_from_checkpoint")
)
self._finalize_training(output_dir, "CSM")
return
elif self._audio_type == "snac":
# Orpheus: language model with SNAC codec tokens — plain HF Trainer
# DataCollatorForSeq2Seq dynamically pads variable-length sequences per batch
# (text + audio codes vary in length) and pads labels with -100.
from transformers import (
Trainer as HFTrainer,
TrainingArguments,
DataCollatorForSeq2Seq,
)
config = self._build_audio_training_args(training_args, output_dir)
self.trainer = HFTrainer(
model = self.model,
train_dataset = dataset,
args = TrainingArguments(**config),
data_collator = DataCollatorForSeq2Seq(
tokenizer = self.tokenizer,
padding = True,
pad_to_multiple_of = 8,
),
)
self.trainer.add_callback(self._create_progress_callback())
batch_size = training_args.get("batch_size", 2)
total = self._calculate_total_steps(
len(dataset),
batch_size,
training_args.get("gradient_accumulation_steps", 4),
training_args.get("num_epochs", 3),
training_args.get("max_steps", 0),
)
self._update_progress(
total_steps = total, status_message = "Starting SNAC training..."
)
logger.info(f"SNAC training config: {config}\n")
self.trainer.train(
resume_from_checkpoint = training_args.get("resume_from_checkpoint")
)
self._finalize_training(output_dir, "SNAC")
return
elif self._audio_type == "whisper":
# Whisper: Seq2SeqTrainer with custom speech collator
from transformers import Seq2SeqTrainer, Seq2SeqTrainingArguments
from utils.datasets import DataCollatorSpeechSeq2SeqWithPadding
eval_dataset = training_args.get("eval_dataset", None)
extra = {"remove_unused_columns": False, "label_names": ["labels"]}
if eval_dataset:
extra["eval_strategy"] = "steps"
extra["eval_steps"] = training_args.get("eval_steps", 5)
config = self._build_audio_training_args(
training_args, output_dir, extra_args = extra
)
trainer_kwargs = {
"model": self.model,
"train_dataset": dataset,
"data_collator": DataCollatorSpeechSeq2SeqWithPadding(
processor = self.tokenizer
),
"processing_class": self.tokenizer.feature_extractor,
"args": Seq2SeqTrainingArguments(**config),
}
if eval_dataset:
trainer_kwargs["eval_dataset"] = eval_dataset
self.trainer = Seq2SeqTrainer(**trainer_kwargs)
self.trainer.add_callback(self._create_progress_callback())
batch_size = training_args.get("batch_size", 2)
total = self._calculate_total_steps(
len(dataset),
batch_size,
training_args.get("gradient_accumulation_steps", 4),
training_args.get("num_epochs", 3),
training_args.get("max_steps", 0),
)
self._update_progress(
total_steps = total, status_message = "Starting Whisper training..."
)
logger.info(f"Whisper training config: {config}\n")
self.trainer.train(
resume_from_checkpoint = training_args.get("resume_from_checkpoint")
)
self._finalize_training(output_dir, "Whisper")
return
elif self._audio_type is not None and self._audio_type not in (
"bicodec",
"dac",
):
# bicodec/dac use the standard SFTTrainer text path below
raise NotImplementedError(
f"Audio training for '{self._audio_type}' not yet implemented"
)
# ========== DATA COLLATOR SELECTION ==========
# Detect special model types
model_name_lower = self.model_name.lower()
is_deepseek_ocr = (
"deepseek" in model_name_lower and "ocr" in model_name_lower
)
logger.info("Configuring data collator...\n")
dataset_final_format = (
str(dataset.get("final_format", "")).lower()
if isinstance(dataset, dict)
else ""
)
raw_text_mode = dataset_final_format == "raw_text"
data_collator = None # Default to built-in data collator
if is_deepseek_ocr:
# Special DeepSeek OCR collator - auto-install if needed
logger.info("Detected DeepSeek OCR model\n")
# Ensure DeepSeek OCR module is installed
if not _ensure_deepseek_ocr_installed():
error_msg = (
"Failed to install DeepSeek OCR module. "
"Please install manually: "
"from huggingface_hub import snapshot_download; "
"snapshot_download('unsloth/DeepSeek-OCR', local_dir='deepseek_ocr')"
)
logger.error(error_msg)
self._update_progress(error = error_msg, is_training = False)
return
try:
from backend.data_utils import DeepSeekOCRDataCollator
logger.info("Configuring DeepSeek OCR data collator...\n")
FastVisionModel.for_training(self.model)
# DeepSeek OCR's (image_size, base_size, crop_mode) is a
# coupled preset; changing image_size alone desyncs the
# per-crop pixel grid from num_queries. Use Gundam.
if training_args.get("vision_image_size") is not None:
logger.info(
"Vision image resize ignored for DeepSeek OCR "
"(uses fixed Gundam preset).\n"
)
data_collator = DeepSeekOCRDataCollator(
tokenizer = self.tokenizer,
model = self.model,
image_size = 640,
base_size = 1024,
crop_mode = True,
train_on_responses_only = training_args.get(
"train_on_completions", False
),
)
logger.info("DeepSeek OCR data collator configured successfully\n")
except Exception as e:
logger.error(f"Failed to configure DeepSeek OCR collator: {e}")
error_msg = f"Error configuring DeepSeek OCR: {str(e)}"
self._update_progress(error = error_msg, is_training = False)
return
elif self.is_audio_vlm and not raw_text_mode:
# Audio VLM collator (e.g. Gemma 3N with audio data)
# Mirrors the collate_fn from Gemma3N_(4B)-Audio notebook
logger.info("Configuring audio VLM data collator...\n")
processor = self.tokenizer # FastModel returns processor as tokenizer
audio_col_name = getattr(self, "_audio_vlm_audio_col", "audio")
def audio_vlm_collate_fn(examples):
texts = []
audios = []
for example in examples:
text = processor.apply_chat_template(
example["messages"],
tokenize = False,
add_generation_prompt = False,
).strip()
texts.append(text)
audios.append(example[audio_col_name]["array"])
batch = processor(
text = texts, audio = audios, return_tensors = "pt", padding = True
)
# Labels = input_ids with special tokens masked
labels = batch["input_ids"].clone()
labels[labels == processor.tokenizer.pad_token_id] = -100
for attr in (
"audio_token_id",
"image_token_id",
"boi_token_id",
"eoi_token_id",
):
token_id = getattr(processor.tokenizer, attr, None)
if token_id is not None:
labels[labels == token_id] = -100
batch["labels"] = labels
return batch
data_collator = audio_vlm_collate_fn
logger.info("Audio VLM data collator configured\n")
elif self.is_vlm and not raw_text_mode:
# Standard VLM collator (images)
logger.info("Using UnslothVisionDataCollator for vision model\n")
from unsloth.trainer import UnslothVisionDataCollator
FastVisionModel.for_training(self.model)
vision_image_size = training_args.get("vision_image_size")
if vision_image_size is None:
data_collator = UnslothVisionDataCollator(
self.model, self.tokenizer
)
else:
logger.info(
f"Vision image resize: {vision_image_size} (max dimension)\n"
)
data_collator = UnslothVisionDataCollator(
self.model,
self.tokenizer,
resize = vision_image_size,
resize_dimension = "max",
)
logger.info("Vision data collator configured\n")
# ========== TRAINING CONFIGURATION ==========
# Handle warmup_steps vs warmup_ratio
warmup_steps_val = training_args.get("warmup_steps", None)
warmup_ratio_val = training_args.get("warmup_ratio", None)
lr_value = training_args.get("learning_rate", 2e-4)
logger.info(
f"[DEBUG] learning_rate from training_args: {lr_value} (type: {type(lr_value).__name__})\n"
)
config_args = {
"per_device_train_batch_size": training_args.get("batch_size", 2),
"gradient_accumulation_steps": training_args.get(
"gradient_accumulation_steps", 4
),
"num_train_epochs": training_args.get(
"num_epochs", 3
), # Default to epochs
"learning_rate": lr_value,
"fp16": not is_bfloat16_supported(),
"bf16": is_bfloat16_supported(),
"logging_steps": 1,
"weight_decay": training_args.get("weight_decay", 0.001),
"seed": training_args.get("random_seed", 3407),
"output_dir": output_dir,
"report_to": _build_report_targets(training_args),
"include_num_input_tokens_seen": True, # Enable token counting
"dataset_num_proc": dataset_map_num_proc(
1
if (self.is_audio or self.is_audio_vlm or self._cuda_audio_used)
else max(1, (os.cpu_count() or 1) // 4)
),
"max_seq_length": training_args.get("max_seq_length", 2048),
}
if training_args.get("enable_tensorboard", False):
config_args["logging_dir"] = str(
resolve_tensorboard_dir(training_args.get("tensorboard_dir"))
)
logger.info(
f"[DEBUG] dataset_num_proc={config_args['dataset_num_proc']} (is_audio={self.is_audio}, is_audio_vlm={self.is_audio_vlm}, _cuda_audio_used={self._cuda_audio_used})"
)
# On spawn-based platforms (Windows, macOS) with transformers 5.x,
# disable DataLoader multiprocessing to avoid issues with modified
# sys.path (.venv_t5) in spawned workers.
if sys.platform in ("win32", "darwin"):
import transformers as _tf
if _tf.__version__.startswith("5."):
config_args["dataloader_num_workers"] = 0
# Add warmup parameter - use warmup_ratio if provided, otherwise warmup_steps
if warmup_ratio_val is not None:
config_args["warmup_ratio"] = warmup_ratio_val
logger.info(f"Using warmup_ratio: {warmup_ratio_val}\n")
elif warmup_steps_val is not None:
config_args["warmup_steps"] = warmup_steps_val
logger.info(f"Using warmup_steps: {warmup_steps_val}\n")
else:
# Default to warmup_steps if neither provided
config_args["warmup_steps"] = 5
logger.info(f"Using default warmup_steps: 5\n")
# Add save_steps if specified
save_steps_val = training_args.get("save_steps", 0)
if save_steps_val and save_steps_val > 0:
config_args["save_steps"] = save_steps_val
config_args["save_strategy"] = "steps"
# If max_steps is specified, use it instead of epochs
max_steps_val = training_args.get("max_steps", 0)
if max_steps_val and max_steps_val > 0:
del config_args["num_train_epochs"] # Remove epochs
config_args["max_steps"] = max_steps_val # Use steps instead
logger.info(f"Training for {max_steps_val} steps\n")
else:
logger.info(f"Training for {config_args['num_train_epochs']} epochs\n")
# ========== EVAL CONFIGURATION ==========
eval_dataset = training_args.get("eval_dataset", None)
eval_steps_val = training_args.get("eval_steps", 0.00)
if eval_dataset is not None:
if eval_steps_val > 0:
config_args["eval_strategy"] = "steps"
config_args["eval_steps"] = eval_steps_val
config_args["per_device_eval_batch_size"] = config_args[
"per_device_train_batch_size"
]
logger.info(
f"✅ Evaluation enabled: eval_steps={eval_steps_val} (fraction of total steps)\n"
)
logger.info(f"Eval dataset: {len(eval_dataset)} rows\n")
else:
logger.info(
f"⚠️ Eval dataset provided but eval_steps={eval_steps_val} (disabled)\n"
)
logger.info("To enable evaluation, set eval_steps > 0.0\n")
else:
logger.info("No eval dataset — evaluation disabled\n")
# Add model-specific parameters
# Use optim and lr_scheduler_type from training_args if provided, otherwise use defaults
optim_value = training_args.get("optim", "adamw_8bit")
lr_scheduler_type_value = training_args.get("lr_scheduler_type", "linear")
if (self.is_vlm or self.is_audio_vlm) and not raw_text_mode:
# Vision / audio VLM config (both need skip_prepare_dataset + remove_unused_columns)
# Raw-text runs on VLM-capable models are routed to the text path below.
label = "audio VLM" if self.is_audio_vlm else "vision"
logger.info(f"Configuring {label} model training parameters\n")
# Use provided values or defaults for vision models
optim_value = training_args.get("optim", "adamw_torch_fused")
lr_scheduler_type_value = training_args.get(
"lr_scheduler_type", "cosine"
)
config_args.update(
{
"optim": optim_value,
"lr_scheduler_type": lr_scheduler_type_value,
"gradient_checkpointing": True,
"gradient_checkpointing_kwargs": {"use_reentrant": False},
"max_grad_norm": 0.3,
"remove_unused_columns": False,
"dataset_text_field": "",
"dataset_kwargs": {"skip_prepare_dataset": True},
"max_length": training_args.get("max_seq_length", 2048),
}
)
else:
is_cpt = training_args.get("is_cpt", False)
self.is_cpt = is_cpt
if is_cpt:
logger.info("Configuring Continued Pretraining (CPT) parameters\n")
elif raw_text_mode:
logger.info("Configuring raw-text training parameters\n")
else:
logger.info("Configuring text model training parameters\n")
config_args.update(
{
"optim": optim_value,
"lr_scheduler_type": lr_scheduler_type_value,
"dataset_text_field": "text",
}
)
# Only add packing for text models (not DeepSeek OCR which is VLM)
if not is_deepseek_ocr:
packing_enabled = training_args.get("packing", False)
config_args["packing"] = packing_enabled
logger.info(
f"Sequence packing: {'enabled' if packing_enabled else 'disabled'}\n"
)
# Audio codec overrides — BiCodec/DAC use the text SFTTrainer path
if self._audio_type == "bicodec":
config_args["packing"] = False
logger.info("Applied BiCodec overrides: packing=False\n")
elif self._audio_type == "dac":
config_args["packing"] = False
logger.info("Applied DAC overrides: packing=False\n")
logger.info(f"The configuration is: {config_args}")
logger.info("Training configuration prepared\n")
# ========== TRAINER INITIALIZATION ==========
if self.is_audio_vlm and not raw_text_mode:
# Audio VLM (e.g. Gemma 3N + audio): raw Dataset from _format_audio_vlm_dataset
# Notebook uses processing_class=processor.tokenizer (text tokenizer only)
# Raw-text runs are routed to the text path below.
train_dataset = (
dataset if isinstance(dataset, Dataset) else dataset["dataset"]
)
processing_class = (
self.tokenizer.tokenizer
if hasattr(self.tokenizer, "tokenizer")
else self.tokenizer
)
trainer_kwargs = {
"model": self.model,
"train_dataset": train_dataset,
"processing_class": processing_class,
"data_collator": data_collator,
"args": SFTConfig(**config_args),
}
if eval_dataset is not None:
trainer_kwargs["eval_dataset"] = eval_dataset
self.trainer = SFTTrainer(**trainer_kwargs)
elif self.is_vlm and not raw_text_mode:
# Image VLM: dataset is dict wrapper from format_and_template_dataset
# Raw-text runs are routed to the text path below.
train_dataset = (
dataset["dataset"] if isinstance(dataset, dict) else dataset
)
trainer_kwargs = {
"model": self.model,
"train_dataset": train_dataset,
"processing_class": self.tokenizer,
"data_collator": data_collator,
"args": SFTConfig(**config_args),
}
if eval_dataset is not None:
trainer_kwargs["eval_dataset"] = eval_dataset
self.trainer = SFTTrainer(**trainer_kwargs)
else:
# For text-only training, if the tokenizer is actually a Processor
# (e.g., Gemma-3 returns a ProcessorMixin even for text), we must
# unwrap to the raw tokenizer. Otherwise Unsloth's SFTTrainer detects
# ProcessorMixin → sets _is_vlm=True → skips _prepare_dataset entirely,
# and the 'text' column never gets tokenized to 'input_ids'.
from transformers import ProcessorMixin
sft_tokenizer = self.tokenizer
if isinstance(self.tokenizer, ProcessorMixin) and hasattr(
self.tokenizer, "tokenizer"
):
logger.info(
f" ⚠️ Unwrapping Processor → raw tokenizer for text-only SFTTrainer"
)
sft_tokenizer = self.tokenizer.tokenizer
if is_cpt:
try:
from unsloth import (
UnslothTrainer as _UnslothCPTTrainer,
UnslothTrainingArguments as _UnslothTrainingArguments,
)
except ImportError as exc:
raise RuntimeError(
"CPT requires a newer Unsloth install that exports "
"`UnslothTrainer` and `UnslothTrainingArguments` "
"(for embedding_learning_rate support). "
"Upgrade with: `pip install -U unsloth unsloth_zoo`."
) from exc
embedding_lr = training_args.get("embedding_learning_rate")
logger.info(
f"CPT: using UnslothTrainer with embedding_learning_rate={embedding_lr}\n"
)
trainer_kwargs = {
"model": self.model,
"tokenizer": sft_tokenizer,
"train_dataset": dataset["dataset"],
"data_collator": data_collator,
"args": _UnslothTrainingArguments(
embedding_learning_rate = embedding_lr,
**config_args,
),
}
if eval_dataset is not None:
trainer_kwargs["eval_dataset"] = eval_dataset
self.trainer = _UnslothCPTTrainer(**trainer_kwargs)
else:
trainer_kwargs = {
"model": self.model,
"tokenizer": sft_tokenizer,
"train_dataset": dataset["dataset"],
"data_collator": data_collator,
"args": SFTConfig(**config_args),
}
if eval_dataset is not None:
trainer_kwargs["eval_dataset"] = eval_dataset
self.trainer = SFTTrainer(**trainer_kwargs)
# Restore the full processor as processing_class so checkpoint
# saves include preprocessor_config.json (needed for GGUF export).
if sft_tokenizer is not self.tokenizer:
self.trainer.processing_class = self.tokenizer
logger.info("Trainer initialized\n")
# ========== TRAIN ON RESPONSES ONLY ==========
# Determine if we should train on responses only
# Raw-text datasets always train on all tokens.
instruction_part = None
response_part = None
is_cpt = training_args.get("is_cpt", False)
train_on_responses_enabled = (
False
if (is_cpt or raw_text_mode)
else training_args.get("train_on_completions", False)
)
if is_cpt:
logger.info(
"CPT mode: skipping train_on_responses_only — training on all tokens\n"
)
elif raw_text_mode:
logger.info(
"Raw-text mode: skipping train_on_responses_only — training on all tokens\n"
)
# DeepSeek OCR handles this internally in its collator, so skip
# Audio VLM handles label masking in its collator, so skip
if (
train_on_responses_enabled
and not self.is_audio_vlm
and not self.is_audio
and not (is_deepseek_ocr or dataset_final_format == "alpaca")
):
try:
logger.info("Configuring train on responses only...\n")
# Get the template mapping for this model
model_name_lower = self.model_name.lower()
if model_name_lower in MODEL_TO_TEMPLATE_MAPPER:
template_name = MODEL_TO_TEMPLATE_MAPPER[model_name_lower]
logger.info(f"Detected template: {template_name}\n")
if template_name in TEMPLATE_TO_RESPONSES_MAPPER:
instruction_part = TEMPLATE_TO_RESPONSES_MAPPER[
template_name
]["instruction"]
response_part = TEMPLATE_TO_RESPONSES_MAPPER[template_name][
"response"
]
logger.info(
f"Instruction marker: {instruction_part[:50]}...\n"
)
logger.info(f"Response marker: {response_part[:50]}...\n")
else:
logger.info(
f"No response mapping found for template: {template_name}\n"
)
train_on_responses_enabled = False
else:
logger.info(
f"No template mapping found for model: {self.model_name}\n"
)
train_on_responses_enabled = False
except Exception as e:
logger.warning(f"Could not configure train on responses: {e}")
train_on_responses_enabled = False
# Apply train on responses only if we have valid parts
if (
train_on_responses_enabled
and instruction_part
and response_part
and not self.is_audio_vlm
and not self.is_audio
and not (is_deepseek_ocr or dataset_final_format == "alpaca")
):
try:
from unsloth.chat_templates import train_on_responses_only
self.trainer = train_on_responses_only(
self.trainer,
instruction_part = instruction_part,
response_part = response_part,
num_proc = config_args["dataset_num_proc"],
)
logger.info("Train on responses only configured successfully\n")
# ── Safety net: check if all samples were filtered out ──
# Unsloth's train_on_responses_only masks non-response
# tokens with -100. If max_seq_length is too short and the
# response portion gets truncated away, EVERY sample ends
# up with all labels == -100 and Unsloth removes them,
# leaving 0 usable training samples.
filtered_len = len(self.trainer.train_dataset)
original_len = len(dataset["dataset"])
dropped = original_len - filtered_len
drop_pct = (
round(100 * dropped / original_len, 1)
if original_len > 0
else 0
)
if filtered_len == 0 or drop_pct > 30:
max_seq = training_args.get("max_seq_length", 2048)
error_msg = (
f"{dropped}/{original_len} samples ({drop_pct}%) "
f"were dropped after applying 'train on responses "
f"only' — only {filtered_len} remain. This usually "
f"means max_seq_length ({max_seq}) is too short "
f"and the response portion is being truncated "
f"away. Try increasing max_seq_length (e.g. 8192) "
f"or disabling 'Train on completions'."
)
logger.error(error_msg)
self._update_progress(error = error_msg, is_training = False)
return
if dropped > 0:
logger.info(
f"⚠️ {dropped}/{original_len} samples "
f"({drop_pct}%) were dropped (all labels "
f"masked). {filtered_len} samples remain.\n"
)
logger.info(f"Post-filter dataset size: {filtered_len} samples\n")
# [DEBUG] Decode first sample AFTER train_on_completions applied
# try:
# _row = self.trainer.train_dataset[0]
# _space = self.tokenizer(
# " ", add_special_tokens = False
# ).input_ids[0]
# print("[DEBUG] === After train_on_completions ===", flush = True)
# print(
# f"[DEBUG] input_ids decoded:\n{self.tokenizer.decode(_row['input_ids'])}\n",
# flush = True,
# )
# print(
# f"[DEBUG] labels decoded (-100 → space):\n{self.tokenizer.decode([_space if x == -100 else x for x in _row['labels']])}\n",
# flush = True,
# )
# except Exception as _dbg_e:
# print(
# f"[DEBUG] Could not decode post-completions sample: {_dbg_e}",
# flush = True,
# )
except Exception as e:
logger.warning(f"Failed to apply train on responses only: {e}")
train_on_responses_enabled = False
else:
if train_on_responses_enabled and is_deepseek_ocr:
logger.info("Train on responses handled by DeepSeek OCR collator\n")
else:
logger.info("Training on full sequences (including prompts)\n")
# ========== PROGRESS TRACKING ==========
self.trainer.add_callback(self._create_progress_callback())
num_samples = len(
dataset["dataset"] if isinstance(dataset, dict) else dataset
)
batch_size = training_args.get("batch_size", 2)
total_steps = self._calculate_total_steps(
num_samples,
batch_size,
training_args.get("gradient_accumulation_steps", 4),
training_args.get("num_epochs", 3),
training_args.get("max_steps", 0),
)
self._update_progress(total_steps = total_steps)
# ========== START TRAINING ==========
self._update_progress(status_message = "Starting training...")
logger.info("Starting training...\n")
self.trainer.train(
resume_from_checkpoint = training_args.get("resume_from_checkpoint")
)
# ========== SAVE MODEL ==========
self._finalize_training(output_dir)
except Exception as e:
import traceback
logger.error(f"Training error: {e}")
logger.error(f"Full traceback:\n{traceback.format_exc()}")
self._update_progress(is_training = False, error = str(e))
finally:
self.is_training = False
def _patch_adapter_config(self, output_dir: str) -> None:
"""Patch adapter_config.json with unsloth_training_method.
Values: 'qlora', 'lora', 'FT', 'CPT', 'DPO', 'GRPO', etc.
For LoRA/QLoRA, the distinction comes from load_in_4bit.
"""
config_path = os.path.join(output_dir, "adapter_config.json")
if not os.path.exists(config_path):
logger.info("No adapter_config.json found — skipping training method patch")
return
try:
with open(config_path, "r") as f:
config = json.load(f)
# Determine the training method
if self.is_cpt:
method = "CPT"
elif self.load_in_4bit:
method = "qlora"
else:
method = "lora"
config["unsloth_training_method"] = method
logger.info(
f"Patching adapter_config.json with unsloth_training_method='{method}'"
)
with open(config_path, "w") as f:
json.dump(config, f, indent = 2)
except Exception as e:
logger.warning(f"Failed to patch adapter_config.json: {e}")
def stop_training(self, save: bool = True):
"""Stop ongoing training"""
logger.info(f"\nStopping training (save={save})...")
self.should_stop = True
self.save_on_stop = save
stop_msg = (
"Stopping training and saving checkpoint..."
if save
else "Cancelling training..."
)
self._update_progress(status_message = stop_msg)
# If trainer exists, try to stop it gracefully
if self.trainer:
try:
# The callback will catch should_stop flag and stop the training loop
logger.info("Training will stop at next step...\n")
except Exception as e:
logger.error(f"Error stopping trainer: {e}")
def get_training_progress(self) -> TrainingProgress:
"""Get current training progress"""
with self._lock:
return self.training_progress
def cleanup(self):
"""Cleanup resources"""
if self.trainer:
self.trainer = None
if self.model:
self.model = None
if self.tokenizer:
self.tokenizer = None
# Clear GPU memory
clear_gpu_cache()
def _ensure_deepseek_ocr_installed():
"""
Auto-install DeepSeek OCR module if not available.
Downloads from HuggingFace hub as a local module.
Returns:
bool: True if available (either already installed or just installed)
"""
try:
# Try importing to see if already available
from deepseek_ocr.modeling_deepseekocr import format_messages
logger.info("DeepSeek OCR module already available")
return True
except ImportError:
pass
try:
logger.info(
"DeepSeek OCR module not found. Auto-installing from HuggingFace..."
)
logger.info("\n Downloading DeepSeek OCR module from HuggingFace...\n")
from huggingface_hub import snapshot_download
import sys
import os
# Get the script directory to install locally
script_dir = os.path.dirname(os.path.abspath(__file__))
parent_dir = os.path.dirname(script_dir) # Go up to project root
# Download to project root as 'deepseek_ocr' folder
local_dir = os.path.join(parent_dir, "deepseek_ocr")
snapshot_download(
"unsloth/DeepSeek-OCR", local_dir = local_dir, local_dir_use_symlinks = False
)
# Add to sys.path if not already there
if parent_dir not in sys.path:
sys.path.insert(0, parent_dir)
# Try importing again
from deepseek_ocr.modeling_deepseekocr import format_messages
logger.info("DeepSeek OCR module installed successfully")
logger.info("DeepSeek OCR module installed successfully!\n")
return True
except Exception as e:
logger.error(f"Failed to install DeepSeek OCR module: {e}")
logger.info(f"\n❌ Failed to install DeepSeek OCR module: {e}\n")
return False
# Global trainer instance
_trainer_instance = None
def get_trainer() -> UnslothTrainer:
"""Get global trainer instance"""
global _trainer_instance
if _trainer_instance is None:
_trainer_instance = UnslothTrainer()
return _trainer_instance