unsloth/studio/backend/tests/test_diffusion_device.py
Daniel Han 96940d87b7
Studio diffusion (Phase 4): native stable-diffusion.cpp engine for CPU/Mac (#6679)
* Studio diffusion: cross-platform device policy, fp16 guard, lock split, validate-before-evict

Phase 1 of porting the richer diffusion stack onto the image-generation backend.

- Add a compartmentalized device/dtype policy module (diffusion_device.py)
  resolving CUDA/ROCm/XPU/MPS/CPU with capability flags. Keeps the NVIDIA
  capability-based bf16 choice; ROCm and XPU are isolated; MPS uses bf16 or
  fp32, never a silent fp16 that renders a black image.
- Add a per-family fp16_incompatible flag (Z-Image) and promote a resolved
  float16 to float32 for those families so they do not produce black images.
- Split the backend locks: a generation holds only _generate_lock, so status,
  unload, and a new load are never blocked by a long denoise. Add per-generation
  cancellation via callback_on_step_end so an eviction or a superseding load
  preempts a running generation; a replacement load waits for it to stop before
  allocating, so two pipelines never sit in VRAM at once.
- Validate a load request before the GPU handoff so an unloadable pick never
  evicts a working chat model, and reject missing local paths up front.
- Add CPU-only tests for the device policy, dtype guard, lock split and
  cancellation, and validate-before-evict, plus a GPU benchmark/regression
  script (scripts/diffusion_bench.py) measuring latency, peak VRAM, and PSNR
  against a saved reference.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio diffusion (Phase 2A): measured-budget memory planner + offload/VAE policy

Add a lean, backend-agnostic memory policy that picks a CPU-offload policy and
VAE tiling/slicing from measured free device memory vs the model's estimated
resident footprint, then applies it to the built pipeline. auto stays resident
when the model fits (byte-identical to the prior resident path), and falls to
whole-module offload when tight; fast/balanced/low_vram are explicit overrides.
Sequential submodule offload is unreliable for GGUF transformers on diffusers
0.38, so it falls back to whole-module offload and status reports the policy
actually engaged.

Verified on Z-Image-Turbo Q4_K_M (B200): auto reproduces the resident image with
no VRAM/latency regression (PSNR inf); balanced/low_vram cut generation peak VRAM
47.9% (15951 -> 8318 MB) with byte-identical output, at the expected latency cost.

73 prior + 35 new CPU tests pass.

* Studio diffusion (Phase 2D): streamed block-level offload + functional VAE tiling

Add a streamed 'group' offload tier (diffusers apply_group_offloading, block_level,
use_stream) that keeps the transformer flowing through the GPU a few blocks at a
time while the text encoder / VAE stay resident, and fix VAE tiling to drive the
VAE submodule (pipelines like Z-Image expose enable_tiling on pipe.vae, not the
pipeline). apply_memory_plan now returns the (policy, tiling) actually engaged so
status never overstates either, and group falls back to whole-module offload when
the transformer can't be streamed.

Measured on Z-Image (B200), all lossless (PSNR inf vs resident): balanced/group
cuts generation peak VRAM 32% (15951 -> 10840 MB) at near-resident speed (2.07 ->
2.99s); low_vram/model cuts it 48% (-> 8318 MB) but is slower (7.99s). Mode names
now match that tradeoff: balanced = stream the transformer, low_vram = offload
every component. auto picks group when the companions fit resident, else model.

112 CPU tests pass.

* Studio diffusion (Phase 5): image quality-vs-quant accuracy harness

Add scripts/diffusion_quality.py, the accuracy analogue of the KLD workflow: hold
prompt + seed fixed, render a grid with a reference quant (default BF16), then render
each candidate quant and measure drift from the reference. Records mean PSNR + SSIM
(pure-numpy, no skimage/scipy) and optional CLIP text-alignment + image-similarity
(transformers, --clip), plus file size, latency, and peak VRAM, then prints a
quality-vs-cost table and recommends the smallest quant within a quality budget.
--selftest validates the metrics on synthetic images with no GPU or model.

Verified on Z-Image (B200): the table degrades monotonically with quant size
(Q8 -> Q4 -> Q2: PSNR 21.7 -> 15.5, SSIM 0.82 -> 0.61), while CLIP-text stays flat
(~0.34) -- quantization erodes fine detail far more than prompt adherence.

* Studio diffusion (Phase 3): opt-in speed layer (channels_last / compile / TF32)

Add a speed_mode knob (off by default, so the render path stays bit-identical):
default applies channels_last VAE + regional torch.compile of the denoiser's
repeated block where eligible; max also enables TF32 matmul and fused QKV. Regional
compile is gated off for the GGUF transformer (dequantises per-op) and for families
flagged not compile-friendly (a new supports_torch_compile flag, False for Z-Image),
so it activates automatically only once a non-GGUF bf16 transformer is loaded. Speed
optims run before placement/offload, per the diffusers composition order. status now
reports speed_mode + the optims actually engaged.

Verified on Z-Image (B200): default -> ['channels_last'], max -> ['channels_last',
'tf32'], compile correctly skipped for GGUF; generation works in every mode.

121 CPU tests pass.

* Studio diffusion (Phase 2B): opt-in fp8 text-encoder layerwise casting

Add a text_encoder_fp8 knob that casts the companion text encoder(s) to fp8 (e4m3)
storage via diffusers apply_layerwise_casting, upcasting per layer to the bf16
compute dtype while normalisations and embeddings stay full precision. Applied
before placement, gated to CUDA + bf16, best-effort (a failure leaves the encoder
dense). status reports which encoders were cast.

Verified on Z-Image (B200, balanced/group mode where the encoder stays resident):
generation peak VRAM dropped 37% (10840 -> 6791 MB, below the lowest-VRAM offload)
at near-resident speed. It is a memory-vs-quality tradeoff, not free -- ~20 dB PSNR
vs the bf16 encoder, a larger shift than one transformer quant step -- so it is off
by default and documented as such, with the Phase 5 harness to size the cost.

127 CPU tests pass.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio diffusion (Phase 2C): NVFP4 text-encoder quant (+ generalise fp8 knob)

Generalise the text-encoder precision knob from a fp8 bool to text_encoder_quant
(fp8 | nvfp4). nvfp4 quantises the companion text encoder to 4-bit via torchao
NVFP4 weight-only (two-level microscaling) on Blackwell's FP4 tensor cores; fp8
stays the broader-hardware path (cc>=8.9). Both are gated, best-effort, and run
before placement; status reports the mode actually engaged. This is the lean
realisation of GGUF-native text-encoder quant: 4-bit on the encoder without the
3045-line port.

Verified on Z-Image (B200, balanced/group where the encoder stays resident), vs the
bf16 encoder: nvfp4 cut generation peak VRAM 48% (10840 -> 5593 MB, the lowest TE
option, below whole-model offload) at near-fp8 quality (16.4 vs 17.1 dB PSNR), and
both quants ran faster than bf16. A memory-vs-quality tradeoff (off by default);
size it per model with the Phase 5 quality harness. diffusion_bench gains
--text-encoder-quant.

129 CPU tests pass.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio diffusion (Phase 4): native stable-diffusion.cpp engine for CPU/Mac

Adds the CPU / Apple-Silicon tier of the two-engine strategy, mirroring the
chat backend's llama.cpp shell-out. Diffusers stays the default on CUDA / ROCm
/ XPU; this covers the hardware diffusers serves poorly, consuming the same
split GGUF assets Studio already curates.

- sd_cpp_args.py: pure sd-cli command builder. Maps the family to its
  text-encoder flag (Z-Image Qwen3 to --llm, Qwen-Image to --qwen2vl, FLUX.1
  CLIP-L + T5), and the diffusers memory policy (none/group/model/sequential)
  to sd.cpp's offload flags (--offload-to-cpu / --clip-on-cpu / --vae-on-cpu /
  --vae-tiling / --diffusion-fa), so one user knob drives both engines.
- sd_cpp_engine.py: SdCppEngine over a located sd-cli. find_sd_cpp_binary()
  with the same precedence as the llama finder (env override, then the Studio
  install root, then in-tree, then PATH), an is_available/version probe, and a
  one-shot subprocess generate that streams progress and returns the PNG.
  runtime_env() prepends the binary's directory to the platform library path
  so a prebuilt's bundled libstable-diffusion.so resolves.
  select_diffusion_engine() is the pure routing decision (GPU backends to
  diffusers, CPU/MPS to native when present).
- install_sd_cpp_prebuilt.py: resolve + download the per-host prebuilt
  (macOS-arm64/Metal, Linux x86_64 CPU, Vulkan/ROCm/Windows variants) into the
  Studio install root. resolve_release_asset() is a pure, unit-tested
  host-to-asset matrix.
- scripts/sd_cpp_smoke.py: end-to-end native generation harness.

Tests (CPU-only, subprocess/filesystem stubbed): 49 new across args, engine,
routing, runtime env, and the installer resolver. Full diffusion suite 166
passing.

Verified on a B200 box: built sd-cli (CUDA) and the prebuilt (CPU) both
generate Z-Image-Turbo Q4_K end to end through SdCppEngine: balanced (group
offload, 5.0s gen), low_vram (full CPU offload + VAE tiling, 13.4s), and the
dynamically-linked CPU prebuilt (50.4s on CPU), all producing coherent images.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio diffusion (Phase 4): enforce the sd-cli timeout while reading output

Iterating proc.stdout directly blocks until the stream closes, so a sd-cli that hangs
without producing output (or without closing stdout) would never reach proc.wait and the
wall-clock timeout was silently bypassed. Drain stdout on a daemon thread and wait on the
PROCESS, so the main thread always enforces the timeout and kills a hung process (which
closes the pipe and ends the reader). Add a test that times out even when stdout blocks,
and make the no-binary test hermetic so a host-installed sd-cli can't leak in.

* Studio diffusion (Phase 4) review fixes: sd.cpp installer + engine hardening

- install_sd_cpp_prebuilt: download the release archive with urlopen + an explicit
  timeout + copyfileobj (urlretrieve has no timeout and hangs on a stalled socket);
  extract through a per-member containment check (Zip-Slip guard); expanduser the
  --install-dir so a tilde path is not taken literally; and on Windows CUDA also fetch
  the separately-published cudart runtime DLL archive so sd-cli.exe can start.
- sd_cpp_engine: find_sd_cpp_binary honors UNSLOTH_STUDIO_HOME / STUDIO_HOME like the
  installer, so a custom-root install is discovered without UNSLOTH_SD_CPP_PATH; start
  sd-cli with the parent-death child_popen_kwargs so it is not orphaned on a backend
  crash; reap the SIGKILLed child (proc.wait) so a cancel/timeout does not leave a zombie.
- tests: Zip-Slip rejection, normal extraction, studio-home discovery.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio diffusion (Phase 4) review round 2: collect sd-cli batch outputs

Codex review: when batch_count > 1, stable-diffusion.cpp's save_results() writes
the numbered files <stem>_<idx><suffix> (base_0.png, base_1.png, ...) instead of
the literal --output path. SdCppEngine.generate checked only the literal path, so
a batch generation would exit 0 and then raise 'no image' (or return a stale
file). generate now returns the literal path when present and otherwise falls
back to the numbered siblings; single-image behavior is unchanged.

Test: a fake sd-cli that writes img_0.png/img_1.png (not img.png) is collected
without error.

---------

Co-authored-by: oobabooga <112222186+oobabooga@users.noreply.github.com>
2026-07-01 15:03:53 -03:00

324 lines
11 KiB
Python

# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""Hermetic, CPU-only tests for the diffusion device/dtype resolver.
`torch` is stubbed via a fake module so no GPU/torch is needed, and
`utils.hardware` is either stubbed (studio-layer path) or forced to fail
(torch-probe fallback path). Both paths are asserted.
"""
from __future__ import annotations
import sys
import types
import pytest
from core.inference import diffusion_device as dd
# ── Fakes ─────────────────────────────────────────────────────────────
class _FakeDtype:
def __init__(self, name: str) -> None:
self.name = name
def __eq__(self, other: object) -> bool:
return isinstance(other, _FakeDtype) and other.name == self.name
def __hash__(self) -> int:
return hash(self.name)
def __repr__(self) -> str: # str(dtype) -> "torch.bfloat16"
return f"torch.{self.name}"
BF16 = _FakeDtype("bfloat16")
FP16 = _FakeDtype("float16")
FP32 = _FakeDtype("float32")
class _FiniteResult:
def __init__(self, finite: bool) -> None:
self._finite = finite
def all(self) -> "_FiniteResult":
return self
def item(self) -> bool:
return self._finite
class _FakeTensor:
def __init__(self, finite: bool = True) -> None:
self._finite = finite
def __add__(self, other: object) -> "_FakeTensor":
return self
def float(self) -> "_FakeTensor":
return self
def _make_torch(
*,
cuda_available: bool = False,
capability = (8, 0),
capability_raises: bool = False,
bf16_supported: bool = False,
hip = None,
mps_available: bool = False,
mps_probe: str = "pass", # "pass" | "raise" | "nonfinite"
xpu_available = None, # None -> no xpu attr; True/False -> present
xpu_bf16: bool = False,
) -> types.ModuleType:
torch = types.ModuleType("torch")
torch.bfloat16 = BF16
torch.float16 = FP16
torch.float32 = FP32
torch.version = types.SimpleNamespace(hip = hip)
def _get_cap():
if capability_raises:
raise RuntimeError("no capability")
return capability
torch.cuda = types.SimpleNamespace(
is_available = lambda: cuda_available,
get_device_capability = _get_cap,
is_bf16_supported = lambda: bf16_supported,
)
mps_ns = types.SimpleNamespace(is_available = lambda: mps_available)
torch.backends = types.SimpleNamespace(mps = mps_ns)
def _ones(*_a, **_k):
if mps_probe == "raise":
raise RuntimeError("bf16 unsupported on this MPS")
return _FakeTensor(finite = (mps_probe == "pass"))
torch.ones = _ones
torch.isfinite = lambda t: _FiniteResult(getattr(t, "_finite", True))
if xpu_available is not None:
torch.xpu = types.SimpleNamespace(
is_available = lambda: xpu_available,
is_bf16_supported = lambda: xpu_bf16,
)
return torch
def _install(
monkeypatch,
torch,
*,
studio_device = None,
is_rocm = False,
hardware_fails = False,
):
"""Install the fake torch and either a fake or failing `utils.hardware`."""
monkeypatch.setitem(sys.modules, "torch", torch)
if hardware_fails:
# Force `from utils.hardware import ...` to raise -> torch-probe fallback.
monkeypatch.setitem(sys.modules, "utils.hardware", None)
return
class _DT:
CUDA = "cuda"
XPU = "xpu"
MLX = "mlx"
CPU = "cpu"
fake_uh = types.ModuleType("utils.hardware")
fake_uh.DeviceType = _DT
fake_uh.get_device = lambda: studio_device
fake_uh.hardware = types.SimpleNamespace(IS_ROCM = is_rocm)
monkeypatch.setitem(sys.modules, "utils.hardware", fake_uh)
# ── Studio-layer path ─────────────────────────────────────────────────
def test_cuda_ampere_bf16(monkeypatch):
torch = _make_torch(cuda_available = True, capability = (8, 0))
_install(monkeypatch, torch, studio_device = "cuda")
t = dd.resolve_diffusion_device_target()
assert (t.device, t.dtype, t.backend, t.vendor) == ("cuda", BF16, "cuda", "nvidia")
assert (
t.supports_model_cpu_offload
and t.supports_default_torch_compile
and t.supports_pinned_transfer
)
def test_cuda_pre_ampere_fp16(monkeypatch):
torch = _make_torch(cuda_available = True, capability = (7, 5), bf16_supported = True)
_install(monkeypatch, torch, studio_device = "cuda")
t = dd.resolve_diffusion_device_target()
# is_bf16_supported() is True (emulated) but capability < 8 -> fp16.
assert t.dtype == FP16 and t.backend == "cuda"
def test_cuda_capability_raises_falls_back_fp16(monkeypatch):
torch = _make_torch(cuda_available = True, capability_raises = True)
_install(monkeypatch, torch, studio_device = "cuda")
t = dd.resolve_diffusion_device_target()
assert t.dtype == FP16 and t.device == "cuda"
def test_cuda_studio_says_cuda_but_unavailable_is_cpu(monkeypatch):
torch = _make_torch(cuda_available = False)
_install(monkeypatch, torch, studio_device = "cuda")
t = dd.resolve_diffusion_device_target()
assert t.device == "cpu" and t.dtype == FP32
def test_rocm_target(monkeypatch):
torch = _make_torch(cuda_available = True, bf16_supported = True)
_install(monkeypatch, torch, studio_device = "cuda", is_rocm = True)
t = dd.resolve_diffusion_device_target()
assert (t.device, t.backend, t.vendor) == ("cuda", "rocm", "amd")
assert t.dtype == BF16
assert t.supports_default_torch_compile is False # ROCm disables default compile
def test_rocm_without_bf16_uses_fp16(monkeypatch):
torch = _make_torch(cuda_available = True, bf16_supported = False)
_install(monkeypatch, torch, studio_device = "cuda", is_rocm = True)
t = dd.resolve_diffusion_device_target()
assert t.dtype == FP16 and t.backend == "rocm"
def test_xpu_bf16(monkeypatch):
torch = _make_torch(xpu_available = True, xpu_bf16 = True)
_install(monkeypatch, torch, studio_device = "xpu")
t = dd.resolve_diffusion_device_target()
assert (t.device, t.backend, t.vendor, t.dtype) == ("xpu", "xpu", "intel", BF16)
assert (
t.supports_model_cpu_offload
and not t.supports_default_torch_compile
and not t.supports_pinned_transfer
)
def test_xpu_without_bf16_fp16(monkeypatch):
torch = _make_torch(xpu_available = True, xpu_bf16 = False)
_install(monkeypatch, torch, studio_device = "xpu")
t = dd.resolve_diffusion_device_target()
assert t.device == "xpu" and t.dtype == FP16
def test_mps_probe_pass_bf16(monkeypatch):
torch = _make_torch(mps_available = True, mps_probe = "pass")
_install(monkeypatch, torch, studio_device = "mlx")
t = dd.resolve_diffusion_device_target()
assert (t.device, t.backend, t.vendor, t.dtype) == ("mps", "mps", "apple", BF16)
assert not t.supports_model_cpu_offload
def test_mps_probe_raises_uses_fp32_not_fp16(monkeypatch):
torch = _make_torch(mps_available = True, mps_probe = "raise")
_install(monkeypatch, torch, studio_device = "mlx")
t = dd.resolve_diffusion_device_target()
assert t.device == "mps" and t.dtype == FP32 # strict: never silent fp16
def test_mps_probe_nonfinite_uses_fp32(monkeypatch):
torch = _make_torch(mps_available = True, mps_probe = "nonfinite")
_install(monkeypatch, torch, studio_device = "mlx")
t = dd.resolve_diffusion_device_target()
assert t.device == "mps" and t.dtype == FP32
def test_studio_cpu_on_apple_prefers_mps(monkeypatch):
torch = _make_torch(mps_available = True, mps_probe = "pass")
_install(monkeypatch, torch, studio_device = "cpu") # Studio reports CPU (no mlx pkg)
t = dd.resolve_diffusion_device_target()
assert t.device == "mps" and t.dtype == BF16
def test_cpu_when_nothing_available(monkeypatch):
torch = _make_torch(mps_available = False)
_install(monkeypatch, torch, studio_device = "cpu")
t = dd.resolve_diffusion_device_target()
assert (t.device, t.backend, t.vendor, t.dtype) == ("cpu", "cpu", None, FP32)
assert not any(
(t.supports_model_cpu_offload, t.supports_default_torch_compile, t.supports_pinned_transfer)
)
# ── torch-probe fallback path (utils.hardware import fails) ────────────
def test_fallback_cuda(monkeypatch):
torch = _make_torch(cuda_available = True, capability = (9, 0))
_install(monkeypatch, torch, hardware_fails = True)
t = dd.resolve_diffusion_device_target()
assert t.device == "cuda" and t.dtype == BF16 and t.backend == "cuda"
def test_fallback_rocm_via_torch_hip(monkeypatch):
torch = _make_torch(cuda_available = True, bf16_supported = True, hip = "6.2")
_install(monkeypatch, torch, hardware_fails = True)
t = dd.resolve_diffusion_device_target()
assert t.backend == "rocm" and t.vendor == "amd"
def test_fallback_xpu(monkeypatch):
torch = _make_torch(cuda_available = False, xpu_available = True, xpu_bf16 = True)
_install(monkeypatch, torch, hardware_fails = True)
t = dd.resolve_diffusion_device_target()
assert t.device == "xpu" and t.dtype == BF16
def test_fallback_mps(monkeypatch):
torch = _make_torch(cuda_available = False, mps_available = True, mps_probe = "pass")
_install(monkeypatch, torch, hardware_fails = True)
t = dd.resolve_diffusion_device_target()
assert t.device == "mps" and t.dtype == BF16
def test_fallback_cpu(monkeypatch):
torch = _make_torch(cuda_available = False, mps_available = False)
_install(monkeypatch, torch, hardware_fails = True)
t = dd.resolve_diffusion_device_target()
assert t.device == "cpu" and t.dtype == FP32
# ── from-torch-device reconstruction + public dict ────────────────────
def test_from_torch_device_cuda(monkeypatch):
torch = _make_torch()
monkeypatch.setitem(sys.modules, "torch", torch)
t = dd.diffusion_device_target_from_torch_device("cuda:0", FP32)
assert (t.device, t.backend, t.vendor, t.dtype) == ("cuda", "cuda", "nvidia", FP32)
assert t.is_cuda_torch_device
def test_from_torch_device_mps_and_cpu(monkeypatch):
torch = _make_torch()
monkeypatch.setitem(sys.modules, "torch", torch)
mps = dd.diffusion_device_target_from_torch_device("mps", FP16)
assert mps.device == "mps" and not mps.supports_model_cpu_offload
cpu = dd.diffusion_device_target_from_torch_device("cpu", FP32)
assert cpu.device == "cpu" and cpu.vendor is None
@pytest.mark.parametrize(
"dtype,expected", [(BF16, "bfloat16"), (FP16, "float16"), (FP32, "float32")]
)
def test_public_dict_dtype_string(dtype, expected):
t = dd.DiffusionDeviceTarget(
device = "cuda",
dtype = dtype,
backend = "cuda",
vendor = "nvidia",
supports_model_cpu_offload = True,
supports_default_torch_compile = True,
supports_pinned_transfer = True,
)
d = t.as_public_dict()
assert d["dtype"] == expected and "torch." not in d["dtype"]