* Studio diffusion: cross-platform device policy, fp16 guard, lock split, validate-before-evict Phase 1 of porting the richer diffusion stack onto the image-generation backend. - Add a compartmentalized device/dtype policy module (diffusion_device.py) resolving CUDA/ROCm/XPU/MPS/CPU with capability flags. Keeps the NVIDIA capability-based bf16 choice; ROCm and XPU are isolated; MPS uses bf16 or fp32, never a silent fp16 that renders a black image. - Add a per-family fp16_incompatible flag (Z-Image) and promote a resolved float16 to float32 for those families so they do not produce black images. - Split the backend locks: a generation holds only _generate_lock, so status, unload, and a new load are never blocked by a long denoise. Add per-generation cancellation via callback_on_step_end so an eviction or a superseding load preempts a running generation; a replacement load waits for it to stop before allocating, so two pipelines never sit in VRAM at once. - Validate a load request before the GPU handoff so an unloadable pick never evicts a working chat model, and reject missing local paths up front. - Add CPU-only tests for the device policy, dtype guard, lock split and cancellation, and validate-before-evict, plus a GPU benchmark/regression script (scripts/diffusion_bench.py) measuring latency, peak VRAM, and PSNR against a saved reference. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio diffusion (Phase 2A): measured-budget memory planner + offload/VAE policy Add a lean, backend-agnostic memory policy that picks a CPU-offload policy and VAE tiling/slicing from measured free device memory vs the model's estimated resident footprint, then applies it to the built pipeline. auto stays resident when the model fits (byte-identical to the prior resident path), and falls to whole-module offload when tight; fast/balanced/low_vram are explicit overrides. Sequential submodule offload is unreliable for GGUF transformers on diffusers 0.38, so it falls back to whole-module offload and status reports the policy actually engaged. Verified on Z-Image-Turbo Q4_K_M (B200): auto reproduces the resident image with no VRAM/latency regression (PSNR inf); balanced/low_vram cut generation peak VRAM 47.9% (15951 -> 8318 MB) with byte-identical output, at the expected latency cost. 73 prior + 35 new CPU tests pass. * Studio diffusion (Phase 2D): streamed block-level offload + functional VAE tiling Add a streamed 'group' offload tier (diffusers apply_group_offloading, block_level, use_stream) that keeps the transformer flowing through the GPU a few blocks at a time while the text encoder / VAE stay resident, and fix VAE tiling to drive the VAE submodule (pipelines like Z-Image expose enable_tiling on pipe.vae, not the pipeline). apply_memory_plan now returns the (policy, tiling) actually engaged so status never overstates either, and group falls back to whole-module offload when the transformer can't be streamed. Measured on Z-Image (B200), all lossless (PSNR inf vs resident): balanced/group cuts generation peak VRAM 32% (15951 -> 10840 MB) at near-resident speed (2.07 -> 2.99s); low_vram/model cuts it 48% (-> 8318 MB) but is slower (7.99s). Mode names now match that tradeoff: balanced = stream the transformer, low_vram = offload every component. auto picks group when the companions fit resident, else model. 112 CPU tests pass. * Studio diffusion (Phase 5): image quality-vs-quant accuracy harness Add scripts/diffusion_quality.py, the accuracy analogue of the KLD workflow: hold prompt + seed fixed, render a grid with a reference quant (default BF16), then render each candidate quant and measure drift from the reference. Records mean PSNR + SSIM (pure-numpy, no skimage/scipy) and optional CLIP text-alignment + image-similarity (transformers, --clip), plus file size, latency, and peak VRAM, then prints a quality-vs-cost table and recommends the smallest quant within a quality budget. --selftest validates the metrics on synthetic images with no GPU or model. Verified on Z-Image (B200): the table degrades monotonically with quant size (Q8 -> Q4 -> Q2: PSNR 21.7 -> 15.5, SSIM 0.82 -> 0.61), while CLIP-text stays flat (~0.34) -- quantization erodes fine detail far more than prompt adherence. * Studio diffusion (Phase 3): opt-in speed layer (channels_last / compile / TF32) Add a speed_mode knob (off by default, so the render path stays bit-identical): default applies channels_last VAE + regional torch.compile of the denoiser's repeated block where eligible; max also enables TF32 matmul and fused QKV. Regional compile is gated off for the GGUF transformer (dequantises per-op) and for families flagged not compile-friendly (a new supports_torch_compile flag, False for Z-Image), so it activates automatically only once a non-GGUF bf16 transformer is loaded. Speed optims run before placement/offload, per the diffusers composition order. status now reports speed_mode + the optims actually engaged. Verified on Z-Image (B200): default -> ['channels_last'], max -> ['channels_last', 'tf32'], compile correctly skipped for GGUF; generation works in every mode. 121 CPU tests pass. * Studio diffusion (Phase 2B): opt-in fp8 text-encoder layerwise casting Add a text_encoder_fp8 knob that casts the companion text encoder(s) to fp8 (e4m3) storage via diffusers apply_layerwise_casting, upcasting per layer to the bf16 compute dtype while normalisations and embeddings stay full precision. Applied before placement, gated to CUDA + bf16, best-effort (a failure leaves the encoder dense). status reports which encoders were cast. Verified on Z-Image (B200, balanced/group mode where the encoder stays resident): generation peak VRAM dropped 37% (10840 -> 6791 MB, below the lowest-VRAM offload) at near-resident speed. It is a memory-vs-quality tradeoff, not free -- ~20 dB PSNR vs the bf16 encoder, a larger shift than one transformer quant step -- so it is off by default and documented as such, with the Phase 5 harness to size the cost. 127 CPU tests pass. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio diffusion (Phase 2C): NVFP4 text-encoder quant (+ generalise fp8 knob) Generalise the text-encoder precision knob from a fp8 bool to text_encoder_quant (fp8 | nvfp4). nvfp4 quantises the companion text encoder to 4-bit via torchao NVFP4 weight-only (two-level microscaling) on Blackwell's FP4 tensor cores; fp8 stays the broader-hardware path (cc>=8.9). Both are gated, best-effort, and run before placement; status reports the mode actually engaged. This is the lean realisation of GGUF-native text-encoder quant: 4-bit on the encoder without the 3045-line port. Verified on Z-Image (B200, balanced/group where the encoder stays resident), vs the bf16 encoder: nvfp4 cut generation peak VRAM 48% (10840 -> 5593 MB, the lowest TE option, below whole-model offload) at near-fp8 quality (16.4 vs 17.1 dB PSNR), and both quants ran faster than bf16. A memory-vs-quality tradeoff (off by default); size it per model with the Phase 5 quality harness. diffusion_bench gains --text-encoder-quant. 129 CPU tests pass. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio diffusion (Phase 4): native stable-diffusion.cpp engine for CPU/Mac Adds the CPU / Apple-Silicon tier of the two-engine strategy, mirroring the chat backend's llama.cpp shell-out. Diffusers stays the default on CUDA / ROCm / XPU; this covers the hardware diffusers serves poorly, consuming the same split GGUF assets Studio already curates. - sd_cpp_args.py: pure sd-cli command builder. Maps the family to its text-encoder flag (Z-Image Qwen3 to --llm, Qwen-Image to --qwen2vl, FLUX.1 CLIP-L + T5), and the diffusers memory policy (none/group/model/sequential) to sd.cpp's offload flags (--offload-to-cpu / --clip-on-cpu / --vae-on-cpu / --vae-tiling / --diffusion-fa), so one user knob drives both engines. - sd_cpp_engine.py: SdCppEngine over a located sd-cli. find_sd_cpp_binary() with the same precedence as the llama finder (env override, then the Studio install root, then in-tree, then PATH), an is_available/version probe, and a one-shot subprocess generate that streams progress and returns the PNG. runtime_env() prepends the binary's directory to the platform library path so a prebuilt's bundled libstable-diffusion.so resolves. select_diffusion_engine() is the pure routing decision (GPU backends to diffusers, CPU/MPS to native when present). - install_sd_cpp_prebuilt.py: resolve + download the per-host prebuilt (macOS-arm64/Metal, Linux x86_64 CPU, Vulkan/ROCm/Windows variants) into the Studio install root. resolve_release_asset() is a pure, unit-tested host-to-asset matrix. - scripts/sd_cpp_smoke.py: end-to-end native generation harness. Tests (CPU-only, subprocess/filesystem stubbed): 49 new across args, engine, routing, runtime env, and the installer resolver. Full diffusion suite 166 passing. Verified on a B200 box: built sd-cli (CUDA) and the prebuilt (CPU) both generate Z-Image-Turbo Q4_K end to end through SdCppEngine: balanced (group offload, 5.0s gen), low_vram (full CPU offload + VAE tiling, 13.4s), and the dynamically-linked CPU prebuilt (50.4s on CPU), all producing coherent images. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio diffusion (Phase 4): enforce the sd-cli timeout while reading output Iterating proc.stdout directly blocks until the stream closes, so a sd-cli that hangs without producing output (or without closing stdout) would never reach proc.wait and the wall-clock timeout was silently bypassed. Drain stdout on a daemon thread and wait on the PROCESS, so the main thread always enforces the timeout and kills a hung process (which closes the pipe and ends the reader). Add a test that times out even when stdout blocks, and make the no-binary test hermetic so a host-installed sd-cli can't leak in. * Studio diffusion (Phase 4) review fixes: sd.cpp installer + engine hardening - install_sd_cpp_prebuilt: download the release archive with urlopen + an explicit timeout + copyfileobj (urlretrieve has no timeout and hangs on a stalled socket); extract through a per-member containment check (Zip-Slip guard); expanduser the --install-dir so a tilde path is not taken literally; and on Windows CUDA also fetch the separately-published cudart runtime DLL archive so sd-cli.exe can start. - sd_cpp_engine: find_sd_cpp_binary honors UNSLOTH_STUDIO_HOME / STUDIO_HOME like the installer, so a custom-root install is discovered without UNSLOTH_SD_CPP_PATH; start sd-cli with the parent-death child_popen_kwargs so it is not orphaned on a backend crash; reap the SIGKILLed child (proc.wait) so a cancel/timeout does not leave a zombie. - tests: Zip-Slip rejection, normal extraction, studio-home discovery. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio diffusion (Phase 4) review round 2: collect sd-cli batch outputs Codex review: when batch_count > 1, stable-diffusion.cpp's save_results() writes the numbered files <stem>_<idx><suffix> (base_0.png, base_1.png, ...) instead of the literal --output path. SdCppEngine.generate checked only the literal path, so a batch generation would exit 0 and then raise 'no image' (or return a stale file). generate now returns the literal path when present and otherwise falls back to the numbered siblings; single-image behavior is unchanged. Test: a fake sd-cli that writes img_0.png/img_1.png (not img.png) is collected without error. --------- Co-authored-by: oobabooga <112222186+oobabooga@users.noreply.github.com>
166 lines
6.2 KiB
Python
166 lines
6.2 KiB
Python
# SPDX-License-Identifier: AGPL-3.0-only
|
|
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
|
|
|
"""Unit tests for the prebuilt sd-cli asset resolver (``install_sd_cpp_prebuilt``).
|
|
|
|
Pure: the host -> release-asset matrix is exercised against a fixed asset list
|
|
(a real stable-diffusion.cpp release), no network. The installer lives under
|
|
``studio/`` (not ``studio/backend``), so the test puts that dir on the path.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import sys
|
|
from pathlib import Path
|
|
|
|
_STUDIO = Path(__file__).resolve().parents[2]
|
|
if str(_STUDIO) not in sys.path:
|
|
sys.path.insert(0, str(_STUDIO))
|
|
|
|
import zipfile # noqa: E402
|
|
|
|
import pytest # noqa: E402
|
|
|
|
from install_sd_cpp_prebuilt import ( # noqa: E402
|
|
_safe_extractall,
|
|
default_install_dir,
|
|
resolve_release_asset,
|
|
)
|
|
|
|
# A real stable-diffusion.cpp latest-release asset list.
|
|
_ASSETS = [
|
|
"cudart-sd-bin-win-cu12-x64.zip",
|
|
"sd-master-8caa3f9-bin-Darwin-macOS-15.7.7-arm64.zip",
|
|
"sd-master-8caa3f9-bin-Linux-Ubuntu-24.04-x86_64-rocm-7.13.0.zip",
|
|
"sd-master-8caa3f9-bin-Linux-Ubuntu-24.04-x86_64-rocm-7.2.1.zip",
|
|
"sd-master-8caa3f9-bin-Linux-Ubuntu-24.04-x86_64-vulkan.zip",
|
|
"sd-master-8caa3f9-bin-Linux-Ubuntu-24.04-x86_64.zip",
|
|
"sd-master-8caa3f9-bin-win-avx-x64.zip",
|
|
"sd-master-8caa3f9-bin-win-avx2-x64.zip",
|
|
"sd-master-8caa3f9-bin-win-avx512-x64.zip",
|
|
"sd-master-8caa3f9-bin-win-cuda12-x64.zip",
|
|
"sd-master-8caa3f9-bin-win-noavx-x64.zip",
|
|
"sd-master-8caa3f9-bin-win-rocm-7.13.0-x64.zip",
|
|
"sd-master-8caa3f9-bin-win-vulkan-x64.zip",
|
|
]
|
|
|
|
|
|
def _resolve(
|
|
system,
|
|
machine,
|
|
accelerator = "auto",
|
|
):
|
|
return resolve_release_asset(_ASSETS, system = system, machine = machine, accelerator = accelerator)
|
|
|
|
|
|
# ── macOS (the key Apple-Silicon target) ────────────────────────────────────
|
|
|
|
|
|
def test_macos_arm64_picks_darwin_arm64():
|
|
assert _resolve("Darwin", "arm64") == "sd-master-8caa3f9-bin-Darwin-macOS-15.7.7-arm64.zip"
|
|
# aarch64 spelling resolves the same
|
|
assert _resolve("Darwin", "aarch64").startswith("sd-master") and "arm64" in _resolve(
|
|
"Darwin", "aarch64"
|
|
)
|
|
|
|
|
|
def test_macos_intel_has_no_prebuilt():
|
|
# only an arm64 Darwin asset exists -> Intel Macs must build from source
|
|
assert _resolve("Darwin", "x86_64") is None
|
|
|
|
|
|
# ── Linux (CPU is the default tier) ─────────────────────────────────────────
|
|
|
|
|
|
def test_linux_x86_64_auto_picks_plain_cpu_build():
|
|
# the plain x86_64 zip, NOT a rocm/vulkan one
|
|
assert _resolve("Linux", "x86_64") == "sd-master-8caa3f9-bin-Linux-Ubuntu-24.04-x86_64.zip"
|
|
|
|
|
|
def test_linux_vulkan_and_rocm_select_accelerator_builds():
|
|
assert (
|
|
_resolve("Linux", "x86_64", "vulkan")
|
|
== "sd-master-8caa3f9-bin-Linux-Ubuntu-24.04-x86_64-vulkan.zip"
|
|
)
|
|
assert "rocm" in _resolve("Linux", "x86_64", "rocm")
|
|
|
|
|
|
def test_linux_arm64_has_no_prebuilt():
|
|
assert _resolve("Linux", "aarch64") is None
|
|
|
|
|
|
# ── Windows ─────────────────────────────────────────────────────────────────
|
|
|
|
|
|
def test_windows_auto_picks_avx2():
|
|
assert _resolve("Windows", "AMD64") == "sd-master-8caa3f9-bin-win-avx2-x64.zip"
|
|
|
|
|
|
def test_windows_cuda_picks_cuda12():
|
|
assert _resolve("Windows", "AMD64", "cuda") == "sd-master-8caa3f9-bin-win-cuda12-x64.zip"
|
|
|
|
|
|
def test_windows_vulkan_picks_vulkan():
|
|
assert _resolve("Windows", "AMD64", "vulkan") == "sd-master-8caa3f9-bin-win-vulkan-x64.zip"
|
|
|
|
|
|
# ── cudart helper archive is never chosen as the engine ─────────────────────
|
|
|
|
|
|
def test_cudart_runtime_archive_never_selected():
|
|
for accel in ("auto", "cuda", "vulkan", "rocm"):
|
|
chosen = _resolve("Windows", "AMD64", accel)
|
|
assert chosen is None or not chosen.startswith("cudart")
|
|
|
|
|
|
# ── install dir ─────────────────────────────────────────────────────────────
|
|
|
|
|
|
def test_default_install_dir_is_sibling_of_llama(monkeypatch):
|
|
monkeypatch.delenv("UNSLOTH_STUDIO_HOME", raising = False)
|
|
monkeypatch.delenv("STUDIO_HOME", raising = False)
|
|
d = default_install_dir()
|
|
assert d.name == "stable-diffusion.cpp"
|
|
assert d.parent.name == ".unsloth"
|
|
|
|
|
|
# ── safe extraction (Zip-Slip guard) ─────────────────────────────────────────
|
|
|
|
|
|
def test_safe_extractall_rejects_path_traversal(tmp_path):
|
|
target = tmp_path / "install"
|
|
target.mkdir()
|
|
archive = tmp_path / "evil.zip"
|
|
with zipfile.ZipFile(archive, "w") as zf:
|
|
zf.writestr("sd-cli", b"ok")
|
|
zf.writestr("../escape.txt", b"pwned") # escapes the install dir
|
|
with zipfile.ZipFile(archive) as zf:
|
|
with pytest.raises(RuntimeError, match = "unsafe path"):
|
|
_safe_extractall(zf, target)
|
|
assert not (tmp_path / "escape.txt").exists()
|
|
|
|
|
|
def test_safe_extractall_extracts_normal_members(tmp_path):
|
|
target = tmp_path / "install"
|
|
target.mkdir()
|
|
archive = tmp_path / "ok.zip"
|
|
with zipfile.ZipFile(archive, "w") as zf:
|
|
zf.writestr("build/bin/sd-cli", b"ok")
|
|
with zipfile.ZipFile(archive) as zf:
|
|
_safe_extractall(zf, target)
|
|
assert (target / "build" / "bin" / "sd-cli").read_bytes() == b"ok"
|
|
|
|
|
|
def test_find_sd_cpp_binary_honors_studio_home(tmp_path, monkeypatch):
|
|
# A binary installed under a custom Studio root must be discovered without also
|
|
# setting UNSLOTH_SD_CPP_PATH (matches default_install_dir's env handling).
|
|
from core.inference import sd_cpp_engine as eng
|
|
|
|
monkeypatch.delenv("SD_CLI_PATH", raising = False)
|
|
monkeypatch.delenv("UNSLOTH_SD_CPP_PATH", raising = False)
|
|
studio_home = tmp_path / "studio_root"
|
|
monkeypatch.setenv("UNSLOTH_STUDIO_HOME", str(studio_home))
|
|
binary = tmp_path / "stable-diffusion.cpp" / "build" / "bin" / "sd-cli"
|
|
binary.parent.mkdir(parents = True)
|
|
binary.write_bytes(b"x")
|
|
assert eng.find_sd_cpp_binary() == str(binary)
|