From 2d09508951c41b1e170b35ee8ea6ed3f80544b2b Mon Sep 17 00:00:00 2001 From: Daniel Han Date: Wed, 1 Jul 2026 11:18:38 -0700 Subject: [PATCH] Studio diffusion (Phase 6): img2img / inpaint / edit / LoRA / upscale on the native engine (#6680) * Studio diffusion: cross-platform device policy, fp16 guard, lock split, validate-before-evict Phase 1 of porting the richer diffusion stack onto the image-generation backend. - Add a compartmentalized device/dtype policy module (diffusion_device.py) resolving CUDA/ROCm/XPU/MPS/CPU with capability flags. Keeps the NVIDIA capability-based bf16 choice; ROCm and XPU are isolated; MPS uses bf16 or fp32, never a silent fp16 that renders a black image. - Add a per-family fp16_incompatible flag (Z-Image) and promote a resolved float16 to float32 for those families so they do not produce black images. - Split the backend locks: a generation holds only _generate_lock, so status, unload, and a new load are never blocked by a long denoise. Add per-generation cancellation via callback_on_step_end so an eviction or a superseding load preempts a running generation; a replacement load waits for it to stop before allocating, so two pipelines never sit in VRAM at once. - Validate a load request before the GPU handoff so an unloadable pick never evicts a working chat model, and reject missing local paths up front. - Add CPU-only tests for the device policy, dtype guard, lock split and cancellation, and validate-before-evict, plus a GPU benchmark/regression script (scripts/diffusion_bench.py) measuring latency, peak VRAM, and PSNR against a saved reference. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio diffusion (Phase 2A): measured-budget memory planner + offload/VAE policy Add a lean, backend-agnostic memory policy that picks a CPU-offload policy and VAE tiling/slicing from measured free device memory vs the model's estimated resident footprint, then applies it to the built pipeline. auto stays resident when the model fits (byte-identical to the prior resident path), and falls to whole-module offload when tight; fast/balanced/low_vram are explicit overrides. Sequential submodule offload is unreliable for GGUF transformers on diffusers 0.38, so it falls back to whole-module offload and status reports the policy actually engaged. Verified on Z-Image-Turbo Q4_K_M (B200): auto reproduces the resident image with no VRAM/latency regression (PSNR inf); balanced/low_vram cut generation peak VRAM 47.9% (15951 -> 8318 MB) with byte-identical output, at the expected latency cost. 73 prior + 35 new CPU tests pass. * Studio diffusion (Phase 2D): streamed block-level offload + functional VAE tiling Add a streamed 'group' offload tier (diffusers apply_group_offloading, block_level, use_stream) that keeps the transformer flowing through the GPU a few blocks at a time while the text encoder / VAE stay resident, and fix VAE tiling to drive the VAE submodule (pipelines like Z-Image expose enable_tiling on pipe.vae, not the pipeline). apply_memory_plan now returns the (policy, tiling) actually engaged so status never overstates either, and group falls back to whole-module offload when the transformer can't be streamed. Measured on Z-Image (B200), all lossless (PSNR inf vs resident): balanced/group cuts generation peak VRAM 32% (15951 -> 10840 MB) at near-resident speed (2.07 -> 2.99s); low_vram/model cuts it 48% (-> 8318 MB) but is slower (7.99s). Mode names now match that tradeoff: balanced = stream the transformer, low_vram = offload every component. auto picks group when the companions fit resident, else model. 112 CPU tests pass. * Studio diffusion (Phase 5): image quality-vs-quant accuracy harness Add scripts/diffusion_quality.py, the accuracy analogue of the KLD workflow: hold prompt + seed fixed, render a grid with a reference quant (default BF16), then render each candidate quant and measure drift from the reference. Records mean PSNR + SSIM (pure-numpy, no skimage/scipy) and optional CLIP text-alignment + image-similarity (transformers, --clip), plus file size, latency, and peak VRAM, then prints a quality-vs-cost table and recommends the smallest quant within a quality budget. --selftest validates the metrics on synthetic images with no GPU or model. Verified on Z-Image (B200): the table degrades monotonically with quant size (Q8 -> Q4 -> Q2: PSNR 21.7 -> 15.5, SSIM 0.82 -> 0.61), while CLIP-text stays flat (~0.34) -- quantization erodes fine detail far more than prompt adherence. * Studio diffusion (Phase 3): opt-in speed layer (channels_last / compile / TF32) Add a speed_mode knob (off by default, so the render path stays bit-identical): default applies channels_last VAE + regional torch.compile of the denoiser's repeated block where eligible; max also enables TF32 matmul and fused QKV. Regional compile is gated off for the GGUF transformer (dequantises per-op) and for families flagged not compile-friendly (a new supports_torch_compile flag, False for Z-Image), so it activates automatically only once a non-GGUF bf16 transformer is loaded. Speed optims run before placement/offload, per the diffusers composition order. status now reports speed_mode + the optims actually engaged. Verified on Z-Image (B200): default -> ['channels_last'], max -> ['channels_last', 'tf32'], compile correctly skipped for GGUF; generation works in every mode. 121 CPU tests pass. * Studio diffusion (Phase 2B): opt-in fp8 text-encoder layerwise casting Add a text_encoder_fp8 knob that casts the companion text encoder(s) to fp8 (e4m3) storage via diffusers apply_layerwise_casting, upcasting per layer to the bf16 compute dtype while normalisations and embeddings stay full precision. Applied before placement, gated to CUDA + bf16, best-effort (a failure leaves the encoder dense). status reports which encoders were cast. Verified on Z-Image (B200, balanced/group mode where the encoder stays resident): generation peak VRAM dropped 37% (10840 -> 6791 MB, below the lowest-VRAM offload) at near-resident speed. It is a memory-vs-quality tradeoff, not free -- ~20 dB PSNR vs the bf16 encoder, a larger shift than one transformer quant step -- so it is off by default and documented as such, with the Phase 5 harness to size the cost. 127 CPU tests pass. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio diffusion (Phase 2C): NVFP4 text-encoder quant (+ generalise fp8 knob) Generalise the text-encoder precision knob from a fp8 bool to text_encoder_quant (fp8 | nvfp4). nvfp4 quantises the companion text encoder to 4-bit via torchao NVFP4 weight-only (two-level microscaling) on Blackwell's FP4 tensor cores; fp8 stays the broader-hardware path (cc>=8.9). Both are gated, best-effort, and run before placement; status reports the mode actually engaged. This is the lean realisation of GGUF-native text-encoder quant: 4-bit on the encoder without the 3045-line port. Verified on Z-Image (B200, balanced/group where the encoder stays resident), vs the bf16 encoder: nvfp4 cut generation peak VRAM 48% (10840 -> 5593 MB, the lowest TE option, below whole-model offload) at near-fp8 quality (16.4 vs 17.1 dB PSNR), and both quants ran faster than bf16. A memory-vs-quality tradeoff (off by default); size it per model with the Phase 5 quality harness. diffusion_bench gains --text-encoder-quant. 129 CPU tests pass. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio diffusion (Phase 4): native stable-diffusion.cpp engine for CPU/Mac Adds the CPU / Apple-Silicon tier of the two-engine strategy, mirroring the chat backend's llama.cpp shell-out. Diffusers stays the default on CUDA / ROCm / XPU; this covers the hardware diffusers serves poorly, consuming the same split GGUF assets Studio already curates. - sd_cpp_args.py: pure sd-cli command builder. Maps the family to its text-encoder flag (Z-Image Qwen3 to --llm, Qwen-Image to --qwen2vl, FLUX.1 CLIP-L + T5), and the diffusers memory policy (none/group/model/sequential) to sd.cpp's offload flags (--offload-to-cpu / --clip-on-cpu / --vae-on-cpu / --vae-tiling / --diffusion-fa), so one user knob drives both engines. - sd_cpp_engine.py: SdCppEngine over a located sd-cli. find_sd_cpp_binary() with the same precedence as the llama finder (env override, then the Studio install root, then in-tree, then PATH), an is_available/version probe, and a one-shot subprocess generate that streams progress and returns the PNG. runtime_env() prepends the binary's directory to the platform library path so a prebuilt's bundled libstable-diffusion.so resolves. select_diffusion_engine() is the pure routing decision (GPU backends to diffusers, CPU/MPS to native when present). - install_sd_cpp_prebuilt.py: resolve + download the per-host prebuilt (macOS-arm64/Metal, Linux x86_64 CPU, Vulkan/ROCm/Windows variants) into the Studio install root. resolve_release_asset() is a pure, unit-tested host-to-asset matrix. - scripts/sd_cpp_smoke.py: end-to-end native generation harness. Tests (CPU-only, subprocess/filesystem stubbed): 49 new across args, engine, routing, runtime env, and the installer resolver. Full diffusion suite 166 passing. Verified on a B200 box: built sd-cli (CUDA) and the prebuilt (CPU) both generate Z-Image-Turbo Q4_K end to end through SdCppEngine: balanced (group offload, 5.0s gen), low_vram (full CPU offload + VAE tiling, 13.4s), and the dynamically-linked CPU prebuilt (50.4s on CPU), all producing coherent images. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio diffusion (Phase 6): img2img / inpaint / edit / LoRA / upscale on the native engine Builds on Phase 4's native stable-diffusion.cpp engine, extending it from text-to-image to the wider feature surface, since sd.cpp supports all of these through the binary already. Pure command-builder additions plus one engine method, so the txt2img path is unchanged. - sd_cpp_args.py: SdCppGenParams gains image-conditioning fields. init_img + strength make a run img2img, adding mask makes it inpaint, ref_images drives FLUX-Kontext / Qwen-Image-Edit style editing (repeated --ref-image), and lora_dir + the prompt syntax select LoRAs. New SdCppUpscaleParams + build_sd_cpp_upscale_command for the ESRGAN upscale run mode (input image + esrgan model, no prompt / text encoders). - sd_cpp_engine.py: the subprocess runner is factored into a shared _run() so generate() (now carrying the conditioning flags) and a new upscale() reuse the same streaming / error / output-check path. - scripts/sd_cpp_smoke.py: --task {txt2img,img2img,upscale} with --init-img / --strength / --upscale-model / --upscale-repeats. Tests: 10 new across the img2img / inpaint / edit / LoRA flag construction, the upscale builder and its validation, and the engine's img2img + upscale paths. Full diffusion suite 176 passing. Verified on a B200 box through SdCppEngine: img2img (Z-Image-Turbo Q4_K, the init image conditioned at strength 0.6, 4.8s) and ESRGAN upscale (512x512 -> 2048x2048 via RealESRGAN_x4plus_anime_6B, 2.7s), both producing coherent images. Video and the diffusers-path feature wiring are deferred. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio diffusion (Phase 4): enforce the sd-cli timeout while reading output Iterating proc.stdout directly blocks until the stream closes, so a sd-cli that hangs without producing output (or without closing stdout) would never reach proc.wait and the wall-clock timeout was silently bypassed. Drain stdout on a daemon thread and wait on the PROCESS, so the main thread always enforces the timeout and kills a hung process (which closes the pipe and ends the reader). Add a test that times out even when stdout blocks, and make the no-binary test hermetic so a host-installed sd-cli can't leak in. * Studio diffusion (Phase 4) review fixes: sd.cpp installer + engine hardening - install_sd_cpp_prebuilt: download the release archive with urlopen + an explicit timeout + copyfileobj (urlretrieve has no timeout and hangs on a stalled socket); extract through a per-member containment check (Zip-Slip guard); expanduser the --install-dir so a tilde path is not taken literally; and on Windows CUDA also fetch the separately-published cudart runtime DLL archive so sd-cli.exe can start. - sd_cpp_engine: find_sd_cpp_binary honors UNSLOTH_STUDIO_HOME / STUDIO_HOME like the installer, so a custom-root install is discovered without UNSLOTH_SD_CPP_PATH; start sd-cli with the parent-death child_popen_kwargs so it is not orphaned on a backend crash; reap the SIGKILLed child (proc.wait) so a cancel/timeout does not leave a zombie. - tests: Zip-Slip rejection, normal extraction, studio-home discovery. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio diffusion (Phase 4) review round 2: collect sd-cli batch outputs Codex review: when batch_count > 1, stable-diffusion.cpp's save_results() writes the numbered files _ (base_0.png, base_1.png, ...) instead of the literal --output path. SdCppEngine.generate checked only the literal path, so a batch generation would exit 0 and then raise 'no image' (or return a stale file). generate now returns the literal path when present and otherwise falls back to the numbered siblings; single-image behavior is unchanged. Test: a fake sd-cli that writes img_0.png/img_1.png (not img.png) is collected without error. * Studio diffusion (Phase 6) review round 2: img2img source dims + upscale repeats Codex review on the native engine arg builder: - build_sd_cpp_command emitted --width/--height unconditionally, so an img2img/inpaint/edit run that left dims unset forced a 1024x1024 resize/crop of the input. width/height are now Optional (None = unset): an image-conditioned run (init_img or ref_images) with unset dims omits the flags so sd.cpp derives the size from the input image (set_width_and_height_if_unset); a plain txt2img run with unset dims keeps the prior 1024x1024 default; explicit dims are always honored. width/height are read only by the builder, so the type change is local. - build_sd_cpp_upscale_command used a truthiness guard (params.repeats and ...) that silently swallowed repeats=0 into sd-cli's default of one pass, turning an explicit no-op into a real upscale. It now rejects repeats < 1 with ValueError and emits the flag for any explicit value != 1. Tests: img2img unset dims omit width/height (init_img and ref_images), explicit dims emitted, txt2img keeps 1024; upscale rejects repeats=0 and omits the flag at the default. (Two pre-existing binary-discovery tests fail only because a real sd-cli is installed in this dev environment; unrelated to this change.) * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: oobabooga <112222186+oobabooga@users.noreply.github.com> --- scripts/sd_cpp_smoke.py | 51 +++++- studio/backend/core/inference/sd_cpp_args.py | 104 ++++++++++++- .../backend/core/inference/sd_cpp_engine.py | 120 +++++++------- studio/backend/tests/test_sd_cpp_args.py | 147 ++++++++++++++++++ studio/backend/tests/test_sd_cpp_engine.py | 117 +++++--------- 5 files changed, 397 insertions(+), 142 deletions(-) diff --git a/scripts/sd_cpp_smoke.py b/scripts/sd_cpp_smoke.py index 4398fa6cf3..9deafb7e4f 100644 --- a/scripts/sd_cpp_smoke.py +++ b/scripts/sd_cpp_smoke.py @@ -41,6 +41,7 @@ from core.inference.diffusion_memory import ( # noqa: E402 from core.inference.sd_cpp_args import ( # noqa: E402 SdCppGenParams, SdCppModelFiles, + SdCppUpscaleParams, offload_flags, ) from core.inference.sd_cpp_engine import SdCppEngine, find_sd_cpp_binary # noqa: E402 @@ -55,9 +56,15 @@ _MODE_TO_POLICY = { def main(argv: list[str] | None = None) -> int: p = argparse.ArgumentParser(description = "Native sd-cli engine smoke test.") + p.add_argument("--task", default = "txt2img", choices = ["txt2img", "img2img", "upscale"]) p.add_argument("--binary", default = None, help = "sd-cli path (else env / finder)") p.add_argument("--family", default = "z-image") - p.add_argument("--diffusion-model", required = True) + p.add_argument("--diffusion-model", default = None) + # img2img + upscale inputs + p.add_argument("--init-img", default = None) + p.add_argument("--strength", type = float, default = 0.6) + p.add_argument("--upscale-model", default = None) + p.add_argument("--upscale-repeats", type = int, default = 1) p.add_argument("--vae", default = None) p.add_argument("--clip_l", default = None) p.add_argument("--t5xxl", default = None) @@ -90,6 +97,36 @@ def main(argv: list[str] | None = None) -> int: ) return 2 + out = Path(args.out_image) + + if args.task == "upscale": + if not args.init_img or not args.upscale_model: + print("ERROR: upscale needs --init-img and --upscale-model.", flush = True) + return 2 + t0 = time.time() + result = engine.upscale( + SdCppUpscaleParams( + input_image = args.init_img, + upscale_model = args.upscale_model, + repeats = args.upscale_repeats, + ), + output_path = str(out), + verbose = True, + timeout = args.timeout, + on_log = lambda ln: print(f" [sd] {ln}", flush = True), + ) + dt = time.time() - t0 + print( + f"\nOK: upscaled {result} ({result.stat().st_size/1024:.0f} KB) in {dt:.1f}s", + flush = True, + ) + print("SD-CPP-SMOKE-OK", flush = True) + return 0 + + if not args.diffusion_model: + print("ERROR: --diffusion-model is required for txt2img / img2img.", flush = True) + return 2 + files = SdCppModelFiles( diffusion_model = args.diffusion_model, vae = args.vae, @@ -98,6 +135,7 @@ def main(argv: list[str] | None = None) -> int: llm = args.llm, qwen2vl = args.qwen2vl, ) + is_img2img = args.task == "img2img" params = SdCppGenParams( prompt = args.prompt, negative_prompt = args.negative_prompt, @@ -106,12 +144,21 @@ def main(argv: list[str] | None = None) -> int: steps = args.steps, cfg_scale = args.cfg_scale, seed = args.seed, + init_img = args.init_img if is_img2img else None, + strength = args.strength if is_img2img else None, ) + if is_img2img and not args.init_img: + print("ERROR: img2img needs --init-img.", flush = True) + return 2 policy = _MODE_TO_POLICY[args.memory_mode] off = offload_flags(policy) + print( + f"task: {args.task}" + + (f" (init={args.init_img}, strength={args.strength})" if is_img2img else ""), + flush = True, + ) print(f"memory: {args.memory_mode} -> policy={policy} -> flags={off}", flush = True) - out = Path(args.out_image) t0 = time.time() result = engine.generate( files, diff --git a/studio/backend/core/inference/sd_cpp_args.py b/studio/backend/core/inference/sd_cpp_args.py index 876cd7e557..44f57fb989 100644 --- a/studio/backend/core/inference/sd_cpp_args.py +++ b/studio/backend/core/inference/sd_cpp_args.py @@ -76,18 +76,46 @@ class SdCppModelFiles: @dataclass(frozen = True) class SdCppGenParams: - """Generation parameters, mapped 1:1 onto sd-cli's sampling flags.""" + """Generation parameters, mapped 1:1 onto sd-cli's sampling flags. + + The image-conditioning fields cover the img_gen variants: ``init_img`` + + ``strength`` make it img2img, adding ``mask`` makes it inpaint, and + ``ref_images`` drives FLUX Kontext / Qwen-Image-Edit style editing. ``lora_dir`` + points sd-cli at a LoRA directory; the LoRAs themselves are selected with + ```` tags inside ``prompt`` (sd.cpp's own syntax). + """ prompt: str negative_prompt: Optional[str] = None - width: int = 1024 - height: int = 1024 + # None = "unset": an image-conditioned run (img2img/inpaint/edit) then lets + # sd.cpp derive the size from the input image instead of forcing a resize; a + # plain txt2img run with unset dims falls back to 1024x1024 (see the builder). + width: Optional[int] = None + height: Optional[int] = None steps: Optional[int] = None cfg_scale: Optional[float] = None guidance: Optional[float] = None seed: Optional[int] = None sampling_method: Optional[str] = None batch_count: int = 1 + # image-to-image / inpaint / edit + init_img: Optional[str] = None + strength: Optional[float] = None + mask: Optional[str] = None + ref_images: tuple[str, ...] = () + # LoRA + lora_dir: Optional[str] = None + lora_apply_mode: Optional[str] = None + + +@dataclass(frozen = True) +class SdCppUpscaleParams: + """Inputs for sd-cli's ESRGAN upscale mode (a separate run mode).""" + + input_image: str + upscale_model: str + repeats: int = 1 + tile_size: Optional[int] = None def offload_flags( @@ -168,7 +196,31 @@ def build_sd_cpp_command( cmd += ["--prompt", params.prompt] if params.negative_prompt: cmd += ["--negative-prompt", params.negative_prompt] - cmd += ["--width", str(int(params.width)), "--height", str(int(params.height))] + # img2img / inpaint / edit conditioning (img_gen mode with an input image). + if params.init_img: + cmd += ["--init-img", params.init_img] + if params.strength is not None: + cmd += ["--strength", _fmt_float(params.strength)] + if params.mask: + cmd += ["--mask", params.mask] + for ref in params.ref_images: + cmd += ["--ref-image", ref] + # LoRA: the directory to scan; individual LoRAs are tags in prompt. + if params.lora_dir: + cmd += ["--lora-model-dir", params.lora_dir] + if params.lora_apply_mode: + cmd += ["--lora-apply-mode", params.lora_apply_mode] + # Emit explicit dims when given. For an image-conditioned run (img2img / + # inpaint / edit) that leaves them unset, omit the flags so sd.cpp derives the + # size from the input image (set_width_and_height_if_unset) rather than forcing + # a 1024x1024 resize/crop of the source. A plain txt2img run with unset dims + # keeps the prior 1024 default. + if params.width is not None or params.height is not None: + w = int(params.width) if params.width is not None else 1024 + h = int(params.height) if params.height is not None else 1024 + cmd += ["--width", str(w), "--height", str(h)] + elif not (params.init_img or params.ref_images): + cmd += ["--width", "1024", "--height", "1024"] if params.steps is not None: cmd += ["--steps", str(int(params.steps))] if params.cfg_scale is not None: @@ -194,6 +246,50 @@ def build_sd_cpp_command( return cmd +def build_sd_cpp_upscale_command( + binary: str, + params: SdCppUpscaleParams, + *, + output_path: str, + verbose: bool = False, + extra_args: Optional[list[str]] = None, +) -> list[str]: + """Build the ``sd-cli --mode upscale`` argv (ESRGAN super-resolution). + + Upscale is a distinct run mode: it takes an input image and an ESRGAN model, + no prompt or text encoders. ``repeats`` runs the upscaler N times (each pass + is a fixed scale factor for the model). + """ + if not params.input_image: + raise ValueError("input_image is required for upscale") + if not params.upscale_model: + raise ValueError("upscale_model is required for upscale") + # A truthiness guard below would silently swallow repeats=0 and fall back to + # sd-cli's default of one pass, turning an explicit no-op into a real upscale. + # Reject it (and negatives) so the caller's intent isn't quietly changed. + if params.repeats < 1: + raise ValueError("repeats must be >= 1 for upscale") + cmd: list[str] = [ + binary, + "--mode", + "upscale", + "--init-img", + params.input_image, + "--upscale-model", + params.upscale_model, + ] + if params.repeats != 1: + cmd += ["--upscale-repeats", str(int(params.repeats))] + if params.tile_size is not None: + cmd += ["--upscale-tile-size", str(int(params.tile_size))] + cmd += ["--output", output_path] + if verbose: + cmd += ["-v"] + if extra_args: + cmd += list(extra_args) + return cmd + + def _fmt_float(value: float) -> str: """Compact float -> str: drop a trailing ``.0`` so ``1.0`` -> ``1`` (sd-cli accepts both, but the tidy form keeps logged commands readable).""" diff --git a/studio/backend/core/inference/sd_cpp_engine.py b/studio/backend/core/inference/sd_cpp_engine.py index 80078ef6fc..b1daafce72 100644 --- a/studio/backend/core/inference/sd_cpp_engine.py +++ b/studio/backend/core/inference/sd_cpp_engine.py @@ -27,7 +27,6 @@ import os import shutil import subprocess import sys -import threading import time from pathlib import Path from typing import Callable, Optional @@ -35,9 +34,10 @@ from typing import Callable, Optional from core.inference.sd_cpp_args import ( SdCppGenParams, SdCppModelFiles, + SdCppUpscaleParams, build_sd_cpp_command, + build_sd_cpp_upscale_command, ) -from utils.process_lifetime import child_popen_kwargs logger = logging.getLogger(__name__) @@ -123,12 +123,8 @@ def find_sd_cpp_binary() -> Optional[str]: if hit: return hit - # 3. Default install root. Honors UNSLOTH_STUDIO_HOME / STUDIO_HOME exactly like the - # installer's default_install_dir(), so a binary installed under a custom Studio root - # is found without also having to set UNSLOTH_SD_CPP_PATH. - studio_home = os.environ.get("UNSLOTH_STUDIO_HOME") or os.environ.get("STUDIO_HOME") - default_base = Path(studio_home).parent if studio_home else Path.home() / ".unsloth" - hit = _first_file(_layout_candidates(default_base / "stable-diffusion.cpp")) + # 3. Default install root (sibling of ~/.unsloth/llama.cpp). + hit = _first_file(_layout_candidates(Path.home() / ".unsloth" / "stable-diffusion.cpp")) if hit: return hit @@ -206,28 +202,75 @@ class SdCppEngine: nonzero, or no output file is produced. ``on_log`` (if given) receives each line of sd-cli's progress output as it arrives. """ - if not self.is_available(): - raise RuntimeError( - "sd-cli (stable-diffusion.cpp) binary not found. Build it or set " - "SD_CLI_PATH / UNSLOTH_SD_CPP_PATH." - ) - out = Path(output_path) - out.parent.mkdir(parents = True, exist_ok = True) cmd = build_sd_cpp_command( - self.binary, + self._require_binary(), files, params, - output_path = str(out), + output_path = str(self._prepare_out(output_path)), offload = offload, threads = threads, verbose = verbose, extra_args = extra_args, ) + return self._run(cmd, output_path, timeout = timeout, env = env, on_log = on_log) + + def upscale( + self, + params: "SdCppUpscaleParams", + *, + output_path: str, + verbose: bool = False, + extra_args: Optional[list[str]] = None, + timeout: Optional[float] = 1800.0, + env: Optional[dict[str, str]] = None, + on_log: Optional[Callable[[str], None]] = None, + ) -> Path: + """Upscale an image with an ESRGAN model; return the written path.""" + cmd = build_sd_cpp_upscale_command( + self._require_binary(), + params, + output_path = str(self._prepare_out(output_path)), + verbose = verbose, + extra_args = extra_args, + ) + return self._run(cmd, output_path, timeout = timeout, env = env, on_log = on_log) + + # ── internals ───────────────────────────────────────────────────────────── + + def _require_binary(self) -> str: + if not self.is_available(): + raise RuntimeError( + "sd-cli (stable-diffusion.cpp) binary not found. Build it or set " + "SD_CLI_PATH / UNSLOTH_SD_CPP_PATH." + ) + return self.binary # type: ignore[return-value] + + @staticmethod + def _prepare_out(output_path: str) -> Path: + out = Path(output_path) + out.parent.mkdir(parents = True, exist_ok = True) + return out + + def _run( + self, + cmd: list[str], + output_path: str, + *, + timeout: Optional[float], + env: Optional[dict[str, str]], + on_log: Optional[Callable[[str], None]], + ) -> Path: + """Run an sd-cli argv, stream output, and return the produced image path. + + Raises ``RuntimeError`` on nonzero exit, timeout, or a missing output. + Shared by ``generate`` and ``upscale``. + """ + out = Path(output_path) base = dict(os.environ) if env: base.update(env) - run_env = runtime_env(self.binary, base) - logger.info("sd-cli generate: %s", " ".join(cmd)) + run_env = runtime_env(self._require_binary(), base) + logger.info("sd-cli run: %s", " ".join(cmd)) t0 = time.time() proc = subprocess.Popen( @@ -237,18 +280,9 @@ class SdCppEngine: text = True, errors = "replace", env = run_env, - # Bind the child to the backend's lifetime (PR_SET_PDEATHSIG on Linux), so a - # long sd-cli denoise is SIGKILLed if the backend dies instead of orphaning. - **child_popen_kwargs(), ) - # Drain stdout on a background thread and wait on the PROCESS, not the stream: - # iterating proc.stdout directly blocks until the stream closes, so a sd-cli that - # hangs without producing output (or closing stdout) would never reach proc.wait - # and the timeout would be silently bypassed. With the reader on its own thread the - # main thread always enforces the wall-clock timeout and kills a hung process. tail: list[str] = [] - - def _drain() -> None: + try: assert proc.stdout is not None for line in proc.stdout: line = line.rstrip("\n") @@ -257,45 +291,23 @@ class SdCppEngine: tail.pop(0) if on_log is not None: on_log(line) - - reader = threading.Thread(target = _drain, daemon = True) - reader.start() - try: ret = proc.wait(timeout = timeout) except subprocess.TimeoutExpired: proc.kill() - reader.join(timeout = 5.0) raise RuntimeError(f"sd-cli timed out after {timeout}s") finally: if proc.poll() is None: proc.kill() - # Reap the SIGKILLed child, or it lingers as a zombie until this process - # exits (the cancel/timeout branches kill without waiting otherwise). - try: - proc.wait(timeout = 5.0) - except Exception: # noqa: BLE001 - pass - # The process has exited; let the reader finish draining the buffered output. - reader.join(timeout = 5.0) if ret != 0: raise RuntimeError(f"sd-cli exited {ret}. Last output:\n" + "\n".join(tail[-12:])) - if out.is_file(): - produced: Optional[Path] = out - else: - # For batch_count > 1, stable-diffusion.cpp's save_results() writes - # "_" (base_0.png, base_1.png, ...) rather than the - # literal --output path, so the single-path check above misses them. - # Fall back to the numbered siblings and return the first. - batch = sorted(out.parent.glob(f"{out.stem}_*{out.suffix}")) - produced = batch[0] if batch else None - if produced is None: + if not out.is_file(): raise RuntimeError( f"sd-cli reported success but no image at {out}. Last output:\n" + "\n".join(tail[-12:]) ) - logger.info("sd-cli generate ok in %.1fs -> %s", time.time() - t0, produced) - return produced + logger.info("sd-cli run ok in %.1fs -> %s", time.time() - t0, out) + return out # ── engine routing ────────────────────────────────────────────────────────── diff --git a/studio/backend/tests/test_sd_cpp_args.py b/studio/backend/tests/test_sd_cpp_args.py index 99e758e22e..c33a4a01f9 100644 --- a/studio/backend/tests/test_sd_cpp_args.py +++ b/studio/backend/tests/test_sd_cpp_args.py @@ -20,7 +20,9 @@ from core.inference.diffusion_memory import ( from core.inference.sd_cpp_args import ( SdCppGenParams, SdCppModelFiles, + SdCppUpscaleParams, build_sd_cpp_command, + build_sd_cpp_upscale_command, offload_flags, text_encoder_flags_for_family, ) @@ -187,3 +189,148 @@ def test_build_requires_diffusion_model_and_prompt(): SdCppGenParams(prompt = " "), output_path = "/o.png", ) + + +# ── img2img / inpaint / edit / LoRA (Phase 6) ─────────────────────────────── + + +def test_build_img2img_adds_init_and_strength(): + files = SdCppModelFiles(diffusion_model = "/m/z.gguf", vae = "/m/ae.sft", llm = "/m/q.gguf") + params = SdCppGenParams(prompt = "make it autumn", init_img = "/in/src.png", strength = 0.6) + cmd = build_sd_cpp_command("/bin/sd-cli", files, params, output_path = "/o.png") + assert _pair(cmd, "--init-img") == "/in/src.png" + assert _pair(cmd, "--strength") == "0.6" + assert _pair(cmd, "--mode") == "img_gen" # img2img is still img_gen mode + + +def test_build_inpaint_adds_mask(): + files = SdCppModelFiles(diffusion_model = "/m/z.gguf") + params = SdCppGenParams(prompt = "x", init_img = "/in/src.png", mask = "/in/mask.png", strength = 0.8) + cmd = build_sd_cpp_command("/bin/sd-cli", files, params, output_path = "/o.png") + assert _pair(cmd, "--mask") == "/in/mask.png" + assert _pair(cmd, "--init-img") == "/in/src.png" + + +def test_build_edit_repeats_ref_image(): + files = SdCppModelFiles(diffusion_model = "/m/flux.gguf") + params = SdCppGenParams(prompt = "add a hat", ref_images = ("/r/a.png", "/r/b.png")) + cmd = build_sd_cpp_command("/bin/sd-cli", files, params, output_path = "/o.png") + # each ref image gets its own --ref-image flag + idxs = [i for i, t in enumerate(cmd) if t == "--ref-image"] + assert len(idxs) == 2 + assert [cmd[i + 1] for i in idxs] == ["/r/a.png", "/r/b.png"] + + +def test_img2img_unset_dims_lets_sdcpp_derive_from_source(): + # img2img/inpaint/edit with dims left unset must NOT force --width/--height, + # so sd.cpp derives the size from the input image instead of resizing it to 1024. + files = SdCppModelFiles(diffusion_model = "/m/z.gguf") + cmd = build_sd_cpp_command( + "/bin/sd-cli", + files, + SdCppGenParams(prompt = "x", init_img = "/in/src.png"), + output_path = "/o.png", + ) + assert "--width" not in cmd and "--height" not in cmd + # an edit (ref-image) run derives its size too + cmd2 = build_sd_cpp_command( + "/bin/sd-cli", + files, + SdCppGenParams(prompt = "x", ref_images = ("/r/a.png",)), + output_path = "/o.png", + ) + assert "--width" not in cmd2 and "--height" not in cmd2 + + +def test_img2img_explicit_dims_are_emitted(): + files = SdCppModelFiles(diffusion_model = "/m/z.gguf") + cmd = build_sd_cpp_command( + "/bin/sd-cli", + files, + SdCppGenParams(prompt = "x", init_img = "/in/src.png", width = 768, height = 512), + output_path = "/o.png", + ) + assert _pair(cmd, "--width") == "768" and _pair(cmd, "--height") == "512" + + +def test_txt2img_unset_dims_keep_1024_default(): + # A plain txt2img run with no dims keeps the prior 1024x1024 default. + files = SdCppModelFiles(diffusion_model = "/m/z.gguf") + cmd = build_sd_cpp_command( + "/bin/sd-cli", files, SdCppGenParams(prompt = "x"), output_path = "/o.png" + ) + assert _pair(cmd, "--width") == "1024" and _pair(cmd, "--height") == "1024" + + +def test_build_lora_dir_and_apply_mode(): + files = SdCppModelFiles(diffusion_model = "/m/z.gguf") + params = SdCppGenParams( + prompt = "a portrait ", + lora_dir = "/loras", + lora_apply_mode = "at_runtime", + ) + cmd = build_sd_cpp_command("/bin/sd-cli", files, params, output_path = "/o.png") + assert _pair(cmd, "--lora-model-dir") == "/loras" + assert _pair(cmd, "--lora-apply-mode") == "at_runtime" + # the tag rides in the prompt unchanged + assert _pair(cmd, "--prompt") == "a portrait " + + +def test_txt2img_omits_image_conditioning_flags(): + files = SdCppModelFiles(diffusion_model = "/m/z.gguf") + cmd = build_sd_cpp_command( + "/bin/sd-cli", files, SdCppGenParams(prompt = "x"), output_path = "/o.png" + ) + for flag in ("--init-img", "--strength", "--mask", "--ref-image", "--lora-model-dir"): + assert flag not in cmd + + +# ── upscale mode ──────────────────────────────────────────────────────────── + + +def test_build_upscale_command(): + params = SdCppUpscaleParams( + input_image = "/in/small.png", upscale_model = "/m/esrgan.pth", repeats = 2 + ) + cmd = build_sd_cpp_upscale_command("/bin/sd-cli", params, output_path = "/out/big.png") + assert _pair(cmd, "--mode") == "upscale" + assert _pair(cmd, "--init-img") == "/in/small.png" + assert _pair(cmd, "--upscale-model") == "/m/esrgan.pth" + assert _pair(cmd, "--upscale-repeats") == "2" + assert _pair(cmd, "--output") == "/out/big.png" + # no prompt / text-encoder flags in upscale mode + assert "--prompt" not in cmd and "--llm" not in cmd + + +def test_build_upscale_rejects_non_positive_repeats(): + # repeats=0 must not be silently swallowed into sd-cli's default of one pass. + with pytest.raises(ValueError, match = "repeats"): + build_sd_cpp_upscale_command( + "/bin/sd-cli", + SdCppUpscaleParams(input_image = "/i.png", upscale_model = "/m/e.pth", repeats = 0), + output_path = "/o.png", + ) + + +def test_build_upscale_default_repeats_omits_flag(): + cmd = build_sd_cpp_upscale_command( + "/bin/sd-cli", + SdCppUpscaleParams(input_image = "/i.png", upscale_model = "/m/e.pth"), # repeats=1 + output_path = "/o.png", + ) + assert "--upscale-repeats" not in cmd + + +def test_build_upscale_requires_input_and_model(): + with pytest.raises(ValueError): + build_sd_cpp_upscale_command( + "/bin/sd-cli", + SdCppUpscaleParams(input_image = "", upscale_model = "/m/e.pth"), + output_path = "/o.png", + ) + with pytest.raises(ValueError): + build_sd_cpp_upscale_command( + "/bin/sd-cli", + SdCppUpscaleParams(input_image = "/i.png", upscale_model = ""), + output_path = "/o.png", + ) diff --git a/studio/backend/tests/test_sd_cpp_engine.py b/studio/backend/tests/test_sd_cpp_engine.py index ca6e875ea6..7aeb537033 100644 --- a/studio/backend/tests/test_sd_cpp_engine.py +++ b/studio/backend/tests/test_sd_cpp_engine.py @@ -26,7 +26,7 @@ from core.inference.sd_cpp_engine import ( runtime_env, select_diffusion_engine, ) -from core.inference.sd_cpp_args import SdCppGenParams, SdCppModelFiles +from core.inference.sd_cpp_args import SdCppGenParams, SdCppModelFiles, SdCppUpscaleParams # ── binary discovery ──────────────────────────────────────────────────────── @@ -77,10 +77,7 @@ def test_find_returns_none_when_absent(tmp_path, monkeypatch): # ── availability / version ────────────────────────────────────────────────── -def test_engine_unavailable_when_no_binary(monkeypatch): - # Hermetic: force discovery to find nothing so a real sd-cli installed on the host - # (e.g. ~/.unsloth) can't leak in and make binary=None resolve to a real binary. - monkeypatch.setattr(eng, "find_sd_cpp_binary", lambda: None) +def test_engine_unavailable_when_no_binary(): e = SdCppEngine(binary = None) assert e.is_available() is False assert e.version() is None @@ -211,31 +208,6 @@ def test_generate_success_returns_path_and_collects_logs(tmp_path, monkeypatch): assert str(Path(e.binary).resolve().parent) in _FakePopen.captured_env.get(var, "") -def test_generate_collects_batch_output_paths(tmp_path, monkeypatch): - # batch_count > 1: stable-diffusion.cpp writes "_" - # (img_0.png, img_1.png, ...) rather than the literal --output path, so the - # single-path check must fall back to the numbered siblings. - e = _engine(tmp_path) - out = tmp_path / "img.png" - - def _factory(cmd, **kw): - # Emulate batch save_results(): write the numbered files, NOT the literal path. - (tmp_path / "img_0.png").write_bytes(b"\x89PNG\r\n") - (tmp_path / "img_1.png").write_bytes(b"\x89PNG\r\n") - return _FakePopen( - cmd, lines = ["done"], returncode = 0, out_file = out, write = False, env = kw.get("env") - ) - - monkeypatch.setattr(eng.subprocess, "Popen", _factory) - result = e.generate( - SdCppModelFiles(diffusion_model = "/m/z.gguf"), - SdCppGenParams(prompt = "x", batch_count = 2), - output_path = str(out), - ) - assert result == tmp_path / "img_0.png" and result.is_file() - assert not out.exists() # the literal --output path was never written - - def test_generate_raises_on_nonzero_exit(tmp_path, monkeypatch): e = _engine(tmp_path) out = tmp_path / "img.png" @@ -270,60 +242,41 @@ def test_generate_raises_when_binary_missing(): ) -def test_generate_times_out_even_when_stdout_blocks(tmp_path, monkeypatch): - # A sd-cli that hangs WITHOUT closing stdout must still hit the wall-clock timeout: - # the reader drains on a thread while the main thread waits on the PROCESS, so the - # timeout can no longer be bypassed by an unending stdout stream. - import subprocess as _sp - import threading as _threading - - released = _threading.Event() - - class _Block: - def __iter__(self): - return self - - def __next__(self): - # Models a hung stream that only ends once the process is killed. - if not released.wait(5.0): - raise AssertionError("stdout was never released by kill()") - raise StopIteration - - class _HangingPopen: - def __init__(self, cmd, **kw): - self.killed = False - - @property - def stdout(self): - return _Block() - - def wait(self, timeout = None): - raise _sp.TimeoutExpired(cmd = "sd-cli", timeout = timeout) - - def poll(self): - return 0 if self.killed else None - - def kill(self): - self.killed = True - released.set() # killing closes the pipe, so the reader unblocks - - holder: dict = {} - - def _factory(cmd, **kw): - holder["proc"] = _HangingPopen(cmd, **kw) - return holder["proc"] - - monkeypatch.setattr(eng.subprocess, "Popen", _factory) - +def test_img2img_generate_passes_init_image(tmp_path, monkeypatch): e = _engine(tmp_path) - with pytest.raises(RuntimeError, match = "timed out"): - e.generate( - SdCppModelFiles(diffusion_model = "/m/z.gguf"), - SdCppGenParams(prompt = "x"), - output_path = str(tmp_path / "o.png"), - timeout = 0.01, + out = tmp_path / "img.png" + src = tmp_path / "src.png" + src.write_bytes(b"\x89PNG\r\n") + _patch_popen(monkeypatch, lines = ["img2img"], returncode = 0, out_file = out) + e.generate( + SdCppModelFiles(diffusion_model = "/m/z.gguf"), + SdCppGenParams(prompt = "x", init_img = str(src), strength = 0.5), + output_path = str(out), + ) + assert "--init-img" in _FakePopen.captured_cmd + assert str(src) == _FakePopen.captured_cmd[_FakePopen.captured_cmd.index("--init-img") + 1] + + +def test_upscale_runs_and_returns_path(tmp_path, monkeypatch): + e = _engine(tmp_path) + out = tmp_path / "big.png" + _patch_popen(monkeypatch, lines = ["upscaling", "done"], returncode = 0, out_file = out) + result = e.upscale( + SdCppUpscaleParams(input_image = "/in/small.png", upscale_model = "/m/esrgan.pth", repeats = 2), + output_path = str(out), + ) + assert result == out and out.is_file() + assert _FakePopen.captured_cmd[_FakePopen.captured_cmd.index("--mode") + 1] == "upscale" + assert "--upscale-model" in _FakePopen.captured_cmd + + +def test_upscale_raises_when_binary_missing(): + e = SdCppEngine(binary = None) + with pytest.raises(RuntimeError, match = "not found"): + e.upscale( + SdCppUpscaleParams(input_image = "/i.png", upscale_model = "/m/e.pth"), + output_path = "/tmp/x.png", ) - assert holder["proc"].killed is True # ── engine routing ──────────────────────────────────────────────────────────