Codex review on the native engine arg builder:
- build_sd_cpp_command emitted --width/--height unconditionally, so an
img2img/inpaint/edit run that left dims unset forced a 1024x1024 resize/crop of
the input. width/height are now Optional (None = unset): an image-conditioned
run (init_img or ref_images) with unset dims omits the flags so sd.cpp derives
the size from the input image (set_width_and_height_if_unset); a plain txt2img
run with unset dims keeps the prior 1024x1024 default; explicit dims are always
honored. width/height are read only by the builder, so the type change is local.
- build_sd_cpp_upscale_command used a truthiness guard (params.repeats and ...)
that silently swallowed repeats=0 into sd-cli's default of one pass, turning an
explicit no-op into a real upscale. It now rejects repeats < 1 with ValueError
and emits the flag for any explicit value != 1.
Tests: img2img unset dims omit width/height (init_img and ref_images), explicit
dims emitted, txt2img keeps 1024; upscale rejects repeats=0 and omits the flag at
the default. (Two pre-existing binary-discovery tests fail only because a real
sd-cli is installed in this dev environment; unrelated to this change.)
Re-review of the diffusion stack (#6675/#6679/#6680) surfaced one real accuracy
bug and a dead-on-arrival speed path; this fixes both and adds the lossless /
near-lossless wins, all measured on a B200.
Correctness:
- TF32 global-state leak (fix). speed_mode=max flipped torch.backends.*.allow_tf32
process-wide and never restored them, so a later `off` load silently inherited
TF32 and was no longer bit-identical. Added snapshot_backend_flags /
restore_backend_flags (TF32 + cudnn.benchmark), captured before the speed layer
runs and restored on unload. Verified: load max -> unload -> load off is now
byte-identical (PSNR inf) to a fresh off.
- sd-cli timeout could hang forever. _run() blocked in `for line in stdout` and
only checked the timeout after EOF, so a child stuck in model load / GPU init
with no output ignored the timeout. Drained stdout on a reader thread with a
wall-clock deadline. Added a silent-hang regression test.
Speed (diffusers path), near-lossless, opt-in tiers:
- Regional torch.compile now runs on the GGUF transformer. The is_gguf gate (and
Z-Image's supports_torch_compile=False) were stale: compile_repeated_blocks
compiles and runs ~2.2x faster on the GGUF Z-Image transformer on
torch 2.9.1 / diffusers 0.38 (the per-op dequant stays eager, the rest of the
block compiles). Measured: off 1.80s -> default 0.82s/gen (+54.7%), PSNR 37.7 dB
vs eager -- far above the Q4 quant noise floor (~21 dB), so it does not move
output quality. Gate relaxed; default tier delivers it.
- cudnn.benchmark added to the default tier (autotunes the fixed-shape VAE convs).
- torch.inference_mode() around the pipeline call (lossless, strictly faster than
the no_grad diffusers uses internally).
Memory path:
- VAE tiling (not bit-identical >1MP) restricted to the model/sequential/CPU tiers;
the balanced (group) tier keeps exact slicing only, so it is now bit-identical to
the resident image (verified PSNR inf) and slightly faster.
- Group offload adds non_blocking + record_stream on the CUDA stream path to
overlap each block's H2D copy with compute (lossless; gated on the installed
diffusers signature so older versions still work).
Native (sd.cpp) path:
- native_speed_flags: a first-class speed knob (default -> --diffusion-fa, a
near-lossless CUDA win that was previously only added on offload tiers; max also
-> --diffusion-conv-direct). conv-direct stays opt-in: measured +45% on CUDA, so
it is never auto-on. Engine generate() merges it, de-duped against offload flags.
Default profile: a GGUF model with no explicit speed_mode now resolves to the
`default` profile (resolve_speed_mode), since compile's perturbation sits below the
quantisation noise floor and so does not reduce quality versus the dense reference;
out of the box a GGUF Z-Image generation drops from 1.80s to 0.81s. Dense models
stay `off` / bit-identical, and an explicit speed_mode -- including "off" -- is
always honored, so the byte-identical path remains one flag away and is the
regression reference.
Tooling: scripts/compile_probe.py (eager vs compiled GGUF probe), scripts/
perf_verify.py (the B200 verification above), and diffusion_bench.py gains
--speed-mode so the speed tiers are benchmarkable.
Tests: 183 passing (was 166); new coverage for the backend-flag snapshot/restore,
GGUF compile eligibility, the balanced tiling/slicing split, native_speed_flags +
the engine de-dup, and the sd-cli silent-hang timeout.
Builds on Phase 4's native stable-diffusion.cpp engine, extending it from
text-to-image to the wider feature surface, since sd.cpp supports all of these
through the binary already. Pure command-builder additions plus one engine
method, so the txt2img path is unchanged.
- sd_cpp_args.py: SdCppGenParams gains image-conditioning fields. init_img +
strength make a run img2img, adding mask makes it inpaint, ref_images drives
FLUX-Kontext / Qwen-Image-Edit style editing (repeated --ref-image), and
lora_dir + the <lora:name:weight> prompt syntax select LoRAs. New
SdCppUpscaleParams + build_sd_cpp_upscale_command for the ESRGAN upscale run
mode (input image + esrgan model, no prompt / text encoders).
- sd_cpp_engine.py: the subprocess runner is factored into a shared _run() so
generate() (now carrying the conditioning flags) and a new upscale() reuse
the same streaming / error / output-check path.
- scripts/sd_cpp_smoke.py: --task {txt2img,img2img,upscale} with --init-img /
--strength / --upscale-model / --upscale-repeats.
Tests: 10 new across the img2img / inpaint / edit / LoRA flag construction, the
upscale builder and its validation, and the engine's img2img + upscale paths.
Full diffusion suite 176 passing.
Verified on a B200 box through SdCppEngine: img2img (Z-Image-Turbo Q4_K, the
init image conditioned at strength 0.6, 4.8s) and ESRGAN upscale
(512x512 -> 2048x2048 via RealESRGAN_x4plus_anime_6B, 2.7s), both producing
coherent images. Video and the diffusers-path feature wiring are deferred.
Adds the CPU / Apple-Silicon tier of the two-engine strategy, mirroring the
chat backend's llama.cpp shell-out. Diffusers stays the default on CUDA / ROCm
/ XPU; this covers the hardware diffusers serves poorly, consuming the same
split GGUF assets Studio already curates.
- sd_cpp_args.py: pure sd-cli command builder. Maps the family to its
text-encoder flag (Z-Image Qwen3 to --llm, Qwen-Image to --qwen2vl, FLUX.1
CLIP-L + T5), and the diffusers memory policy (none/group/model/sequential)
to sd.cpp's offload flags (--offload-to-cpu / --clip-on-cpu / --vae-on-cpu /
--vae-tiling / --diffusion-fa), so one user knob drives both engines.
- sd_cpp_engine.py: SdCppEngine over a located sd-cli. find_sd_cpp_binary()
with the same precedence as the llama finder (env override, then the Studio
install root, then in-tree, then PATH), an is_available/version probe, and a
one-shot subprocess generate that streams progress and returns the PNG.
runtime_env() prepends the binary's directory to the platform library path
so a prebuilt's bundled libstable-diffusion.so resolves.
select_diffusion_engine() is the pure routing decision (GPU backends to
diffusers, CPU/MPS to native when present).
- install_sd_cpp_prebuilt.py: resolve + download the per-host prebuilt
(macOS-arm64/Metal, Linux x86_64 CPU, Vulkan/ROCm/Windows variants) into the
Studio install root. resolve_release_asset() is a pure, unit-tested
host-to-asset matrix.
- scripts/sd_cpp_smoke.py: end-to-end native generation harness.
Tests (CPU-only, subprocess/filesystem stubbed): 49 new across args, engine,
routing, runtime env, and the installer resolver. Full diffusion suite 166
passing.
Verified on a B200 box: built sd-cli (CUDA) and the prebuilt (CPU) both
generate Z-Image-Turbo Q4_K end to end through SdCppEngine: balanced (group
offload, 5.0s gen), low_vram (full CPU offload + VAE tiling, 13.4s), and the
dynamically-linked CPU prebuilt (50.4s on CPU), all producing coherent images.