* Studio diffusion: cross-platform device policy, fp16 guard, lock split, validate-before-evict Phase 1 of porting the richer diffusion stack onto the image-generation backend. - Add a compartmentalized device/dtype policy module (diffusion_device.py) resolving CUDA/ROCm/XPU/MPS/CPU with capability flags. Keeps the NVIDIA capability-based bf16 choice; ROCm and XPU are isolated; MPS uses bf16 or fp32, never a silent fp16 that renders a black image. - Add a per-family fp16_incompatible flag (Z-Image) and promote a resolved float16 to float32 for those families so they do not produce black images. - Split the backend locks: a generation holds only _generate_lock, so status, unload, and a new load are never blocked by a long denoise. Add per-generation cancellation via callback_on_step_end so an eviction or a superseding load preempts a running generation; a replacement load waits for it to stop before allocating, so two pipelines never sit in VRAM at once. - Validate a load request before the GPU handoff so an unloadable pick never evicts a working chat model, and reject missing local paths up front. - Add CPU-only tests for the device policy, dtype guard, lock split and cancellation, and validate-before-evict, plus a GPU benchmark/regression script (scripts/diffusion_bench.py) measuring latency, peak VRAM, and PSNR against a saved reference. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio diffusion (Phase 2A): measured-budget memory planner + offload/VAE policy Add a lean, backend-agnostic memory policy that picks a CPU-offload policy and VAE tiling/slicing from measured free device memory vs the model's estimated resident footprint, then applies it to the built pipeline. auto stays resident when the model fits (byte-identical to the prior resident path), and falls to whole-module offload when tight; fast/balanced/low_vram are explicit overrides. Sequential submodule offload is unreliable for GGUF transformers on diffusers 0.38, so it falls back to whole-module offload and status reports the policy actually engaged. Verified on Z-Image-Turbo Q4_K_M (B200): auto reproduces the resident image with no VRAM/latency regression (PSNR inf); balanced/low_vram cut generation peak VRAM 47.9% (15951 -> 8318 MB) with byte-identical output, at the expected latency cost. 73 prior + 35 new CPU tests pass. * Studio diffusion (Phase 2D): streamed block-level offload + functional VAE tiling Add a streamed 'group' offload tier (diffusers apply_group_offloading, block_level, use_stream) that keeps the transformer flowing through the GPU a few blocks at a time while the text encoder / VAE stay resident, and fix VAE tiling to drive the VAE submodule (pipelines like Z-Image expose enable_tiling on pipe.vae, not the pipeline). apply_memory_plan now returns the (policy, tiling) actually engaged so status never overstates either, and group falls back to whole-module offload when the transformer can't be streamed. Measured on Z-Image (B200), all lossless (PSNR inf vs resident): balanced/group cuts generation peak VRAM 32% (15951 -> 10840 MB) at near-resident speed (2.07 -> 2.99s); low_vram/model cuts it 48% (-> 8318 MB) but is slower (7.99s). Mode names now match that tradeoff: balanced = stream the transformer, low_vram = offload every component. auto picks group when the companions fit resident, else model. 112 CPU tests pass. * Studio diffusion (Phase 5): image quality-vs-quant accuracy harness Add scripts/diffusion_quality.py, the accuracy analogue of the KLD workflow: hold prompt + seed fixed, render a grid with a reference quant (default BF16), then render each candidate quant and measure drift from the reference. Records mean PSNR + SSIM (pure-numpy, no skimage/scipy) and optional CLIP text-alignment + image-similarity (transformers, --clip), plus file size, latency, and peak VRAM, then prints a quality-vs-cost table and recommends the smallest quant within a quality budget. --selftest validates the metrics on synthetic images with no GPU or model. Verified on Z-Image (B200): the table degrades monotonically with quant size (Q8 -> Q4 -> Q2: PSNR 21.7 -> 15.5, SSIM 0.82 -> 0.61), while CLIP-text stays flat (~0.34) -- quantization erodes fine detail far more than prompt adherence. * Studio diffusion (Phase 3): opt-in speed layer (channels_last / compile / TF32) Add a speed_mode knob (off by default, so the render path stays bit-identical): default applies channels_last VAE + regional torch.compile of the denoiser's repeated block where eligible; max also enables TF32 matmul and fused QKV. Regional compile is gated off for the GGUF transformer (dequantises per-op) and for families flagged not compile-friendly (a new supports_torch_compile flag, False for Z-Image), so it activates automatically only once a non-GGUF bf16 transformer is loaded. Speed optims run before placement/offload, per the diffusers composition order. status now reports speed_mode + the optims actually engaged. Verified on Z-Image (B200): default -> ['channels_last'], max -> ['channels_last', 'tf32'], compile correctly skipped for GGUF; generation works in every mode. 121 CPU tests pass. * Studio diffusion (Phase 2B): opt-in fp8 text-encoder layerwise casting Add a text_encoder_fp8 knob that casts the companion text encoder(s) to fp8 (e4m3) storage via diffusers apply_layerwise_casting, upcasting per layer to the bf16 compute dtype while normalisations and embeddings stay full precision. Applied before placement, gated to CUDA + bf16, best-effort (a failure leaves the encoder dense). status reports which encoders were cast. Verified on Z-Image (B200, balanced/group mode where the encoder stays resident): generation peak VRAM dropped 37% (10840 -> 6791 MB, below the lowest-VRAM offload) at near-resident speed. It is a memory-vs-quality tradeoff, not free -- ~20 dB PSNR vs the bf16 encoder, a larger shift than one transformer quant step -- so it is off by default and documented as such, with the Phase 5 harness to size the cost. 127 CPU tests pass. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio diffusion (Phase 2C): NVFP4 text-encoder quant (+ generalise fp8 knob) Generalise the text-encoder precision knob from a fp8 bool to text_encoder_quant (fp8 | nvfp4). nvfp4 quantises the companion text encoder to 4-bit via torchao NVFP4 weight-only (two-level microscaling) on Blackwell's FP4 tensor cores; fp8 stays the broader-hardware path (cc>=8.9). Both are gated, best-effort, and run before placement; status reports the mode actually engaged. This is the lean realisation of GGUF-native text-encoder quant: 4-bit on the encoder without the 3045-line port. Verified on Z-Image (B200, balanced/group where the encoder stays resident), vs the bf16 encoder: nvfp4 cut generation peak VRAM 48% (10840 -> 5593 MB, the lowest TE option, below whole-model offload) at near-fp8 quality (16.4 vs 17.1 dB PSNR), and both quants ran faster than bf16. A memory-vs-quality tradeoff (off by default); size it per model with the Phase 5 quality harness. diffusion_bench gains --text-encoder-quant. 129 CPU tests pass. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio diffusion (Phase 4): native stable-diffusion.cpp engine for CPU/Mac Adds the CPU / Apple-Silicon tier of the two-engine strategy, mirroring the chat backend's llama.cpp shell-out. Diffusers stays the default on CUDA / ROCm / XPU; this covers the hardware diffusers serves poorly, consuming the same split GGUF assets Studio already curates. - sd_cpp_args.py: pure sd-cli command builder. Maps the family to its text-encoder flag (Z-Image Qwen3 to --llm, Qwen-Image to --qwen2vl, FLUX.1 CLIP-L + T5), and the diffusers memory policy (none/group/model/sequential) to sd.cpp's offload flags (--offload-to-cpu / --clip-on-cpu / --vae-on-cpu / --vae-tiling / --diffusion-fa), so one user knob drives both engines. - sd_cpp_engine.py: SdCppEngine over a located sd-cli. find_sd_cpp_binary() with the same precedence as the llama finder (env override, then the Studio install root, then in-tree, then PATH), an is_available/version probe, and a one-shot subprocess generate that streams progress and returns the PNG. runtime_env() prepends the binary's directory to the platform library path so a prebuilt's bundled libstable-diffusion.so resolves. select_diffusion_engine() is the pure routing decision (GPU backends to diffusers, CPU/MPS to native when present). - install_sd_cpp_prebuilt.py: resolve + download the per-host prebuilt (macOS-arm64/Metal, Linux x86_64 CPU, Vulkan/ROCm/Windows variants) into the Studio install root. resolve_release_asset() is a pure, unit-tested host-to-asset matrix. - scripts/sd_cpp_smoke.py: end-to-end native generation harness. Tests (CPU-only, subprocess/filesystem stubbed): 49 new across args, engine, routing, runtime env, and the installer resolver. Full diffusion suite 166 passing. Verified on a B200 box: built sd-cli (CUDA) and the prebuilt (CPU) both generate Z-Image-Turbo Q4_K end to end through SdCppEngine: balanced (group offload, 5.0s gen), low_vram (full CPU offload + VAE tiling, 13.4s), and the dynamically-linked CPU prebuilt (50.4s on CPU), all producing coherent images. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio diffusion (Phase 4): enforce the sd-cli timeout while reading output Iterating proc.stdout directly blocks until the stream closes, so a sd-cli that hangs without producing output (or without closing stdout) would never reach proc.wait and the wall-clock timeout was silently bypassed. Drain stdout on a daemon thread and wait on the PROCESS, so the main thread always enforces the timeout and kills a hung process (which closes the pipe and ends the reader). Add a test that times out even when stdout blocks, and make the no-binary test hermetic so a host-installed sd-cli can't leak in. * Studio diffusion (Phase 4) review fixes: sd.cpp installer + engine hardening - install_sd_cpp_prebuilt: download the release archive with urlopen + an explicit timeout + copyfileobj (urlretrieve has no timeout and hangs on a stalled socket); extract through a per-member containment check (Zip-Slip guard); expanduser the --install-dir so a tilde path is not taken literally; and on Windows CUDA also fetch the separately-published cudart runtime DLL archive so sd-cli.exe can start. - sd_cpp_engine: find_sd_cpp_binary honors UNSLOTH_STUDIO_HOME / STUDIO_HOME like the installer, so a custom-root install is discovered without UNSLOTH_SD_CPP_PATH; start sd-cli with the parent-death child_popen_kwargs so it is not orphaned on a backend crash; reap the SIGKILLed child (proc.wait) so a cancel/timeout does not leave a zombie. - tests: Zip-Slip rejection, normal extraction, studio-home discovery. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio diffusion (Phase 4) review round 2: collect sd-cli batch outputs Codex review: when batch_count > 1, stable-diffusion.cpp's save_results() writes the numbered files <stem>_<idx><suffix> (base_0.png, base_1.png, ...) instead of the literal --output path. SdCppEngine.generate checked only the literal path, so a batch generation would exit 0 and then raise 'no image' (or return a stale file). generate now returns the literal path when present and otherwise falls back to the numbered siblings; single-image behavior is unchanged. Test: a fake sd-cli that writes img_0.png/img_1.png (not img.png) is collected without error. --------- Co-authored-by: oobabooga <112222186+oobabooga@users.noreply.github.com>
360 lines
14 KiB
Python
360 lines
14 KiB
Python
# SPDX-License-Identifier: AGPL-3.0-only
|
|
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
|
|
|
"""Install a prebuilt ``sd-cli`` (stable-diffusion.cpp) for the native diffusion
|
|
engine.
|
|
|
|
The chat backend ships a prebuilt llama-server; this is the diffusion analogue,
|
|
kept deliberately small. stable-diffusion.cpp publishes per-platform release
|
|
zips (macOS-arm64/Metal, Linux x86_64 CPU, plus Vulkan / ROCm / Windows
|
|
variants), so on the Phase-4 targets (Apple Silicon and CPU) there is nothing to
|
|
compile: resolve the right asset, download, extract into
|
|
``~/.unsloth/stable-diffusion.cpp``, and the engine's finder picks it up.
|
|
|
|
``resolve_release_asset`` -- the host -> asset choice -- is a pure function so the
|
|
matching matrix is unit-tested without any network. CUDA / ROCm / XPU hosts stay
|
|
on diffusers and never need this; it exists for the engines diffusers serves
|
|
poorly.
|
|
|
|
Usage:
|
|
python studio/install_sd_cpp_prebuilt.py # auto-detect host
|
|
python studio/install_sd_cpp_prebuilt.py --accelerator vulkan
|
|
python studio/install_sd_cpp_prebuilt.py --print-asset # resolve only
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import argparse
|
|
import hashlib
|
|
import json
|
|
import os
|
|
import platform
|
|
import shutil
|
|
import stat
|
|
import sys
|
|
import urllib.error
|
|
import urllib.request
|
|
import zipfile
|
|
from pathlib import Path
|
|
from typing import Optional, Sequence
|
|
|
|
# Default upstream source. Overridable with UNSLOTH_SD_CPP_REPO so a pinned unslothai
|
|
# mirror (built the same way as unslothai/llama.cpp's prebuilts) can be used without a
|
|
# code change once it exists; otherwise this falls back to leejet upstream.
|
|
DEFAULT_REPO = "leejet/stable-diffusion.cpp"
|
|
# Pinned release tag for REPRODUCIBILITY: "releases/latest" silently swaps the binary
|
|
# under users on every upstream push. Override with UNSLOTH_SD_CPP_TAG; set it empty to
|
|
# track latest. If the pinned tag is gone upstream, install falls back to latest.
|
|
DEFAULT_TAG = "master-737-3b6c9ca"
|
|
|
|
# Back-compat alias (some callers/tests import REPO).
|
|
REPO = DEFAULT_REPO
|
|
|
|
|
|
def _repo() -> str:
|
|
return (os.environ.get("UNSLOTH_SD_CPP_REPO") or DEFAULT_REPO).strip() or DEFAULT_REPO
|
|
|
|
|
|
def _pinned_tag() -> Optional[str]:
|
|
"""The release tag to install: env override, else the pinned default; '' = latest."""
|
|
val = os.environ.get("UNSLOTH_SD_CPP_TAG", DEFAULT_TAG).strip()
|
|
return val or None
|
|
|
|
|
|
# accelerator -> the token that must appear in a Linux/Windows asset name.
|
|
_LINUX_ACCEL_TOKEN = {"rocm": "rocm", "vulkan": "vulkan"}
|
|
_WINDOWS_ACCEL_TOKEN = {
|
|
"cuda": "cuda12",
|
|
"vulkan": "vulkan",
|
|
"rocm": "rocm",
|
|
"cpu": "avx2",
|
|
"auto": "avx2",
|
|
}
|
|
# Tokens that mark an accelerator-specific Linux build; "auto"/"cpu" want none of them.
|
|
_LINUX_ACCEL_MARKERS = ("rocm", "vulkan", "cuda", "sycl", "musa")
|
|
|
|
_ARCH_TOKENS = {
|
|
"x86_64": ("x86_64", "x64", "amd64"),
|
|
"amd64": ("x86_64", "x64", "amd64"),
|
|
"arm64": ("arm64", "aarch64"),
|
|
"aarch64": ("arm64", "aarch64"),
|
|
}
|
|
|
|
|
|
def _arch_tokens(machine: str) -> tuple[str, ...]:
|
|
return _ARCH_TOKENS.get(machine.lower(), (machine.lower(),))
|
|
|
|
|
|
def resolve_release_asset(
|
|
asset_names: Sequence[str],
|
|
*,
|
|
system: str,
|
|
machine: str,
|
|
accelerator: str = "auto",
|
|
) -> Optional[str]:
|
|
"""Pick the best release asset for a host, or None if none matches.
|
|
|
|
``system`` / ``machine`` are ``platform.system()`` / ``platform.machine()``
|
|
values; ``accelerator`` is ``auto`` (CPU/Metal default), ``vulkan``,
|
|
``rocm``, or ``cuda`` (Windows only). Pure -- the caller passes the release's
|
|
asset name list.
|
|
"""
|
|
system = system.lower()
|
|
accel = accelerator.lower()
|
|
arch = _arch_tokens(machine)
|
|
zips = [
|
|
a for a in asset_names if a.lower().endswith(".zip") and not a.lower().startswith("cudart")
|
|
]
|
|
|
|
if system == "darwin":
|
|
pool = [
|
|
a
|
|
for a in zips
|
|
if ("darwin" in a.lower() or "macos" in a.lower()) and any(t in a.lower() for t in arch)
|
|
]
|
|
return pool[0] if pool else None
|
|
|
|
if system == "windows":
|
|
pool = [a for a in zips if "bin-win" in a.lower()]
|
|
token = _WINDOWS_ACCEL_TOKEN.get(accel, accel)
|
|
sel = [a for a in pool if token in a.lower()]
|
|
if not sel: # fall back to a plain avx2 CPU build
|
|
sel = [a for a in pool if "avx2" in a.lower()]
|
|
return sel[0] if sel else (pool[0] if pool else None)
|
|
|
|
# linux (and anything else unix-like)
|
|
pool = [a for a in zips if "linux" in a.lower() and any(t in a.lower() for t in arch)]
|
|
if accel in _LINUX_ACCEL_TOKEN:
|
|
sel = [a for a in pool if _LINUX_ACCEL_TOKEN[accel] in a.lower()]
|
|
else: # auto / cpu -> the plain build with no accelerator marker
|
|
sel = [a for a in pool if not any(m in a.lower() for m in _LINUX_ACCEL_MARKERS)]
|
|
return sel[0] if sel else None
|
|
|
|
|
|
def _fetch_release(
|
|
tag: Optional[str] = None,
|
|
*,
|
|
repo: Optional[str] = None,
|
|
token: Optional[str] = None,
|
|
timeout: float = 30.0,
|
|
) -> dict:
|
|
"""GET a release JSON from GitHub. With ``tag`` set, fetch that exact release (and fall
|
|
back to latest if the tag is gone upstream); otherwise fetch latest. ``token`` is
|
|
optional and lifts the API rate limit."""
|
|
repo = repo or _repo()
|
|
token = token or os.environ.get("GH_TOKEN") or os.environ.get("GITHUB_TOKEN")
|
|
|
|
def _get(url: str) -> dict:
|
|
req = urllib.request.Request(url, headers = {"Accept": "application/vnd.github+json"})
|
|
if token:
|
|
req.add_header("Authorization", f"Bearer {token}")
|
|
with urllib.request.urlopen(req, timeout = timeout) as resp: # noqa: S310 (fixed https host)
|
|
return json.loads(resp.read().decode("utf-8"))
|
|
|
|
base = f"https://api.github.com/repos/{repo}/releases"
|
|
if tag:
|
|
try:
|
|
return _get(f"{base}/tags/{tag}")
|
|
except urllib.error.HTTPError as exc: # pinned tag removed upstream -> latest
|
|
if exc.code != 404:
|
|
raise
|
|
print(
|
|
f"sd-cli: pinned tag {tag} not found on {repo}; falling back to latest", flush = True
|
|
)
|
|
return _get(f"{base}/latest")
|
|
|
|
|
|
# Back-compat alias: the old name fetched latest.
|
|
def _fetch_latest_release(*, token: Optional[str] = None, timeout: float = 30.0) -> dict:
|
|
return _fetch_release(None, token = token, timeout = timeout)
|
|
|
|
|
|
def _verify_sha256(path: Path, expected_digest: Optional[str]) -> None:
|
|
"""Verify ``path`` against a GitHub asset ``digest`` ('sha256:<hex>'). Integrity check
|
|
against a corrupted/tampered download before we extract + execute the binary. When the
|
|
release publishes no digest (older releases), warn and proceed rather than hard-fail."""
|
|
if not expected_digest:
|
|
print(f"sd-cli: WARNING no digest for {path.name}; cannot verify integrity", flush = True)
|
|
return
|
|
algo, _, want = expected_digest.partition(":")
|
|
if algo.lower() != "sha256" or not want:
|
|
print(
|
|
f"sd-cli: WARNING unrecognised digest {expected_digest!r}; skipping check", flush = True
|
|
)
|
|
return
|
|
h = hashlib.sha256()
|
|
with open(path, "rb") as f:
|
|
for chunk in iter(lambda: f.read(1 << 20), b""):
|
|
h.update(chunk)
|
|
got = h.hexdigest()
|
|
if got != want.lower():
|
|
raise RuntimeError(f"sha256 mismatch for {path.name}: expected {want.lower()}, got {got}")
|
|
|
|
|
|
def default_install_dir() -> Path:
|
|
"""``~/.unsloth/stable-diffusion.cpp`` (or under ``UNSLOTH_STUDIO_HOME`` /
|
|
``STUDIO_HOME`` if set), the sibling of the llama.cpp install the finder
|
|
probes."""
|
|
home = os.environ.get("UNSLOTH_STUDIO_HOME") or os.environ.get("STUDIO_HOME")
|
|
base = Path(home).parent if home else Path.home() / ".unsloth"
|
|
return base / "stable-diffusion.cpp"
|
|
|
|
|
|
def _make_executable(path: Path) -> None:
|
|
mode = path.stat().st_mode
|
|
path.chmod(mode | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
|
|
|
|
|
|
def _locate_sd_cli(root: Path) -> Optional[Path]:
|
|
name = "sd-cli.exe" if sys.platform == "win32" else "sd-cli"
|
|
for p in root.rglob(name):
|
|
if p.is_file():
|
|
return p
|
|
return None
|
|
|
|
|
|
def _download(
|
|
url: str,
|
|
dest: Path,
|
|
*,
|
|
timeout: float = 300.0,
|
|
) -> None:
|
|
"""Stream ``url`` to ``dest`` with an explicit timeout. ``urlretrieve`` takes no
|
|
timeout and can hang forever on a stalled socket. A User-Agent is set because the
|
|
GitHub asset CDN can reject header-less requests; the API fetch carries any token."""
|
|
import shutil
|
|
|
|
req = urllib.request.Request(url, headers = {"User-Agent": "unsloth-sd-cpp-installer"})
|
|
with urllib.request.urlopen(req, timeout = timeout) as resp, open(dest, "wb") as f: # noqa: S310
|
|
shutil.copyfileobj(resp, f)
|
|
|
|
|
|
def _safe_extractall(zf: zipfile.ZipFile, target: Path) -> None:
|
|
"""``extractall`` with a per-member containment check, so an archive carrying an
|
|
absolute path or a ``..`` entry can't write outside ``target`` (Zip-Slip)."""
|
|
base = target.resolve()
|
|
for member in zf.infolist():
|
|
dest = (base / member.filename).resolve()
|
|
if dest != base and base not in dest.parents:
|
|
raise RuntimeError(f"unsafe path in archive: {member.filename!r}")
|
|
zf.extractall(target)
|
|
|
|
|
|
def _maybe_fetch_windows_cudart(release: dict, chosen: str, target: Path) -> None:
|
|
"""On Windows + a CUDA build, also fetch the separate CUDA-runtime DLL archive.
|
|
|
|
Upstream ships the runtime as ``cudart-sd-...-win-cu12-...zip`` (which
|
|
``resolve_release_asset`` filters out); without those DLLs ``sd-cli.exe`` cannot start
|
|
on a machine that does not already have the CUDA runtime installed."""
|
|
if platform.system().lower() != "windows" or "cuda" not in chosen.lower():
|
|
return
|
|
cudart = next(
|
|
(
|
|
a
|
|
for a in release.get("assets", [])
|
|
if a["name"].lower().startswith("cudart") and "win" in a["name"].lower()
|
|
),
|
|
None,
|
|
)
|
|
if cudart is None:
|
|
return
|
|
dest = target / cudart["name"]
|
|
print(f"downloading CUDA runtime {cudart['name']} ...", flush = True)
|
|
try:
|
|
_download(cudart["browser_download_url"], dest)
|
|
with zipfile.ZipFile(dest) as zf:
|
|
_safe_extractall(zf, target)
|
|
finally:
|
|
dest.unlink(missing_ok = True)
|
|
|
|
|
|
def install(
|
|
*,
|
|
install_dir: Optional[Path] = None,
|
|
accelerator: str = "auto",
|
|
token: Optional[str] = None,
|
|
) -> Path:
|
|
"""Download + extract the prebuilt for this host. Returns the sd-cli path.
|
|
|
|
Raises ``RuntimeError`` if no asset matches the host (the caller should then
|
|
build from source) or the archive has no ``sd-cli``.
|
|
"""
|
|
target = install_dir or default_install_dir()
|
|
release = _fetch_release(_pinned_tag(), token = token)
|
|
print(f"sd-cli: source {_repo()} release {release.get('tag_name', '?')}", flush = True)
|
|
names = [a["name"] for a in release.get("assets", [])]
|
|
chosen = resolve_release_asset(
|
|
names,
|
|
system = platform.system(),
|
|
machine = platform.machine(),
|
|
accelerator = accelerator,
|
|
)
|
|
if not chosen:
|
|
raise RuntimeError(
|
|
f"No prebuilt sd-cli for {platform.system()}/{platform.machine()} "
|
|
f"(accelerator={accelerator}). Build from source: "
|
|
f"https://github.com/{_repo()}"
|
|
)
|
|
asset = next(a for a in release["assets"] if a["name"] == chosen)
|
|
url = asset["browser_download_url"]
|
|
target.mkdir(parents = True, exist_ok = True)
|
|
archive = target / chosen
|
|
print(f"downloading {chosen} -> {archive}", flush = True)
|
|
try:
|
|
_download(url, archive)
|
|
# Verify integrity BEFORE extracting + executing.
|
|
_verify_sha256(archive, asset.get("digest"))
|
|
print("extracting ...", flush = True)
|
|
with zipfile.ZipFile(archive) as zf:
|
|
_safe_extractall(zf, target)
|
|
# Windows CUDA builds need the separately-published cudart runtime DLLs.
|
|
_maybe_fetch_windows_cudart(release, chosen, target)
|
|
finally:
|
|
# Always drop the archive: on a sha256 mismatch / corrupt zip / network error it
|
|
# must not linger (and a stale partial would defeat a later retry).
|
|
archive.unlink(missing_ok = True)
|
|
sd_cli = _locate_sd_cli(target)
|
|
if not sd_cli:
|
|
raise RuntimeError(f"archive {chosen} contained no sd-cli binary")
|
|
if sys.platform != "win32":
|
|
_make_executable(sd_cli)
|
|
print(f"installed sd-cli -> {sd_cli}", flush = True)
|
|
return sd_cli
|
|
|
|
|
|
def main(argv: Optional[list[str]] = None) -> int:
|
|
p = argparse.ArgumentParser(description = "Install a prebuilt sd-cli (stable-diffusion.cpp).")
|
|
p.add_argument(
|
|
"--accelerator", default = "auto", choices = ["auto", "cpu", "vulkan", "rocm", "cuda"]
|
|
)
|
|
p.add_argument("--install-dir", default = None)
|
|
p.add_argument(
|
|
"--print-asset", action = "store_true", help = "resolve + print the asset, don't download"
|
|
)
|
|
args = p.parse_args(argv)
|
|
|
|
if args.print_asset:
|
|
release = _fetch_release(_pinned_tag())
|
|
names = [a["name"] for a in release.get("assets", [])]
|
|
chosen = resolve_release_asset(
|
|
names,
|
|
system = platform.system(),
|
|
machine = platform.machine(),
|
|
accelerator = args.accelerator,
|
|
)
|
|
print(chosen or "(no matching prebuilt; build from source)")
|
|
return 0 if chosen else 2
|
|
|
|
try:
|
|
install(
|
|
install_dir = Path(args.install_dir).expanduser() if args.install_dir else None,
|
|
accelerator = args.accelerator,
|
|
)
|
|
except RuntimeError as exc:
|
|
print(f"error: {exc}", file = sys.stderr)
|
|
return 1
|
|
return 0
|
|
|
|
|
|
if __name__ == "__main__":
|
|
raise SystemExit(main())
|