unsloth/studio/install_sd_cpp_prebuilt.py
Daniel Han 96940d87b7
Studio diffusion (Phase 4): native stable-diffusion.cpp engine for CPU/Mac (#6679)
* Studio diffusion: cross-platform device policy, fp16 guard, lock split, validate-before-evict

Phase 1 of porting the richer diffusion stack onto the image-generation backend.

- Add a compartmentalized device/dtype policy module (diffusion_device.py)
  resolving CUDA/ROCm/XPU/MPS/CPU with capability flags. Keeps the NVIDIA
  capability-based bf16 choice; ROCm and XPU are isolated; MPS uses bf16 or
  fp32, never a silent fp16 that renders a black image.
- Add a per-family fp16_incompatible flag (Z-Image) and promote a resolved
  float16 to float32 for those families so they do not produce black images.
- Split the backend locks: a generation holds only _generate_lock, so status,
  unload, and a new load are never blocked by a long denoise. Add per-generation
  cancellation via callback_on_step_end so an eviction or a superseding load
  preempts a running generation; a replacement load waits for it to stop before
  allocating, so two pipelines never sit in VRAM at once.
- Validate a load request before the GPU handoff so an unloadable pick never
  evicts a working chat model, and reject missing local paths up front.
- Add CPU-only tests for the device policy, dtype guard, lock split and
  cancellation, and validate-before-evict, plus a GPU benchmark/regression
  script (scripts/diffusion_bench.py) measuring latency, peak VRAM, and PSNR
  against a saved reference.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio diffusion (Phase 2A): measured-budget memory planner + offload/VAE policy

Add a lean, backend-agnostic memory policy that picks a CPU-offload policy and
VAE tiling/slicing from measured free device memory vs the model's estimated
resident footprint, then applies it to the built pipeline. auto stays resident
when the model fits (byte-identical to the prior resident path), and falls to
whole-module offload when tight; fast/balanced/low_vram are explicit overrides.
Sequential submodule offload is unreliable for GGUF transformers on diffusers
0.38, so it falls back to whole-module offload and status reports the policy
actually engaged.

Verified on Z-Image-Turbo Q4_K_M (B200): auto reproduces the resident image with
no VRAM/latency regression (PSNR inf); balanced/low_vram cut generation peak VRAM
47.9% (15951 -> 8318 MB) with byte-identical output, at the expected latency cost.

73 prior + 35 new CPU tests pass.

* Studio diffusion (Phase 2D): streamed block-level offload + functional VAE tiling

Add a streamed 'group' offload tier (diffusers apply_group_offloading, block_level,
use_stream) that keeps the transformer flowing through the GPU a few blocks at a
time while the text encoder / VAE stay resident, and fix VAE tiling to drive the
VAE submodule (pipelines like Z-Image expose enable_tiling on pipe.vae, not the
pipeline). apply_memory_plan now returns the (policy, tiling) actually engaged so
status never overstates either, and group falls back to whole-module offload when
the transformer can't be streamed.

Measured on Z-Image (B200), all lossless (PSNR inf vs resident): balanced/group
cuts generation peak VRAM 32% (15951 -> 10840 MB) at near-resident speed (2.07 ->
2.99s); low_vram/model cuts it 48% (-> 8318 MB) but is slower (7.99s). Mode names
now match that tradeoff: balanced = stream the transformer, low_vram = offload
every component. auto picks group when the companions fit resident, else model.

112 CPU tests pass.

* Studio diffusion (Phase 5): image quality-vs-quant accuracy harness

Add scripts/diffusion_quality.py, the accuracy analogue of the KLD workflow: hold
prompt + seed fixed, render a grid with a reference quant (default BF16), then render
each candidate quant and measure drift from the reference. Records mean PSNR + SSIM
(pure-numpy, no skimage/scipy) and optional CLIP text-alignment + image-similarity
(transformers, --clip), plus file size, latency, and peak VRAM, then prints a
quality-vs-cost table and recommends the smallest quant within a quality budget.
--selftest validates the metrics on synthetic images with no GPU or model.

Verified on Z-Image (B200): the table degrades monotonically with quant size
(Q8 -> Q4 -> Q2: PSNR 21.7 -> 15.5, SSIM 0.82 -> 0.61), while CLIP-text stays flat
(~0.34) -- quantization erodes fine detail far more than prompt adherence.

* Studio diffusion (Phase 3): opt-in speed layer (channels_last / compile / TF32)

Add a speed_mode knob (off by default, so the render path stays bit-identical):
default applies channels_last VAE + regional torch.compile of the denoiser's
repeated block where eligible; max also enables TF32 matmul and fused QKV. Regional
compile is gated off for the GGUF transformer (dequantises per-op) and for families
flagged not compile-friendly (a new supports_torch_compile flag, False for Z-Image),
so it activates automatically only once a non-GGUF bf16 transformer is loaded. Speed
optims run before placement/offload, per the diffusers composition order. status now
reports speed_mode + the optims actually engaged.

Verified on Z-Image (B200): default -> ['channels_last'], max -> ['channels_last',
'tf32'], compile correctly skipped for GGUF; generation works in every mode.

121 CPU tests pass.

* Studio diffusion (Phase 2B): opt-in fp8 text-encoder layerwise casting

Add a text_encoder_fp8 knob that casts the companion text encoder(s) to fp8 (e4m3)
storage via diffusers apply_layerwise_casting, upcasting per layer to the bf16
compute dtype while normalisations and embeddings stay full precision. Applied
before placement, gated to CUDA + bf16, best-effort (a failure leaves the encoder
dense). status reports which encoders were cast.

Verified on Z-Image (B200, balanced/group mode where the encoder stays resident):
generation peak VRAM dropped 37% (10840 -> 6791 MB, below the lowest-VRAM offload)
at near-resident speed. It is a memory-vs-quality tradeoff, not free -- ~20 dB PSNR
vs the bf16 encoder, a larger shift than one transformer quant step -- so it is off
by default and documented as such, with the Phase 5 harness to size the cost.

127 CPU tests pass.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio diffusion (Phase 2C): NVFP4 text-encoder quant (+ generalise fp8 knob)

Generalise the text-encoder precision knob from a fp8 bool to text_encoder_quant
(fp8 | nvfp4). nvfp4 quantises the companion text encoder to 4-bit via torchao
NVFP4 weight-only (two-level microscaling) on Blackwell's FP4 tensor cores; fp8
stays the broader-hardware path (cc>=8.9). Both are gated, best-effort, and run
before placement; status reports the mode actually engaged. This is the lean
realisation of GGUF-native text-encoder quant: 4-bit on the encoder without the
3045-line port.

Verified on Z-Image (B200, balanced/group where the encoder stays resident), vs the
bf16 encoder: nvfp4 cut generation peak VRAM 48% (10840 -> 5593 MB, the lowest TE
option, below whole-model offload) at near-fp8 quality (16.4 vs 17.1 dB PSNR), and
both quants ran faster than bf16. A memory-vs-quality tradeoff (off by default);
size it per model with the Phase 5 quality harness. diffusion_bench gains
--text-encoder-quant.

129 CPU tests pass.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio diffusion (Phase 4): native stable-diffusion.cpp engine for CPU/Mac

Adds the CPU / Apple-Silicon tier of the two-engine strategy, mirroring the
chat backend's llama.cpp shell-out. Diffusers stays the default on CUDA / ROCm
/ XPU; this covers the hardware diffusers serves poorly, consuming the same
split GGUF assets Studio already curates.

- sd_cpp_args.py: pure sd-cli command builder. Maps the family to its
  text-encoder flag (Z-Image Qwen3 to --llm, Qwen-Image to --qwen2vl, FLUX.1
  CLIP-L + T5), and the diffusers memory policy (none/group/model/sequential)
  to sd.cpp's offload flags (--offload-to-cpu / --clip-on-cpu / --vae-on-cpu /
  --vae-tiling / --diffusion-fa), so one user knob drives both engines.
- sd_cpp_engine.py: SdCppEngine over a located sd-cli. find_sd_cpp_binary()
  with the same precedence as the llama finder (env override, then the Studio
  install root, then in-tree, then PATH), an is_available/version probe, and a
  one-shot subprocess generate that streams progress and returns the PNG.
  runtime_env() prepends the binary's directory to the platform library path
  so a prebuilt's bundled libstable-diffusion.so resolves.
  select_diffusion_engine() is the pure routing decision (GPU backends to
  diffusers, CPU/MPS to native when present).
- install_sd_cpp_prebuilt.py: resolve + download the per-host prebuilt
  (macOS-arm64/Metal, Linux x86_64 CPU, Vulkan/ROCm/Windows variants) into the
  Studio install root. resolve_release_asset() is a pure, unit-tested
  host-to-asset matrix.
- scripts/sd_cpp_smoke.py: end-to-end native generation harness.

Tests (CPU-only, subprocess/filesystem stubbed): 49 new across args, engine,
routing, runtime env, and the installer resolver. Full diffusion suite 166
passing.

Verified on a B200 box: built sd-cli (CUDA) and the prebuilt (CPU) both
generate Z-Image-Turbo Q4_K end to end through SdCppEngine: balanced (group
offload, 5.0s gen), low_vram (full CPU offload + VAE tiling, 13.4s), and the
dynamically-linked CPU prebuilt (50.4s on CPU), all producing coherent images.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio diffusion (Phase 4): enforce the sd-cli timeout while reading output

Iterating proc.stdout directly blocks until the stream closes, so a sd-cli that hangs
without producing output (or without closing stdout) would never reach proc.wait and the
wall-clock timeout was silently bypassed. Drain stdout on a daemon thread and wait on the
PROCESS, so the main thread always enforces the timeout and kills a hung process (which
closes the pipe and ends the reader). Add a test that times out even when stdout blocks,
and make the no-binary test hermetic so a host-installed sd-cli can't leak in.

* Studio diffusion (Phase 4) review fixes: sd.cpp installer + engine hardening

- install_sd_cpp_prebuilt: download the release archive with urlopen + an explicit
  timeout + copyfileobj (urlretrieve has no timeout and hangs on a stalled socket);
  extract through a per-member containment check (Zip-Slip guard); expanduser the
  --install-dir so a tilde path is not taken literally; and on Windows CUDA also fetch
  the separately-published cudart runtime DLL archive so sd-cli.exe can start.
- sd_cpp_engine: find_sd_cpp_binary honors UNSLOTH_STUDIO_HOME / STUDIO_HOME like the
  installer, so a custom-root install is discovered without UNSLOTH_SD_CPP_PATH; start
  sd-cli with the parent-death child_popen_kwargs so it is not orphaned on a backend
  crash; reap the SIGKILLed child (proc.wait) so a cancel/timeout does not leave a zombie.
- tests: Zip-Slip rejection, normal extraction, studio-home discovery.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio diffusion (Phase 4) review round 2: collect sd-cli batch outputs

Codex review: when batch_count > 1, stable-diffusion.cpp's save_results() writes
the numbered files <stem>_<idx><suffix> (base_0.png, base_1.png, ...) instead of
the literal --output path. SdCppEngine.generate checked only the literal path, so
a batch generation would exit 0 and then raise 'no image' (or return a stale
file). generate now returns the literal path when present and otherwise falls
back to the numbered siblings; single-image behavior is unchanged.

Test: a fake sd-cli that writes img_0.png/img_1.png (not img.png) is collected
without error.

---------

Co-authored-by: oobabooga <112222186+oobabooga@users.noreply.github.com>
2026-07-01 15:03:53 -03:00

360 lines
14 KiB
Python

# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""Install a prebuilt ``sd-cli`` (stable-diffusion.cpp) for the native diffusion
engine.
The chat backend ships a prebuilt llama-server; this is the diffusion analogue,
kept deliberately small. stable-diffusion.cpp publishes per-platform release
zips (macOS-arm64/Metal, Linux x86_64 CPU, plus Vulkan / ROCm / Windows
variants), so on the Phase-4 targets (Apple Silicon and CPU) there is nothing to
compile: resolve the right asset, download, extract into
``~/.unsloth/stable-diffusion.cpp``, and the engine's finder picks it up.
``resolve_release_asset`` -- the host -> asset choice -- is a pure function so the
matching matrix is unit-tested without any network. CUDA / ROCm / XPU hosts stay
on diffusers and never need this; it exists for the engines diffusers serves
poorly.
Usage:
python studio/install_sd_cpp_prebuilt.py # auto-detect host
python studio/install_sd_cpp_prebuilt.py --accelerator vulkan
python studio/install_sd_cpp_prebuilt.py --print-asset # resolve only
"""
from __future__ import annotations
import argparse
import hashlib
import json
import os
import platform
import shutil
import stat
import sys
import urllib.error
import urllib.request
import zipfile
from pathlib import Path
from typing import Optional, Sequence
# Default upstream source. Overridable with UNSLOTH_SD_CPP_REPO so a pinned unslothai
# mirror (built the same way as unslothai/llama.cpp's prebuilts) can be used without a
# code change once it exists; otherwise this falls back to leejet upstream.
DEFAULT_REPO = "leejet/stable-diffusion.cpp"
# Pinned release tag for REPRODUCIBILITY: "releases/latest" silently swaps the binary
# under users on every upstream push. Override with UNSLOTH_SD_CPP_TAG; set it empty to
# track latest. If the pinned tag is gone upstream, install falls back to latest.
DEFAULT_TAG = "master-737-3b6c9ca"
# Back-compat alias (some callers/tests import REPO).
REPO = DEFAULT_REPO
def _repo() -> str:
return (os.environ.get("UNSLOTH_SD_CPP_REPO") or DEFAULT_REPO).strip() or DEFAULT_REPO
def _pinned_tag() -> Optional[str]:
"""The release tag to install: env override, else the pinned default; '' = latest."""
val = os.environ.get("UNSLOTH_SD_CPP_TAG", DEFAULT_TAG).strip()
return val or None
# accelerator -> the token that must appear in a Linux/Windows asset name.
_LINUX_ACCEL_TOKEN = {"rocm": "rocm", "vulkan": "vulkan"}
_WINDOWS_ACCEL_TOKEN = {
"cuda": "cuda12",
"vulkan": "vulkan",
"rocm": "rocm",
"cpu": "avx2",
"auto": "avx2",
}
# Tokens that mark an accelerator-specific Linux build; "auto"/"cpu" want none of them.
_LINUX_ACCEL_MARKERS = ("rocm", "vulkan", "cuda", "sycl", "musa")
_ARCH_TOKENS = {
"x86_64": ("x86_64", "x64", "amd64"),
"amd64": ("x86_64", "x64", "amd64"),
"arm64": ("arm64", "aarch64"),
"aarch64": ("arm64", "aarch64"),
}
def _arch_tokens(machine: str) -> tuple[str, ...]:
return _ARCH_TOKENS.get(machine.lower(), (machine.lower(),))
def resolve_release_asset(
asset_names: Sequence[str],
*,
system: str,
machine: str,
accelerator: str = "auto",
) -> Optional[str]:
"""Pick the best release asset for a host, or None if none matches.
``system`` / ``machine`` are ``platform.system()`` / ``platform.machine()``
values; ``accelerator`` is ``auto`` (CPU/Metal default), ``vulkan``,
``rocm``, or ``cuda`` (Windows only). Pure -- the caller passes the release's
asset name list.
"""
system = system.lower()
accel = accelerator.lower()
arch = _arch_tokens(machine)
zips = [
a for a in asset_names if a.lower().endswith(".zip") and not a.lower().startswith("cudart")
]
if system == "darwin":
pool = [
a
for a in zips
if ("darwin" in a.lower() or "macos" in a.lower()) and any(t in a.lower() for t in arch)
]
return pool[0] if pool else None
if system == "windows":
pool = [a for a in zips if "bin-win" in a.lower()]
token = _WINDOWS_ACCEL_TOKEN.get(accel, accel)
sel = [a for a in pool if token in a.lower()]
if not sel: # fall back to a plain avx2 CPU build
sel = [a for a in pool if "avx2" in a.lower()]
return sel[0] if sel else (pool[0] if pool else None)
# linux (and anything else unix-like)
pool = [a for a in zips if "linux" in a.lower() and any(t in a.lower() for t in arch)]
if accel in _LINUX_ACCEL_TOKEN:
sel = [a for a in pool if _LINUX_ACCEL_TOKEN[accel] in a.lower()]
else: # auto / cpu -> the plain build with no accelerator marker
sel = [a for a in pool if not any(m in a.lower() for m in _LINUX_ACCEL_MARKERS)]
return sel[0] if sel else None
def _fetch_release(
tag: Optional[str] = None,
*,
repo: Optional[str] = None,
token: Optional[str] = None,
timeout: float = 30.0,
) -> dict:
"""GET a release JSON from GitHub. With ``tag`` set, fetch that exact release (and fall
back to latest if the tag is gone upstream); otherwise fetch latest. ``token`` is
optional and lifts the API rate limit."""
repo = repo or _repo()
token = token or os.environ.get("GH_TOKEN") or os.environ.get("GITHUB_TOKEN")
def _get(url: str) -> dict:
req = urllib.request.Request(url, headers = {"Accept": "application/vnd.github+json"})
if token:
req.add_header("Authorization", f"Bearer {token}")
with urllib.request.urlopen(req, timeout = timeout) as resp: # noqa: S310 (fixed https host)
return json.loads(resp.read().decode("utf-8"))
base = f"https://api.github.com/repos/{repo}/releases"
if tag:
try:
return _get(f"{base}/tags/{tag}")
except urllib.error.HTTPError as exc: # pinned tag removed upstream -> latest
if exc.code != 404:
raise
print(
f"sd-cli: pinned tag {tag} not found on {repo}; falling back to latest", flush = True
)
return _get(f"{base}/latest")
# Back-compat alias: the old name fetched latest.
def _fetch_latest_release(*, token: Optional[str] = None, timeout: float = 30.0) -> dict:
return _fetch_release(None, token = token, timeout = timeout)
def _verify_sha256(path: Path, expected_digest: Optional[str]) -> None:
"""Verify ``path`` against a GitHub asset ``digest`` ('sha256:<hex>'). Integrity check
against a corrupted/tampered download before we extract + execute the binary. When the
release publishes no digest (older releases), warn and proceed rather than hard-fail."""
if not expected_digest:
print(f"sd-cli: WARNING no digest for {path.name}; cannot verify integrity", flush = True)
return
algo, _, want = expected_digest.partition(":")
if algo.lower() != "sha256" or not want:
print(
f"sd-cli: WARNING unrecognised digest {expected_digest!r}; skipping check", flush = True
)
return
h = hashlib.sha256()
with open(path, "rb") as f:
for chunk in iter(lambda: f.read(1 << 20), b""):
h.update(chunk)
got = h.hexdigest()
if got != want.lower():
raise RuntimeError(f"sha256 mismatch for {path.name}: expected {want.lower()}, got {got}")
def default_install_dir() -> Path:
"""``~/.unsloth/stable-diffusion.cpp`` (or under ``UNSLOTH_STUDIO_HOME`` /
``STUDIO_HOME`` if set), the sibling of the llama.cpp install the finder
probes."""
home = os.environ.get("UNSLOTH_STUDIO_HOME") or os.environ.get("STUDIO_HOME")
base = Path(home).parent if home else Path.home() / ".unsloth"
return base / "stable-diffusion.cpp"
def _make_executable(path: Path) -> None:
mode = path.stat().st_mode
path.chmod(mode | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
def _locate_sd_cli(root: Path) -> Optional[Path]:
name = "sd-cli.exe" if sys.platform == "win32" else "sd-cli"
for p in root.rglob(name):
if p.is_file():
return p
return None
def _download(
url: str,
dest: Path,
*,
timeout: float = 300.0,
) -> None:
"""Stream ``url`` to ``dest`` with an explicit timeout. ``urlretrieve`` takes no
timeout and can hang forever on a stalled socket. A User-Agent is set because the
GitHub asset CDN can reject header-less requests; the API fetch carries any token."""
import shutil
req = urllib.request.Request(url, headers = {"User-Agent": "unsloth-sd-cpp-installer"})
with urllib.request.urlopen(req, timeout = timeout) as resp, open(dest, "wb") as f: # noqa: S310
shutil.copyfileobj(resp, f)
def _safe_extractall(zf: zipfile.ZipFile, target: Path) -> None:
"""``extractall`` with a per-member containment check, so an archive carrying an
absolute path or a ``..`` entry can't write outside ``target`` (Zip-Slip)."""
base = target.resolve()
for member in zf.infolist():
dest = (base / member.filename).resolve()
if dest != base and base not in dest.parents:
raise RuntimeError(f"unsafe path in archive: {member.filename!r}")
zf.extractall(target)
def _maybe_fetch_windows_cudart(release: dict, chosen: str, target: Path) -> None:
"""On Windows + a CUDA build, also fetch the separate CUDA-runtime DLL archive.
Upstream ships the runtime as ``cudart-sd-...-win-cu12-...zip`` (which
``resolve_release_asset`` filters out); without those DLLs ``sd-cli.exe`` cannot start
on a machine that does not already have the CUDA runtime installed."""
if platform.system().lower() != "windows" or "cuda" not in chosen.lower():
return
cudart = next(
(
a
for a in release.get("assets", [])
if a["name"].lower().startswith("cudart") and "win" in a["name"].lower()
),
None,
)
if cudart is None:
return
dest = target / cudart["name"]
print(f"downloading CUDA runtime {cudart['name']} ...", flush = True)
try:
_download(cudart["browser_download_url"], dest)
with zipfile.ZipFile(dest) as zf:
_safe_extractall(zf, target)
finally:
dest.unlink(missing_ok = True)
def install(
*,
install_dir: Optional[Path] = None,
accelerator: str = "auto",
token: Optional[str] = None,
) -> Path:
"""Download + extract the prebuilt for this host. Returns the sd-cli path.
Raises ``RuntimeError`` if no asset matches the host (the caller should then
build from source) or the archive has no ``sd-cli``.
"""
target = install_dir or default_install_dir()
release = _fetch_release(_pinned_tag(), token = token)
print(f"sd-cli: source {_repo()} release {release.get('tag_name', '?')}", flush = True)
names = [a["name"] for a in release.get("assets", [])]
chosen = resolve_release_asset(
names,
system = platform.system(),
machine = platform.machine(),
accelerator = accelerator,
)
if not chosen:
raise RuntimeError(
f"No prebuilt sd-cli for {platform.system()}/{platform.machine()} "
f"(accelerator={accelerator}). Build from source: "
f"https://github.com/{_repo()}"
)
asset = next(a for a in release["assets"] if a["name"] == chosen)
url = asset["browser_download_url"]
target.mkdir(parents = True, exist_ok = True)
archive = target / chosen
print(f"downloading {chosen} -> {archive}", flush = True)
try:
_download(url, archive)
# Verify integrity BEFORE extracting + executing.
_verify_sha256(archive, asset.get("digest"))
print("extracting ...", flush = True)
with zipfile.ZipFile(archive) as zf:
_safe_extractall(zf, target)
# Windows CUDA builds need the separately-published cudart runtime DLLs.
_maybe_fetch_windows_cudart(release, chosen, target)
finally:
# Always drop the archive: on a sha256 mismatch / corrupt zip / network error it
# must not linger (and a stale partial would defeat a later retry).
archive.unlink(missing_ok = True)
sd_cli = _locate_sd_cli(target)
if not sd_cli:
raise RuntimeError(f"archive {chosen} contained no sd-cli binary")
if sys.platform != "win32":
_make_executable(sd_cli)
print(f"installed sd-cli -> {sd_cli}", flush = True)
return sd_cli
def main(argv: Optional[list[str]] = None) -> int:
p = argparse.ArgumentParser(description = "Install a prebuilt sd-cli (stable-diffusion.cpp).")
p.add_argument(
"--accelerator", default = "auto", choices = ["auto", "cpu", "vulkan", "rocm", "cuda"]
)
p.add_argument("--install-dir", default = None)
p.add_argument(
"--print-asset", action = "store_true", help = "resolve + print the asset, don't download"
)
args = p.parse_args(argv)
if args.print_asset:
release = _fetch_release(_pinned_tag())
names = [a["name"] for a in release.get("assets", [])]
chosen = resolve_release_asset(
names,
system = platform.system(),
machine = platform.machine(),
accelerator = args.accelerator,
)
print(chosen or "(no matching prebuilt; build from source)")
return 0 if chosen else 2
try:
install(
install_dir = Path(args.install_dir).expanduser() if args.install_dir else None,
accelerator = args.accelerator,
)
except RuntimeError as exc:
print(f"error: {exc}", file = sys.stderr)
return 1
return 0
if __name__ == "__main__":
raise SystemExit(main())