* feat(studio): use lemonade-sdk/llamacpp-rocm per-GPU prebuilts for ROCm hosts
For AMD GPUs that rocminfo/hipinfo reports a recognised gfx target
(gfx103X / gfx110X / gfx1150 / gfx1151 / gfx120X), resolve_lemonade_rocm_choice()
now fetches the latest lemonade-sdk/llamacpp-rocm release and returns the
matching per-architecture zip, bundling all required ROCm runtime libs.
This runs before the existing upstream ggml-org combined-ROCm tarball fallback
on Linux and before the upstream HIP zip on Windows, so both platforms benefit
from the more targeted build when available.
Changes:
- Add LEMONADE_ROCM_REPO / LEMONADE_ROCM_RELEASES_API constants
- Add HostInfo.rocm_gfx_target populated from rocminfo (Linux) / hipinfo (Windows)
- Add _LEMONADE_GFX_FAMILIES prefix map and _lemonade_gfx_family() helper
- Add resolve_lemonade_rocm_choice() that fetches latest lemonade release and
constructs the llama-{tag}-{os}-rocm-{gfxFamily}-x64.zip asset URL
- Wire into resolve_upstream_asset_choice() for both Linux (ubuntu) and Windows paths
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Honor pinned llama.cpp tag in lemonade ROCm resolver
The resolver always fetched lemonade-sdk/llamacpp-rocm's /releases/latest,
ignoring the upstream llama.cpp tag the caller had pinned. On a reproducible
install where the user requested 'b1260' that meant we would silently pick
up whatever lemonade had published as latest at install time, with no way
to roll back to the matching tag.
Lemonade tags llama.cpp upstream tags 1:1, so:
- When llama_tag is unset or 'latest', keep hitting /releases/latest.
- When llama_tag is pinned (e.g. 'b1260'), hit /releases/tags/b1260.
- When the pinned tag is not published by lemonade (404), skip silently
and let the caller fall through to the upstream tarball -- this keeps
pinned installs reproducible instead of drifting.
resolve_lemonade_rocm_choice now takes llama_tag (default 'latest' for
backward compatibility) and both call sites in resolve_upstream_asset_choice
forward the upstream llama_tag to it.
Note: this PR still has open integration concerns flagged in review --
the simple-policy planner and approved-checksum manifest don't yet route
or accept lemonade assets. Those are larger changes and not in scope for
this commit; addressing the pinned-tag drift independently because it is
small, localized, and self-contained.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Add mock test for lemonade ROCm prebuilt asset resolution
Validates GPU family mapping and that resolve_lemonade_rocm_choice
returns real lemonade release URLs for all supported gfx targets on
both Linux and Windows, without requiring AMD hardware.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Wire lemonade ROCm prebuilts into the simple-policy install path
setup.sh invokes install_llama_prebuilt.py with --simple-policy, which
dispatches through resolve_simple_install_release_plans -> direct_linux_release_plan
(or direct_upstream_release_plan on Windows). Those planners only
handled CUDA + CPU attempts, so ROCm-only hosts (e.g. gfx1151 Strix
Halo) had no compatible prebuilt asset and silently fell through to
source build, even though resolve_lemonade_rocm_choice already knew
how to fetch a per-GPU lemonade-sdk binary.
Add a lemonade ROCm/HIP attempt to both simple-policy planners for
ROCm-only hosts, and add regression tests that drive the dispatchers
end-to-end so this can't be skipped silently again.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Document that lemonade ROCm prebuilts work on any glibc Linux
The lemonade-sdk asset filename uses "ubuntu" as a label, but the binary
is a manylinux-style glibc build with no Ubuntu-specific dependencies.
It runs on Arch, Fedora, openSUSE, Debian, etc. as long as the host
glibc is recent enough.
No behavior change -- the dispatch already runs for any Linux ROCm
host. This commit only clarifies the comment, docstring, and log
message so users on non-Ubuntu distros (e.g. Strix Halo on Arch) don't
mistake the asset name for distro gating.
* Pattern matching fix for libggml-cpu*.so*
* fix(lemonade): pass resolved tag to lemonade resolver; add upstream HIP fallback; stub API in tests
- direct_linux_release_plan: pass bundle.upstream_tag (not requested_tag)
to resolve_lemonade_rocm_choice so a "latest" request doesn't mix a
newer lemonade binary with an older planned unsloth release (Codex P2)
- direct_upstream_release_plan: same fix on the Windows path (release_tag
instead of requested_tag); also add the upstream HIP asset
(llama-<tag>-bin-win-hip-radeon-x64.zip) as a fallback between lemonade
and CPU so unsupported GPUs or transient lemonade failures don't silently
downgrade to CPU when an upstream ROCm prebuilt exists (Codex P2)
- test file: stub fetch_json with a synthetic lemonade release payload so
the suite is hermetic and not subject to GitHub API rate limits (Codex P1);
add test_simple_policy_windows_hip_falls_back_to_upstream_when_lemonade_unavailable
to cover the new HIP fallback path
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix(lemonade): revert bad tag fix; keep upstream HIP fallback + hermetic tests
The previous commit wrongly passed bundle.upstream_tag / release_tag to
resolve_lemonade_rocm_choice. Lemonade uses its own versioning (b1262,
b1264, …) completely independent of unslothai's tags (b9186, …), so
passing a resolved unslothai tag caused a 404 and silently skipped the
lemonade binary entirely. Revert both call sites to requested_tag.
Keep the two valid fixes from the prior commit:
- Upstream HIP fallback (llama-<tag>-bin-win-hip-radeon-x64.zip) between
lemonade and CPU in direct_upstream_release_plan, so unsupported GPUs
or transient lemonade failures don't silently downgrade to CPU (Codex P2)
- Stub fetch_json in tests so the suite is hermetic (Codex P1)
* fix(studio/rocm): respect HIP_VISIBLE_DEVICES when picking lemonade gfx target
The rocminfo / hipinfo regex took the first gfx match in the agent listing.
On mixed APU + dGPU hosts (e.g. Strix Halo gfx1151 + discrete RX 7900 gfx1100)
this picked whichever GPU appeared first in the tool's stdout, not the one
HIP actually runs on. The downloaded lemonade asset could then be a binary
for a different arch than the active device.
Extracted a module-level _pick_rocm_gfx_target() helper that:
- collects every gfx token in order via re.findall (skips gfx000 / generic ISAs)
- if HIP_VISIBLE_DEVICES or ROCR_VISIBLE_DEVICES is set, parses the first
comma-separated entry as an integer index into that list
- falls back to the first GPU for non-integer (UUID-style) or out-of-range
values, matching the previous default behaviour
Both Linux (rocminfo) and Windows (hipinfo) branches use the helper.
Existing 18 lemonade tests pass; no behavioural change for single-GPU hosts.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix(studio/rocm): dedup rocminfo gfx tokens and honor disabled visibility
Follow-up to 8793aef0. rocminfo / hipinfo emit each gfx target multiple
times per GPU (Name, ISA triple, marketing name), so the prior re.findall
indexing returned the wrong device when HIP_VISIBLE_DEVICES picked GPU 1
on a mixed-arch host -- the helper picked the second occurrence of GPU 0
instead. Collapse to unique tokens (insertion-ordered) before indexing.
Also handle HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES values of '' and
'-1' as "no AMD visible" (matches the rest of Studio's visibility code),
returning None so the planner does not pick a Lemonade asset for a
hidden GPU.
* fix(studio/rocm): memoise lemonade release lookup + HIP_PATH hipinfo fallback
Robustness pass on top of 25d4ab63:
1. resolve_lemonade_rocm_choice() is called twice per install (planner +
resolve_upstream_asset_choice), so every install previously hit
api.github.com twice with identical args -- doubling the 403/rate-limit
failure surface on busy CI runners. Extract the fetch into a
functools.lru_cache(maxsize=8) helper keyed on (api_url, llama_tag).
The cached helper also owns the error-path logging the resolver was
doing inline. Tests that need to vary fetch_json output across
invocations should call _fetch_lemonade_release_cached.cache_clear().
2. Windows detect_host probe for hipinfo / amd-smi only used
shutil.which(), so HIP SDK installs that set HIP_PATH but do not put
%HIP_PATH%\bin on system PATH classified the host as non-ROCm. The
PowerShell installer and studio/install_python_stack.py already
resolve HIP_PATH\bin\hipinfo.exe as a fallback; mirror that here so
the install planner agrees with the rest of Studio on what counts as
a ROCm host.
Tests: 18 passed (shipped); 52 extra sim cases pass (added 2 for the
lru_cache deduplication path).
* fix(studio/rocm): lemonade URL trust pinning, runtime overlay covers HIP libs, opt-out env
Robustness pass driven by 5 parallel reviewers of head c08b15e6:
1. URL trust pinning. AssetChoice.url comes from the GitHub API response's
browser_download_url field. Lemonade attempts are not in the approved-hash
manifest, so a compromised API response could redirect the download to an
attacker-chosen host without the integrity gate catching it. New
_is_trusted_github_release_url() validates https + github.com/<expected_repo>
release path OR objects.githubusercontent.com (GitHub's CDN). Resolver
refuses to download otherwise.
2. UNSLOTH_DISABLE_LEMONADE_ROCM opt-out. Users who prefer the upstream HIP
build path can set this env var to skip lemonade outright (e.g. for
air-gapped installs or stricter trust requirements). The install log
already prints a NOTE explaining that lemonade lacks approved-hash
coverage so users know the trust model.
3. linux-rocm runtime overlay patterns extended to cover lemonade's bundled
HIP/ROCm runtime libs (libamdhip64.so*, libhsa-runtime64.so*, libhipblas*,
librocblas*, librocsolver*, librocsparse*, librocrand*, libMIOpen*,
libmagma*). The upstream tarball does not ship these (links against
system /opt/rocm), so the new glob entries are no-op for upstream and
load-bearing for lemonade. Without this, install_from_archives would
drop the bundled runtime libs from the lemonade ZIP and llama-server's
RPATH would fail to load amdhip64 at first inference.
4. _lemonade_release_api_for now URL-encodes llama_tag with quote(safe="").
Defence in depth: a tag containing /, ?, #, or whitespace cannot reshape
the request URL. Tags come from internal resolution today, but this
removes the risk if a future caller passes user-controlled input.
5. Empty browser_download_url skipped explicitly in the resolver. The
release_asset_map helper defaults missing URLs to "". Previously this
would have been passed to download_file("") which raises a less obvious
error than the new clean log + return None.
6. Docstring on _lemonade_release_api_for clarified: lemonade tags match
ggml-org/llama.cpp tags 1:1, NOT unslothai/llama.cpp fork tags. The
earlier P1 review report misread this and re-asked for the tag-drift
fix that the author intentionally reverted in f256cea950.
7. Autouse pytest fixture in test_lemonade_llamacpp_rocm_bins_mock.py
clears _fetch_lemonade_release_cached between tests. Today's tests all
mock the same payload so no pollution surfaces, but the lru_cache
becomes a footgun the moment any future test parametrises return values.
Tests: 28 passed (was 18, added 10 covering URL pinning, opt-out env,
pinned-tag helper, URL encoding, empty URL, runtime patterns).
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix(studio/rocm): complete lemonade runtime overlay + honor CUDA_VISIBLE_DEVICES
runtime_patterns_for_choice("linux-rocm") was missing libamd_comgr.so*,
librocm_kpack.so*, and librocm_sysdeps_*.so*, which are direct NEEDED
entries of libamdhip64.so.7 in every lemonade bundle. install_from_archives
did not copy them, so the preflight ldd walk failed on llama-server and every
lemonade attempt fell back to source build. Confirmed on gfx1151 b1272 by
h34v3nzc0dex; manually staging the full bundle with the three missing
patterns restored the install and preserved bench parity.
Also extend _pick_rocm_gfx_target to check CUDA_VISIBLE_DEVICES after
HIP_VISIBLE_DEVICES and ROCR_VISIBLE_DEVICES: AMD's HIP runtime honours
all three with identical semantics, so mixed-arch hosts where users set
only CUDA_VISIBLE_DEVICES were still picking the first gfx token instead
of the user-selected device.
Tests: 30 passed (was 28); added 2 for CUDA_VISIBLE_DEVICES (multi-GPU
pick + -1 opt-out) and extended the runtime-patterns test to assert the
three newly added lib globs are present.
* fix: tighten CDN trust check and fix multi-GPU arch selection
- _is_trusted_github_release_url: require /github-production-release-asset-
path prefix so only real GitHub release CDN URLs are accepted
- _pick_rocm_gfx_target: parse rocminfo Agent N section boundaries to build
a per-physical-GPU arch list; same arch across multiple GPUs no longer
collapses to a single token, so HIP_VISIBLE_DEVICES indexing works correctly
- tests: update CDN test to use realistic path prefix, add rejection test for
arbitrary CDN path, add regression test for same-arch multi-GPU case
* fix: use broad lib*.so* glob for linux-rocm runtime overlay
Replace the explicit lib allowlist in runtime_patterns_for_choice with
lib*.so* for the linux-rocm path.
The lemonade ROCm ZIPs carry a full HIP/ROCm runtime including transitive
deps like libLLVM.so.23.0git and libclang-cpp.so.23.0git (pulled in by
libamd_comgr.so.3). These names change across ROCm releases and were not
in the allowlist, so preflight would see them as unresolved NEEDED entries
and fall back to a source build. The broad glob catches everything in the
bundle now and in future releases without needing to enumerate each library.
* fix: show lemonade binary tag in install summary log line
Store binary_repo and binary_release_tag in UNSLOTH_PREBUILT_INFO.json
so the setup.sh summary can distinguish the source tree (unslothai/llama.cpp)
from the actual binary origin (lemonade-sdk/llamacpp-rocm).
Before: 'installed release: unslothai/llama.cpp@b9334'
After: 'installed release: unslothai/llama.cpp@b9334 + lemonade@b1280'
Upstream installs are unchanged (binary_repo == published_repo).
* fix(merge-compat): align Windows ROCm guard and helper name with strix branch
- elif host.has_rocm instead of if...not to match fix/rocm-strix-halo-unified-memory
- _resolve_exe instead of _resolve_amd_exe with identical body/docstring
Eliminates the two conflict hunks that would otherwise block a clean bot merge
of feature/lemonade-rocm-prebuilts on top of fix/rocm-strix-halo-unified-memory.
* fix: three PR review corrections for lemonade ROCm prebuilt integration
- _lemonade_release_api_for: fix docstring claiming lemonade tags match
ggml-org 1:1 -- lemonade may be several builds behind ggml-org (noted
by oobabooga, confirmed: lemonade b1281 vs ggml-org b9370)
- direct_linux_release_plan: move cpu_choice into else-branch so ROCm-only
hosts never get a CPU fallback in the attempts list -- a failed lemonade
binary now raises PrebuiltFallback and triggers the HIP source build
instead of silently installing a CPU-only binary
- apply_approved_hashes: pass lemonade attempts through without requiring
a manifest entry; lemonade is explicitly documented as relying on
functional validation only, so rejecting it here caused PrebuiltFallback
on the non-simple-policy path before the upstream ROCm/HIP fallback
could be considered
* fix: copy hipblaslt/rocblas library subdirs from lemonade ROCm archives
copy_globs matches filenames only and copies flat, so it cannot
preserve the hipblaslt/library/<gfx>/ and rocblas/library/<gfx>/
Tensile kernel catalog trees that lemonade ROCm zips ship alongside
the .so files. Add runtime_subdirs_for_choice() and call shutil.copytree
for each named subdir after the copy_globs pass in install_from_archives.
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Co-authored-by: Jaeic Lee <jaeiclee@users.noreply.github.com>
1402 lines
60 KiB
Bash
Executable file
1402 lines
60 KiB
Bash
Executable file
#!/usr/bin/env bash
|
|
# SPDX-License-Identifier: AGPL-3.0-only
|
|
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
|
|
|
set -euo pipefail
|
|
|
|
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
REPO_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
|
|
RULE=$(printf '\342\224\200%.0s' {1..52})
|
|
|
|
# ── Parse flags ──
|
|
# --local: install from the local repo checkout (overlays unsloth as editable
|
|
# and unsloth-zoo from git main). Mirrors install.sh --local for the Colab
|
|
# path that runs setup.sh directly without going through install.sh.
|
|
if [ "$#" -gt 0 ]; then
|
|
for _arg in "$@"; do
|
|
case "$_arg" in
|
|
--local)
|
|
export STUDIO_LOCAL_INSTALL=1
|
|
export STUDIO_LOCAL_REPO="$REPO_ROOT"
|
|
;;
|
|
esac
|
|
done
|
|
fi
|
|
|
|
# ── Maintainer-editable defaults ──────────────────────────────────────────
|
|
# Change these in the GitHub-hosted script so all users get updated defaults.
|
|
# User environment variables always override these baked-in values.
|
|
#
|
|
# _DEFAULT_LLAMA_PR_FORCE : PR number to build by default ("" = normal path)
|
|
# _DEFAULT_LLAMA_SOURCE : git clone URL for source builds
|
|
# _DEFAULT_LLAMA_TAG : llama.cpp ref to build ("latest" = newest release,
|
|
# "master" = bleeding-edge, "bNNNN" = specific tag)
|
|
# Prefer "latest" over "master" -- "master" bypasses
|
|
# the prebuilt resolver (no matching GitHub release),
|
|
# forces a source build, and causes HTTP 422 errors.
|
|
# Only use "master" temporarily when the latest release
|
|
# is missing support for a new model architecture.
|
|
# ──────────────────────────────────────────────────────────────────────────
|
|
_DEFAULT_LLAMA_PR_FORCE=""
|
|
_DEFAULT_LLAMA_SOURCE="https://github.com/ggml-org/llama.cpp"
|
|
_DEFAULT_LLAMA_TAG="latest"
|
|
_DEFAULT_LLAMA_FORCE_COMPILE_REF="master"
|
|
|
|
# ── Colors (same palette as startup_banner / install_python_stack) ──
|
|
if [ -n "${NO_COLOR:-}" ]; then
|
|
C_TITLE= C_DIM= C_OK= C_WARN= C_ERR= C_RST=
|
|
elif [ -t 1 ] || [ -n "${FORCE_COLOR:-}" ]; then
|
|
C_TITLE=$'\033[38;5;150m'
|
|
C_DIM=$'\033[38;5;245m'
|
|
C_OK=$'\033[38;5;108m'
|
|
C_WARN=$'\033[38;5;136m'
|
|
C_ERR=$'\033[91m'
|
|
C_RST=$'\033[0m'
|
|
else
|
|
C_TITLE= C_DIM= C_OK= C_WARN= C_ERR= C_RST=
|
|
fi
|
|
|
|
# ── Output helpers ──
|
|
# Consistent column layout: 2-space indent, 15-char label (fits llama-quantize), then value.
|
|
# Usage: step <label> <message> [color] (color defaults to C_OK)
|
|
step() { printf " ${C_DIM}%-15.15s${C_RST}${3:-$C_OK}%s${C_RST}\n" "$1" "$2"; }
|
|
substep() { printf " ${C_DIM}%-15s%s${C_RST}\n" "" "$1"; }
|
|
|
|
_is_verbose() {
|
|
[ "${UNSLOTH_VERBOSE:-0}" = "1" ]
|
|
}
|
|
|
|
verbose_substep() {
|
|
if _is_verbose; then
|
|
substep "$1"
|
|
fi
|
|
return 0
|
|
}
|
|
|
|
run_maybe_quiet() {
|
|
if _is_verbose; then
|
|
"$@"
|
|
else
|
|
"$@" > /dev/null 2>&1
|
|
fi
|
|
}
|
|
|
|
# ── Helper: run command quietly, show output only on failure ──
|
|
_run_quiet() {
|
|
local on_fail=$1
|
|
local label=$2
|
|
shift 2
|
|
|
|
if _is_verbose; then
|
|
local exit_code
|
|
"$@" && return 0
|
|
exit_code=$?
|
|
step "error" "$label failed (exit code $exit_code)" "$C_ERR" >&2
|
|
if [ "$on_fail" = "exit" ]; then
|
|
exit "$exit_code"
|
|
else
|
|
return "$exit_code"
|
|
fi
|
|
fi
|
|
|
|
local tmplog
|
|
tmplog=$(mktemp) || {
|
|
step "error" "Failed to create temporary file" "$C_ERR" >&2
|
|
[ "$on_fail" = "exit" ] && exit 1 || return 1
|
|
}
|
|
|
|
if "$@" >"$tmplog" 2>&1; then
|
|
rm -f "$tmplog"
|
|
return 0
|
|
else
|
|
local exit_code=$?
|
|
step "error" "$label failed (exit code $exit_code)" "$C_ERR" >&2
|
|
cat "$tmplog" >&2
|
|
rm -f "$tmplog"
|
|
|
|
if [ "$on_fail" = "exit" ]; then
|
|
exit "$exit_code"
|
|
else
|
|
return "$exit_code"
|
|
fi
|
|
fi
|
|
}
|
|
|
|
run_quiet() {
|
|
_run_quiet exit "$@"
|
|
}
|
|
|
|
run_quiet_no_exit() {
|
|
_run_quiet return "$@"
|
|
}
|
|
|
|
_nvcc_meets_llama_minimum() {
|
|
# Echo "ok|too_old|unknown" then the parsed "X.Y" version, one per line.
|
|
# llama.cpp needs CUDA toolkit >= 12.4 (#4437; setup.ps1 aborts via #4517).
|
|
_nvcc_bin=$1
|
|
[ -n "$_nvcc_bin" ] || { echo "unknown"; echo ""; return 0; }
|
|
_raw=$("$_nvcc_bin" --version 2>/dev/null \
|
|
| sed -n 's/.*release \([0-9][0-9]*\.[0-9][0-9]*\).*/\1/p' \
|
|
| head -1)
|
|
if [ -z "$_raw" ]; then
|
|
echo "unknown"; echo ""; return 0
|
|
fi
|
|
_maj=${_raw%%.*}
|
|
_min_raw=${_raw#*.}
|
|
_min=${_min_raw%%.*}
|
|
if [ "$_maj" -lt 12 ] 2>/dev/null; then
|
|
echo "too_old"
|
|
elif [ "$_maj" -eq 12 ] && [ "$_min" -lt 4 ] 2>/dev/null; then
|
|
echo "too_old"
|
|
else
|
|
echo "ok"
|
|
fi
|
|
echo "$_raw"
|
|
}
|
|
|
|
print_llama_error_log() {
|
|
local log_file=$1
|
|
[ -s "$log_file" ] || return 0
|
|
substep "llama.cpp diagnostics (last 120 lines):"
|
|
tail -n 120 "$log_file" | sed 's/^/ | /' >&2
|
|
}
|
|
|
|
installed_llama_prebuilt_release() {
|
|
local install_dir=${1:-}
|
|
local metadata_path="$install_dir/UNSLOTH_PREBUILT_INFO.json"
|
|
[ -f "$metadata_path" ] || return 0
|
|
python - "$metadata_path" <<'PY' 2>/dev/null || true
|
|
import json
|
|
import sys
|
|
from pathlib import Path
|
|
|
|
try:
|
|
payload = json.loads(Path(sys.argv[1]).read_text(encoding="utf-8"))
|
|
except Exception:
|
|
raise SystemExit(0)
|
|
|
|
if not isinstance(payload, dict):
|
|
raise SystemExit(0)
|
|
|
|
repo = str(payload.get("published_repo") or "").strip()
|
|
release_tag = str(payload.get("release_tag") or "").strip()
|
|
llama_tag = str(payload.get("tag") or "").strip()
|
|
source = str(payload.get("source") or "").strip()
|
|
binary_repo = str(payload.get("binary_repo") or "").strip()
|
|
binary_tag = str(payload.get("binary_release_tag") or "").strip()
|
|
if not repo or not release_tag:
|
|
raise SystemExit(0)
|
|
|
|
# For non-upstream sources (e.g. lemonade) the published_repo/release_tag
|
|
# refer to the unsloth source tree while the actual binaries came from a
|
|
# different repo. Show both so the log is unambiguous.
|
|
if source and source != "upstream" and binary_repo and binary_tag and binary_repo != repo:
|
|
message = f"installed release: {repo}@{release_tag} + {source}@{binary_tag}"
|
|
else:
|
|
message = f"installed release: {repo}@{release_tag}"
|
|
if llama_tag and llama_tag != release_tag:
|
|
message += f" (tag {llama_tag})"
|
|
print(message)
|
|
PY
|
|
}
|
|
|
|
print_installed_llama_prebuilt_release() {
|
|
local install_dir=${1:-}
|
|
local installed_release
|
|
installed_release="$(installed_llama_prebuilt_release "$install_dir")"
|
|
if [ -n "$installed_release" ]; then
|
|
substep "$installed_release"
|
|
fi
|
|
}
|
|
|
|
# ── Banner ──
|
|
echo ""
|
|
printf " ${C_TITLE}%s${C_RST}\n" "🦥 Unsloth Studio Setup"
|
|
printf " ${C_DIM}%s${C_RST}\n" "$RULE"
|
|
verbose_substep "verbose diagnostics enabled"
|
|
_LLAMA_ONLY="${UNSLOTH_STUDIO_LLAMA_ONLY:-0}"
|
|
if [ "$_LLAMA_ONLY" = "1" ]; then
|
|
substep "llama.cpp only mode"
|
|
fi
|
|
if [ "${STUDIO_LOCAL_INSTALL:-0}" = "1" ]; then
|
|
substep "local mode: overlaying $REPO_ROOT (editable) + unsloth-zoo from git main"
|
|
fi
|
|
# ── Clean up stale caches ──
|
|
rm -rf "$REPO_ROOT/unsloth_compiled_cache"
|
|
rm -rf "$SCRIPT_DIR/backend/unsloth_compiled_cache"
|
|
rm -rf "$SCRIPT_DIR/tmp/unsloth_compiled_cache"
|
|
|
|
# ── Detect Colab ──
|
|
IS_COLAB=false
|
|
keynames=$'\n'$(printenv | cut -d= -f1)
|
|
if [[ "$keynames" == *$'\nCOLAB_'* ]]; then
|
|
IS_COLAB=true
|
|
fi
|
|
|
|
if [ "$_LLAMA_ONLY" != "1" ]; then
|
|
# ── Detect whether frontend needs building ──
|
|
# Skip if SKIP_STUDIO_FRONTEND=1 (Tauri desktop app bundles its own frontend),
|
|
# or if dist/ exists AND no tracked input is newer than dist/.
|
|
if [ "${SKIP_STUDIO_FRONTEND:-0}" = "1" ]; then
|
|
_NEED_FRONTEND_BUILD=false
|
|
step "frontend" "bundled (Tauri)"
|
|
else
|
|
_NEED_FRONTEND_BUILD=true
|
|
if [ -d "$SCRIPT_DIR/frontend/dist" ]; then
|
|
_changed=$(find "$SCRIPT_DIR/frontend" -maxdepth 1 -type f \
|
|
! -name 'bun.lock' \
|
|
-newer "$SCRIPT_DIR/frontend/dist" -print -quit 2>/dev/null)
|
|
if [ -z "$_changed" ]; then
|
|
_changed=$(find "$SCRIPT_DIR/frontend/src" "$SCRIPT_DIR/frontend/public" \
|
|
-type f -newer "$SCRIPT_DIR/frontend/dist" -print -quit 2>/dev/null) || true
|
|
fi
|
|
[ -z "$_changed" ] && _NEED_FRONTEND_BUILD=false
|
|
fi
|
|
fi # end SKIP_STUDIO_FRONTEND guard
|
|
|
|
if [ "$_NEED_FRONTEND_BUILD" = false ]; then
|
|
step "frontend" "up to date"
|
|
verbose_substep "frontend dist is newer than source inputs"
|
|
else
|
|
|
|
# ── Node ──
|
|
NEED_NODE=true
|
|
if command -v node &>/dev/null && command -v npm &>/dev/null; then
|
|
NODE_MAJOR=$(node -v | sed 's/v//' | cut -d. -f1)
|
|
NODE_MINOR=$(node -v | sed 's/v//' | cut -d. -f2)
|
|
NPM_MAJOR=$(npm -v | cut -d. -f1)
|
|
# Vite 8 requires Node ^20.19.0 || >=22.12.0
|
|
NODE_OK=false
|
|
if [ "$NODE_MAJOR" -eq 20 ] && [ "$NODE_MINOR" -ge 19 ]; then NODE_OK=true; fi
|
|
if [ "$NODE_MAJOR" -eq 22 ] && [ "$NODE_MINOR" -ge 12 ]; then NODE_OK=true; fi
|
|
if [ "$NODE_MAJOR" -ge 23 ]; then NODE_OK=true; fi
|
|
if [ "$NODE_OK" = true ] && [ "$NPM_MAJOR" -ge 11 ]; then
|
|
NEED_NODE=false
|
|
else
|
|
if [ "$IS_COLAB" = true ] && [ "$NODE_OK" = true ]; then
|
|
# In Colab, just upgrade npm directly - nvm doesn't work well
|
|
if [ "$NPM_MAJOR" -lt 11 ]; then
|
|
substep "upgrading npm..."
|
|
run_maybe_quiet npm install -g npm@latest
|
|
fi
|
|
NEED_NODE=false
|
|
fi
|
|
fi
|
|
fi
|
|
|
|
if [ "$NEED_NODE" = true ]; then
|
|
substep "installing nvm..."
|
|
export NODE_OPTIONS=--dns-result-order=ipv4first
|
|
if _is_verbose; then
|
|
curl -so- https://raw.githubusercontent.com/nvm-sh/nvm/v0.40.1/install.sh | bash
|
|
else
|
|
curl -so- https://raw.githubusercontent.com/nvm-sh/nvm/v0.40.1/install.sh | bash > /dev/null 2>&1
|
|
fi
|
|
|
|
export NVM_DIR="$HOME/.nvm"
|
|
set +u
|
|
[ -s "$NVM_DIR/nvm.sh" ] && \. "$NVM_DIR/nvm.sh"
|
|
|
|
if [ -f "$HOME/.npmrc" ]; then
|
|
if grep -qE '^\s*(prefix|globalconfig)\s*=' "$HOME/.npmrc"; then
|
|
sed -i.bak '/^\s*\(prefix\|globalconfig\)\s*=/d' "$HOME/.npmrc"
|
|
fi
|
|
fi
|
|
|
|
substep "installing Node LTS..."
|
|
run_quiet "nvm install" nvm install --lts
|
|
if _is_verbose; then
|
|
nvm use --lts
|
|
else
|
|
nvm use --lts > /dev/null 2>&1
|
|
fi
|
|
set -u
|
|
|
|
NODE_MAJOR=$(node -v | sed 's/v//' | cut -d. -f1)
|
|
NPM_MAJOR=$(npm -v | cut -d. -f1)
|
|
|
|
if [ "$NODE_MAJOR" -lt 20 ]; then
|
|
step "node" "FAILED -- version must be >= 20 (got $(node -v))" "$C_ERR"
|
|
exit 1
|
|
fi
|
|
if [ "$NPM_MAJOR" -lt 11 ]; then
|
|
substep "upgrading npm..."
|
|
run_quiet "npm update" npm install -g npm@latest
|
|
fi
|
|
fi
|
|
|
|
step "node" "$(node -v) | npm $(npm -v)"
|
|
verbose_substep "node check: NEED_NODE=$NEED_NODE NODE_OK=${NODE_OK:-unknown} NPM_MAJOR=${NPM_MAJOR:-unknown}"
|
|
|
|
# ── Install bun (optional, faster package installs) ──
|
|
# Uses npm to install bun globally -- Node is already guaranteed above,
|
|
# avoids platform-specific installers, PATH issues, and admin requirements.
|
|
if ! command -v bun &>/dev/null; then
|
|
substep "installing bun..."
|
|
if run_maybe_quiet npm install -g bun && command -v bun &>/dev/null; then
|
|
substep "bun installed ($(bun --version))"
|
|
else
|
|
substep "bun install skipped (npm will be used instead)"
|
|
fi
|
|
else
|
|
substep "bun already installed ($(bun --version))"
|
|
fi
|
|
|
|
# ── Build frontend ──
|
|
substep "building frontend..."
|
|
cd "$SCRIPT_DIR/frontend"
|
|
_HIDDEN_GITIGNORES=()
|
|
_dir="$(pwd)"
|
|
while [ "$_dir" != "/" ]; do
|
|
_dir="$(dirname "$_dir")"
|
|
if [ -f "$_dir/.gitignore" ] && grep -qx '\*' "$_dir/.gitignore" 2>/dev/null; then
|
|
mv "$_dir/.gitignore" "$_dir/.gitignore._twbuild"
|
|
_HIDDEN_GITIGNORES+=("$_dir/.gitignore")
|
|
fi
|
|
done
|
|
|
|
_restore_gitignores() {
|
|
for _gi in "${_HIDDEN_GITIGNORES[@]+"${_HIDDEN_GITIGNORES[@]}"}"; do
|
|
mv "${_gi}._twbuild" "$_gi" 2>/dev/null || true
|
|
done
|
|
}
|
|
trap _restore_gitignores EXIT
|
|
|
|
# Use bun for install if available (faster), fall back to npm.
|
|
# Build always uses npm (Node runtime -- avoids bun runtime issues on some platforms).
|
|
# NOTE: We intentionally avoid run_quiet for the bun install attempt because
|
|
# run_quiet calls exit on failure, which would kill the script before the npm
|
|
# fallback can run. Instead we capture output manually and only show it on failure.
|
|
#
|
|
# IMPORTANT: bun's package cache can become corrupt -- packages get stored
|
|
# with only metadata (package.json, README) but no actual content (bin/,
|
|
# lib/). When this happens bun install exits 0 but leaves binaries missing.
|
|
# We verify critical binaries after install. If missing, we clear the cache
|
|
# and retry once before falling back to npm.
|
|
_try_bun_install() {
|
|
local _log _exit_code=0
|
|
_log=$(mktemp)
|
|
bun install >"$_log" 2>&1 || _exit_code=$?
|
|
|
|
# bun may create .exe shims on Windows (Git Bash / MSYS2) instead of plain scripts
|
|
if [ "$_exit_code" -eq 0 ] \
|
|
&& { [ -x node_modules/.bin/tsc ] || [ -f node_modules/.bin/tsc.exe ] || [ -f node_modules/.bin/tsc.bunx ]; } \
|
|
&& { [ -x node_modules/.bin/vite ] || [ -f node_modules/.bin/vite.exe ] || [ -f node_modules/.bin/vite.bunx ]; }; then
|
|
rm -f "$_log"
|
|
return 0
|
|
fi
|
|
|
|
# Either bun install failed or it exited 0 but left packages missing
|
|
if [ "$_exit_code" -ne 0 ]; then
|
|
echo " bun install failed (exit code $_exit_code):"
|
|
else
|
|
echo " bun install exited 0 but critical binaries are missing:"
|
|
fi
|
|
sed 's/^/ | /' "$_log" >&2
|
|
rm -f "$_log"
|
|
rm -rf node_modules
|
|
return 1
|
|
}
|
|
|
|
_bun_install_ok=false
|
|
if command -v bun &>/dev/null; then
|
|
substep "using bun for package install (faster)"
|
|
if _try_bun_install; then
|
|
_bun_install_ok=true
|
|
else
|
|
# First attempt failed, likely due to corrupt cache entries.
|
|
# Clear the cache and retry once.
|
|
echo " Clearing bun cache and retrying..."
|
|
run_maybe_quiet bun pm cache rm || true
|
|
if _try_bun_install; then
|
|
_bun_install_ok=true
|
|
fi
|
|
fi
|
|
fi
|
|
if [ "$_bun_install_ok" = false ]; then
|
|
run_quiet_no_exit "npm install" npm install --no-fund --no-audit --loglevel=error
|
|
_npm_install_rc=$?
|
|
if [ "$_npm_install_rc" -ne 0 ]; then
|
|
exit "$_npm_install_rc"
|
|
fi
|
|
fi
|
|
run_quiet "npm run build" npm run build
|
|
|
|
_restore_gitignores
|
|
trap - EXIT
|
|
|
|
_MAX_CSS=$(find "$SCRIPT_DIR/frontend/dist/assets" -name '*.css' -exec wc -c {} + 2>/dev/null | sort -n | tail -1 | awk '{print $1}')
|
|
if [ -z "$_MAX_CSS" ]; then
|
|
step "frontend" "built (warning: no CSS emitted)" "$C_WARN"
|
|
elif [ "$_MAX_CSS" -lt 100000 ]; then
|
|
step "frontend" "built (warning: CSS may be truncated)" "$C_WARN"
|
|
else
|
|
step "frontend" "built"
|
|
fi
|
|
|
|
cd "$SCRIPT_DIR"
|
|
|
|
fi # end frontend build check
|
|
|
|
# ── oxc-validator runtime ──
|
|
if [ -d "$SCRIPT_DIR/backend/core/data_recipe/oxc-validator" ] && command -v npm &>/dev/null; then
|
|
cd "$SCRIPT_DIR/backend/core/data_recipe/oxc-validator"
|
|
run_quiet_no_exit "npm install (oxc validator runtime)" npm install --no-fund --no-audit --loglevel=error
|
|
_oxc_install_rc=$?
|
|
if [ "$_oxc_install_rc" -ne 0 ]; then
|
|
exit "$_oxc_install_rc"
|
|
fi
|
|
cd "$SCRIPT_DIR"
|
|
fi
|
|
|
|
# ── Python venv + deps ──
|
|
# UNSLOTH_STUDIO_HOME (or STUDIO_HOME alias) overrides the install root
|
|
# (mirrors install.sh). UNSLOTH_STUDIO_HOME wins when both are set.
|
|
_studio_override_var=""
|
|
_studio_override="${UNSLOTH_STUDIO_HOME:-}"
|
|
if [ -n "$_studio_override" ]; then
|
|
_studio_override_var="UNSLOTH_STUDIO_HOME"
|
|
else
|
|
_studio_override="${STUDIO_HOME:-}"
|
|
[ -n "$_studio_override" ] && _studio_override_var="STUDIO_HOME"
|
|
fi
|
|
# Strip whitespace so " " is treated as unset (matches Python .strip()).
|
|
_studio_override=$(printf '%s' "$_studio_override" | sed -e 's/^[[:space:]]*//' -e 's/[[:space:]]*$//')
|
|
case "$_studio_override" in
|
|
"~") _studio_override="$HOME" ;;
|
|
"~/"*) _studio_override="$HOME/${_studio_override#'~/'}" ;;
|
|
esac
|
|
if [ -n "$_studio_override" ]; then
|
|
# setup.sh runs against an existing install (via 'unsloth studio update');
|
|
# a typo in the override must fail fast instead of materializing an
|
|
# empty workspace dir. Mirrors setup.ps1 behavior.
|
|
if [ ! -d "$_studio_override" ]; then
|
|
echo "ERROR: $_studio_override_var=$_studio_override does not exist." >&2
|
|
echo " Run install.sh to create the install root before 'unsloth studio update'." >&2
|
|
exit 1
|
|
fi
|
|
[ -w "$_studio_override" ] || { echo "ERROR: $_studio_override_var=$_studio_override is not writable." >&2; exit 1; }
|
|
STUDIO_HOME="$(CDPATH= cd -P -- "$_studio_override" && pwd -P)" || exit 1
|
|
else
|
|
STUDIO_HOME="$HOME/.unsloth/studio"
|
|
fi
|
|
VENV_DIR="$STUDIO_HOME/unsloth_studio"
|
|
VENV_T5_530_DIR="$STUDIO_HOME/.venv_t5_530"
|
|
VENV_T5_550_DIR="$STUDIO_HOME/.venv_t5_550"
|
|
|
|
[ -d "$REPO_ROOT/.venv" ] && rm -rf "$REPO_ROOT/.venv"
|
|
[ -d "$REPO_ROOT/.venv_overlay" ] && rm -rf "$REPO_ROOT/.venv_overlay"
|
|
[ -d "$REPO_ROOT/.venv_t5" ] && rm -rf "$REPO_ROOT/.venv_t5"
|
|
[ -d "$REPO_ROOT/.venv_t5_530" ] && rm -rf "$REPO_ROOT/.venv_t5_530"
|
|
[ -d "$REPO_ROOT/.venv_t5_550" ] && rm -rf "$REPO_ROOT/.venv_t5_550"
|
|
# Note: do NOT delete $STUDIO_HOME/.venv here — install.sh handles migration
|
|
|
|
_COLAB_NO_VENV=false
|
|
if [ ! -x "$VENV_DIR/bin/python" ]; then
|
|
if [ "$IS_COLAB" = true ]; then
|
|
# On Colab there is no Studio venv -- install backend deps into system Python.
|
|
# Strip all version constraints so pip keeps Colab's pre-installed
|
|
# packages (huggingface-hub, datasets, transformers) and only pulls
|
|
# in genuinely missing ones (structlog, fastapi, etc.).
|
|
substep "Colab detected, installing Studio backend dependencies..."
|
|
_COLAB_REQS_TMP="$(mktemp)"
|
|
sed 's/[><=!~;].*//' "$SCRIPT_DIR/backend/requirements/studio.txt" \
|
|
| grep -v '^#' | grep -v '^$' > "$_COLAB_REQS_TMP"
|
|
if [ -s "$_COLAB_REQS_TMP" ]; then
|
|
if ! run_quiet_no_exit "install Colab backend deps" pip install -q -r "$_COLAB_REQS_TMP"; then
|
|
rm -f "$_COLAB_REQS_TMP"
|
|
step "python" "Colab backend dependency install failed" "$C_ERR"
|
|
exit 1
|
|
fi
|
|
else
|
|
step "python" "no Colab backend dependencies resolved from requirements file" "$C_WARN"
|
|
fi
|
|
rm -f "$_COLAB_REQS_TMP"
|
|
_COLAB_NO_VENV=true
|
|
else
|
|
step "python" "venv not found at $VENV_DIR" "$C_ERR"
|
|
substep "Run install.sh first to create the environment:"
|
|
substep "curl -fsSL https://unsloth.ai/install.sh | sh"
|
|
exit 1
|
|
fi
|
|
else
|
|
source "$VENV_DIR/bin/activate"
|
|
fi
|
|
|
|
install_python_stack() {
|
|
python "$SCRIPT_DIR/install_python_stack.py"
|
|
}
|
|
|
|
USE_UV=false
|
|
if command -v uv &>/dev/null; then
|
|
USE_UV=true
|
|
elif {
|
|
if _is_verbose; then
|
|
curl -LsSf https://astral.sh/uv/install.sh | sh
|
|
else
|
|
curl -LsSf https://astral.sh/uv/install.sh | sh > /dev/null 2>&1
|
|
fi
|
|
}; then
|
|
export PATH="$HOME/.local/bin:$PATH"
|
|
command -v uv &>/dev/null && USE_UV=true
|
|
fi
|
|
|
|
fast_install() {
|
|
if [ "$USE_UV" = true ]; then
|
|
uv pip install --python "$(command -v python)" "$@" && return 0
|
|
fi
|
|
python -m pip install "$@"
|
|
}
|
|
|
|
cd "$SCRIPT_DIR"
|
|
|
|
# On Colab without a venv, skip venv-dependent Python deps sections but
|
|
# continue to llama.cpp install so GGUF inference is available.
|
|
if [ "$_COLAB_NO_VENV" = true ]; then
|
|
step "python" "backend deps installed into system Python"
|
|
substep "continuing to llama.cpp install for GGUF inference support"
|
|
fi
|
|
|
|
# ── Check if Python deps need updating ──
|
|
# Compare installed package version against PyPI latest.
|
|
# Skip all Python dependency work if versions match (fast update path).
|
|
# On Colab (no venv), skip this version check (it needs $VENV_DIR/bin/python)
|
|
# but still run install_python_stack below (it uses sys.executable).
|
|
_SKIP_PYTHON_DEPS=false
|
|
_SKIP_VERSION_CHECK=false
|
|
if [ "$_COLAB_NO_VENV" = true ]; then
|
|
_SKIP_VERSION_CHECK=true
|
|
fi
|
|
_PKG_NAME="${STUDIO_PACKAGE_NAME:-unsloth}"
|
|
if [ "$_SKIP_VERSION_CHECK" != true ] && [ "${SKIP_STUDIO_BASE:-0}" != "1" ] && [ "${STUDIO_LOCAL_INSTALL:-0}" != "1" ]; then
|
|
# Only check when NOT called from install.sh (which just installed the package)
|
|
INSTALLED_VER=$("$VENV_DIR/bin/python" -c "
|
|
import sys; from importlib.metadata import version
|
|
print(version(sys.argv[1]))
|
|
" "$_PKG_NAME" 2>/dev/null || echo "")
|
|
|
|
LATEST_VER=$(curl -fsSL --max-time 5 "https://pypi.org/pypi/$_PKG_NAME/json" 2>/dev/null \
|
|
| "$VENV_DIR/bin/python" -c "import sys,json; print(json.load(sys.stdin)['info']['version'])" 2>/dev/null \
|
|
|| echo "")
|
|
|
|
if [ -n "$INSTALLED_VER" ] && [ -n "$LATEST_VER" ] && [ "$INSTALLED_VER" = "$LATEST_VER" ]; then
|
|
step "python" "$_PKG_NAME $INSTALLED_VER is up to date"
|
|
_SKIP_PYTHON_DEPS=true
|
|
elif [ -n "$INSTALLED_VER" ] && [ -n "$LATEST_VER" ]; then
|
|
substep "$_PKG_NAME $INSTALLED_VER -> $LATEST_VER available, updating..."
|
|
elif [ -z "$LATEST_VER" ]; then
|
|
substep "could not reach PyPI, updating to be safe..."
|
|
fi
|
|
fi
|
|
|
|
if [ "$_SKIP_PYTHON_DEPS" = false ]; then
|
|
install_python_stack
|
|
else
|
|
step "python" "dependencies up to date"
|
|
verbose_substep "python deps check: installed=$_PKG_NAME@${INSTALLED_VER:-unknown} latest=${LATEST_VER:-unknown}"
|
|
fi
|
|
|
|
# ── 6b. Pre-install transformers 5.x into .venv_t5_530/ and .venv_t5_550/ ──
|
|
# Models like GLM-4.7-Flash, Qwen3 MoE need transformers>=5.3.0.
|
|
# Gemma 4 models need transformers>=5.5.0.
|
|
# Pre-install into separate directories to avoid runtime pip overhead.
|
|
# The training subprocess prepends the appropriate dir to sys.path.
|
|
#
|
|
# Runs outside the _SKIP_PYTHON_DEPS gate so that upgrades from legacy
|
|
# single .venv_t5 are always migrated to the tiered layout.
|
|
# why: in env-override mode $STUDIO_HOME is user-chosen; require the
|
|
# ownership marker before rm -rf so unrelated dirs survive. Gated on the
|
|
# canonical comparison so an override pointing at the legacy default still
|
|
# behaves like a default install.
|
|
_STUDIO_OWNED_MARKER=".unsloth-studio-owned"
|
|
_LEGACY_STUDIO_HOME="$HOME/.unsloth/studio"
|
|
_studio_home_canon="$STUDIO_HOME"
|
|
if [ -d "$_studio_home_canon" ]; then
|
|
_studio_home_canon=$(CDPATH= cd -P -- "$_studio_home_canon" 2>/dev/null && pwd -P) \
|
|
|| _studio_home_canon="$STUDIO_HOME"
|
|
fi
|
|
if [ -d "$_LEGACY_STUDIO_HOME" ]; then
|
|
_LEGACY_STUDIO_HOME=$(CDPATH= cd -P -- "$_LEGACY_STUDIO_HOME" 2>/dev/null && pwd -P) \
|
|
|| _LEGACY_STUDIO_HOME="$HOME/.unsloth/studio"
|
|
fi
|
|
_STUDIO_HOME_IS_CUSTOM=false
|
|
if [ "$_studio_home_canon" != "$_LEGACY_STUDIO_HOME" ]; then
|
|
_STUDIO_HOME_IS_CUSTOM=true
|
|
fi
|
|
_assert_studio_owned_or_absent() {
|
|
_aso_dir="$1"
|
|
_aso_label="$2"
|
|
[ -d "$_aso_dir" ] || return 0
|
|
if [ "$_STUDIO_HOME_IS_CUSTOM" = true ] && [ ! -f "$_aso_dir/$_STUDIO_OWNED_MARKER" ]; then
|
|
echo "ERROR: $_aso_dir already exists and is not marked as a Studio-owned $_aso_label." >&2
|
|
echo " Move it aside or choose an empty UNSLOTH_STUDIO_HOME before re-running." >&2
|
|
exit 1
|
|
fi
|
|
}
|
|
_NEED_T5_INSTALL=false
|
|
if [ -d "$STUDIO_HOME/.venv_t5" ]; then
|
|
# Legacy layout — migrate
|
|
_assert_studio_owned_or_absent "$STUDIO_HOME/.venv_t5" "legacy transformers sidecar venv"
|
|
rm -rf "$STUDIO_HOME/.venv_t5"
|
|
_NEED_T5_INSTALL=true
|
|
fi
|
|
[ ! -d "$VENV_T5_530_DIR" ] && _NEED_T5_INSTALL=true
|
|
[ ! -d "$VENV_T5_550_DIR" ] && _NEED_T5_INSTALL=true
|
|
# Also reinstall when python deps were updated (packages may need rebuild)
|
|
[ "$_SKIP_PYTHON_DEPS" = false ] && _NEED_T5_INSTALL=true
|
|
|
|
if [ "$_NEED_T5_INSTALL" = true ]; then
|
|
_assert_studio_owned_or_absent "$VENV_T5_530_DIR" "transformers 5.3 sidecar venv"
|
|
[ -d "$VENV_T5_530_DIR" ] && rm -rf "$VENV_T5_530_DIR"
|
|
mkdir -p "$VENV_T5_530_DIR"
|
|
: > "$VENV_T5_530_DIR/$_STUDIO_OWNED_MARKER" 2>/dev/null || true
|
|
run_quiet "install transformers 5.3.0" fast_install --target "$VENV_T5_530_DIR" --no-deps "transformers==5.3.0"
|
|
run_quiet "install huggingface_hub for t5_530" fast_install --target "$VENV_T5_530_DIR" --no-deps "huggingface_hub==1.8.0"
|
|
run_quiet "install hf_xet for t5_530" fast_install --target "$VENV_T5_530_DIR" --no-deps "hf_xet==1.4.2"
|
|
run_quiet "install tiktoken for t5_530" fast_install --target "$VENV_T5_530_DIR" "tiktoken"
|
|
step "transformers" "5.3.0 pre-installed"
|
|
|
|
_assert_studio_owned_or_absent "$VENV_T5_550_DIR" "transformers 5.5 sidecar venv"
|
|
[ -d "$VENV_T5_550_DIR" ] && rm -rf "$VENV_T5_550_DIR"
|
|
mkdir -p "$VENV_T5_550_DIR"
|
|
: > "$VENV_T5_550_DIR/$_STUDIO_OWNED_MARKER" 2>/dev/null || true
|
|
run_quiet "install transformers 5.5.0" fast_install --target "$VENV_T5_550_DIR" --no-deps "transformers==5.5.0"
|
|
run_quiet "install huggingface_hub for t5_550" fast_install --target "$VENV_T5_550_DIR" --no-deps "huggingface_hub==1.8.0"
|
|
run_quiet "install hf_xet for t5_550" fast_install --target "$VENV_T5_550_DIR" --no-deps "hf_xet==1.4.2"
|
|
run_quiet "install tiktoken for t5_550" fast_install --target "$VENV_T5_550_DIR" "tiktoken"
|
|
step "transformers" "5.5.0 pre-installed"
|
|
fi
|
|
fi
|
|
|
|
# ── GPU detection summary (mirrors setup.ps1 step "gpu" block) ──
|
|
_setup_amd_detected=false
|
|
_setup_gfx_all=""
|
|
_setup_mkt=""
|
|
if command -v rocminfo >/dev/null 2>&1 && \
|
|
rocminfo 2>/dev/null | awk '/Name:[[:space:]]*gfx[1-9][0-9]/{found=1} END{exit !found}'; then
|
|
_setup_amd_detected=true
|
|
_setup_gfx_all=$(rocminfo 2>/dev/null | grep -oE 'gfx[1-9][0-9a-z]{2,3}' || true)
|
|
_setup_mkt=$(rocminfo 2>/dev/null | awk -F': ' \
|
|
'/Marketing Name:/{gsub(/^[[:space:]]+|[[:space:]]+$/,"", $2); if($2){print $2; exit}}' || true)
|
|
elif command -v amd-smi >/dev/null 2>&1 && \
|
|
amd-smi list 2>/dev/null | awk '/^GPU[[:space:]]*[:\[][[:space:]]*[0-9]/{ found=1 } END{ exit !found }'; then
|
|
_setup_amd_detected=true
|
|
_setup_gfx_all=$(amd-smi list 2>/dev/null | grep -oE 'gfx[1-9][0-9a-z]{2,3}' || true)
|
|
[ -z "$_setup_gfx_all" ] && \
|
|
_setup_gfx_all=$(amd-smi static --asic 2>/dev/null | grep -oE 'gfx[1-9][0-9a-z]{2,3}' || true)
|
|
_setup_mkt=$(amd-smi static --asic 2>/dev/null | awk -F'[:|]' \
|
|
'/[Mm]arket.?[Nn]ame/{gsub(/^[[:space:]]+|[[:space:]]+$/,"", $2); if($2){print $2; exit}}' || true)
|
|
fi
|
|
|
|
if command -v nvidia-smi >/dev/null 2>&1 && \
|
|
nvidia-smi -L 2>/dev/null | awk '/^GPU[[:space:]]+[0-9]+:/{found=1} END{exit !found}'; then
|
|
step "gpu" "NVIDIA GPU detected"
|
|
elif [ "$_setup_amd_detected" = true ]; then
|
|
_setup_vis="${HIP_VISIBLE_DEVICES:-${ROCR_VISIBLE_DEVICES:-}}"
|
|
_setup_vis_idx=0
|
|
if [ -n "$_setup_vis" ] && [ "$_setup_vis" != "-1" ]; then
|
|
_setup_first="${_setup_vis%%,*}"
|
|
case "$_setup_first" in ''|*[!0-9]*) ;; *) _setup_vis_idx=$_setup_first ;; esac
|
|
fi
|
|
_setup_gfx=$(printf '%s\n' "$_setup_gfx_all" | awk -v idx="$_setup_vis_idx" \
|
|
'NF && !seen[$0]++ { a[n++]=$0 } END { if(idx>=n) idx=0; if(n>0) print a[idx] }')
|
|
# UNSLOTH_ROCM_GFX_ARCH env override (mirrors setup.ps1)
|
|
if [ -n "${UNSLOTH_ROCM_GFX_ARCH:-}" ]; then
|
|
_setup_gfx="${UNSLOTH_ROCM_GFX_ARCH}"
|
|
substep "gfx arch from UNSLOTH_ROCM_GFX_ARCH env override: $_setup_gfx"
|
|
# Name-based arch inference when tools don't report gfx (mirrors setup.ps1 nameArchTable)
|
|
elif [ -z "$_setup_gfx" ] && [ -n "$_setup_mkt" ]; then
|
|
case "$_setup_mkt" in
|
|
*"9070 XT"*|*9080*) _setup_gfx="gfx1201" ;; # RDNA 4
|
|
*9070*|*9060*) _setup_gfx="gfx1200" ;; # RDNA 4
|
|
*"8060S"*|*"890M"*|*"Strix Halo"*|*"HX 37"*|*"HX 38"*|*"AI 9 HX"*) _setup_gfx="gfx1151" ;; # RDNA 3.5 iGPU
|
|
*"880M"*|*"Strix Point"*|*"AI 9 36"*|*"AI 7 35"*|*"AI 5 34"*) _setup_gfx="gfx1150" ;; # RDNA 3.5 iGPU
|
|
*"RX 7900"*|*"RX 7800"*|*"RX 7700"*) _setup_gfx="gfx1100" ;; # RDNA 3 desktop
|
|
*"RX 7600"*) _setup_gfx="gfx1102" ;; # RDNA 3
|
|
*"780M"*|*"760M"*|*"740M"*|*"Phoenix"*) _setup_gfx="gfx1103" ;; # RDNA 3 iGPU
|
|
esac
|
|
if [ -n "$_setup_gfx" ]; then
|
|
substep "gfx arch inferred from GPU name: $_setup_gfx"
|
|
substep "Tip: set UNSLOTH_ROCM_GFX_ARCH=$_setup_gfx to skip inference next time"
|
|
fi
|
|
fi
|
|
# ROCm version via hipconfig, then amd-smi
|
|
_setup_rocm_ver=""
|
|
if command -v hipconfig >/dev/null 2>&1; then
|
|
_setup_rocm_ver=$(hipconfig --version 2>/dev/null | awk 'NR==1 && /^[0-9]/{print; exit}' || true)
|
|
fi
|
|
if [ -z "$_setup_rocm_ver" ] && command -v amd-smi >/dev/null 2>&1; then
|
|
_setup_rocm_ver=$(amd-smi version 2>/dev/null | awk -F'ROCm version: ' \
|
|
'NF>1{gsub(/[[:space:]]/,"", $2); print $2; exit}' || true)
|
|
fi
|
|
if [ -n "$_setup_gfx" ]; then
|
|
step "gpu" "AMD ROCm ($_setup_gfx)"
|
|
else
|
|
step "gpu" "AMD ROCm"
|
|
fi
|
|
_setup_rocm_root="${ROCM_PATH:-${HIP_PATH:-/opt/rocm}}"
|
|
substep "ROCm: $_setup_rocm_root"
|
|
[ -n "$_setup_rocm_ver" ] && substep "hipconfig: $_setup_rocm_ver"
|
|
[ -n "$_setup_mkt" ] && [ -n "$_setup_gfx" ] && substep "GPU: $_setup_mkt"
|
|
else
|
|
step "gpu" "none (chat-only / GGUF)" "$C_WARN"
|
|
substep "Training and GPU inference require an NVIDIA or AMD ROCm GPU."
|
|
fi
|
|
|
|
# ── 7. Prefer prebuilt llama.cpp bundles before any source build path ──
|
|
# Nest llama.cpp under $STUDIO_HOME only for real env-overrides; legacy
|
|
# default keeps ~/.unsloth/llama.cpp so pre-PR builds are still discovered.
|
|
if [ "$_STUDIO_HOME_IS_CUSTOM" = true ]; then
|
|
UNSLOTH_HOME="$STUDIO_HOME"
|
|
else
|
|
UNSLOTH_HOME="$HOME/.unsloth"
|
|
fi
|
|
mkdir -p "$UNSLOTH_HOME"
|
|
LLAMA_CPP_DIR="$UNSLOTH_HOME/llama.cpp"
|
|
LLAMA_SERVER_BIN="$LLAMA_CPP_DIR/build/bin/llama-server"
|
|
_NEED_LLAMA_SOURCE_BUILD=false
|
|
_LLAMA_CPP_DEGRADED=false
|
|
_LLAMA_FORCE_COMPILE="${UNSLOTH_LLAMA_FORCE_COMPILE:-0}"
|
|
_REQUESTED_LLAMA_TAG="${UNSLOTH_LLAMA_TAG:-${_DEFAULT_LLAMA_TAG}}"
|
|
_HOST_SYSTEM="$(uname -s 2>/dev/null || true)"
|
|
_HOST_MACHINE="$(uname -m 2>/dev/null || true)"
|
|
|
|
# Pick the release repo install_llama_prebuilt.py plans against.
|
|
# unslothai/llama.cpp ships only Linux CUDA bundles, so CPU-only Linux
|
|
# x86_64 routes to ggml-org for bin-ubuntu-x64.tar.gz. Anything with a
|
|
# GPU tool installed stays on unslothai (CUDA bundle / ROCm source build).
|
|
_LINUX_HAS_GPU=false
|
|
for _GPU_TOOL in nvidia-smi rocminfo amd-smi hipconfig hipinfo; do
|
|
if command -v "$_GPU_TOOL" >/dev/null 2>&1; then
|
|
_LINUX_HAS_GPU=true
|
|
break
|
|
fi
|
|
done
|
|
|
|
if [ "$_HOST_SYSTEM" = "Darwin" ]; then
|
|
_HELPER_RELEASE_REPO="ggml-org/llama.cpp"
|
|
elif [ "$_HOST_SYSTEM" = "Linux" ] \
|
|
&& [ "$_HOST_MACHINE" = "x86_64" ] \
|
|
&& [ "$_LINUX_HAS_GPU" = false ]; then
|
|
_HELPER_RELEASE_REPO="ggml-org/llama.cpp"
|
|
elif [ "$_HOST_SYSTEM" = "Linux" ] \
|
|
&& { [ "$_HOST_MACHINE" = "aarch64" ] || [ "$_HOST_MACHINE" = "arm64" ]; } \
|
|
&& [ "$_LINUX_HAS_GPU" = false ]; then
|
|
# Linux ARM64 (Ampere Altra, Raspberry Pi 5, GitHub `ubuntu-24.04-arm`,
|
|
# CPU-only Jetson rescue mode, ...). unslothai/llama.cpp only ships
|
|
# the Linux CUDA bundles, so without this branch the prebuilt
|
|
# resolver returns 0 attempts on every release and the installer
|
|
# falls all the way back to a source build. Upstream ggml-org ships
|
|
# llama-bNNNN-bin-ubuntu-arm64.tar.gz from at least b9072 onward.
|
|
_HELPER_RELEASE_REPO="ggml-org/llama.cpp"
|
|
else
|
|
_HELPER_RELEASE_REPO="unslothai/llama.cpp"
|
|
fi
|
|
unset _GPU_TOOL
|
|
_LLAMA_PR="${UNSLOTH_LLAMA_PR:-}"
|
|
_SKIP_PREBUILT_INSTALL=false
|
|
_LLAMA_PR_FORCE="${UNSLOTH_LLAMA_PR_FORCE:-${_DEFAULT_LLAMA_PR_FORCE}}"
|
|
_LLAMA_SOURCE="${_DEFAULT_LLAMA_SOURCE}"
|
|
_LLAMA_SOURCE="${_LLAMA_SOURCE%.git}" # normalize: strip trailing .git
|
|
_RESOLVED_SOURCE_URL="$_LLAMA_SOURCE"
|
|
_RESOLVED_SOURCE_REF="$_REQUESTED_LLAMA_TAG"
|
|
_RESOLVED_SOURCE_REF_KIND="tag"
|
|
_RESOLVED_LLAMA_TAG="$_REQUESTED_LLAMA_TAG"
|
|
|
|
if [ "$_LLAMA_FORCE_COMPILE" = "1" ]; then
|
|
_NEED_LLAMA_SOURCE_BUILD=true
|
|
_SKIP_PREBUILT_INSTALL=true
|
|
fi
|
|
|
|
# Baked-in PR_FORCE promotes to _LLAMA_PR when user hasn't set one.
|
|
if [ -z "$_LLAMA_PR" ] && [ -n "$_LLAMA_PR_FORCE" ] && \
|
|
[[ "$_LLAMA_PR_FORCE" =~ ^[0-9]+$ ]] && [ "$_LLAMA_PR_FORCE" -gt 0 ]; then
|
|
_LLAMA_PR="$_LLAMA_PR_FORCE"
|
|
step "llama.cpp" "baked-in PR_FORCE=$_LLAMA_PR_FORCE" "$C_WARN"
|
|
fi
|
|
|
|
if [ -n "$_LLAMA_PR" ]; then
|
|
if ! [[ "$_LLAMA_PR" =~ ^[0-9]+$ ]] || [ "$_LLAMA_PR" -le 0 ]; then
|
|
step "llama.cpp" "UNSLOTH_LLAMA_PR=$_LLAMA_PR is not a valid PR number" "$C_ERR"
|
|
exit 1
|
|
fi
|
|
step "llama.cpp" "UNSLOTH_LLAMA_PR=$_LLAMA_PR -- will build from PR head" "$C_WARN"
|
|
_RESOLVED_LLAMA_TAG="pr-$_LLAMA_PR"
|
|
_RESOLVED_SOURCE_URL="$_LLAMA_SOURCE"
|
|
_RESOLVED_SOURCE_REF="pr-$_LLAMA_PR"
|
|
_RESOLVED_SOURCE_REF_KIND="pull"
|
|
_NEED_LLAMA_SOURCE_BUILD=true
|
|
_SKIP_PREBUILT_INSTALL=true
|
|
fi
|
|
|
|
verbose_substep "requested llama.cpp tag: $_REQUESTED_LLAMA_TAG (repo: $_HELPER_RELEASE_REPO)"
|
|
|
|
if [ "$_LLAMA_FORCE_COMPILE" = "1" ]; then
|
|
step "llama.cpp" "UNSLOTH_LLAMA_FORCE_COMPILE=1 -- skipping prebuilt" "$C_WARN"
|
|
_NEED_LLAMA_SOURCE_BUILD=true
|
|
elif [ "${_SKIP_PREBUILT_INSTALL:-false}" = true ]; then
|
|
substep "prebuilt install skipped -- falling back to source build"
|
|
else
|
|
substep "installing prebuilt llama.cpp..."
|
|
if [ -d "$LLAMA_CPP_DIR" ]; then
|
|
substep "existing install detected -- validating update"
|
|
fi
|
|
# why: install_llama_prebuilt.py uses os.replace(), which would displace
|
|
# an unrelated $UNSLOTH_STUDIO_HOME/llama.cpp before the source-build
|
|
# ownership check below ever runs.
|
|
if [ "$_STUDIO_HOME_IS_CUSTOM" = true ]; then
|
|
_assert_studio_owned_or_absent "$LLAMA_CPP_DIR" "llama.cpp install"
|
|
fi
|
|
_PREBUILT_CMD=(
|
|
python "$SCRIPT_DIR/install_llama_prebuilt.py"
|
|
--install-dir "$LLAMA_CPP_DIR"
|
|
--llama-tag "$_REQUESTED_LLAMA_TAG"
|
|
--published-repo "$_HELPER_RELEASE_REPO"
|
|
--simple-policy
|
|
)
|
|
if [ -n "${UNSLOTH_LLAMA_RELEASE_TAG:-}" ]; then
|
|
_PREBUILT_CMD+=(--published-release-tag "$UNSLOTH_LLAMA_RELEASE_TAG")
|
|
fi
|
|
_PREBUILT_LOG="$(mktemp)"
|
|
set +e
|
|
if _is_verbose; then
|
|
"${_PREBUILT_CMD[@]}" 2>&1 | tee "$_PREBUILT_LOG"
|
|
_PREBUILT_STATUS=${PIPESTATUS[0]}
|
|
else
|
|
"${_PREBUILT_CMD[@]}" >"$_PREBUILT_LOG" 2>&1
|
|
_PREBUILT_STATUS=$?
|
|
fi
|
|
set -e
|
|
|
|
if [ "$_PREBUILT_STATUS" -eq 0 ]; then
|
|
if grep -Fq "already matches" "$_PREBUILT_LOG"; then
|
|
step "llama.cpp" "prebuilt up to date and validated"
|
|
else
|
|
step "llama.cpp" "prebuilt installed and validated"
|
|
fi
|
|
if [ "$_STUDIO_HOME_IS_CUSTOM" = true ] && [ -d "$LLAMA_CPP_DIR" ]; then
|
|
: > "$LLAMA_CPP_DIR/$_STUDIO_OWNED_MARKER" 2>/dev/null || true
|
|
fi
|
|
print_installed_llama_prebuilt_release "$LLAMA_CPP_DIR"
|
|
verbose_substep "llama.cpp install dir: $LLAMA_CPP_DIR"
|
|
rm -f "$_PREBUILT_LOG"
|
|
elif [ "$_PREBUILT_STATUS" -eq 3 ]; then
|
|
step "llama.cpp" "install blocked by active llama.cpp process" "$C_WARN"
|
|
print_llama_error_log "$_PREBUILT_LOG"
|
|
rm -f "$_PREBUILT_LOG"
|
|
if [ -d "$LLAMA_CPP_DIR" ]; then
|
|
substep "existing install was restored"
|
|
fi
|
|
substep "close Studio or other llama.cpp users and retry"
|
|
exit 3
|
|
else
|
|
step "llama.cpp" "prebuilt install failed (continuing)" "$C_WARN"
|
|
print_llama_error_log "$_PREBUILT_LOG"
|
|
rm -f "$_PREBUILT_LOG"
|
|
if [ -d "$LLAMA_CPP_DIR" ]; then
|
|
substep "prebuilt update failed; existing install restored"
|
|
fi
|
|
substep "falling back to source build"
|
|
_NEED_LLAMA_SOURCE_BUILD=true
|
|
fi
|
|
fi
|
|
|
|
# Source-built llama.cpp installs do not have the prebuilt metadata used above
|
|
# for exact release matching. Reuse a complete local source build unless the
|
|
# caller explicitly requested a rebuild or a PR-specific llama.cpp checkout.
|
|
if [ "$_NEED_LLAMA_SOURCE_BUILD" = true ] && \
|
|
[ "$_LLAMA_FORCE_COMPILE" != "1" ] && \
|
|
[ -z "$_LLAMA_PR" ] && \
|
|
[ -x "$LLAMA_CPP_DIR/build/bin/llama-server" ] && \
|
|
[ -x "$LLAMA_CPP_DIR/build/bin/llama-quantize" ]; then
|
|
step "llama.cpp" "existing source build found; skipping rebuild"
|
|
ln -sf build/bin/llama-quantize "$LLAMA_CPP_DIR/llama-quantize"
|
|
if [ "$_STUDIO_HOME_IS_CUSTOM" = true ]; then
|
|
: > "$LLAMA_CPP_DIR/$_STUDIO_OWNED_MARKER" 2>/dev/null || true
|
|
fi
|
|
_NEED_LLAMA_SOURCE_BUILD=false
|
|
fi
|
|
|
|
# ── 8. WSL: pre-install GGUF build dependencies for fallback source builds ──
|
|
# On WSL, sudo requires a password and can't be entered during GGUF export
|
|
# (runs in a non-interactive subprocess). Install build deps here instead.
|
|
if [ "$_NEED_LLAMA_SOURCE_BUILD" = true ] && grep -qi microsoft /proc/version 2>/dev/null; then
|
|
_GGUF_DEPS="pciutils build-essential cmake curl git libcurl4-openssl-dev"
|
|
apt-get update -y >/dev/null 2>&1 || true
|
|
apt-get install -y $_GGUF_DEPS >/dev/null 2>&1 || true
|
|
|
|
_STILL_MISSING=""
|
|
for _pkg in $_GGUF_DEPS; do
|
|
case "$_pkg" in
|
|
build-essential) command -v gcc >/dev/null 2>&1 || _STILL_MISSING="$_STILL_MISSING $_pkg" ;;
|
|
pciutils) command -v lspci >/dev/null 2>&1 || _STILL_MISSING="$_STILL_MISSING $_pkg" ;;
|
|
libcurl4-openssl-dev) command -v curl-config >/dev/null 2>&1 || _STILL_MISSING="$_STILL_MISSING $_pkg" ;;
|
|
*) command -v "$_pkg" >/dev/null 2>&1 || _STILL_MISSING="$_STILL_MISSING $_pkg" ;;
|
|
esac
|
|
done
|
|
_STILL_MISSING=$(echo "$_STILL_MISSING" | sed 's/^ *//')
|
|
|
|
if [ -z "$_STILL_MISSING" ]; then
|
|
step "gguf deps" "installed"
|
|
elif command -v sudo >/dev/null 2>&1; then
|
|
step "gguf deps" "sudo required for: $_STILL_MISSING" "$C_WARN"
|
|
printf " %-15s" ""
|
|
printf "accept? [Y/n] "
|
|
if [ -r /dev/tty ]; then
|
|
read -r REPLY </dev/tty || REPLY="y"
|
|
else
|
|
REPLY="y"
|
|
fi
|
|
case "$REPLY" in
|
|
[nN]*)
|
|
substep "skipped -- run manually:"
|
|
substep "sudo apt-get install -y $_STILL_MISSING"
|
|
_SKIP_GGUF_BUILD=true
|
|
;;
|
|
*)
|
|
sudo apt-get update -y
|
|
sudo apt-get install -y $_STILL_MISSING
|
|
step "gguf deps" "installed"
|
|
;;
|
|
esac
|
|
else
|
|
step "gguf deps" "missing (no sudo) -- install manually:" "$C_WARN"
|
|
substep "apt-get install -y $_STILL_MISSING"
|
|
_SKIP_GGUF_BUILD=true
|
|
fi
|
|
fi
|
|
|
|
# ── 9. Build llama.cpp binaries for GGUF inference + export when prebuilt install fails ──
|
|
# Builds at ~/.unsloth/llama.cpp — a single shared location under the user's
|
|
# home directory. This is used by both the inference server and the GGUF
|
|
# export pipeline (unsloth-zoo).
|
|
# - llama-server: for GGUF model inference
|
|
# - llama-quantize: for GGUF export quantization (symlinked to root for check_llama_cpp())
|
|
if [ "$_NEED_LLAMA_SOURCE_BUILD" = false ]; then
|
|
:
|
|
elif [ "${_SKIP_GGUF_BUILD:-}" = true ]; then
|
|
step "llama.cpp" "skipped (missing build deps)" "$C_WARN"
|
|
[ -f "$LLAMA_SERVER_BIN" ] || _LLAMA_CPP_DEGRADED=true
|
|
else
|
|
{
|
|
if ! command -v cmake &>/dev/null; then
|
|
step "llama.cpp" "skipped (cmake not found)" "$C_WARN"
|
|
[ -f "$LLAMA_SERVER_BIN" ] || _LLAMA_CPP_DEGRADED=true
|
|
elif ! command -v git &>/dev/null; then
|
|
step "llama.cpp" "skipped (git not found)" "$C_WARN"
|
|
[ -f "$LLAMA_SERVER_BIN" ] || _LLAMA_CPP_DEGRADED=true
|
|
else
|
|
if [ -z "$_LLAMA_PR" ]; then
|
|
_RESOLVED_SOURCE_URL="$_LLAMA_SOURCE"
|
|
if [ "$_LLAMA_FORCE_COMPILE" = "1" ]; then
|
|
if [ "$_REQUESTED_LLAMA_TAG" = "latest" ]; then
|
|
_RESOLVED_SOURCE_REF="${UNSLOTH_LLAMA_FORCE_COMPILE_REF:-${_DEFAULT_LLAMA_FORCE_COMPILE_REF}}"
|
|
_RESOLVED_SOURCE_REF_KIND="branch"
|
|
else
|
|
_RESOLVED_SOURCE_REF="$_REQUESTED_LLAMA_TAG"
|
|
_RESOLVED_SOURCE_REF_KIND="tag"
|
|
fi
|
|
elif [ "$_REQUESTED_LLAMA_TAG" = "latest" ]; then
|
|
_RESOLVE_TAG_ARGS=(--resolve-llama-tag latest --published-repo "ggml-org/llama.cpp" --output-format json)
|
|
set +e
|
|
_RESOLVE_TAG_JSON="$(python "$SCRIPT_DIR/install_llama_prebuilt.py" "${_RESOLVE_TAG_ARGS[@]}" 2>/dev/null)"
|
|
_RESOLVE_TAG_STATUS=$?
|
|
set -e
|
|
if [ "$_RESOLVE_TAG_STATUS" -eq 0 ] && [ -n "${_RESOLVE_TAG_JSON:-}" ]; then
|
|
_RESOLVED_SOURCE_REF="$(
|
|
printf '%s' "$_RESOLVE_TAG_JSON" | python -c 'import json,sys; print(json.load(sys.stdin).get("llama_tag",""))' 2>/dev/null || true
|
|
)"
|
|
else
|
|
_RESOLVED_SOURCE_REF=""
|
|
fi
|
|
if [ -z "$_RESOLVED_SOURCE_REF" ]; then
|
|
_RESOLVED_SOURCE_REF="latest"
|
|
fi
|
|
_RESOLVED_SOURCE_REF_KIND="tag"
|
|
else
|
|
_RESOLVED_SOURCE_REF="$_REQUESTED_LLAMA_TAG"
|
|
_RESOLVED_SOURCE_REF_KIND="tag"
|
|
fi
|
|
if [ -z "$_RESOLVED_SOURCE_URL" ]; then
|
|
_RESOLVED_SOURCE_URL="$_LLAMA_SOURCE"
|
|
fi
|
|
if [ -z "$_RESOLVED_SOURCE_REF" ]; then
|
|
_RESOLVED_SOURCE_REF="$_REQUESTED_LLAMA_TAG"
|
|
fi
|
|
fi
|
|
verbose_substep "source build repo: $_RESOLVED_SOURCE_URL"
|
|
verbose_substep "source build ref: ${_RESOLVED_SOURCE_REF:-latest} (${_RESOLVED_SOURCE_REF_KIND})"
|
|
BUILD_OK=true
|
|
mkdir -p "$(dirname "$LLAMA_CPP_DIR")"
|
|
_BUILD_TMP="${LLAMA_CPP_DIR}.build.$$"
|
|
rm -rf "$_BUILD_TMP"
|
|
if [ -n "$_LLAMA_PR" ]; then
|
|
run_quiet_no_exit "clone llama.cpp" \
|
|
git clone --depth 1 "${_LLAMA_SOURCE}.git" "$_BUILD_TMP" || BUILD_OK=false
|
|
if [ "$BUILD_OK" = true ]; then
|
|
run_quiet_no_exit "fetch PR #$_LLAMA_PR" \
|
|
git -C "$_BUILD_TMP" fetch --depth 1 origin "pull/$_LLAMA_PR/head:pr-$_LLAMA_PR" || BUILD_OK=false
|
|
fi
|
|
if [ "$BUILD_OK" = true ]; then
|
|
run_quiet_no_exit "checkout PR #$_LLAMA_PR" \
|
|
git -C "$_BUILD_TMP" checkout "pr-$_LLAMA_PR" || BUILD_OK=false
|
|
fi
|
|
elif [ "$_RESOLVED_SOURCE_REF_KIND" = "pull" ] && [ -n "$_RESOLVED_SOURCE_REF" ]; then
|
|
run_quiet_no_exit "clone llama.cpp" \
|
|
git clone --depth 1 "${_RESOLVED_SOURCE_URL}.git" "$_BUILD_TMP" || BUILD_OK=false
|
|
if [ "$BUILD_OK" = true ]; then
|
|
run_quiet_no_exit "fetch source PR ref" \
|
|
git -C "$_BUILD_TMP" fetch --depth 1 origin "$_RESOLVED_SOURCE_REF" || BUILD_OK=false
|
|
fi
|
|
if [ "$BUILD_OK" = true ]; then
|
|
run_quiet_no_exit "checkout source PR ref" \
|
|
git -C "$_BUILD_TMP" checkout -B unsloth-llama-build FETCH_HEAD || BUILD_OK=false
|
|
fi
|
|
elif [ "$_RESOLVED_SOURCE_REF_KIND" = "commit" ] && [ -n "$_RESOLVED_SOURCE_REF" ]; then
|
|
run_quiet_no_exit "clone llama.cpp" \
|
|
git clone --depth 1 "${_RESOLVED_SOURCE_URL}.git" "$_BUILD_TMP" || BUILD_OK=false
|
|
if [ "$BUILD_OK" = true ]; then
|
|
run_quiet_no_exit "fetch source commit" \
|
|
git -C "$_BUILD_TMP" fetch --depth 1 origin "$_RESOLVED_SOURCE_REF" || BUILD_OK=false
|
|
fi
|
|
if [ "$BUILD_OK" = true ]; then
|
|
run_quiet_no_exit "checkout source commit" \
|
|
git -C "$_BUILD_TMP" checkout -B unsloth-llama-build FETCH_HEAD || BUILD_OK=false
|
|
fi
|
|
else
|
|
_CLONE_ARGS=(git clone --depth 1)
|
|
if [ "$_RESOLVED_SOURCE_REF" != "latest" ] && [ -n "$_RESOLVED_SOURCE_REF" ]; then
|
|
_CLONE_ARGS+=(--branch "$_RESOLVED_SOURCE_REF")
|
|
fi
|
|
_CLONE_ARGS+=("${_RESOLVED_SOURCE_URL}.git" "$_BUILD_TMP")
|
|
run_quiet_no_exit "clone llama.cpp" \
|
|
"${_CLONE_ARGS[@]}" || BUILD_OK=false
|
|
fi
|
|
|
|
if [ "$BUILD_OK" = true ]; then
|
|
# Set Release explicitly (llama.cpp only defaults to it on non-MSVC/Xcode).
|
|
CMAKE_ARGS="-DCMAKE_BUILD_TYPE=Release -DLLAMA_BUILD_TESTS=OFF -DLLAMA_BUILD_EXAMPLES=OFF -DLLAMA_BUILD_SERVER=ON -DGGML_NATIVE=ON"
|
|
_TRY_METAL_CPU_FALLBACK=false
|
|
_HOST_SYSTEM="$(uname -s 2>/dev/null || true)"
|
|
_HOST_MACHINE="$(uname -m 2>/dev/null || true)"
|
|
_IS_MACOS_ARM64=false
|
|
if [ "$_HOST_SYSTEM" = "Darwin" ] && { [ "$_HOST_MACHINE" = "arm64" ] || [ "$_HOST_MACHINE" = "aarch64" ]; }; then
|
|
_IS_MACOS_ARM64=true
|
|
fi
|
|
|
|
if command -v ccache &>/dev/null; then
|
|
CMAKE_ARGS="$CMAKE_ARGS -DCMAKE_C_COMPILER_LAUNCHER=ccache -DCMAKE_CXX_COMPILER_LAUNCHER=ccache -DCMAKE_CUDA_COMPILER_LAUNCHER=ccache"
|
|
fi
|
|
CPU_FALLBACK_CMAKE_ARGS="$CMAKE_ARGS"
|
|
|
|
GPU_BACKEND=""
|
|
NVCC_PATH=""
|
|
if command -v nvcc &>/dev/null; then
|
|
NVCC_PATH="$(command -v nvcc)"
|
|
GPU_BACKEND="cuda"
|
|
elif [ -x /usr/local/cuda/bin/nvcc ]; then
|
|
NVCC_PATH="/usr/local/cuda/bin/nvcc"
|
|
export PATH="/usr/local/cuda/bin:$PATH"
|
|
GPU_BACKEND="cuda"
|
|
elif ls /usr/local/cuda-*/bin/nvcc &>/dev/null 2>&1; then
|
|
# Pick the newest cuda-XX.X directory
|
|
NVCC_PATH="$(ls -d /usr/local/cuda-*/bin/nvcc 2>/dev/null | sort -V | tail -1)"
|
|
export PATH="$(dirname "$NVCC_PATH"):$PATH"
|
|
GPU_BACKEND="cuda"
|
|
fi
|
|
|
|
# Check for ROCm (AMD) only if CUDA was not already selected
|
|
ROCM_HIPCC=""
|
|
if [ -z "$GPU_BACKEND" ]; then
|
|
if command -v hipcc &>/dev/null; then
|
|
ROCM_HIPCC="$(command -v hipcc)"
|
|
GPU_BACKEND="rocm"
|
|
elif [ -x /opt/rocm/bin/hipcc ]; then
|
|
ROCM_HIPCC="/opt/rocm/bin/hipcc"
|
|
export PATH="/opt/rocm/bin:$PATH"
|
|
GPU_BACKEND="rocm"
|
|
elif ls /opt/rocm-*/bin/hipcc &>/dev/null 2>&1; then
|
|
ROCM_HIPCC="$(ls -d /opt/rocm-*/bin/hipcc 2>/dev/null | sort -V | tail -1)"
|
|
export PATH="$(dirname "$ROCM_HIPCC"):$PATH"
|
|
GPU_BACKEND="rocm"
|
|
fi
|
|
fi
|
|
|
|
_BUILD_DESC="building"
|
|
if [ "$_IS_MACOS_ARM64" = true ]; then
|
|
# Metal takes precedence on Apple Silicon (CUDA/ROCm not functional on macOS)
|
|
_BUILD_DESC="building (Metal)"
|
|
CMAKE_ARGS="$CMAKE_ARGS -DGGML_METAL=ON -DGGML_METAL_EMBED_LIBRARY=ON -DGGML_METAL_USE_BF16=ON -DCMAKE_INSTALL_RPATH=@loader_path -DCMAKE_BUILD_WITH_INSTALL_RPATH=ON"
|
|
CPU_FALLBACK_CMAKE_ARGS="$CPU_FALLBACK_CMAKE_ARGS -DGGML_METAL=OFF"
|
|
_TRY_METAL_CPU_FALLBACK=true
|
|
elif [ -n "$NVCC_PATH" ]; then
|
|
# Returns "ok|too_old|unknown\nX.Y" on stdout.
|
|
_NVCC_CHECK="$(_nvcc_meets_llama_minimum "$NVCC_PATH")"
|
|
_NVCC_STATUS="$(printf '%s\n' "$_NVCC_CHECK" | sed -n '1p')"
|
|
_NVCC_VER="$(printf '%s\n' "$_NVCC_CHECK" | sed -n '2p')"
|
|
|
|
if [ "$_NVCC_STATUS" = "too_old" ]; then
|
|
substep "CUDA toolkit $_NVCC_VER is below llama.cpp minimum (12.4)." "$C_ERR"
|
|
substep "install a newer CUDA toolkit: https://developer.nvidia.com/cuda-toolkit-archive" "$C_WARN"
|
|
substep "falling back to CPU llama.cpp build for this run." "$C_WARN"
|
|
NVCC_PATH=""
|
|
GPU_BACKEND=""
|
|
_BUILD_DESC="building (CPU, CUDA toolkit < 12.4)"
|
|
else
|
|
CMAKE_ARGS="$CMAKE_ARGS -DGGML_CUDA=ON"
|
|
|
|
CUDA_ARCHS=""
|
|
if command -v nvidia-smi &>/dev/null; then
|
|
_raw_caps=$(nvidia-smi --query-gpu=compute_cap --format=csv,noheader 2>/dev/null || true)
|
|
while IFS= read -r _cap; do
|
|
_cap=$(echo "$_cap" | tr -d '[:space:]')
|
|
if [[ "$_cap" =~ ^([0-9]+)\.([0-9]+)$ ]]; then
|
|
_arch="${BASH_REMATCH[1]}${BASH_REMATCH[2]}"
|
|
# Append if not already present
|
|
case ";$CUDA_ARCHS;" in
|
|
*";$_arch;"*) ;;
|
|
*) CUDA_ARCHS="${CUDA_ARCHS:+$CUDA_ARCHS;}$_arch" ;;
|
|
esac
|
|
fi
|
|
done <<< "$_raw_caps"
|
|
fi
|
|
|
|
if [ -n "$CUDA_ARCHS" ]; then
|
|
CMAKE_ARGS="$CMAKE_ARGS -DCMAKE_CUDA_ARCHITECTURES=${CUDA_ARCHS}"
|
|
_BUILD_DESC="building (CUDA, sm_${CUDA_ARCHS//;/+sm_})"
|
|
else
|
|
_BUILD_DESC="building (CUDA)"
|
|
fi
|
|
|
|
CMAKE_ARGS="$CMAKE_ARGS -DCMAKE_CUDA_FLAGS=--threads=0"
|
|
|
|
# Accept a host gcc/clang newer than nvcc's whitelist; a fresh
|
|
# toolkit (e.g. CUDA 13.3) otherwise aborts with "#error --
|
|
# unsupported GNU version". Via env, not CMAKE_ARGS, to avoid
|
|
# word-splitting.
|
|
export NVCC_PREPEND_FLAGS="${NVCC_PREPEND_FLAGS:+$NVCC_PREPEND_FLAGS }-allow-unsupported-compiler"
|
|
fi
|
|
elif [ "$GPU_BACKEND" = "rocm" ]; then
|
|
# Resolve hipcc symlinks to find the real ROCm root
|
|
_HIPCC_REAL="$(readlink -f "$ROCM_HIPCC" 2>/dev/null || printf '%s' "$ROCM_HIPCC")"
|
|
ROCM_ROOT=""
|
|
if command -v hipconfig &>/dev/null; then
|
|
ROCM_ROOT="$(hipconfig -R 2>/dev/null || true)"
|
|
fi
|
|
if [ -z "$ROCM_ROOT" ]; then
|
|
ROCM_ROOT="$(cd "$(dirname "$_HIPCC_REAL")/.." 2>/dev/null && pwd)"
|
|
fi
|
|
|
|
_BUILD_DESC="building (ROCm)"
|
|
CMAKE_ARGS="$CMAKE_ARGS -DGGML_HIP=ON"
|
|
|
|
# ROCm 7.x ships clang-20 which on Ubuntu 24.04+ defaults to the
|
|
# highest-numbered gcc lib dir (/usr/lib/gcc/x86_64-linux-gnu/14/)
|
|
# which contains runtime objects but NOT C++ headers, causing:
|
|
# fatal error: 'cstdlib' file not found
|
|
# Find the newest gcc install dir that actually has both the
|
|
# runtime dir AND /usr/include/c++/<ver> headers, then pass it
|
|
# to clang via --gcc-install-dir so HIP builds succeed.
|
|
_GCC_INSTALL_DIR=""
|
|
_gcc_pm="$(gcc -print-multiarch 2>/dev/null)"
|
|
case "$_gcc_pm" in
|
|
*-linux-gnu*) _GCC_MULTIARCH="$_gcc_pm" ;;
|
|
*) _GCC_MULTIARCH="$(uname -m)-linux-gnu" ;;
|
|
esac
|
|
for _gcc_ver in 14 13 12 11; do
|
|
if [ -d "/usr/lib/gcc/$_GCC_MULTIARCH/$_gcc_ver/include" ] && \
|
|
[ -d "/usr/include/c++/$_gcc_ver" ]; then
|
|
_GCC_INSTALL_DIR="/usr/lib/gcc/$_GCC_MULTIARCH/$_gcc_ver"
|
|
break
|
|
fi
|
|
done
|
|
if [ -n "$_GCC_INSTALL_DIR" ]; then
|
|
CMAKE_ARGS="$CMAKE_ARGS -DCMAKE_HIP_FLAGS=--gcc-install-dir=\"$_GCC_INSTALL_DIR\""
|
|
substep "ROCm HIP gcc install dir: $_GCC_INSTALL_DIR"
|
|
fi
|
|
|
|
export ROCM_PATH="$ROCM_ROOT"
|
|
export HIP_PATH="$ROCM_ROOT"
|
|
|
|
# Use upstream-recommended HIP compiler (not legacy hipcc-as-CXX)
|
|
if command -v hipconfig &>/dev/null; then
|
|
_HIP_CLANG_DIR="$(hipconfig -l 2>/dev/null || true)"
|
|
[ -n "$_HIP_CLANG_DIR" ] && export HIPCXX="$_HIP_CLANG_DIR/clang"
|
|
fi
|
|
|
|
# Detect AMD GPU architecture (gfx target)
|
|
GPU_TARGETS=""
|
|
if command -v rocminfo &>/dev/null; then
|
|
_gfx_list=$(rocminfo 2>/dev/null | grep -oE 'gfx[0-9]{2,4}[a-z]?' | sort -u || true)
|
|
_valid_gfx=""
|
|
for _gfx in $_gfx_list; do
|
|
if [[ "$_gfx" =~ ^gfx[0-9]{2,4}[a-z]?$ ]]; then
|
|
# Drop bare family-level targets (gfx10, gfx11, gfx12, ...)
|
|
# when a specific sibling is present in the same list.
|
|
# rocminfo on ROCm 6.1+ emits both the specific GPU and
|
|
# the LLVM generic family line (e.g. gfx1100 alongside
|
|
# gfx11-generic), and the outer grep above captures the
|
|
# bare family prefix from the generic line. Passing that
|
|
# bare prefix to -DGPU_TARGETS breaks the HIP/llama.cpp
|
|
# build because clang only accepts specific gfxNNN ids.
|
|
# No real AMD GPU has a 2-digit gfx id, so this filter
|
|
# can only ever drop family prefixes, never real targets.
|
|
if [[ "$_gfx" =~ ^gfx[0-9]{2}$ ]] \
|
|
&& echo "$_gfx_list" | grep -qE "^${_gfx}[0-9][0-9a-z]?$"; then
|
|
continue
|
|
fi
|
|
_valid_gfx="${_valid_gfx}${_valid_gfx:+;}$_gfx"
|
|
fi
|
|
done
|
|
[ -n "$_valid_gfx" ] && GPU_TARGETS="$_valid_gfx"
|
|
fi
|
|
|
|
if [ -n "$GPU_TARGETS" ]; then
|
|
CMAKE_ARGS="$CMAKE_ARGS -DGPU_TARGETS=${GPU_TARGETS}"
|
|
_BUILD_DESC="building (ROCm, ${GPU_TARGETS//;/+})"
|
|
fi
|
|
elif [ -d /usr/local/cuda ] || nvidia-smi &>/dev/null; then
|
|
_BUILD_DESC="building (CPU, CUDA driver found but nvcc missing)"
|
|
elif [ -d /opt/rocm ] || command -v rocm-smi &>/dev/null; then
|
|
_BUILD_DESC="building (CPU, ROCm driver found but hipcc missing)"
|
|
else
|
|
_BUILD_DESC="building (CPU)"
|
|
fi
|
|
|
|
substep "$_BUILD_DESC..."
|
|
|
|
NCPU=$(nproc 2>/dev/null || sysctl -n hw.ncpu 2>/dev/null || echo 4)
|
|
CMAKE_GENERATOR_ARGS=""
|
|
if command -v ninja &>/dev/null; then
|
|
CMAKE_GENERATOR_ARGS="-G Ninja"
|
|
fi
|
|
|
|
# GPU label for the CPU-fallback message: Metal, else GPU_BACKEND
|
|
# (cuda/rocm). Empty on a bare CPU build (nothing to fall back from).
|
|
_gpu_fallback_label() {
|
|
if [ "$_TRY_METAL_CPU_FALLBACK" = true ]; then
|
|
echo "Metal"
|
|
elif [ -n "$GPU_BACKEND" ]; then
|
|
printf '%s' "$GPU_BACKEND" | tr '[:lower:]' '[:upper:]'
|
|
fi
|
|
}
|
|
|
|
if ! run_quiet_no_exit "cmake llama.cpp" cmake $CMAKE_GENERATOR_ARGS -S "$_BUILD_TMP" -B "$_BUILD_TMP/build" $CMAKE_ARGS; then
|
|
_FB_LABEL="$(_gpu_fallback_label)"
|
|
if [ -n "$_FB_LABEL" ]; then
|
|
_TRY_METAL_CPU_FALLBACK=false
|
|
substep "$_FB_LABEL configure failed; retrying CPU build..." "$C_WARN"
|
|
rm -rf "$_BUILD_TMP/build"
|
|
if run_quiet_no_exit "cmake llama.cpp (cpu fallback)" cmake $CMAKE_GENERATOR_ARGS -S "$_BUILD_TMP" -B "$_BUILD_TMP/build" $CPU_FALLBACK_CMAKE_ARGS; then
|
|
_BUILD_DESC="building (CPU fallback after $_FB_LABEL configure failed)"
|
|
# Now configured for CPU; clear GPU_BACKEND so a later
|
|
# build-step failure won't re-enter fallback on this config.
|
|
GPU_BACKEND=""
|
|
else
|
|
BUILD_OK=false
|
|
fi
|
|
else
|
|
BUILD_OK=false
|
|
fi
|
|
fi
|
|
fi
|
|
|
|
if [ "$BUILD_OK" = true ]; then
|
|
if ! run_quiet_no_exit "build llama-server" cmake --build "$_BUILD_TMP/build" --config Release --target llama-server -j"$NCPU"; then
|
|
_FB_LABEL="$(_gpu_fallback_label)"
|
|
if [ -n "$_FB_LABEL" ]; then
|
|
_TRY_METAL_CPU_FALLBACK=false
|
|
substep "$_FB_LABEL build failed; retrying CPU build..." "$C_WARN"
|
|
rm -rf "$_BUILD_TMP/build"
|
|
if run_quiet_no_exit "cmake llama.cpp (cpu fallback)" cmake $CMAKE_GENERATOR_ARGS -S "$_BUILD_TMP" -B "$_BUILD_TMP/build" $CPU_FALLBACK_CMAKE_ARGS; then
|
|
_BUILD_DESC="building (CPU fallback after $_FB_LABEL build failed)"
|
|
GPU_BACKEND=""
|
|
run_quiet_no_exit "build llama-server (cpu fallback)" cmake --build "$_BUILD_TMP/build" --config Release --target llama-server -j"$NCPU" || BUILD_OK=false
|
|
else
|
|
BUILD_OK=false
|
|
fi
|
|
else
|
|
BUILD_OK=false
|
|
fi
|
|
fi
|
|
fi
|
|
|
|
if [ "$BUILD_OK" = true ]; then
|
|
run_quiet_no_exit "build llama-quantize" cmake --build "$_BUILD_TMP/build" --config Release --target llama-quantize -j"$NCPU" || true
|
|
fi
|
|
|
|
# Swap only after build succeeds -- preserves existing install on failure
|
|
if [ "$BUILD_OK" = true ]; then
|
|
_assert_studio_owned_or_absent "$LLAMA_CPP_DIR" "llama.cpp install"
|
|
rm -rf "$LLAMA_CPP_DIR"
|
|
mv "$_BUILD_TMP" "$LLAMA_CPP_DIR"
|
|
: > "$LLAMA_CPP_DIR/$_STUDIO_OWNED_MARKER" 2>/dev/null || true
|
|
# Symlink to llama.cpp root -- check_llama_cpp() looks for the binary there
|
|
QUANTIZE_BIN="$LLAMA_CPP_DIR/build/bin/llama-quantize"
|
|
if [ -f "$QUANTIZE_BIN" ]; then
|
|
ln -sf build/bin/llama-quantize "$LLAMA_CPP_DIR/llama-quantize"
|
|
fi
|
|
else
|
|
rm -rf "$_BUILD_TMP"
|
|
fi
|
|
|
|
if [ "$BUILD_OK" = true ] && [ -f "$LLAMA_SERVER_BIN" ]; then
|
|
step "llama.cpp" "built"
|
|
[ -f "$LLAMA_CPP_DIR/llama-quantize" ] && step "llama-quantize" "built"
|
|
elif [ "$BUILD_OK" = true ]; then
|
|
step "llama.cpp" "binary not found after build" "$C_WARN"
|
|
_LLAMA_CPP_DEGRADED=true
|
|
else
|
|
step "llama.cpp" "build failed" "$C_ERR"
|
|
[ -f "$LLAMA_SERVER_BIN" ] || _LLAMA_CPP_DEGRADED=true
|
|
fi
|
|
fi
|
|
}
|
|
fi # end _SKIP_GGUF_BUILD check
|
|
|
|
# ── Footer ──
|
|
if [ "$_LLAMA_ONLY" = "1" ]; then
|
|
echo ""
|
|
printf " ${C_DIM}%s${C_RST}\n" "$RULE"
|
|
if [ "$_LLAMA_CPP_DEGRADED" = true ]; then
|
|
printf " ${C_WARN}%s${C_RST}\n" "llama.cpp update finished (limited: llama.cpp unavailable)"
|
|
else
|
|
printf " ${C_TITLE}%s${C_RST}\n" "llama.cpp update finished"
|
|
fi
|
|
printf " ${C_DIM}%s${C_RST}\n" "$RULE"
|
|
elif [ "$IS_COLAB" = true ]; then
|
|
echo ""
|
|
printf " ${C_DIM}%s${C_RST}\n" "$RULE"
|
|
if [ "$_LLAMA_CPP_DEGRADED" = true ]; then
|
|
printf " ${C_WARN}%s${C_RST}\n" "Unsloth Studio Setup Complete (limited: llama.cpp unavailable)"
|
|
else
|
|
printf " ${C_TITLE}%s${C_RST}\n" "Unsloth Studio Setup Complete"
|
|
fi
|
|
printf " ${C_DIM}%s${C_RST}\n" "$RULE"
|
|
substep "from colab import start"
|
|
substep "start()"
|
|
else
|
|
printf " ${C_DIM}%s${C_RST}\n" "$RULE"
|
|
if [ "$_LLAMA_CPP_DEGRADED" = true ]; then
|
|
printf " ${C_WARN}%s${C_RST}\n" "Unsloth Studio Installed (limited: llama.cpp unavailable)"
|
|
else
|
|
printf " ${C_TITLE}%s${C_RST}\n" "Unsloth Studio Installed"
|
|
fi
|
|
printf " ${C_DIM}%s${C_RST}\n" "$RULE"
|
|
if [ "$_LLAMA_CPP_DEGRADED" = true ]; then
|
|
printf " ${C_DIM}%-15s${C_WARN}%s${C_RST}\n" "launch" "unsloth studio -p 8888"
|
|
else
|
|
printf " ${C_DIM}%-15s${C_OK}%s${C_RST}\n" "launch" "unsloth studio -p 8888"
|
|
fi
|
|
printf " ${C_DIM}%-15s%s${C_RST}\n" "" "(add -H 0.0.0.0 to allow network / cloud access)"
|
|
fi
|
|
echo ""
|
|
|
|
# When called from install.sh (SKIP_STUDIO_BASE=1), exit non-zero so the
|
|
# installer can report the GGUF failure after finishing PATH/shortcut setup.
|
|
# When called directly via 'unsloth studio update', keep the install
|
|
# successful -- the footer above already reports the limitation and Studio
|
|
# is still usable for non-GGUF workflows.
|
|
if [ "$_LLAMA_CPP_DEGRADED" = true ] && [ "${SKIP_STUDIO_BASE:-0}" = "1" ]; then
|
|
exit 1
|
|
fi
|