* Studio: detect an interrupted dependency install instead of launching a backend that cannot import An installer killed part-way leaves a venv with a working CLI but without studio.txt's dependencies. Nothing recorded that, so three separate places all reported it healthy: - the desktop preflight probed only `unsloth -h` (typer + rich) and a hardcoded desktop-capabilities dict, neither of which touches studio.backend, so it returned ManagedReady and spawned a backend that died on `import structlog`; - setup.sh's fast path compared the installed unsloth version against PyPI, which matches on a half-built venv because unsloth is installed early, so `unsloth studio update` printed "up to date" and repaired nothing; - start_managed_repair calls that update and then re-checks with the same blind probes, so Repair reported success without fixing anything. install_python_stack.py now clears a completion manifest before the dependency pass and writes it only after the final step. `unsloth studio verify-install` and desktop-capabilities' new studio_install_ok field read it, the preflight turns a false answer into ManagedStale so auto-repair runs, and setup.sh / setup.ps1 gain an escape hatch next to the existing anyio one. Separately, the wheel ships studio/ and studio.backend* but declared none of their dependencies, so `unsloth train`, `export`, `chat`, `inference` and `studio` all ended in a rich traceback after a plain pip install. structlog is the only hard module-level import that chain reaches once starlette's annotation-only import moves under TYPE_CHECKING, so it becomes a core dependency and the rest of the server stack becomes a [studio] extra mirroring studio.txt. The CLI import sites now report missing dependencies as a sentence with two remedies. Fixes #4701, #5260, #7147 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Match the trimmed comments merged on the pip branch * Put the install manifest in the preflight fingerprint for PR #7492 The capability cache keyed the venv on pyvenv.cfg, uv.lock, requirements.txt, the interpreter and site-packages/unsloth_cli/commands/studio.py, none of which a repair touches when it only reinstalls studio.txt. So an entry cached while the install was healthy stayed valid after the manifest was dropped, and the probe returned Ready on exactly the half-built venv this is meant to catch. * Address the review findings on PR #7492 Fail the install when the completion manifest cannot be written, instead of exiting 0 without the record every later check requires, which is a repair loop by construction. Compare the version of the package the manifest names, so `studio update --package X` does not read as a permanent version change. Read the manifest from the venv that owns it when the CLI runs outside the managed venv, and drop the dependency verdict in that case: the walk ran against the wrong interpreter and says nothing about that venv. Name the import that actually failed. `unsloth train` reaches torch through the same guard, and the studio extra does not carry it, so recommending that extra alone left the command failing in the same place. * Declare click, which typer stopped providing, for PR #7492 unsloth_cli/commands/start.py imports click at module scope and unsloth_cli/__init__.py imports that module, so every unsloth command needs it. typer carried click through 0.19 and dropped it in 0.27, and the declared floor is typer>=0.12.0, so a fresh resolve gets no click. On the published wheel it still arrives because huggingface_hub requires click<9,>=8.4.2, which is luck rather than a declaration. A wheel built from this branch's dependency list has neither, and every command dies at import. Verified: before, `unsloth --help` on a fresh venv raised ModuleNotFoundError for click; after, it exits 0. The drift test now covers it. * Keep a running backend from the previous app version manageable The manageability bump gated two unrelated things through one constant. For the managed CLI probe 2 is right: a CLI reporting 1 cannot answer studio_install_ok. For a RUNNING backend it is wrong, because a process already started cannot change what it reports, so bumping studio/backend/main.py in lockstep does not help one the previous app version spawned. That backend is proven ours by root id and ownership token, but lifecycle_control_block_reason returned Unmanageable, and that branch never calls adopt_verified_backend. has_owned_backend() stays false, so Repair falls into block_external_conflict, which finds the same process and refuses: the app could no longer stop a backend it owns the token for. The same regression in backend.rs turned a terminal-launched same-root server from AttachedReady into ExternalConflict. Split the constant: DESKTOP_BACKEND_MANAGEABILITY_VERSION = 1 for the two live-backend probes, DESKTOP_MANAGEABILITY_VERSION = 2 for the CLI probe. Every real gate (protocol, auth, ownership, desktop-login, MIN_DESKTOP_BACKEND_VERSION) is untouched, so an old backend still reaches OwnedStale, adopt, stop, repair. Also stop the installer when the stale manifest cannot be removed. Windows raises on a read-only or locked file, and the pass would then run behind a marker that still names this version and these digests, so a run killed part-way would verify as complete. * Answer for the managed venv, not the one the CLI happens to run in The guard matched ModuleNotFoundError.name, an import name, against missing_requirements(), which returns distribution names. So a missing PyJWT printed 'pip install jwt', and jwt, docx and fitz are each a real but unrelated PyPI project (fitz is a neuroimaging workflow tool), so following the advice installed the wrong package and left the backend just as broken. Map the import to its distribution before deciding, and never offer the import itself. install_state() verified the caller's own prefix. The wheel ships studio/, so a CLI installed outside the managed venv always finds its own copy of the helper first, and a healthy managed install reported studio_install_incomplete with a missing list copied from the wrong venv. Selecting the root is not enough: _installed_version() reads the running interpreter and req_root defaults to the caller's studio.txt, so both checks still answered for the wrong venv. Hand verify_install() that venv's own metadata, enumerated through Distribution.discover(context = ...path), which does not fall back to sys.path. The candidate order is untouched, so shadowed-tree detection is unchanged. setup.ps1 replaces pip, torch and triton before install_python_stack.py runs, so the manifest it drops is not dropped before the first mutation. A run killed in between kept a marker that still verifies while torch was half-replaced; drop it at the top of the dependency pass instead. setup.sh is unaffected, the stack is the first thing its pass runs, and a test now pins both. pip uninstall rewrites nothing that was fingerprinted, and cache_matches re-reads the cached studio_install_ok rather than re-checking, so a venv that lost a studio.txt package kept being served the healthy verdict. Fold a sorted hash of the installed dist-info names into the marker hash. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * A missing manifest helper is a torn install, not an old one studio/install_manifest.py ships in the same wheel as _studio_deps.py, so nothing legitimately has one without the other: a CLI predating both never reaches this code, and the desktop already calls such a CLI stale on desktop_manageability_version. Returning ok=true there reported a healthy install for a tree the package update had half replaced, and the preflight then launched a backend whose own run.py could be just as absent. Report it incomplete so repair runs. * Tighten comments across the install-detection changes * Validate Studio dependency readiness --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com> Co-authored-by: Wasim Yousef Said <wasimysdev@gmail.com>
237 lines
10 KiB
YAML
237 lines
10 KiB
YAML
# SPDX-License-Identifier: AGPL-3.0-only
|
|
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved.
|
|
|
|
# Verifies that `unsloth studio update --local` is idempotent: a fresh
|
|
# install via install.sh, followed by `unsloth studio update --local`,
|
|
# succeeds and is a no-op for the llama.cpp prebuilt (it should report
|
|
# "prebuilt up to date and validated", not re-run the source build).
|
|
#
|
|
# This catches regressions in setup.sh's update path that the existing
|
|
# GGUF / wheel jobs would miss because they only invoke install.sh once.
|
|
|
|
name: Unsloth Update CI
|
|
|
|
on:
|
|
pull_request:
|
|
paths:
|
|
- 'install.sh'
|
|
- 'scripts/uninstall.sh'
|
|
- 'studio/setup.sh'
|
|
- 'studio/install_python_stack.py'
|
|
- 'studio/install_llama_prebuilt.py'
|
|
- 'studio/backend/requirements/**'
|
|
- 'unsloth_cli/commands/studio.py'
|
|
- 'pyproject.toml'
|
|
- '.github/workflows/studio-update-smoke.yml'
|
|
push:
|
|
branches: [main, pip]
|
|
workflow_dispatch:
|
|
|
|
concurrency:
|
|
group: ${{ github.workflow }}-${{ github.ref }}
|
|
cancel-in-progress: true
|
|
|
|
permissions:
|
|
contents: read
|
|
|
|
jobs:
|
|
update-idempotency:
|
|
name: Unsloth Updating Tests
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 15
|
|
steps:
|
|
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
|
with:
|
|
persist-credentials: false
|
|
|
|
- name: Linux deps for llama.cpp prebuilt
|
|
run: |
|
|
sudo apt-get update
|
|
sudo apt-get install -y --no-install-recommends \
|
|
libcurl4-openssl-dev libssl-dev jq
|
|
|
|
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
|
|
with:
|
|
node-version: '22'
|
|
|
|
- uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6.2.0
|
|
with:
|
|
python-version: '3.12'
|
|
# Don't cache pip: this job runs `bash install.sh` and
|
|
# `unsloth studio update --local` which both go through
|
|
# `uv` and never populate ~/.cache/pip. setup-python's
|
|
# post-step then fatal-errors with "Cache folder path is
|
|
# retrieved for pip but doesn't exist on disk".
|
|
|
|
- name: Install Unsloth (--local, --no-torch)
|
|
# Pass the workflow token so the llama.cpp prebuilt installer's
|
|
# GitHub-API call to list releases isn't rate-limited (60/hr
|
|
# unauthenticated). Without this, three consecutive install +
|
|
# update + update calls in this job exceed the limit and the
|
|
# prebuilt path falls back to source build.
|
|
env:
|
|
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
|
# Withheld on PR: this step runs checked-out PR code; public GGUF still downloads.
|
|
HF_TOKEN: ${{ github.event_name != 'pull_request' && secrets.HF_TOKEN || '' }}
|
|
run: |
|
|
mkdir -p logs
|
|
set -o pipefail
|
|
bash install.sh --local --no-torch 2>&1 | tee logs/install.log
|
|
|
|
- name: First update should be a no-op (prebuilt already validated)
|
|
# `unsloth studio update --local` runs studio/setup.sh against
|
|
# the local repo. Right after install.sh the llama.cpp prebuilt
|
|
# has just been installed and validated, so the second run must
|
|
# take the "prebuilt up to date and validated" code path. Any
|
|
# source-build fallback or re-download here means setup.sh's
|
|
# idempotency regressed.
|
|
env:
|
|
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
|
# Withheld on PR: this step runs checked-out PR code; public GGUF still downloads.
|
|
HF_TOKEN: ${{ github.event_name != 'pull_request' && secrets.HF_TOKEN || '' }}
|
|
run: |
|
|
set -o pipefail
|
|
unsloth studio update --local 2>&1 | tee logs/update.log
|
|
if grep -q "falling back to source build" logs/update.log; then
|
|
echo "::error::studio update fell back to source-build llama.cpp on a fresh install. setup.sh idempotency regressed."
|
|
grep -E "llama-prebuilt|llama.cpp" logs/update.log | tail -60
|
|
exit 1
|
|
fi
|
|
if ! grep -qE "prebuilt up to date and validated|prebuilt installed and validated" logs/update.log; then
|
|
echo "::error::no prebuilt up-to-date marker in update.log. Did setup.sh skip the prebuilt path on update?"
|
|
grep -E "llama-prebuilt|llama.cpp" logs/update.log | tail -60
|
|
exit 1
|
|
fi
|
|
echo "update path took the prebuilt fast path"
|
|
|
|
- name: Second update must also be a no-op
|
|
# Two consecutive `update`s back-to-back is the usual desktop
|
|
# flow (auto-update, then user-triggered update). Asserting the
|
|
# second run is also clean rules out hidden state changes from
|
|
# the first one.
|
|
env:
|
|
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
|
# Withheld on PR: this step runs checked-out PR code; public GGUF still downloads.
|
|
HF_TOKEN: ${{ github.event_name != 'pull_request' && secrets.HF_TOKEN || '' }}
|
|
run: |
|
|
set -o pipefail
|
|
unsloth studio update --local 2>&1 | tee logs/update2.log
|
|
grep -q "falling back to source build" logs/update2.log && {
|
|
echo "::error::second update fell back to source build"
|
|
tail -60 logs/update2.log; exit 1; } || true
|
|
grep -qE "prebuilt up to date and validated|prebuilt installed and validated" logs/update2.log
|
|
echo "second update was clean"
|
|
|
|
- name: Boot Unsloth briefly to confirm the install is still usable
|
|
# If `update --local` accidentally broke the venv or wiped the
|
|
# llama-server binary, the server would fail to start here.
|
|
run: |
|
|
mkdir -p logs
|
|
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p 18891 \
|
|
> logs/studio.log 2>&1 &
|
|
PID=$!
|
|
for i in $(seq 1 60); do
|
|
if curl -fs http://127.0.0.1:18891/api/health > /tmp/health.json; then
|
|
jq -e '.status == "healthy"' /tmp/health.json
|
|
break
|
|
fi
|
|
sleep 1
|
|
done
|
|
if ! jq -e '.status == "healthy"' /tmp/health.json 2>/dev/null; then
|
|
echo "Unsloth failed to come up after `update`"
|
|
tail -200 logs/studio.log
|
|
kill "$PID" 2>/dev/null || true
|
|
exit 1
|
|
fi
|
|
kill "$PID" 2>/dev/null || true
|
|
echo "post-update Unsloth /api/health OK"
|
|
|
|
- name: A complete install reports itself complete
|
|
run: |
|
|
set -o pipefail
|
|
unsloth studio verify-install
|
|
unsloth studio desktop-capabilities --json | tee /tmp/caps.json
|
|
jq -e '.studio_install_ok == true' /tmp/caps.json
|
|
jq -e '.desktop_manageability_version >= 2' /tmp/caps.json
|
|
|
|
- name: An incomplete install must not report itself ready
|
|
# An installer killed part-way leaves a working CLI but no studio.txt
|
|
# deps, which the old preflight called ManagedReady. The manifest is
|
|
# written last, so removing it reproduces that state.
|
|
run: |
|
|
set -o pipefail
|
|
# install.sh's default root, resolved explicitly: `python` on PATH
|
|
# here is setup-python's, not the managed venv.
|
|
MANIFEST="$HOME/.unsloth/studio/unsloth_studio/unsloth_install_manifest.json"
|
|
test -f "$MANIFEST" || { echo "::error::installer never wrote $MANIFEST"; exit 1; }
|
|
rm -f "$MANIFEST"
|
|
unsloth studio desktop-capabilities --json | tee /tmp/caps_bad.json
|
|
jq -e '.studio_install_ok == false' /tmp/caps_bad.json
|
|
if unsloth studio verify-install; then
|
|
echo "::error::verify-install passed on an install with no manifest"
|
|
exit 1
|
|
fi
|
|
echo "incomplete install correctly reported not-ready"
|
|
|
|
- name: Update repairs an incomplete install
|
|
# `--local` bypasses setup.sh's PyPI version compare, so this asserts
|
|
# the repair OUTCOME. The non-local fast path the desktop Repair button
|
|
# uses is covered by tests/studio/install/test_setup_fast_path_guard.py.
|
|
env:
|
|
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
|
run: |
|
|
set -o pipefail
|
|
unsloth studio update --local 2>&1 | tee logs/update_repair.log
|
|
unsloth studio verify-install
|
|
unsloth studio desktop-capabilities --json | jq -e '.studio_install_ok == true'
|
|
echo "update repaired the incomplete install"
|
|
|
|
- name: Uninstall and verify clean
|
|
# Round-trip the installer through scripts/uninstall.sh: confirms the
|
|
# uninstaller actually finds and removes everything install.sh +
|
|
# update wrote. Safety-guard scenarios (refuse-$HOME etc.) belong
|
|
# in a separate fast smoke job; this is the happy-path cleanup
|
|
# assertion that catches regressions where install.sh starts
|
|
# writing to a new location and scripts/uninstall.sh hasn't caught up.
|
|
# Skips gracefully if scripts/uninstall.sh has not landed yet (lets
|
|
# this workflow merge before #5497).
|
|
run: |
|
|
set -o pipefail
|
|
if [ ! -f scripts/uninstall.sh ]; then
|
|
echo "scripts/uninstall.sh not present in this tree; skipping round-trip"
|
|
: > logs/uninstall.log
|
|
exit 0
|
|
fi
|
|
sh scripts/uninstall.sh 2>&1 | tee logs/uninstall.log
|
|
leak=0
|
|
for p in \
|
|
"$HOME/.unsloth/studio" \
|
|
"$HOME/.local/share/unsloth" \
|
|
"$HOME/Desktop/Unsloth Studio.desktop" \
|
|
"$HOME/.local/bin/unsloth"; do
|
|
if [ -e "$p" ] || [ -L "$p" ]; then
|
|
echo "::error::leak: $p"
|
|
ls -la "$p" 2>&1 | head -3
|
|
leak=$((leak + 1))
|
|
fi
|
|
done
|
|
[ "$leak" -eq 0 ] || exit 1
|
|
# Idempotent: re-runs exit 0 on an empty $HOME.
|
|
sh scripts/uninstall.sh 2>&1 | tail -5
|
|
sh scripts/uninstall.sh 2>&1 | tail -5
|
|
echo "PASS: install -> update -> uninstall round-trip clean"
|
|
|
|
- name: Upload update logs
|
|
# Always upload so a green run still leaves the install + two
|
|
# update logs + uninstall log reviewable.
|
|
if: always()
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
|
with:
|
|
name: studio-update-log
|
|
path: |
|
|
logs/install.log
|
|
logs/update.log
|
|
logs/update2.log
|
|
logs/studio.log
|
|
logs/uninstall.log
|
|
retention-days: 7
|