Replace standalone Studio wording with Unsloth (#7221)
* Replace standalone Studio wording with Unsloth Replace the single word Studio with Unsloth wherever it is used as shorthand for Unsloth Studio in docs, CLI output, UI strings, i18n locales, workflow display names, comments and docstrings. Kept unchanged: the full name Unsloth Studio, third party product names (LM Studio, Visual Studio, Mac Studio), feature names (Recipe Studio, Fine-tuning Studio and its translations), and all identifiers such as env vars, commands, paths and filenames. * Address review feedback on the Studio wording rename Use "an" before Unsloth where the rename left the article as "a". Restore the split brand where Unsloth and Studio render as two halves of the full product name: the onboarding sidebar subtitle and the IPv6 localhost warning. Scope two messages to the full name Unsloth Studio where plain Unsloth was misleading: the AMD README bullet and the CLI studio setup error.
This commit is contained in:
parent
e9ef2ac60f
commit
6d8c18cd1a
264 changed files with 1095 additions and 1095 deletions
2
.gitattributes
vendored
2
.gitattributes
vendored
|
|
@ -6,7 +6,7 @@
|
|||
# them when run in WSL/Linux (e.g. `set -e` -> "set: Illegal option -").
|
||||
*.sh text eol=lf
|
||||
|
||||
# Normalize Studio frontend sources to LF. Scoped to the frontend tree (rather
|
||||
# Normalize Unsloth frontend sources to LF. Scoped to the frontend tree (rather
|
||||
# than repo-wide *.ts/*.tsx/... rules) so the policy can't force LF on files
|
||||
# elsewhere. text=auto lets Git detect and leave binary assets (logos, fonts)
|
||||
# untouched while text files (.ts/.tsx/.json/.html/.svg/...) are stored as LF.
|
||||
|
|
|
|||
4
.github/scripts/agent-guides-drive.sh
vendored
4
.github/scripts/agent-guides-drive.sh
vendored
|
|
@ -166,8 +166,8 @@ parse_connect() {
|
|||
echo "[$AGENT] connect --no-launch printed:"; cat_redacted "$raw"
|
||||
CONNECT_ENV="$(grep -E '^(export |unset )' "$raw" || true)"
|
||||
# The launch command is the last non-export, non-status line. start.py
|
||||
# prints "Studio <url> · model <id>" and "Updated ..." status lines first.
|
||||
CONNECT_CMD="$(grep -vE '^(export |unset |Studio |Updated |Disabled |Warning|Loading)' "$raw" \
|
||||
# prints "Unsloth <url> · model <id>" and "Updated ..." status lines first.
|
||||
CONNECT_CMD="$(grep -vE '^(export |unset |Unsloth |Updated |Disabled |Warning|Loading)' "$raw" \
|
||||
| grep -E '[^[:space:]]' | tail -1)"
|
||||
[ -n "$CONNECT_CMD" ] || guide_fail "could not parse a launch command from connect --no-launch output"
|
||||
redact "$raw"
|
||||
|
|
|
|||
2
.github/scripts/assert-llama-loads.sh
vendored
2
.github/scripts/assert-llama-loads.sh
vendored
|
|
@ -2,7 +2,7 @@
|
|||
# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved.
|
||||
#
|
||||
# Assert Studio installed a llama.cpp that loads and runs on THIS macOS. Tests
|
||||
# Assert Unsloth installed a llama.cpp that loads and runs on THIS macOS. Tests
|
||||
# the contract that matters (binaries load and their minimum-OS is <= this host)
|
||||
# instead of the old "did install.sh fall back to a source build?" grep, since a
|
||||
# source build with a correct deployment target is a valid outcome.
|
||||
|
|
|
|||
2
.github/scripts/assert-prompt-cache.sh
vendored
2
.github/scripts/assert-prompt-cache.sh
vendored
|
|
@ -31,7 +31,7 @@
|
|||
# (llama_cpp.py:337-340). So default: ~/.unsloth/studio/logs/llama-server/.
|
||||
#
|
||||
# <P> is the INTERNAL llama-server port (self._find_free_port(),
|
||||
# llama_cpp.py:3489 / :4641) -- a RANDOM port, NOT the Studio port. So we must
|
||||
# llama_cpp.py:3489 / :4641) -- a RANDOM port, NOT the Unsloth port. So we must
|
||||
# NOT filter the log glob by STUDIO_PORT (the brief's `port-<STUDIO_PORT>`
|
||||
# glob would never match). We pick the newest llama-*.log instead.
|
||||
#
|
||||
|
|
|
|||
4
.github/scripts/hf-download-with-retry.sh
vendored
4
.github/scripts/hf-download-with-retry.sh
vendored
|
|
@ -3,7 +3,7 @@
|
|||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
||||
#
|
||||
# Download a single file from a Hugging Face repo with a stall-retry
|
||||
# watchdog. Used by the Studio CI workflows so a hung hf-xet transfer
|
||||
# watchdog. Used by the Unsloth CI workflows so a hung hf-xet transfer
|
||||
# kills + retries instead of silently consuming the job's timeout.
|
||||
#
|
||||
# Usage: hf-download-with-retry.sh REPO FILE LOCAL_DIR
|
||||
|
|
@ -35,7 +35,7 @@ REPO="${1:?usage: hf-download-with-retry.sh REPO FILE [LOCAL_DIR]}"
|
|||
FILE="${2:?usage: hf-download-with-retry.sh REPO FILE [LOCAL_DIR]}"
|
||||
# LOCAL_DIR is optional. If empty, hf falls back to HF_HUB_CACHE
|
||||
# (~/.cache/huggingface/hub) which is the desired path for callers
|
||||
# that populate HF_HOME for a downstream Studio model load.
|
||||
# that populate HF_HOME for a downstream Unsloth model load.
|
||||
LOCAL_DIR="${3:-}"
|
||||
|
||||
# Stall threshold per attempt, in seconds. Override with
|
||||
|
|
|
|||
4
.github/workflows/lint-ci.yml
vendored
4
.github/workflows/lint-ci.yml
vendored
|
|
@ -13,10 +13,10 @@
|
|||
# committed YAML / JSON config.
|
||||
#
|
||||
# TypeScript and Rust are NOT duplicated here on purpose:
|
||||
# - Studio Frontend CI runs `npm run typecheck` (= `tsc --noEmit`)
|
||||
# - Unsloth Frontend CI runs `npm run typecheck` (= `tsc --noEmit`)
|
||||
# and `npm run build` (vite/swc) on every studio/frontend/**
|
||||
# change, which is a full TS AST + type check.
|
||||
# - Studio Tauri CI runs `tauri build --debug --no-bundle` on
|
||||
# - Unsloth Tauri CI runs `tauri build --debug --no-bundle` on
|
||||
# every studio/src-tauri/** or studio/frontend/** change, which
|
||||
# compiles the Rust crate (= cargo check + cargo build).
|
||||
# Each is a stricter check than a parse-only step would be, so a
|
||||
|
|
|
|||
16
.github/workflows/local-agent-guides-ci.yml
vendored
16
.github/workflows/local-agent-guides-ci.yml
vendored
|
|
@ -154,7 +154,7 @@ jobs:
|
|||
path: gguf-cache
|
||||
key: ${{ runner.os }}-gguf-${{ env.GGUF_REPO }}-${{ env.GGUF_FILE }}-v1
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
# Gated off PR (see note above); public GGUF still downloads.
|
||||
|
|
@ -256,7 +256,7 @@ jobs:
|
|||
done
|
||||
fi
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
# Guard the PID: an unset/zero UNSLOTH_SERVER_PID would make
|
||||
|
|
@ -359,7 +359,7 @@ jobs:
|
|||
path: gguf-cache
|
||||
key: ${{ runner.os }}-gguf-${{ env.GGUF_REPO }}-${{ env.GGUF_FILE }}-v1
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
# Gated off PR (see note above); public GGUF still downloads.
|
||||
|
|
@ -448,7 +448,7 @@ jobs:
|
|||
done
|
||||
fi
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
# Guard the PID: an unset/zero UNSLOTH_SERVER_PID would make
|
||||
|
|
@ -543,7 +543,7 @@ jobs:
|
|||
path: gguf-cache
|
||||
key: ${{ runner.os }}-gguf-${{ env.GGUF_REPO }}-${{ env.GGUF_FILE }}-v1
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
HF_TOKEN: ${{ secrets.HF_TOKEN }}
|
||||
|
|
@ -620,7 +620,7 @@ jobs:
|
|||
done
|
||||
fi
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
if [ -n "${UNSLOTH_SERVER_PID:-}" ] && [ "${UNSLOTH_SERVER_PID}" != "0" ]; then
|
||||
|
|
@ -706,7 +706,7 @@ jobs:
|
|||
path: hf-cache
|
||||
key: ${{ runner.os }}-hf-${{ env.GGUF_REPO }}-${{ env.GGUF_VARIANT }}-v2
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
# Gated off PR (see note above); public GGUF still downloads.
|
||||
|
|
@ -764,7 +764,7 @@ jobs:
|
|||
done
|
||||
fi
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
# Guard the PID: an unset/zero UNSLOTH_SERVER_PID would make
|
||||
|
|
|
|||
12
.github/workflows/mlx-ci.yml
vendored
12
.github/workflows/mlx-ci.yml
vendored
|
|
@ -130,7 +130,7 @@ jobs:
|
|||
# MLX support landed after the most recent unsloth-zoo PyPI
|
||||
# release; the wheel still raises NotImplementedError on
|
||||
# Apple Silicon when device_type.get_device_type() runs
|
||||
# unguarded. Studio's own install.sh overlays unsloth-zoo
|
||||
# unguarded. Unsloth's own install.sh overlays unsloth-zoo
|
||||
# from git main for the same reason. Pulling deps lets pip
|
||||
# resolve the platform-conditional MLX-only wheels (mlx,
|
||||
# mlx-lm, mlx-vlm gated on darwin+arm64 in unsloth-zoo's
|
||||
|
|
@ -317,13 +317,13 @@ jobs:
|
|||
echo
|
||||
done
|
||||
|
||||
# Validates the macOS prebuilt path Studio's setup.sh uses (#5963): install the
|
||||
# Validates the macOS prebuilt path Unsloth's setup.sh uses (#5963): install the
|
||||
# unslothai/llama.cpp fork's latest release, download a small public GGUF, and
|
||||
# check llama-server /completion end to end. Split and placed last so the
|
||||
# untrusted binary runs only in the final smoke step, after every HF_TOKEN step,
|
||||
# leaving no token-bearing step or shared workspace for a tampered prebuilt to
|
||||
# corrupt. GH_TOKEN: releases API; HF_TOKEN (withheld on PR): probe + GGUF fetch.
|
||||
- name: Studio prebuilt llama.cpp install + GGUF download (Mac M1)
|
||||
- name: Unsloth prebuilt llama.cpp install + GGUF download (Mac M1)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
|
|
@ -344,12 +344,12 @@ jobs:
|
|||
|
||||
# Final step: runs the downloaded binaries with no secrets present, and clears
|
||||
# the GitHub Actions command files so a tampered prebuilt cannot influence the job.
|
||||
- name: Studio prebuilt llama.cpp GGUF inference smoke (Mac M1)
|
||||
- name: Unsloth prebuilt llama.cpp GGUF inference smoke (Mac M1)
|
||||
run: |
|
||||
set -euo pipefail
|
||||
unset GITHUB_ENV GITHUB_PATH GITHUB_OUTPUT GITHUB_STEP_SUMMARY
|
||||
INSTALL_DIR="$HOME/.unsloth-studio-prebuilt-test/llama.cpp"
|
||||
# Studio bundles only llama-server + llama-quantize (not llama-cli);
|
||||
# Unsloth bundles only llama-server + llama-quantize (not llama-cli);
|
||||
# inference goes through llama-server's HTTP /completion endpoint.
|
||||
LLAMA_SERVER="$INSTALL_DIR/build/bin/llama-server"
|
||||
LLAMA_QUANT="$INSTALL_DIR/build/bin/llama-quantize"
|
||||
|
|
@ -400,4 +400,4 @@ jobs:
|
|||
tail -40 /tmp/llama-server.log
|
||||
exit 1
|
||||
fi
|
||||
echo "OK: Studio prebuilt llama.cpp on Mac M1 + GGUF /completion works"
|
||||
echo "OK: Unsloth prebuilt llama.cpp on Mac M1 + GGUF /completion works"
|
||||
|
|
|
|||
8
.github/workflows/release-desktop.yml
vendored
8
.github/workflows/release-desktop.yml
vendored
|
|
@ -4,7 +4,7 @@ on:
|
|||
workflow_dispatch:
|
||||
inputs:
|
||||
studio_version:
|
||||
description: 'Studio version tag to release (for example, v0.1.39-beta)'
|
||||
description: 'Unsloth version tag to release (for example, v0.1.39-beta)'
|
||||
type: string
|
||||
required: true
|
||||
pypi_version:
|
||||
|
|
@ -69,7 +69,7 @@ jobs:
|
|||
if not studio_version:
|
||||
sys.exit('studio_version is required, for example v0.1.39-beta')
|
||||
if re.fullmatch(r'v?20\d{2}\.\d+\.\d+(?:[-+][0-9A-Za-z.-]+)?', studio_version):
|
||||
sys.exit(f'studio_version must be a Studio SemVer tag, not a date-style backend version: {studio_version}')
|
||||
sys.exit(f'studio_version must be an Unsloth SemVer tag, not a date-style backend version: {studio_version}')
|
||||
|
||||
semver_tag = re.compile(
|
||||
r'^v(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)'
|
||||
|
|
@ -146,7 +146,7 @@ jobs:
|
|||
print(f'pypi_version={pypi_version}', file=output)
|
||||
PY
|
||||
|
||||
- name: Verify PyPI package and Studio stamp
|
||||
- name: Verify PyPI package and Unsloth stamp
|
||||
shell: bash
|
||||
env:
|
||||
STUDIO_VERSION: ${{ steps.prepare.outputs.studio_version }}
|
||||
|
|
@ -211,7 +211,7 @@ jobs:
|
|||
fi
|
||||
python3 scripts/stamp_studio_release.py --verify-dist "$RUNNER_TEMP/pypi-unsloth-dist" --expected "$STUDIO_VERSION"
|
||||
else
|
||||
echo "scripts/stamp_studio_release.py not found; release-desktop requires #5308 to verify the PyPI Studio stamp." >&2
|
||||
echo "scripts/stamp_studio_release.py not found; release-desktop requires #5308 to verify the PyPI Unsloth stamp." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
|
|
|
|||
30
.github/workflows/security-audit.yml
vendored
30
.github/workflows/security-audit.yml
vendored
|
|
@ -36,8 +36,8 @@
|
|||
# - unsloth `huggingfacenotorch` extras (the canonical install path
|
||||
# for fine-tuning users; pulls transformers / peft / accelerate /
|
||||
# trl / datasets / diffusers / sentence-transformers / etc.)
|
||||
# - all six Studio backend requirements files
|
||||
# - Studio frontend (npm) and Tauri shell (cargo)
|
||||
# - all six Unsloth backend requirements files
|
||||
# - Unsloth frontend (npm) and Tauri shell (cargo)
|
||||
# Each Python step builds a filtered dep list from pyproject.toml +
|
||||
# requirements/*.txt before auditing. We do NOT install any of these
|
||||
# -- pip-audit resolves through PyPI metadata, scan_packages.py
|
||||
|
|
@ -218,7 +218,7 @@ jobs:
|
|||
# on the runner). A comment line is left in place so the
|
||||
# skipped specs are obvious in the artifact.
|
||||
# The `huggingface` extra is `huggingfacenotorch` plus torch /
|
||||
# torchvision / triton, deliberately skipped: Studio backend
|
||||
# torchvision / triton, deliberately skipped: Unsloth backend
|
||||
# already pins a torch and the +cu* / +cpu local-version tags
|
||||
# trip up the PyPI resolver in `-r` mode.
|
||||
run: |
|
||||
|
|
@ -253,7 +253,7 @@ jobs:
|
|||
# `-r requirements.txt` resolves the requirements through pip's
|
||||
# dependency resolver against PyPI metadata and audits the
|
||||
# resolved tree without ever executing setup.py / install
|
||||
# hooks. Way faster than installing the full Studio runtime
|
||||
# hooks. Way faster than installing the full Unsloth runtime
|
||||
# and -- critically -- safer: an attacker who has compromised
|
||||
# a transitive dep cannot run code in this job.
|
||||
#
|
||||
|
|
@ -326,9 +326,9 @@ jobs:
|
|||
} >> "$GITHUB_STEP_SUMMARY"
|
||||
|
||||
# ─────────────────────────────────────────────────────────────
|
||||
# npm: Studio frontend
|
||||
# npm: Unsloth frontend
|
||||
# ─────────────────────────────────────────────────────────────
|
||||
- name: npm audit (Studio frontend)
|
||||
- name: npm audit (Unsloth frontend)
|
||||
# `npm audit` resolves the lockfile through the npmjs.com
|
||||
# advisory DB. `--audit-level=high` filters the noise floor
|
||||
# to only HIGH and CRITICAL. We do NOT pass --omit=dev: a
|
||||
|
|
@ -342,7 +342,7 @@ jobs:
|
|||
# Always also write the full JSON for grep-ability.
|
||||
npm audit --json > ../../logs-npm-audit.json || true
|
||||
{
|
||||
echo "## npm audit (Studio frontend)"
|
||||
echo "## npm audit (Unsloth frontend)"
|
||||
echo
|
||||
echo '```'
|
||||
tail -200 ../../logs-npm-audit.txt
|
||||
|
|
@ -350,9 +350,9 @@ jobs:
|
|||
} >> "$GITHUB_STEP_SUMMARY"
|
||||
|
||||
# ─────────────────────────────────────────────────────────────
|
||||
# cargo: Studio Tauri shell
|
||||
# cargo: Unsloth Tauri shell
|
||||
# ─────────────────────────────────────────────────────────────
|
||||
- name: cargo audit (Studio Tauri)
|
||||
- name: cargo audit (Unsloth Tauri)
|
||||
# `--deny warnings` would make the job fail on any advisory.
|
||||
# Keep non-blocking initially; drop continue-on-error after
|
||||
# the baseline closes.
|
||||
|
|
@ -362,7 +362,7 @@ jobs:
|
|||
set +e
|
||||
cargo audit | tee ../../logs-cargo-audit.txt
|
||||
{
|
||||
echo "## cargo audit (Studio Tauri)"
|
||||
echo "## cargo audit (Unsloth Tauri)"
|
||||
echo
|
||||
echo '```'
|
||||
tail -200 ../../logs-cargo-audit.txt
|
||||
|
|
@ -559,7 +559,7 @@ jobs:
|
|||
|
||||
# ─────────────────────────────────────────────────────────────
|
||||
# CycloneDX SBOM. Lets downstream consumers audit what's
|
||||
# actually shipped in unsloth wheels and the Studio backend
|
||||
# actually shipped in unsloth wheels and the Unsloth backend
|
||||
# runtime. Generates one JSON file per requirements input plus
|
||||
# a combined SBOM keyed off pyproject.toml; uploads as a build
|
||||
# artifact (and a future step can attest it via SLSA).
|
||||
|
|
@ -740,7 +740,7 @@ jobs:
|
|||
# `--with-deps` makes the scan transitive: every package the
|
||||
# declared set resolves to gets fetched and pattern-scanned, not
|
||||
# just the top-level pins. Resolving the full transitive closure
|
||||
# of the unsloth + Studio dep tree downloads several hundred
|
||||
# of the unsloth + Unsloth dep tree downloads several hundred
|
||||
# archives, hence the longer timeout.
|
||||
#
|
||||
# Sharded across runners for wall-clock parallelism. Each shard
|
||||
|
|
@ -749,7 +749,7 @@ jobs:
|
|||
# composition tries to balance load:
|
||||
# - hf-stack: pyproject extras + no-torch-runtime
|
||||
# (~150 archives, transformers/peft/accelerate/...)
|
||||
# - studio: FastAPI/Studio backend + overrides + extras-no-deps
|
||||
# - studio: FastAPI/Unsloth backend + overrides + extras-no-deps
|
||||
# (~150 archives, smaller scientific stack)
|
||||
# - extras: the heavy openai-whisper / scikit-learn / librosa
|
||||
# stack (~250 archives, dominant cost)
|
||||
|
|
@ -964,7 +964,7 @@ jobs:
|
|||
# documented at scripts/scan_npm_packages.py top-of-file. The
|
||||
# script is stdlib-only so adding it does not increase the
|
||||
# transitive supply-chain surface.
|
||||
name: npm scan-packages (Studio frontend tarballs)
|
||||
name: npm scan-packages (Unsloth frontend tarballs)
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
needs: []
|
||||
|
|
@ -1173,7 +1173,7 @@ jobs:
|
|||
with:
|
||||
python-version: '3.12'
|
||||
|
||||
- name: Install Studio frontend deps (--ignore-scripts)
|
||||
- name: Install Unsloth frontend deps (--ignore-scripts)
|
||||
# `npm audit signatures` requires node_modules to be populated.
|
||||
# `--ignore-scripts` is mandatory: this is exactly the lever the
|
||||
# new-install-script gate below protects against, and we must
|
||||
|
|
|
|||
14
.github/workflows/studio-api-smoke.yml
vendored
14
.github/workflows/studio-api-smoke.yml
vendored
|
|
@ -1,7 +1,7 @@
|
|||
# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved.
|
||||
|
||||
# Studio API & Auth Tests -- HTTP-level integration tests for the
|
||||
# Unsloth API & Auth Tests -- HTTP-level integration tests for the
|
||||
# FastAPI surface. No Playwright, no model UI; tests/studio/test_studio_api_smoke.py
|
||||
# runs ~30 s and asserts:
|
||||
# - CORS hardening (no wildcard + credentials, no bootstrap leak)
|
||||
|
|
@ -15,7 +15,7 @@
|
|||
# Reuses the GGUF cache key from studio-ui-smoke.yml so the model
|
||||
# download is one cache-hit on the second job.
|
||||
|
||||
name: Studio API CI
|
||||
name: Unsloth API CI
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
|
|
@ -40,7 +40,7 @@ permissions:
|
|||
|
||||
jobs:
|
||||
api-smoke:
|
||||
name: Studio API & Auth Tests
|
||||
name: Unsloth API & Auth Tests
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 12
|
||||
env:
|
||||
|
|
@ -98,7 +98,7 @@ jobs:
|
|||
path: hf-cache
|
||||
key: ${{ runner.os }}-hf-${{ env.GGUF_REPO }}-${{ env.GGUF_VARIANT }}-v2
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
# Withheld on PR: this step runs checked-out PR code; public GGUF still downloads.
|
||||
|
|
@ -111,7 +111,7 @@ jobs:
|
|||
- name: Install pyjwt for the JWT-expiry forge test
|
||||
run: pip install 'pyjwt>=2.6'
|
||||
|
||||
- name: Reset auth + boot Studio (API-only)
|
||||
- name: Reset auth + boot Unsloth (API-only)
|
||||
run: |
|
||||
unsloth studio reset-password
|
||||
mkdir -p logs
|
||||
|
|
@ -144,7 +144,7 @@ jobs:
|
|||
echo "STUDIO_NEW_PW=$NEW" >> "$GITHUB_ENV"
|
||||
echo "STUDIO_NEW2_PW=$NEW2" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Run Studio API & Auth tests
|
||||
- name: Run Unsloth API & Auth tests
|
||||
# The script is named WITHOUT a `test_` prefix so it isn't
|
||||
# auto-collected by pytest in Backend CI's `tests/` walk
|
||||
# (which doesn't set BASE_URL and would crash at import).
|
||||
|
|
@ -153,7 +153,7 @@ jobs:
|
|||
STUDIO_AUTH_DIR: /home/runner/.unsloth/studio/auth
|
||||
run: python tests/studio/studio_api_smoke.py
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
kill "${STUDIO_PID}" 2>/dev/null || true
|
||||
|
|
|
|||
2
.github/workflows/studio-backend-ci.yml
vendored
2
.github/workflows/studio-backend-ci.yml
vendored
|
|
@ -64,7 +64,7 @@ jobs:
|
|||
- name: Install backend test dependencies (CPU only)
|
||||
run: |
|
||||
python -m pip install --upgrade pip
|
||||
# Studio's declared backend deps:
|
||||
# Unsloth's declared backend deps:
|
||||
pip install -r studio/backend/requirements/studio.txt
|
||||
# Extras that studio.txt does not list but the import chain needs
|
||||
# (python-multipart for FastAPI form/file uploads, sqlalchemy/cryptography
|
||||
|
|
|
|||
|
|
@ -9,7 +9,7 @@
|
|||
# export is validated separately. No GPU / model / llama.cpp: the tests mock the probes and block
|
||||
# torch/unsloth, so the job installs only a CPU PyTorch plus import deps.
|
||||
|
||||
name: Studio export capability
|
||||
name: Unsloth export capability
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
|
|
|
|||
4
.github/workflows/studio-frontend-ci.yml
vendored
4
.github/workflows/studio-frontend-ci.yml
vendored
|
|
@ -136,7 +136,7 @@ jobs:
|
|||
- name: Build
|
||||
run: npm run build
|
||||
|
||||
- name: Built bundle must not contain Studio's unstable_Provider call site
|
||||
- name: Built bundle must not contain Unsloth's unstable_Provider call site
|
||||
run: |
|
||||
set -e
|
||||
JS=$(ls dist/assets/index-*.js | head -1)
|
||||
|
|
@ -144,7 +144,7 @@ jobs:
|
|||
echo "main bundle: $JS"
|
||||
echo "unstable_Provider: hits=$HITS (assistant-ui internals contribute up to 3)"
|
||||
if [ "$HITS" -gt 3 ]; then
|
||||
echo "::error file=studio/frontend/src/features/chat/runtime-provider.tsx::Studio bundle still passes unstable_Provider through useRemoteThreadListRuntime; this is the 2026.5.1 chat-history regression. Pass adapters directly into useLocalRuntime instead."
|
||||
echo "::error file=studio/frontend/src/features/chat/runtime-provider.tsx::Unsloth bundle still passes unstable_Provider through useRemoteThreadListRuntime; this is the 2026.5.1 chat-history regression. Pass adapters directly into useLocalRuntime instead."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
|
|
|
|||
56
.github/workflows/studio-inference-smoke.yml
vendored
56
.github/workflows/studio-inference-smoke.yml
vendored
|
|
@ -1,7 +1,7 @@
|
|||
# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved.
|
||||
|
||||
# Three end-to-end smoke jobs that boot a freshly-installed Studio and
|
||||
# Three end-to-end smoke jobs that boot a freshly-installed Unsloth and
|
||||
# exercise the surfaces real users hit through the OpenAI / Anthropic
|
||||
# SDKs and curl. Each job picks the smallest model that exercises the
|
||||
# behaviour under test, primes HF_HOME via actions/cache, and shares
|
||||
|
|
@ -27,7 +27,7 @@
|
|||
# All three jobs run in parallel. Total wall time is dominated by job 3
|
||||
# on a cold cache; warm cache cuts that to ~3 min.
|
||||
|
||||
name: Studio GGUF CI
|
||||
name: Unsloth GGUF CI
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
|
|
@ -112,7 +112,7 @@ jobs:
|
|||
path: hf-cache
|
||||
key: ${{ runner.os }}-hf-${{ env.GGUF_REPO }}-${{ env.GGUF_VARIANT }}-v2
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
# Withheld on PR: this step runs checked-out PR code; public GGUF still downloads.
|
||||
|
|
@ -125,7 +125,7 @@ jobs:
|
|||
- name: Install OpenAI + Anthropic Python SDKs
|
||||
run: pip install 'openai>=1.50' 'anthropic>=0.40'
|
||||
|
||||
- name: Reset auth + boot Studio (API-only)
|
||||
- name: Reset auth + boot Unsloth (API-only)
|
||||
run: |
|
||||
unsloth studio reset-password
|
||||
mkdir -p logs
|
||||
|
|
@ -142,7 +142,7 @@ jobs:
|
|||
fi
|
||||
sleep 1
|
||||
done
|
||||
echo "Studio did not become healthy in 180s"
|
||||
echo "Unsloth did not become healthy in 180s"
|
||||
tail -200 logs/studio.log
|
||||
exit 1
|
||||
|
||||
|
|
@ -229,11 +229,11 @@ jobs:
|
|||
return replies
|
||||
|
||||
def run_anthropic():
|
||||
# Two SDK quirks vs. Studio:
|
||||
# Two SDK quirks vs. Unsloth:
|
||||
# 1. base_url must NOT include /v1 -- the SDK appends
|
||||
# /v1/messages itself; otherwise the request hits
|
||||
# /v1/v1/messages and 405s.
|
||||
# 2. The SDK sends `x-api-key` by default, but Studio's
|
||||
# 2. The SDK sends `x-api-key` by default, but Unsloth's
|
||||
# auth layer is HTTPBearer-only. Override via
|
||||
# default_headers so Authorization: Bearer ... is
|
||||
# sent instead.
|
||||
|
|
@ -276,7 +276,7 @@ jobs:
|
|||
print(
|
||||
f"[{label}] WARN non-determinism at temperature=0.0 across "
|
||||
f"{len(determinism_failures)} of {len(first)} turn(s); "
|
||||
f"small-quant model drift, not a Studio regression. "
|
||||
f"small-quant model drift, not an Unsloth regression. "
|
||||
f"Details: " + " | ".join(determinism_failures)
|
||||
)
|
||||
# Sanity: turn-2 reply should mention the earlier question, and
|
||||
|
|
@ -290,7 +290,7 @@ jobs:
|
|||
print(f"[{label}] {status_word} -- 4 turns, history grounded ('paris' present)")
|
||||
PY
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
kill "${STUDIO_PID}" 2>/dev/null || true
|
||||
|
|
@ -323,7 +323,7 @@ jobs:
|
|||
# store xet chunks + blobs + snapshots = ~4 GiB compressed --
|
||||
# 4-5x file-size inflation, dominated by xet chunks. Use main's
|
||||
# `--local-dir gguf-cache` pattern to cache the flat .gguf only.
|
||||
# Studio's /api/inference/load accepts either a HF repo (which
|
||||
# Unsloth's /api/inference/load accepts either a HF repo (which
|
||||
# uses HF_HOME) or an absolute file path; passing the absolute
|
||||
# path keeps the test off HF_HOME entirely so the cache size
|
||||
# tracks the GGUF file 1:1. The OpenAI/Anth and JSON+images
|
||||
|
|
@ -380,7 +380,7 @@ jobs:
|
|||
path: gguf-cache
|
||||
key: ${{ runner.os }}-gguf-${{ env.GGUF_REPO }}-${{ env.GGUF_FILE }}-v1
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
# Withheld on PR: this step runs checked-out PR code; public GGUF still downloads.
|
||||
|
|
@ -390,7 +390,7 @@ jobs:
|
|||
set -o pipefail
|
||||
bash install.sh --local --no-torch 2>&1 | tee logs/install.log
|
||||
|
||||
- name: Reset auth + boot Studio (API-only, default tool policy)
|
||||
- name: Reset auth + boot Unsloth (API-only, default tool policy)
|
||||
# We deliberately use the API-only mode rather than
|
||||
# `unsloth studio run` because the latter calls
|
||||
# `set_tool_policy(...)` with a resolved bool: on loopback the
|
||||
|
|
@ -503,7 +503,7 @@ jobs:
|
|||
that the tool path executed.
|
||||
|
||||
A shared CI runner can stall the stream transport (the
|
||||
connection opening, or a mid-stream read) even when Studio
|
||||
connection opening, or a mid-stream read) even when Unsloth
|
||||
is healthy, so retry a stall once with a fresh request
|
||||
capped at 300s. A stall means the stream did NOT complete,
|
||||
so partial events are normally NOT returned (an early
|
||||
|
|
@ -575,11 +575,11 @@ jobs:
|
|||
|
||||
def _tool_invoked(events):
|
||||
"""Structural check: True iff some SSE payload is a real
|
||||
tool envelope (Studio tool_start/tool_end, Anthropic
|
||||
tool envelope (Unsloth tool_start/tool_end, Anthropic
|
||||
tool_use/tool_result, OpenAI non-empty delta.tool_calls /
|
||||
message.tool_calls / finish_reason='tool_calls' /
|
||||
role:'tool' / function_call). tool_status is NOT
|
||||
evidence: Studio emits empty tool_status events on
|
||||
evidence: Unsloth emits empty tool_status events on
|
||||
iteration boundaries even when no tool ran.
|
||||
"""
|
||||
for raw in events:
|
||||
|
|
@ -698,7 +698,7 @@ jobs:
|
|||
attempt has structural invocation evidence. WARN (not
|
||||
FAIL) if invoked but no attempt produces the expected
|
||||
literal in tool_end.result -- small-quant Qwen3.5-2B can
|
||||
emit OpenAI tool_calls deltas without Studio's GGUF
|
||||
emit OpenAI tool_calls deltas without Unsloth's GGUF
|
||||
agentic loop intercepting them, and that GGUF-vs-OpenAI
|
||||
format mismatch is out of scope for #5642.
|
||||
"""
|
||||
|
|
@ -811,7 +811,7 @@ jobs:
|
|||
# because (a) the search may legitimately return no results,
|
||||
# and (b) DuckDuckGo upstream blocks GHA IP ranges often
|
||||
# enough that requiring a tool_call marker would create
|
||||
# red-herring failures from infra rather than from Studio.
|
||||
# red-herring failures from infra rather than from Unsloth.
|
||||
try:
|
||||
# Best-effort and bounded: a single 180s attempt keeps a stall
|
||||
# from eating the job's timeout-minutes (it already WARNs, so a
|
||||
|
|
@ -834,7 +834,7 @@ jobs:
|
|||
print(f"[tools] WARN web_search probe failed (non-blocking): {exc}")
|
||||
|
||||
# ── 5. Thinking on / off ─────────────────────────────────────
|
||||
# Studio strips think blocks from message.content for tools-mode
|
||||
# Unsloth strips think blocks from message.content for tools-mode
|
||||
# responses, so we toggle plain chat (no enable_tools) and look
|
||||
# at the surfaced reasoning_content / message.thinking field.
|
||||
def thinking_call(enable):
|
||||
|
|
@ -848,7 +848,7 @@ jobs:
|
|||
})
|
||||
assert status == 200
|
||||
msg = data["choices"][0]["message"]
|
||||
# Studio surfaces thinking via reasoning_content (OpenAI
|
||||
# Unsloth surfaces thinking via reasoning_content (OpenAI
|
||||
# extension). Fall back to inline <think> markers for
|
||||
# robustness across template versions.
|
||||
raw = (msg.get("content") or "") + (msg.get("reasoning_content") or "")
|
||||
|
|
@ -868,7 +868,7 @@ jobs:
|
|||
print(f"[tools] PASS thinking on/off (on={len(on_text)} chars, off={len(off_text)} chars)")
|
||||
PY
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
kill "${STUDIO_PID}" 2>/dev/null || true
|
||||
|
|
@ -960,7 +960,7 @@ jobs:
|
|||
path: hf-cache
|
||||
key: ${{ runner.os }}-hf-${{ env.GGUF_REPO }}-${{ env.GGUF_VARIANT }}-${{ env.MMPROJ_FILE }}-v2
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
# Withheld on PR: this step runs checked-out PR code; public GGUF still downloads.
|
||||
|
|
@ -973,7 +973,7 @@ jobs:
|
|||
- name: Install OpenAI + Anthropic Python SDKs
|
||||
run: pip install 'openai>=1.50' 'anthropic>=0.40'
|
||||
|
||||
- name: Reset auth + boot Studio (API-only)
|
||||
- name: Reset auth + boot Unsloth (API-only)
|
||||
# See Job 2's comment: API-only mode keeps tool_policy=None so
|
||||
# response_format requests aren't routed through the agentic
|
||||
# tool loop.
|
||||
|
|
@ -1076,13 +1076,13 @@ jobs:
|
|||
# llama.cpp's HTTP server supports OpenAI-compatible JSON
|
||||
# mode: `response_format: {"type": "json_object"}` constrains
|
||||
# the model to emit syntactically-valid JSON. We use raw HTTP
|
||||
# rather than the OpenAI SDK so that the field shape Studio
|
||||
# rather than the OpenAI SDK so that the field shape Unsloth
|
||||
# forwards to llama-server is unambiguous (the SDK rewrites
|
||||
# response_format depending on which variant it recognises).
|
||||
# We deliberately do NOT pass a strict JSON schema -- on
|
||||
# small Gemma-4 quants the GBNF-from-schema path occasionally
|
||||
# produces empty output, and JSON mode is the surface we care
|
||||
# about exposing through Studio.
|
||||
# about exposing through Unsloth.
|
||||
status, data = post("/v1/chat/completions", {
|
||||
"model": "default",
|
||||
"messages": [
|
||||
|
|
@ -1112,7 +1112,7 @@ jobs:
|
|||
print(f"[json] PASS json_object -> {parsed}")
|
||||
|
||||
# ── 2. OpenAI image_url (data URI base64) ───────────────────
|
||||
# 64x64 solid-red PNG. stb_image (used by Studio's image
|
||||
# 64x64 solid-red PNG. stb_image (used by Unsloth's image
|
||||
# normaliser at routes/inference.py:3410) rejects 4x4 or
|
||||
# smaller PNGs as truncated, so we go up to 64x64 -- still
|
||||
# tiny in token cost. The assertion is loose: any non-empty
|
||||
|
|
@ -1148,9 +1148,9 @@ jobs:
|
|||
print("[image/openai] PASS image_url accepted, non-empty response")
|
||||
|
||||
# ── 3. Anthropic source/base64 image ────────────────────────
|
||||
# Two SDK quirks vs. Studio: base_url must NOT include /v1
|
||||
# Two SDK quirks vs. Unsloth: base_url must NOT include /v1
|
||||
# (the SDK appends it itself; otherwise /v1/v1/messages -> 405),
|
||||
# and Studio's auth is HTTPBearer-only so the SDK's default
|
||||
# and Unsloth's auth is HTTPBearer-only so the SDK's default
|
||||
# x-api-key header is ignored -- send Authorization: Bearer
|
||||
# via default_headers.
|
||||
anthropic = Anthropic(
|
||||
|
|
@ -1184,7 +1184,7 @@ jobs:
|
|||
print("[image/anthropic] PASS source/base64 accepted, non-empty response")
|
||||
PY
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
kill "${STUDIO_PID}" 2>/dev/null || true
|
||||
|
|
|
|||
|
|
@ -1,7 +1,7 @@
|
|||
# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved.
|
||||
#
|
||||
# Event-loop regression test for the Studio model-load orchestrator.
|
||||
# Event-loop regression test for the Unsloth model-load orchestrator.
|
||||
# Pins down issue #5642 (Win10 UI freeze on model load): the /load
|
||||
# route calls LlamaCppBackend.detect_audio_type synchronously, blocking
|
||||
# the FastAPI event loop on a chain of sync httpx.Client.post() probes.
|
||||
|
|
@ -14,7 +14,7 @@
|
|||
# danielhanchen/unsloth-staging-2 (Ubuntu / macOS / Windows all
|
||||
# green at PR time).
|
||||
|
||||
name: Studio load-orchestrator CI
|
||||
name: Unsloth load-orchestrator CI
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
|
|
|
|||
10
.github/workflows/studio-mac-api-smoke.yml
vendored
10
.github/workflows/studio-mac-api-smoke.yml
vendored
|
|
@ -33,7 +33,7 @@ permissions:
|
|||
|
||||
jobs:
|
||||
api-smoke:
|
||||
name: Studio API & Auth Tests
|
||||
name: Unsloth API & Auth Tests
|
||||
runs-on: macos-14
|
||||
timeout-minutes: 25
|
||||
env:
|
||||
|
|
@ -83,7 +83,7 @@ jobs:
|
|||
path: hf-cache
|
||||
key: ${{ runner.os }}-hf-${{ env.GGUF_REPO }}-${{ env.GGUF_VARIANT }}-v2
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
# Withheld on PR: this step runs checked-out PR code; public GGUF still downloads.
|
||||
|
|
@ -99,7 +99,7 @@ jobs:
|
|||
- name: Install pyjwt for the JWT-expiry forge test
|
||||
run: pip install 'pyjwt>=2.6'
|
||||
|
||||
- name: Reset auth + boot Studio (API-only)
|
||||
- name: Reset auth + boot Unsloth (API-only)
|
||||
run: |
|
||||
unsloth studio reset-password
|
||||
mkdir -p logs
|
||||
|
|
@ -129,13 +129,13 @@ jobs:
|
|||
echo "STUDIO_NEW_PW=$NEW" >> "$GITHUB_ENV"
|
||||
echo "STUDIO_NEW2_PW=$NEW2" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Run Studio API & Auth tests
|
||||
- name: Run Unsloth API & Auth tests
|
||||
env:
|
||||
BASE_URL: http://127.0.0.1:18895
|
||||
STUDIO_AUTH_DIR: /Users/runner/.unsloth/studio/auth
|
||||
run: python tests/studio/studio_api_smoke.py
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
kill "${STUDIO_PID}" 2>/dev/null || true
|
||||
|
|
|
|||
58
.github/workflows/studio-mac-inference-smoke.yml
vendored
58
.github/workflows/studio-mac-inference-smoke.yml
vendored
|
|
@ -1,7 +1,7 @@
|
|||
# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved.
|
||||
|
||||
# Three end-to-end smoke jobs that boot a freshly-installed Studio and
|
||||
# Three end-to-end smoke jobs that boot a freshly-installed Unsloth and
|
||||
# exercise the surfaces real users hit through the OpenAI / Anthropic
|
||||
# SDKs and curl. Each job picks the smallest model that exercises the
|
||||
# behaviour under test, primes a model cache via actions/cache, and
|
||||
|
|
@ -108,7 +108,7 @@ jobs:
|
|||
path: hf-cache
|
||||
key: ${{ runner.os }}-hf-${{ env.GGUF_REPO }}-${{ env.GGUF_VARIANT }}-v2
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
# Withheld on PR: this step runs checked-out PR code; public GGUF still downloads.
|
||||
|
|
@ -124,7 +124,7 @@ jobs:
|
|||
- name: Install OpenAI + Anthropic Python SDKs
|
||||
run: pip install 'openai>=1.50' 'anthropic>=0.40'
|
||||
|
||||
- name: Reset auth + boot Studio (API-only)
|
||||
- name: Reset auth + boot Unsloth (API-only)
|
||||
run: |
|
||||
unsloth studio reset-password
|
||||
mkdir -p logs
|
||||
|
|
@ -141,7 +141,7 @@ jobs:
|
|||
fi
|
||||
sleep 1
|
||||
done
|
||||
echo "Studio did not become healthy in 180s"
|
||||
echo "Unsloth did not become healthy in 180s"
|
||||
tail -200 logs/studio.log
|
||||
exit 1
|
||||
|
||||
|
|
@ -228,11 +228,11 @@ jobs:
|
|||
return replies
|
||||
|
||||
def run_anthropic():
|
||||
# Two SDK quirks vs. Studio:
|
||||
# Two SDK quirks vs. Unsloth:
|
||||
# 1. base_url must NOT include /v1 -- the SDK appends
|
||||
# /v1/messages itself; otherwise the request hits
|
||||
# /v1/v1/messages and 405s.
|
||||
# 2. The SDK sends `x-api-key` by default, but Studio's
|
||||
# 2. The SDK sends `x-api-key` by default, but Unsloth's
|
||||
# auth layer is HTTPBearer-only. Override via
|
||||
# default_headers so Authorization: Bearer ... is
|
||||
# sent instead.
|
||||
|
|
@ -283,7 +283,7 @@ jobs:
|
|||
print(f"[{label}] OK -- 4 turns, run1 == run2, history grounded")
|
||||
PY
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
kill "${STUDIO_PID}" 2>/dev/null || true
|
||||
|
|
@ -363,7 +363,7 @@ jobs:
|
|||
path: gguf-cache
|
||||
key: ${{ runner.os }}-gguf-${{ env.GGUF_REPO }}-${{ env.GGUF_FILE }}-v1
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
# Withheld on PR: this step runs checked-out PR code; public GGUF still downloads.
|
||||
|
|
@ -376,7 +376,7 @@ jobs:
|
|||
- name: Assert llama.cpp loads on this macOS
|
||||
run: bash .github/scripts/assert-llama-loads.sh
|
||||
|
||||
- name: Reset auth + boot Studio (API-only, default tool policy)
|
||||
- name: Reset auth + boot Unsloth (API-only, default tool policy)
|
||||
# We deliberately use the API-only mode rather than
|
||||
# `unsloth studio run` because the latter calls
|
||||
# `set_tool_policy(...)` with a resolved bool: on loopback the
|
||||
|
|
@ -478,7 +478,7 @@ jobs:
|
|||
call with enable_tools=true must use this helper.
|
||||
|
||||
A shared CI runner can stall the stream transport (the
|
||||
connection opening, or a mid-stream read) even when Studio
|
||||
connection opening, or a mid-stream read) even when Unsloth
|
||||
is healthy, so harden the read three ways: retry a stall
|
||||
once with a fresh request capped at 300s; return any text
|
||||
already streamed before a stall (a stall on the trailing
|
||||
|
|
@ -574,11 +574,11 @@ jobs:
|
|||
assert status == 200, f"tool call status {status}: {data}"
|
||||
choice = data["choices"][0]
|
||||
tool_calls = (choice.get("message") or {}).get("tool_calls") or []
|
||||
# Studio's contract: when tool_choice='required', llama.cpp's
|
||||
# Unsloth's contract: when tool_choice='required', llama.cpp's
|
||||
# grammar should force a tool_calls payload. On Mac that
|
||||
# contract is sometimes broken by the underlying quant; the
|
||||
# PASS path is "tool_calls present + correct schema", the
|
||||
# WARN path documents Studio still returned 200 with a
|
||||
# WARN path documents Unsloth still returned 200 with a
|
||||
# well-formed choices[] envelope.
|
||||
if tool_calls:
|
||||
tc = tool_calls[0]
|
||||
|
|
@ -660,7 +660,7 @@ jobs:
|
|||
print(f"[tools] WARN web_search probe failed (non-blocking): {exc}")
|
||||
|
||||
# ── 4. Thinking on / off ─────────────────────────────────────
|
||||
# Studio strips think blocks from message.content for tools-mode
|
||||
# Unsloth strips think blocks from message.content for tools-mode
|
||||
# responses, so we toggle plain chat (no enable_tools) and look
|
||||
# at the surfaced reasoning_content / message.thinking field.
|
||||
def thinking_call(enable):
|
||||
|
|
@ -678,7 +678,7 @@ jobs:
|
|||
}, timeout = 180)
|
||||
assert status == 200
|
||||
msg = data["choices"][0]["message"]
|
||||
# Studio surfaces thinking via reasoning_content (OpenAI
|
||||
# Unsloth surfaces thinking via reasoning_content (OpenAI
|
||||
# extension). Fall back to inline <think> markers for
|
||||
# robustness across template versions.
|
||||
raw = (msg.get("content") or "") + (msg.get("reasoning_content") or "")
|
||||
|
|
@ -704,7 +704,7 @@ jobs:
|
|||
print(f"[tools] PASS thinking on/off (on={len(on_text)} chars, off={len(off_text)} chars)")
|
||||
PY
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
kill "${STUDIO_PID}" 2>/dev/null || true
|
||||
|
|
@ -810,7 +810,7 @@ jobs:
|
|||
path: gguf-cache
|
||||
key: ${{ runner.os }}-gguf-${{ env.GGUF_REPO }}-${{ env.GGUF_FILE }}-${{ env.MMPROJ_FILE }}-v2
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
# Withheld on PR: this step runs checked-out PR code; public GGUF still downloads.
|
||||
|
|
@ -826,7 +826,7 @@ jobs:
|
|||
- name: Install OpenAI + Anthropic Python SDKs
|
||||
run: pip install 'openai>=1.50' 'anthropic>=0.40'
|
||||
|
||||
- name: Reset auth + boot Studio (API-only)
|
||||
- name: Reset auth + boot Unsloth (API-only)
|
||||
# See Job 2's comment: API-only mode keeps tool_policy=None so
|
||||
# response_format requests aren't routed through the agentic
|
||||
# tool loop.
|
||||
|
|
@ -929,13 +929,13 @@ jobs:
|
|||
# llama.cpp's HTTP server supports OpenAI-compatible JSON
|
||||
# mode: `response_format: {"type": "json_object"}` constrains
|
||||
# the model to emit syntactically-valid JSON. We use raw HTTP
|
||||
# rather than the OpenAI SDK so that the field shape Studio
|
||||
# rather than the OpenAI SDK so that the field shape Unsloth
|
||||
# forwards to llama-server is unambiguous (the SDK rewrites
|
||||
# response_format depending on which variant it recognises).
|
||||
# We deliberately do NOT pass a strict JSON schema -- on
|
||||
# small Gemma-4 quants the GBNF-from-schema path occasionally
|
||||
# produces empty output, and JSON mode is the surface we care
|
||||
# about exposing through Studio.
|
||||
# about exposing through Unsloth.
|
||||
status, data = post("/v1/chat/completions", {
|
||||
"model": "default",
|
||||
"messages": [
|
||||
|
|
@ -1007,7 +1007,7 @@ jobs:
|
|||
)
|
||||
|
||||
# ── 2. OpenAI image_url (data URI base64) ───────────────────
|
||||
# 64x64 solid-red PNG. stb_image (used by Studio's image
|
||||
# 64x64 solid-red PNG. stb_image (used by Unsloth's image
|
||||
# normaliser at routes/inference.py:3410) rejects 4x4 or
|
||||
# smaller PNGs as truncated, so we go up to 64x64 -- still
|
||||
# tiny in token cost. The assertion is loose: any non-empty
|
||||
|
|
@ -1023,11 +1023,11 @@ jobs:
|
|||
# The Mac prebuilt llama.cpp server has a known crash when
|
||||
# processing image inputs alongside the gemma-4-E2B mmproj
|
||||
# (server disconnects mid-completion). This is upstream
|
||||
# llama.cpp behaviour, not Studio. Wrap both SDK calls in
|
||||
# llama.cpp behaviour, not Unsloth. Wrap both SDK calls in
|
||||
# try/except so an upstream crash registers as a WARN rather
|
||||
# than failing the whole job. Studio's contract (OpenAI/
|
||||
# than failing the whole job. Unsloth's contract (OpenAI/
|
||||
# Anthropic image fields are accepted and forwarded) is
|
||||
# validated by the request body Studio constructs, not by
|
||||
# validated by the request body Unsloth constructs, not by
|
||||
# whether llama.cpp can decode it on Mac Metal.
|
||||
client = OpenAI(base_url = f"{BASE}/v1", api_key = KEY)
|
||||
try:
|
||||
|
|
@ -1053,14 +1053,14 @@ jobs:
|
|||
except Exception as exc:
|
||||
print(
|
||||
f"[image/openai] WARN image_url SDK call raised: {type(exc).__name__}: "
|
||||
f"{exc}. Likely upstream llama.cpp Mac+vision crash, NOT a Studio "
|
||||
f"regression. Studio successfully forwarded the request."
|
||||
f"{exc}. Likely upstream llama.cpp Mac+vision crash, NOT an Unsloth "
|
||||
f"regression. Unsloth successfully forwarded the request."
|
||||
)
|
||||
|
||||
# ── 3. Anthropic source/base64 image ────────────────────────
|
||||
# Two SDK quirks vs. Studio: base_url must NOT include /v1
|
||||
# Two SDK quirks vs. Unsloth: base_url must NOT include /v1
|
||||
# (the SDK appends it itself; otherwise /v1/v1/messages -> 405),
|
||||
# and Studio's auth is HTTPBearer-only so the SDK's default
|
||||
# and Unsloth's auth is HTTPBearer-only so the SDK's default
|
||||
# x-api-key header is ignored -- send Authorization: Bearer
|
||||
# via default_headers.
|
||||
anthropic = Anthropic(
|
||||
|
|
@ -1099,11 +1099,11 @@ jobs:
|
|||
print(
|
||||
f"[image/anthropic] WARN anthropic image SDK call raised: "
|
||||
f"{type(exc).__name__}: {exc}. Likely upstream llama.cpp Mac+vision "
|
||||
f"crash, NOT a Studio regression."
|
||||
f"crash, NOT an Unsloth regression."
|
||||
)
|
||||
PY
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
kill "${STUDIO_PID}" 2>/dev/null || true
|
||||
|
|
|
|||
|
|
@ -1,7 +1,7 @@
|
|||
# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved.
|
||||
|
||||
# Proves Studio's llama.cpp install loads on every supported macOS. The heavy
|
||||
# Proves Unsloth's llama.cpp install loads on every supported macOS. The heavy
|
||||
# app smokes stay single-OS; this matrix covers the OS-version dimension cheaply
|
||||
# (install.sh + binary-load assert). Regression guard for the macOS-version
|
||||
# selection in studio/install_llama_prebuilt.py.
|
||||
|
|
@ -60,7 +60,7 @@ jobs:
|
|||
with:
|
||||
python-version: '3.12'
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
# Withheld on PR: this step runs checked-out PR code; public GGUF still downloads.
|
||||
|
|
|
|||
18
.github/workflows/studio-mac-ui-smoke.yml
vendored
18
.github/workflows/studio-mac-ui-smoke.yml
vendored
|
|
@ -83,7 +83,7 @@ jobs:
|
|||
path: hf-cache
|
||||
key: ${{ runner.os }}-hf-${{ env.GGUF_REPO }}-${{ env.GGUF_VARIANT }}-v2
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
# Withheld on PR: this step runs checked-out PR code; public GGUF still downloads.
|
||||
|
|
@ -143,7 +143,7 @@ jobs:
|
|||
print(f"pipeTransport.js: patched JSON.parse calls in {path}")
|
||||
PY
|
||||
|
||||
- name: Reset auth + boot Studio
|
||||
- name: Reset auth + boot Unsloth
|
||||
run: |
|
||||
unsloth studio reset-password
|
||||
mkdir -p logs
|
||||
|
|
@ -188,7 +188,7 @@ jobs:
|
|||
# dies mid-test, (2) Chromium net::ERR_NO_BUFFER_SPACE when the
|
||||
# runner's kernel briefly runs out of socket buffers, and (3) a
|
||||
# goto 'interrupted by another navigation' when the SPA auth
|
||||
# guard redirects mid-navigation. The retry FULLY resets Studio
|
||||
# guard redirects mid-navigation. The retry FULLY resets Unsloth
|
||||
# (kill, reset-password, reboot, wait /api/health, re-export
|
||||
# bootstrap pw) before re-running the script. A real test failure
|
||||
# (assertion / timeout) does NOT match any pattern so it bypasses
|
||||
|
|
@ -209,7 +209,7 @@ jobs:
|
|||
|| grep -q "ERR_NO_BUFFER_SPACE" logs/playwright_attempt_${attempt}.log \
|
||||
|| grep -q "interrupted by another navigation" logs/playwright_attempt_${attempt}.log; } \
|
||||
&& [ "$attempt" -lt "$max_attempts" ]; then
|
||||
echo "::warning::Playwright flake on attempt ${attempt}; resetting Studio and retrying..."
|
||||
echo "::warning::Playwright flake on attempt ${attempt}; resetting Unsloth and retrying..."
|
||||
kill "${STUDIO_PID}" 2>/dev/null || true
|
||||
sleep 2
|
||||
unsloth studio reset-password
|
||||
|
|
@ -238,13 +238,13 @@ jobs:
|
|||
exit "$rc"
|
||||
done
|
||||
|
||||
- name: Stop Studio (chat-ui ends with Shutdown click; this is belt-and-suspenders)
|
||||
- name: Stop Unsloth (chat-ui ends with Shutdown click; this is belt-and-suspenders)
|
||||
if: always()
|
||||
run: |
|
||||
kill "${STUDIO_PID}" 2>/dev/null || true
|
||||
sleep 2
|
||||
|
||||
- name: Reset auth + boot Studio for extra UI tests (port 18897)
|
||||
- name: Reset auth + boot Unsloth for extra UI tests (port 18897)
|
||||
run: |
|
||||
unsloth studio reset-password
|
||||
mkdir -p logs
|
||||
|
|
@ -271,7 +271,7 @@ jobs:
|
|||
echo "STUDIO_EXTRA_OLD_PW=$OLD" >> "$GITHUB_ENV"
|
||||
echo "STUDIO_EXTRA_NEW_PW=$NEW" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Drive Compare/Recipes/Export/Studio/Settings with Playwright
|
||||
- name: Drive Compare/Recipes/Export/Unsloth/Settings with Playwright
|
||||
env:
|
||||
BASE_URL: http://127.0.0.1:18897
|
||||
STUDIO_OLD_PW: ${{ env.STUDIO_EXTRA_OLD_PW }}
|
||||
|
|
@ -300,7 +300,7 @@ jobs:
|
|||
|| grep -q "ERR_NO_BUFFER_SPACE" logs/playwright_extra_attempt_${attempt}.log \
|
||||
|| grep -q "interrupted by another navigation" logs/playwright_extra_attempt_${attempt}.log; } \
|
||||
&& [ "$attempt" -lt "$max_attempts" ]; then
|
||||
echo "::warning::Playwright flake on attempt ${attempt}; resetting Studio and retrying..."
|
||||
echo "::warning::Playwright flake on attempt ${attempt}; resetting Unsloth and retrying..."
|
||||
kill "${STUDIO_EXTRA_PID}" 2>/dev/null || true
|
||||
sleep 2
|
||||
unsloth studio reset-password
|
||||
|
|
@ -327,7 +327,7 @@ jobs:
|
|||
exit "$rc"
|
||||
done
|
||||
|
||||
- name: Stop second Studio
|
||||
- name: Stop second Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
kill "${STUDIO_EXTRA_PID}" 2>/dev/null || true
|
||||
|
|
|
|||
16
.github/workflows/studio-mac-update-smoke.yml
vendored
16
.github/workflows/studio-mac-update-smoke.yml
vendored
|
|
@ -4,15 +4,15 @@
|
|||
# Mac counterpart to studio-update-smoke.yml. Verifies that on a real
|
||||
# Apple Silicon (macos-14, M1) runner:
|
||||
#
|
||||
# 1. install.sh --local --no-torch installs Studio AND auto-fetches
|
||||
# 1. install.sh --local --no-torch installs Unsloth AND auto-fetches
|
||||
# the prebuilt llama.cpp Mac binary (llama-bNNNN-bin-macos-arm64
|
||||
# from ggml-org/llama.cpp). Hitting the source-build fallback is
|
||||
# treated as an Unsloth bug -- Studio must always pick the
|
||||
# treated as an Unsloth bug -- Unsloth must always pick the
|
||||
# prebuilt on Mac.
|
||||
# 2. unsloth studio update --local is idempotent. Two consecutive
|
||||
# runs both report "prebuilt up to date and validated", no
|
||||
# source-build fallback.
|
||||
# 3. The installed Studio still boots and /api/health returns
|
||||
# 3. The installed Unsloth still boots and /api/health returns
|
||||
# healthy after the update path.
|
||||
|
||||
name: Mac Studio Update CI
|
||||
|
|
@ -42,7 +42,7 @@ permissions:
|
|||
|
||||
jobs:
|
||||
update-idempotency:
|
||||
name: Studio Updating Tests
|
||||
name: Unsloth Updating Tests
|
||||
runs-on: macos-14
|
||||
timeout-minutes: 30
|
||||
steps:
|
||||
|
|
@ -59,7 +59,7 @@ jobs:
|
|||
python-version: '3.12'
|
||||
cache: 'pip'
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
# Withheld on PR: this step runs checked-out PR code; public GGUF still downloads.
|
||||
|
|
@ -106,7 +106,7 @@ jobs:
|
|||
grep -qE "prebuilt up to date and validated|prebuilt installed and validated" logs/update2.log
|
||||
echo "second update was clean"
|
||||
|
||||
- name: Boot Studio briefly to confirm the install is still usable
|
||||
- name: Boot Unsloth briefly to confirm the install is still usable
|
||||
run: |
|
||||
mkdir -p logs
|
||||
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p 18891 \
|
||||
|
|
@ -123,13 +123,13 @@ jobs:
|
|||
sleep 1
|
||||
done
|
||||
if [ -z "$HEALTHY" ]; then
|
||||
echo "Studio failed to come up after \`update\`"
|
||||
echo "Unsloth failed to come up after \`update\`"
|
||||
tail -200 logs/studio.log
|
||||
kill "$PID" 2>/dev/null || true
|
||||
exit 1
|
||||
fi
|
||||
kill "$PID" 2>/dev/null || true
|
||||
echo "post-update Studio /api/health OK"
|
||||
echo "post-update Unsloth /api/health OK"
|
||||
|
||||
- name: Uninstall and verify clean
|
||||
# Round-trip through scripts/uninstall.sh on real macOS. As a side
|
||||
|
|
|
|||
2
.github/workflows/studio-tauri-smoke.yml
vendored
2
.github/workflows/studio-tauri-smoke.yml
vendored
|
|
@ -12,7 +12,7 @@
|
|||
# stay in release-desktop.yml (manual `workflow_dispatch`) because they need
|
||||
# code-signing secrets and ~30 min of runner time each.
|
||||
|
||||
name: Studio Tauri CI
|
||||
name: Unsloth Tauri CI
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
|
|
|
|||
34
.github/workflows/studio-ui-smoke.yml
vendored
34
.github/workflows/studio-ui-smoke.yml
vendored
|
|
@ -1,8 +1,8 @@
|
|||
# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved.
|
||||
|
||||
# End-to-end Studio chat UI smoke via Playwright + Chromium against a
|
||||
# headless Linux runner. Boots Studio with the smallest GGUF
|
||||
# End-to-end Unsloth chat UI smoke via Playwright + Chromium against a
|
||||
# headless Linux runner. Boots Unsloth with the smallest GGUF
|
||||
# (gemma-3-270m-it UD-Q4_K_XL, ~254 MiB), drives the actual frontend
|
||||
# bundle, and asserts the full bootstrap-password / change-password /
|
||||
# send-message / persist-on-reload journey works end to end.
|
||||
|
|
@ -14,7 +14,7 @@
|
|||
# frontend-only CI happily pass while the actual user-visible UI is
|
||||
# broken (cf. the 2026.5.1 chat-history release).
|
||||
|
||||
name: Studio UI CI
|
||||
name: Unsloth UI CI
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
|
|
@ -97,7 +97,7 @@ jobs:
|
|||
path: hf-cache
|
||||
key: ${{ runner.os }}-hf-${{ env.GGUF_REPO }}-${{ env.GGUF_VARIANT }}-v2
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
# Withheld on PR: this step runs checked-out PR code; public GGUF still downloads.
|
||||
|
|
@ -115,7 +115,7 @@ jobs:
|
|||
# warm runner.
|
||||
python -m playwright install --with-deps chromium
|
||||
|
||||
- name: Reset auth + boot Studio
|
||||
- name: Reset auth + boot Unsloth
|
||||
run: |
|
||||
unsloth studio reset-password
|
||||
mkdir -p logs
|
||||
|
|
@ -147,7 +147,7 @@ jobs:
|
|||
# NEW + NEW2 are generated freshly per CI run via secrets.token_urlsafe
|
||||
# rather than hardcoded. If a workflow gets compromised, the
|
||||
# attacker can't replay a known-good rotated password against
|
||||
# any future / parallel Studio install -- the rotated value
|
||||
# any future / parallel Unsloth install -- the rotated value
|
||||
# only ever exists for the lifetime of this single job, masked
|
||||
# in the log via ::add-mask::.
|
||||
run: |
|
||||
|
|
@ -165,18 +165,18 @@ jobs:
|
|||
env:
|
||||
BASE_URL: http://127.0.0.1:18892
|
||||
# The test file lives in the repo so it can be run locally
|
||||
# against a freshly-installed Studio (BASE_URL=...; STUDIO_OLD_PW=
|
||||
# against a freshly-installed Unsloth (BASE_URL=...; STUDIO_OLD_PW=
|
||||
# $(cat ~/.unsloth/studio/auth/.bootstrap_password); python ...).
|
||||
PW_ART_DIR: logs/playwright
|
||||
# Strict mode: in CI a missing button / nav / dialog must
|
||||
# FAIL the test. Locally the test still runs against partial
|
||||
# Studio installs without STUDIO_UI_STRICT.
|
||||
# Unsloth installs without STUDIO_UI_STRICT.
|
||||
STUDIO_UI_STRICT: '1'
|
||||
run: |
|
||||
mkdir -p logs/playwright
|
||||
python tests/studio/playwright_chat_ui.py
|
||||
|
||||
- name: Stop Studio (chat-ui ends with Shutdown click; this is belt-and-suspenders)
|
||||
- name: Stop Unsloth (chat-ui ends with Shutdown click; this is belt-and-suspenders)
|
||||
if: always()
|
||||
run: |
|
||||
kill "${STUDIO_PID}" 2>/dev/null || true
|
||||
|
|
@ -184,10 +184,10 @@ jobs:
|
|||
|
||||
# The chat UI test ends by clicking the Shutdown menuitem, which
|
||||
# leaves the server dead. The extra UI test (Compare / Recipes /
|
||||
# Export / Studio / Settings) needs a fresh Studio, so we boot a
|
||||
# Export / Unsloth / Settings) needs a fresh Unsloth, so we boot a
|
||||
# second one on a different port. Boot is fast (~3-5s on the
|
||||
# warm install we already did) so this adds little wall time.
|
||||
- name: Reset auth + boot Studio for extra UI tests (port 18894)
|
||||
- name: Reset auth + boot Unsloth for extra UI tests (port 18894)
|
||||
run: |
|
||||
unsloth studio reset-password
|
||||
mkdir -p logs
|
||||
|
|
@ -214,7 +214,7 @@ jobs:
|
|||
echo "STUDIO_EXTRA_OLD_PW=$OLD" >> "$GITHUB_ENV"
|
||||
echo "STUDIO_EXTRA_NEW_PW=$NEW" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Drive Compare/Recipes/Export/Studio/Settings with Playwright
|
||||
- name: Drive Compare/Recipes/Export/Unsloth/Settings with Playwright
|
||||
env:
|
||||
BASE_URL: http://127.0.0.1:18894
|
||||
STUDIO_OLD_PW: ${{ env.STUDIO_EXTRA_OLD_PW }}
|
||||
|
|
@ -227,16 +227,16 @@ jobs:
|
|||
mkdir -p logs/playwright_extra
|
||||
python tests/studio/playwright_extra_ui.py
|
||||
|
||||
- name: Stop second Studio
|
||||
- name: Stop second Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
kill "${STUDIO_EXTRA_PID}" 2>/dev/null || true
|
||||
sleep 2
|
||||
|
||||
# IME + multilingual paste regression (issue #5318 / PR #5327).
|
||||
# Third Studio on its own port so a hang here cannot poison the
|
||||
# Third Unsloth on its own port so a hang here cannot poison the
|
||||
# earlier UI tests. No GGUF -- the bug surface is the composer.
|
||||
- name: Reset auth + boot Studio for IME / i18n tests (port 18896)
|
||||
- name: Reset auth + boot Unsloth for IME / i18n tests (port 18896)
|
||||
run: |
|
||||
unsloth studio reset-password
|
||||
mkdir -p logs
|
||||
|
|
@ -256,7 +256,7 @@ jobs:
|
|||
|
||||
- name: Pass bootstrap pw for IME / i18n test
|
||||
# IME smoke does the change-password against the bootstrap that
|
||||
# Studio's frontend injects into the page, so it only needs the
|
||||
# Unsloth's frontend injects into the page, so it only needs the
|
||||
# NEW password.
|
||||
run: |
|
||||
NEW="CIIme-$(python -c 'import secrets; print(secrets.token_urlsafe(16))')"
|
||||
|
|
@ -273,7 +273,7 @@ jobs:
|
|||
mkdir -p logs/playwright_ime
|
||||
python tests/studio/playwright_chat_ime_i18n.py
|
||||
|
||||
- name: Stop third Studio
|
||||
- name: Stop third Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
kill "${STUDIO_IME_PID}" 2>/dev/null || true
|
||||
|
|
|
|||
12
.github/workflows/studio-update-smoke.yml
vendored
12
.github/workflows/studio-update-smoke.yml
vendored
|
|
@ -9,7 +9,7 @@
|
|||
# This catches regressions in setup.sh's update path that the existing
|
||||
# GGUF / wheel jobs would miss because they only invoke install.sh once.
|
||||
|
||||
name: Studio Update CI
|
||||
name: Unsloth Update CI
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
|
|
@ -36,7 +36,7 @@ permissions:
|
|||
|
||||
jobs:
|
||||
update-idempotency:
|
||||
name: Studio Updating Tests
|
||||
name: Unsloth Updating Tests
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 15
|
||||
steps:
|
||||
|
|
@ -63,7 +63,7 @@ jobs:
|
|||
# post-step then fatal-errors with "Cache folder path is
|
||||
# retrieved for pip but doesn't exist on disk".
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
# Pass the workflow token so the llama.cpp prebuilt installer's
|
||||
# GitHub-API call to list releases isn't rate-limited (60/hr
|
||||
# unauthenticated). Without this, three consecutive install +
|
||||
|
|
@ -122,7 +122,7 @@ jobs:
|
|||
grep -qE "prebuilt up to date and validated|prebuilt installed and validated" logs/update2.log
|
||||
echo "second update was clean"
|
||||
|
||||
- name: Boot Studio briefly to confirm the install is still usable
|
||||
- name: Boot Unsloth briefly to confirm the install is still usable
|
||||
# If `update --local` accidentally broke the venv or wiped the
|
||||
# llama-server binary, the server would fail to start here.
|
||||
run: |
|
||||
|
|
@ -138,13 +138,13 @@ jobs:
|
|||
sleep 1
|
||||
done
|
||||
if ! jq -e '.status == "healthy"' /tmp/health.json 2>/dev/null; then
|
||||
echo "Studio failed to come up after `update`"
|
||||
echo "Unsloth failed to come up after `update`"
|
||||
tail -200 logs/studio.log
|
||||
kill "$PID" 2>/dev/null || true
|
||||
exit 1
|
||||
fi
|
||||
kill "$PID" 2>/dev/null || true
|
||||
echo "post-update Studio /api/health OK"
|
||||
echo "post-update Unsloth /api/health OK"
|
||||
|
||||
- name: Uninstall and verify clean
|
||||
# Round-trip the installer through scripts/uninstall.sh: confirms the
|
||||
|
|
|
|||
16
.github/workflows/studio-windows-api-smoke.yml
vendored
16
.github/workflows/studio-windows-api-smoke.yml
vendored
|
|
@ -9,7 +9,7 @@
|
|||
# (Section 6) is Linux-only and short-circuits on non-POSIX; the rest
|
||||
# is platform-portable.
|
||||
|
||||
name: Windows Studio API CI
|
||||
name: Windows Unsloth API CI
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
|
|
@ -34,7 +34,7 @@ permissions:
|
|||
|
||||
jobs:
|
||||
api-smoke:
|
||||
name: Studio API & Auth Tests
|
||||
name: Unsloth API & Auth Tests
|
||||
runs-on: windows-latest
|
||||
timeout-minutes: 30
|
||||
defaults:
|
||||
|
|
@ -105,7 +105,7 @@ jobs:
|
|||
# studio-windows-update-smoke.yml for the full rationale --
|
||||
# creating an empty studio/frontend/dist trips setup.ps1's
|
||||
# mtime-based staleness check into "frontend up to date, skip
|
||||
# rebuild" and Studio boots with an empty dist directory.
|
||||
# rebuild" and Unsloth boots with an empty dist directory.
|
||||
# Add-MpPreference accepts paths that do not yet exist.
|
||||
foreach ($p in @(
|
||||
"$env:USERPROFILE\.unsloth",
|
||||
|
|
@ -121,7 +121,7 @@ jobs:
|
|||
}
|
||||
}
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
shell: pwsh
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
|
|
@ -161,7 +161,7 @@ jobs:
|
|||
echo "install.ps1 installed the Windows prebuilt llama.cpp:"
|
||||
cat "$INFO"
|
||||
|
||||
- name: Add Studio shim to GITHUB_PATH
|
||||
- name: Add Unsloth shim to GITHUB_PATH
|
||||
# install.ps1's User-PATH update doesn't propagate to a
|
||||
# running Git Bash session; export the shim dir so the
|
||||
# next `unsloth ...` invocation finds it.
|
||||
|
|
@ -177,7 +177,7 @@ jobs:
|
|||
- name: Install pyjwt for the JWT-expiry forge test
|
||||
run: python -m pip install 'pyjwt>=2.6'
|
||||
|
||||
- name: Reset auth + boot Studio (API-only)
|
||||
- name: Reset auth + boot Unsloth (API-only)
|
||||
run: |
|
||||
unsloth studio reset-password
|
||||
mkdir -p logs
|
||||
|
|
@ -207,7 +207,7 @@ jobs:
|
|||
echo "STUDIO_NEW_PW=$NEW" >> "$GITHUB_ENV"
|
||||
echo "STUDIO_NEW2_PW=$NEW2" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Run Studio API & Auth tests
|
||||
- name: Run Unsloth API & Auth tests
|
||||
# Do NOT pin STUDIO_AUTH_DIR here. The Mac/Linux mirrors
|
||||
# hardcode runner-specific paths (/Users/runner/...,
|
||||
# /home/runner/...), but on Windows the path is
|
||||
|
|
@ -219,7 +219,7 @@ jobs:
|
|||
BASE_URL: http://127.0.0.1:18895
|
||||
run: python tests/studio/studio_api_smoke.py
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
kill "${STUDIO_PID}" 2>/dev/null || true
|
||||
|
|
|
|||
|
|
@ -1,7 +1,7 @@
|
|||
# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved.
|
||||
|
||||
# Three end-to-end smoke jobs that boot a freshly-installed Studio and
|
||||
# Three end-to-end smoke jobs that boot a freshly-installed Unsloth and
|
||||
# exercise the surfaces real users hit through the OpenAI / Anthropic
|
||||
# SDKs and curl, on the FREE windows-latest runner. Each job picks the
|
||||
# smallest model that exercises the behaviour under test, primes
|
||||
|
|
@ -16,7 +16,7 @@
|
|||
# Qwen3-VL-2B-Instruct UD-IQ2_XXS + mmproj-F16 (~1.4 GiB total).
|
||||
# Within the 14 GB windows-latest SSD budget.
|
||||
|
||||
name: Windows Studio GGUF CI
|
||||
name: Windows Unsloth GGUF CI
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
|
|
@ -57,7 +57,7 @@ jobs:
|
|||
STUDIO_PORT: '18888'
|
||||
HF_HOME: ${{ github.workspace }}/hf-cache
|
||||
# Force UTF-8 for stdio (Windows defaults to cp1252; hf
|
||||
# download / Studio CLI print "✓" checkmarks and crash
|
||||
# download / Unsloth CLI print "✓" checkmarks and crash
|
||||
# otherwise).
|
||||
PYTHONIOENCODING: utf-8
|
||||
PYTHONUTF8: '1'
|
||||
|
|
@ -160,7 +160,7 @@ jobs:
|
|||
# studio-windows-update-smoke.yml for the full rationale --
|
||||
# creating an empty studio/frontend/dist trips setup.ps1's
|
||||
# mtime-based staleness check into "frontend up to date, skip
|
||||
# rebuild" and Studio boots with an empty dist directory.
|
||||
# rebuild" and Unsloth boots with an empty dist directory.
|
||||
# Add-MpPreference accepts paths that do not yet exist.
|
||||
foreach ($p in @(
|
||||
"$env:USERPROFILE\.unsloth",
|
||||
|
|
@ -176,7 +176,7 @@ jobs:
|
|||
}
|
||||
}
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
shell: pwsh
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
|
|
@ -214,7 +214,7 @@ jobs:
|
|||
echo "install.ps1 installed the Windows prebuilt llama.cpp:"
|
||||
cat "$INFO"
|
||||
|
||||
- name: Add Studio shim to GITHUB_PATH
|
||||
- name: Add Unsloth shim to GITHUB_PATH
|
||||
run: |
|
||||
SHIM_DIR=~/.unsloth/studio/bin
|
||||
if [ ! -f "$SHIM_DIR/unsloth.exe" ]; then
|
||||
|
|
@ -227,7 +227,7 @@ jobs:
|
|||
- name: Install OpenAI + Anthropic Python SDKs
|
||||
run: python -m pip install 'openai>=1.50' 'anthropic>=0.40'
|
||||
|
||||
- name: Reset auth + boot Studio (API-only)
|
||||
- name: Reset auth + boot Unsloth (API-only)
|
||||
run: |
|
||||
unsloth studio reset-password
|
||||
mkdir -p logs
|
||||
|
|
@ -244,7 +244,7 @@ jobs:
|
|||
fi
|
||||
sleep 1
|
||||
done
|
||||
echo "Studio did not become healthy in 180s"
|
||||
echo "Unsloth did not become healthy in 180s"
|
||||
tail -200 logs/studio.log
|
||||
exit 1
|
||||
|
||||
|
|
@ -281,7 +281,7 @@ jobs:
|
|||
# Retry the load step a few times so a transient TCP RST during
|
||||
# llama-server warm-up (Windows runner image churn,
|
||||
# windows-latest -> windows-2025-vs2026 rollout) doesn't fail
|
||||
# the whole job. The Studio backend's _wait_for_health now
|
||||
# the whole job. The Unsloth backend's _wait_for_health now
|
||||
# catches httpx.ReadError too; this retry layer covers the
|
||||
# cases the backend can't recover from on its own.
|
||||
LOAD_OK=0
|
||||
|
|
@ -382,15 +382,15 @@ jobs:
|
|||
print(f"[{label}] OK -- 4 turns, run1 == run2, history grounded")
|
||||
PY
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
# Run as cmd so we are not running through the Git Bash shell;
|
||||
# Git Bash on windows-latest has been observed to exit 143
|
||||
# (SIGTERM) from any inline kill/sleep block, masking a green
|
||||
# test run. The runner reclaims the Studio child process at
|
||||
# test run. The runner reclaims the Unsloth child process at
|
||||
# job end either way, so just emit a marker and exit 0.
|
||||
shell: cmd
|
||||
run: echo Stop Studio (no-op; runner reclaims STUDIO_PID=%STUDIO_PID% at job end)
|
||||
run: echo Stop Unsloth (no-op; runner reclaims STUDIO_PID=%STUDIO_PID% at job end)
|
||||
|
||||
- name: Collect llama-server logs
|
||||
if: always()
|
||||
|
|
@ -398,10 +398,10 @@ jobs:
|
|||
# copy must not fail an otherwise-green job.
|
||||
continue-on-error: true
|
||||
shell: bash
|
||||
# Copy llama-server's own stdout/stderr (teed by Studio under
|
||||
# Copy llama-server's own stdout/stderr (teed by Unsloth under
|
||||
# ~/.unsloth/studio/logs/llama-server/) into the workspace so
|
||||
# upload-artifact can pick it up. Crucial for diagnosing a
|
||||
# subprocess crash where Studio's traceback only shows the
|
||||
# subprocess crash where Unsloth's traceback only shows the
|
||||
# symptom (httpx ReadError) but not the cause.
|
||||
run: |
|
||||
mkdir -p logs/llama-server
|
||||
|
|
@ -439,14 +439,14 @@ jobs:
|
|||
# (211 s on first run; subsequent runs hit the cache, but the
|
||||
# one-time cost recurs every time the cache key bumps). Use
|
||||
# main's `--local-dir gguf-cache` pattern: cache the flat .gguf
|
||||
# only, pass an absolute path to Studio's /api/inference/load.
|
||||
# only, pass an absolute path to Unsloth's /api/inference/load.
|
||||
# The OpenAI/Anth and JSON+images jobs still cover the
|
||||
# gguf_variant resolution path.
|
||||
GGUF_REPO: unsloth/Qwen3.5-2B-GGUF
|
||||
GGUF_FILE: Qwen3.5-2B-UD-Q4_K_XL.gguf
|
||||
STUDIO_PORT: '18898'
|
||||
# Force UTF-8 for stdio (Windows defaults to cp1252; hf
|
||||
# download / Studio CLI print "✓" checkmarks and crash
|
||||
# download / Unsloth CLI print "✓" checkmarks and crash
|
||||
# otherwise).
|
||||
PYTHONIOENCODING: utf-8
|
||||
PYTHONUTF8: '1'
|
||||
|
|
@ -507,7 +507,7 @@ jobs:
|
|||
# studio-windows-update-smoke.yml for the full rationale --
|
||||
# creating an empty studio/frontend/dist trips setup.ps1's
|
||||
# mtime-based staleness check into "frontend up to date, skip
|
||||
# rebuild" and Studio boots with an empty dist directory.
|
||||
# rebuild" and Unsloth boots with an empty dist directory.
|
||||
# Add-MpPreference accepts paths that do not yet exist.
|
||||
foreach ($p in @(
|
||||
"$env:USERPROFILE\.unsloth",
|
||||
|
|
@ -523,7 +523,7 @@ jobs:
|
|||
}
|
||||
}
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
shell: pwsh
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
|
|
@ -561,7 +561,7 @@ jobs:
|
|||
echo "install.ps1 installed the Windows prebuilt llama.cpp:"
|
||||
cat "$INFO"
|
||||
|
||||
- name: Add Studio shim to GITHUB_PATH
|
||||
- name: Add Unsloth shim to GITHUB_PATH
|
||||
run: |
|
||||
SHIM_DIR=~/.unsloth/studio/bin
|
||||
if [ ! -f "$SHIM_DIR/unsloth.exe" ]; then
|
||||
|
|
@ -571,7 +571,7 @@ jobs:
|
|||
fi
|
||||
cygpath -w "$SHIM_DIR" >> "$GITHUB_PATH"
|
||||
|
||||
- name: Reset auth + boot Studio (API-only, default tool policy)
|
||||
- name: Reset auth + boot Unsloth (API-only, default tool policy)
|
||||
run: |
|
||||
unsloth studio reset-password
|
||||
mkdir -p logs
|
||||
|
|
@ -607,7 +607,7 @@ jobs:
|
|||
# raw string, but we cannot embed `\a` etc. in JSON without
|
||||
# JSON-string-escaping every backslash. Replace `\` with `/`
|
||||
# via bash parameter expansion -- pathlib.Path on Windows
|
||||
# accepts forward slashes natively, so Studio's loader sees
|
||||
# accepts forward slashes natively, so Unsloth's loader sees
|
||||
# a normal path.
|
||||
GGUF_PATH="${GITHUB_WORKSPACE//\\//}/gguf-cache/${GGUF_FILE}"
|
||||
ls -lh "$GGUF_PATH"
|
||||
|
|
@ -680,7 +680,7 @@ jobs:
|
|||
def post_sse(path, body, *, timeout = 600, retries = 1, soft = False):
|
||||
# The server-side agentic loop always answers over SSE. A
|
||||
# shared CI runner can stall the stream transport (the
|
||||
# connection opening, or a mid-stream read) even when Studio
|
||||
# connection opening, or a mid-stream read) even when Unsloth
|
||||
# is healthy, so harden the read three ways:
|
||||
# * retry a transport stall once with a fresh request,
|
||||
# capped at 300s (a healthy server answers a retry
|
||||
|
|
@ -882,15 +882,15 @@ jobs:
|
|||
print(f"[tools] PASS thinking on/off (on={len(on_text)} chars, off={len(off_text)} chars)")
|
||||
PY
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
# Run as cmd so we are not running through the Git Bash shell;
|
||||
# Git Bash on windows-latest has been observed to exit 143
|
||||
# (SIGTERM) from any inline kill/sleep block, masking a green
|
||||
# test run. The runner reclaims the Studio child process at
|
||||
# test run. The runner reclaims the Unsloth child process at
|
||||
# job end either way, so just emit a marker and exit 0.
|
||||
shell: cmd
|
||||
run: echo Stop Studio (no-op; runner reclaims STUDIO_PID=%STUDIO_PID% at job end)
|
||||
run: echo Stop Unsloth (no-op; runner reclaims STUDIO_PID=%STUDIO_PID% at job end)
|
||||
|
||||
- name: Collect llama-server logs
|
||||
if: always()
|
||||
|
|
@ -898,10 +898,10 @@ jobs:
|
|||
# copy must not fail an otherwise-green job.
|
||||
continue-on-error: true
|
||||
shell: bash
|
||||
# Copy llama-server's own stdout/stderr (teed by Studio under
|
||||
# Copy llama-server's own stdout/stderr (teed by Unsloth under
|
||||
# ~/.unsloth/studio/logs/llama-server/) into the workspace so
|
||||
# upload-artifact can pick it up. Crucial for diagnosing a
|
||||
# subprocess crash where Studio's traceback only shows the
|
||||
# subprocess crash where Unsloth's traceback only shows the
|
||||
# symptom (httpx ReadError) but not the cause.
|
||||
run: |
|
||||
mkdir -p logs/llama-server
|
||||
|
|
@ -939,7 +939,7 @@ jobs:
|
|||
STUDIO_PORT: '18899'
|
||||
HF_HOME: ${{ github.workspace }}/hf-cache
|
||||
# Force UTF-8 for stdio (Windows defaults to cp1252; hf
|
||||
# download / Studio CLI print "✓" checkmarks and crash
|
||||
# download / Unsloth CLI print "✓" checkmarks and crash
|
||||
# otherwise).
|
||||
PYTHONIOENCODING: utf-8
|
||||
PYTHONUTF8: '1'
|
||||
|
|
@ -1005,7 +1005,7 @@ jobs:
|
|||
# studio-windows-update-smoke.yml for the full rationale --
|
||||
# creating an empty studio/frontend/dist trips setup.ps1's
|
||||
# mtime-based staleness check into "frontend up to date, skip
|
||||
# rebuild" and Studio boots with an empty dist directory.
|
||||
# rebuild" and Unsloth boots with an empty dist directory.
|
||||
# Add-MpPreference accepts paths that do not yet exist.
|
||||
foreach ($p in @(
|
||||
"$env:USERPROFILE\.unsloth",
|
||||
|
|
@ -1021,7 +1021,7 @@ jobs:
|
|||
}
|
||||
}
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
shell: pwsh
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
|
|
@ -1059,7 +1059,7 @@ jobs:
|
|||
echo "install.ps1 installed the Windows prebuilt llama.cpp:"
|
||||
cat "$INFO"
|
||||
|
||||
- name: Add Studio shim to GITHUB_PATH
|
||||
- name: Add Unsloth shim to GITHUB_PATH
|
||||
run: |
|
||||
SHIM_DIR=~/.unsloth/studio/bin
|
||||
if [ ! -f "$SHIM_DIR/unsloth.exe" ]; then
|
||||
|
|
@ -1072,7 +1072,7 @@ jobs:
|
|||
- name: Install OpenAI + Anthropic Python SDKs
|
||||
run: python -m pip install 'openai>=1.50' 'anthropic>=0.40'
|
||||
|
||||
- name: Reset auth + boot Studio (API-only)
|
||||
- name: Reset auth + boot Unsloth (API-only)
|
||||
run: |
|
||||
unsloth studio reset-password
|
||||
mkdir -p logs
|
||||
|
|
@ -1262,7 +1262,7 @@ jobs:
|
|||
except Exception as exc:
|
||||
print(
|
||||
f"[image/openai] WARN image_url SDK call raised: {type(exc).__name__}: "
|
||||
f"{exc}. Studio successfully forwarded the request; failure here is "
|
||||
f"{exc}. Unsloth successfully forwarded the request; failure here is "
|
||||
f"upstream llama.cpp vision behaviour."
|
||||
)
|
||||
|
||||
|
|
@ -1303,19 +1303,19 @@ jobs:
|
|||
print(
|
||||
f"[image/anthropic] WARN anthropic image SDK call raised: "
|
||||
f"{type(exc).__name__}: {exc}. Likely upstream llama.cpp vision "
|
||||
f"behaviour, NOT a Studio regression."
|
||||
f"behaviour, NOT an Unsloth regression."
|
||||
)
|
||||
PY
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
# Run as cmd so we are not running through the Git Bash shell;
|
||||
# Git Bash on windows-latest has been observed to exit 143
|
||||
# (SIGTERM) from any inline kill/sleep block, masking a green
|
||||
# test run. The runner reclaims the Studio child process at
|
||||
# test run. The runner reclaims the Unsloth child process at
|
||||
# job end either way, so just emit a marker and exit 0.
|
||||
shell: cmd
|
||||
run: echo Stop Studio (no-op; runner reclaims STUDIO_PID=%STUDIO_PID% at job end)
|
||||
run: echo Stop Unsloth (no-op; runner reclaims STUDIO_PID=%STUDIO_PID% at job end)
|
||||
|
||||
- name: Collect llama-server logs
|
||||
if: always()
|
||||
|
|
@ -1323,10 +1323,10 @@ jobs:
|
|||
# copy must not fail an otherwise-green job.
|
||||
continue-on-error: true
|
||||
shell: bash
|
||||
# Copy llama-server's own stdout/stderr (teed by Studio under
|
||||
# Copy llama-server's own stdout/stderr (teed by Unsloth under
|
||||
# ~/.unsloth/studio/logs/llama-server/) into the workspace so
|
||||
# upload-artifact can pick it up. Crucial for diagnosing a
|
||||
# subprocess crash where Studio's traceback only shows the
|
||||
# subprocess crash where Unsloth's traceback only shows the
|
||||
# symptom (httpx ReadError) but not the cause.
|
||||
run: |
|
||||
mkdir -p logs/llama-server
|
||||
|
|
@ -1348,7 +1348,7 @@ jobs:
|
|||
|
||||
# ── folded from studio-windows-no-vs-smoke.yml: install + run with no Visual Studio ──
|
||||
no-vs-cpu:
|
||||
name: Studio install + inference without Visual Studio
|
||||
name: Unsloth install + inference without Visual Studio
|
||||
runs-on: windows-latest
|
||||
timeout-minutes: 35
|
||||
defaults:
|
||||
|
|
@ -1502,7 +1502,7 @@ jobs:
|
|||
python -m pip install torch --index-url https://download.pytorch.org/whl/cpu --extra-index-url https://pypi.org/simple
|
||||
python -c "import torch; print('torch', torch.__version__, 'cuda?', torch.cuda.is_available())"
|
||||
|
||||
- name: Install Studio (--local, --no-torch) with no build tools present
|
||||
- name: Install Unsloth (--local, --no-torch) with no build tools present
|
||||
shell: pwsh
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
|
|
@ -1538,13 +1538,13 @@ jobs:
|
|||
echo "Prebuilt installed with no build tools:"
|
||||
cat "$INFO"
|
||||
|
||||
- name: Add Studio shim to GITHUB_PATH
|
||||
- name: Add Unsloth shim to GITHUB_PATH
|
||||
run: |
|
||||
SHIM_DIR=~/.unsloth/studio/bin
|
||||
[ -f "$SHIM_DIR/unsloth.exe" ] || { echo "::error::unsloth.exe shim not found"; ls -la ~/.unsloth/studio/ || true; exit 1; }
|
||||
cygpath -w "$SHIM_DIR" >> "$GITHUB_PATH"
|
||||
|
||||
- name: Reset auth + boot Studio (API-only)
|
||||
- name: Reset auth + boot Unsloth (API-only)
|
||||
run: |
|
||||
unsloth studio reset-password
|
||||
mkdir -p logs
|
||||
|
|
@ -1613,10 +1613,10 @@ jobs:
|
|||
}
|
||||
Remove-Item -LiteralPath $root -Recurse -Force -ErrorAction SilentlyContinue
|
||||
|
||||
- name: Stop Studio
|
||||
- name: Stop Unsloth
|
||||
if: always()
|
||||
shell: cmd
|
||||
run: echo Stop Studio (no-op; runner reclaims STUDIO_PID=%STUDIO_PID% at job end)
|
||||
run: echo Stop Unsloth (no-op; runner reclaims STUDIO_PID=%STUDIO_PID% at job end)
|
||||
|
||||
- name: Collect llama-server logs
|
||||
if: always()
|
||||
|
|
|
|||
32
.github/workflows/studio-windows-ui-smoke.yml
vendored
32
.github/workflows/studio-windows-ui-smoke.yml
vendored
|
|
@ -4,11 +4,11 @@
|
|||
# Windows counterpart to studio-ui-smoke.yml / studio-mac-ui-smoke.yml.
|
||||
# Same Playwright + Chromium end-to-end chat UI flow + extra UI flow,
|
||||
# but on the FREE windows-latest runner so we catch Windows-specific
|
||||
# regressions in the install path (install.ps1), the Studio CLI's
|
||||
# regressions in the install path (install.ps1), the Unsloth CLI's
|
||||
# Windows process-management branches, and the llama.cpp prebuilt's
|
||||
# Windows HTTP layer.
|
||||
|
||||
name: Windows Studio UI CI
|
||||
name: Windows Unsloth UI CI
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
|
|
@ -49,7 +49,7 @@ jobs:
|
|||
GGUF_FILE: gemma-3-270m-it-UD-Q4_K_XL.gguf
|
||||
STUDIO_PORT: '18896'
|
||||
HF_HOME: ${{ github.workspace }}/hf-cache
|
||||
# Force UTF-8 for stdio so Python tools (hf download, Studio
|
||||
# Force UTF-8 for stdio so Python tools (hf download, Unsloth
|
||||
# CLI, etc.) can print Unicode characters like the success
|
||||
# checkmark "✓". Windows defaults to cp1252 / charmap and
|
||||
# any tool that prints "OK ✓" hits a UnicodeEncodeError.
|
||||
|
|
@ -121,7 +121,7 @@ jobs:
|
|||
# studio-windows-update-smoke.yml for the full rationale --
|
||||
# creating an empty studio/frontend/dist trips setup.ps1's
|
||||
# mtime-based staleness check into "frontend up to date, skip
|
||||
# rebuild" and Studio boots with an empty dist directory.
|
||||
# rebuild" and Unsloth boots with an empty dist directory.
|
||||
# Add-MpPreference accepts paths that do not yet exist.
|
||||
foreach ($p in @(
|
||||
"$env:USERPROFILE\.unsloth",
|
||||
|
|
@ -148,7 +148,7 @@ jobs:
|
|||
Set-Content -LiteralPath (Join-Path $appDir 'launch-studio.vbs') -Value 'WScript.Echo "legacy"' -Encoding Unicode
|
||||
Write-Host "seeded legacy launch-studio.vbs at $appDir"
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
# install.ps1 is the supported Windows installer. install.sh
|
||||
# has no Windows branch (apt-get / brew calls). The PS1
|
||||
# script's `Install-UnslothStudio @args` line at the bottom
|
||||
|
|
@ -205,7 +205,7 @@ jobs:
|
|||
echo "install.ps1 installed the Windows prebuilt llama.cpp:"
|
||||
cat "$INFO"
|
||||
|
||||
- name: Assert Studio launcher chain (no VBS, hidden PowerShell shortcut)
|
||||
- name: Assert Unsloth launcher chain (no VBS, hidden PowerShell shortcut)
|
||||
# The shortcut launch path is otherwise untested here (the steps below
|
||||
# boot `unsloth studio` directly). Guard against re-introducing the VBS
|
||||
# that tripped Kaspersky HEUR:Trojan.VBS.Agent.gen and against the .lnk
|
||||
|
|
@ -234,7 +234,7 @@ jobs:
|
|||
}
|
||||
Write-Host "launcher chain OK (no VBS; hidden powershell over launch-studio.ps1)"
|
||||
|
||||
- name: Launch Studio via the shortcut and assert health
|
||||
- name: Launch Unsloth via the shortcut and assert health
|
||||
# Run the exact command the .lnk stores (hidden PowerShell over
|
||||
# launch-studio.ps1) and confirm it brings the backend up. This is the
|
||||
# only step that proves the shortcut launch is not silently broken.
|
||||
|
|
@ -265,10 +265,10 @@ jobs:
|
|||
$owner = (Get-NetTCPConnection -LocalPort $foundPort -State Listen -ErrorAction Stop | Select-Object -First 1).OwningProcess
|
||||
if ($owner) { taskkill /PID $owner /T /F 2>$null | Out-Null }
|
||||
} catch {}
|
||||
if (-not $foundPort) { throw "Studio did not become healthy when launched via the shortcut" }
|
||||
Write-Host "Studio healthy on port $foundPort (launched via the shortcut)"
|
||||
if (-not $foundPort) { throw "Unsloth did not become healthy when launched via the shortcut" }
|
||||
Write-Host "Unsloth healthy on port $foundPort (launched via the shortcut)"
|
||||
|
||||
- name: Add Studio shim to GITHUB_PATH
|
||||
- name: Add Unsloth shim to GITHUB_PATH
|
||||
# install.ps1 puts unsloth.exe at $StudioHome\bin\unsloth.exe
|
||||
# and adds that dir to the User PATH via the Windows registry.
|
||||
# Registry-level PATH updates don't propagate to a running
|
||||
|
|
@ -284,7 +284,7 @@ jobs:
|
|||
fi
|
||||
# GITHUB_PATH wants Windows-style paths; convert via cygpath.
|
||||
cygpath -w "$SHIM_DIR" >> "$GITHUB_PATH"
|
||||
echo "Added Studio shim dir to PATH: $(cygpath -w "$SHIM_DIR")"
|
||||
echo "Added Unsloth shim dir to PATH: $(cygpath -w "$SHIM_DIR")"
|
||||
|
||||
- name: Install Playwright + Chromium
|
||||
# No --with-deps on Windows: that flag installs Linux apt
|
||||
|
|
@ -294,7 +294,7 @@ jobs:
|
|||
python -m pip install 'playwright>=1.45'
|
||||
python -m playwright install chromium
|
||||
|
||||
- name: Reset auth + boot Studio
|
||||
- name: Reset auth + boot Unsloth
|
||||
run: |
|
||||
unsloth studio reset-password
|
||||
mkdir -p logs
|
||||
|
|
@ -339,13 +339,13 @@ jobs:
|
|||
mkdir -p logs/playwright
|
||||
python tests/studio/playwright_chat_ui.py
|
||||
|
||||
- name: Stop Studio (chat-ui ends with Shutdown click; this is belt-and-suspenders)
|
||||
- name: Stop Unsloth (chat-ui ends with Shutdown click; this is belt-and-suspenders)
|
||||
if: always()
|
||||
run: |
|
||||
kill "${STUDIO_PID}" 2>/dev/null || true
|
||||
sleep 2
|
||||
|
||||
- name: Reset auth + boot Studio for extra UI tests (port 18897)
|
||||
- name: Reset auth + boot Unsloth for extra UI tests (port 18897)
|
||||
run: |
|
||||
unsloth studio reset-password
|
||||
mkdir -p logs
|
||||
|
|
@ -372,7 +372,7 @@ jobs:
|
|||
echo "STUDIO_EXTRA_OLD_PW=$OLD" >> "$GITHUB_ENV"
|
||||
echo "STUDIO_EXTRA_NEW_PW=$NEW" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Drive Compare/Recipes/Export/Studio/Settings with Playwright
|
||||
- name: Drive Compare/Recipes/Export/Unsloth/Settings with Playwright
|
||||
env:
|
||||
BASE_URL: http://127.0.0.1:18897
|
||||
STUDIO_OLD_PW: ${{ env.STUDIO_EXTRA_OLD_PW }}
|
||||
|
|
@ -386,7 +386,7 @@ jobs:
|
|||
mkdir -p logs/playwright_extra
|
||||
python tests/studio/playwright_extra_ui.py
|
||||
|
||||
- name: Stop second Studio
|
||||
- name: Stop second Unsloth
|
||||
if: always()
|
||||
run: |
|
||||
kill "${STUDIO_EXTRA_PID}" 2>/dev/null || true
|
||||
|
|
|
|||
|
|
@ -5,19 +5,19 @@
|
|||
# studio-mac-update-smoke.yml. Verifies that on the FREE
|
||||
# windows-latest runner:
|
||||
#
|
||||
# 1. install.ps1 --local --no-torch installs Studio AND auto-fetches
|
||||
# 1. install.ps1 --local --no-torch installs Unsloth AND auto-fetches
|
||||
# the prebuilt llama.cpp Windows binary (app-<tag>-windows-x64-cpu
|
||||
# from unslothai/llama.cpp). Hitting the source-build fallback is
|
||||
# treated as an Unsloth bug -- Studio must always pick the
|
||||
# treated as an Unsloth bug -- Unsloth must always pick the
|
||||
# prebuilt on Windows.
|
||||
# 2. unsloth studio update --local is idempotent. Two consecutive
|
||||
# runs both report "prebuilt up to date and validated", no
|
||||
# source-build fallback. The CLI's _find_setup_script picks
|
||||
# setup.ps1 on Windows automatically.
|
||||
# 3. The installed Studio still boots and /api/health returns
|
||||
# 3. The installed Unsloth still boots and /api/health returns
|
||||
# healthy after the update path.
|
||||
|
||||
name: Windows Studio Update CI
|
||||
name: Windows Unsloth Update CI
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
|
|
@ -45,7 +45,7 @@ permissions:
|
|||
|
||||
jobs:
|
||||
update-idempotency:
|
||||
name: Studio Updating Tests
|
||||
name: Unsloth Updating Tests
|
||||
runs-on: windows-latest
|
||||
timeout-minutes: 30
|
||||
defaults:
|
||||
|
|
@ -53,7 +53,7 @@ jobs:
|
|||
shell: bash
|
||||
env:
|
||||
# Force UTF-8 for stdio (Windows defaults to cp1252; hf
|
||||
# download / Studio CLI print "✓" checkmarks and crash
|
||||
# download / Unsloth CLI print "✓" checkmarks and crash
|
||||
# otherwise).
|
||||
PYTHONIOENCODING: utf-8
|
||||
PYTHONUTF8: '1'
|
||||
|
|
@ -90,7 +90,7 @@ jobs:
|
|||
# reuses the existing Node with no download.
|
||||
#
|
||||
# (2) Defender. windows-latest's real-time scan opens / hashes
|
||||
# every file Studio writes during install (Vite output =
|
||||
# every file Unsloth writes during install (Vite output =
|
||||
# thousands of small chunks, uv pip = wheel-extraction =
|
||||
# thousands of small files). The latency dominates the
|
||||
# 200 s frontend build and the 90 s deps install. Adding
|
||||
|
|
@ -109,7 +109,7 @@ jobs:
|
|||
# setup.ps1 line 1281-1296's mtime-based "is the frontend
|
||||
# stale?" check into "up to date, skip rebuild", because the
|
||||
# newly-created dist's mtime is younger than every source
|
||||
# file. Studio then boots with an empty dist and 500s on
|
||||
# file. Unsloth then boots with an empty dist and 500s on
|
||||
# GET / with FileNotFoundError: dist\index.html. See run
|
||||
# 25546676715 / job 74984469728.
|
||||
# Add-MpPreference accepts paths that do not yet exist; the
|
||||
|
|
@ -129,7 +129,7 @@ jobs:
|
|||
}
|
||||
}
|
||||
|
||||
- name: Install Studio (--local, --no-torch)
|
||||
- name: Install Unsloth (--local, --no-torch)
|
||||
shell: pwsh
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
|
|
@ -168,7 +168,7 @@ jobs:
|
|||
echo "install.ps1 installed the Windows prebuilt llama.cpp:"
|
||||
cat "$INFO"
|
||||
|
||||
- name: Add Studio shim to GITHUB_PATH
|
||||
- name: Add Unsloth shim to GITHUB_PATH
|
||||
run: |
|
||||
SHIM_DIR=~/.unsloth/studio/bin
|
||||
if [ ! -f "$SHIM_DIR/unsloth.exe" ]; then
|
||||
|
|
@ -212,7 +212,7 @@ jobs:
|
|||
grep -qE "prebuilt up to date and validated|prebuilt installed and validated" logs/update2.log
|
||||
echo "second update was clean"
|
||||
|
||||
- name: Boot Studio briefly to confirm the install is still usable
|
||||
- name: Boot Unsloth briefly to confirm the install is still usable
|
||||
run: |
|
||||
mkdir -p logs
|
||||
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p 18891 \
|
||||
|
|
@ -239,13 +239,13 @@ jobs:
|
|||
sleep 1
|
||||
done
|
||||
if [ -z "$HEALTHY" ]; then
|
||||
echo "Studio failed to come up after \`update\`"
|
||||
echo "Unsloth failed to come up after \`update\`"
|
||||
tail -200 logs/studio.log
|
||||
kill "$PID" 2>/dev/null || true
|
||||
exit 1
|
||||
fi
|
||||
kill "$PID" 2>/dev/null || true
|
||||
echo "post-update Studio /api/health OK"
|
||||
echo "post-update Unsloth /api/health OK"
|
||||
|
||||
- name: Uninstall and verify clean
|
||||
# Round-trip through scripts/uninstall.ps1 against the default
|
||||
|
|
|
|||
10
.github/workflows/wheel-smoke.yml
vendored
10
.github/workflows/wheel-smoke.yml
vendored
|
|
@ -3,7 +3,7 @@
|
|||
|
||||
# Builds the PyPI wheel from the PR branch, then verifies the built wheel
|
||||
# actually contains what we expect to ship and does NOT contain the broken
|
||||
# Studio bundle that 2026.5.1 published. This is the single workflow that
|
||||
# Unsloth bundle that 2026.5.1 published. This is the single workflow that
|
||||
# would have blocked the 2026.5.1 release before twine upload.
|
||||
#
|
||||
# Verified locally end-to-end against this branch:
|
||||
|
|
@ -12,7 +12,7 @@
|
|||
# lockfile shipped, frontend dist shipped,
|
||||
# no node_modules in wheel, no bun.lock in wheel,
|
||||
# main bundle has unstable_Provider hits=1 (assistant-ui internals only).
|
||||
# - Studio backend imports cleanly from the installed wheel with the
|
||||
# - Unsloth backend imports cleanly from the installed wheel with the
|
||||
# lightweight dep set below.
|
||||
|
||||
name: Wheel CI
|
||||
|
|
@ -101,7 +101,7 @@ jobs:
|
|||
hits = data.count("unstable_Provider:")
|
||||
print(f"main bundle: {js[0]}")
|
||||
print(f"unstable_Provider hits: {hits} (>=4 indicates 2026.5.1 regression)")
|
||||
checks["bundle has no Studio unstable_Provider call site"] = (hits < 4)
|
||||
checks["bundle has no Unsloth unstable_Provider call site"] = (hits < 4)
|
||||
|
||||
print()
|
||||
for k, v in checks.items():
|
||||
|
|
@ -109,7 +109,7 @@ jobs:
|
|||
sys.exit(0 if all(checks.values()) else 1)
|
||||
PY
|
||||
|
||||
- name: Studio backend import smoke
|
||||
- name: Unsloth backend import smoke
|
||||
# Imports `studio.backend.main:app` from the freshly-installed wheel in
|
||||
# a clean venv. This catches the class of bug that 2026.5.1 shipped with:
|
||||
# frontend dist missing, package-lock.json missing, or the wheel's Python
|
||||
|
|
@ -125,7 +125,7 @@ jobs:
|
|||
/tmp/v/bin/pip install --no-deps dist/unsloth-*.whl
|
||||
# Run from /tmp so Python imports the installed package, not the source tree.
|
||||
cd /tmp
|
||||
/tmp/v/bin/python -c "from studio.backend.main import app; print('Studio backend OK:', app.title)"
|
||||
/tmp/v/bin/python -c "from studio.backend.main import app; print('Unsloth backend OK:', app.title)"
|
||||
|
||||
- name: Upload wheel on failure
|
||||
if: failure()
|
||||
|
|
|
|||
16
README.md
16
README.md
|
|
@ -65,7 +65,7 @@ Unsloth Studio (Beta) works on **Windows, Linux, WSL** and **macOS**.
|
|||
* **CPU:** Supported for Chat and Data Recipes currently
|
||||
* **NVIDIA:** Training works on RTX 30/40/50, Blackwell, DGX Spark, Station and more
|
||||
* **macOS:** Training, MLX and GGUF inference are ALL supported.
|
||||
* **AMD:** Chat + Data works. Train with [Unsloth Core](#unsloth-core-code-based). Studio support is out soon.
|
||||
* **AMD:** Chat + Data works. Train with [Unsloth Core](#unsloth-core-code-based). Unsloth Studio support is out soon.
|
||||
* **Multi-GPU:** Available now, with a major upgrade on the way
|
||||
|
||||
#### macOS, Linux, WSL:
|
||||
|
|
@ -86,7 +86,7 @@ unsloth studio -p 8888
|
|||
```
|
||||
For LAN or cloud access, add `-H 0.0.0.0` (raw port only; add `--cloudflare` for a public URL). By default, Unsloth is accessible only locally.
|
||||
|
||||
To reach Studio over HTTPS, use `unsloth studio --secure`. Studio stays bound to localhost and is reached only through a free Cloudflare tunnel, which publishes it at a public `https://*.trycloudflare.com` URL (it fails closed if the tunnel can't start, so the raw port is never exposed). This makes Studio reachable from the internet, so anyone with the link and API key can use it and run code: keep your API key private (see Remote access below).
|
||||
To reach Unsloth over HTTPS, use `unsloth studio --secure`. Unsloth stays bound to localhost and is reached only through a free Cloudflare tunnel, which publishes it at a public `https://*.trycloudflare.com` URL (it fails closed if the tunnel can't start, so the raw port is never exposed). This makes Unsloth reachable from the internet, so anyone with the link and API key can use it and run code: keep your API key private (see Remote access below).
|
||||
|
||||
#### Docker
|
||||
Use our [Docker image](https://hub.docker.com/r/unsloth/unsloth) ```unsloth/unsloth``` container. Run:
|
||||
|
|
@ -208,7 +208,7 @@ unsloth studio -p 8888
|
|||
#### Remote access: `--secure` (HTTPS tunnel) vs raw port
|
||||
By default `unsloth studio` binds to `127.0.0.1` (this machine only). To reach it from another device, pick one of:
|
||||
|
||||
- `--secure` (recommended): serve **only** through a free Cloudflare HTTPS link. Studio stays bound to localhost and the tunnel provides the public URL; it fails closed (does not start) if the tunnel can't come up, so the raw port is never exposed.
|
||||
- `--secure` (recommended): serve **only** through a free Cloudflare HTTPS link. Unsloth stays bound to localhost and the tunnel provides the public URL; it fails closed (does not start) if the tunnel can't come up, so the raw port is never exposed.
|
||||
```bash
|
||||
unsloth studio --secure -p 8888
|
||||
```
|
||||
|
|
@ -218,7 +218,7 @@ unsloth studio -H 0.0.0.0 -p 8888
|
|||
```
|
||||
The Cloudflare tunnel is **off by default**: `-H 0.0.0.0` exposes the raw port only, not a public internet URL. Pair the wildcard bind with `--cloudflare` (`unsloth studio -H 0.0.0.0 --cloudflare`) to also publish a public `https://*.trycloudflare.com` link, or prefer `--secure` (above), which keeps the raw port private. `--cloudflare` has no effect on a loopback bind.
|
||||
|
||||
The first time Studio is published on a public URL (`--secure` or `--cloudflare`) with the auto-generated admin password still in place, it asks for a new admin password in the terminal (masked input with confirmation) before the public link goes up. Without an attached terminal it warns instead and keeps the bootstrap deadline: Studio shuts down after `UNSLOTH_STUDIO_BOOTSTRAP_TIMEOUT` (default 1 hour) unless the password is changed in the web UI.
|
||||
The first time Unsloth is published on a public URL (`--secure` or `--cloudflare`) with the auto-generated admin password still in place, it asks for a new admin password in the terminal (masked input with confirmation) before the public link goes up. Without an attached terminal it warns instead and keeps the bootstrap deadline: Unsloth shuts down after `UNSLOTH_STUDIO_BOOTSTRAP_TIMEOUT` (default 1 hour) unless the password is changed in the web UI.
|
||||
|
||||
For headless setups that cannot answer that prompt, set the initial admin password non-interactively with `--password` (only takes effect when no password is set yet; if one already exists it is a hard error, so rotate later with `unsloth studio reset-password`):
|
||||
|
||||
|
|
@ -230,7 +230,7 @@ printf '%s\n' 'your-strong-password' | unsloth studio --secure --password - #
|
|||
|
||||
A literal `--password VALUE` is visible in the process list and shell history, so prefer the `UNSLOTH_STUDIO_PASSWORD` env var or `--password -` (stdin) for automation. This applies to any launch (public or a headless `-H 0.0.0.0` bind), and the password is set in the parent before the server binds, so it never reaches a re-executed child process.
|
||||
|
||||
Server-side tools (web search, Python and terminal code execution) run as your user and are on by default. Anyone who can reach the server with the API key can run code on this machine, so keep your API key private and pass `--disable-tools` when exposing Studio.
|
||||
Server-side tools (web search, Python and terminal code execution) run as your user and are on by default. Anyone who can reach the server with the API key can run code on this machine, so keep your API key private and pass `--disable-tools` when exposing Unsloth.
|
||||
|
||||
#### Advanced launch options
|
||||
Installer options can be passed as environment variables. On macOS, Linux and WSL place the variable after the pipe so the shell passes it to `sh`; on Windows set it with `$env:` before piping to `iex`.
|
||||
|
|
@ -243,7 +243,7 @@ curl -fsSL https://unsloth.ai/install.sh | UNSLOTH_NO_TORCH=1 sh
|
|||
$env:UNSLOTH_NO_TORCH=1; irm https://unsloth.ai/install.ps1 | iex
|
||||
```
|
||||
|
||||
Skip the post-install prompt that starts Studio (useful for automated installs):
|
||||
Skip the post-install prompt that starts Unsloth (useful for automated installs):
|
||||
```bash
|
||||
curl -fsSL https://unsloth.ai/install.sh | UNSLOTH_SKIP_AUTOSTART=1 sh
|
||||
```
|
||||
|
|
@ -279,9 +279,9 @@ UNSLOTH_NPM_REGISTRY=https://artifactory.example.com/api/npm/npm/ ./install.sh -
|
|||
```powershell
|
||||
$env:UNSLOTH_NPM_REGISTRY='https://artifactory.example.com/api/npm/npm/'; .\install.ps1 --local
|
||||
```
|
||||
It is threaded as `--registry` into the Studio frontend `npm`/`bun` installs; the supply-chain locks (7-day `min-release-age`, exact version pins) stay in force.
|
||||
It is threaded as `--registry` into the Unsloth frontend `npm`/`bun` installs; the supply-chain locks (7-day `min-release-age`, exact version pins) stay in force.
|
||||
|
||||
Cap Studio's native CPU thread pools on high-core hosts: `UNSLOTH_CPU_THREADS=8 unsloth studio -p 8888`.
|
||||
Cap Unsloth's native CPU thread pools on high-core hosts: `UNSLOTH_CPU_THREADS=8 unsloth studio -p 8888`.
|
||||
|
||||
#### Uninstall
|
||||
The recommended way to fully remove Unsloth Studio is the matching uninstall script for your OS. It stops any running servers, removes the install dir, the launcher data dir, the desktop shortcut, and any platform-specific entries (macOS `.app` bundle + Launch Services on Mac; Start Menu, `HKCU\Software\Unsloth` registry key and user `PATH` entries on Windows):
|
||||
|
|
|
|||
8
build.sh
8
build.sh
|
|
@ -4,9 +4,9 @@
|
|||
|
||||
set -euo pipefail
|
||||
|
||||
# PyPI/Studio release publishing must use `./build.sh publish` (or an
|
||||
# equivalent stamp -> build -> verify-dist -> upload flow) so packaged Studio
|
||||
# artifacts include the display-only Studio release version.
|
||||
# PyPI/Unsloth release publishing must use `./build.sh publish` (or an
|
||||
# equivalent stamp -> build -> verify-dist -> upload flow) so packaged Unsloth
|
||||
# artifacts include the display-only Unsloth release version.
|
||||
|
||||
# 1. Build frontend (Vite outputs to dist/)
|
||||
cd studio/frontend
|
||||
|
|
@ -87,7 +87,7 @@ cd ../..
|
|||
# 2. Clean old artifacts
|
||||
rm -rf build dist *.egg-info
|
||||
|
||||
# 3. Stamp display-only Studio release metadata for packaged builds.
|
||||
# 3. Stamp display-only Unsloth release metadata for packaged builds.
|
||||
_STUDIO_BUILD_INFO="studio/backend/utils/_studio_release_build.py"
|
||||
_STUDIO_BUILD_INFO_BACKUP="$(mktemp)"
|
||||
cp "$_STUDIO_BUILD_INFO" "$_STUDIO_BUILD_INFO_BACKUP"
|
||||
|
|
|
|||
36
install.ps1
36
install.ps1
|
|
@ -176,7 +176,7 @@ function Install-UnslothStudio {
|
|||
$envOverride = $env:STUDIO_HOME.Trim()
|
||||
}
|
||||
|
||||
# Custom Studio roots are not supported with --tauri (desktop app still
|
||||
# Custom Unsloth roots are not supported with --tauri (desktop app still
|
||||
# resolves %USERPROFILE%\.unsloth\studio). Pass through if override == legacy.
|
||||
if ($TauriMode -and $envOverride) {
|
||||
$_tauriOverride = $envOverride
|
||||
|
|
@ -756,7 +756,7 @@ function Find-FreeLaunchPort {
|
|||
return `$null
|
||||
}
|
||||
|
||||
# If Studio is already healthy on any expected port, just open it and exit.
|
||||
# If Unsloth is already healthy on any expected port, just open it and exit.
|
||||
`$existingPort = Find-HealthyStudioPort
|
||||
if (`$existingPort) {
|
||||
Start-Process "http://localhost:`$existingPort"
|
||||
|
|
@ -772,7 +772,7 @@ try {
|
|||
`$haveMutex = `$true
|
||||
}
|
||||
if (-not `$haveMutex) {
|
||||
# Another launcher is already running; wait for it to bring Studio up
|
||||
# Another launcher is already running; wait for it to bring Unsloth up
|
||||
`$deadline = (Get-Date).AddSeconds(`$timeoutSec)
|
||||
while ((Get-Date) -lt `$deadline) {
|
||||
`$port = Find-HealthyStudioPort
|
||||
|
|
@ -1438,7 +1438,7 @@ exit 0
|
|||
if (Test-Path -LiteralPath $VenvPython) {
|
||||
# why: matching guard to the .venv branch below -- in env-mode
|
||||
# $StudioHome is a user-chosen workspace, so refuse to nuke an
|
||||
# existing $StudioHome\unsloth_studio that lacks Studio sentinels.
|
||||
# existing $StudioHome\unsloth_studio that lacks Unsloth sentinels.
|
||||
# -PathType Leaf rejects a directory at the sentinel path. Accept the
|
||||
# in-VENV ownership marker so partial-install retries are not blocked.
|
||||
if (
|
||||
|
|
@ -1449,7 +1449,7 @@ exit 0
|
|||
) {
|
||||
Write-Host "[ERROR] $VenvDir already exists but does not look like an Unsloth Studio install." -ForegroundColor Red
|
||||
Write-Host " Move it aside or choose an empty UNSLOTH_STUDIO_HOME." -ForegroundColor Yellow
|
||||
throw "Refusing to delete non-Studio venv at $VenvDir"
|
||||
throw "Refusing to delete non-Unsloth venv at $VenvDir"
|
||||
}
|
||||
# New layout already exists -- replace only after preserving rollback copy.
|
||||
substep "preserving existing environment for rollback..."
|
||||
|
|
@ -1468,7 +1468,7 @@ exit 0
|
|||
# workspace root (e.g. user's existing project Python venv).
|
||||
$OldVenv = Join-Path $StudioHome ".venv"
|
||||
$OldPy = Join-Path $OldVenv "Scripts\python.exe"
|
||||
substep "found legacy Studio environment, validating..."
|
||||
substep "found legacy Unsloth environment, validating..."
|
||||
$prevEAP2 = $ErrorActionPreference
|
||||
$ErrorActionPreference = "Continue"
|
||||
try {
|
||||
|
|
@ -1498,7 +1498,7 @@ exit 0
|
|||
# Skip in env-mode so we don't relocate the default-install venv into
|
||||
# the workspace root.
|
||||
$CwdVenv = Join-Path $env:USERPROFILE "unsloth_studio"
|
||||
substep "found CWD-relative Studio environment, migrating to $VenvDir..."
|
||||
substep "found CWD-relative Unsloth environment, migrating to $VenvDir..."
|
||||
Move-Item -LiteralPath $CwdVenv -Destination $VenvDir -Force
|
||||
substep "moved ~/unsloth_studio -> ~/.unsloth/studio/unsloth_studio"
|
||||
$_Migrated = $true
|
||||
|
|
@ -1517,7 +1517,7 @@ exit 0
|
|||
substep "$VenvDir"
|
||||
}
|
||||
|
||||
# Mark the freshly-created venv as Studio-owned so a partial install can be
|
||||
# Mark the freshly-created venv as Unsloth-owned so a partial install can be
|
||||
# repaired by re-running install.ps1; the env-mode deletion guard above
|
||||
# accepts this marker as the primary sentinel.
|
||||
if (Test-Path -LiteralPath $VenvDir -PathType Container) {
|
||||
|
|
@ -1526,7 +1526,7 @@ exit 0
|
|||
|
||||
# ── Helper: run amd-smi without triggering a UAC elevation prompt ──
|
||||
# amd-smi on Windows auto-elevates to read GPU/APU memory, surfacing a confusing
|
||||
# DiskPart UAC prompt mid-install (Studio backend amd.py hits the same).
|
||||
# DiskPart UAC prompt mid-install (Unsloth backend amd.py hits the same).
|
||||
# __COMPAT_LAYER=RunAsInvoker forces it (and helpers it spawns) to run
|
||||
# un-elevated; on failure the WMI name -> gfx fallback still resolves the arch.
|
||||
function Invoke-AmdSmiNoElevate {
|
||||
|
|
@ -1653,7 +1653,7 @@ exit 0
|
|||
function Test-HipinfoIsVenvInternal {
|
||||
param([AllowNull()][string]$HipinfoPath)
|
||||
if ([string]::IsNullOrWhiteSpace($HipinfoPath)) { return $false }
|
||||
# Also derive the venv from the setup python + default Studio home, so
|
||||
# Also derive the venv from the setup python + default Unsloth home, so
|
||||
# the venv hipInfo is caught when VenvDir/VIRTUAL_ENV are unset.
|
||||
$venvRoots = @()
|
||||
if ($env:VIRTUAL_ENV) { $venvRoots += $env:VIRTUAL_ENV }
|
||||
|
|
@ -1663,7 +1663,7 @@ exit 0
|
|||
try { $venvRoots += (Split-Path -Parent (Split-Path -Parent $env:UNSLOTH_SETUP_PYTHON)) } catch {}
|
||||
}
|
||||
if ($env:USERPROFILE) { $venvRoots += (Join-Path $env:USERPROFILE ".unsloth\studio\unsloth_studio") }
|
||||
# A custom Studio home (UNSLOTH_STUDIO_HOME / STUDIO_HOME alias) moves the
|
||||
# A custom Unsloth home (UNSLOTH_STUDIO_HOME / STUDIO_HOME alias) moves the
|
||||
# venv off the default path; seed it too or its hipInfo escapes the filter.
|
||||
$studioHomeEnv = if (-not [string]::IsNullOrWhiteSpace($env:UNSLOTH_STUDIO_HOME)) { $env:UNSLOTH_STUDIO_HOME.Trim() } elseif (-not [string]::IsNullOrWhiteSpace($env:STUDIO_HOME)) { $env:STUDIO_HOME.Trim() } else { $null }
|
||||
if ($studioHomeEnv) {
|
||||
|
|
@ -1942,7 +1942,7 @@ exit 0
|
|||
substep " Ensure the ROCm compute driver is installed alongside the display driver:" "Yellow"
|
||||
substep " https://rocm.docs.amd.com/en/latest/deploy/windows/index.html" "Yellow"
|
||||
} elseif ($ROCmGfxArch) {
|
||||
# Known arch: Studio setup installs AMD's bundled-runtime ROCm PyTorch wheels
|
||||
# Known arch: Unsloth setup installs AMD's bundled-runtime ROCm PyTorch wheels
|
||||
# (repo.amd.com), which ship their own runtime -- HIP SDK optional.
|
||||
step "gpu" "AMD ROCm ($ROCmGfxArch)" "Cyan"
|
||||
substep "Detected: $ROCmGpuLabel" "Cyan"
|
||||
|
|
@ -2219,8 +2219,8 @@ exit 0
|
|||
$torchInstallExit = Invoke-InstallCommandRetry -Label "install PyTorch (AMD ROCm)" { uv pip install --python $VenvPython --force-reinstall --default-index $ROCmIndexUrl $torchSpec $visionSpec $audioSpec }
|
||||
if ($torchInstallExit -ne 0) {
|
||||
# Transient AMD-index failure: fall back to a CPU base so the install
|
||||
# still completes; Studio setup retries ROCm afterwards.
|
||||
substep "ROCm PyTorch install failed (exit $torchInstallExit); using a CPU base, Studio setup retries ROCm." "Yellow"
|
||||
# still completes; Unsloth setup retries ROCm afterwards.
|
||||
substep "ROCm PyTorch install failed (exit $torchInstallExit); using a CPU base, Unsloth setup retries ROCm." "Yellow"
|
||||
# --force-reinstall: a failed ROCm install can leave an unpinned ROCm
|
||||
# torch (e.g. 2.10.0+rocm on gfx110X/gfx90a) that still satisfies the CPU
|
||||
# torch>= range, so without it uv would keep the ROCm build and only swap
|
||||
|
|
@ -2422,7 +2422,7 @@ exit 0
|
|||
Write-TauriLog "ERROR" "unsloth CLI was not installed correctly"
|
||||
Write-Host "[ERROR] unsloth CLI was not installed correctly." -ForegroundColor Red
|
||||
Write-Host " Expected: $UnslothExe" -ForegroundColor Yellow
|
||||
Write-Host " This usually means an older unsloth version was installed that does not include the Studio CLI." -ForegroundColor Yellow
|
||||
Write-Host " This usually means an older unsloth version was installed that does not include the Unsloth CLI." -ForegroundColor Yellow
|
||||
Write-Host " Try re-running the installer or see: https://github.com/unslothai/unsloth?tab=readme-ov-file#-quickstart" -ForegroundColor Yellow
|
||||
return (Exit-InstallFailure "unsloth CLI was not installed correctly")
|
||||
}
|
||||
|
|
@ -2533,7 +2533,7 @@ exit 0
|
|||
Write-Host " Move or remove it manually, then re-run the installer." -ForegroundColor Yellow
|
||||
throw "Cannot create unsloth launcher: $ShimExe is a directory."
|
||||
}
|
||||
# try/catch: if unsloth.exe is locked (Studio running), keep the old shim.
|
||||
# try/catch: if unsloth.exe is locked (Unsloth running), keep the old shim.
|
||||
$shimUpdated = $false
|
||||
try {
|
||||
if (Test-Path -LiteralPath $ShimExe) { Remove-Item -LiteralPath $ShimExe -Force -ErrorAction Stop }
|
||||
|
|
@ -2551,7 +2551,7 @@ exit 0
|
|||
if (Test-Path -LiteralPath $ShimExe) {
|
||||
Write-Host "[WARN] Could not refresh unsloth launcher at $ShimExe." -ForegroundColor Yellow
|
||||
Write-Host " This usually means a running 'unsloth studio' process still holds the file open." -ForegroundColor Yellow
|
||||
Write-Host " Close Studio and re-run the installer to pick up the latest launcher." -ForegroundColor Yellow
|
||||
Write-Host " Close Unsloth and re-run the installer to pick up the latest launcher." -ForegroundColor Yellow
|
||||
Write-Host " Continuing with the existing launcher." -ForegroundColor Yellow
|
||||
} else {
|
||||
Write-Host "[WARN] Could not create unsloth launcher at $ShimExe" -ForegroundColor Yellow
|
||||
|
|
@ -2616,7 +2616,7 @@ exit 0
|
|||
# Diagnostic only; never block install on a probe failure.
|
||||
}
|
||||
|
||||
# In interactive terminals, ask the user before starting Studio unless the
|
||||
# In interactive terminals, ask the user before starting Unsloth unless the
|
||||
# caller explicitly disabled the post-install prompt.
|
||||
# In non-interactive environments (CI, Docker) just print instructions.
|
||||
$IsInteractive = (-not $SkipAutostart) -and [Environment]::UserInteractive -and (-not [Console]::IsInputRedirected)
|
||||
|
|
|
|||
30
install.sh
30
install.sh
|
|
@ -97,7 +97,7 @@ if [ "$_VERBOSE" = true ]; then
|
|||
export UNSLOTH_VERBOSE=1
|
||||
fi
|
||||
|
||||
# Custom Studio roots are not supported with --tauri (desktop app still
|
||||
# Custom Unsloth roots are not supported with --tauri (desktop app still
|
||||
# resolves ~/.unsloth/studio). Pass through if the override == legacy default.
|
||||
if [ "$TAURI_MODE" = true ]; then
|
||||
_tauri_override_var=""
|
||||
|
|
@ -663,7 +663,7 @@ POLL_INTERVAL_SEC=0.25
|
|||
LOG_FILE="$DATA_DIR/studio.log"
|
||||
# why: in env-override mode multiple installs share an OS user; namespace the
|
||||
# lock and remember our own healthy port so we never attach to an unrelated
|
||||
# Studio listening on the global 8888..8908 range.
|
||||
# Unsloth listening on the global 8888..8908 range.
|
||||
LOCK_DIR="${XDG_RUNTIME_DIR:-/tmp}/unsloth-studio-launcher-$(id -u).lock"
|
||||
PORT_FILE=""
|
||||
# why: gate on the install-time mode (baked above) instead of the runtime env
|
||||
|
|
@ -734,7 +734,7 @@ _candidate_ports() {
|
|||
_find_healthy_port() {
|
||||
if [ -n "$PORT_FILE" ] && [ -f "$PORT_FILE" ]; then
|
||||
# why: env-mode installs only attach to a port we previously launched
|
||||
# ourselves; never to a sibling Studio that happens to be healthy.
|
||||
# ourselves; never to a sibling Unsloth that happens to be healthy.
|
||||
_p=$(cat "$PORT_FILE" 2>/dev/null || true)
|
||||
case "$_p" in
|
||||
''|*[!0-9]*) ;;
|
||||
|
|
@ -901,7 +901,7 @@ _acquire_lock() {
|
|||
# Lock dir exists -- check if owner is still alive
|
||||
_old_pid=$(cat "$LOCK_DIR/pid" 2>/dev/null || true)
|
||||
if [ -n "$_old_pid" ] && kill -0 "$_old_pid" 2>/dev/null; then
|
||||
# Another launcher is running; wait for it to bring Studio up
|
||||
# Another launcher is running; wait for it to bring Unsloth up
|
||||
_deadline=$(($(date +%s) + TIMEOUT_SEC))
|
||||
while [ "$(date +%s)" -lt "$_deadline" ]; do
|
||||
_port=$(_find_healthy_port) && {
|
||||
|
|
@ -1371,7 +1371,7 @@ WSLPS1_EOF
|
|||
# shortcut wasn't created; tell the user how to launch / re-enable it.
|
||||
if [ "$_css_created" -ne 1 ]; then
|
||||
substep "Couldn't create the Windows shortcut (WSL interop may be disabled)." "$C_WARN"
|
||||
substep " Launch Studio from Windows: wsl -d \"$_css_distro\" -- bash -lc 'unsloth studio'" "$C_WARN"
|
||||
substep " Launch Unsloth from Windows: wsl -d \"$_css_distro\" -- bash -lc 'unsloth studio'" "$C_WARN"
|
||||
substep " (re-enable shortcuts: turn WSL interop back on, e.g. run 'wsl --shutdown' then reopen WSL.)" "$C_WARN"
|
||||
fi
|
||||
fi
|
||||
|
|
@ -1439,7 +1439,7 @@ if [ "$MAC_INTEL" = true ]; then
|
|||
echo ""
|
||||
echo " NOTE: Intel Mac (x86_64) detected."
|
||||
echo " PyTorch is unavailable for this platform (dropped Jan 2024)."
|
||||
echo " Studio will install in GGUF-only mode."
|
||||
echo " Unsloth will install in GGUF-only mode."
|
||||
echo " Chat, inference via GGUF, and data recipes will work."
|
||||
echo " Training requires Apple Silicon or Linux with GPU."
|
||||
echo ""
|
||||
|
|
@ -1671,7 +1671,7 @@ _maybe_reroute_strixhalo_to_2404() {
|
|||
_maybe_reroute_strixhalo_to_2404 || true
|
||||
|
||||
# ── Check system dependencies ──
|
||||
# cmake/git are only needed to *build* llama.cpp from source. Studio downloads a
|
||||
# cmake/git are only needed to *build* llama.cpp from source. Unsloth downloads a
|
||||
# prebuilt by default, and setup.sh self-skips the source build when they're
|
||||
# absent -- so macOS doesn't block on cmake (requiring it would force a manual
|
||||
# Homebrew install). Linux keeps requiring them; its package manager has them.
|
||||
|
|
@ -1825,7 +1825,7 @@ _MIGRATED=false
|
|||
if [ -x "$VENV_DIR/bin/python" ]; then
|
||||
# why: matching guard to the .venv branch below -- in env-mode
|
||||
# $STUDIO_HOME is a user-chosen workspace, so refuse to nuke an
|
||||
# existing $STUDIO_HOME/unsloth_studio that lacks Studio sentinels.
|
||||
# existing $STUDIO_HOME/unsloth_studio that lacks Unsloth sentinels.
|
||||
# Accept the in-VENV ownership marker so partial-install retries are
|
||||
# not blocked. Sentinels must be regular files: -f follows symlinks
|
||||
# to files (the legitimate ln -s shim shape) but rejects directories
|
||||
|
|
@ -1846,7 +1846,7 @@ elif [ "$_STUDIO_HOME_REDIRECT" != "env" ] && [ -x "$STUDIO_HOME/.venv/bin/pytho
|
|||
# Skip in env-mode so we don't rm -rf an unrelated .venv at the
|
||||
# workspace root (e.g. user's existing project Python venv).
|
||||
# In no-torch mode, a missing torch package is expected; validate Python only.
|
||||
substep "found legacy Studio environment, validating..."
|
||||
substep "found legacy Unsloth environment, validating..."
|
||||
_legacy_ok=false
|
||||
if [ "$SKIP_TORCH" = true ]; then
|
||||
if "$STUDIO_HOME/.venv/bin/python" -c "import sys; print(sys.executable)" >/dev/null 2>&1; then
|
||||
|
|
@ -1903,7 +1903,7 @@ if [ ! -x "$VENV_DIR/bin/python" ]; then
|
|||
fi
|
||||
fi
|
||||
|
||||
# Mark the freshly-created venv as Studio-owned so a partial install can be
|
||||
# Mark the freshly-created venv as Unsloth-owned so a partial install can be
|
||||
# repaired by re-running install.sh; the env-mode deletion guard above accepts
|
||||
# this marker as the primary sentinel.
|
||||
if [ -x "$VENV_DIR/bin/python" ]; then
|
||||
|
|
@ -2335,7 +2335,7 @@ _pick_radeon_wheel() {
|
|||
# the installer -- always returns 0. Runs the idempotent helper (ROCm 7.2 +
|
||||
# librocdxg), then sources the env it persisted so detection finds the GPU.
|
||||
# Export the ROCm-on-WSL env into this process and persist it to /etc/profile.d
|
||||
# so non-login Studio/llama launches inherit it. Idempotent (writes only when
|
||||
# so non-login Unsloth/llama launches inherit it. Idempotent (writes only when
|
||||
# the drop-in is missing); no-op without librocdxg, so never fires off WSL.
|
||||
# /etc/profile.d is root-owned -- sudo-tee when not root, else ROCm vanishes
|
||||
# after this shell on a non-root reinstall. Best-effort either way.
|
||||
|
|
@ -2380,7 +2380,7 @@ _maybe_bootstrap_rocm_wsl() {
|
|||
rocminfo 2>/dev/null | awk '/Name:[[:space:]]*gfx[1-9]/ && !/generic/{found=1} END{exit !found}'; then
|
||||
# rocminfo may work only via the transient env _ensure_rocm_probe_env
|
||||
# just set, which dies with the installer. Persist the drop-in so login
|
||||
# shells (Studio, llama.cpp) inherit it -- else a reinstall over an
|
||||
# shells (Unsloth, llama.cpp) inherit it -- else a reinstall over an
|
||||
# existing /opt/rocm (uninstall keeps ROCm but drops it) loses the GPU.
|
||||
_persist_rocm_wsl_dropin
|
||||
return 0
|
||||
|
|
@ -2402,7 +2402,7 @@ _maybe_bootstrap_rocm_wsl() {
|
|||
# shellcheck disable=SC1091
|
||||
. /etc/profile.d/unsloth-rocm-wsl.sh || true
|
||||
else
|
||||
# librocdxg present but the env drop-in is gone (e.g. a Studio
|
||||
# librocdxg present but the env drop-in is gone (e.g. an Unsloth
|
||||
# uninstall removed it while keeping shared ROCm). Restore the env.
|
||||
_persist_rocm_wsl_dropin
|
||||
fi
|
||||
|
|
@ -3033,7 +3033,7 @@ if [ "$SKIP_TORCH" = false ] && [ -n "${TORCH_INDEX_URL:-}" ]; then
|
|||
fi
|
||||
|
||||
# ── Run studio setup ──
|
||||
tauri_log "STEP" "Running Studio setup"
|
||||
tauri_log "STEP" "Running Unsloth setup"
|
||||
# When --local, use the repo's own setup.sh directly.
|
||||
# Otherwise, find it inside the installed package.
|
||||
SETUP_SH=""
|
||||
|
|
@ -3227,7 +3227,7 @@ printf " ${C_TITLE}%s${C_RST}\n" "Unsloth Studio installed!"
|
|||
printf " ${C_DIM}%s${C_RST}\n" "$RULE"
|
||||
echo ""
|
||||
|
||||
# In interactive terminals, ask the user before starting Studio unless the
|
||||
# In interactive terminals, ask the user before starting Unsloth unless the
|
||||
# caller explicitly disabled the post-install prompt.
|
||||
# In non-interactive environments (Docker, CI, cloud-init) just print instructions.
|
||||
if [ "$_SKIP_AUTOSTART" != true ] && [ -t 1 ]; then
|
||||
|
|
|
|||
|
|
@ -219,7 +219,7 @@ fi
|
|||
echo "${ROCM_DIR}/lib" | $SUDO tee /etc/ld.so.conf.d/rocm.conf >/dev/null
|
||||
$SUDO ldconfig
|
||||
|
||||
# ── Step 4: persist environment (system-wide so Studio's worker inherits it) ──
|
||||
# ── Step 4: persist environment (system-wide so Unsloth's worker inherits it) ──
|
||||
say "Persisting ROCm-on-WSL environment"
|
||||
_envfile="/etc/profile.d/unsloth-rocm-wsl.sh"
|
||||
$SUDO tee "$_envfile" >/dev/null <<EOF
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved.
|
||||
|
||||
"""Lockfile supply-chain audit for the Studio frontend and Tauri shell.
|
||||
"""Lockfile supply-chain audit for the Unsloth frontend and Tauri shell.
|
||||
|
||||
Runs BEFORE `npm ci` / `cargo fetch` in CI. Refuses to proceed when a
|
||||
lockfile contains patterns indicating supply-chain injection (npm
|
||||
|
|
@ -294,7 +294,7 @@ CARGO_REGISTRY_SOURCE = "registry+https://github.com/rust-lang/crates.io-index"
|
|||
|
||||
# Cargo non-registry source allowlist: `(crate_name, exact_source_string)`.
|
||||
# Both must match verbatim; bumping the pinned SHA forces a re-review.
|
||||
# Studio's Tauri shell pulls `fix-path-env` from git because it is not
|
||||
# Unsloth's Tauri shell pulls `fix-path-env` from git because it is not
|
||||
# published to crates.io; commit c4c45d5 was reviewed when it landed.
|
||||
CARGO_SOURCE_ALLOWLIST: tuple[tuple[str, str], ...] = (
|
||||
(
|
||||
|
|
|
|||
|
|
@ -62,7 +62,7 @@ REPO_ROOT = Path(__file__).resolve().parents[1]
|
|||
# Hard caps (deliberately conservative; npm tarballs in this repo are
|
||||
# all well under these limits, so a packaging spike is noticeable).
|
||||
# ─────────────────────────────────────────────────────────────────────
|
||||
# Caps calibrated against the real Studio frontend transitive closure:
|
||||
# Caps calibrated against the real Unsloth frontend transitive closure:
|
||||
# - typescript.js is 9.1 MB (TS compiler bundled into one file)
|
||||
# - mermaid 11.x dist/mermaid.js.map is ~12 MB (sourcemap)
|
||||
# - lightningcss-linux-x64-{gnu,musl}.node is 10 MB
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
||||
|
||||
"""Stamp and verify display-only Studio release metadata for builds."""
|
||||
"""Stamp and verify display-only Unsloth release metadata for builds."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
|
|
@ -50,7 +50,7 @@ MAX_VERSION_LENGTH = 64
|
|||
PLACEHOLDER = """# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
||||
|
||||
\"\"\"Build-stamped Studio release metadata.
|
||||
\"\"\"Build-stamped Unsloth release metadata.
|
||||
|
||||
Release builds may rewrite this module in the build workspace before creating
|
||||
Python artifacts. Keep the committed value neutral so source checkouts do not
|
||||
|
|
@ -145,7 +145,7 @@ def build_info_source(version: str | None) -> str:
|
|||
return f'''# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
||||
|
||||
"""Build-stamped Studio release metadata."""
|
||||
"""Build-stamped Unsloth release metadata."""
|
||||
|
||||
STUDIO_RELEASE_VERSION = {literal}
|
||||
'''
|
||||
|
|
@ -168,7 +168,7 @@ def stamp(require_release: bool) -> int:
|
|||
version, source = resolve_version()
|
||||
if version is not None and not is_valid_version(version):
|
||||
print(
|
||||
f"Invalid Studio release version from {source}: {version!r}",
|
||||
f"Invalid Unsloth release version from {source}: {version!r}",
|
||||
file = sys.stderr,
|
||||
)
|
||||
return 2
|
||||
|
|
@ -196,9 +196,9 @@ def stamp(require_release: bool) -> int:
|
|||
if version is None:
|
||||
if require_release:
|
||||
print(
|
||||
"No Studio release version available. Set "
|
||||
"No Unsloth release version available. Set "
|
||||
"UNSLOTH_STUDIO_RELEASE_VERSION, build from a GitHub tag, "
|
||||
"or run from an exact local Studio release tag.",
|
||||
"or run from an exact local Unsloth release tag.",
|
||||
file = sys.stderr,
|
||||
)
|
||||
return 2
|
||||
|
|
@ -207,7 +207,7 @@ def stamp(require_release: bool) -> int:
|
|||
return 0
|
||||
|
||||
_atomic_write_text(BUILD_INFO_PATH, build_info_source(version), encoding = "utf-8")
|
||||
print(f"Stamping Studio release version {version} from {source}", file = sys.stderr)
|
||||
print(f"Stamping Unsloth release version {version} from {source}", file = sys.stderr)
|
||||
print(version)
|
||||
return 0
|
||||
|
||||
|
|
@ -233,7 +233,7 @@ def _read_sdist_member(path: Path) -> str | None:
|
|||
|
||||
def verify_dist(expected: str, dist_dir: Path) -> int:
|
||||
if not is_valid_version(expected):
|
||||
print(f"Invalid expected Studio release version: {expected!r}", file = sys.stderr)
|
||||
print(f"Invalid expected Unsloth release version: {expected!r}", file = sys.stderr)
|
||||
return 2
|
||||
|
||||
artifacts = list(dist_dir.glob("*.whl")) + list(dist_dir.glob("*.tar.gz"))
|
||||
|
|
@ -251,14 +251,14 @@ def verify_dist(expected: str, dist_dir: Path) -> int:
|
|||
if content is None:
|
||||
failures.append(f"{artifact.name}: missing {BUILD_INFO_SUFFIX}")
|
||||
elif expected_line not in content:
|
||||
failures.append(f"{artifact.name}: Studio release version mismatch")
|
||||
failures.append(f"{artifact.name}: Unsloth release version mismatch")
|
||||
|
||||
if failures:
|
||||
for failure in failures:
|
||||
print(failure, file = sys.stderr)
|
||||
return 2
|
||||
|
||||
print(f"Verified Studio release version {expected} in {len(artifacts)} artifact(s)")
|
||||
print(f"Verified Unsloth release version {expected} in {len(artifacts)} artifact(s)")
|
||||
return 0
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -83,7 +83,7 @@ function Uninstall-UnslothStudio {
|
|||
}
|
||||
}
|
||||
|
||||
# A path is a Studio-owned root iff one of install.ps1's sentinels exists:
|
||||
# A path is an Unsloth-owned root iff one of install.ps1's sentinels exists:
|
||||
# <root>\share\studio.conf, <root>\unsloth_studio\.unsloth-studio-owned,
|
||||
# or <root>\bin\unsloth.exe.
|
||||
function _IsStudioRoot {
|
||||
|
|
@ -164,7 +164,7 @@ function Uninstall-UnslothStudio {
|
|||
return $p
|
||||
}
|
||||
|
||||
# Discover non-default Studio roots from env vars + studio.conf files.
|
||||
# Discover non-default Unsloth roots from env vars + studio.conf files.
|
||||
# Mirrors install.ps1's precedence: UNSLOTH_STUDIO_HOME wins, STUDIO_HOME
|
||||
# is ignored when both are set, so uninstalling install A doesn't also
|
||||
# delete install B if the user has a stale STUDIO_HOME pointing at B.
|
||||
|
|
@ -207,7 +207,7 @@ function Uninstall-UnslothStudio {
|
|||
|
||||
# Return $true iff the PID's image path lives under one of $KnownRoots.
|
||||
# Prevents killing an unrelated process that happens to listen on a stale
|
||||
# Studio port.
|
||||
# Unsloth port.
|
||||
function _PidUnderKnownRoot {
|
||||
param([int]$Pid_, [string[]]$KnownRoots)
|
||||
if (-not $KnownRoots -or $KnownRoots.Count -eq 0) { return $false }
|
||||
|
|
@ -223,8 +223,8 @@ function Uninstall-UnslothStudio {
|
|||
return $false
|
||||
}
|
||||
|
||||
# Stop a Studio backend whose port is recorded in <DataDir>\studio.port.
|
||||
# Only kills if the listening PID's exe path is under a known Studio root.
|
||||
# Stop an Unsloth backend whose port is recorded in <DataDir>\studio.port.
|
||||
# Only kills if the listening PID's exe path is under a known Unsloth root.
|
||||
function _StopByPortFile {
|
||||
param([string]$PortFile, [string[]]$KnownRoots)
|
||||
if (-not (Test-Path -LiteralPath $PortFile -PathType Leaf)) { return }
|
||||
|
|
@ -372,7 +372,7 @@ function Uninstall-UnslothStudio {
|
|||
continue
|
||||
}
|
||||
if (-not (_IsStudioRoot $r)) {
|
||||
_Substep "refusing to remove non-Studio path: $r" "Yellow"
|
||||
_Substep "refusing to remove non-Unsloth path: $r" "Yellow"
|
||||
continue
|
||||
}
|
||||
_RemovePath $r
|
||||
|
|
@ -436,7 +436,7 @@ function Uninstall-UnslothStudio {
|
|||
$entries = $rawPath -split ';'
|
||||
$kept = New-Object System.Collections.ArrayList
|
||||
$removedAny = $false
|
||||
# Only remove PATH entries that live inside a Studio root we
|
||||
# Only remove PATH entries that live inside an Unsloth root we
|
||||
# actually own (default or env-mode). A literal substring
|
||||
# match on `unsloth_studio` would clobber unrelated user
|
||||
# virtualenvs that happen to share the name.
|
||||
|
|
|
|||
|
|
@ -12,7 +12,7 @@
|
|||
|
||||
set -e
|
||||
|
||||
# Stop a Studio server via its PID file (written by install.sh's _spawn_terminal).
|
||||
# Stop an Unsloth server via its PID file (written by install.sh's _spawn_terminal).
|
||||
_kill_pid_file() {
|
||||
_pid_file="$1"
|
||||
[ -f "$_pid_file" ] || return 0
|
||||
|
|
@ -47,7 +47,7 @@ _pkill_studio() {
|
|||
command -v pkill >/dev/null 2>&1 || return 0
|
||||
|
||||
# Scope fallback patterns to the install roots we are removing so a
|
||||
# different Studio install (different UNSLOTH_STUDIO_HOME) is not touched.
|
||||
# different Unsloth install (different UNSLOTH_STUDIO_HOME) is not touched.
|
||||
_kill_roots="$HOME/.unsloth/studio"
|
||||
_roots_from_conf=$(_custom_studio_roots 2>/dev/null || true)
|
||||
[ -n "$_roots_from_conf" ] && _kill_roots="$_kill_roots
|
||||
|
|
@ -89,7 +89,7 @@ _remove_path() {
|
|||
fi
|
||||
}
|
||||
|
||||
# Accept as Studio root only if Studio sentinels exist (matches install.sh's
|
||||
# Accept as Unsloth root only if Unsloth sentinels exist (matches install.sh's
|
||||
# env-mode ownership guard at install.sh:1358-1361). A bare unsloth_studio/
|
||||
# directory is NOT enough -- require the install-time owner marker so a user
|
||||
# directory that happens to contain a folder named "unsloth_studio" is safe.
|
||||
|
|
@ -175,8 +175,8 @@ _custom_studio_roots() {
|
|||
_from_conf "$HOME/.local/share/unsloth/studio.conf"
|
||||
}
|
||||
|
||||
# Remove $HOME/.local/bin/unsloth only if it's a Studio-managed symlink.
|
||||
# Studio's install.sh writes this as a symlink into the studio venv
|
||||
# Remove $HOME/.local/bin/unsloth only if it's an Unsloth-managed symlink.
|
||||
# Unsloth's install.sh writes this as a symlink into the studio venv
|
||||
# (install.sh: `ln -sfn "$VENV_DIR/bin/unsloth" "$_shim_path"`). A
|
||||
# pip-installed `unsloth` CLI is a regular file — leave it alone to avoid
|
||||
# wiping an unrelated install.
|
||||
|
|
@ -206,7 +206,7 @@ _custom_studio_roots | while IFS= read -r _custom_root; do
|
|||
continue
|
||||
fi
|
||||
if ! _is_studio_root "$_custom_root"; then
|
||||
echo " refusing to remove non-Studio path: $_custom_root" >&2
|
||||
echo " refusing to remove non-Unsloth path: $_custom_root" >&2
|
||||
continue
|
||||
fi
|
||||
_remove_path "$_custom_root"
|
||||
|
|
@ -234,7 +234,7 @@ _remove_path "$HOME/.unsloth/rocm-smoketest"
|
|||
# Drop ~/.unsloth only if now empty (rmdir refuses non-empty, so user content is kept).
|
||||
rmdir "$HOME/.unsloth" 2>/dev/null || true
|
||||
_remove_path "$HOME/.local/share/unsloth"
|
||||
# CLI shim: only the symlink Studio created, never a pip-installed file.
|
||||
# CLI shim: only the symlink Unsloth created, never a pip-installed file.
|
||||
_remove_cli_shim
|
||||
|
||||
echo "Removing desktop shortcut and launcher lock..."
|
||||
|
|
|
|||
|
|
@ -1,10 +1,10 @@
|
|||
# Unsloth Studio MCP server
|
||||
|
||||
Studio can expose a local MCP server so an MCP client can inspect models and
|
||||
Unsloth can expose a local MCP server so an MCP client can inspect models and
|
||||
GPU state, validate recipes, start or stop training, inspect recipe output, and
|
||||
export a loaded model.
|
||||
|
||||
The server is disabled by default. Enable it for a local Studio process with:
|
||||
The server is disabled by default. Enable it for a local Unsloth process with:
|
||||
|
||||
```bash
|
||||
UNSLOTH_STUDIO_ENABLE_MCP=1 \
|
||||
|
|
@ -12,8 +12,8 @@ UNSLOTH_STUDIO_MCP_TOKEN='use-a-local-secret' \
|
|||
unsloth studio
|
||||
```
|
||||
|
||||
The endpoint is `http://127.0.0.1:8888/mcp/` when Studio uses its default port
|
||||
(a request to `/mcp` redirects to the canonical `/mcp/`). Use the actual Studio
|
||||
The endpoint is `http://127.0.0.1:8888/mcp/` when Unsloth uses its default port
|
||||
(a request to `/mcp` redirects to the canonical `/mcp/`). Use the actual Unsloth
|
||||
port when it is configured differently.
|
||||
|
||||
The high-impact tools are:
|
||||
|
|
@ -23,9 +23,9 @@ The high-impact tools are:
|
|||
- `validate_recipe`, `get_recipe_job_status`, and `get_recipe_job_dataset`
|
||||
- `load_checkpoint` and `export_gguf`
|
||||
|
||||
`start_training` accepts the same fields as the Studio `TrainingStartRequest`.
|
||||
`start_training` accepts the same fields as the Unsloth `TrainingStartRequest`.
|
||||
The request is validated by the existing Pydantic model before a subprocess is
|
||||
started. Export paths use the existing Studio validation as well.
|
||||
started. Export paths use the existing Unsloth validation as well.
|
||||
|
||||
The endpoint always requires `UNSLOTH_STUDIO_MCP_TOKEN` and checks an exact
|
||||
Bearer token for both HTTP and WebSocket connections. Keep it on localhost
|
||||
|
|
|
|||
|
|
@ -33,7 +33,7 @@
|
|||
"\n",
|
||||
"We are actively working on making Unsloth Studio install on Colab T4 GPUs faster.\n",
|
||||
"\n",
|
||||
"[Features](https://unsloth.ai/docs/new/unsloth-studio#features) • [Quickstart](https://unsloth.ai/docs/new/unsloth-studio/start) • [Data Recipes](https://unsloth.ai/docs/new/unsloth-studio/data-recipe) • [Studio Chat](https://unsloth.ai/docs/new/unsloth-studio/chat) • [Export](https://unsloth.ai/docs/new/unsloth-studio/export)"
|
||||
"[Features](https://unsloth.ai/docs/new/unsloth-studio#features) • [Quickstart](https://unsloth.ai/docs/new/unsloth-studio/start) • [Data Recipes](https://unsloth.ai/docs/new/unsloth-studio/data-recipe) • [Unsloth Chat](https://unsloth.ai/docs/new/unsloth-studio/chat) • [Export](https://unsloth.ai/docs/new/unsloth-studio/export)"
|
||||
]
|
||||
},
|
||||
{
|
||||
|
|
|
|||
|
|
@ -3,7 +3,7 @@
|
|||
Source: google/gemma-4-31B-it HF discussion/PR #118 (adds the preserve_thinking
|
||||
flag plus null-rendering, string-arguments validation, balanced turn tags, empty
|
||||
messages handling, and OpenAI image_url/input_audio aliases).
|
||||
Studio-local changes vs PR #118:
|
||||
Unsloth-local changes vs PR #118:
|
||||
1. preserve_thinking defaults to false (see SETUP block below).
|
||||
2. The empty "<|channel>thought\n<channel|>" block on enable_thinking=false is
|
||||
NOT emitted. Google ships a distinct template for E2B/E4B (google/gemma-4-E2B-it,
|
||||
|
|
|
|||
|
|
@ -3,7 +3,7 @@
|
|||
Source: google/gemma-4-31B-it HF discussion/PR #118 (adds the preserve_thinking
|
||||
flag plus null-rendering, string-arguments validation, balanced turn tags, empty
|
||||
messages handling, and OpenAI image_url/input_audio aliases).
|
||||
Studio-local change: preserve_thinking defaults to false (see SETUP block below).
|
||||
Unsloth-local change: preserve_thinking defaults to false (see SETUP block below).
|
||||
Applied to unsloth/gemma-4-*-GGUF models so the embedded GGUF template does not
|
||||
need re-downloading. Keep in sync with upstream if PR #118 changes.
|
||||
-#}
|
||||
|
|
|
|||
|
|
@ -148,7 +148,7 @@ async def authenticated_via_api_key(
|
|||
) -> bool:
|
||||
"""True when the caller used an sk-unsloth API key, not a UI session JWT.
|
||||
|
||||
Lets routes treat programmatic API callers differently from the Studio UI
|
||||
Lets routes treat programmatic API callers differently from the Unsloth UI
|
||||
(e.g. refuse a teardown the UI would allow).
|
||||
"""
|
||||
return bool(credentials and credentials.credentials.startswith(API_KEY_PREFIX))
|
||||
|
|
|
|||
|
|
@ -1,13 +1,13 @@
|
|||
# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
||||
|
||||
"""Auto-shutdown for an exposed first-run Studio whose admin password is unchanged.
|
||||
"""Auto-shutdown for an exposed first-run Unsloth whose admin password is unchanged.
|
||||
|
||||
On a fresh install the seeded bootstrap admin password stays a valid login
|
||||
credential until first login changes it. When the web UI is put on the network
|
||||
(``--secure`` / ``0.0.0.0``) and nobody completes that first-login change within
|
||||
a deadline, tear Studio down so a fresh, unconfigured instance does not stay
|
||||
publicly reachable indefinitely. If the password was changed, Studio keeps
|
||||
a deadline, tear Unsloth down so a fresh, unconfigured instance does not stay
|
||||
publicly reachable indefinitely. If the password was changed, Unsloth keeps
|
||||
running.
|
||||
|
||||
Scope: web UI launches only (never ``--api-only``, which authenticates by API
|
||||
|
|
@ -98,7 +98,7 @@ def enforce_bootstrap_password_deadline(
|
|||
) -> bool:
|
||||
"""Deadline handler: shut down iff the seeded admin password is still unchanged.
|
||||
|
||||
Returns True if it shut Studio down, False if it left it running (the
|
||||
Returns True if it shut Unsloth down, False if it left it running (the
|
||||
password was changed in time).
|
||||
"""
|
||||
try:
|
||||
|
|
@ -106,7 +106,7 @@ def enforce_bootstrap_password_deadline(
|
|||
except Exception:
|
||||
return False
|
||||
if not still_default:
|
||||
return False # password changed in time -> leave Studio running
|
||||
return False # password changed in time -> leave Unsloth running
|
||||
|
||||
message = (
|
||||
"\nUnsloth Studio was exposed on the network but its default admin "
|
||||
|
|
|
|||
|
|
@ -146,7 +146,7 @@ def get_connection() -> sqlite3.Connection:
|
|||
pass
|
||||
conn.row_factory = sqlite3.Row
|
||||
# WAL lets token reads run concurrently with refresh-token writes;
|
||||
# busy_timeout bounds lock waits. Matches the other Studio SQLite stores.
|
||||
# busy_timeout bounds lock waits. Matches the other Unsloth SQLite stores.
|
||||
# Set busy_timeout first: switching journal_mode needs a lock, so if a
|
||||
# refresh-token write already holds one, journal_mode=WAL raises SQLITE_BUSY;
|
||||
# with busy_timeout already in effect it waits instead of failing and leaving
|
||||
|
|
@ -305,8 +305,8 @@ def get_or_create_identity_secret() -> bytes:
|
|||
def compute_identity_proof(nonce: bytes, host: str, port: int) -> str:
|
||||
"""HMAC-SHA256 proof that the caller holds this install's identity secret,
|
||||
bound to the loopback address and port the connection landed on. A proof
|
||||
relayed from a Studio on a different address/port (a squatter proxying to the
|
||||
real one, e.g. localhost resolving to ::1 while Studio is on 127.0.0.1) was
|
||||
relayed from an Unsloth on a different address/port (a squatter proxying to the
|
||||
real one, e.g. localhost resolving to ::1 while Unsloth is on 127.0.0.1) was
|
||||
computed for that other endpoint and won't match the one the client dialed."""
|
||||
try:
|
||||
host = ipaddress.ip_address(host).compressed # normalise 127.0.0.1 / ::1 forms
|
||||
|
|
|
|||
|
|
@ -2,14 +2,14 @@
|
|||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
||||
|
||||
"""Interactive terminal prompt that forces a bootstrap password change before
|
||||
Studio is exposed on a public Cloudflare URL (``--secure`` / ``--cloudflare``).
|
||||
Unsloth is exposed on a public Cloudflare URL (``--secure`` / ``--cloudflare``).
|
||||
|
||||
Masked input echoes one ``*`` per keystroke (unlike ``getpass``). Works on
|
||||
Windows (``msvcrt``) and Linux/macOS (``termios``). All output goes to stderr so
|
||||
redirected stdout never swallows the prompt.
|
||||
|
||||
Mirrored for the CLI at ``unsloth_cli/commands/_password_prompt.py`` (the CLI
|
||||
cannot import the Studio backend package); keep the two in sync.
|
||||
cannot import the Unsloth backend package); keep the two in sync.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
|
@ -252,7 +252,7 @@ def prompt_for_password_change(
|
|||
out.flush()
|
||||
return True
|
||||
except (KeyboardInterrupt, EOFError):
|
||||
out.write("Password change aborted; not exposing Studio.\n")
|
||||
out.write("Password change aborted; not exposing Unsloth.\n")
|
||||
out.flush()
|
||||
return False
|
||||
|
||||
|
|
|
|||
|
|
@ -1,13 +1,13 @@
|
|||
# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
||||
|
||||
"""Free Cloudflare quick tunnel for Studio's 0.0.0.0 launches.
|
||||
"""Free Cloudflare quick tunnel for Unsloth's 0.0.0.0 launches.
|
||||
|
||||
The raw http://<ip>:<port> is often unreachable (https-vs-http, blocked ports,
|
||||
closed security groups); a cloudflared quick tunnel gives a free
|
||||
https://*.trycloudflare.com URL that works anywhere, with no account or domain.
|
||||
|
||||
Best-effort throughout: any failure collapses to "no URL" and Studio keeps
|
||||
Best-effort throughout: any failure collapses to "no URL" and Unsloth keeps
|
||||
running. Stdlib only (back-end imports are lazy) so it is safe to import early.
|
||||
"""
|
||||
|
||||
|
|
@ -95,7 +95,7 @@ def _cache_path() -> Optional[Path]:
|
|||
|
||||
|
||||
def find_cloudflared() -> Optional[str]:
|
||||
"""Locate an existing cloudflared: PATH first, then the Studio bin cache."""
|
||||
"""Locate an existing cloudflared: PATH first, then the Unsloth bin cache."""
|
||||
on_path = shutil.which("cloudflared")
|
||||
if on_path:
|
||||
return on_path
|
||||
|
|
@ -309,7 +309,7 @@ class CloudflareTunnel:
|
|||
pass
|
||||
|
||||
|
||||
# Single serving process per Studio launch, so one module-level tunnel handle is
|
||||
# Single serving process per Unsloth launch, so one module-level tunnel handle is
|
||||
# enough; the lock guards the start/stop/shutdown races.
|
||||
_active_tunnel: Optional[CloudflareTunnel] = None
|
||||
_active_lock = threading.Lock()
|
||||
|
|
|
|||
|
|
@ -129,7 +129,7 @@ def start_cloudflare_tunnel(port: int) -> "str | None":
|
|||
logger.warning(
|
||||
"Cloudflare link not started: the admin account still has its temporary "
|
||||
"bootstrap password, which is exposed to anyone who can load the page. "
|
||||
"Open Studio in this tab, log in and change the admin password, then re-run "
|
||||
"Open Unsloth in this tab, log in and change the admin password, then re-run "
|
||||
"start(cloudflare=True) to get the shareable link."
|
||||
)
|
||||
return None
|
||||
|
|
@ -203,7 +203,7 @@ def _shareable_link_html(cloudflare_url: str) -> str:
|
|||
display: flex; align-items: center; gap: 12px;">
|
||||
<img src="https://github.com/unslothai/unsloth/raw/main/studio/frontend/public/unsloth-gem.png"
|
||||
height="48" style="display:block;">
|
||||
Shareable Studio Link is Ready!
|
||||
Shareable Unsloth Link is Ready!
|
||||
</h2>
|
||||
<a href="{cloudflare_url}" onclick="var w=window.open(this.href,'_blank');if(!w){{return true;}}return false;"
|
||||
style="display: inline-flex; align-items: center; gap: 10px; padding: 14px 28px;
|
||||
|
|
@ -223,7 +223,7 @@ def _shareable_link_html(cloudflare_url: str) -> str:
|
|||
|
||||
|
||||
def _show_and_embed(port: int, *, cloudflare_url: "str | None" = None):
|
||||
"""Render the Studio header + iframe for *port*, with a shareable-link card above
|
||||
"""Render the Unsloth header + iframe for *port*, with a shareable-link card above
|
||||
when *cloudflare_url* is set. Falls back to serve_kernel_port_as_iframe."""
|
||||
url = get_colab_url(port)
|
||||
logger.info(f"🌐 Unsloth Studio URL: {url}")
|
||||
|
|
@ -281,7 +281,7 @@ def start(port: int = 8888, *, cloudflare: bool = False):
|
|||
Args:
|
||||
port: Port to bind/serve on.
|
||||
cloudflare: Opt in to a shareable Cloudflare HTTPS link reachable from any
|
||||
device (default OFF). It exposes Studio's login page beyond Colab, so it
|
||||
device (default OFF). It exposes Unsloth's login page beyond Colab, so it
|
||||
stays an explicit opt-in; the default shows only the in-tab proxy iframe.
|
||||
|
||||
Usage:
|
||||
|
|
@ -292,10 +292,10 @@ def start(port: int = 8888, *, cloudflare: bool = False):
|
|||
|
||||
logger.info("🦥 Starting Unsloth Studio...")
|
||||
|
||||
# Fast path: Studio already running (cell re-run). Re-launching would collide on
|
||||
# Fast path: Unsloth already running (cell re-run). Re-launching would collide on
|
||||
# the port, so just re-show the link and iframe.
|
||||
if _is_studio_healthy(port):
|
||||
logger.info(f" Studio is already running on port {port} — reusing existing server.")
|
||||
logger.info(f" Unsloth is already running on port {port} — reusing existing server.")
|
||||
# try/finally: tear the tunnel down even if interrupted mid-start/render.
|
||||
try:
|
||||
cf_url = start_cloudflare_tunnel(port) if cloudflare else None
|
||||
|
|
|
|||
|
|
@ -133,7 +133,7 @@ def parse_log_message(msg: str) -> ParsedUpdate | None:
|
|||
source = "github",
|
||||
status = "rate_limited",
|
||||
retry_after_sec = seconds,
|
||||
message = ("Waiting for GitHub rate limit. Studio will resume automatically."),
|
||||
message = ("Waiting for GitHub rate limit. Unsloth will resume automatically."),
|
||||
),
|
||||
)
|
||||
|
||||
|
|
@ -147,7 +147,7 @@ def parse_log_message(msg: str) -> ParsedUpdate | None:
|
|||
status = "rate_limited",
|
||||
retry_after_sec = seconds,
|
||||
message = (
|
||||
"Waiting for GitHub secondary rate limit. Studio will resume automatically."
|
||||
"Waiting for GitHub secondary rate limit. Unsloth will resume automatically."
|
||||
),
|
||||
),
|
||||
)
|
||||
|
|
@ -161,7 +161,7 @@ def parse_log_message(msg: str) -> ParsedUpdate | None:
|
|||
source = "github",
|
||||
status = "rate_limited",
|
||||
retry_after_sec = seconds,
|
||||
message = ("Waiting for GitHub rate limit. Studio will resume automatically."),
|
||||
message = ("Waiting for GitHub rate limit. Unsloth will resume automatically."),
|
||||
),
|
||||
)
|
||||
|
||||
|
|
|
|||
|
|
@ -238,7 +238,7 @@ def _run_oxc_batch(
|
|||
if not node_executable:
|
||||
return _fallback_results(
|
||||
len(code_values),
|
||||
"Node.js not found (install Node >= 20.19, or re-run Studio setup to provision it).",
|
||||
"Node.js not found (install Node >= 20.19, or re-run Unsloth setup to provision it).",
|
||||
)
|
||||
try:
|
||||
tmp_dir = ensure_dir(oxc_validator_tmp_root())
|
||||
|
|
|
|||
|
|
@ -280,8 +280,8 @@ def create_data_designer(recipe: dict[str, Any], *, artifact_path: str | None =
|
|||
from data_designer.interface.data_designer import DataDesigner # pyright: ignore[reportMissingImports]
|
||||
|
||||
if artifact_path is None:
|
||||
# DataDesigner defaults to cwd/artifacts; packaged Studio can run with
|
||||
# cwd=/, so keep default callers on Studio's writable recipe artifact root.
|
||||
# DataDesigner defaults to cwd/artifacts; packaged Unsloth can run with
|
||||
# cwd=/, so keep default callers on Unsloth's writable recipe artifact root.
|
||||
artifact_path = str(recipe_datasets_root())
|
||||
|
||||
recipe = _strip_frontend_model_config_metadata(recipe)
|
||||
|
|
|
|||
|
|
@ -11,7 +11,7 @@ subprocess and can be imported directly from .inference when needed.
|
|||
Public names are resolved lazily (PEP 562): importing this package -- or a
|
||||
dependency-light leaf like ``core.inference.chat_eos`` -- must NOT eagerly pull
|
||||
the orchestrator / llama_cpp import chain (httpx, subprocess plumbing, the ML
|
||||
backend and its Studio dependencies). Those load only when a public name is
|
||||
backend and its Unsloth dependencies). Those load only when a public name is
|
||||
actually accessed, so standalone helpers stay unit-testable without the full
|
||||
inference stack.
|
||||
"""
|
||||
|
|
|
|||
|
|
@ -539,7 +539,7 @@ class AnthropicPassthroughEmitter:
|
|||
|
||||
Only calls naming a tool in ``allowed_tools`` (the client's declared
|
||||
tools) are promoted; everything else streams as text exactly as before.
|
||||
Never enabled for Studio's own tool loop.
|
||||
Never enabled for Unsloth's own tool loop.
|
||||
"""
|
||||
from core.inference.passthrough_healing import StreamToolCallHealer
|
||||
|
||||
|
|
|
|||
|
|
@ -150,7 +150,7 @@ def _split_partial_marker(text: str, marker: str) -> tuple[str, str]:
|
|||
class ReasoningChannelNormalizer:
|
||||
"""Incrementally convert one native reasoning channel to ``<think>``.
|
||||
|
||||
The parser follows mlx-vlm's streaming boundary behavior but emits Studio's
|
||||
The parser follows mlx-vlm's streaming boundary behavior but emits Unsloth's
|
||||
established canonical text contract. Only the configured opening and
|
||||
closing markers are consumed; tool-call and other control markers remain
|
||||
available to downstream parsers.
|
||||
|
|
|
|||
|
|
@ -4,13 +4,13 @@
|
|||
"""Bundled chat-template selection for GGUF inference.
|
||||
|
||||
Some shipped GGUF quants embed an older chat template. Rather than re-cutting and
|
||||
asking users to re-download every quant, Studio can override the embedded template
|
||||
asking users to re-download every quant, Unsloth can override the embedded template
|
||||
at llama-server launch time with a bundled, up-to-date Jinja template for known
|
||||
model families. The override is wired through the existing ``chat_template_override``
|
||||
-> ``--chat-template-file`` path in ``LlamaCppBackend.load_model``.
|
||||
|
||||
Currently this covers ``unsloth/gemma-4-*-GGUF``, which gains the upstream PR #118
|
||||
``preserve_thinking`` flag (defaulted OFF here) so the Studio "Preserve thinking"
|
||||
``preserve_thinking`` flag (defaulted OFF here) so the Unsloth "Preserve thinking"
|
||||
toggle appears while staying disabled by default.
|
||||
"""
|
||||
|
||||
|
|
|
|||
|
|
@ -473,7 +473,7 @@ def _apply_mistral_reasoning_controls(
|
|||
# handles every provider without storing credentials.
|
||||
def _create_shared_http_client() -> httpx.AsyncClient:
|
||||
# Unsupported env proxy schemes (socks:// etc) raise at construction and
|
||||
# would crash Studio startup (#6090); retry ignoring env proxies instead.
|
||||
# would crash Unsloth startup (#6090); retry ignoring env proxies instead.
|
||||
try:
|
||||
return httpx.AsyncClient()
|
||||
except (ImportError, ValueError) as exc:
|
||||
|
|
@ -858,7 +858,7 @@ class ExternalProviderClient:
|
|||
if not self._is_openai_compatible():
|
||||
# Gemini speaks its own native REST shape (contents/parts);
|
||||
# `_stream_gemini` translates request/response into the OpenAI
|
||||
# Chat Completions chunk format the rest of Studio expects.
|
||||
# Chat Completions chunk format the rest of Unsloth expects.
|
||||
# API ref: https://ai.google.dev/gemini-api/docs
|
||||
if self.provider_type == "gemini":
|
||||
async for line in self._stream_gemini(
|
||||
|
|
@ -1706,7 +1706,7 @@ class ExternalProviderClient:
|
|||
# Translate OpenAI multimodal parts -> Anthropic native shapes.
|
||||
# - `image_url` -> `{type:"image", source:...}`
|
||||
# - `input_document` -> `{type:"document", source:...}`
|
||||
# (Studio extension; mirrors Anthropic's document block,
|
||||
# (Unsloth extension; mirrors Anthropic's document block,
|
||||
# which supports PDFs as base64 or URL per
|
||||
# https://platform.claude.com/docs/en/build-with-claude/vision)
|
||||
anthropic_parts: list[dict[str, Any]] = []
|
||||
|
|
@ -1749,7 +1749,7 @@ class ExternalProviderClient:
|
|||
}
|
||||
)
|
||||
elif part.get("type") == "input_document":
|
||||
# Studio's normalised PDF/doc type (file_data data-URI or
|
||||
# Unsloth's normalised PDF/doc type (file_data data-URI or
|
||||
# file_url) -> Anthropic's native `document` block.
|
||||
url = part.get("file_url") or ""
|
||||
data_uri = part.get("file_data") or ""
|
||||
|
|
@ -4704,7 +4704,7 @@ class ExternalProviderClient:
|
|||
{"type": "image_generation_call", "id": call_id}
|
||||
)
|
||||
elif part_type == "input_document":
|
||||
# Map Studio's `input_document` onto Responses' `input_file`.
|
||||
# Map Unsloth's `input_document` onto Responses' `input_file`.
|
||||
# https://developers.openai.com/api/docs/guides/images-vision
|
||||
file_url = part.get("file_url")
|
||||
file_data = part.get("file_data")
|
||||
|
|
@ -6010,7 +6010,7 @@ class ExternalProviderClient:
|
|||
if not models and self.provider_type == "ollama":
|
||||
models = await self._list_ollama_native_models()
|
||||
# Gemini's native /v1beta/models uses a different shape; repackage
|
||||
# into the OpenAI-compatible one Studio expects.
|
||||
# into the OpenAI-compatible one Unsloth expects.
|
||||
if not models and self.provider_type == "gemini":
|
||||
models = self._parse_gemini_models(data)
|
||||
return models
|
||||
|
|
@ -6213,7 +6213,7 @@ def _friendly_provider_error_text(
|
|||
*,
|
||||
model: str | None = None,
|
||||
) -> str:
|
||||
"""Rewrite common provider errors into actionable Studio copy."""
|
||||
"""Rewrite common provider errors into actionable Unsloth copy."""
|
||||
if status_code == 404 and model:
|
||||
lowered = raw_message.lower()
|
||||
if "not found" in lowered or "not_found" in lowered:
|
||||
|
|
|
|||
|
|
@ -116,7 +116,7 @@ LLAMA_SERVER_NOT_FOUND_DETAIL = (
|
|||
|
||||
# llama-server can serve HTTP 200 while running a model entirely on CPU when a
|
||||
# GPU backend fails to init (#5807 / #5106 / #5830). Classify the startup log so
|
||||
# Studio can warn. Priority: explicit "offloaded N/M layers to GPU" counts
|
||||
# Unsloth can warn. Priority: explicit "offloaded N/M layers to GPU" counts
|
||||
# (authoritative), then GPU "model buffer size" lines (host-pinned _Host
|
||||
# excluded), then the "device_info:" device table (disconfirm only).
|
||||
_GPU_OFFLOAD_MARKERS = (
|
||||
|
|
@ -1363,7 +1363,7 @@ def _kv_bytes_per_elem(cache_type: Optional[str]) -> float:
|
|||
|
||||
def _env_main_cache_type_for_budget(env: Optional[Mapping[str, str]] = None) -> Optional[str]:
|
||||
"""Heavier of the inherited LLAMA_ARG_CACHE_TYPE_K/_V env types when it
|
||||
exceeds the f16 default, else None. Studio emits --cache-type only for the
|
||||
exceeds the f16 default, else None. Unsloth emits --cache-type only for the
|
||||
param/extras path, so a heavier env (f32) would otherwise reach the child
|
||||
unbudgeted; quantized env types stay over-reserved by f16 (-> None)."""
|
||||
e = os.environ if env is None else env
|
||||
|
|
@ -1682,7 +1682,7 @@ def _build_ngram_mod_flags(
|
|||
return []
|
||||
|
||||
|
||||
# Canonical Speculative Decoding modes exposed by the Studio chat UI.
|
||||
# Canonical Speculative Decoding modes exposed by the Unsloth chat UI.
|
||||
# Dropdown renders five (auto, mtp, ngram, mtp+ngram, off); the load API
|
||||
# also accepts legacy values the original Switch and external callers emit
|
||||
# (default, draft-mtp, ngram-mod, ngram-simple).
|
||||
|
|
@ -1731,7 +1731,7 @@ def _backfill_usage_from_timings(usage, timings):
|
|||
"""Synthesize ``usage`` from llama-server's ``timings`` when the
|
||||
OpenAI-style usage block is missing or reports zero tokens.
|
||||
|
||||
The Studio chat UI computes generation t/s from
|
||||
The Unsloth chat UI computes generation t/s from
|
||||
``meta.usage.completion_tokens / totalStreamTime``. llama-server always
|
||||
populates ``timings.predicted_n`` (true decoded count) and
|
||||
``timings.prompt_n``, but the final SSE chunk's ``usage`` can be absent
|
||||
|
|
@ -1804,7 +1804,7 @@ def _llama_lib_dir(binary: str) -> Path:
|
|||
def _is_external_link(path: Path) -> bool:
|
||||
"""True when ``path`` is a --with-llama-cpp-dir local link: a POSIX symlink
|
||||
or a Windows directory junction / reparse point. Such a link resolves into
|
||||
the user's own llama.cpp checkout, which Studio does not own."""
|
||||
the user's own llama.cpp checkout, which Unsloth does not own."""
|
||||
try:
|
||||
if os.path.islink(path):
|
||||
return True
|
||||
|
|
@ -1960,7 +1960,7 @@ class LlamaCppBackend:
|
|||
# observes it (direct proxy endpoints, or nothing in flight).
|
||||
self._mtp_watchdog_thread: Optional[threading.Thread] = None
|
||||
self._mtp_watchdog_stop = threading.Event()
|
||||
# True when the launch actually runs MTP+tensor (Studio- or user/env-driven);
|
||||
# True when the launch actually runs MTP+tensor (Unsloth- or user/env-driven);
|
||||
# gates the probe, watchdog, and recovery so pass-through MTP is covered.
|
||||
self._mtp_runtime_fallback_active = False
|
||||
self._stdout_lines: list[str] = []
|
||||
|
|
@ -2353,7 +2353,7 @@ class LlamaCppBackend:
|
|||
|
||||
@staticmethod
|
||||
def _resolved_studio_root_and_is_legacy() -> "tuple[Optional[Path], bool]":
|
||||
"""Resolve the Studio install root and classify it as the legacy
|
||||
"""Resolve the Unsloth install root and classify it as the legacy
|
||||
~/.unsloth/studio root vs. a custom (env/venv-inferred) root.
|
||||
|
||||
Returns (resolved_root, is_legacy). On any import/resolution failure the
|
||||
|
|
@ -3241,7 +3241,7 @@ class LlamaCppBackend:
|
|||
return
|
||||
prev = curr
|
||||
|
||||
# Free-VRAM fraction at which Studio pins the GPU directly instead of
|
||||
# Free-VRAM fraction at which Unsloth pins the GPU directly instead of
|
||||
# deferring to ``--fit on``. 3% headroom: the compute buffer is now modelled in
|
||||
# the fit, so this only guards fragmentation + multi-GPU per-device CUDA context
|
||||
# (~2-3%); kept >= 3% as a floor (0.90 dropped 91-94% fits to CPU offload, #5106).
|
||||
|
|
@ -3800,7 +3800,7 @@ class LlamaCppBackend:
|
|||
return total if total > 0 else None
|
||||
return draft_kv + weights + target_ctx_copy
|
||||
|
||||
_DEFAULT_N_UBATCH = 512 # llama.cpp --ubatch default; Studio does not override it
|
||||
_DEFAULT_N_UBATCH = 512 # llama.cpp --ubatch default; Unsloth does not override it
|
||||
_COMPUTE_BUFFER_SAFETY = 1.15 # upper-bound margin on the compute-buffer estimate
|
||||
# Soft VRAM the modeled terms omit; charged to the fit budget on tight tiers (#6682).
|
||||
_CUDA_CONTEXT_RESERVE_BYTES = 320 * 1024 * 1024 # CUDA ctx + cuBLAS workspace (~330 MiB)
|
||||
|
|
@ -3940,7 +3940,7 @@ class LlamaCppBackend:
|
|||
n_ubatch: Optional[int] = None,
|
||||
) -> tuple[Optional[list[int]], bool, int]:
|
||||
"""Largest serving-slot count in [1, n_parallel) whose fully-on-GPU footprint fits,
|
||||
so Studio keeps the model on GPU (-ngl -1) instead of --fit on, which offloads layers
|
||||
so Unsloth keeps the model on GPU (-ngl -1) instead of --fit on, which offloads layers
|
||||
to host and collapses decode ~3x (oobabooga #6718). ``base_footprint_bytes`` is the
|
||||
slot-independent footprint (weights + soft overhead + MTP + context-linear compute,
|
||||
minus the folded compute buffer); each candidate re-adds the slot-sized compute buffer
|
||||
|
|
@ -4416,7 +4416,7 @@ class LlamaCppBackend:
|
|||
]
|
||||
|
||||
# Otherwise hand off to the resolver (cache / bootstrap / transformers / HF). Diffusion models
|
||||
# skip it: they do not use Studio's SWA pattern and the resolver can raise for them.
|
||||
# skip it: they do not use Unsloth's SWA pattern and the resolver can raise for them.
|
||||
if (
|
||||
self._sliding_window_pattern is None
|
||||
and self._sliding_window
|
||||
|
|
@ -4536,7 +4536,7 @@ class LlamaCppBackend:
|
|||
) -> bool:
|
||||
"""Launch the OpenAI-compat diffusion shim (which drives the on-device
|
||||
visual decoder) and wait for health. Presents the same /v1 + /health
|
||||
interface as llama-server, so the rest of Studio is unchanged.
|
||||
interface as llama-server, so the rest of Unsloth is unchanged.
|
||||
"""
|
||||
assets = self._find_diffusion_assets()
|
||||
if assets is None:
|
||||
|
|
@ -4608,7 +4608,7 @@ class LlamaCppBackend:
|
|||
logger.debug(f"Could not open diffusion runner log file: {e}")
|
||||
|
||||
# The shim (and its visual server) die with this backend process, so a
|
||||
# Studio crash/restart never orphans a GPU process.
|
||||
# Unsloth crash/restart never orphans a GPU process.
|
||||
self._process = subprocess.Popen(
|
||||
cmd,
|
||||
stdout = subprocess.PIPE,
|
||||
|
|
@ -5242,7 +5242,7 @@ class LlamaCppBackend:
|
|||
return (
|
||||
f"'{arch}' is a diffusion (image-generation) GGUF, which "
|
||||
"llama-server cannot run as a chat/completion model. Use "
|
||||
"Studio's Images page to generate with local diffusion "
|
||||
"Unsloth's Images page to generate with local diffusion "
|
||||
"GGUFs such as FLUX and Qwen-Image."
|
||||
)
|
||||
if is_ollama:
|
||||
|
|
@ -6103,7 +6103,7 @@ class LlamaCppBackend:
|
|||
and not bool(mtp_draft_path)
|
||||
)
|
||||
# LLAMA_ARG_SPEC_TYPE only reaches the child when neither extras
|
||||
# nor Studio emit a spec flag (mode "off", no user --spec-type),
|
||||
# nor Unsloth emit a spec flag (mode "off", no user --spec-type),
|
||||
# since _build_speculative_flags emits one for every other mode.
|
||||
# Consult the env for the reserve only then, else a stale MTP env
|
||||
# would over-reserve.
|
||||
|
|
@ -6112,7 +6112,7 @@ class LlamaCppBackend:
|
|||
if (not _extra_args_set_spec_type(extra_args) and _mtp_canonical == "off")
|
||||
else {}
|
||||
)
|
||||
# Extras can run MTP even when Studio suppresses its own emission.
|
||||
# Extras can run MTP even when Unsloth suppresses its own emission.
|
||||
_user_mtp_via_extras = _extra_args_requests_mtp(extra_args, env = _spec_env)
|
||||
# A non-MTP model-based draft mode (draft-simple/draft-eagle3) in
|
||||
# extras also loads a separate draft model that needs reserving;
|
||||
|
|
@ -6178,7 +6178,7 @@ class LlamaCppBackend:
|
|||
_mtp_eff_n_max = 2 if gpus else 3
|
||||
# Separate-drafter weights live on GPU (an embedded head is
|
||||
# already in model_size). Size the drafter the launch loads, by
|
||||
# precedence: extras --model-draft (last-wins), else Studio's
|
||||
# precedence: extras --model-draft (last-wins), else Unsloth's
|
||||
# emitted mtp_draft_path, else the env drafter. Sizing the wrong
|
||||
# one would under-reserve and OOM.
|
||||
_cli_draft_for_budget = _extra_args_mtp_draft_path(extra_args, env = {})
|
||||
|
|
@ -7097,12 +7097,12 @@ class LlamaCppBackend:
|
|||
|
||||
# Vulkan pins via --device (a cmd arg, unlike the env-based
|
||||
# CUDA/ROCm pin below), emitted BEFORE user extras so llama.cpp's
|
||||
# last-wins parsing lets a user --device override Studio's pick.
|
||||
# last-wins parsing lets a user --device override Unsloth's pick.
|
||||
if is_vulkan_backend and gpu_indices is not None:
|
||||
cmd += LlamaCppBackend._vulkan_pin_args(gpu_indices)
|
||||
|
||||
# User pass-through args go last so llama.cpp's last-wins parsing
|
||||
# lets the user override Studio's auto-set flags. Already
|
||||
# lets the user override Unsloth's auto-set flags. Already
|
||||
# validated by the route via validate_extra_args().
|
||||
if extra_args:
|
||||
cmd.extend(str(a) for a in extra_args)
|
||||
|
|
@ -7118,9 +7118,9 @@ class LlamaCppBackend:
|
|||
if "--threads" not in cmd:
|
||||
env.pop("LLAMA_ARG_THREADS", None)
|
||||
|
||||
# Reconcile the inherited LLAMA_ARG_* env with Studio's final
|
||||
# Reconcile the inherited LLAMA_ARG_* env with Unsloth's final
|
||||
# decision: stripping CLI extras on a tensor->layer downgrade
|
||||
# can't remove env vars, so the child could run a mode/KV Studio
|
||||
# can't remove env vars, so the child could run a mode/KV Unsloth
|
||||
# didn't budget.
|
||||
if not tensor_parallel:
|
||||
# Layer split: clear a non-layer inherited split mode (and any
|
||||
|
|
@ -7130,7 +7130,7 @@ class LlamaCppBackend:
|
|||
env.pop("LLAMA_ARG_SPLIT_MODE", None)
|
||||
env.pop("LLAMA_ARG_TENSOR_SPLIT", None)
|
||||
else:
|
||||
# Studio owns the tensor split: it emits --tensor-split when it
|
||||
# Unsloth owns the tensor split: it emits --tensor-split when it
|
||||
# picks an uneven one (CLI wins) and nothing when an even split
|
||||
# is safe. Clear any inherited LLAMA_ARG_TENSOR_SPLIT so the even
|
||||
# case can't be overridden by a stale env (the layer branch above
|
||||
|
|
@ -7201,7 +7201,7 @@ class LlamaCppBackend:
|
|||
# 'on') even when -ngl is explicit. That step has aborted on
|
||||
# some ROCm hosts (ggml-cuda.cu ROCm error during worst-case
|
||||
# estimation, e.g. MTP + mmproj models on gfx1151). When
|
||||
# Studio's own VRAM math already placed the model
|
||||
# Unsloth's own VRAM math already placed the model
|
||||
# (use_fit=False), the step is redundant second-guessing --
|
||||
# retry once with --fit off before declaring the load failed.
|
||||
# Never retry when fit was requested (use_fit) or the caller
|
||||
|
|
@ -7284,7 +7284,7 @@ class LlamaCppBackend:
|
|||
and _startup_crashed
|
||||
and not _split_axis_crash
|
||||
):
|
||||
# We forced --fit off because Studio's (conservative) VRAM
|
||||
# We forced --fit off because Unsloth's (conservative) VRAM
|
||||
# math placed the model fully on GPU. A startup crash here
|
||||
# means that estimate was optimistic, so fall back to --fit
|
||||
# on and let llama.cpp offload rather than fail the load.
|
||||
|
|
@ -7296,7 +7296,7 @@ class LlamaCppBackend:
|
|||
self._process.returncode,
|
||||
self._llama_log_path,
|
||||
)
|
||||
# Flip Studio's own --fit off (added first, before any
|
||||
# Flip Unsloth's own --fit off (added first, before any
|
||||
# user extra args) to on; a user's later --fit still wins
|
||||
# by last-arg. Defensive: if absent, the default is already
|
||||
# --fit on, so leave it.
|
||||
|
|
@ -7313,7 +7313,7 @@ class LlamaCppBackend:
|
|||
):
|
||||
logger.warning(
|
||||
"llama-server crashed during startup (exit code %s) "
|
||||
"with the default memory-fit step enabled; Studio "
|
||||
"with the default memory-fit step enabled; Unsloth "
|
||||
"already verified the model fits, retrying once "
|
||||
"with --fit off. Crash log: %s",
|
||||
self._process.returncode,
|
||||
|
|
@ -7393,7 +7393,7 @@ class LlamaCppBackend:
|
|||
cmd = _fa_cmd
|
||||
healthy = _spawn_and_wait(_fa_cmd, label = "-noflash")
|
||||
|
||||
# MTP from Studio's spec flags or the user's (extra_args
|
||||
# MTP from Unsloth's spec flags or the user's (extra_args
|
||||
# --spec-type / LLAMA_ARG_SPEC_TYPE). The env reaches the child
|
||||
# only when neither emits a spec flag, so consult it only then.
|
||||
_launch_spec_env: Mapping[str, str] = (
|
||||
|
|
@ -7587,11 +7587,11 @@ class LlamaCppBackend:
|
|||
if self._gpu_offload_active is False:
|
||||
logger.warning(
|
||||
"llama-server appears to have loaded the model entirely "
|
||||
"on CPU even though Studio detected at least one GPU. "
|
||||
"on CPU even though Unsloth detected at least one GPU. "
|
||||
"This usually means the prebuilt binary's GPU backend "
|
||||
"failed to load -- on Windows, cudart64_X.dll / "
|
||||
"cublas64_X.dll could not be resolved. Reinstall the "
|
||||
"Studio llama.cpp prebuilt or install a matching CUDA "
|
||||
"Unsloth llama.cpp prebuilt or install a matching CUDA "
|
||||
"toolkit (issue unslothai/unsloth#5106).",
|
||||
)
|
||||
|
||||
|
|
@ -7888,7 +7888,7 @@ class LlamaCppBackend:
|
|||
logger.info(
|
||||
"Auto: MLA embedded-MTP model detected; llama.cpp's MLA/DSA "
|
||||
"MTP path is slower than no speculation, so using ngram-mod "
|
||||
"instead. Override via the Studio Speculative Decoding "
|
||||
"instead. Override via the Unsloth Speculative Decoding "
|
||||
"dropdown or UNSLOTH_MLA_MTP_ENABLED=1."
|
||||
)
|
||||
_emit_ngram_mod()
|
||||
|
|
@ -7916,7 +7916,7 @@ class LlamaCppBackend:
|
|||
f"MTP GGUF detected but model size {_mtp_size_b:.1f}B "
|
||||
"is below the 3B speedup threshold; using ngram-mod "
|
||||
"only (zero-VRAM, no draft head). Override via "
|
||||
"--spec-type or the Studio Speculative Decoding "
|
||||
"--spec-type or the Unsloth Speculative Decoding "
|
||||
"dropdown."
|
||||
)
|
||||
_emit_ngram_mod()
|
||||
|
|
@ -8294,7 +8294,7 @@ class LlamaCppBackend:
|
|||
def _pid_parent_is_alive(pid: int) -> bool:
|
||||
"""True if the recorded server's parent is still running, i.e. the server is
|
||||
NOT orphaned. Lets the cross-session reap kill only a true orphan (parent
|
||||
gone) and never a live server owned by a running Studio, regardless of which
|
||||
gone) and never a live server owned by a running Unsloth, regardless of which
|
||||
process performs the sweep. Biased toward "alive" on uncertainty so a live
|
||||
server is never mistakenly reaped."""
|
||||
try:
|
||||
|
|
@ -8334,9 +8334,9 @@ class LlamaCppBackend:
|
|||
@classmethod
|
||||
def _reap_recorded_pid(cls) -> int:
|
||||
"""Kill the exact llama-server PID recorded at spawn, but only when it is a
|
||||
genuine orphan -- its parent (the Studio that spawned it) is gone. This is
|
||||
genuine orphan -- its parent (the Unsloth that spawned it) is gone. This is
|
||||
the cross-session backstop the parent-death reaper (Job Object /
|
||||
PR_SET_PDEATHSIG) cannot cover: an orphan left by an already-dead Studio
|
||||
PR_SET_PDEATHSIG) cannot cover: an orphan left by an already-dead Unsloth
|
||||
(macOS, a best-effort failure, or a pre-existing orphan). Path-independent,
|
||||
so it also catches an orphan the install-root match would miss.
|
||||
|
||||
|
|
@ -8393,7 +8393,7 @@ class LlamaCppBackend:
|
|||
"""Kill orphaned llama-server processes started by studio.
|
||||
|
||||
Only kills processes whose resolved binary lives under a known
|
||||
Studio install dir (or matches an exact env-var override), to avoid
|
||||
Unsloth install dir (or matches an exact env-var override), to avoid
|
||||
terminating unrelated llama-server instances. Mirrors every location
|
||||
_find_llama_server_binary() can return, so orphans from any
|
||||
supported install path are cleaned up.
|
||||
|
|
@ -8413,7 +8413,7 @@ class LlamaCppBackend:
|
|||
try:
|
||||
# -- Build the ownership allowlist --------------------------------
|
||||
# exact_binaries -- env var overrides (exact path match).
|
||||
# install_roots -- Studio-owned dir trees (binary must be under one).
|
||||
# install_roots -- Unsloth-owned dir trees (binary must be under one).
|
||||
install_roots: list[Path] = []
|
||||
|
||||
# Env-mode custom root (mirrors _find_llama_server_binary).
|
||||
|
|
@ -8423,7 +8423,7 @@ class LlamaCppBackend:
|
|||
install_roots.append(_resolved_sr / "llama.cpp")
|
||||
|
||||
# Primary install dir (default mode only). Env-mode skips this so a
|
||||
# custom-root Studio can't kill a default-install Studio's server.
|
||||
# custom-root Unsloth can't kill a default-install Unsloth's server.
|
||||
if not _is_custom_root:
|
||||
install_roots.append(Path.home() / ".unsloth" / "llama.cpp")
|
||||
|
||||
|
|
@ -8497,7 +8497,7 @@ class LlamaCppBackend:
|
|||
if not is_ours:
|
||||
continue
|
||||
|
||||
# A live parent means a running Studio (or the user's
|
||||
# A live parent means a running Unsloth (or the user's
|
||||
# shell) still owns it -- not an orphan.
|
||||
if LlamaCppBackend._pid_parent_is_alive(proc.info["pid"]):
|
||||
continue
|
||||
|
|
@ -8577,7 +8577,7 @@ class LlamaCppBackend:
|
|||
def _fit_off_retry_eligible(cmd: "list[str]", use_fit: bool) -> bool:
|
||||
"""Whether a llama-server startup crash may be retried with --fit off.
|
||||
|
||||
Only when Studio's own VRAM math placed the model (use_fit=False)
|
||||
Only when Unsloth's own VRAM math placed the model (use_fit=False)
|
||||
and nothing on the command line set the fit mode explicitly
|
||||
(-fit / --fit, space- or equals-form). --fit-ctx / --fit-target /
|
||||
-fitc / -fitt tune the fit step but do not select the mode, so
|
||||
|
|
@ -8821,7 +8821,7 @@ class LlamaCppBackend:
|
|||
return None
|
||||
|
||||
def _reconcile_effective_ctx_with_server(self) -> None:
|
||||
"""Adopt the server's real ``n_ctx`` when it is below Studio's value.
|
||||
"""Adopt the server's real ``n_ctx`` when it is below Unsloth's value.
|
||||
|
||||
Keeps ``context_length`` (load response, status route, passthrough
|
||||
``max_tokens`` ceiling) honest; clients sized to the requested value
|
||||
|
|
|
|||
|
|
@ -59,7 +59,7 @@ _INFERENCE_SUFFIXES = (
|
|||
"/messages/count_tokens", # counts via the loaded tokenizer; protect like /messages
|
||||
"/embeddings",
|
||||
"/responses",
|
||||
"/generate/stream", # Studio's own streaming route on the same llama-server
|
||||
"/generate/stream", # Unsloth's own streaming route on the same llama-server
|
||||
"/audio/generate", # direct GGUF TTS; can outlive the idle TTL
|
||||
)
|
||||
|
||||
|
|
|
|||
|
|
@ -3,10 +3,10 @@
|
|||
|
||||
"""Boundary validator for user-supplied llama-server pass-through args.
|
||||
|
||||
Reject only flags Studio manages (model identity, auth, network, parallel
|
||||
Reject only flags Unsloth manages (model identity, auth, network, parallel
|
||||
slots). Everything else (sampling, ``-c``, ``-ngl``, ``--flash-attn``,
|
||||
``--cache-type-*``, ``--spec-*``, ``--jinja``, ...) is appended after
|
||||
Studio's auto-set flags so llama.cpp's last-wins parser lets the user override.
|
||||
Unsloth's auto-set flags so llama.cpp's last-wins parser lets the user override.
|
||||
|
||||
Ref: https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
|
||||
"""
|
||||
|
|
@ -22,12 +22,12 @@ _DENYLIST_GROUPS: tuple[frozenset[str], ...] = (
|
|||
# Parallel slots: owned by typer --parallel; a pass-through would desync
|
||||
# app.state.llama_parallel_slots from llama-server.
|
||||
frozenset({"-np", "--parallel", "--n-parallel"}),
|
||||
# Model identity: Studio resolves it from LoadRequest; a second -m would
|
||||
# load a different model than Studio thinks it loaded.
|
||||
# Model identity: Unsloth resolves it from LoadRequest; a second -m would
|
||||
# load a different model than Unsloth thinks it loaded.
|
||||
frozenset({"-m", "--model"}),
|
||||
# Public model id: Studio sets a sanitized --alias so the OpenAI API never
|
||||
# Public model id: Unsloth sets a sanitized --alias so the OpenAI API never
|
||||
# exposes the local .gguf path. A user-supplied alias is appended after
|
||||
# Studio's and, with llama.cpp's last-wins parsing, would reintroduce the
|
||||
# Unsloth's and, with llama.cpp's last-wins parsing, would reintroduce the
|
||||
# path leak this is meant to prevent.
|
||||
frozenset({"-a", "--alias"}),
|
||||
frozenset({"-mu", "--model-url"}),
|
||||
|
|
@ -39,14 +39,14 @@ _DENYLIST_GROUPS: tuple[frozenset[str], ...] = (
|
|||
frozenset({"-hft", "--hf-token"}),
|
||||
frozenset({"-mm", "--mmproj"}),
|
||||
frozenset({"-mmu", "--mmproj-url"}),
|
||||
# Networking: Studio binds + proxies; retargeting orphans the proxy.
|
||||
# Networking: Unsloth binds + proxies; retargeting orphans the proxy.
|
||||
frozenset({"--host"}),
|
||||
frozenset({"--port"}),
|
||||
frozenset({"--path"}),
|
||||
frozenset({"--api-prefix"}),
|
||||
frozenset({"--reuse-port"}),
|
||||
# Auth / TLS: Studio terminates auth; upstream --api-key / TLS shadows
|
||||
# Studio's key and breaks the proxy hop.
|
||||
# Auth / TLS: Unsloth terminates auth; upstream --api-key / TLS shadows
|
||||
# Unsloth's key and breaks the proxy hop.
|
||||
frozenset({"--api-key"}),
|
||||
frozenset({"--api-key-file"}),
|
||||
frozenset({"--ssl-key-file"}),
|
||||
|
|
@ -64,11 +64,11 @@ _DENYLIST_GROUPS: tuple[frozenset[str], ...] = (
|
|||
frozenset({"--models-max"}),
|
||||
frozenset({"--models-autoload", "--no-models-autoload"}),
|
||||
# Server-mode flips: --embedding / --rerank restrict llama-server to
|
||||
# those endpoints, breaking Studio's /v1/chat/completions hop.
|
||||
# those endpoints, breaking Unsloth's /v1/chat/completions hop.
|
||||
frozenset({"--embedding", "--embeddings"}),
|
||||
frozenset({"--rerank", "--reranking"}),
|
||||
# llama-server's own built-in tools flag would silently stack on top of
|
||||
# Studio's --enable-tools / --disable-tools policy resolver.
|
||||
# Unsloth's --enable-tools / --disable-tools policy resolver.
|
||||
frozenset({"--tools"}),
|
||||
)
|
||||
|
||||
|
|
@ -120,7 +120,7 @@ def validate_extra_args(args: Optional[Iterable[str]]) -> list[str]:
|
|||
|
||||
|
||||
def is_managed_flag(flag: str) -> bool:
|
||||
"""True if ``flag`` is Studio-managed. Normalises via ``_flag_name`` so
|
||||
"""True if ``flag`` is Unsloth-managed. Normalises via ``_flag_name`` so
|
||||
`-np8` / `--parallel=8` classify like the canonical tokens."""
|
||||
normalised = _flag_name(flag)
|
||||
return normalised is not None and normalised in _DENYLIST
|
||||
|
|
@ -142,7 +142,7 @@ _SPEC_FLAGS: frozenset[str] = frozenset(
|
|||
"--draft-min",
|
||||
"--draft-max",
|
||||
# MTP path (llama.cpp #22673). The drafter selectors (local --model-draft
|
||||
# and HF --spec-draft-hf aliases) are Studio-managed since the separate-
|
||||
# and HF --spec-draft-hf aliases) are Unsloth-managed since the separate-
|
||||
# drafter support (Gemma 4): an inherited copy must not last-wins-override
|
||||
# the auto-detected drafter. Explicit extras for the current load are never
|
||||
# stripped. The per-drafter tuning knobs (--spec-draft-type-*, -ngld,
|
||||
|
|
@ -179,9 +179,9 @@ _TEMPLATE_FLAGS: frozenset[str] = frozenset(
|
|||
# (--split-mode tensor). Pass-through stays allowed so users keep the
|
||||
# row/none/layer modes the toggle doesn't expose, but it's stripped on
|
||||
# inherit and reconciled into the round-tripped tensor_parallel state.
|
||||
# --tensor-split is coupled to the split mode and is stripped with it: Studio
|
||||
# --tensor-split is coupled to the split mode and is stripped with it: Unsloth
|
||||
# owns the tensor-mode split ratios, so an inherited/stale --tensor-split must
|
||||
# not last-wins-override Studio's computed asymmetric split.
|
||||
# not last-wins-override Unsloth's computed asymmetric split.
|
||||
_SPLIT_MODE_FLAGS: frozenset[str] = frozenset({"-sm", "--split-mode"})
|
||||
_TENSOR_SPLIT_FLAGS: frozenset[str] = frozenset({"-ts", "--tensor-split"})
|
||||
_SPLIT_SHADOWING_FLAGS: frozenset[str] = _SPLIT_MODE_FLAGS | _TENSOR_SPLIT_FLAGS
|
||||
|
|
@ -197,7 +197,7 @@ _BOOLEAN_SHADOWING_FLAGS: frozenset[str] = frozenset({"--spec-default", "--jinja
|
|||
def parse_ctx_override(args: Optional[Iterable[str]]) -> Optional[int]:
|
||||
"""Return the last user-supplied ``-c`` / ``--ctx-size`` value.
|
||||
|
||||
Mirrors llama.cpp's last-wins parsing for the one numeric knob Studio's
|
||||
Mirrors llama.cpp's last-wins parsing for the one numeric knob Unsloth's
|
||||
load-time fit logic needs.
|
||||
"""
|
||||
if not args:
|
||||
|
|
@ -286,7 +286,7 @@ def parse_cache_override(args: Optional[Iterable[str]]) -> Optional[str]:
|
|||
Mirrors parse_ctx_override but for cache type. Recognises both -ctk
|
||||
(key) and -ctv (value). When both flags appear, returns the last-wins
|
||||
value, treating key and value cache flags as the same setting because
|
||||
Studio's KV estimate has a single cache_type_kv knob.
|
||||
Unsloth's KV estimate has a single cache_type_kv knob.
|
||||
"""
|
||||
return _last_flag_value(args, _CACHE_FLAGS)
|
||||
|
||||
|
|
@ -341,7 +341,7 @@ def resolve_tensor_parallel(args: Optional[Iterable[str]], fallback_tensor_paral
|
|||
|
||||
|
||||
def _env_split_mode_is_tensor(env: Optional[Mapping[str, str]] = None) -> bool:
|
||||
"""True when the inherited LLAMA_ARG_SPLIT_MODE env selects tensor. Studio
|
||||
"""True when the inherited LLAMA_ARG_SPLIT_MODE env selects tensor. Unsloth
|
||||
emits --split-mode only on its tensor branch, so a tensor env on the layer
|
||||
path would run the child tensor-parallel unbudgeted; this flips the budget
|
||||
to tensor. Only tensor is heavier, so other modes are ignored."""
|
||||
|
|
@ -425,7 +425,7 @@ def strip_shadowing_flags(
|
|||
strip_template: bool = True,
|
||||
strip_split_mode: bool = True,
|
||||
) -> list[str]:
|
||||
"""Strip flags that shadow first-class Studio settings.
|
||||
"""Strip flags that shadow first-class Unsloth settings.
|
||||
|
||||
Used when inheriting a previous load's ``llama_extra_args`` so an
|
||||
inherited `-c 4096` can't override the current `max_seq_length`
|
||||
|
|
|
|||
|
|
@ -5,7 +5,7 @@
|
|||
engine-stats log line (generation/prompt throughput, requests in flight).
|
||||
|
||||
llama-server already computes these (it needs `--metrics`); this lifts them
|
||||
into Studio's structured log so the terminal shows serving health, not just
|
||||
into Unsloth's structured log so the terminal shows serving health, not just
|
||||
per-request access lines. Emitted only while there is activity.
|
||||
"""
|
||||
|
||||
|
|
|
|||
|
|
@ -130,7 +130,7 @@ def info_has_local_gguf(info) -> bool:
|
|||
def _build_index() -> dict[str, _LocalGgufEntry]:
|
||||
"""Map normalized id/model_id/display_name -> local GGUF entry.
|
||||
|
||||
Scans the same roots Studio's model picker lists (./models, the active plus
|
||||
Scans the same roots Unsloth's model picker lists (./models, the active plus
|
||||
legacy/default HF caches, LM Studio dirs, and user scan folders) so a named
|
||||
local model is never missed and silently served as the loaded one. Ollama's
|
||||
scanner is skipped: it creates symlinks as a side effect and this runs on the
|
||||
|
|
@ -199,7 +199,7 @@ def _build_index() -> dict[str, _LocalGgufEntry]:
|
|||
raw_id = getattr(info, "id", None)
|
||||
if not raw_id:
|
||||
continue
|
||||
# Skip what Studio hides from its pickers (validation probe, RAG embed
|
||||
# Skip what Unsloth hides from its pickers (validation probe, RAG embed
|
||||
# weights): not chat models, so never an auto-switch target.
|
||||
if _is_hidden_model(raw_id, getattr(info, "path", None)):
|
||||
continue
|
||||
|
|
|
|||
|
|
@ -906,7 +906,7 @@ def _call_stdio_tool(
|
|||
def _remaining() -> Optional[float]:
|
||||
return None if deadline is None else max(0.0, deadline - time.monotonic())
|
||||
|
||||
# Callers without a Studio session id must retain the former one-shot
|
||||
# Callers without an Unsloth session id must retain the former one-shot
|
||||
# behavior: no browser/cookie/tool state can leak into another request.
|
||||
# Use an ephemeral key (and close it below) rather than the shared empty
|
||||
# scope that the persistent-session cache used previously.
|
||||
|
|
|
|||
|
|
@ -5,7 +5,7 @@
|
|||
|
||||
With server-side tools disabled (``unsloth run --disable-tools``, every
|
||||
``unsloth start`` coding agent), requests carrying the client's own ``tools``
|
||||
bypass Studio's tool loop and are relayed to/from llama-server verbatim. Small
|
||||
bypass Unsloth's tool loop and are relayed to/from llama-server verbatim. Small
|
||||
GGUF models often emit their tool calls as TEXT (``<tool_call>{...}</tool_call>``,
|
||||
Gemma ``<|tool_call>...``, ``<function=...>`` XML) instead of structured
|
||||
``tool_calls`` -- on the passthrough that text reaches the agent as prose and
|
||||
|
|
@ -18,7 +18,7 @@ promotes calls whose function name exactly matches a declared tool. Promotion
|
|||
removes EXACTLY the promoted calls' markup spans (the parser reports them):
|
||||
undeclared calls, unparseable blocks, and suppressed alternate formats keep
|
||||
every byte and relay as text, so healing can never silently delete model
|
||||
output. Responses without a tool signal, requests without tools, and Studio's
|
||||
output. Responses without a tool signal, requests without tools, and Unsloth's
|
||||
own enable-tools loop are untouched. Per-request opt-out:
|
||||
``auto_heal_tool_calls: false``. Process kill-switch:
|
||||
``UNSLOTH_DISABLE_TOOL_CALL_HEALING=1``.
|
||||
|
|
|
|||
|
|
@ -122,12 +122,12 @@ def calculate_cost(provider: str, model: str, usage: dict[str, Any]) -> dict[str
|
|||
"priced": bool(prices),
|
||||
}
|
||||
|
||||
# Accept raw (input_tokens/output_tokens) and Studio chat-style
|
||||
# Accept raw (input_tokens/output_tokens) and Unsloth chat-style
|
||||
# (prompt_tokens/completion_tokens) envelopes. Cache buckets differ:
|
||||
# raw Anthropic: input_tokens EXCLUDES cache buckets
|
||||
# raw OpenAI: input_tokens INCLUDES cache_read
|
||||
# Studio Anthropic: prompt_tokens INCLUDES cache_creation + cache_read
|
||||
# Studio OpenAI: prompt_tokens == raw input_tokens
|
||||
# Unsloth Anthropic: prompt_tokens INCLUDES cache_creation + cache_read
|
||||
# Unsloth OpenAI: prompt_tokens == raw input_tokens
|
||||
# Clamp >=0 so corrupted payloads can't produce a negative bill.
|
||||
cache_creation = max(0, int(usage.get("cache_creation_input_tokens") or 0))
|
||||
cache_read_native_present = (
|
||||
|
|
@ -160,7 +160,7 @@ def calculate_cost(provider: str, model: str, usage: dict[str, Any]) -> dict[str
|
|||
output_tokens = max(0, int(usage.get("completion_tokens") or 0))
|
||||
if provider == "openai":
|
||||
# Cached tokens land on input_tokens_details (raw Responses) or
|
||||
# prompt_tokens_details (Studio chat-style).
|
||||
# prompt_tokens_details (Unsloth chat-style).
|
||||
for key in ("input_tokens_details", "prompt_tokens_details"):
|
||||
details = usage.get(key) or {}
|
||||
if isinstance(details, dict):
|
||||
|
|
|
|||
|
|
@ -995,7 +995,7 @@ def run_safetensors_tool_loop(
|
|||
if not safety_tc:
|
||||
# Re-prompt once on plan-without-action, before any tool runs
|
||||
# (GGUF loop parity). The retry is gated on nudge_tool_calls so
|
||||
# Studio callers (which send True) always nudge, while API callers
|
||||
# Unsloth callers (which send True) always nudge, while API callers
|
||||
# who omit the flag keep today's no-reprompt behavior (opt-in).
|
||||
intent_text = _reprompt_intent_text(
|
||||
content_accum,
|
||||
|
|
|
|||
|
|
@ -4,7 +4,7 @@
|
|||
"""Sandbox-side compatibility shim for ChatGPT code-interpreter paths.
|
||||
|
||||
Models habitually write to /mnt/data (or /mnt/outputs, /home/sandbox,
|
||||
/workspace), none of which exist in the Studio sandbox. This module sits on the
|
||||
/workspace), none of which exist in the Unsloth sandbox. This module sits on the
|
||||
sandbox subprocess PYTHONPATH (see ``tools._build_safe_env``), so it loads at
|
||||
interpreter startup in every sandboxed ``python`` run and any Python the
|
||||
``terminal`` tool launches.
|
||||
|
|
|
|||
|
|
@ -1,7 +1,7 @@
|
|||
# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
||||
|
||||
"""Shared controller state for Studio local agentic tool loops.
|
||||
"""Shared controller state for Unsloth local agentic tool loops.
|
||||
|
||||
This module is intentionally dependency-light: it owns only per-response
|
||||
ledger state and value objects used by the GGUF and safetensors loops.
|
||||
|
|
|
|||
|
|
@ -2502,7 +2502,7 @@ def _build_safe_env(workdir: str) -> dict[str, str]:
|
|||
shim directory.
|
||||
"""
|
||||
# Start from the running interpreter's dir so 'python'/'pip' resolve to the
|
||||
# same environment the Studio server runs in.
|
||||
# same environment the Unsloth server runs in.
|
||||
exe_dir = os.path.dirname(sys.executable)
|
||||
path_entries = [exe_dir] if exe_dir else []
|
||||
|
||||
|
|
@ -2792,7 +2792,7 @@ def _bypass_preexec():
|
|||
"""Minimal pre-exec for bypass exec: os.setsid() only.
|
||||
|
||||
Required, not a restriction: _kill_process_tree does killpg(getpgid(child)),
|
||||
so without a new session a timeout/cancel would kill the Studio server too.
|
||||
so without a new session a timeout/cancel would kill the Unsloth server too.
|
||||
"""
|
||||
try:
|
||||
os.setsid()
|
||||
|
|
@ -2800,13 +2800,13 @@ def _bypass_preexec():
|
|||
pass
|
||||
|
||||
|
||||
# Hardening the Studio parent is done once (PR_SET_DUMPABLE is process-global
|
||||
# Hardening the Unsloth parent is done once (PR_SET_DUMPABLE is process-global
|
||||
# and sticky); guarded so repeated bypass calls do not re-issue the prctl.
|
||||
_parent_proc_hardened = False
|
||||
|
||||
|
||||
def _harden_parent_against_proc_env_leak() -> bool:
|
||||
"""Make the Studio process's /proc/<pid>/environ unreadable to its children.
|
||||
"""Make the Unsloth process's /proc/<pid>/environ unreadable to its children.
|
||||
|
||||
Stripping the child env is not enough on Linux: a bypassed same-UID child
|
||||
can read /proc/<getppid()>/environ to recover the parent's unfiltered
|
||||
|
|
@ -5482,7 +5482,7 @@ def _truncate(text: str, limit: int = _MAX_OUTPUT_CHARS) -> str:
|
|||
|
||||
|
||||
# ChatGPT code-interpreter path conventions models write out of habit; none
|
||||
# exist in the Studio sandbox, so a failure on one earns the retry hint.
|
||||
# exist in the Unsloth sandbox, so a failure on one earns the retry hint.
|
||||
_MISSING_PATH_PREFIXES = (
|
||||
"/mnt/data",
|
||||
"/mnt/outputs",
|
||||
|
|
@ -5688,7 +5688,7 @@ def _python_exec(
|
|||
# Close the /proc/<parent>/environ secret-recovery path first; if it
|
||||
# cannot be applied, fail closed rather than leak the parent environ.
|
||||
return (
|
||||
"Execution error: could not harden the Studio process against "
|
||||
"Execution error: could not harden the Unsloth process against "
|
||||
"/proc environment reads; refusing bypass execution."
|
||||
)
|
||||
|
||||
|
|
@ -5833,7 +5833,7 @@ def _bash_exec(
|
|||
# Close the /proc/<parent>/environ secret-recovery path first; if it
|
||||
# cannot be applied, fail closed rather than leak the parent environ.
|
||||
return (
|
||||
"Execution error: could not harden the Studio process against "
|
||||
"Execution error: could not harden the Unsloth process against "
|
||||
"/proc environment reads; refusing bypass execution."
|
||||
)
|
||||
|
||||
|
|
|
|||
|
|
@ -6,7 +6,7 @@
|
|||
Both turn pixels into indexable text and are a no-op (never raise) without a loaded
|
||||
vision model. They reuse the chat model's vision endpoint, so it must be served with
|
||||
``--ubatch-size`` >= one image's tokens (some encoders, e.g. Gemma, attend
|
||||
non-causally and abort otherwise); Studio's vision chat already requires this."""
|
||||
non-causally and abort otherwise); Unsloth's vision chat already requires this."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
|
|
|
|||
|
|
@ -10,7 +10,7 @@ Opt-in (``RAG_EMBED_BACKEND=llama-server``). Runs a dedicated
|
|||
Device is ``auto`` (GPU when present, else CPU, falling back to CPU if a GPU start
|
||||
fails); ``RAG_EMBED_DEVICE`` forces it. We call only llama_cpp's *static* helpers
|
||||
(no torch), copying the instance-coupled bits locally, since constructing a
|
||||
``LlamaCppBackend`` runs an ``__init__`` reaper that kills any Studio llama-server
|
||||
``LlamaCppBackend`` runs an ``__init__`` reaper that kills any Unsloth llama-server
|
||||
-- so each request re-spawns ours if it died (self-heal).
|
||||
"""
|
||||
|
||||
|
|
|
|||
|
|
@ -39,7 +39,7 @@ _model = None
|
|||
_name: str | None = None
|
||||
|
||||
|
||||
# Studio device -> torch device string. Apple has no torch device -> CPU.
|
||||
# Unsloth device -> torch device string. Apple has no torch device -> CPU.
|
||||
_TORCH_DEVICE = {DeviceType.CUDA: "cuda", DeviceType.XPU: "xpu"}
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -53,7 +53,7 @@ def get_resume_checkpoint_path(path_value: str) -> Optional[str]:
|
|||
def normalize_resume_output_dir(path_value: str) -> str:
|
||||
path = resolve_output_dir(path_value)
|
||||
if not _is_under_outputs(path):
|
||||
raise ValueError("Resume checkpoint must be inside Studio outputs.")
|
||||
raise ValueError("Resume checkpoint must be inside Unsloth outputs.")
|
||||
return str(path)
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -797,7 +797,7 @@ class UnslothTrainer:
|
|||
)
|
||||
logger.info("Loaded text model")
|
||||
|
||||
raise_if_offloaded(self.model, device_map, "Studio training")
|
||||
raise_if_offloaded(self.model, device_map, "Unsloth training")
|
||||
|
||||
if self.should_stop:
|
||||
return False
|
||||
|
|
|
|||
|
|
@ -140,7 +140,7 @@ def should_use_mlx_training_backend(*, device: Optional[Any] = None) -> bool:
|
|||
|
||||
|
||||
def _build_training_worker_config(values: dict[str, Any]) -> dict[str, Any]:
|
||||
"""Build the normalized worker config shared by Studio and the CLI adapter."""
|
||||
"""Build the normalized worker config shared by Unsloth and the CLI adapter."""
|
||||
config = {
|
||||
"model_name": values["model_name"],
|
||||
"project_name": values.get("project_name"),
|
||||
|
|
@ -307,7 +307,7 @@ PLOT_HEIGHT = 3.5
|
|||
|
||||
@dataclass
|
||||
class TrainingProgress:
|
||||
"""Shared training progress payload for Studio and backend-aware trainers."""
|
||||
"""Shared training progress payload for Unsloth and backend-aware trainers."""
|
||||
|
||||
epoch: float = 0
|
||||
step: int = 0
|
||||
|
|
@ -328,7 +328,7 @@ class TrainingProgress:
|
|||
|
||||
|
||||
class _MLXTrainerAdapter:
|
||||
"""Adapts the legacy UnslothTrainer API to the shared Studio MLX worker path."""
|
||||
"""Adapts the legacy UnslothTrainer API to the shared Unsloth MLX worker path."""
|
||||
|
||||
def __init__(self):
|
||||
self.model = None
|
||||
|
|
|
|||
|
|
@ -1100,7 +1100,7 @@ _MLX_VLM_RESIZED_IMAGE_LAYOUT_CACHE = {}
|
|||
|
||||
|
||||
def _mlx_vlm_resized_image_layout(processor = None) -> str | None:
|
||||
"""Return the numpy image layout expected after Studio-side VLM resizing."""
|
||||
"""Return the numpy image layout expected after Unsloth-side VLM resizing."""
|
||||
image_processor = getattr(processor, "image_processor", None)
|
||||
if image_processor is None:
|
||||
return None
|
||||
|
|
@ -1257,7 +1257,7 @@ _MLX_STUDIO_LR_SCHEDULERS = {"linear", "cosine", "constant"}
|
|||
|
||||
|
||||
# Fallback alias map mirroring unsloth_zoo._normalize_mlx_optimizer_name, used
|
||||
# only when mlx (Apple Silicon) is not importable so Studio config validation
|
||||
# only when mlx (Apple Silicon) is not importable so Unsloth config validation
|
||||
# still works on non-MLX hosts. The zoo function stays the source of truth.
|
||||
_MLX_STUDIO_ADAMW_ALIASES = frozenset(
|
||||
(
|
||||
|
|
@ -1309,7 +1309,7 @@ def _normalize_mlx_studio_scheduler(value):
|
|||
|
||||
|
||||
def _resolve_mlx_local_dataset_files(file_paths: list) -> list[str]:
|
||||
"""Resolve CLI paths and Studio local dataset uploads without importing the GPU trainer."""
|
||||
"""Resolve CLI paths and Unsloth local dataset uploads without importing the GPU trainer."""
|
||||
from utils.paths import resolve_dataset_path
|
||||
|
||||
all_files: list[str] = []
|
||||
|
|
@ -1912,7 +1912,7 @@ def _run_mlx_training(event_queue, stop_queue, config):
|
|||
if "max_grad_leaf_norm" in _supported_fields:
|
||||
mlx_config_kwargs["max_grad_leaf_norm"] = max_grad_leaf_norm
|
||||
if "append_eos" in _supported_fields:
|
||||
# Studio SFT formatting owns rendered examples; raw/CPT text still
|
||||
# Unsloth SFT formatting owns rendered examples; raw/CPT text still
|
||||
# needs MLX to append EOS like the CUDA raw-text path.
|
||||
mlx_config_kwargs["append_eos"] = bool(raw_text_mode)
|
||||
|
||||
|
|
@ -2121,7 +2121,7 @@ def run_mlx_training_process(
|
|||
config: dict,
|
||||
transformers_activated: bool = False,
|
||||
) -> None:
|
||||
"""MLX worker entrypoint shared by Studio subprocesses and the CLI adapter."""
|
||||
"""MLX worker entrypoint shared by Unsloth subprocesses and the CLI adapter."""
|
||||
model_name = config["model_name"]
|
||||
|
||||
backend_path = str(Path(__file__).resolve().parent.parent.parent)
|
||||
|
|
@ -2780,7 +2780,7 @@ def run_training_process(*, event_queue: Any, stop_queue: Any, config: dict) ->
|
|||
)
|
||||
# Unified Windows APUs: the WDDM budget is user-raisable, but
|
||||
# nothing on the box says so -- users see "48 GB VRAM" on a
|
||||
# 96 GB machine and assume a Studio bug. Say where the limit
|
||||
# 96 GB machine and assume an Unsloth bug. Say where the limit
|
||||
# comes from and how to raise it.
|
||||
if _is_unified and sys.platform == "win32":
|
||||
try:
|
||||
|
|
|
|||
|
|
@ -76,7 +76,7 @@ def spawn_worker(
|
|||
env["HF_HUB_DISABLE_PROGRESS_BARS"] = "1"
|
||||
env["HF_HUB_DISABLE_TELEMETRY"] = "1"
|
||||
env["HF_HUB_DISABLE_XET"] = "0" if use_xet else "1"
|
||||
# No token in Studio settings: fall back to the backend's own HF_TOKEN so
|
||||
# No token in Unsloth settings: fall back to the backend's own HF_TOKEN so
|
||||
# private repos stay downloadable (needed while inkling repos are private).
|
||||
if not hf_token:
|
||||
hf_token = os.environ.get("HF_TOKEN") or None
|
||||
|
|
|
|||
|
|
@ -165,7 +165,7 @@ def _looks_like_model_dir(directory: Path) -> bool:
|
|||
def _build_browse_allowlist(
|
||||
media_roots: Optional[list[Path]] = None, drive_roots: Optional[list[Path]] = None
|
||||
) -> list[Path]:
|
||||
"""Root directories the browser may walk (also seeds the suggestion chips): HOME, resolved HF cache dirs, Studio outputs/exports/root, registered scan folders, and well-known local-LLM dirs. Each is added only if it resolves to a real directory so the sandbox has no dead boundary.
|
||||
"""Root directories the browser may walk (also seeds the suggestion chips): HOME, resolved HF cache dirs, Unsloth outputs/exports/root, registered scan folders, and well-known local-LLM dirs. Each is added only if it resolves to a real directory so the sandbox has no dead boundary.
|
||||
|
||||
*media_roots* / *drive_roots* let the caller pass already-probed
|
||||
removable-media and Windows drive roots so they aren't scanned again (a
|
||||
|
|
|
|||
|
|
@ -85,7 +85,7 @@ def _contained_link_path(link_dir: Path, link_name: str) -> Optional[Path]:
|
|||
|
||||
|
||||
def _ollama_links_dir(ollama_dir: Path) -> Optional[Path]:
|
||||
"""Writable directory for Ollama ``.gguf`` symlinks. Prefers ``<ollama_dir>/.studio_links/`` next to the blobs; falls back to Studio's cache (read-only system installs), then the temp dir (sandboxed installs)."""
|
||||
"""Writable directory for Ollama ``.gguf`` symlinks. Prefers ``<ollama_dir>/.studio_links/`` next to the blobs; falls back to Unsloth's cache (read-only system installs), then the temp dir (sandboxed installs)."""
|
||||
|
||||
def _ensure_writable_dir(path: Path) -> Optional[Path]:
|
||||
try:
|
||||
|
|
|
|||
|
|
@ -3,7 +3,7 @@
|
|||
|
||||
"""Filesystem layout for Hub download state.
|
||||
|
||||
State directory sits beside HF's cache (under Studio's own cache root)
|
||||
State directory sits beside HF's cache (under Unsloth's own cache root)
|
||||
so it survives ``huggingface-cli delete-cache`` and any other HF-side
|
||||
cache lifecycle. Two subdirectories:
|
||||
|
||||
|
|
|
|||
|
|
@ -19,7 +19,7 @@ os.environ["PYTHONWARNINGS"] = "ignore"
|
|||
|
||||
# Pin GPU index ordering to PCI bus id before any torch import creates a CUDA
|
||||
# context. Without this, torch/CUDA default to FASTEST_FIRST while nvidia-smi
|
||||
# (and Studio's VRAM probes) use PCI-bus order, so a GPU index chosen from
|
||||
# (and Unsloth's VRAM probes) use PCI-bus order, so a GPU index chosen from
|
||||
# nvidia-smi data can resolve to a different physical card via
|
||||
# CUDA_VISIBLE_DEVICES. setdefault so an explicit user override wins. See
|
||||
# utils/hardware/hardware.py for the full rationale; set here too so the entry
|
||||
|
|
@ -93,7 +93,7 @@ if sys.platform == "win32":
|
|||
# ── Windows AMD ROCm: make hipInfo.exe resolvable for subprocess probes ──
|
||||
# bitsandbytes' get_rocm_gpu_arch() runs `hipinfo.exe` via PATH at import
|
||||
# time; the AMD torch wheel ships it in the venv Scripts dir, which is on
|
||||
# PATH only when the venv is activated -- Studio launches python directly.
|
||||
# PATH only when the venv is activated -- Unsloth launches python directly.
|
||||
# Without this, every bitsandbytes import logs a scary (but harmless)
|
||||
# "Could not detect ROCm GPU architecture: [WinError 2]" ERROR + WARNING.
|
||||
# Gated on the file existing: only AMD ROCm wheels ship hipInfo.exe, so
|
||||
|
|
@ -252,7 +252,7 @@ def _read_studio_install_id() -> str:
|
|||
|
||||
Returns "" when absent or not a 64-char lowercase-hex token; then
|
||||
/api/health emits "" and the launcher accepts any healthy backend.
|
||||
Carries no install-path info (matters when Studio runs -H 0.0.0.0)."""
|
||||
Carries no install-path info (matters when Unsloth runs -H 0.0.0.0)."""
|
||||
try:
|
||||
token = (_STUDIO_ROOT_RESOLVED / "share" / "studio_install_id").read_text().strip()
|
||||
except (OSError, ValueError):
|
||||
|
|
@ -573,7 +573,7 @@ async def lifespan(app: FastAPI):
|
|||
print("DEFAULT ADMIN ACCOUNT CREATED")
|
||||
print(f" username: {storage.DEFAULT_ADMIN_USERNAME}")
|
||||
print(f" password saved to: {bootstrap_path}")
|
||||
print(" Open the Studio UI to sign in and change it.")
|
||||
print(" Open the Unsloth UI to sign in and change it.")
|
||||
print("=" * 60 + "\n")
|
||||
else:
|
||||
app.state.bootstrap_password = (
|
||||
|
|
@ -613,7 +613,7 @@ app = FastAPI(
|
|||
)
|
||||
|
||||
# The MCP surface is opt-in because it can start GPU jobs and write model
|
||||
# artifacts. Mount it only when explicitly enabled by the Studio process.
|
||||
# artifacts. Mount it only when explicitly enabled by the Unsloth process.
|
||||
if os.environ.get("UNSLOTH_STUDIO_ENABLE_MCP") == "1":
|
||||
from fastmcp.utilities.lifespan import combine_lifespans
|
||||
|
||||
|
|
@ -973,7 +973,7 @@ app.include_router(training_router, prefix = "/api/train", tags = ["training"])
|
|||
app.include_router(models_router, prefix = "/api/models", tags = ["models"])
|
||||
app.include_router(chat_history_router, prefix = "/api/chat", tags = ["chat"])
|
||||
app.include_router(inference_router, prefix = "/api/inference", tags = ["inference"])
|
||||
# Studio-only inference endpoints (cancel, etc.) are NOT exposed on the /v1
|
||||
# Unsloth-only inference endpoints (cancel, etc.) are NOT exposed on the /v1
|
||||
# OpenAI-compat prefix below.
|
||||
app.include_router(inference_studio_router, prefix = "/api/inference", tags = ["inference"])
|
||||
|
||||
|
|
@ -1080,7 +1080,7 @@ def studio_install_source(_current_subject: str = Depends(get_current_subject)):
|
|||
|
||||
@app.get("/api/studio/update-status")
|
||||
def studio_update_status(_current_subject: str = Depends(get_current_subject)):
|
||||
"""Return source-aware manual update status for browser-served Studio."""
|
||||
"""Return source-aware manual update status for browser-served Unsloth."""
|
||||
return get_studio_update_status(UNSLOTH_VERSION)
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -3,7 +3,7 @@
|
|||
|
||||
"""Curated MCP tools for driving an Unsloth Studio instance.
|
||||
|
||||
The MCP surface deliberately wraps the existing Studio services instead of
|
||||
The MCP surface deliberately wraps the existing Unsloth services instead of
|
||||
duplicating training or export logic. It is opt-in because several tools can
|
||||
start GPU work or write model artifacts.
|
||||
"""
|
||||
|
|
@ -17,14 +17,14 @@ from fastmcp import FastMCP
|
|||
|
||||
|
||||
class BearerTokenMiddleware:
|
||||
"""Require an exact bearer token when Studio MCP is exposed remotely."""
|
||||
"""Require an exact bearer token when Unsloth MCP is exposed remotely."""
|
||||
|
||||
def __init__(self, app: Any, token: str) -> None:
|
||||
if not token or not token.strip():
|
||||
raise ValueError("Studio MCP bearer token must be a non-empty value")
|
||||
raise ValueError("Unsloth MCP bearer token must be a non-empty value")
|
||||
if not token.isascii():
|
||||
# A non-ASCII token cannot be sent in an HTTP header; reject it here.
|
||||
raise ValueError("Studio MCP bearer token must contain ASCII characters only")
|
||||
raise ValueError("Unsloth MCP bearer token must contain ASCII characters only")
|
||||
self.app = app
|
||||
# Compare on raw header bytes: str hmac.compare_digest raises on non-ASCII
|
||||
# input, which would surface as a 500 instead of a clean 401.
|
||||
|
|
@ -76,18 +76,18 @@ def _dump(value: Any) -> Any:
|
|||
def _clamp(value: int, low: int, high: int) -> int:
|
||||
"""Clamp an MCP-supplied integer into an inclusive range.
|
||||
|
||||
MCP tools call the Studio route functions directly, which skips FastAPI's
|
||||
MCP tools call the Unsloth route functions directly, which skips FastAPI's
|
||||
Query(ge=, le=) validation, so we re-apply the same bounds here.
|
||||
"""
|
||||
return max(low, min(value, high))
|
||||
|
||||
|
||||
def create_studio_mcp() -> FastMCP:
|
||||
"""Create the Studio MCP server and register the high-value tools."""
|
||||
"""Create the Unsloth MCP server and register the high-value tools."""
|
||||
mcp = FastMCP(
|
||||
"Unsloth Studio",
|
||||
instructions = (
|
||||
"Use read tools to inspect the local Studio state before starting GPU work. "
|
||||
"Use read tools to inspect the local Unsloth state before starting GPU work. "
|
||||
"Training and export tools can consume substantial VRAM and write files. "
|
||||
"Never expose tokens or local paths from tool results unless the user asks."
|
||||
),
|
||||
|
|
@ -116,7 +116,7 @@ def create_studio_mcp() -> FastMCP:
|
|||
|
||||
@mcp.tool
|
||||
async def list_local_models(models_dir: str = "./models") -> dict[str, Any]:
|
||||
"""List local and cached models available to Studio."""
|
||||
"""List local and cached models available to Unsloth."""
|
||||
from routes.models import list_local_models as list_models
|
||||
return _dump(await list_models(models_dir = models_dir, current_subject = "mcp"))
|
||||
|
||||
|
|
@ -128,9 +128,9 @@ def create_studio_mcp() -> FastMCP:
|
|||
|
||||
@mcp.tool
|
||||
async def start_training(config: dict[str, Any]) -> dict[str, Any]:
|
||||
"""Start a validated Studio training job from a TrainingStartRequest-shaped object.
|
||||
"""Start a validated Unsloth training job from a TrainingStartRequest-shaped object.
|
||||
|
||||
The config is validated by the same Pydantic model used by the Studio UI.
|
||||
The config is validated by the same Pydantic model used by the Unsloth UI.
|
||||
Call get_training_status first and do not start work while another job runs.
|
||||
"""
|
||||
from models import TrainingStartRequest
|
||||
|
|
@ -138,7 +138,7 @@ def create_studio_mcp() -> FastMCP:
|
|||
|
||||
request = TrainingStartRequest.model_validate(config)
|
||||
# Pass via_api_key explicitly (a direct call leaves it a Depends object).
|
||||
# MCP drives Studio like the UI session, so it coexists and frees VRAM.
|
||||
# MCP drives Unsloth like the UI session, so it coexists and frees VRAM.
|
||||
return _dump(await start(request, current_subject = "mcp", via_api_key = False))
|
||||
|
||||
@mcp.tool
|
||||
|
|
@ -159,7 +159,7 @@ def create_studio_mcp() -> FastMCP:
|
|||
|
||||
@mcp.tool
|
||||
def validate_recipe(recipe: dict[str, Any]) -> dict[str, Any]:
|
||||
"""Validate a Data Recipe with the same validator used by Studio."""
|
||||
"""Validate a Data Recipe with the same validator used by Unsloth."""
|
||||
from models.data_recipe import RecipePayload
|
||||
from routes.data_recipe.validate import validate
|
||||
|
||||
|
|
@ -225,7 +225,7 @@ def create_studio_mcp() -> FastMCP:
|
|||
imatrix: bool = False,
|
||||
imatrix_path: str | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""Export the loaded model to GGUF using Studio's existing path validation.
|
||||
"""Export the loaded model to GGUF using Unsloth's existing path validation.
|
||||
|
||||
quantization_method may be a single method or a list to produce several
|
||||
GGUFs from one load. Pass hf_token when push_to_hub is set (the backend
|
||||
|
|
|
|||
|
|
@ -105,7 +105,7 @@ class LoadRequest(BaseModel):
|
|||
description = (
|
||||
"Extra arguments forwarded verbatim to llama-server for GGUF models. "
|
||||
"One token per list entry, e.g. ['--top-k', '20', '--seed', '42']. "
|
||||
"Studio-managed flags (model identity, port, context length, GPU placement, "
|
||||
"Unsloth-managed flags (model identity, port, context length, GPU placement, "
|
||||
"auth, UI/server mode) are rejected. Ignored for non-GGUF models."
|
||||
),
|
||||
)
|
||||
|
|
@ -151,13 +151,13 @@ class TransformersUpgradeInfo(BaseModel):
|
|||
)
|
||||
supported_in_pypi: bool = Field(
|
||||
False,
|
||||
description = "True if the latest PyPI release ships this model_type; Studio can "
|
||||
description = "True if the latest PyPI release ships this model_type; Unsloth can "
|
||||
"install it into a persistent sidecar after user consent.",
|
||||
)
|
||||
supported_in_main: bool = Field(
|
||||
False,
|
||||
description = "True if transformers GitHub main ships this model_type (dev-only; "
|
||||
"not installable through Studio yet).",
|
||||
"not installable through Unsloth yet).",
|
||||
)
|
||||
|
||||
|
||||
|
|
@ -533,7 +533,7 @@ class ImageContentPart(BaseModel):
|
|||
class InputDocumentContentPart(BaseModel):
|
||||
"""Document (PDF / file) content part in a multimodal message.
|
||||
|
||||
Studio-normalised shape (file_data or file_url, plus optional filename/media_type).
|
||||
Unsloth-normalised shape (file_data or file_url, plus optional filename/media_type).
|
||||
Mapped onto Anthropic ``document`` / OpenAI ``input_file`` for vision providers;
|
||||
dropped for non-vision providers.
|
||||
"""
|
||||
|
|
@ -689,7 +689,7 @@ class ThinkingConfig(BaseModel):
|
|||
"""Anthropic-compatible thinking/reasoning configuration.
|
||||
Use type='disabled' to turn off thinking, or type='enabled' to turn it on.
|
||||
Only type is read; extra fields (e.g. budget_tokens) are ignored, since
|
||||
Studio sets provider thinking budgets itself.
|
||||
Unsloth sets provider thinking budgets itself.
|
||||
"""
|
||||
|
||||
type: Literal["disabled", "enabled"] = "disabled"
|
||||
|
|
@ -748,7 +748,7 @@ class ChatCompletionRequest(BaseModel):
|
|||
None,
|
||||
description = (
|
||||
"OpenAI function-tool definitions. When provided without `enable_tools=true`, "
|
||||
"Studio forwards the tools to the backend so the model returns structured "
|
||||
"Unsloth forwards the tools to the backend so the model returns structured "
|
||||
"tool_calls for the client to execute (standard OpenAI function calling)."
|
||||
),
|
||||
)
|
||||
|
|
@ -1160,7 +1160,7 @@ class ChatCompletionRequest(BaseModel):
|
|||
and (self.enable_tools is True or bool(self.mcp_enabled))
|
||||
):
|
||||
# "Ask" gates every call, so a direct API caller that omits the legacy
|
||||
# confirm flag must still hit the confirmation gate for Studio's own
|
||||
# confirm flag must still hit the confirmation gate for Unsloth's own
|
||||
# tool loop. An explicit confirm_tool_calls=False wins over the mode
|
||||
# (mirrors _permission_mode_confirm and the Anthropic pre-switch guard),
|
||||
# so only self-enable when the flag is unset. Only self-enable when that
|
||||
|
|
@ -1168,7 +1168,7 @@ class ChatCompletionRequest(BaseModel):
|
|||
# (enable_tools / mcp_enabled) -- the router enters the loop on those
|
||||
# signals, not on enabled_tools alone (which merely filters which tools
|
||||
# run). A plain client-tool passthrough (client-supplied `tools` that
|
||||
# Studio does not execute) must route verbatim, and external-provider
|
||||
# Unsloth does not execute) must route verbatim, and external-provider
|
||||
# routing rejects confirm_tool_calls with tools, so skip the fold there.
|
||||
#
|
||||
# "auto" is deliberately NOT folded: it only prompts for a call the
|
||||
|
|
|
|||
|
|
@ -446,7 +446,7 @@ class TrainingStartRequest(BaseModel):
|
|||
random_seed: int = Field(
|
||||
3407,
|
||||
description = (
|
||||
"Random seed; matches the Studio backend / MLX worker default "
|
||||
"Random seed; matches the Unsloth backend / MLX worker default "
|
||||
"and unsloth's historical recommended value."
|
||||
),
|
||||
)
|
||||
|
|
|
|||
|
|
@ -4,7 +4,7 @@ A Data Designer seed-reader plugin for **Unsloth Studio** that scrapes real
|
|||
GitHub data (issues, pull requests, commits) from one or more repositories
|
||||
and hands it to the recipe pipeline as a seed dataset.
|
||||
|
||||
Designed to ship with Studio as a default seed source so any user with a
|
||||
Designed to ship with Unsloth as a default seed source so any user with a
|
||||
GitHub token can build training datasets straight from live repos.
|
||||
|
||||
## What it does
|
||||
|
|
@ -64,7 +64,7 @@ sleeps until reset when the budget drops below a safety threshold.
|
|||
|
||||
## Install
|
||||
|
||||
Shipped as a default Studio plugin. For development:
|
||||
Shipped as a default Unsloth plugin. For development:
|
||||
|
||||
```bash
|
||||
pip install -e .
|
||||
|
|
|
|||
|
|
@ -3,4 +3,4 @@
|
|||
|
||||
# Intentionally empty. Data-designer loads submodules lazily via qualified names
|
||||
# in plugin.py, so importing this package must not touch data_designer.engine.*
|
||||
# during Studio bootstrap (circular import).
|
||||
# during Unsloth bootstrap (circular import).
|
||||
|
|
|
|||
|
|
@ -1,7 +1,7 @@
|
|||
# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
||||
|
||||
"""Multi-repo GitHub scraper for the Studio seed plugin.
|
||||
"""Multi-repo GitHub scraper for the Unsloth seed plugin.
|
||||
|
||||
Drives the GraphQL scraper in `scraper_impl/` per repo, capped via trial_limits
|
||||
to stop at `limit` items per resource. Then reads the per-resource JSONL shards
|
||||
|
|
|
|||
|
|
@ -5,7 +5,7 @@ julius
|
|||
torchcodec==0.10.0
|
||||
snac
|
||||
|
||||
# peft 0.19.0 causes export subprocess shutdown issues in Studio;
|
||||
# peft 0.19.0 causes export subprocess shutdown issues in Unsloth;
|
||||
# installing with --no-deps to avoid pulling in torch>=0.11.0
|
||||
peft==0.18.1
|
||||
|
||||
|
|
|
|||
|
|
@ -70,7 +70,7 @@ cut_cross_entropy
|
|||
pillow
|
||||
|
||||
# RAG store + document parsing, mirroring studio.txt. Pinned here because
|
||||
# this file installs --no-deps; without them Studio runs with RAG disabled.
|
||||
# this file installs --no-deps; without them Unsloth runs with RAG disabled.
|
||||
sqlite-vec==0.1.9
|
||||
pymupdf==1.27.2.3
|
||||
# 0.3.x keeps pymupdf-layout (which pulls onnxruntime) an optional extra; the
|
||||
|
|
|
|||
|
|
@ -4,7 +4,7 @@ transformers==4.57.6
|
|||
trl==0.23.1
|
||||
huggingface-hub==0.36.2
|
||||
|
||||
# Studio stack
|
||||
# Unsloth stack
|
||||
datasets==4.3.0
|
||||
pyarrow==23.0.1
|
||||
|
||||
|
|
|
|||
|
|
@ -1,4 +1,4 @@
|
|||
# Studio UI backend dependencies
|
||||
# Unsloth UI backend dependencies
|
||||
typer
|
||||
fastapi
|
||||
uvicorn
|
||||
|
|
@ -9,7 +9,7 @@ pandas
|
|||
nest_asyncio
|
||||
datasets==4.3.0
|
||||
pyjwt
|
||||
# gradio>=4.0.0 # 148 MB - Studio uses React + FastAPI, not Gradio
|
||||
# gradio>=4.0.0 # 148 MB - Unsloth uses React + FastAPI, not Gradio
|
||||
huggingface-hub==0.36.2
|
||||
structlog>=24.1.0
|
||||
diceware
|
||||
|
|
|
|||
|
|
@ -338,11 +338,11 @@ def _clear_login_bucket(key: tuple[str, str]) -> None:
|
|||
# so FastAPI runs it in the threadpool rather than blocking the event loop.
|
||||
@router.get("/identity")
|
||||
def identity(nonce: str, request: Request) -> dict:
|
||||
"""Challenge-response proof this is the real local Studio: caller sends a nonce,
|
||||
"""Challenge-response proof this is the real local Unsloth: caller sends a nonce,
|
||||
gets HMAC(install identity secret, nonce, connection address + port).
|
||||
Unauthenticated and side-effect free; a process that can't read the same-user
|
||||
secret can't forge a proof, and binding to the address/port the connection
|
||||
landed on stops a squatter relaying a proof from the real Studio elsewhere."""
|
||||
landed on stops a squatter relaying a proof from the real Unsloth elsewhere."""
|
||||
try:
|
||||
raw = base64.urlsafe_b64decode(nonce)
|
||||
except Exception:
|
||||
|
|
|
|||
|
|
@ -37,7 +37,7 @@ def _resolve_local_v1_endpoint(request: Request) -> str:
|
|||
|
||||
Resolution order:
|
||||
1. ``app.state.server_port`` (run.py, post-bind) - survives proxies/tunnels.
|
||||
2. ``request.scope["server"]`` - when Studio starts outside ``run_server``.
|
||||
2. ``request.scope["server"]`` - when Unsloth starts outside ``run_server``.
|
||||
3. parsed ``request.base_url`` - last resort for test fixtures.
|
||||
"""
|
||||
port: Any = getattr(request.app.state, "server_port", None)
|
||||
|
|
|
|||
|
|
@ -485,7 +485,7 @@ async def upload_dataset(
|
|||
|
||||
# Stream to disk in chunks to avoid holding the whole file in memory. The
|
||||
# route-level cap gives a clear training-dataset error and avoids leaving
|
||||
# oversized partial files in the Studio uploads directory.
|
||||
# oversized partial files in the Unsloth uploads directory.
|
||||
upload_limit_bytes = get_upload_limit_bytes()
|
||||
total_bytes = 0
|
||||
upload_complete = False
|
||||
|
|
|
|||
|
|
@ -92,7 +92,7 @@ def _mlx_distributed_launch_detected() -> bool:
|
|||
def _install_httpcore_asyncgen_silencer() -> None:
|
||||
"""Silence benign httpx/httpcore asyncgen GC noise on Python 3.13.
|
||||
|
||||
When Studio proxies a llama-server stream via httpx, the innermost
|
||||
When Unsloth proxies a llama-server stream via httpx, the innermost
|
||||
``HTTP11ConnectionByteStream.__aiter__`` async generator is finalised by
|
||||
the asyncgen GC hook on a task different from the one that opened it. Its
|
||||
``aclose`` calls ``anyio.Lock.acquire`` → ``cancel_shielded_checkpoint``,
|
||||
|
|
@ -229,14 +229,14 @@ def _friendly_upstream_error(text: str) -> str:
|
|||
parse grammar" / "failed to initialize samplers"). This surfaces to coding agents as
|
||||
a hard 400 on every tool-bearing turn. It is a llama-server limitation with some
|
||||
model/quant + tool-schema combinations, and recent llama.cpp builds handle the common
|
||||
coding-agent tools, so point the user at updating Studio rather than the raw body.
|
||||
coding-agent tools, so point the user at updating Unsloth rather than the raw body.
|
||||
"""
|
||||
lowered = text.lower()
|
||||
if "failed to parse grammar" in lowered or "failed to initialize samplers" in lowered:
|
||||
return (
|
||||
"The model couldn't compile a tool-calling grammar for this request. This is a "
|
||||
"llama-server limitation with some model/quant and tool-schema combinations. "
|
||||
"Update Studio (it installs the latest llama.cpp, which handles the common "
|
||||
"Update Unsloth (it installs the latest llama.cpp, which handles the common "
|
||||
"coding-agent tools) or try a different GGUF model."
|
||||
)
|
||||
return f"llama-server error: {text}"
|
||||
|
|
@ -731,7 +731,7 @@ def _openai_passthrough_sse_line_terminal_state(raw_line: str) -> Optional[str]:
|
|||
|
||||
Some llama-server builds can emit the logical final chunk (``finish_reason``)
|
||||
and optional usage chunk, then keep the HTTP stream open without sending the
|
||||
OpenAI ``data: [DONE]`` sentinel. Classifying those chunks lets Studio close
|
||||
OpenAI ``data: [DONE]`` sentinel. Classifying those chunks lets Unsloth close
|
||||
the client stream promptly while preserving an optional trailing usage chunk.
|
||||
"""
|
||||
if not raw_line.startswith("data:"):
|
||||
|
|
@ -1786,7 +1786,7 @@ import numpy as np
|
|||
from datetime import date as _date
|
||||
|
||||
router = APIRouter()
|
||||
# Studio-only router (not mounted on /v1 OpenAI-compat).
|
||||
# Unsloth-only router (not mounted on /v1 OpenAI-compat).
|
||||
studio_router = APIRouter()
|
||||
|
||||
|
||||
|
|
@ -2108,9 +2108,9 @@ def _effective_enable_tools(payload) -> Optional[bool]:
|
|||
|
||||
|
||||
def _explicit_studio_tool_loop_requested(payload) -> bool:
|
||||
"""True when the request itself asks Studio to execute local tools.
|
||||
"""True when the request itself asks Unsloth to execute local tools.
|
||||
|
||||
Process-wide CLI policy can default Studio's tool loop on for ordinary chat,
|
||||
Process-wide CLI policy can default Unsloth's tool loop on for ordinary chat,
|
||||
but it must not steal OpenAI-compatible client tools or response_format
|
||||
requests from the llama-server passthrough path. A policy of ``False``
|
||||
(--disable-tools) vetoes even an explicit ``enable_tools: true`` ask.
|
||||
|
|
@ -2122,7 +2122,7 @@ def _explicit_studio_tool_loop_requested(payload) -> bool:
|
|||
|
||||
|
||||
def _permission_mode_confirm(payload) -> bool:
|
||||
"""Effective confirm-gate intent for Studio's own local tool loop.
|
||||
"""Effective confirm-gate intent for Unsloth's own local tool loop.
|
||||
|
||||
Honors the documented default that an unset permission_mode behaves as
|
||||
"ask". An explicit confirm_tool_calls (True or False) wins; explicit
|
||||
|
|
@ -2144,7 +2144,7 @@ def _permission_mode_confirm(payload) -> bool:
|
|||
|
||||
|
||||
def _confirm_gate_needs_stream(payload) -> bool:
|
||||
"""Whether Studio's local tool-loop confirm gate still requires stream=true.
|
||||
"""Whether Unsloth's local tool-loop confirm gate still requires stream=true.
|
||||
|
||||
The gate can only prompt while streaming, so a non-streaming request that will
|
||||
prompt must 400 up front. auto ("Approve for me") only prompts for a call the
|
||||
|
|
@ -3143,7 +3143,7 @@ def _is_explicit_tensor_drop(request: LoadRequest) -> bool:
|
|||
"""True only when the request explicitly selects a non-tensor --split-mode (e.g.
|
||||
layer/row/none), a deliberate departure from a preserved tensor->layer fallback.
|
||||
|
||||
A bare tensor_parallel field is NOT a drop: the Studio UI always sends it and echoes
|
||||
A bare tensor_parallel field is NOT a drop: the Unsloth UI always sends it and echoes
|
||||
the /load response's resolved value back, so after a fallback every reload carries
|
||||
tensor_parallel=false even though the user never changed it -- treating that as a drop
|
||||
would collapse the preserved multi-GPU placement on the next ctx/settings reload. An
|
||||
|
|
@ -4197,8 +4197,8 @@ async def _load_model_impl(request: LoadRequest, fastapi_request: Request, curre
|
|||
raise HTTPException(
|
||||
status_code = 400,
|
||||
detail = (
|
||||
"Studio does not support distributed MLX inference under "
|
||||
"mlx.launch. Use `mlx.launch ... unsloth chat` or run Studio "
|
||||
"Unsloth does not support distributed MLX inference under "
|
||||
"mlx.launch. Use `mlx.launch ... unsloth chat` or run Unsloth "
|
||||
"without the distributed launcher."
|
||||
),
|
||||
)
|
||||
|
|
@ -4288,7 +4288,7 @@ async def _load_model_impl(request: LoadRequest, fastapi_request: Request, curre
|
|||
# omits chat_template_override, so strip the inherited
|
||||
# --chat-template-file in that case too -- otherwise the stale
|
||||
# extra arg (appended last) shadows the bundled template while
|
||||
# Studio reports the bundled template's capabilities.
|
||||
# Unsloth reports the bundled template's capabilities.
|
||||
fields_set = getattr(request, "model_fields_set", set())
|
||||
stripped = strip_shadowing_flags(
|
||||
llama_backend.extra_args,
|
||||
|
|
@ -4727,7 +4727,7 @@ def _requires_trust_remote_code_for_model(
|
|||
model_identifier: str, hf_token: Optional[str] = None
|
||||
) -> bool:
|
||||
"""Whether loading this model would execute custom repo code, so the consent
|
||||
dialog must run first. True if the Studio YAML default enables
|
||||
dialog must run first. True if the Unsloth YAML default enables
|
||||
``trust_remote_code`` OR the raw config declares an ``auto_map`` (Hub/local,
|
||||
config.json or tokenizer_config.json). Reads raw JSON only; never imports
|
||||
model code."""
|
||||
|
|
@ -5360,7 +5360,7 @@ async def confirm_tool_call(
|
|||
|
||||
@studio_router.get("/monitor")
|
||||
async def get_api_monitor(current_subject: str = Depends(get_current_subject)):
|
||||
"""Return recent OpenAI-compatible API activity for Studio."""
|
||||
"""Return recent OpenAI-compatible API activity for Unsloth."""
|
||||
active_model = _monitor_active_model()
|
||||
active_requests = api_monitor.active_count(subject = current_subject)
|
||||
if active_requests:
|
||||
|
|
@ -5548,7 +5548,7 @@ async def get_status(current_subject: str = Depends(get_current_subject)):
|
|||
_display_model_id = os.path.basename(_model_id)
|
||||
_inference_cfg = load_inference_config(_model_id) if _model_id else None
|
||||
_audio_type = getattr(llama_backend, "_audio_type", None)
|
||||
# Don't surface Studio's auto-applied bundled family template (e.g. the
|
||||
# Don't surface Unsloth's auto-applied bundled family template (e.g. the
|
||||
# gemma-4 override) as a user-authored override: the frontend adopts
|
||||
# status.chat_template_override as editable state and would otherwise
|
||||
# re-send it as an explicit override for a later, unrelated model. Only
|
||||
|
|
@ -6173,7 +6173,7 @@ def _build_external_messages(
|
|||
metadata; strip it for providers that can't parse the unknown key.
|
||||
2. Marked server-side builtin cards (`_server_tool: true` on a
|
||||
canonical builtin name, or a Gemini `native_part` payload) are
|
||||
Studio-internal tool cards from a prior native Gemini turn;
|
||||
Unsloth-internal tool cards from a prior native Gemini turn;
|
||||
forwarding them to OpenAI / Anthropic / custom OAI-compat gateways
|
||||
sends an orphan `tool_calls` entry (no matching tool declaration,
|
||||
often no matching `role="tool"` reply) that can be rejected. We
|
||||
|
|
@ -6859,7 +6859,7 @@ async def openai_chat_completions(
|
|||
# is invalid and must not evict the resident model first.
|
||||
#
|
||||
# Enter the local-loop arm exactly when the passthrough router below would
|
||||
# run Studio's own tool loop. That gate is `_tools_on or _mcp_allowed`
|
||||
# run Unsloth's own tool loop. That gate is `_tools_on or _mcp_allowed`
|
||||
# (see the use_tools block): _effective_enable_tools (which lets a
|
||||
# process-wide --enable-tools policy force the loop on) plus mcp_enabled
|
||||
# honoring --disable-tools, and tool_choice="none" disabling it unless the
|
||||
|
|
@ -6884,7 +6884,7 @@ async def openai_chat_completions(
|
|||
or bool(payload.openai_code_exec_container_id)
|
||||
or bool(payload.anthropic_code_exec_container_id)
|
||||
# A JSON-schema response_format is guided-decoding structured output the
|
||||
# router forwards to the llama-server passthrough, not Studio's tool
|
||||
# router forwards to the llama-server passthrough, not Unsloth's tool
|
||||
# loop, so a --enable-tools policy must not 400 it as a local-confirm
|
||||
# request under ask/auto.
|
||||
or bool(_extract_response_format(payload))
|
||||
|
|
@ -6961,7 +6961,7 @@ async def openai_chat_completions(
|
|||
using_gguf = llama_backend.is_loaded
|
||||
|
||||
# OpenAI-SDK clients send ``chat_template_kwargs`` via ``extra_body``, which
|
||||
# the SDK spreads into the request body at the top level. Studio's
|
||||
# the SDK spreads into the request body at the top level. Unsloth's
|
||||
# ChatCompletionRequest has ``extra="allow"`` so pydantic stashes them in
|
||||
# ``model_extra``, but downstream generators consume the typed
|
||||
# ``payload.enable_thinking``. Lift ``enable_thinking`` from the extra-body
|
||||
|
|
@ -7215,7 +7215,7 @@ async def openai_chat_completions(
|
|||
|
||||
# ── Standard OpenAI function-calling pass-through (GGUF only) ────
|
||||
# When a client (opencode / Claude Code via OpenAI compat / Cursor /
|
||||
# Continue / ...) sends standard OpenAI `tools` without Studio's
|
||||
# Continue / ...) sends standard OpenAI `tools` without Unsloth's
|
||||
# `enable_tools` shorthand, forward the request to llama-server
|
||||
# verbatim so structured `tool_calls` flow back to the client. This
|
||||
# branch runs BEFORE `_extract_content_parts` because that helper is
|
||||
|
|
@ -7238,7 +7238,7 @@ async def openai_chat_completions(
|
|||
_has_tool_catalog = bool(payload.tools and len(payload.tools) > 0)
|
||||
_has_active_tool_catalog = _has_tool_catalog and payload.tool_choice != "none"
|
||||
_has_client_tool_contract = _has_active_tool_catalog or _has_tool_messages
|
||||
# The Studio tool loop needs a tool-capable backend, so a request that asks
|
||||
# The Unsloth tool loop needs a tool-capable backend, so a request that asks
|
||||
# for it on a backend that can't run it (DiffusionGemma forces supports_tools
|
||||
# off) must not steal client tools from the passthrough (#6851).
|
||||
_studio_tool_loop_requested = (
|
||||
|
|
@ -7434,7 +7434,7 @@ async def openai_chat_completions(
|
|||
use_tools = False
|
||||
|
||||
if use_tools:
|
||||
# permission_mode ask/auto require the confirm gate for Studio's own
|
||||
# permission_mode ask/auto require the confirm gate for Unsloth's own
|
||||
# tool loop. The request validator self-enables confirm only for
|
||||
# request-level tool signals (enable_tools/enabled_tools/mcp_enabled);
|
||||
# when a CLI policy (--enable-tools) forces the loop on without those,
|
||||
|
|
@ -8711,7 +8711,7 @@ async def openai_chat_completions(
|
|||
_sf_model_info = backend.models.get(backend.active_model_name, {})
|
||||
_sf_tpl = (_sf_model_info.get("chat_template_info") or {}).get("template")
|
||||
# Named templates may expose native reasoning only in their ``tool_use``
|
||||
# branch. Use a truthy placeholder for Studio-managed tools, whose concrete
|
||||
# branch. Use a truthy placeholder for Unsloth-managed tools, whose concrete
|
||||
# schemas are selected below, and the request schemas for client passthrough.
|
||||
_sf_server_tool_intent = bool(
|
||||
_effective_enable_tools(payload) or _explicit_studio_tool_loop_requested(payload)
|
||||
|
|
@ -8790,7 +8790,7 @@ async def openai_chat_completions(
|
|||
_sf_use_tools = False
|
||||
|
||||
if _sf_use_tools:
|
||||
# permission_mode ask/auto require the confirm gate for Studio's own tool
|
||||
# permission_mode ask/auto require the confirm gate for Unsloth's own tool
|
||||
# loop; when a CLI policy (--enable-tools) forces the loop on without a
|
||||
# request-level tool signal, derive confirm here so the mode still gates
|
||||
# the call (matching the GGUF path). off/full never prompt.
|
||||
|
|
@ -12050,7 +12050,7 @@ def _anthropic_requested_studio_tools(tools: Optional[list]) -> set[str]:
|
|||
def _select_anthropic_server_tools(
|
||||
all_tools: list[dict], requested_studio_tools: set[str], enabled_tools: Optional[list[str]]
|
||||
) -> list[dict]:
|
||||
"""Select Studio tools requested through Anthropic tools and extensions."""
|
||||
"""Select Unsloth tools requested through Anthropic tools and extensions."""
|
||||
if not requested_studio_tools and enabled_tools is None:
|
||||
return all_tools
|
||||
|
||||
|
|
@ -12289,7 +12289,7 @@ async def anthropic_messages(
|
|||
),
|
||||
)
|
||||
|
||||
# Reject an unsupported confirm-gated permission mode for Studio's own
|
||||
# Reject an unsupported confirm-gated permission mode for Unsloth's own
|
||||
# ("server") Anthropic tools before the switch, mirroring the malformed- and
|
||||
# mixed-tool checks above. ask always wants a per-call pause this passthrough
|
||||
# cannot offer, so it 400s whenever server tools are selected. auto only needs
|
||||
|
|
@ -12778,11 +12778,11 @@ async def _anthropic_tool_stream(
|
|||
ends_on_tool_use = True
|
||||
elif etype == "tool_end":
|
||||
tool_blocks_emitted += 1
|
||||
# A tool_end means Studio executed the tool server-side, so
|
||||
# A tool_end means Unsloth executed the tool server-side, so
|
||||
# the response no longer ends on a pending client action.
|
||||
# Without this, a server tool that produces no trailing text
|
||||
# would be mislabeled stop_reason "tool_use", telling the
|
||||
# client to run a tool Studio already ran.
|
||||
# client to run a tool Unsloth already ran.
|
||||
ends_on_tool_use = False
|
||||
elif etype == "content" and event.get("text"):
|
||||
ends_on_tool_use = False
|
||||
|
|
@ -13698,7 +13698,7 @@ def _openai_messages_for_passthrough(payload) -> list[dict]:
|
|||
structured ``tool_calls``. Content-parts images already in the list are
|
||||
left untouched.
|
||||
|
||||
When a client uses Studio's legacy ``image_base64`` top-level field, the
|
||||
When a client uses Unsloth's legacy ``image_base64`` top-level field, the
|
||||
image is re-encoded to PNG (llama-server's stb_image has limited format
|
||||
support) and spliced into the last user message as an OpenAI ``image_url``
|
||||
content part so vision + function-calling requests work transparently.
|
||||
|
|
@ -13832,7 +13832,7 @@ def _build_openai_passthrough_body(
|
|||
) -> dict:
|
||||
"""Assemble the llama-server request body from a ChatCompletionRequest.
|
||||
|
||||
Only known OpenAI / llama-server fields are forwarded, so Studio-specific
|
||||
Only known OpenAI / llama-server fields are forwarded, so Unsloth-specific
|
||||
extensions (``enable_tools``, ``enabled_tools``, ``session_id``, ...) never
|
||||
leak to the backend.
|
||||
"""
|
||||
|
|
@ -14082,7 +14082,7 @@ async def _openai_passthrough_stream_admitted(
|
|||
admission_lease: LlamaAdmissionLease,
|
||||
tracker,
|
||||
):
|
||||
"""Streaming client-side pass-through after Studio granted an upstream slot.
|
||||
"""Streaming client-side pass-through after Unsloth granted an upstream slot.
|
||||
|
||||
Forwards the client's OpenAI function-calling request to llama-server and
|
||||
relays the SSE stream back with minimal normalization (reasoning-only
|
||||
|
|
|
|||
|
|
@ -82,7 +82,7 @@ def _validate_url(url: str) -> str:
|
|||
if _looks_like_command(trimmed):
|
||||
detail = (
|
||||
"Local commands aren't enabled on this server. To allow them, "
|
||||
"set UNSLOTH_STUDIO_ALLOW_STDIO_MCP=1 and restart Studio, or use "
|
||||
"set UNSLOTH_STUDIO_ALLOW_STDIO_MCP=1 and restart Unsloth, or use "
|
||||
"an http:// or https:// URL instead."
|
||||
)
|
||||
else:
|
||||
|
|
|
|||
|
|
@ -544,7 +544,7 @@ def _ollama_links_dir(ollama_dir: Path) -> Optional[Path]:
|
|||
"""Return a writable directory for Ollama ``.gguf`` symlinks.
|
||||
|
||||
Prefers ``<ollama_dir>/.studio_links/`` so links sit next to their
|
||||
blobs; falls back to a per-ollama-dir namespace under Studio's cache
|
||||
blobs; falls back to a per-ollama-dir namespace under Unsloth's cache
|
||||
when the models dir is read-only (common for system installs).
|
||||
"""
|
||||
from utils.paths.storage_roots import cache_root
|
||||
|
|
@ -555,7 +555,7 @@ def _ollama_links_dir(ollama_dir: Path) -> Optional[Path]:
|
|||
return primary
|
||||
except OSError as e:
|
||||
logger.debug(
|
||||
"Ollama dir %s not writable for .studio_links (%s); falling back to Studio cache",
|
||||
"Ollama dir %s not writable for .studio_links (%s); falling back to Unsloth cache",
|
||||
ollama_dir,
|
||||
e,
|
||||
)
|
||||
|
|
@ -594,7 +594,7 @@ def _scan_ollama_dir(ollama_dir: Path, limit: Optional[int] = None) -> List[Loca
|
|||
model, keyed by a short hash of the manifest path, so
|
||||
``detect_mmproj_file`` only sees that model's projector). Links are
|
||||
symlinks when possible, else hardlinks; the link dir is
|
||||
``.studio_links/`` when writable, else Studio's cache.
|
||||
``.studio_links/`` when writable, else Unsloth's cache.
|
||||
"""
|
||||
manifests_root = ollama_dir / "manifests"
|
||||
if not manifests_root.is_dir():
|
||||
|
|
@ -1194,7 +1194,7 @@ def _build_browse_allowlist(
|
|||
"""Return the root directories the folder browser may walk.
|
||||
|
||||
The same list seeds the sidebar suggestion chips, so chip targets are
|
||||
always reachable. Roots: HOME, resolved HF cache dirs, Studio's
|
||||
always reachable. Roots: HOME, resolved HF cache dirs, Unsloth's
|
||||
outputs/exports/studio root, registered scan folders, and well-known
|
||||
local-LLM dirs (LM Studio, Ollama, ``~/models``); each added only if
|
||||
it resolves to a real directory.
|
||||
|
|
@ -1486,7 +1486,7 @@ def browse_folders(
|
|||
"Directory to list. If omitted, defaults to the current user's "
|
||||
"home directory. Tilde (`~`) and relative paths are expanded. "
|
||||
"Must resolve inside the allowlist of browseable roots (HOME, "
|
||||
"HF cache, Studio dirs, registered scan folders, well-known "
|
||||
"HF cache, Unsloth dirs, registered scan folders, well-known "
|
||||
"model dirs)."
|
||||
),
|
||||
),
|
||||
|
|
@ -2251,15 +2251,15 @@ async def delete_finetuned_model(
|
|||
gguf_variant: Optional[str] = Body(None),
|
||||
current_subject: str = Depends(get_current_subject),
|
||||
):
|
||||
"""Delete a Studio-trained or exported model from disk.
|
||||
"""Delete an Unsloth-trained or exported model from disk.
|
||||
|
||||
Only paths under Studio's outputs/exports roots are accepted.
|
||||
Only paths under Unsloth's outputs/exports roots are accepted.
|
||||
Exported GGUF entries can delete one quant variant at a time.
|
||||
"""
|
||||
if source not in {"training", "exported"}:
|
||||
raise HTTPException(
|
||||
status_code = 400,
|
||||
detail = "Only trained or exported Studio models can be deleted",
|
||||
detail = "Only trained or exported Unsloth models can be deleted",
|
||||
)
|
||||
|
||||
if not model_path or not model_path.strip():
|
||||
|
|
@ -2291,14 +2291,14 @@ async def delete_finetuned_model(
|
|||
if not _is_path_under_lexically(delete_path, allowed_root):
|
||||
raise HTTPException(
|
||||
status_code = 400,
|
||||
detail = "Model path is outside Studio storage",
|
||||
detail = "Model path is outside Unsloth storage",
|
||||
)
|
||||
if export_type == "gguf" and gguf_variant:
|
||||
target_path = delete_path.resolve()
|
||||
if not _is_path_under(target_path, allowed_root):
|
||||
raise HTTPException(
|
||||
status_code = 400,
|
||||
detail = "Model path is outside Studio storage",
|
||||
detail = "Model path is outside Unsloth storage",
|
||||
)
|
||||
else:
|
||||
target_path = delete_path
|
||||
|
|
@ -2311,7 +2311,7 @@ async def delete_finetuned_model(
|
|||
if should_check_resolved_path and not _is_path_under(target_path, allowed_root):
|
||||
raise HTTPException(
|
||||
status_code = 400,
|
||||
detail = "Model path is outside Studio storage",
|
||||
detail = "Model path is outside Unsloth storage",
|
||||
)
|
||||
if target_path == allowed_root:
|
||||
raise HTTPException(
|
||||
|
|
@ -3456,7 +3456,7 @@ _EXPORT_SIZE_CACHE: dict[str, tuple[int, int, str]] = {}
|
|||
|
||||
|
||||
def _is_sizable_local_path(model: str) -> bool:
|
||||
"""True only for local paths under a Studio data root.
|
||||
"""True only for local paths under an Unsloth data root.
|
||||
|
||||
Containment is decided lexically (no filesystem access) before the path is
|
||||
touched, then the path is symlink-resolved and re-checked so a symlink
|
||||
|
|
|
|||
|
|
@ -127,9 +127,9 @@ async def start_training(
|
|||
try:
|
||||
logger.info(f"Starting training job with model: {request.model_name}")
|
||||
|
||||
# When Studio is driven as an inference API (API-key auth), refuse to start
|
||||
# When Unsloth is driven as an inference API (API-key auth), refuse to start
|
||||
# training while a request is in flight: training frees VRAM by unloading
|
||||
# the chat model, which would kill the stream. The Studio UI (session auth)
|
||||
# the chat model, which would kill the stream. The Unsloth UI (session auth)
|
||||
# still starts training and coexists/frees VRAM as before. (A mixed UI+API
|
||||
# session is not yet special-cased.)
|
||||
if via_api_key is True:
|
||||
|
|
@ -139,7 +139,7 @@ async def start_training(
|
|||
status_code = 409,
|
||||
detail = (
|
||||
"Cannot start training over the API while an inference request is in "
|
||||
"progress. Wait for it to finish, or start training from the Studio UI."
|
||||
"progress. Wait for it to finish, or start training from the Unsloth UI."
|
||||
),
|
||||
)
|
||||
|
||||
|
|
|
|||
|
|
@ -232,7 +232,7 @@ def _working_local_url(port: int) -> "str | None":
|
|||
def _localhost_ipv6_mismatch_url(bind_host: str, port: int) -> "str | None":
|
||||
"""Return the IPv4 loopback URL when localhost won't reach 127.0.0.1.
|
||||
|
||||
Local Studio binds to 127.0.0.1. Where localhost resolves to IPv6 only (::1),
|
||||
Local Unsloth binds to 127.0.0.1. Where localhost resolves to IPv6 only (::1),
|
||||
http://localhost:<port> fails (or hits a different process on ::1) even though
|
||||
http://127.0.0.1:<port> works. Return the IPv4 URL for the caller to surface.
|
||||
"""
|
||||
|
|
@ -243,7 +243,7 @@ def _localhost_ipv6_mismatch_url(bind_host: str, port: int) -> "str | None":
|
|||
|
||||
ipv4_url = f"http://127.0.0.1:{port}"
|
||||
|
||||
# Only warn once Studio is confirmed answering on IPv4 loopback.
|
||||
# Only warn once Unsloth is confirmed answering on IPv4 loopback.
|
||||
if _working_local_url(port) != ipv4_url:
|
||||
return None
|
||||
|
||||
|
|
@ -265,7 +265,7 @@ def _localhost_ipv6_mismatch_url(bind_host: str, port: int) -> "str | None":
|
|||
if host == "::1":
|
||||
has_ipv6_loopback = True
|
||||
|
||||
# A connection to ::1 is NOT evidence Studio is reachable there: Studio binds
|
||||
# A connection to ::1 is NOT evidence Unsloth is reachable there: Unsloth binds
|
||||
# 127.0.0.1 only, so anything on ::1 is a different process. Dual-stack
|
||||
# localhost is fine (browsers fall back to 127.0.0.1), so only the IPv6-only
|
||||
# case strands the user.
|
||||
|
|
@ -287,7 +287,7 @@ def _stdout_color_ok() -> bool:
|
|||
|
||||
|
||||
def _print_localhost_ipv6_mismatch_warning(local_url: str, port: int) -> None:
|
||||
"""Warn that localhost points at ::1 while Studio is bound to 127.0.0.1."""
|
||||
"""Warn that localhost points at ::1 while Unsloth is bound to 127.0.0.1."""
|
||||
use_color = _stdout_color_ok()
|
||||
warn_c = "\033[38;5;215;1m" if use_color else ""
|
||||
reset = "\033[0m" if use_color else ""
|
||||
|
|
@ -303,7 +303,7 @@ def _print_localhost_ipv6_mismatch_warning(local_url: str, port: int) -> None:
|
|||
def _verify_global_reachability(display_host: str, port: int) -> None:
|
||||
"""Probe check-host.net to confirm display_host:port is reachable from the
|
||||
public internet. Synchronous so output lands between the banner URLs and the
|
||||
stop hint. Bounded at ~15s; failures swallowed (verifier failing != Studio
|
||||
stop hint. Bounded at ~15s; failures swallowed (verifier failing != Unsloth
|
||||
failing). Only meaningful for a wildcard bind."""
|
||||
global _public_reachable
|
||||
# Reset to "unknown" each run; set True/False only when the probe decides.
|
||||
|
|
@ -563,15 +563,15 @@ def _print_cloudflare_line(secure: bool = False, loopback_host: str = "127.0.0.1
|
|||
" Cloudflare tunnel: ON. This Cloudflare URL is PUBLIC, and the "
|
||||
"raw port is also publicly reachable. --no-cloudflare disables "
|
||||
f"only the Cloudflare URL; bind {loopback_host} or close firewall "
|
||||
"access to keep Studio private.",
|
||||
"access to keep Unsloth private.",
|
||||
warn,
|
||||
)
|
||||
else:
|
||||
_emit(
|
||||
" Cloudflare tunnel: ON. This is a PUBLIC internet URL: anyone "
|
||||
"who has it can reach this Studio. Relaunch with --no-cloudflare "
|
||||
"who has it can reach this Unsloth. Relaunch with --no-cloudflare "
|
||||
f"to disable the Cloudflare URL; bind {loopback_host} or close "
|
||||
"firewall access to keep Studio private.",
|
||||
"firewall access to keep Unsloth private.",
|
||||
warn,
|
||||
)
|
||||
return
|
||||
|
|
@ -580,12 +580,12 @@ def _print_cloudflare_line(secure: bool = False, loopback_host: str = "127.0.0.1
|
|||
_emit(
|
||||
" Cloudflare tunnel: requested but failed to start. The raw port is "
|
||||
"still reachable from the public internet (see the reachability check "
|
||||
"above): anyone who can reach it can access this Studio.",
|
||||
"above): anyone who can reach it can access this Unsloth.",
|
||||
warn,
|
||||
)
|
||||
elif _public_reachable is False:
|
||||
_emit(
|
||||
" Cloudflare tunnel: requested but failed to start. Studio is reachable "
|
||||
" Cloudflare tunnel: requested but failed to start. Unsloth is reachable "
|
||||
"on your local network only (no public link).",
|
||||
warn,
|
||||
)
|
||||
|
|
@ -593,7 +593,7 @@ def _print_cloudflare_line(secure: bool = False, loopback_host: str = "127.0.0.1
|
|||
_emit(
|
||||
" Cloudflare tunnel: requested but failed to start. There is no "
|
||||
"Cloudflare public link. Raw port reachability was not verified; "
|
||||
f"bind {loopback_host} or close firewall access to keep Studio private.",
|
||||
f"bind {loopback_host} or close firewall access to keep Unsloth private.",
|
||||
warn,
|
||||
)
|
||||
elif _cloudflare_flag:
|
||||
|
|
@ -601,19 +601,19 @@ def _print_cloudflare_line(secure: bool = False, loopback_host: str = "127.0.0.1
|
|||
_emit(
|
||||
" Cloudflare tunnel: OFF for this mode. The raw port is still "
|
||||
"reachable from the public internet (see the reachability check above): "
|
||||
"anyone who can reach it can access this Studio.",
|
||||
"anyone who can reach it can access this Unsloth.",
|
||||
warn,
|
||||
)
|
||||
elif _public_reachable is False:
|
||||
_emit(
|
||||
" Cloudflare tunnel: OFF for this mode. Studio is reachable on your "
|
||||
" Cloudflare tunnel: OFF for this mode. Unsloth is reachable on your "
|
||||
"local network only (no public link)."
|
||||
)
|
||||
else:
|
||||
_emit(
|
||||
" Cloudflare tunnel: OFF for this mode. There is no Cloudflare public "
|
||||
"link. Raw port reachability was not verified; "
|
||||
f"bind {loopback_host} or close firewall access to keep Studio private.",
|
||||
f"bind {loopback_host} or close firewall access to keep Unsloth private.",
|
||||
warn,
|
||||
)
|
||||
elif _cloudflare_flag is False or _cloudflare_flag is None:
|
||||
|
|
@ -624,12 +624,12 @@ def _print_cloudflare_line(secure: bool = False, loopback_host: str = "127.0.0.1
|
|||
f" Cloudflare tunnel: OFF ({_reason}). The raw port is still "
|
||||
"reachable from the public internet (see the reachability check above): "
|
||||
"pass --cloudflare to also expose a public Cloudflare HTTPS link, or "
|
||||
f"bind {loopback_host} to keep Studio private.",
|
||||
f"bind {loopback_host} to keep Unsloth private.",
|
||||
warn,
|
||||
)
|
||||
elif _public_reachable is False:
|
||||
_emit(
|
||||
f" Cloudflare tunnel: OFF ({_reason}). Studio is reachable on your "
|
||||
f" Cloudflare tunnel: OFF ({_reason}). Unsloth is reachable on your "
|
||||
"local network only. Pass --cloudflare to expose a public "
|
||||
"Cloudflare HTTPS link."
|
||||
)
|
||||
|
|
@ -638,7 +638,7 @@ def _print_cloudflare_line(secure: bool = False, loopback_host: str = "127.0.0.1
|
|||
f" Cloudflare tunnel: OFF ({_reason}). There is no Cloudflare "
|
||||
"public link. Raw port reachability was not verified; pass --cloudflare "
|
||||
"to expose a public Cloudflare HTTPS link, or "
|
||||
f"bind {loopback_host} or close firewall access to keep Studio private.",
|
||||
f"bind {loopback_host} or close firewall access to keep Unsloth private.",
|
||||
warn,
|
||||
)
|
||||
|
||||
|
|
@ -674,7 +674,7 @@ def _is_port_free(host: str, port: int) -> bool:
|
|||
|
||||
For a ``0.0.0.0`` wildcard host, also check whether anything is listening on
|
||||
``127.0.0.1`` (and ``::1`` when IPv6 exists): an SSH tunnel may hold loopback
|
||||
while the wildcard bind succeeds, making Studio unreachable via ``localhost``.
|
||||
while the wildcard bind succeeds, making Unsloth unreachable via ``localhost``.
|
||||
"""
|
||||
import socket
|
||||
|
||||
|
|
@ -1087,7 +1087,7 @@ def _terminal_password_gate(
|
|||
) -> Tuple[bool, bool]:
|
||||
"""Force a terminal password change before the public tunnel goes up.
|
||||
|
||||
When the tunnel is about to publish Studio and the seeded admin password was
|
||||
When the tunnel is about to publish Unsloth and the seeded admin password was
|
||||
never changed, ask for a new one (masked, confirmed) before any public URL
|
||||
exists. The CLI normally does this before re-exec'ing the backend; this is
|
||||
the backstop for direct `python run.py` launches and older-CLI installs.
|
||||
|
|
@ -1147,7 +1147,7 @@ def _terminal_password_gate(
|
|||
)
|
||||
if not deadline_arms:
|
||||
print(
|
||||
"Refusing to publish Studio on a public Cloudflare URL: the "
|
||||
"Refusing to publish Unsloth on a public Cloudflare URL: the "
|
||||
"default admin password was never changed, no terminal is "
|
||||
"attached to change it here, and the bootstrap shutdown "
|
||||
"deadline does not apply to this launch (api-only, or "
|
||||
|
|
@ -1163,11 +1163,11 @@ def _terminal_password_gate(
|
|||
# terminal-attached run / reset-password instead of reading it from disk.
|
||||
print(
|
||||
" WARNING: the default admin password is still active while "
|
||||
"Studio is about to be published on a public Cloudflare URL, and "
|
||||
"Unsloth is about to be published on a public Cloudflare URL, and "
|
||||
"no terminal is attached to change it here. The public page will "
|
||||
"NOT auto-fill the bootstrap credential. Set a new password by "
|
||||
"running `unsloth studio` locally with a terminal attached, or "
|
||||
"`unsloth studio reset-password`. Studio shuts down after the "
|
||||
"`unsloth studio reset-password`. Unsloth shuts down after the "
|
||||
"bootstrap deadline (UNSLOTH_STUDIO_BOOTSTRAP_TIMEOUT, default 1h) "
|
||||
"unless the password is changed.",
|
||||
file = sys.stderr,
|
||||
|
|
@ -1222,7 +1222,7 @@ def _apply_supplied_password(password_value: "Optional[str]") -> None:
|
|||
_auth_storage.ensure_default_admin()
|
||||
if not _auth_storage.requires_password_change(_admin):
|
||||
print(
|
||||
"Error: a Studio admin password is already set; --password only sets "
|
||||
"Error: an Unsloth admin password is already set; --password only sets "
|
||||
"the initial password. Run `unsloth studio reset-password` first.",
|
||||
file = sys.stderr,
|
||||
flush = True,
|
||||
|
|
@ -1337,7 +1337,7 @@ def run_server(
|
|||
pass
|
||||
|
||||
# Persist a session log + native-crash stacks BEFORE importing main, so
|
||||
# even import-time failures leave evidence on disk. Field report: Studio
|
||||
# even import-time failures leave evidence on disk. Field report: Unsloth
|
||||
# "terminates without a warning" -- a native crash in the GPU runtime
|
||||
# kills the process with no Python traceback, and a desktop-shortcut
|
||||
# console closes before anything can be read. Console-only logging made
|
||||
|
|
@ -1406,7 +1406,7 @@ def run_server(
|
|||
ensure_studio_directories()
|
||||
|
||||
logger.info(
|
||||
"Ensured Studio directories in %.1fms",
|
||||
"Ensured Unsloth directories in %.1fms",
|
||||
(time.perf_counter() - boot_started) * 1000,
|
||||
)
|
||||
|
||||
|
|
@ -1455,7 +1455,7 @@ def run_server(
|
|||
installer_bin = home / "unsloth_studio" / "bin" / "unsloth"
|
||||
tried_lines = "\n".join(f" - {p}" for p in attempted) or " (none)"
|
||||
raise SystemExit(
|
||||
"[ERROR] Studio frontend build not found.\n"
|
||||
"[ERROR] Unsloth frontend build not found.\n"
|
||||
f"Tried:\n{tried_lines}\n"
|
||||
"\n"
|
||||
"Likely cause: another 'unsloth' on PATH is shadowing the "
|
||||
|
|
@ -1557,7 +1557,7 @@ def run_server(
|
|||
)
|
||||
if not _pw_proceed:
|
||||
print(
|
||||
"Not starting Studio; set a new admin password first, or launch "
|
||||
"Not starting Unsloth; set a new admin password first, or launch "
|
||||
"without --secure/--cloudflare.",
|
||||
file = sys.stderr,
|
||||
flush = True,
|
||||
|
|
@ -1695,7 +1695,7 @@ def run_server(
|
|||
logger = logger,
|
||||
)
|
||||
logger.info(
|
||||
"Studio will shut down in %ds unless the default admin password is changed.",
|
||||
"Unsloth will shut down in %ds unless the default admin password is changed.",
|
||||
_bootstrap_timeout,
|
||||
)
|
||||
except Exception as e: # best-effort: never block startup on the timeout
|
||||
|
|
@ -1753,11 +1753,11 @@ def _build_arg_parser():
|
|||
"--cloudflare",
|
||||
action = argparse.BooleanOptionalAction,
|
||||
default = None,
|
||||
help = "Expose Studio on a PUBLIC internet URL via a free Cloudflare HTTPS "
|
||||
help = "Expose Unsloth on a PUBLIC internet URL via a free Cloudflare HTTPS "
|
||||
"tunnel, for non-api-only wildcard binds (0.0.0.0 or ::). Off by default; "
|
||||
"pass --cloudflare to enable it (--secure implies it), --no-cloudflare to "
|
||||
"force it off. It does not change a raw wildcard bind. If the admin "
|
||||
"password was never changed, Studio asks for a new one in the terminal "
|
||||
"password was never changed, Unsloth asks for a new one in the terminal "
|
||||
"before publishing the URL.",
|
||||
)
|
||||
parser.add_argument(
|
||||
|
|
@ -1767,7 +1767,7 @@ def _build_arg_parser():
|
|||
help = "Expose ONLY a Cloudflare HTTPS link: bind localhost and fail closed "
|
||||
"if the tunnel can't start. Without it, --no-secure also serves the raw "
|
||||
"0.0.0.0 port, which is reachable from anywhere on the network. If the "
|
||||
"admin password was never changed, Studio asks for a new one in the "
|
||||
"admin password was never changed, Unsloth asks for a new one in the "
|
||||
"terminal before publishing the URL.",
|
||||
)
|
||||
# Back-compat: accept --not-secure as a hidden alias for --no-secure.
|
||||
|
|
|
|||
|
|
@ -1,7 +1,7 @@
|
|||
# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
||||
|
||||
"""Terminal banner for Studio startup.
|
||||
"""Terminal banner for Unsloth startup.
|
||||
|
||||
Stdlib only -- safe to import without the rest of the backend.
|
||||
"""
|
||||
|
|
@ -172,7 +172,7 @@ def print_studio_access_banner(
|
|||
secondary,
|
||||
),
|
||||
style(
|
||||
" Only on trusted networks -- anyone who reaches this machine can use Studio.",
|
||||
" Only on trusted networks -- anyone who reaches this machine can use Unsloth.",
|
||||
secondary,
|
||||
),
|
||||
]
|
||||
|
|
|
|||
Some files were not shown because too many files have changed in this diff Show more
Loading…
Add table
Add a link
Reference in a new issue