* Studio: expose --parallel / -np on `unsloth studio run`
The CLI was hardcoding `llama_parallel_slots=4` in `run_kwargs` at
`unsloth_cli/commands/studio.py`, leaving users unable to tune the
concurrent decode slot count even though the engine, KV-cache math,
and `studio.backend.run.run_server(llama_parallel_slots=...)`
plumbing all already accepted any N. This change adds a `--parallel`
/ `--n-parallel` / `-np` typer option (default 4 -- matches the
previous hardcoded value), forwards it into `run_kwargs`, and pins
the new surface with 4 unit tests.
Per-request state in `routes/inference.py` is already isolated
(`cancel_event` and `prev_text` are per-request locals in every
streaming handler; the `_lock` / `_serial_load_lock` only wrap
load/unload, not chat completions), so no concurrency refactor is
needed alongside this -- the engine layer already handles N
concurrent requests on one loaded model when llama-server is told
to.
Range guards: 1 <= N <= 64. With higher N each slot gets ctx/N KV
cache; users tuning this should be aware that per-call context
shrinks proportionally.
`unsloth studio` (the bare default command, no subcommand) still
defaults to llama_parallel_slots=1 via `run_server`'s own default;
this PR does not change that path -- it only exposes the knob on the
one-liner `studio run` command that already silently used 4.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Forward --parallel through venv re-exec and drop colliding short aliases
`unsloth studio run` re-execs into the Studio venv when invoked from
outside it (the common path). The arg-builder forwards every typer
option but the new --parallel, so the child re-execs at the default 4
and any user value is silently dropped. Worse: pre-PR users who
already pass `-np N` as a pass-through extra (where llama.cpp's
last-wins parsing made it stick) silently lose N after this PR lands.
Forward --parallel explicitly in the re-exec arg list.
While auditing the re-exec path, also drop the colliding 1-char
short aliases -m (--model) and -f (--frontend) plus the redundant
-hfr. Click's short-option clustering had been silently mis-parsing
~11 llama-server short flags via the pass-through path: -fa as
`-f a`, -mg 0 as `-m g` + stray 0, -fitt 1024 as `-f itt` + stray
1024, -hff path as `-f f` + stray `-h path`, -cmoe / -cram / -sm /
-ncmoe etc. The docstring promise ("any flag this command does not
recognize is forwarded verbatim") was silently violated.
-hf (2-char) is kept because Click treats multi-char shorts atomically
(no clustering of -hff / -hfv / -hffv / -hft) and -hf is documented
in basics/api/README.md. --model / --hf-repo / --frontend long forms
all unchanged. studio_default keeps -f because it has no pass-through.
Tests:
- test_studio_run_parallel_flag.py: 8 new re-exec coverage cases
(all 3 aliases, 3 platforms via sys.platform mock, pre-PR `-np`
regression, mixed with pass-through extras).
- test_studio_run_short_alias_clashes.py (new): surface checks that
the removed shorts cannot reappear, plus 11 parametrized cases
proving each previously-broken llama-server short flag now passes
through verbatim, plus a happy-path test that documented -hf still
works for `org/repo:variant` syntax.
All 27 tests pass. Negative test (revert either fix) shows the new
tests catch the regression.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Fix stale studio run docstring describing rejected llama-server flags
The pre-PR docstring listed --port, -c / --ctx-size, --api-key, -ngl,
--jinja, --flash-attn, --no-context-shift as "rejected with HTTP 400",
but only --port and --api-key (plus other networking / auth / model
identity / single-model UI flags) are actually in
studio/backend/core/inference/llama_server_args.py's denylist. -c /
-ngl / --jinja / --flash-attn / --no-context-shift are pass-through
and last-wins-override Studio's auto-set value.
Rewrite the docstring to match the real denylist groups and point at
the canonical source. Also add --parallel to one of the examples now
that it is a first-class flag.
* ci: broaden Linux + narrow Windows llama.cpp runtime patterns + trim #5741 comments (#5746)
* ci: broaden Linux llama.cpp runtime pattern to lib*.so*
#5741 patched the explicit Linux pattern list to add
``libllama-*-impl.so*`` after ggml-org/llama.cpp#23462 (between
b9279 and b9283) split each binary's entry code into a paired
``lib<binary>-impl.so`` shared library. Same class of upstream
repackaging will hit us again whenever a new shared lib is added.
Mirror what macOS already does and replace the per-lib list with a
single ``lib*.so*`` glob. ``copy_globs`` (line 3614) unions
patterns, so the per-variant ``libggml-cuda.so*`` / ``libggml-hip.so*``
entries were never filtering anything; the spec lives in
``runtime_payload_health_groups`` (line 5209) which keeps the
explicit minimum-required list per variant.
Dry-run against b9296-bin-ubuntu-x64.tar.gz: 40 files copied (all
ggml, llama, mtmd, impl variants + the two binaries we ship), 22
skipped (other CLIs, rpc-server, LICENSE). Functionally equal to
the post-#5741 set.
* cleanup: trim #5741 comments on the pydantic split
Comments added in #5741 explained the original bug in full each
time. They are mostly redundant with the commit message and the PR.
Trim them to one short paragraph per site.
No behavior change.
* ci: narrow Windows runtime pattern to llama-server.exe + llama-quantize.exe
Studio only invokes llama-server and llama-quantize. Mac and Linux
already filter to those two binaries; Windows was the odd one out
with ``*.exe`` copying every CLI upstream ships (llama-cli,
llama-bench, llama-mtmd-cli, ...).
Dry-run on b9296 (win cpu-x64, cpu-arm64, cuda-13.1, hip-radeon):
20 unused EXEs skipped per variant, all DLLs (incl. the new
llama-*-impl.dll family) still copied via ``*.dll``.
``existing_install_matches_choice`` already checks llama-server.exe
exists explicitly (line 5297), so the health gate is unchanged.
* Lower default weight_decay in RL config from 0.01 to 0.001 (#5747)
In full FT, AdamW weight decay shrinks the parameter directly so the
implicit prior is W -> 0. In LoRA the trained parameters are A and B
while the effective weight is W = W_init + (alpha/r) * B @ A; decaying
A and B separately drives BA -> 0, hence W -> W_init rather than 0.
The previous default of 0.01 inherited from full-FT recipes adds a
measurable pull on the merged adapter back toward the base model over
a few thousand steps. 0.001 keeps a small Frobenius-norm prior on
||A||^2 + ||B||^2 for numerical stability without meaningfully biasing
the merged weight toward init, and aligns with the value used across
the unsloth notebook templates.
* Studio: strip orphan tool_call XML leaking into visible content (#5735)
* Studio: strip orphan tool_call XML from streamed visible content
The speculative-buffer state machine in
`studio/backend/core/inference/llama_cpp.py` can slice a tool_call XML
block between the silent DRAINING path and the user-visible
content_accum, depending on when in the model's emission the BUFFERING
-> STREAMING -> DRAINING transitions fire. Three leak shapes were
observed in a 2026-05-22 sweep of 900 Qwen3.5 / Qwen3.6 GGUF runs:
Pre-fix XML leak rate: 20/900 (2.22%), concentrated 6.7% on the
larger Q8 / MTP configs:
Qwen3.6-35B-A3B Q8_0 4/60 (6.7%)
Qwen3.6-35B-A3B-MTP Q4 4/60 (6.7%)
Qwen3.5-35B-A3B Q8_0 3/60 (5.0%)
Qwen3.6-27B Q8_0 3/60 (5.0%)
The existing `_TOOL_XML_RE` only matched well-formed
`<tool_call>...</tool_call>` and `<function=...></function>` pairs, so
unterminated openings (close was DRAINED) and orphan closes (opening
was DRAINED) survived the strip and reached the user.
Fix relaxes the regex to also strip:
1. Orphan opening up to end-of-string: `(?:</tool_call>|\Z)`
2. Orphan closing tag: bare `</tool_call>` / `</function>`
Verified on the full sweep: 20/900 -> 0/900 (100% of detected leaks
eliminated). 16 unit tests in `test_tool_xml_strip.py` pin all three
leak shapes plus the well-formed cases, plus parametrised checks on
the 5 actual real-world leak samples from the sweep data.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Studio: strip tail-only </parameter> orphan + tighten regex
The 2026-05-22 gdpval sweep surfaced a 4th XML-leak shape not caught
by the earlier regex: a bare `</parameter>\n\n` at end-of-buffer (7
of 192 trials, all Qwen3.5-27B + a few Qwen3.6-27B). The model emits
the full `<tool_call><function=...><parameter=...>...content...
</parameter></function></tool_call>` envelope, the speculative buffer
DRAINS the opening tags as intended, but EOS (max_tokens cutoff)
truncates the outer `</function></tool_call>` close, leaving just
`</parameter>` as the visible tail.
We strip this ONLY when end-anchored (`\s*\Z`) so legitimate
mid-text uses (user code samples, documentation discussing the
Qwen tool-call XML shape) survive. Verified on the 192-trial
gdpval corpus: before=7, after=0.
While at it, fold the five top-level alternations into three by
sharing tag-name and prefix subgroups:
<tool_call>... + <function=\w+>... + --> <(?:tool_call|function=\w+)>...
</tool_call> | </function> --> </(?:tool_call|function)>
Semantically identical (verified by replay over the 192-trial
corpus + adversarial inputs, 0 diffs) and 1.34x faster on real
workloads. Backtracking-safety pinned by two new perf guards
(256KB '<' spam, 1000x orphan opens).
Tests: 16 -> 28 (6 new functional + 4 well-formed-vs-orphan +
2 perf guards).
* Tighten comments in XML-strip regex and tests
Code says what it does; comments were repeating it. Strip the verbose
explanations down to the WHY-only bits (engine quirk, tail-anchor
rationale, real-world source of each test sample). No code changes.
inference.py: 21 -> 12 lines around _TOOL_XML_RE
test_tool_xml_strip.py: 343 -> 259 lines (-84)
Tests: 28/28 still pass.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* Address review: deny pass-through --parallel, preserve legacy short aliases, fix test harness
Round 1 review fixes for #5737:
1. Deny --parallel / --n-parallel / -np in the pass-through validator.
Without this, `unsloth studio run --model X --parallel 8 -- --parallel
999` would last-win-override the running llama-server slot count while
Studio's app.state.llama_parallel_slots and KV-cache fitting stay at
the typer value (8), so the resource plan and the running process
disagree. Also bypasses the typer 1..64 range guard. Reject so the
only path is the first-class typer flag.
2. Backwards-compat shim for -m / -hfr / -f. Dropping the short aliases
from typer broke any script using `unsloth studio run -m X` or
`-hfr Y` or `-f dist`. Add _consume_legacy_short_aliases which pops
EXACT whole-token matches (or `-x=value` inline form) from ctx.args
into the corresponding typer parameter. Clustered tokens (`-fa`,
`-mg`, `-fitt`, ...) are left in the pass-through tail unchanged.
--model becomes Optional with an explicit missing-required check
after the preprocessor so legacy `-m X` still satisfies the
"must specify a model" requirement.
3. Drop mix_stderr from CliRunner. Typer 0.25.1 / Click 8.4.1 removed
the kwarg; the test harness raised TypeError before exercising the
PR behaviour. Tests run cleanly on current and older Typer/Click.
4. Correct the -np regression test docstring. Pre-PR `-np 8` was
clustered by Click as `-p 8` (port=8) + stray `-n`, silently
breaking the port binding -- not "passed through as 8 slots". The
post-PR assertion (child gets --parallel 8) is unchanged.
5. Update studio run docstring listing rejected flags so it now
correctly includes --parallel / -np / --n-parallel.
New tests:
- test_llama_server_args.py: parametrized denylist coverage for
--parallel / --n-parallel / -np including equals-form, including
out-of-range bypass attempts (999, 0). is_managed_flag flips True.
- test_studio_run_short_alias_clashes.py: legacy -m / -hfr / -f
promote to typer params; --model X + -m Y conflict errors; clustered
-mg / -fa / -fitt still pass through (the original bug fix holds).
132 tests pass (98 backend + 34 cli).
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Extend legacy-alias shim tests for repo:variant, inline value form, and missing model
Three additional edge cases for the -m / -hfr / -f preprocessor:
- `-m unsloth/foo:UD-Q4_K_XL` round-trips through both the preprocessor
and _split_repo_variant so the child sees --model + --gguf-variant.
- `-m=foo` inline value form is promoted just like `-m foo`.
- Missing --model after the preprocessor raises typer.Exit(2) cleanly
(replacing typer's pre-PR required-flag enforcement now that --model
is Optional to allow the legacy promotion path).
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Scrub .github/workflows for staging push (matches staging base)
* Fix studio CLI argv handling and pass-through docstring drift
- studio/backend/core/inference/llama_server_args.py: drop the stale
``-np``/``--parallel`` entry from the docstring's pass-through tunable
list. These flags moved into _DENYLIST_GROUPS so the docstring now
contradicts the validator and would mislead future maintainers
debugging the ValueError from validate_extra_args(["--parallel","8"]).
The deleted wording was introduced by dbea77e34 ("Studio: forward
llama-server args from `unsloth studio run`, activate `unsloth run`,
and allow passing model:quant to load models") when --parallel was
still a documented pass-through; the same commit's "quant" reference
is about the model:quant syntax, unrelated to the parallel slot
wording being deleted here.
- unsloth_cli/commands/studio.py: add _expand_attached_np_short next to
_consume_legacy_short_aliases. Both work around Click's short-option
clustering for this command -- the legacy preprocessor for `-m` / `-f`
/ `-hfr` and this one for the attached `-np<N>` form. Click clusters
`-np8` as `-n -p 8` because `-p` is the typer short for `--port`,
silently setting port=8 and dropping the parallel value; rewriting the
attached form into separated `-np <N>` in sys.argv before Click
parses preserves the user's value. Space/equals forms (`-np 8`,
`-np=8`) already work and are left alone.
- unsloth_cli/__init__.py: import _expand_attached_np_short from the
studio command and run it only when argv[0] looks like the unsloth
console-script or workspace cli.py, so importing this module from a
notebook or pytest run does not mutate the caller's argv.
* Tighten the -np canonicaliser comments
Drop the helper's co-location sentence (location is self-evident from
grep) and shorten the entry-gate rationale to one short sentence
covering the why.
* Sync .github/workflows with upstream author branch
* Sync .github/workflows with upstream author branch
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Bump install.sh / install.ps1 pin to unsloth>=2026.5.7 (#5753)
PyPI release unsloth 2026.5.7 is now live. Bumps the pinned floor in
install.sh and install.ps1 from unsloth>=2026.5.6 to unsloth>=2026.5.7
so fresh installs resolve to the new wheel.
Tagged on main as v0.1.416-beta.
* Catch attached `-np<N>` form in backend pass-through validator
The CLI-side `_expand_attached_np_short` rewrites `-np8` to `-np 8`
before Click parses, but HTTP /load `llama_extra_args=["-np8"]` goes
straight to `validate_extra_args` which only matched the exact token.
Reproducer: `validate_extra_args(["-np8"])` previously returned
`["-np8"]` instead of raising; once forwarded to llama-server it
last-win-overrode Studio's slot count while
`app.state.llama_parallel_slots` stayed at the typer value.
Normalise `-np<digits>` to `-np` in `_flag_name` so the denylist
catches the attached form alongside `-np`, `-np=8`, `--parallel`,
`--parallel=8`, and `--n-parallel`. Tests parametrize the new form
including out-of-range values.
* Restore _consume_legacy_short_aliases unit tests + _expand_attached_np_short tests
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Restore .github/workflows from origin/main
Earlier merge from claude_review's staging-scrub commits accidentally
deleted production CI workflows. Restore them to main's state.
* Scrub .github/workflows for staging push (matches staging base)
* Sync .github/workflows with upstream author branch
* Round 5+6: broaden -np gate to exact basenames + runtime parallel test
Reviewer-flagged improvements squashed into one commit so the auto-push
review bot doesn't keep stomping the branch:
- unsloth_cli/__init__.py: exact-basename match instead of
endswith('cli.py'). Covers unsloth, unsloth.exe, unsloth-cli,
unsloth-cli.exe, cli.py, unsloth-cli.py. A third-party mycli.py that
happens to import unsloth_cli no longer has its argv mutated.
- unsloth_cli/tests/test_studio_run_parallel_flag.py: parametrised
runtime test (N in {1, 4, 8, 64}) that fakes the in-venv path and
asserts run_server is invoked with llama_parallel_slots=N.
Complements the existing source-text check so refactors that preserve
runtime semantics don't trip a false failure.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Round 7: respect '--' end-of-options and reject flag-as-value
Round 7 reviewer flagged three legitimate edge cases:
- _expand_attached_np_short rewrote post-'--' tokens. Convention: '--'
ends option processing; payload after it is raw. Stop the loop there.
- _consume_legacy_short_aliases promoted post-'--' legacy aliases for
the same reason. Treat post-'--' tail as raw.
- Legacy '-m -fa' silently consumed '-fa' as the model name, hiding
the real CLI shape error. Reject any next-token that starts with '-'
(except the lone '-' stdin/path sentinel) with a clear BadParameter.
Also expanded the missing-model error string to mention the still-
supported legacy '-m' / '-hfr' aliases so users hitting that diagnostic
on legacy scripts get the right migration hint.
Added four regression tests covering each new behaviour.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Round 8: soften flag-as-value to long-form only + normalise is_managed_flag
Round 8 reviewer flagged two cleanups:
- _consume_legacy_short_aliases rejected any next token starting with
'-' as a flag, which would break legitimate values like '-foo'
(path or model name with leading dash). Narrow the rejection to
'--long' tokens only; '-x' short forms still pass through.
- is_managed_flag did raw _DENYLIST membership while validate_extra_args
goes through _flag_name first, so '-np8' / '--parallel=8' /
'--port=9000' classified as not-managed by the helper but rejected
by the validator. Route is_managed_flag through _flag_name so the
two helpers agree on every form callers might use.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Round 9: also catch -np-1 / -np+1 signed attached forms in denylist
Round 9 reviewer noticed _flag_name normalised -np<digits> but missed
signed variants -np-1 and -np+1, so validate_extra_args waved them
through while rejecting --parallel -1. llama.cpp would error out on
negative slot counts anyway, but the validator should classify every
form of the managed flag identically so the boundary is consistent.
* Round 10: signed -np in CLI canonicaliser + reject empty inline aliases
Round 10 reviewer flagged two real issues:
- _expand_attached_np_short rewrote only -np<digits>; signed forms
-np-1 / -np+1 fell through. Backend _flag_name already classifies
them as managed, so the CLI rewriter must too -- otherwise Click
clusters -np-1 into -n -p -1 (port=-1) and never reaches the
backend validator at all.
- -m= / -hfr= / -f= empty inline forms were accepted and produced
--model '' / --frontend '' (then Path('') silently became '.') on
re-exec. Reject empty inline values at the preprocessor with a
clear BadParameter so the malformed input fails fast.
Both behaviours pinned with parametrised regression tests.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Expose --parallel on plain `unsloth studio` for API-path parity
The PR added --parallel to `unsloth studio run` but the plain
`unsloth studio` callback (used for API-only / bare-server launches)
still hardcoded llama_parallel_slots to its run_server default. With
--parallel now denied as a llama_extra_args pass-through, that flow
had no first-class way to raise concurrency.
- unsloth_cli/commands/studio.py: add --parallel / --n-parallel typer
Option (default 4, range 1..64) to studio_default, forward through
the venv re-exec, and pass llama_parallel_slots= to run_server in
the in-venv path.
- studio/backend/run.py: argparse --parallel / --n-parallel with the
same range guard so the spawned child accepts the forwarded flag.
- unsloth_cli/tests/test_studio_run_parallel_flag.py: test pins the
new option presence, aliases, default and range guards.
* Round 12: narrow entry-point gate, preserve pre-PR plain-studio default, drop brittle source-text test
Three Opus subagent reviewers (security / backcompat / code-quality)
flagged the same handful of real issues. Consensus fixes:
- unsloth_cli/__init__.py: narrow the -np canonicaliser gate to just
{unsloth, unsloth.exe} (the only pyproject-declared console_script).
The previous cli.py / unsloth-cli.py entries would silently rewrite
sys.argv for any third-party myproj/cli.py that happens to import
unsloth_cli. Dev users running python cli.py ... -np N still work
via the space form, which parses without the rewrite.
- unsloth_cli/commands/studio.py + studio/backend/run.py: restore the
pre-PR llama_parallel_slots default of 1 on plain unsloth studio and
python studio/backend/run.py. unsloth studio run keeps its
hardcoded-pre-PR default of 4. Without this, my earlier API-path
parity commit silently dropped per-call context to ctx/4 for the
plain-studio flow.
- unsloth_cli/tests/test_studio_run_parallel_flag.py: drop the brittle
source-text grep test (test_run_kwargs_use_parallel_value). The
parametrised runtime test test_in_venv_path_passes_parallel_to_run_server
already pins the same intent against actual behaviour.
- unsloth_cli/tests/test_studio_run_short_alias_clashes.py: pin the
narrow entry-point gate with a parametrised negative test covering
seven third-party argv[0] basenames (cli.py, /path/myproj/cli.py,
pytest, unsloth-cli, etc.). Re-broadening the gate now trips a
test instead of silently mutating an unrelated CLI's argv.
* Round 13: shared parallel constants, denylist invariant test, defence-in-depth
Three Opus subagent reviewers (adversarial-user / maintenance /
cross-file consistency) flagged a consistent set of cleanups; folded
into one commit to avoid the pre-commit.ci force-push race.
unsloth_cli/commands/studio.py:
- Extract _PARALLEL_MIN / _PARALLEL_MAX / _PARALLEL_DEFAULT_RUN /
_PARALLEL_DEFAULT_PLAIN module-level constants and use them in both
typer Options (plain studio_default = 1, studio run = 4).
- _expand_attached_np_short now rewrites -np<junk> when the suffix
starts with a digit (or signed digit) so '-np8x' surfaces as a
clean '-np takes an int' typer error instead of a baffling
'--port invalid' complaint after Click clusters '-n -p 8x'.
- Re-exec forwarding emits --load-in-4bit / --no-load-in-4bit
explicitly in both directions; previously the True default relied
on both layers sharing the same default forever.
- run() docstring now explicitly says --parallel / -np pass-through
via llama_extra_args is denied (use the typer flag above).
studio/backend/run.py:
- Mirror the parallel constants and route the argparse default,
range check, and error message through them. Help text mentions
the asymmetry with 'unsloth studio run' so direct-launch dev users
aren't confused by Default 1 in isolation.
studio/backend/core/inference/llama_server_args.py:
- _flag_name strips surrounding whitespace before denylist lookup so
a caller can't slip a managed flag past the boundary with a
trailing space (the trimmed form is what downstream parsers see).
Tests:
- New typer-aliases-subset-of-denylist invariant: every alias the
typer Option claims as --parallel on run() MUST be in the backend
parallel denylist group. Catches the failure mode where someone
adds a new alias and forgets the boundary.
- Extended denylist parametrize to cover ~14 previously untested
aliases (-mu, -dr, -hfv/-hfrv/-hffv family, -mmu, full --ui group,
--models-preset / --models-autoload / --no-models-autoload).
- Whitespace-padded denylist rejection (' --parallel', '-np ', etc).
- --load-in-4bit re-exec test pinning both polarities + default.
- -np<junk> argv rewriter regression tests.
- Cross-reference headers between the two test files.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix: repair mlx studio base export save_method (#5727)
* Round 14: align backend -np recogniser with CLI rewriter + reject parent --parallel
Round 14 (reviewer.py --parallel 20 with gpt-5.3-codex-spark) flagged
two real P1s and a stale-rebase warning. All three addressed.
- studio/backend/core/inference/llama_server_args.py: widen
_flag_name so -np<digit-prefix> with trailing junk (-np8x,
-np-1foo, -np+1bar, -np9zzz) classifies as managed flag -np,
matching the CLI _expand_attached_np_short rewriter. Without this,
POST /api/inference/load with llama_extra_args=['-np8x'] slipped
past the boundary while the CLI canonicalised the same form. The
two sides now agree on every digit-prefix form.
- unsloth_cli/commands/studio.py: reject --parallel on the
studio group when a subcommand is invoked. Pre-PR the studio
callback had no --parallel; my Round 12 addition made
'unsloth studio --parallel 8 run ...' silently drop the 8
because typer doesn't propagate parent options into subcommand
kwargs. Now errors with exit 2 and a message pointing the
operator at the correct invocation
('unsloth studio run --parallel 8 ...').
- Picked up origin/main via merge (parent commit 0caf0526): the
pre-flight stale-rebase detector found 2 lines on main in
studio/backend/core/export/export.py missing from PR HEAD.
Merged cleanly with no conflicts.
Tests:
- Parametrised denylist coverage for -np<digit-prefix>+junk forms.
- New runtime test confirms exit 2 + helpful error when the group
--parallel is supplied alongside an invoked subcommand.
- Test that the default group --parallel value still lets a
subcommand resolve (no false-positive rejection).
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Studio: tighten code comments across --parallel PR
Comment-only pass over the seven PR-touched files; trim verbose
docstrings, collapse multi-line section dividers, and drop
redundant prose that the code already conveys. No behaviour change.
* Studio: trim remaining verbose docstrings missed in last pass
Shorten the test_studio_run_parallel_flag.py module docstring and
the `Re-exec arg-builder coverage` block. No behaviour change.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Studio: second comment-tightening pass across PR-touched code
Trim docstrings and inline comments in studio.py, run.py,
llama_server_args.py, and unsloth_cli/__init__.py. No behaviour change;
all 215 tests still pass.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Studio: deny --embedding / --rerank / --tools pass-through
`--embedding` and `--rerank` flip llama-server into single-endpoint
mode, which breaks Studio's /v1/chat/completions hop. llama-server's
own `--tools` flag silently stacks on top of Studio's tool policy
resolved by `--enable-tools` / `--disable-tools`.
Add all three (plus the `--embeddings` / `--reranking` plural aliases)
to the boundary denylist so HTTP /load and pass-through extras both
reject them cleanly instead of silently desyncing the server surface.
Test added to the existing `test_denylist_rejects_all_aliases`
parametrize. 220 tests pass.
* Studio: make PR-touched tests robust to minimal envs + Windows
Two cross-OS CI findings:
1. `test_typer_parallel_aliases_are_subset_of_backend_denylist` was
doing `from core.inference.llama_server_args import _DENYLIST_GROUPS`
which triggers `core/inference/__init__.py` and pulls in the full
backend chain (fastapi / structlog / loggers / utils.hardware).
The invariant only needs the constants tuple, so load the module
directly via `importlib.util.spec_from_file_location` -- the test
now runs with just typer + pytest installed.
2. `test_legacy_frontend_alias_still_promotes_to_frontend` asserted
the literal string `"/tmp/dist"` after the value round-trips through
`Path()`. On Windows `str(Path("/tmp/dist"))` is `"\tmp\dist"`, so
the assertion tripped on the same logical path. Compare via
`Path(x) == Path("/tmp/dist")` so the test passes on every OS.
Both surfaced by the staging-4 cross-OS CI; no production-code change.
220 tests still pass locally.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Studio: load llama_server_args.py directly in its unit tests
Same fix as the previous CLI-test commit: import the module via
`importlib.util.spec_from_file_location` instead of
`from core.inference.llama_server_args import ...`, so the test no
longer needs the full backend chain (fastapi / structlog / loggers /
utils.hardware) installed via `core/inference/__init__.py`.
The boundary validator is intentionally dependency-free; its unit
tests should reflect that.
* Fix test_main_composer_has_dir_auto anchor after PR #5784
PR #5784 ("Improve image generation UI") rewrote the message-input
textarea's static `aria-label="Message input"` into a JSX conditional
`aria-label={overlay ? "Image edit instructions" : "Message input"}`
but did not update the RTL bidi-attribute regression test, leaving
the literal-string `find('aria-label="Message input"')` anchor with
no match. The `Repo tests (CPU)` job has been red on main since.
Anchor on the inner `"Message input"` string literal instead -- it
survives both spellings and still pins the same textarea element so
the `dir="auto"` assertion has the right block to inspect.
Verified by re-running the exact CI command:
954 passed, 3 skipped, 23 deselected (was 948 passed, 1 failed).
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Long Yixing <longyixing331@gmail.com>
912 lines
33 KiB
Python
912 lines
33 KiB
Python
# SPDX-License-Identifier: AGPL-3.0-only
|
|
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
|
|
|
"""
|
|
Run script for Unsloth UI Backend.
|
|
Works independently and can be moved to any directory.
|
|
"""
|
|
|
|
import os
|
|
import sys
|
|
from pathlib import Path
|
|
from typing import Optional
|
|
|
|
# Suppress annoying C-level dependency warnings globally (e.g. SwigPyPacked)
|
|
os.environ["PYTHONWARNINGS"] = "ignore"
|
|
|
|
# Add the backend directory to Python path early so local modules are importable
|
|
backend_dir = Path(__file__).parent
|
|
if str(backend_dir) not in sys.path:
|
|
sys.path.insert(0, str(backend_dir))
|
|
|
|
# Fix for Anaconda/conda-forge Python: seed platform._sys_version_cache before
|
|
# any library imports that trigger attrs -> rich -> structlog -> platform crash.
|
|
# See: https://github.com/python/cpython/issues/102396
|
|
import _platform_compat # noqa: F401
|
|
|
|
from loggers import get_logger
|
|
from startup_banner import print_studio_access_banner, print_studio_stop_hint
|
|
|
|
logger = get_logger(__name__)
|
|
|
|
|
|
def _resolve_external_ip() -> str:
|
|
"""
|
|
Resolve the machine's external IP address.
|
|
|
|
Tries (in order):
|
|
1. GCE metadata server (instant, works on Google Cloud VMs)
|
|
2. ifconfig.me (works anywhere with internet)
|
|
3. LAN IP via UDP socket trick (fallback)
|
|
"""
|
|
import urllib.request
|
|
import socket
|
|
|
|
# 1. Try GCE metadata server (responds in <10ms on GCE, times out fast elsewhere)
|
|
try:
|
|
req = urllib.request.Request(
|
|
"http://metadata.google.internal/computeMetadata/v1/instance/network-interfaces/0/access-configs/0/external-ip",
|
|
headers = {"Metadata-Flavor": "Google"},
|
|
)
|
|
with urllib.request.urlopen(req, timeout = 1) as resp:
|
|
ip = resp.read().decode().strip()
|
|
if ip:
|
|
return ip
|
|
except Exception:
|
|
pass
|
|
|
|
# 2. Try public IP service
|
|
try:
|
|
with urllib.request.urlopen("https://ifconfig.me", timeout = 3) as resp:
|
|
ip = resp.read().decode().strip()
|
|
if ip:
|
|
return ip
|
|
except Exception:
|
|
pass
|
|
|
|
# 3. Fallback: LAN IP via UDP socket trick
|
|
try:
|
|
s = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
|
|
s.connect(("8.8.8.8", 80))
|
|
ip = s.getsockname()[0]
|
|
s.close()
|
|
return ip
|
|
except Exception:
|
|
return "0.0.0.0"
|
|
|
|
|
|
def _install_uvicorn_startup_log_rewrite(bind_host: str, display_host: str) -> None:
|
|
"""Rewrite Uvicorn's startup log line: swap wildcard bind for the
|
|
externally-reachable address, replace the CTRL+C suffix with our Mac-aware
|
|
stop hint, and rename the prefix to "Unsloth Studio running on"."""
|
|
import logging
|
|
import re
|
|
|
|
rewrite_host = (
|
|
bind_host in ("0.0.0.0", "::")
|
|
and bool(display_host)
|
|
and display_host != bind_host
|
|
)
|
|
new_suffix = "(To stop: press Ctrl+C -- on macOS, Control+C not Command+C)"
|
|
old_suffix_re = re.compile(r"\(Press CTRL\+C to quit\)")
|
|
old_prefix = "Uvicorn running on "
|
|
new_prefix = "Unsloth Studio running on "
|
|
|
|
def _rewrite(text: str) -> str:
|
|
if text.startswith(old_prefix):
|
|
text = new_prefix + text[len(old_prefix) :]
|
|
return old_suffix_re.sub(new_suffix, text)
|
|
|
|
class _UvicornStartupRewrite(logging.Filter):
|
|
def filter(self, record: logging.LogRecord) -> bool:
|
|
try:
|
|
msg = record.msg if isinstance(record.msg, str) else ""
|
|
if (
|
|
msg.startswith(old_prefix)
|
|
and isinstance(record.args, tuple)
|
|
and len(record.args) >= 3
|
|
):
|
|
if rewrite_host and record.args[1] == bind_host:
|
|
record.args = (
|
|
record.args[0],
|
|
display_host,
|
|
record.args[2],
|
|
*record.args[3:],
|
|
)
|
|
record.msg = _rewrite(msg)
|
|
cmsg = getattr(record, "color_message", None)
|
|
if isinstance(cmsg, str):
|
|
record.color_message = _rewrite(cmsg)
|
|
except Exception:
|
|
pass
|
|
return True
|
|
|
|
f = _UvicornStartupRewrite()
|
|
for name in ("uvicorn", "uvicorn.error"):
|
|
logging.getLogger(name).addFilter(f)
|
|
|
|
|
|
def _local_port_open(host: str, port: int, timeout: float = 1.0) -> bool:
|
|
"""Return True iff a TCP connection to (host, port) succeeds within timeout."""
|
|
import socket
|
|
|
|
try:
|
|
with socket.create_connection((host, port), timeout = timeout):
|
|
return True
|
|
except OSError:
|
|
return False
|
|
|
|
|
|
def _working_local_url(port: int) -> "str | None":
|
|
"""Return a working loopback URL on this machine, or None if neither
|
|
127.0.0.1 nor ::1 responds. Used as a fallback when external reachability fails."""
|
|
if _local_port_open("127.0.0.1", port):
|
|
return f"http://127.0.0.1:{port}"
|
|
if _local_port_open("::1", port):
|
|
return f"http://[::1]:{port}"
|
|
return None
|
|
|
|
|
|
def _stdout_color_ok() -> bool:
|
|
"""Whether to emit ANSI color codes on stdout. Mirrors startup_banner."""
|
|
if os.environ.get("NO_COLOR", "").strip():
|
|
return False
|
|
if os.environ.get("FORCE_COLOR", "").strip():
|
|
return True
|
|
try:
|
|
return sys.stdout.isatty()
|
|
except (AttributeError, OSError, ValueError):
|
|
return False
|
|
|
|
|
|
def _verify_global_reachability(display_host: str, port: int) -> None:
|
|
"""Probe check-host.net to confirm display_host:port is reachable from the
|
|
public internet. Synchronous so the caller can render output between the
|
|
banner URL section and the trailing stop hint. Bounded at ~15s; failures
|
|
are swallowed (the verifier failing is not Studio failing). Only meaningful
|
|
when bound to a wildcard host."""
|
|
import ipaddress
|
|
import json
|
|
import time
|
|
import urllib.error
|
|
import urllib.parse
|
|
import urllib.request
|
|
|
|
if not display_host or display_host in ("0.0.0.0", "::"):
|
|
return
|
|
|
|
use_color = _stdout_color_ok()
|
|
dim = "\033[38;5;245m" if use_color else ""
|
|
ok_c = "\033[38;5;120;1m" if use_color else ""
|
|
err_c = "\033[38;5;203;1m" if use_color else ""
|
|
warn_c = "\033[38;5;215;1m" if use_color else ""
|
|
local_url_c = "\033[38;5;108;1m" if use_color else "" # matches banner's URL color
|
|
reset = "\033[0m" if use_color else ""
|
|
|
|
url = f"http://{display_host}:{port}"
|
|
|
|
# Private / loopback / link-local addresses are not globally routable.
|
|
try:
|
|
addr = ipaddress.ip_address(display_host)
|
|
if addr.is_loopback or addr.is_private or addr.is_link_local:
|
|
print(
|
|
f"{dim} Note: {display_host} is a private/LAN address -- "
|
|
f"reachable on this network only, not from the public internet."
|
|
f"{reset}",
|
|
flush = True,
|
|
)
|
|
return
|
|
except ValueError:
|
|
# Not an IP literal; probe by hostname.
|
|
pass
|
|
|
|
try:
|
|
qs = urllib.parse.urlencode({"host": f"{display_host}:{port}", "max_nodes": 3})
|
|
req = urllib.request.Request(
|
|
f"https://check-host.net/check-tcp?{qs}",
|
|
headers = {
|
|
"Accept": "application/json",
|
|
"User-Agent": "unsloth-studio-reachability/1",
|
|
},
|
|
)
|
|
with urllib.request.urlopen(req, timeout = 5) as resp:
|
|
init = json.loads(resp.read().decode("utf-8", errors = "replace"))
|
|
req_id = init.get("request_id")
|
|
if not req_id:
|
|
return
|
|
|
|
results = {}
|
|
deadline = time.monotonic() + 15.0
|
|
poll_req = urllib.request.Request(
|
|
f"https://check-host.net/check-result/{req_id}",
|
|
headers = {
|
|
"Accept": "application/json",
|
|
"User-Agent": "unsloth-studio-reachability/1",
|
|
},
|
|
)
|
|
while time.monotonic() < deadline:
|
|
time.sleep(1.5)
|
|
try:
|
|
with urllib.request.urlopen(poll_req, timeout = 5) as resp:
|
|
results = json.loads(resp.read().decode("utf-8", errors = "replace"))
|
|
except Exception:
|
|
continue
|
|
if results and all(v is not None for v in results.values()):
|
|
break
|
|
# Two decisive nodes is enough; stop polling early.
|
|
decisive = [
|
|
v
|
|
for v in results.values()
|
|
if isinstance(v, list)
|
|
and v
|
|
and isinstance(v[0], dict)
|
|
and ("time" in v[0] or "error" in v[0])
|
|
]
|
|
if len(decisive) >= 2:
|
|
break
|
|
|
|
ok_nodes = err_nodes = 0
|
|
for v in results.values():
|
|
if not isinstance(v, list) or not v or not isinstance(v[0], dict):
|
|
continue
|
|
if "time" in v[0]:
|
|
ok_nodes += 1
|
|
elif "error" in v[0]:
|
|
err_nodes += 1
|
|
total = ok_nodes + err_nodes
|
|
|
|
print("", flush = True)
|
|
if ok_nodes:
|
|
print(
|
|
f"{ok_c} Reachability check: {url}/ is reachable from the "
|
|
f"public internet ({ok_nodes}/{total} probe nodes connected).{reset}",
|
|
flush = True,
|
|
)
|
|
elif err_nodes:
|
|
print(
|
|
f"{err_c} Reachability check: {url}/ is NOT reachable from "
|
|
f"the public internet ({err_nodes}/{total} probe nodes failed).{reset}",
|
|
flush = True,
|
|
)
|
|
print(f"{dim} Common causes:{reset}", flush = True)
|
|
print(
|
|
f"{dim} * AWS -- the instance's Security Group doesn't "
|
|
f"allow inbound TCP {port}.{reset}",
|
|
flush = True,
|
|
)
|
|
print(
|
|
f"{dim} * GCP -- no firewall rule allowing TCP {port} "
|
|
f"for the instance's network tag.{reset}",
|
|
flush = True,
|
|
)
|
|
print(
|
|
f"{dim} * Azure / other clouds -- equivalent NSG / "
|
|
f"firewall rule missing.{reset}",
|
|
flush = True,
|
|
)
|
|
print(
|
|
f"{dim} * Home -- your router isn't port-forwarding "
|
|
f"{port} to this machine.{reset}",
|
|
flush = True,
|
|
)
|
|
print(
|
|
f"{dim} Workaround that needs no firewall changes -- "
|
|
f"SSH local-forward from your laptop:{reset}",
|
|
flush = True,
|
|
)
|
|
print(
|
|
f"{dim} ssh -L {port}:localhost:{port} "
|
|
f"<user>@{display_host}{reset}",
|
|
flush = True,
|
|
)
|
|
print(
|
|
f"{dim} then open http://localhost:{port}/ in your browser.{reset}",
|
|
flush = True,
|
|
)
|
|
# Only offer the local URL if loopback actually answers.
|
|
local_url = _working_local_url(port)
|
|
if local_url:
|
|
print(
|
|
f"{local_url_c} You can access Unsloth Studio locally "
|
|
f"in the meantime: {local_url}{reset}",
|
|
flush = True,
|
|
)
|
|
else:
|
|
print(
|
|
f"{warn_c} Reachability check: probe nodes did not respond "
|
|
f"in time -- could not verify {url}/.{reset}",
|
|
flush = True,
|
|
)
|
|
except urllib.error.URLError:
|
|
# Outbound HTTPS blocked; skip silently.
|
|
pass
|
|
except Exception:
|
|
pass
|
|
|
|
|
|
def _get_pid_on_port(port: int) -> "tuple[int, str] | None":
|
|
"""Return (pid, process_name) of the process listening on *port*, or None.
|
|
|
|
Uses psutil when available. Falls back gracefully to None so callers
|
|
can still report the port conflict without process details.
|
|
|
|
Works on Windows, macOS, and Linux wherever psutil is installed.
|
|
"""
|
|
try:
|
|
import psutil
|
|
except ImportError:
|
|
return None
|
|
try:
|
|
for conn in psutil.net_connections(kind = "tcp"):
|
|
if conn.status == "LISTEN" and conn.laddr.port == port:
|
|
if conn.pid is None:
|
|
return None
|
|
try:
|
|
proc = psutil.Process(conn.pid)
|
|
return (conn.pid, proc.name())
|
|
except (psutil.NoSuchProcess, psutil.AccessDenied):
|
|
return (conn.pid, "<unknown>")
|
|
except (psutil.AccessDenied, OSError) as e:
|
|
# psutil.net_connections() needs elevated privileges on some platforms
|
|
logger.debug("Failed to scan network connections for port %s: %s", port, e)
|
|
return None
|
|
|
|
|
|
def _is_port_free(host: str, port: int) -> bool:
|
|
"""Check if a port is available for binding.
|
|
|
|
When *host* is ``0.0.0.0`` (wildcard), we also check whether anything
|
|
is already listening on ``127.0.0.1`` (and ``::1`` when IPv6 is
|
|
available). An SSH tunnel or similar process may hold the loopback
|
|
address while our wildcard bind still succeeds, making Unsloth Studio
|
|
unreachable via ``localhost``.
|
|
|
|
Works on Windows, macOS, and Linux.
|
|
"""
|
|
import socket
|
|
|
|
# 1. Can we bind to the requested address?
|
|
# Use getaddrinfo so both IPv4 ("0.0.0.0") and IPv6 ("::") hosts
|
|
# resolve to the correct address family automatically.
|
|
try:
|
|
addr_info = socket.getaddrinfo(host, port, socket.AF_UNSPEC, socket.SOCK_STREAM)
|
|
family, socktype, proto, _, sockaddr = addr_info[0]
|
|
with socket.socket(family, socktype, proto) as s:
|
|
s.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
|
|
s.bind(sockaddr)
|
|
except OSError:
|
|
return False
|
|
|
|
# 2. When binding to all interfaces, verify that localhost is not
|
|
# already claimed by another process (e.g. an SSH -L tunnel).
|
|
# We attempt a TCP connect -- if it succeeds something is listening.
|
|
if host in ("0.0.0.0", "::"):
|
|
for loopback, family in [
|
|
("127.0.0.1", socket.AF_INET),
|
|
("::1", socket.AF_INET6),
|
|
]:
|
|
try:
|
|
with socket.socket(family, socket.SOCK_STREAM) as s:
|
|
s.settimeout(1)
|
|
if s.connect_ex((loopback, port)) == 0:
|
|
# Connection succeeded -- port is taken on loopback
|
|
return False
|
|
except OSError:
|
|
# IPv6 disabled or other OS-level restriction -- skip
|
|
continue
|
|
|
|
return True
|
|
|
|
|
|
def _find_free_port(host: str, start: int, max_attempts: int = 20) -> int:
|
|
"""Find a free port starting from `start`, trying up to max_attempts ports."""
|
|
for offset in range(max_attempts):
|
|
candidate = start + offset
|
|
if _is_port_free(host, candidate):
|
|
return candidate
|
|
raise RuntimeError(
|
|
f"Could not find a free port in range {start}-{start + max_attempts - 1}"
|
|
)
|
|
|
|
|
|
from utils.paths.storage_roots import studio_root as _studio_root
|
|
|
|
_PID_FILE = _studio_root() / "studio.pid"
|
|
|
|
# Direct backend launches bypass the CLI's env re-export; do it here for
|
|
# real custom roots so unsloth-zoo's import-time LLAMA_CPP_DEFAULT_DIR
|
|
# picks up the custom build. Skip for legacy-default to avoid flipping
|
|
# default-mode installs into env-override.
|
|
try:
|
|
_LEGACY_STUDIO_ROOT = (Path.home() / ".unsloth" / "studio").resolve()
|
|
except (OSError, ValueError):
|
|
_LEGACY_STUDIO_ROOT = Path.home() / ".unsloth" / "studio"
|
|
try:
|
|
_STUDIO_ROOT_RESOLVED = _studio_root().resolve()
|
|
except (OSError, ValueError):
|
|
_STUDIO_ROOT_RESOLVED = _studio_root()
|
|
if _STUDIO_ROOT_RESOLVED != _LEGACY_STUDIO_ROOT:
|
|
if not os.environ.get("UNSLOTH_STUDIO_HOME"):
|
|
os.environ["UNSLOTH_STUDIO_HOME"] = str(_STUDIO_ROOT_RESOLVED)
|
|
if not os.environ.get("UNSLOTH_LLAMA_CPP_PATH"):
|
|
os.environ["UNSLOTH_LLAMA_CPP_PATH"] = str(_STUDIO_ROOT_RESOLVED / "llama.cpp")
|
|
|
|
|
|
def _write_pid_file():
|
|
"""Write the current process PID to the studio PID file."""
|
|
try:
|
|
_PID_FILE.parent.mkdir(parents = True, exist_ok = True)
|
|
_PID_FILE.write_text(str(os.getpid()))
|
|
except OSError:
|
|
pass
|
|
|
|
|
|
def _remove_pid_file():
|
|
"""Remove the PID file if it belongs to this process."""
|
|
try:
|
|
if _PID_FILE.is_file():
|
|
stored = _PID_FILE.read_text().strip()
|
|
if stored == str(os.getpid()):
|
|
_PID_FILE.unlink(missing_ok = True)
|
|
except OSError:
|
|
pass
|
|
|
|
|
|
def _graceful_shutdown(server = None):
|
|
"""Explicitly shut down all subprocess backends and the uvicorn server.
|
|
|
|
Called from signal handlers to ensure child processes are cleaned up
|
|
before the parent exits. This is critical on Windows where atexit
|
|
handlers are unreliable after Ctrl+C.
|
|
"""
|
|
_remove_pid_file()
|
|
logger.info("Graceful shutdown initiated — cleaning up subprocesses...")
|
|
|
|
# 1. Shut down uvicorn server (releases the listening socket)
|
|
if server is not None:
|
|
server.should_exit = True
|
|
|
|
# 2. Clean up inference subprocess (if instantiated)
|
|
try:
|
|
from core.inference.orchestrator import _inference_backend
|
|
|
|
if _inference_backend is not None:
|
|
_inference_backend._shutdown_subprocess(timeout = 5.0)
|
|
except Exception as e:
|
|
logger.warning("Error shutting down inference subprocess: %s", e)
|
|
|
|
# 3. Clean up export subprocess (if instantiated)
|
|
try:
|
|
from core.export.orchestrator import _export_backend
|
|
|
|
if _export_backend is not None:
|
|
_export_backend._shutdown_subprocess(timeout = 5.0)
|
|
except Exception as e:
|
|
logger.warning("Error shutting down export subprocess: %s", e)
|
|
|
|
# 4. Clean up training subprocess (if active)
|
|
try:
|
|
from core.training.training import _training_backend
|
|
|
|
if _training_backend is not None:
|
|
_training_backend.force_terminate()
|
|
except Exception as e:
|
|
logger.warning("Error shutting down training subprocess: %s", e)
|
|
|
|
# 5. Kill llama-server subprocess (if loaded)
|
|
try:
|
|
from routes.inference import _llama_cpp_backend
|
|
|
|
if _llama_cpp_backend is not None:
|
|
_llama_cpp_backend._kill_process()
|
|
except Exception as e:
|
|
logger.warning("Error shutting down llama-server: %s", e)
|
|
|
|
logger.info("All subprocesses cleaned up")
|
|
|
|
|
|
# The uvicorn server instance -- set by run_server(), used by callers
|
|
# that need to tell the server to exit (e.g. signal handlers).
|
|
_server = None
|
|
|
|
# Shutdown event -- used to wake the main loop on signal
|
|
_shutdown_event = None
|
|
|
|
|
|
_DEFAULT_FRONTEND_PATH = Path(__file__).resolve().parent.parent / "frontend" / "dist"
|
|
|
|
|
|
def _iter_frontend_fallback_candidates() -> "list[Path]":
|
|
"""Yield `studio/frontend/dist` paths to try when the default is missing.
|
|
|
|
Covers PATH-shadowed binaries whose __file__ resolves into a
|
|
site-packages tree that never received a vite build (e.g. plain
|
|
`pip install unsloth` from PyPI).
|
|
"""
|
|
import ast
|
|
import re
|
|
|
|
out: list[Path] = []
|
|
home_str = (
|
|
os.environ.get("UNSLOTH_STUDIO_HOME")
|
|
or os.environ.get("STUDIO_HOME")
|
|
or str(Path.home() / ".unsloth" / "studio")
|
|
)
|
|
venv_dir = Path(home_str).expanduser() / "unsloth_studio"
|
|
# Installer venv site-packages.
|
|
for pattern in (
|
|
"lib/python*/site-packages/studio/frontend/dist",
|
|
"Lib/site-packages/studio/frontend/dist",
|
|
):
|
|
out.extend(venv_dir.glob(pattern))
|
|
# Editable source roots referenced from the installer venv.
|
|
for sp_pattern in ("lib/python*/site-packages", "Lib/site-packages"):
|
|
for sp in venv_dir.glob(sp_pattern):
|
|
for finder in sp.glob("__editable___*_finder.py"):
|
|
try:
|
|
src = finder.read_text(encoding = "utf-8")
|
|
except OSError:
|
|
continue
|
|
# Tolerate single- or multi-line dict literals; [^}]* still
|
|
# rejects nested dicts, which the setuptools template never
|
|
# emits for editable installs.
|
|
m = re.search(
|
|
r"^MAPPING\s*(?::[^=]*)?=\s*(\{[^}]*\})", src, re.M | re.S
|
|
)
|
|
if not m:
|
|
continue
|
|
try:
|
|
mapping = ast.literal_eval(m.group(1))
|
|
except (SyntaxError, ValueError):
|
|
continue
|
|
# Defensive: literal_eval can return a set / list / None if the
|
|
# matched literal is not a dict (regex captures `{...}`).
|
|
if not isinstance(mapping, dict):
|
|
continue
|
|
studio_pkg = mapping.get("studio")
|
|
if studio_pkg:
|
|
out.append(Path(studio_pkg) / "frontend" / "dist")
|
|
return out
|
|
|
|
|
|
def _resolve_frontend_path(frontend_path: Path) -> tuple[Optional[Path], list[Path]]:
|
|
"""Pick a frontend dir that actually contains `index.html`.
|
|
|
|
Returns (chosen, attempted). `chosen` is None if nothing servable was
|
|
found; `attempted` is the full ordered list for diagnostics.
|
|
"""
|
|
attempted: list[Path] = []
|
|
seen: set[Path] = set()
|
|
|
|
def _try(p: Path) -> bool:
|
|
try:
|
|
key = p.resolve()
|
|
except OSError:
|
|
key = p
|
|
if key in seen:
|
|
return False
|
|
seen.add(key)
|
|
attempted.append(p)
|
|
return (p / "index.html").is_file()
|
|
|
|
if _try(Path(frontend_path)):
|
|
return attempted[-1], attempted
|
|
for alt in _iter_frontend_fallback_candidates():
|
|
if _try(alt):
|
|
return attempted[-1], attempted
|
|
return None, attempted
|
|
|
|
|
|
def run_server(
|
|
host: str = "127.0.0.1",
|
|
port: int = 8888,
|
|
frontend_path: Path = _DEFAULT_FRONTEND_PATH,
|
|
silent: bool = False,
|
|
api_only: bool = False,
|
|
llama_parallel_slots: int = 1,
|
|
):
|
|
"""
|
|
Start the FastAPI server.
|
|
|
|
Args:
|
|
host: Host to bind to
|
|
port: Port to bind to (auto-increments if in use)
|
|
frontend_path: Path to frontend build directory (optional)
|
|
silent: Suppress startup messages
|
|
api_only: Run API server only, no frontend serving (for Tauri desktop app)
|
|
llama_parallel_slots: Number of parallel slots for llama-server
|
|
|
|
Note:
|
|
Signal handlers are NOT registered here so that embedders
|
|
(e.g. Colab notebooks) keep their own interrupt semantics.
|
|
Standalone callers should register handlers after calling this.
|
|
"""
|
|
global _server, _shutdown_event
|
|
|
|
# On Windows the default console encoding (cp1252) cannot encode emoji.
|
|
# Reconfigure stdout to UTF-8 so startup messages do not crash the server.
|
|
if sys.platform == "win32" and hasattr(sys.stdout, "reconfigure"):
|
|
try:
|
|
sys.stdout.reconfigure(encoding = "utf-8", errors = "replace")
|
|
except Exception:
|
|
pass
|
|
|
|
# Set env var BEFORE importing main so CORS middleware picks it up
|
|
if api_only:
|
|
os.environ["UNSLOTH_API_ONLY"] = "1"
|
|
|
|
import nest_asyncio
|
|
|
|
nest_asyncio.apply()
|
|
|
|
import asyncio
|
|
from threading import Thread, Event
|
|
import uvicorn
|
|
|
|
from main import app, setup_frontend
|
|
from utils.paths import ensure_studio_directories
|
|
|
|
# Create all standard directories on startup
|
|
ensure_studio_directories()
|
|
|
|
# Auto-find free port if requested port is in use
|
|
if not _is_port_free(host, port):
|
|
original_port = port
|
|
blocker = _get_pid_on_port(port)
|
|
port = _find_free_port(host, port + 1)
|
|
if not silent:
|
|
print("")
|
|
print("=" * 50)
|
|
if blocker:
|
|
pid, name = blocker
|
|
print(
|
|
f"Port {original_port} is already in use by " f"{name} (PID {pid})."
|
|
)
|
|
else:
|
|
print(f"Port {original_port} is already in use.")
|
|
print(f"Unsloth Studio will use port {port} instead.")
|
|
print(f"Open http://localhost:{port} in your browser.")
|
|
print("=" * 50)
|
|
print("")
|
|
|
|
# Setup frontend if path provided (skip in api-only mode).
|
|
# Falls back through alternate locations if the default lacks a built
|
|
# dist; errors out loudly rather than silently serving 404 on `/`.
|
|
if frontend_path and not api_only:
|
|
chosen, attempted = _resolve_frontend_path(Path(frontend_path))
|
|
if chosen is not None and setup_frontend(app, chosen):
|
|
if not silent:
|
|
# Resolve so logs always show an absolute path for support.
|
|
try:
|
|
display = chosen.resolve()
|
|
except OSError:
|
|
display = chosen
|
|
print(f"[OK] Frontend loaded from {display}")
|
|
else:
|
|
home_str = (
|
|
os.environ.get("UNSLOTH_STUDIO_HOME")
|
|
or os.environ.get("STUDIO_HOME")
|
|
or str(Path.home() / ".unsloth" / "studio")
|
|
)
|
|
# Windows ships the user-facing shim at $STUDIO_HOME/bin/unsloth.exe
|
|
# (a hardlink to the venv exe); Linux/macOS use the venv binary
|
|
# at $STUDIO_HOME/unsloth_studio/bin/unsloth.
|
|
home = Path(home_str).expanduser()
|
|
if sys.platform == "win32":
|
|
installer_bin = home / "bin" / "unsloth.exe"
|
|
else:
|
|
installer_bin = home / "unsloth_studio" / "bin" / "unsloth"
|
|
tried_lines = "\n".join(f" - {p}" for p in attempted) or " (none)"
|
|
raise SystemExit(
|
|
"[ERROR] Studio frontend build not found.\n"
|
|
f"Tried:\n{tried_lines}\n"
|
|
"\n"
|
|
"Likely cause: another 'unsloth' on PATH is shadowing the "
|
|
"installer's binary and points at a site-packages tree with "
|
|
"no built dist.\n"
|
|
"\n"
|
|
"Fix one of:\n"
|
|
f" - run the installer's binary directly: {installer_bin} studio\n"
|
|
" - pass --frontend <path/to/studio/frontend/dist>\n"
|
|
" - pass --api-only to skip serving the web UI\n"
|
|
" - reinstall: curl -fsSL https://unsloth.ai/install.sh | sh"
|
|
)
|
|
|
|
# Resolve once; shared by the log rewrite and the banner.
|
|
display_host = _resolve_external_ip() if host == "0.0.0.0" else host
|
|
_install_uvicorn_startup_log_rewrite(host, display_host)
|
|
|
|
ready_event = Event()
|
|
startup_failed = Event()
|
|
startup_errors = []
|
|
|
|
class _ReadyServer(uvicorn.Server):
|
|
async def startup(self, *args, **kwargs):
|
|
await super().startup(*args, **kwargs)
|
|
if getattr(self, "started", False) and not self.should_exit:
|
|
ready_event.set()
|
|
|
|
# server_header=False suppresses uvicorn's "Server: uvicorn"; SecurityHeadersMiddleware sets its own.
|
|
config = uvicorn.Config(
|
|
app,
|
|
host = host,
|
|
port = port,
|
|
log_level = "info",
|
|
access_log = False,
|
|
server_header = False,
|
|
)
|
|
_server = _ReadyServer(config)
|
|
_shutdown_event = Event()
|
|
|
|
# Expose the actual bound port so request-handling code can build
|
|
# loopback URLs that point at the real backend, not whatever port a
|
|
# reverse proxy or tunnel exposed in the request URL. Only publish
|
|
# an explicit value when we know the concrete port; for ephemeral
|
|
# binds (port==0) leave it unset and let request handlers fall back
|
|
# to the ASGI request scope or request.base_url.
|
|
app.state.server_port = port if port and port > 0 else None
|
|
app.state.llama_parallel_slots = llama_parallel_slots
|
|
|
|
# Expose a shutdown callable via app.state before the server can accept
|
|
# requests so /api/shutdown is available as soon as readiness is published.
|
|
def _trigger_shutdown():
|
|
_graceful_shutdown(_server)
|
|
if _shutdown_event is not None:
|
|
_shutdown_event.set()
|
|
|
|
app.state.trigger_shutdown = _trigger_shutdown
|
|
|
|
# Run server in a daemon thread
|
|
def _run():
|
|
try:
|
|
asyncio.run(_server.serve())
|
|
except BaseException as exc:
|
|
startup_errors.append(exc)
|
|
startup_failed.set()
|
|
finally:
|
|
if not ready_event.is_set():
|
|
startup_failed.set()
|
|
|
|
thread = Thread(target = _run, daemon = True)
|
|
thread.start()
|
|
|
|
# Wait until uvicorn has completed lifespan startup and bound sockets, or
|
|
# until the server exits/fails before startup. This intentionally has no
|
|
# correctness deadline: a slow but live startup should remain in progress.
|
|
try:
|
|
while not ready_event.is_set():
|
|
if startup_failed.is_set() or not thread.is_alive():
|
|
if startup_errors:
|
|
raise RuntimeError(
|
|
"Uvicorn server failed before startup completed"
|
|
) from startup_errors[0]
|
|
raise RuntimeError("Uvicorn server exited before startup completed")
|
|
ready_event.wait(timeout = 0.1)
|
|
except KeyboardInterrupt:
|
|
_graceful_shutdown(_server)
|
|
_shutdown_event.set()
|
|
raise
|
|
|
|
_write_pid_file()
|
|
import atexit
|
|
|
|
atexit.register(_remove_pid_file)
|
|
|
|
# Output port for Tauri to parse when in api-only mode. Emit only after
|
|
# uvicorn sockets are bound and FastAPI lifespan/startup has completed.
|
|
if api_only:
|
|
print(f"TAURI_PORT={port}", flush = True)
|
|
|
|
if not silent:
|
|
wildcard_bind = host in ("0.0.0.0", "::")
|
|
# For wildcard binds, run the reachability check between the URL
|
|
# section and the stop hint so the stop hint stays last on screen.
|
|
print_studio_access_banner(
|
|
port = port,
|
|
bind_host = host,
|
|
display_host = display_host,
|
|
include_stop_hint = not wildcard_bind,
|
|
)
|
|
if wildcard_bind:
|
|
_verify_global_reachability(display_host, port)
|
|
print_studio_stop_hint()
|
|
|
|
return app
|
|
|
|
|
|
# For direct execution (also invoked by CLI via os.execvp / subprocess)
|
|
if __name__ == "__main__":
|
|
import argparse
|
|
import signal
|
|
import traceback
|
|
|
|
# Ensure stderr can handle Unicode on Windows (tracebacks with non-ASCII paths)
|
|
if sys.platform == "win32" and hasattr(sys.stderr, "reconfigure"):
|
|
try:
|
|
sys.stderr.reconfigure(encoding = "utf-8", errors = "replace")
|
|
except Exception:
|
|
pass
|
|
|
|
parser = argparse.ArgumentParser(description = "Run Unsloth UI Backend server")
|
|
parser.add_argument(
|
|
"--host",
|
|
default = "127.0.0.1",
|
|
help = "Host to bind to (default: 127.0.0.1; use 0.0.0.0 for network/cloud access)",
|
|
)
|
|
parser.add_argument("--port", type = int, default = 8888, help = "Port to bind to")
|
|
parser.add_argument(
|
|
"--frontend",
|
|
type = str,
|
|
default = _DEFAULT_FRONTEND_PATH,
|
|
help = "Path to frontend build",
|
|
)
|
|
parser.add_argument("--silent", action = "store_true", help = "Suppress output")
|
|
parser.add_argument(
|
|
"--api-only",
|
|
action = "store_true",
|
|
help = "API server only, no frontend (for Tauri)",
|
|
)
|
|
# Mirror unsloth_cli/commands/studio.py's _PARALLEL_*. Default 1
|
|
# applies only to direct backend launches; `unsloth studio run`
|
|
# always passes its own value (4) explicitly.
|
|
_PARALLEL_MIN = 1
|
|
_PARALLEL_MAX = 64
|
|
_PARALLEL_DEFAULT_PLAIN = 1
|
|
parser.add_argument(
|
|
"--parallel",
|
|
"--n-parallel",
|
|
type = int,
|
|
default = _PARALLEL_DEFAULT_PLAIN,
|
|
help = (
|
|
f"llama-server parallel decode slots ({_PARALLEL_MIN}..{_PARALLEL_MAX}). "
|
|
f"Default {_PARALLEL_DEFAULT_PLAIN}; `unsloth studio run` uses 4."
|
|
),
|
|
)
|
|
|
|
args = parser.parse_args()
|
|
if not _PARALLEL_MIN <= args.parallel <= _PARALLEL_MAX:
|
|
parser.error(f"--parallel must be between {_PARALLEL_MIN} and {_PARALLEL_MAX}")
|
|
|
|
kwargs = dict(
|
|
host = args.host,
|
|
port = args.port,
|
|
silent = args.silent,
|
|
api_only = args.api_only,
|
|
llama_parallel_slots = args.parallel,
|
|
)
|
|
if args.frontend is not None:
|
|
kwargs["frontend_path"] = Path(args.frontend)
|
|
|
|
try:
|
|
run_server(**kwargs)
|
|
except Exception:
|
|
sys.stderr.write("\n")
|
|
sys.stderr.write("=" * 60 + "\n")
|
|
sys.stderr.write("ERROR: Unsloth Studio failed to start.\n")
|
|
sys.stderr.write("=" * 60 + "\n")
|
|
traceback.print_exc(file = sys.stderr)
|
|
sys.stderr.write("\n")
|
|
sys.stderr.write(
|
|
"If a package is missing, try re-running: unsloth studio setup\n"
|
|
)
|
|
sys.stderr.flush()
|
|
sys.exit(1)
|
|
|
|
# Signal handler -- ensures subprocess cleanup on Ctrl+C
|
|
def _signal_handler(signum, frame):
|
|
_graceful_shutdown(_server)
|
|
_shutdown_event.set()
|
|
|
|
signal.signal(signal.SIGINT, _signal_handler)
|
|
signal.signal(signal.SIGTERM, _signal_handler)
|
|
|
|
# On Windows, some terminals send SIGBREAK for Ctrl+C / Ctrl+Break
|
|
if hasattr(signal, "SIGBREAK"):
|
|
signal.signal(signal.SIGBREAK, _signal_handler)
|
|
|
|
# Keep running until shutdown signal.
|
|
# NOTE: Event.wait() without a timeout blocks at the C level on Linux,
|
|
# which prevents Python from delivering SIGINT (Ctrl+C). Using a
|
|
# short timeout in a loop lets the interpreter process pending signals.
|
|
while not _shutdown_event.is_set():
|
|
_shutdown_event.wait(timeout = 1)
|