Commit graph

732 commits

Author SHA1 Message Date
danielhanchen
eae716675b studio/sandbox: close 7 bypass classes from cross-reviewer round-6 audit
Round-6 sonnet-panel review surfaced seven more concrete bypass
classes. All seven are now closed (629 tests passing, 60 new R6
regression tests):

1. Path traversal under home prefix. `~/../etc/shadow`,
   `~root/../etc/shadow`, `~ubuntu/../../etc/shadow`, and
   `/home/u/../u/.aws/credentials` all slipped because
   `_normalize_path_separators` re-attached the home prefix even
   when `..` had escaped it. Now when the tail of a home-prefixed
   path begins with `..`, the projection is treated as absolute
   (mirroring the runtime resolve when HOME is a single-segment path
   like `/root`). POSIX `~<user>/` tilde-user expansion is
   handled the same way.

2. Pandas / numpy reader keyword arguments. `pd.read_csv(filepath_or_buffer='/etc/shadow')`,
   `pd.read_excel(io=...)`, `np.fromfile(fname=...)` used the actual
   API parameter names that the previous gate's `{"file", "path"}`
   kwarg list missed. The kwarg set is expanded to cover
   `filepath_or_buffer`, `path_or_buf`, `fname`, `filename`, `io`,
   `buf`, `source`, `src` and a few common variants.

3. Bash directory exfil verbs. `cp -r ~/.ssh /tmp/out`, `mv ~/.aws /tmp`,
   `tar czf out.tar.gz ~/.ssh`, `rsync -av ~/.aws/ /tmp/`,
   `zip -r out.zip ~/.ssh` previously slipped because
   `_find_sensitive_paths` only flagged named files, leaving the
   bash side asymmetric to the Python shutil dir-exfil gate. A new
   `_BASH_DIR_EXFIL_RE` matches dir-copy verbs (`cp`, `mv`, `rsync`,
   `tar`, `zip`, `7z`, `scp`, `sftp`, `xz`) followed by a sensitive
   directory. `ls ~/.ssh` and `find ~/.aws -type f` stay allowed.

4. Inner-tree alias walk for eval / exec. `exec("import shutil as sh\nsh.copytree('~/.ssh', dst)")`
   slipped because the inner AST visit did not re-run the
   alias-tracking pre-pass that built `shutil_module_aliases`.
   Extracted both pre-pass loops into helpers (`_run_alias_prepass`,
   already had `_run_string_binding_prepass`) and call both on each
   literal eval / exec payload before the visitor recurses.

5. Chained assignment. `a = b = '/etc/shadow'; open(a).read()`
   slipped because the binding pre-pass only handled
   `len(targets) == 1`. Multi-target Assign nodes now bind every
   Name target to the resolved value.

6. Annotated assignment. `path: str = '/etc/shadow'; open(path)`
   slipped because the binding pre-pass walked `ast.Assign` but
   not `ast.AnnAssign`. Same-shape handler for single Name target.

7. Brace-bomb empty-alt bypass. `cat ~/{,x0,...,x341}/{.ssh/id_rsa,other}`
   exhausts the expansion cap before the empty alt's second-brace
   projection reaches `~/.ssh/id_rsa`. A defensive
   `_SENSITIVE_IN_BRACE_RE` catches sensitive-name fragments inside
   an unexpanded brace group attached to a sensitive root,
   regardless of whether the brace expansion completed. Anchored
   with `(?<=[,{/])` lookbehind and `(?=,|\}|/)` lookahead so
   project-local lookalikes (`./workspace/home/u/{a,b}/...`) stay
   allowed via the `_PATH_TOKEN_START` boundary.

NetworkAndIoVisitor inner-tree pre-pass. The visitor eval / exec
recursion now mirrors SignalEscapeVisitor's call to both
`_run_alias_prepass` and `_run_string_binding_prepass` so it is
independently correct regardless of visitor execution order.

Cumulative bypass closures across rounds 1 through 6: 24 distinct
classes, 629 regression tests, three-OS green.
2026-05-24 15:14:33 +00:00
pre-commit-ci[bot]
178bcf70d1 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 14:47:25 +00:00
danielhanchen
5eb06e4bfe studio/sandbox: close Subscript + UDP/connect_ex metadata bypasses
Two more from the follow-up list closed (569 tests passing):

1. ``ast.Subscript`` resolution. ``open(['/etc/shadow'][0])`` and
   ``open({'k': '/etc/shadow'}['k'])`` previously slipped because
   ``_extract_string_from_node`` had no Subscript handler. List /
   tuple / dict subscripts are now resolved: when the index is a
   static constant we return the indexed value; otherwise any
   sensitive entry in the container surfaces so the gate fires.
   Indexes outside the container's static range fall back to
   sensitive-scan + first-resolvable so adversarial patterns like
   ``open(['safe.txt', '/etc/shadow'][i])`` are still blocked.

2. UDP / ``connect_ex`` metadata destination. The connect-only
   ``NetworkAndIoVisitor`` gate missed ``s.sendto(data, address)`` /
   ``s.sendmsg(buffers, ancdata, flags, address)`` (the destination
   tuple is positional but not at index 0) and ``s.connect_ex(addr)``
   (non-raising connect variant). The visitor now matches the full
   ``{connect, connect_ex, sendto, sendmsg}`` set and scans every
   positional arg for a ``(host, port)`` tuple shape; the first
   resolved host wins.

17 new regression tests cover the Subscript class (8 blocked, 3
allowed) and the UDP / connect_ex class (4 blocked, 2 allowed).

After this commit, ``bypass_hunt.py`` reports zero NEW bypasses;
the only remaining ALLOWs are the documented follow-up list
(``getattr(__builtins__, ...)``, ``vars(__builtins__)[...]``,
``base64.b64decode`` of paths, ``chr()`` / ``str.join`` concat,
trusted-host upload-shape evasion).
2026-05-24 14:47:11 +00:00
pre-commit-ci[bot]
2bf12cd89b [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 14:43:50 +00:00
danielhanchen
e400dac77d studio/sandbox: close bash glob and ternary IfExp bypasses
Two final bypass classes from the round-5 follow-up list are now
closed (552 tests passing):

1. Bash glob under a sensitive root. ``cat /etc/sha*ow``, ``cat
   /etc/sh?dow``, ``cat /etc/*``, ``cat ~/.ssh/id_*`` -- the shell
   expands ``*`` / ``?`` against the filesystem at runtime, so the
   literal-path scan never sees ``/etc/shadow``. A new
   ``_SENSITIVE_ROOT_WITH_GLOB_RE`` mirrors the existing
   ``_SENSITIVE_ROOT_WITH_EXPANSION_RE`` (which gates ``$(...)`` /
   backtick substitutions) for the ``*`` / ``?`` family. The
   ``[^\s'\";&|`$]*`` literal-text-only constraint keeps the match
   attached to the sensitive root token, so ``find /etc/ -name
   '*.conf'`` (whitespace between root and glob) and project-local
   globs like ``./src/*.py`` stay allowed.

2. Ternary ``IfExp`` branches. ``open('/etc/shadow' if cond else
   'data.txt')`` previously slipped because
   ``_extract_string_from_node`` had no ``ast.IfExp`` handler.
   Either branch can execute at runtime; the gate now resolves both
   branches and prefers the sensitive one (via the same
   ``_looks_sensitive`` check that backs the binding-bias) so the
   downstream check fires. Falls back to whichever branch resolved
   when neither is sensitive.

24 new regression tests cover the glob class (10 blocked, 6 allowed)
and ternary (6 blocked, 2 allowed).
2026-05-24 14:43:23 +00:00
danielhanchen
115810eae3 studio/sandbox: include /etc/passwd in pre-pass binding bias
_find_sensitive_paths does not match /etc/passwd (it lives in the
open-call gate's _SENSITIVE_FILE_PREFIXES list, not in _ABSOLUTE_SENSITIVE),
so a chained reassign p='/etc/hosts'; p='/etc/passwd'; open(p) kept
/etc/hosts as the representative and slipped through. Duplicate the
open-call prefix list in the pre-pass scope so _looks_sensitive catches
/etc/passwd and the analogous /proc/<pid> reads too.
2026-05-24 14:41:33 +00:00
danielhanchen
35d7b72832 studio/sandbox: use _find_sensitive_paths for binding-bias instead of substring
Refines the round-5 ``_looks_sensitive`` heuristic that biases
``_record_string_binding`` toward the dangerous value in a chained
reassignment. The substring hint set conflated ``/etc/shadow`` with
``/etc/hosts`` -- both contain ``/etc/`` -- so a payload like
``p = '/etc/hosts'; p = '/etc/shadow'; open(p)`` had ``cur`` already
flagged sensitive, the guard refused to update, and the resolved
value stayed at ``/etc/hosts`` (allow-listed). The chained shadow
binding then slipped through.

Now ``_looks_sensitive`` delegates to ``_find_sensitive_paths``, the
authoritative bash / file gate matcher, so the distinction is exact:
``/etc/hosts`` is allow-listed and ``/etc/shadow`` is sensitive.
``_record_string_binding`` also adopts a clean three-way rule mirroring
Python's last-wins semantics for sensitive values:

  * New sensitive value: always wins (covers the chained shadow case).
  * New benign value, current sensitive: keep current (static gate
    cannot prove the new value executes; err on blocking).
  * Both benign: latest seen wins.

Test suite still 528 passing.
2026-05-24 14:39:16 +00:00
pre-commit-ci[bot]
b038848482 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 14:35:56 +00:00
danielhanchen
17739721da studio/sandbox: close 7 bypass classes from cross-reviewer round-5 audit
Sonnet-panel review of round-4 surfaced seven concrete bypass classes
in the static gate. All seven are now closed (528 tests passing, 55
new R4 / R5 regression tests):

1. ``os.path.join`` alias bypasses. ``import os as o; o.path.join(...)``,
   ``from os.path import join`` (and ``as j``), ``from os import path``
   (and ``as op``), ``import posixpath as pp``, ``from posixpath import
   join`` -- previously the FQ match was literal-only (`os.path.join`,
   `posixpath.join`, `ntpath.join`). A pre-pass walk collects every
   alias of ``os`` / ``os.path`` / ``posixpath`` / ``ntpath`` and every
   from-import of ``join`` / ``expanduser``; the resolver checks
   ``<alias>.join`` and bare aliased names too.

2. ``shutil`` alias bypasses. ``import shutil as sh; sh.copy(...)``,
   ``from shutil import copyfile``, ``from shutil import move as mv``,
   etc. -- the file-copy gate matched only the literal ``shutil.X`` FQ.
   The pre-pass now tracks shutil module aliases and from-import
   aliases for ``copyfile`` / ``copy`` / ``copy2`` / ``copytree`` /
   ``move``; the gate canonicalises any matched alias to ``shutil.X``
   so the error message identifies the operation.

3. First-assignment-wins binding bypass. ``p = '/tmp/safe'; p =
   '/etc/shadow'; open(p)`` previously slipped because the pre-pass
   guard ``_target.id not in string_bindings`` ignored every
   reassignment, and the AST walk picked the safe value while Python
   uses last-wins at runtime. New ``string_bindings_all`` tracks every
   literal ever bound to a name; ``_record_string_binding`` biases the
   representative value toward sensitive-shaped paths via a substring
   hint set covering the credential / process-state root tokens. The
   reverse order (``shadow`` then ``safe``) is also caught.

4. Brace-expansion off-by-one. ``cat ~/.aws/{x0,...,x62,credentials}``
   exploited that ``_expand_brace_projections`` started with
   ``out = {original}`` (1 item) so a cap of 64 only left 63
   alternative slots. The inner loop also broke per-alternative on
   the cap, so the sensitive name at position 64+ was never reached.
   Raised the cap to 1024 and the inner loop now expands all
   alternatives of a brace in one pass before the outer cap can stop
   the queue.

5. ``thread-self`` in shell-expansion regex. ``cat /proc/thread-self/
   $(echo environ)`` was missed because ``_SENSITIVE_ROOT_WITH_EXPANSION_RE``
   only listed ``self|\d+`` in the ``/proc/...`` alternation, while
   ``_ABSOLUTE_SENSITIVE`` correctly included ``thread-self``. One
   alternation entry restores symmetry.

6. Eval / exec pre-pass not re-run. ``exec("p='/etc/shadow'\nopen(p)")``
   slipped because the inner AST visit ran without the string-binding
   pre-pass. Extracted the pre-pass into ``_run_string_binding_prepass``
   and call it on each inner literal payload before the visitor
   recurses, so payload-local variable assignments are visible.

7. Pathlib name binding pre-pass. ``p = Path('/etc/shadow');
   p.read_text()`` slipped because the pre-pass only resolved string
   literals -- pathlib constructor calls returned None and the bound
   name remained unresolved. Pre-pass now falls back to
   ``_extract_pathlib_target`` using per-tree alias sets so
   ``import pathlib as pl; p = pl.Path(...)`` and ``from pathlib import
   Path as P; p = P(...)`` both resolve. ``NamedExpr`` (walrus) is also
   surfaced by the pre-pass so walrus-inside-eval expressions are
   visible.

Pre-pass call order. The initial pre-pass invocation moves to AFTER
``_extract_pathlib_target`` is defined so the closure cell binds
correctly (Python looks up free variables in the enclosing scope at
CALL time, not at function-definition time).

Full sandbox suite: 528 passed (455 prior + 73 R4 / R5 regression tests).
2026-05-24 14:35:17 +00:00
pre-commit-ci[bot]
a6400ffc07 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 14:17:45 +00:00
danielhanchen
02e9e4867d studio/sandbox: close round-4 bypass classes (aliasing, FileIO, walrus, copytree)
Closes additional bypass classes surfaced while exercising the gate:

1. Module / function aliasing (`m = os; m.system(...)`,
   `p = os.popen; p(...)`): `visit_Assign` now propagates the source
   alias when one tracked-module name is bound to another, and tracks
   bound method references into `shell_exec_aliases`. Previously only
   `m = __import__('os')` was handled.

2. Importlib from-alias (`from importlib import import_module as IM;
   IM('os').system(...)`): a new visitor-scope `import_module_aliases`
   set plus a `_resolve_dynamic_module` wrapper recognises the
   bound name in both inline-call and bound-name forms.

3. Shutil directory exfil (`shutil.copytree('~/.ssh', dst)`):
   `_matches_sensitive_dir()` adds a directory-only matcher used by
   the file-copy gate only. The boundary `(?=/?$|/?[\s'\";&|)<>])`
   matches the path AS the directory but NOT a single file inside it,
   so per-file allow-listed reads (`~/.ssh/known_hosts`, `~/.ssh/id_rsa.pub`)
   still pass. Covers `.ssh`, `.aws`, `.config/gcloud`, `.gnupg`,
   `.docker`, `.kube`, `.password-store`, plus `/etc`, `/etc/ssh`,
   `/var/spool/cron`, `/proc/<pid>`.

4. Explicit-reader / aliased file readers (`io.FileIO('/etc/shadow')`,
   `codecs.open('/etc/shadow')`, `from io import FileIO; FileIO(...)`):
   the open-call detector now recognises these qualified forms and
   the visitor tracks `from io|codecs import FileIO|open` aliases.

5. Bytes-literal paths (`open(b'/etc/shadow')`) and walrus
   expressions (`open((p := '/etc/shadow'))`): `_extract_string_literal`
   and `_extract_string_from_node` resolve `bytes` Constants via strict
   UTF-8 decode and `NamedExpr` via RHS extraction (recording the
   binding so later uses of the walrus target resolve too).

6. Tuple / list unpacking destructuring (`(a, b) = ('/etc', 'shadow');
   open(a + '/' + b)` and `p, = ['/etc/shadow']; open(p)`): the
   string-binding pre-pass now folds matched-length Tuple/List
   destructurings element-wise.

7. Pandas / numpy file readers (`pd.read_csv('/etc/shadow')`,
   `np.fromfile('/etc/shadow')`, etc.): suffix-match the common
   reader method names so any alias of the source module flows
   through the same sensitive-path gate as `open()`.

91 new regression tests cover each class, both blocked and legitimate
allow-list cases. Full sandbox suite: 473 passed.
2026-05-24 14:17:20 +00:00
pre-commit-ci[bot]
60e0f5056e [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-22 08:13:56 +00:00
Daniel Han
d64c2a10d4 studio/sandbox: close dynamic-import + /proc/self symlink bypasses
Closes two static-bypass classes flagged during round-3 review:

1. __import__('os').system(...) / importlib.import_module('os').popen(...)
   bypassed the bare os.system / subprocess.* gate because the receiver
   was an ast.Call rather than an ast.Name in os_aliases. Adds
   _resolve_dynamic_module_name() so:
     * inline __import__('os').system(...)
     * inline importlib.import_module('os').system(...)
     * import importlib; mod = importlib.import_module('os'); mod.popen(...)
     * m = __import__('subprocess'); m.run([...], shell=True)
   all flow through the same shell-escape detection as
   import os; os.system(...). Legit dynamic imports of safe modules
   (json, pathlib, ...) remain allowed.

2. /proc/<pid>/cwd and /proc/<pid>/root are symlinks to the process
   working directory and the filesystem root. The form
   open(/proc/self/cwd/../../etc/shadow) bypassed
   _normalize_path_separators because .. was collapsed against the
   literal path, not the symlinked target. The form
   open(/proc/self/root/etc/shadow) bypassed any chroot-style
   defence. Adds matching entries to _ABSOLUTE_SENSITIVE so the bash
   gate and the AST open() gate both block any access via these
   symlink prefixes. Legitimate /proc/self/status etc. introspection
   still flows.

Tests:
  TestFollowup_DynamicImportShellEscape (2 cases, 10 parametrised)
  TestFollowup_ProcSelfSymlinkTraversal (3 cases, 16 parametrised)

  pytest studio/backend/tests/test_sandbox_hardening.py -q
    -> 382 passed in 0.67s (was 356).
2026-05-22 08:13:37 +00:00
pre-commit-ci[bot]
f2bbe27cde [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-19 12:33:45 +00:00
Daniel Han
5b0735308a Merge branch 'studio-sandbox-hardening' of https://github.com/unslothai/unsloth into studio-sandbox-hardening
# Conflicts:
#	studio/backend/core/inference/tools.py
2026-05-19 12:33:29 +00:00
Daniel Han
8176694d94 studio: address round-3 sandbox review findings
Round-4 follow-up on the hardening PR after a third 20-reviewer pass.
Closes the high-impact items from that review while preserving the
"do not regress legitimate tool calling" floor; lower-vote items that
would have measurable regression on legit code paths (broad shell
glob ?/*, $VAR in dynamic paths, ANSI-C $'...') are intentionally
deferred.

Parent-directory traversal: _normalize_path_separators now follows
.. segments through posixpath.normpath and reattaches the tilde or
${HOME} prefix, so cat /etc/apt/../shadow and
Path('/proc/self/fd/../environ').read_text() both reach the
canonical regex.

Built-in open() accepts PathLike: open(Path('/etc/shadow')) and
open(file=Path('/etc/shadow')) now flow through the pathlib resolver
the same way receiver reads do.

Pathlib home and transforms: Path.home() resolves to ~ so
(Path.home() / '.aws/credentials') hits the home regex;
.expanduser() / .resolve() / .absolute() are pass-throughs.

Pathlib semantics: _join_path_parts() now matches pathlib's
absolute-segment reset so Path('/tmp') / '/etc/shadow' resolves to
/etc/shadow as it does at runtime.

from builtins import exec as e: tracked in both visitors via
eval_exec_aliases so the aliased call still routes through the
literal-payload recursion.

Process state extensions: /proc/self/cmdline,
/proc/thread-self/*, and /proc/<pid>/task/<tid>/* are added to
_ABSOLUTE_SENSITIVE.

Numeric f-strings: f'/proc/{1}/environ' folds to a literal because
numeric ast.Constant values inside ast.FormattedValue are now
stringified.

os.path.join / os.path.expanduser: resolved statically by
_extract_string_from_node so the stdlib-helper construction paths
do not hide sensitive targets.

Variable assignment tracking: a pre-pass collects ``name = literal``
and ``name = eval`` / ``name = exec`` bindings; the visitors and the
pathlib resolver consult those bindings. The trusted-host gate
intentionally uses a separate strict literal extractor so legit
patterns like ``url = some_input; requests.get(url)`` still pass.

shutil.copyfile / copy / copy2 / copytree / move: the source argument
is gated the same way open() is, blocking file-copy exfil.

Concrete pathlib classes: PosixPath / WindowsPath / PurePath / etc.
are registered in path_aliases by default.

requests.request positional+keyword: for URL-second APIs, args[0] is
the HTTP method (not the URL); when there is only one positional, the
URL extraction falls through to the url= keyword instead of grabbing
the method.

Tests grow from 281 to 357 hardening cases; combined sweep 487 / 487.
Every fix has positive and negative coverage; legit tool calls
(open(Path('data.csv')), os.path.join('logs', 'today.log'),
url = some_input; requests.get(url), shutil.copyfile('a.txt', 'b.txt'))
continue to pass.
2026-05-19 12:32:29 +00:00
pre-commit-ci[bot]
f2809221d4 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-19 12:02:09 +00:00
Daniel Han
7a0bacdba8 studio: address round-2 sandbox review findings
Round-3 follow-up on the same hardening PR after a second 20-reviewer
pass. Every change is detection-widening; legitimate tool calls remain
allowed.

Pathlib readers (Path.open, Path.read_text, Path.read_bytes) now share
one extraction path. The new _extract_pathlib_target helper resolves
Path(a), Path(a, b, ...), Path(...).joinpath(b), and Path(...) / b
through statically-resolvable string parts. NetworkAndIoVisitor tracks
Path aliases (from pathlib import Path as P) and pathlib module aliases
(import pathlib as pl) so the aliased forms hit the same gate. For
receiver-side reads the path is taken exclusively from the receiver --
Path('/etc/shadow').open('r') no longer mis-reads the mode flag as a
path.

The credential-path regex set now matches the POSIX ~user/ expansion
(cat ~ubuntu/.aws/credentials), and the SSH private-key end anchor
includes > so that redirect-attached forms (cat ~/.ssh/id_rsa>... )
are not split-tokenised through the gate. A new
_SENSITIVE_ROOT_WITH_EXPANSION_RE detects sensitive root prefixes
followed by $(...) or backtick substitution, and _find_sensitive_paths
now enumerates bash brace expansion {a,b} plus small glob char classes,
and runs every projection through path-separator normalisation that
collapses // and /./.

Network host validation reaches keyword arguments (url=, host=,
hostname=, address=), host-first APIs whose first positional arg is the
host (socket.create_connection, socket.getaddrinfo,
http.client.HTTPConnection, http.client.HTTPSConnection), and the
url-second APIs (requests.request, httpx.request).

builtins.exec, builtins.eval, and __builtins__.eval flow through the
same literal-payload recursion as the bare forms, including aliased
import builtins as b. open(file=...) and io.open(file=...) keyword
forms are gated alongside the positional form, and the open() path
candidates run through both backslash normalisation and the
//-collapse projection so equivalent spellings (/etc//shadow,
/etc/./shadow) cannot bypass.

Tests grow from 205 to 281 hardening cases (TestR2Finding1 through
TestR2Finding16) and from 131 + 205 = 336 to 131 + 281 = 412 in the
local sweep. Negative cases for every fix continue to ensure
legitimate tool use (Path('data.csv').open(), open(file='logs/today.log'),
requests.get(url='https://wikipedia.org/'), find src/, etc.) stays
allowed.
2026-05-19 12:00:34 +00:00
Daniel Han
e598ebf1d0 Merge remote-tracking branch 'origin/main' into studio-sandbox-hardening 2026-05-19 11:51:10 +00:00
alkinun
b01a1ba1c2
Fix GGUF multi-image chat handling (#5508)
Preserves per-turn OpenAI image_url content parts in the standard GGUF /v1/chat/completions path so multi-image chat history keeps each image attached to its original turn. Legacy top-level image_base64 is injected as a synthetic image_url part only when no message-level image exists. Tool use is disabled whenever any GGUF image is present. Fixes #5470.
2026-05-19 04:36:20 -07:00
Daniel Han
eda3be4101 Merge branch 'studio-sandbox-hardening' of https://github.com/unslothai/unsloth into studio-sandbox-hardening 2026-05-19 11:11:45 +00:00
Daniel Han
8fd57530d1 studio: always use POSIX shlex for sensitive-path dequote
On Windows runners shlex.split(posix=False) leaves splice quotes in place, so cat /etc/sha''dow tokenises to ['cat', "/etc/sha''dow"] and the dequoted scan projection still misses the credential. The threat model is POSIX-shell quote splicing in either bash invoked on Windows or POSIX shells on Linux/macOS; the dequote always wants POSIX semantics. Pre-normalise backslashes so Windows drive paths survive POSIX shlex's escape handling.
2026-05-19 11:11:20 +00:00
pre-commit-ci[bot]
40d60b05c5 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-19 11:09:53 +00:00
Daniel Han
e06126933e Merge branch 'studio-sandbox-hardening' of https://github.com/unslothai/unsloth into studio-sandbox-hardening
# Conflicts:
#	studio/backend/core/inference/tools.py
2026-05-19 11:08:34 +00:00
Daniel Han
2ae885ce72 studio: address sandbox hardening review findings
Round-2 fixes for ten issues surfaced by a 20-reviewer code review of
the initial hardening patches. Every change is detection-widening or a
false-positive narrowing; legitimate tool calls keep working.

Direct Python open now flows through _find_sensitive_paths so
open('/home/u/.aws/credentials').read() is gated the same as
os.system('cat ~/.aws/credentials'). The previous wiring covered only
the bash and shell-exec sides.

Both SignalEscapeVisitor and NetworkAndIoVisitor fail-closed once the
eval / exec literal recursion cap is reached. Wrapping a payload in
four or more nested literal exec layers no longer silently bypasses
inspection.

_find_sensitive_paths scans three projections of the command (raw,
backslash-normalised, shlex-dequoted) and recurses into nested
bash -c and cmd /c shells. Quote-spliced and Windows-backslash forms
of credential paths are all caught.

_HOME_PREFIX_RE adds Windows-style home prefixes (USERPROFILE,
HOMEDRIVE HOMEPATH, env:USERPROFILE, drive-letter Users) so cross-OS
hardening actually applies on Windows. Both sensitive-path regexes
now have a path-token start anchor so project-local lookalike paths
under workspace, fixtures, and tmp are not blocked.

Network host validation for sock.connect and the requests / urllib
FQ-prefix branch now use _extract_string_from_node instead of raw
ast.Constant checks, so concatenated and f-string literal hosts
resolve the same way the open gate already did.

pathlib.Path('/etc/shadow').open() is now inspected; the path is
extracted from the receiver constructor when node.args is empty.

The static-string resolver depth cap moves from 6 to 64, removing the
single-character literal-concat bypass while leaving the recursion
well inside CPython's default frame limit.

The SSH private-key regex gains a filename-end boundary so reading a
public ".pub" key stays allowed (legit developer action) while the
matching private key is still denied.

New regression tests cover one class per finding (TestFinding1 through
TestFinding10) plus updated nested-depth coverage; the previous
test_nested_depth_does_not_crash assertion was inverted by the
fail-closed change and has been replaced. Local sweep: 336 passed
(131 upstream + 205 hardening).
2026-05-19 11:07:16 +00:00
Daniel Han
6efb0f64ea Merge remote-tracking branch 'origin/main' into studio-sandbox-hardening 2026-05-19 10:58:55 +00:00
Daniel Han
5ce4ab4d54
studio: emit one comma-chained --spec-type for CPU/Mac MTP path (#5575)
* studio: emit one comma-chained --spec-type for CPU/Mac MTP path

llama-server takes a single --spec-type whose value may be
comma-separated to chain implementations (e.g. ngram-mod,draft-mtp).
The CPU/Mac MTP branch in LlamaCppBackend.load_model was passing
--spec-type twice in the same invocation, which is not the documented
chaining mechanism and silently drops one of the two specs depending
on llama.cpp's argv handling.

Collapse the pair to --spec-type ngram-mod,{mtp_token} and update the
stale _extra_args_set_spec_type docstring that claimed llama-server
accumulates repeated --spec-type. Update the matching pass-through
fixture in test_llama_server_args.py.

* studio: align MTP ngram-mod knobs with llama.cpp upstream defaults

Two correctness fixes against the llama.cpp server README:

1. The CPU/Mac comma-chained branch was emitting
   --spec-ngram-mod-n-max 6 with --spec-ngram-mod-n-min 48, which is
   nonsensical (min > max). Per the upstream default the value is 64.

2. The standalone ngram-mod branch was emitting --spec-ngram-size-n,
   --draft-min, --draft-max. llama.cpp removed those arg aliases for
   ngram-mod (they live only on the ngram-simple / map families now);
   the correct knobs are --spec-ngram-mod-n-match / n-min / n-max.

Also refresh the inline comment block to point at the server README
rather than the older docs/speculative.md draft- aliases.
2026-05-19 03:16:05 -07:00
pre-commit-ci[bot]
829698280a [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-19 07:52:08 +00:00
Daniel Han
8a5080f26b studio: regression tests for sandbox hardening patches A / B / D
124 tests across 7 classes:

  * TestPatchA_DynamicPaths — open() with concatenated literals + f-strings.
    Pins 7 attack patterns BLOCKED, 6 legitimate dynamic paths ALLOWED,
    12-level deep concat doesn't crash.

  * TestPatchB_FindSensitivePathsHomeAnchored — ~/.ssh/id_*, ~/.aws/,
    ~/.docker/, ~/.kube/, ~/.pypirc/.npmrc, ~/.netrc, ~/.password-store,
    ~/.gnupg/private-keys-v1.d across ~, $HOME, /home/<u>, /root,
    /Users/<u>. Pins 21 attack paths BLOCKED, 15 legitimate paths
    (~/.gitconfig, ~/.bashrc, ~/.ssh/{config,known_hosts}, ~/.npm,
    project-local rc files, /tmp/.npmrc) ALLOWED.

  * TestPatchB_FindSensitivePathsAbsolute — /etc/shadow, /etc/sudoers,
    /etc/ssh/ssh_host_*, /proc/{self,<pid>}/{environ,mem,maps,auxv},
    /proc/kcore, /proc/kallsyms, /var/spool/cron/. Pins 12 attacks
    BLOCKED, 11 legit paths (/etc/hosts, /etc/resolv.conf, /proc/cpuinfo,
    /proc/meminfo, …) ALLOWED.

  * TestPatchB_PythonShellExec — same surface flows through os.system /
    subprocess.run. 6 attacks BLOCKED, 11 legitimate tool-calls ALLOWED.

  * TestPatchD_EvalExecLiteralPayload — exec/eval with a literal payload
    parsed and re-checked. 5 attack payloads BLOCKED, 6 legit
    expressions (eval('1+2'), exec('print("hi")'), nested innocuous
    exec) ALLOWED.

  * TestPatchD_EvalExecDynamicPayload — non-literal eval/exec args
    flagged as dynamic shell escape. 4 patterns BLOCKED.

  * TestPatchD_NestedDepthCap — 10-level nested exec(exec(...)) caps
    at depth 3, doesn't crash, doesn't false-positive.

  * TestCrossCuttingNoRegression — 6 pre-existing BLOCK patterns still
    fire (sudo, signal tampering, /etc/passwd literal, untrusted host,
    metadata host); 7 pre-existing ALLOW patterns still pass (print,
    json.loads, trusted host, dataclass, legitimate open()).

Result on the rebuilt scaffold:
  131/131 pre-existing tests in test_sandbox_tools.py pass
  124/124 new hardening tests pass
  255/255 combined, zero regressions

The "must remain ALLOWED" cases form the non-regression floor that
prevents the patches from making LLM tool calling dumber.
2026-05-19 07:04:49 +00:00
Daniel Han
6984fa8d7c studio: recurse into eval / exec literal payloads (Patch D)
The AST gate previously had no special handling of eval() / exec(),
so a literal payload would slip past every detector:

  exec("import os; os.system('sudo whoami')")          # ALLOWED before
  exec("open('/etc/shadow').read()")                    # ALLOWED before
  eval("__import__('blocked_mod').dangerous()")         # still allowed (chained-call gap)
  payload = '...'; exec(payload)                        # ALLOWED before

Both visitors (SignalEscapeVisitor and NetworkAndIoVisitor) now share
the same gate at the top of visit_Call: when the call is bare-name
eval / exec, try to resolve the first argument via the shared
_extract_string_from_node helper (Patch A); if it resolves, parse it
and recursively visit so every existing detector runs on the inner
code — signal tampering, shell escape, sensitive-file open, network
policy, upload denylist, etc.

When the payload is not statically resolvable, SignalEscapeVisitor
appends a `shell_escape_dynamic` finding — eval/exec of runtime data
is the textbook code-injection vector and there is no legitimate LLM
tool-call reason to dynamically eval an external string. Static
literals (eval('1 + 2'), exec('x = 1\\ny = 2')) are unchanged because
the recursive visit only flags what the rest of the AST gate would
already flag at top level.

Each visitor caps recursion at depth 3 (own counter on the instance)
so adversarial nested eval('eval(...)') cannot blow the stack.

Closes gaps #7, #8 (partial), #11 (partial) from the 13-gap audit.
Out-of-scope chained-call cases (__import__('os').system(...),
getattr(os, 'sys'+'tem')()) stay documented gaps — the OS sandbox is
the intended backstop, see PR 5468.

Regression: 131/131 studio/backend/tests/test_sandbox_tools.py pass.
Legitimate eval/exec on literal expressions (eval('1+2'),
exec('print("hi")'), exec('exec("print(1)")')) verified ALLOWED.
2026-05-19 07:02:47 +00:00
Daniel Han
2965cd8310 studio: gate credential / process-state paths in bash and Python (Patch B)
Adds _find_sensitive_paths() and wires it into _bash_exec (alongside the
existing _find_blocked_commands check) and into _check_args_for_blocked
(so the Python AST gate catches os.system('cat ~/.ssh/id_rsa') the same
way bash $ cat ~/.ssh/id_rsa is caught).

The pattern set is intentionally narrow — only clear-cut credential and
process-state targets:

  Home-anchored (must be prefixed by ~, $HOME, ${HOME}, /home/<u>,
  /root, /Users/<u>):
    .ssh/id_rsa, .ssh/id_ed25519, .ssh/id_ecdsa, .ssh/id_dsa, .ssh/identity
    .aws/credentials, .docker/config.json, .kube/config
    .config/gcloud/{application_default_credentials,access_tokens,credentials}
    .pypirc, .npmrc, .cargo/credentials
    .netrc, .password-store, .gnupg/private-keys-v1.d

  Absolute system targets (match anywhere):
    /etc/shadow, /etc/sudoers, /etc/ssh/ssh_host_*
    /proc/{self,<pid>}/{environ,mem,maps,auxv}
    /proc/kcore, /proc/kallsyms
    /var/spool/cron/

The home-anchored category uses a regex that requires a HOME-equivalent
prefix, so project-local rc files like ./project/.npmrc remain readable
while ~/.npmrc is denied. Legitimate LLM-developer-tool paths
(~/.gitconfig, ~/.bashrc, ~/.ssh/config, ~/.ssh/known_hosts, /etc/hosts,
~/.cache/, ~/.bash_history, project rc files) are intentionally NOT in
the list and still flow through unchanged.

Closes gaps #1, #2, #3, #12, #13 from the documented 13-gap audit.

Regression sweep:
  * 131/131 studio/backend/tests/test_sandbox_tools.py pass
  * 24 legitimate-use cases verified ALLOWED
  * 17 attack patterns verified BLOCKED
2026-05-19 06:57:53 +00:00
Daniel Han
3e4704a856 studio: resolve concatenated + f-string paths in sensitive-file gate (Patch A)
Static-string resolution in _extract_string_from_node was limited to bare
ast.Constant. Concatenated string literals and f-strings with constant
parts evaluated to ast.BinOp / ast.JoinedStr and slipped past the
open() sensitive-file check, so:

  open('/etc/' + 'shadow')           # ALLOWED before
  open(f'/etc/{"shadow"}')           # ALLOWED before
  open('/etc/passwd')                # BLOCKED before (literal)

The helper now resolves ast.BinOp(Add) of two resolvable strings and
ast.JoinedStr whose parts are themselves resolvable. The open()
sensitive-file check uses the helper instead of an inline ast.Constant
isinstance check, so the same widening covers concatenated/f-string
paths without changing what was already blocked.

Resolution is depth-capped at 6 to keep adversarial deep nesting from
blowing the stack. All other call sites of the helper
(_check_args_for_blocked, dynamic-arg shell-escape detection, HF upload
path-shape inspection) automatically inherit the broader resolution.

Closes open() gaps #4 and #6 from the documented 13-gap audit. Does not
attempt to model variable flow (gap #5 stays open by design — the OS
sandbox is the right layer for runtime flow).

Regression: 131/131 studio/backend/tests/test_sandbox_tools.py pass.
2026-05-19 06:54:26 +00:00
Daniel Han
4699c7e291
studio: engage draft-mtp on vision MTP GGUFs (drop incorrect vision gate) (#5560)
* studio: engage draft-mtp on vision MTP GGUFs

The draft-mtp auto-promotion in LlamaCppBackend.load_model was gated on
not effective_is_vision, and the spec-emit branch repeated the same
guard. Every Unsloth -MTP GGUF repo ships an mmproj projector, so
effective_is_vision was always True for those repos and the MTP speedup
silently never engaged out of the box.

llama.cpp #22673 explicitly states MTP is compatible with vision input.
The bundled b9204 server happily loads both: a manual run with
--mmproj ... --spec-type draft-mtp --spec-draft-n-max 6 logs
"loaded multimodal model" followed by
"adding speculative implementation 'draft-mtp'".

Drop the vision gate from both sites and rewrite the matching short
circuit in _already_in_target_state so reload checks reach the auto
promotion path on vision MTP loads. Add three regression tests covering
vision MTP match (auto and default), and non MTP vision repo unaffected.

Verified on a B200 with unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL:
base decode 179.7 t/s vs MTP decode 253.8 t/s, draft acceptance 0.57,
1.41x speedup on a 255 token completion. mmproj still loads and image
input remains available.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: prefer Qwen3.5 -MTP GGUF variants in default model lists

With the vision gate dropped in the previous commit, draft-mtp now
auto-engages on -MTP GGUF repos out of the box. Swap the four Qwen3.5
recommended entries in DEFAULT_MODELS_GGUF and DEFAULT_MODELS_STANDARD
to their -MTP-GGUF counterparts so new users get the speedup by default:

  unsloth/Qwen3.5-4B-GGUF        -> unsloth/Qwen3.5-4B-MTP-GGUF
  unsloth/Qwen3.5-9B-GGUF        -> unsloth/Qwen3.5-9B-MTP-GGUF
  unsloth/Qwen3.5-35B-A3B-GGUF   -> unsloth/Qwen3.5-35B-A3B-MTP-GGUF
  unsloth/Qwen3.5-0.8B-GGUF      -> unsloth/Qwen3.5-0.8B-MTP-GGUF

All four HF repos exist (HEAD 200) and ship the same UD-Q4_K_XL quant
layout as the non-MTP variants. Non-Qwen3.5 entries are untouched.

* bump version to 2026.5.4

Picks up the studio MTP vision-gate fix and the Qwen3.5 -MTP default
swap in this PR.

* studio: prefer Qwen3.6-35B-A3B-MTP-GGUF in default model lists

Same rationale as the previous Qwen3.5 swap. The Qwen3.6 MTP variant
exists at unsloth/Qwen3.6-35B-A3B-MTP-GGUF (HF HEAD 200) and now
auto-engages draft-mtp out of the box with the gate fix.

* studio: drop --spec-draft-n-max from 6 to 3 for draft-mtp

n=6 is too greedy: on Qwen3.6 the draft has to guess 6 tokens ahead
and acceptance crashes to ~0.45, leaving only ~14% throughput gain.

PR ggml-org/llama.cpp#22673's author benched n=3 at ~0.72 acceptance
and 2 to 3x speedup on the same Qwen3.6 family, and the README sample
command uses n=2 or n=3. Match that.

CPU/Mac branch already uses n=3, so this aligns both paths.

* studio: set --spec-draft-n-max back to 6 for draft-mtp on GPU

Reverts the n=3 tuning. n=6 is the original default; user-side comparisons
hold the larger draft window steady so the toggle (next commit) is the
primary on/off lever.

* studio: add Speculative Decoding toggle under Max Tokens

Adds a top-level kill switch (panel-switch under Max Tokens, mirroring
Auto-Healing Tool Calls) that forces the /load request's
speculative_type to "off" when disabled. The backend "off" branch in
LlamaCppBackend.load_model skips both the draft-mtp auto-promotion and
the spec-emit branch, so neither --spec-type draft-mtp nor
--spec-default reaches llama-server.

Wiring:

- chat-runtime-store: new speculativeDecodingEnabled bool, default
  true, persisted to localStorage under unsloth_speculative_decoding,
  plus a setSpeculativeDecodingEnabled setter.
- chat-settings-sheet: SpeculativeDecodingToggle rendered immediately
  beneath the Max Tokens slider for non-external models.
- use-chat-model-runtime: when speculativeDecodingEnabled is false,
  override speculative_type to "off" in the loadModel call so the
  switch wins over any pre-existing speculativeType state (including
  the existing per-model toggle in Model Settings).

Verified end to end on unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL:
toggle ON emits --spec-type draft-mtp --spec-draft-n-max 6; toggle
OFF emits zero --spec-* flags on the same MTP GGUF.

* studio: relocate Speculative Decoding toggle into Model Settings

Move the toggle out from under Max Tokens and back into the Model
Settings section, directly beneath KV Cache Dtype, where the existing
Apply/Reset workflow already drives a reload on dirty. This way flipping
the switch in the UI actually picks up: the section becomes dirty,
Apply re-runs /load with the new speculative_type.

Drop the !currentModelIsMultimodal gate so vision MTP GGUFs can also
disable speculative decoding from the UI.

Switch the toggle's off-value from null to "off" so the backend's "off"
short-circuit fires for MTP models too (null normalises to None which
re-triggers the draft-mtp auto-promotion).

Tooltip now reads "Faster generation with 0% accuracy hit".

Remove the now-redundant speculativeDecodingEnabled bool + setter from
the runtime store and the load-time override in use-chat-model-runtime;
the toggle binds directly to speculativeType.

* studio: restore OOM/TIGHT badge on recommended GGUF rows

The recommended-list row passed vramStatus=null for any GGUF repo
because the existing useRecommendedModelVram hook reads safetensors
totals from HF model info, which GGUF-only repos do not expose. As a
result, an OOM Q-quant repo would render with only a "GGUF" badge and
no visual signal that nothing in it fits.

Add useGgufRecommendedFit: per repo, fetch the variant list via the
existing /api/models/gguf-variants endpoint, take the smallest
variant's size_bytes, and classify with the same 0.7*GPU + 0.7*RAM
thresholds as GgufVariantExpander. Session-scoped cache + in-flight
dedup so a repo is requested at most once.

Wire the result into the three GGUF row sites in pickers.tsx so OOM
and TIGHT badges show on the collapsed cards.

* Revert "studio: restore OOM/TIGHT badge on recommended GGUF rows"

This reverts commit 07793b1240df72b13e51d6dc15f63c4ee8c6cba9.

The new useGgufRecommendedFit hook was treating the symptom. PR #5561
identified the real root cause: useGpuInfo was calling /api/system
with plain fetch instead of authFetch, so the session-auth check
failed silently and gpu.available stayed false everywhere. With no
GPU info, every fit check (variant expander, recommended carousel)
fell back to "no signal" and dropped the OOM/TIGHT badges.

Reverting the over-engineered hook and applying the authFetch fix
in the next commit, which restores the existing badges with one line.

* chore: replace qwen suggested with MTP variant

* fix: restore GPU info auth for GGUF fit badges

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: imagineer99 <samleejackson0@gmail.com>
2026-05-18 08:42:55 -07:00
Roland Tannous
c0cc975c91
fix(studio): handle expired OpenAI shell-tool containers without surfacing error in chat (#5547)
* fix(studio): transparent retry on expired OpenAI shell container

* fix(studio): drop expired OpenAI containers before send
2026-05-18 05:47:57 -07:00
Daniel Han
80d5acafb4
studio: install flash-linear-attention and tilelang for Qwen3.5 family (#5434)
* studio: install flash-linear-attention and tilelang for Qwen3.5 family

Studio currently only installs causal-conv1d for qwen3.5 / qwen3.6 /
qwen3-next models. Without flash-linear-attention installed alongside
it, transformers' Qwen3.5 fast-path gate stays False and the model
falls back to a pure-PyTorch loop for the GatedDeltaNet layers. In a
60-step run on unsloth/Qwen3.5-2B on B200, this fallback costs ~2.35x
vs the full fast path.

On top of that, FLA dispatches its hottest GDN kernels through a
TileLang backend when tilelang is importable. Adding tilelang plus a
pinned apache-tvm-ffi gives another ~26% on the same workload (4.73
s/step to 3.50 s/step) and is what users have been getting indirectly
when they install mamba-ssm (mamba-ssm transitively pulls tilelang and
pins apache-tvm-ffi<=0.1.9, which is the last working version on
sm_100; 0.1.10 and 0.1.11 crash Triton with misaligned address).

Changes:
  * _ensure_flash_linear_attention: pure-Python PyPI install gated on
    the same model match set as _ensure_causal_conv1d_fast_path.
  * _ensure_tilelang_backend: installs apache-tvm-ffi==0.1.9 and
    tilelang==0.1.8 in one pip resolve so the tvm-ffi pin wins over
    tilelang's >=0.1.2 constraint. Gated on the Qwen3.5 family only;
    SSM models (Nemotron-H, Falcon-H1, Granite-H, LFM2) do not use
    FLA's GDN dispatch.
  * UNSLOTH_STUDIO_SKIP_TILELANG_INSTALL=1 escape hatch matching the
    flash-attn pattern.
  * Orchestration block reordered: causal-conv1d -> fla -> mamba-ssm
    -> tilelang -> flash-attn (long context).
  * 7 new tests covering the new helpers, including SSM-model skip,
    skip-env, full Qwen3 family name variants, and graceful pip
    install failure.

Combined Qwen3.5-2B-Vision step time on B200 in our bench goes from
5.0 s/step (current Studio: causal-conv1d only) to 3.5 s/step
(causal-conv1d + fla + tilelang), a 1.43x speedup with no notebook
or user code changes required.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* tests/studio: accept new grad_norm arg in MLX smoke _on_step callback

The MLX trainer's step callback now passes a ninth positional argument
(grad_norm) per unsloth_zoo/mlx/trainer.py's documented signature
``fn(step, total_steps, loss, lr, tokens_sec, peak_gb, elapsed,
num_tokens, grad_norm=None)``. The smoke's local ``_on_step`` was still
defined with eight, so every per-step invocation raised
``TypeError: _on_step() takes 8 positional arguments but 9 were given``,
``losses_per_step`` never got populated, and the post-train
``assert len(losses_per_step) == 7`` failed.

Add the ninth parameter with a default and surface the gradient norm in
the per-step log line when present.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* ci: retrigger after zoo drift + IPython fixes landed in main

* tests/studio: pin max_grad_value=0 in MLX smoke so max_grad_norm=1.0 wins

unsloth_zoo PR #5340 added per-element gradient clipping to MLXTrainer
and defaulted ``MLXTrainingConfig.max_grad_value = 5.0``. When both
``max_grad_norm`` and ``max_grad_value`` are set, the trainer warns:

  Unsloth: max_grad_norm and max_grad_value are both enabled;
  ignoring max_grad_norm in favor of max_grad_value.

and silently drops the test's ``max_grad_norm=1.0``. +-5.0 per-element
is far too loose for this 270M Gemma-3 LoRA r=8 (attention + MLP) at
bs=2 ga=3 lr=1e-3: the update direction is no longer norm-bounded, so
losses overshoot and the model fails to memorise the training row.

Reproduced on a CUDA mirror (scripts/cuda_mlx_mirror_sim.py):

  norm_1       (max_grad_norm=1.0, no clip): losses 7.64 -> 0.006,
                generation contains 'Unsloth' (the smoke's pass case)
  clip_value_5 (max_grad_norm=0, clip+-5.0): losses 7.29 -> 8.39
                (DIVERGED after step 4), generation gibberish, no
                'Unsloth' -- exactly the failure surfaced on PR 5434
                once the _on_step 9-arg fix let the smoke past the
                training loop.

Pin ``max_grad_value=0.0`` so the smoke uses the same ``max_grad_norm=
1.0`` clipping it was designed against. Leaves the new default in
place for everyone else; only the smoke needs deterministic clipping
to validate the round-trip.

* tests/studio: clarify why MLX smoke pins max_grad_value=0

Refresh the rationale comment to reflect the new default landing in
unslothai/unsloth-zoo#652 (max_grad_value=1.0, not 5.0). The smoke
still needs the explicit pin because neither default value reliably
converges in 7 steps at seed=3407:

  max_grad_value=5.0 -- diverges after step 4 (loss 7.3 -> 8.4)
  max_grad_value=1.0 -- stalls (loss ~3.2 plateau across seeds)
  max_grad_value=0.5/0.25/0.1 -- noisier still
  max_grad_norm=1.0  -- cleanly drops loss to <0.01, emits "Unsloth!"

Mention both the historical 5.0 default and the new 1.0 default in
the comment so future readers do not assume the smoke is dead code
referencing a removed knob, and point to the CUDA mirror scripts
(cuda_mlx_mirror_sim.py + cuda_mlx_clip1_vs_norm1.py) for the
empirical evidence.

No behaviour change; comment-only refresh.

* tests/studio: replace fragile substring gate with loss + round-trip gates

The MLX smoke's three "EXPECT in completion" assertions assume the
trained model will greedy-emit the exact "Unsloth" token after the
prompt. On MLX a single near-zero-loss adamw step at the smoke's
fixed seed=3407 can perturb the final-step logits enough that greedy
decoding picks a wrong first token even while the teacher-forced loss
on the training row stays essentially zero (the smoke captures this
exact state -- step 6 loss=0.049, step 7 grad=36.7, step 7 loss=0.17;
completion goes from "Unsloth!" to "5 lbs!"). Reproduced extensively
on CUDA via scripts/cuda_mlx_step7_*.py: at seed=3407 only one config
in a 9-cell sweep lands inside the "Unsloth"-emitting basin, and only
1/3 seeds at that config pass. This is a property of the assertion,
not of save/reload correctness.

Refactor the three assertions to gate on what the smoke is actually
trying to verify:

  in_memory:
    - hard gate: post_train_loss < 1.0 (training memorised the row).
    - soft check: log whether completion contains EXPECT_IN_OUTPUT
      into metrics["in_memory_generation_has_expected"]; print a
      WARN when missing instead of failing.

  lora / merged reload:
    - hard gate: reload output must equal the in-memory completion
      saved in train_metrics.json. This is the actual save/reload
      invariant -- the reloaded weights have to reproduce whatever
      the in-memory model produced. Falls back to the original
      gibberish gate if train_metrics.json is unavailable.

  gguf reload:
    - hard gate: llama.cpp produced usable, non-empty output after
      the prompt (>=4 chars). llama.cpp's tokenizer + sampling differ
      from mlx_lm so byte-exact match isn't sound. Log
      gguf_has_expected for visibility.

Result: the smoke still gates on the real failure modes (training
didn't memorise, save/reload corrupted weights, llama.cpp produced
no output), without depending on the brittle "Unsloth as first
greedy-decoded token" guarantee that MLX's step-7 numerics can break
without harming any save/reload semantics.

Cross-version constraint: no transformers / trl API touched.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* tests/studio: gate MLX reload on training-row loss, not greedy text

The strict reload assertion (out == in_mem_out) failed on macOS:
in-memory completion was '5 lbs!' and the reloaded completion was
'_________________________'. Both are corrupted by the same MLX
step-7 grad spike (see scripts/cuda_mlx_step7_*), but greedy decoding
can pick a different first token at near-zero teacher-forced loss
even when weights are byte-identical, so exact text equality is not
the right round-trip invariant.

Replace with teacher-forced loss equality on TRAIN_TEXT: the
reloaded model must reach essentially the same post_train_loss the
in-memory model recorded. That is the real save/reload correctness
gate, robust to MLX's near-zero-loss adamw greedy-decode
perturbation. Falls back to a non-empty-body check when
train_metrics.json is missing.

CUDA mirror at this seed converges cleanly to ~0.006 loss; on MLX
post_train_loss < 1.0 still holds via the existing memorisation
gate. The completion text and "matches in-memory" flag are still
recorded in metrics for visibility, just not gated on.

* ci: retrigger Backend CI after transient pwsh-startup timeout

* ci: retrigger MLX dispatch after pytorch CDN DNS flake

* studio: harden FLA + tilelang installers per reviewer feedback

Addresses bot review on #5434:

  * Narrow `_ensure_flash_linear_attention` from `_model_wants_causal_conv1d`
    (which also matches Nemotron-H / Falcon-H1 / Granite-H / LFM2) to
    `_model_wants_tilelang` (Qwen3.5 / Qwen3.6 / Qwen3-Next only). True
    SSM families take the mamba_ssm path and never call FLA's GDN
    kernels, so installing FLA there is wasted bandwidth.

  * Pin both `flash-linear-attention==0.5.0` and `fla-core==0.5.0` and
    install with `--no-deps`. Otherwise pip resolves fla-core's
    declared `torch>=2.7.0` requirement and may silently upgrade the
    Studio venv's torch on environments running torch 2.4/2.5/2.6.

  * Skip both installs on Python <3.10 (FLA, fla-core, and tilelang
    all declare `Requires-Python: >=3.10`). On older interpreters the
    pip install would fail every launch and leave the worker on the
    slow torch fallback while still claiming to have set up the fast
    path.

  * Skip tilelang install on non-Linux platforms. `tilelang==0.1.8`
    only publishes Linux x86_64 / aarch64 and macOS arm64 wheels.
    Falling back to its 93MB sdist on a Studio worker is undesirable.

  * Detect an existing `apache-tvm-ffi` 0.1.10 / 0.1.11 install and
    force a reinstall to 0.1.9 with `--force-reinstall --no-deps`.
    Previously the import-only probe returned early and left the
    broken version in place, which crashes Triton on sm_100.

  * Add a 600s timeout to the tilelang and FLA subprocess.run calls,
    matching the existing flash-attn install pattern, so a network
    hang cannot block the training subprocess indefinitely.

  * 13 new / updated tests covering all six guards plus the
    pinned-spec, timeout, and force-reinstall code paths.

Total: 21 passing tests (8 original + 13 new / updated).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: address reviewer.py P1/P2 findings on FLA + tilelang installers

Twelve-reviewer aggregated review on this PR flagged several real
correctness bugs in the first hardening pass. Fixes:

P1:
  * Add UNSLOTH_STUDIO_SKIP_FLA_INSTALL escape hatch for symmetry
    with UNSLOTH_STUDIO_SKIP_TILELANG_INSTALL and the existing
    UNSLOTH_STUDIO_SKIP_FLASHATTN_INSTALL.
  * Install einops alongside fla-core. `--no-deps` was suppressing
    fla-core's only non-torch runtime dep, so on a clean venv
    `import fla.modules` raised ModuleNotFoundError even though pip
    exited 0.
  * Drop --no-deps from the tilelang force-reinstall path. tilelang
    needs z3-solver, ml-dtypes, cloudpickle, etc. at runtime;
    --force-reinstall --no-deps left libz3.so missing and
    `import tilelang` raised OSError on the next training subprocess.
  * Skip FLA install when installed torch is below 2.7.0
    (fla-core declares torch>=2.7.0). Otherwise users on Studio's
    supported torch 2.4/2.5/2.6 stacks get an incompatible FLA
    installed silently.

P2:
  * Replace bare `except ImportError` probes with helpers that catch
    `Exception` so a broken native package (OSError on missing
    .so, RuntimeError in __init__, ...) does not kill the worker
    before the fallback path can run.
  * Tighten the tilelang platform guard from "any linux" to
    "linux + machine in {x86_64, aarch64, ...}" so ppc64le / s390x /
    armv7 do not fall through and download the 93 MB tilelang sdist.
  * Add --only-binary=:all: to the tilelang install command. The
    comment already said we never want the sdist; now the pip
    invocation enforces it.
  * Verify both FLA and tilelang are importable after pip exits 0;
    if not, report and continue on the fallback path.

6 new tests bring the suite to 27 passing (was 21).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: pin packaging + triton with FLA --no-deps install

An end-to-end install simulation in a fresh venv caught a real
regression: `fla/utils.py` does `from packaging import version` and
`import triton` at module load, but fla-core's METADATA only declares
einops + torch. With `--no-deps` the worker would land FLA in any
runtime that lacks packaging (e.g. minimal torch builds) and the
post-install import probe would fall back to the torch GDN loop
silently.

Add `packaging` and `triton` to `_FLA_RUNTIME_DEPS` so the install
spec list always carries them. Tests updated to assert both are now in
the install command.

* studio: hook transformers' fast-path gates for just-in-time FLA + causal-conv1d install

The substring-based detection in this PR (`_model_wants_tilelang` /
`_model_wants_causal_conv1d`) is brittle: it depends on what the user
typed for the model name, not on what the architecture actually needs.
Users typing custom model paths, future Qwen3.7 / non-Qwen GDN
architectures, and any model whose author renamed it would silently
fall back to the torch loop.

The correct signal is the one transformers itself uses to gate the
fast path. `transformers/models/qwen3_5_moe/modeling_qwen3_5_moe.py`
does at module import time:

    if is_causal_conv1d_available():
        from causal_conv1d import causal_conv1d_fn, causal_conv1d_update
    if is_flash_linear_attention_available():
        from fla.modules import FusedRMSNormGated
        from fla.ops.gated_delta_rule import (
            chunk_gated_delta_rule, fused_recurrent_gated_delta_rule,
        )

Wrap both gates so the first call (always at modeling import, before
any forward pass) installs the matching kernel synchronously and
delegates to the original function. Any model whose architecture
queries those gates auto-triggers the install; models that never
query them (Llama, Gemma, dense Qwen, ...) never pay the cost.

Mechanics:

  - Split `_ensure_flash_linear_attention` and `_ensure_tilelang_backend`
    into `_unconditional` variants (no substring gate, retains python
    / torch / platform / skip-env guards) plus thin substring wrappers
    used by the legacy fallback path.
  - New `_install_fast_path_hooks(event_queue)` patches both gates on
    `transformers.utils.import_utils` AND sweeps `sys.modules` so any
    modeling file that already did `from ... import is_X` sees the
    wrapper (the local binding survives a module-level reassignment).
  - Wrappers clear the original's `lru_cache` before delegating, install
    on False, re-check, and short-circuit on subsequent calls.
  - Set `UNSLOTH_STUDIO_SKIP_FAST_PATH_HOOKS=1` to fall back to the
    substring path.

Verified end-to-end against `transformers.models.qwen3_5_moe`:

  PRE_STATE fla=False tilelang=False causal_conv1d=False
  HOOK_INSTALLED
  Hook fired for is_causal_conv1d_available; installing kernel...
  Installing prebuilt causal-conv1d wheel...
  Hook fired for is_flash_linear_attention_available; installing kernel...
  Installing flash-linear-attention==0.5.0 (with fla-core==0.5.0) for the fast path...
  Installed flash-linear-attention for the FLA fast path
  Installing TileLang backend (apache-tvm-ffi==0.1.9, tilelang==0.1.8)...
  Installed TileLang backend for FLA fast path
  MODELING_IMPORT_OK
  FAST_PATH_SYMBOLS {"chunk_gated_delta_rule": true,
                     "fused_recurrent_gated_delta_rule": true,
                     "FusedRMSNormGated": true,
                     "causal_conv1d_fn": true,
                     "causal_conv1d_update": true}
  POST_STATE fla=True tilelang=True causal_conv1d=True

Adds 9 new tests covering: install-on-False, skip-on-True, idempotency,
install-failure handling, env-disable, lru_cache clear, sys.modules
rebind, missing-transformers fallback, substring fallback. Total
test count is now 36 (was 27).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: address reviewer.py n=12 findings on the FLA hook path

Eight issues reproduced by parallel reviewers against 6ce495a; all
fixed and covered by regression tests. 45 pytest cases pass (was 36);
end-to-end Qwen3.5_MoE modeling-import drill still loads all five
fast-path symbols.

P1 fixes:

1. TileLang loses the Qwen-family guard on the normal FLA hook path
   (10/12 reviewers, reproduced with allenai/OLMo-Hybrid-1B). The
   hook unconditionally installed tilelang for any FLA-using model.
   - Threaded `model_name` through `_install_fast_path_hooks(event_queue,
     model_name)`.
   - `_fla_install` now gates tilelang on
     `_model_wants_tilelang(model_name)` AND a successful FLA install.

2. TileLang repair `--force-reinstall` (without `--no-deps`) could
   replace `torch==2.12.0+cu130` with `torch==2.12.0`. Split repair
   into TWO steps:
     step 1: `--force-reinstall --no-deps apache-tvm-ffi==0.1.9`
     step 2: regular install of tilelang + apache-tvm-ffi
   Step 1 surgically downgrades the broken package; step 2 resolves
   missing transitive deps (z3-solver, ml-dtypes) without
   --force-reinstall, so it never replaces torch.

3. Hook could return True after the installer's deep import probe
   failed: when pip exits 0 but `import fla.modules` raises, the old
   wrapper re-called `original()` (transformers' metadata check) and
   trusted it. Refactored:
     - `_ensure_flash_linear_attention_unconditional(...) -> bool`
     - `_ensure_tilelang_backend_unconditional(...) -> bool`
   The wrapper now uses the installer's bool directly.

4. SSM models (Nemotron-H, Falcon-H1, Granite-H) use
   `lazy_load_kernel("causal-conv1d")` and never call
   `is_causal_conv1d_available()`, so the hook never fires for them.
   The orchestrator now always runs `_ensure_causal_conv1d_fast_path`
   outside the hook-mode if/else.

P2 fixes:

5. `_rebind_in_already_imported_modules` invoked transformers' lazy
   module `__getattr__` (hundreds of "Accessing X from .models..."
   warnings, ~3.4s overhead). Switched to `module.__dict__.get(...)`
   which only sees real module-level bindings.

6. TileLang installed even when FLA was skipped (Torch <2.7) or
   failed (timeout, post-install probe failed). Now gated on the
   installer's bool return.

7. TileLang repair was skipped when FLA was already True but tilelang
   missing or apache-tvm-ffi on the broken list. Added an optional
   `post_available_fn` to the wrapper; the FLA hook's
   `_fla_post_available` runs `_ensure_tilelang_backend_unconditional`
   when (model wants tilelang) AND (tilelang missing OR tvm-ffi broken).

8. `_flash_linear_attention_importable()` only checks deep import,
   not version. Added `_flash_linear_attention_current()` that
   compares against the pinned `flash-linear-attention==0.5.0` /
   `fla-core==0.5.0`; older versions trigger `--force-reinstall
   --no-deps` so torch stays untouched.

Helpers extracted to keep the surface tight:
  - `_pip_install_cmd(*args)` builds `uv pip install` or
    `python -m pip install` depending on uv availability.
  - `_run_pip(cmd, event_queue, label)` runs a pip command with
    timeout / failure handling and a status emission.

Regression tests added:

  - test_hook_does_not_install_tilelang_for_non_qwen_fla_model
  - test_hook_does_install_tilelang_for_qwen35
  - test_tilelang_repair_does_not_touch_torch_cuda_stack
  - test_hook_trusts_installer_bool_not_metadata
  - test_rebind_does_not_trigger_module_getattr
  - test_hook_skips_tilelang_when_fla_install_is_skipped
  - test_hook_runs_tilelang_repair_when_fla_already_true
  - test_fla_installer_force_reinstalls_when_older_version_present
  - test_run_training_process_eagerly_installs_causal_conv1d_in_normal_mode

Existing tests updated for the new `_install_fast_path_hooks` signature
and the two-step tilelang repair flow.

End-to-end re-verified against transformers.models.qwen3_5_moe:
PRE_STATE fla=False, hook fires for both gates, FLA + tilelang +
causal-conv1d install, all 5 fast-path symbols non-None.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: fix double-install of tilelang on the FLA hook install path

Backend CI surfaced a test-isolation bug introduced by the
post_available_fn mechanism for finding #7. The wrapper ran
`post_available_fn` in BOTH paths (install ran AND gate already True),
but `_fla_install` already chains tilelang on the install path, so the
post-available step then called tilelang install AGAIN.

This was masked locally because tilelang was installed in the
workspace venv (post_available short-circuited on
`_tilelang_importable()` returning True). CI starts with no tilelang,
so the second call actually fired and the mock recorded two calls.

Fix: only run `post_available_fn` when the install path did NOT run.
That preserves the finding #7 semantics (tilelang repair when FLA
already True but tilelang missing or tvm-ffi broken) without
duplicating the chained install on the gate-was-False path.

Also tightened `test_hook_skips_install_when_gate_already_true` to
monkeypatch `_tilelang_importable=True` and
`_installed_tvm_ffi_version=0.1.9` so it stays a pure "no install at
all" test regardless of the venv's actual state.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* ci: retrigger Mac Studio GGUF after transient HF DNS resolve flake

* studio: skip tilelang on HIP / ROCm torch (Strix Halo crash report)

h34v3nzc0dex tested PR 5434 on Strix Halo (gfx1151, ROCm 7.13,
torch 2.11.0+rocm7.13.0) and hit a hard regression:

  File ".../fla/ops/common/backends/tilelang/__init__.py", line 92,
    in chunk_bwd_dqkwg
  File ".../tilelang/jit/kernel.py", line 137, in __init__
  File ".../tilelang/tileop/gemm/__init__.py", line 143,
    in _select_gemm_instruction
  tvm.error.InternalError: Check failed: (0) is false:
    Unsupported target for gemm:
    hip -keys=hip,gpu -mcpu=gfx1151 ...

`tilelang==0.1.8` ships no HIP GEMM instruction; `_select_gemm_instruction`
raises at lower-time, not import-time. So:
  - pip install succeeds
  - `import tilelang` succeeds
  - `TileLangBackend.is_available()` returns True
  - FLA's dispatcher picks TileLang for `chunk_bwd_dqkwg`
  - training subprocess dies at first GDN backward, no graceful fallback

The PR's existing platform gate (`_tilelang_platform_supported`)
checked only `sys.platform == "linux"` and `platform.machine()`, both
of which look identical on a ROCm box.

Fix has two layers:

1. INSTALL GATE: new `_torch_has_hip()` helper checks
   `torch.version.hip is not None`. `_tilelang_platform_supported`
   now returns False on HIP torch, so the install never fires.

2. RUNTIME GATE: even with the install skipped, a user could have
   tilelang already present (e.g. venv carried over from a CUDA box).
   `_install_fast_path_hooks` now calls
   `os.environ.setdefault("FLA_TILELANG", "0")` when HIP is detected,
   which is the env-var FLA's `TileLangBackend` already honors. Users
   who know they have a HIP-aware tilelang fork can override by
   setting `FLA_TILELANG=1` explicitly.

This costs nothing on CUDA (the gate is a no-op when
`torch.version.hip is None`), and removes the crash for AMD users.
The benchmark numbers in the PR description (1.43x on B200 sm_100)
are not affected.

The other halves of the PR are confirmed working on gfx1151 by the
same report:
  - `flash-linear-attention 0.5.0` runs at production scale
    (B=1 T=8192 H=16 K=128 V=128 and others) with no patches.
  - `causal-conv1d` runs at the shapes the fast-path gate cares
    about. (A separate Ubuntu 24.04 `--gcc-install-dir` build
    workaround is needed for the source-build path; that mirrors
    bbf004c's llama.cpp fix and is out of scope here.)

Tests added:
  - test_tilelang_platform_unsupported_on_hip_torch
  - test_tilelang_install_skipped_on_hip_torch
  - test_install_fast_path_hooks_sets_fla_tilelang_zero_on_hip
  - test_install_fast_path_hooks_respects_user_fla_tilelang_override
  - test_install_fast_path_hooks_does_not_set_fla_tilelang_on_cuda

Total 50 passing (was 45).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* ci: retrigger Windows Studio UI after transient Playwright tab-lookup flake

* studio: auto-discover FLA-using model types from installed transformers

Drop the hand-maintained `_TILELANG_MODEL_SUBSTRINGS` tuple
(qwen3.5 / qwen3_5 / qwen3.6 / qwen3_6 / qwen3-next / qwen3_next)
and derive the allowlist by scanning the installed
`transformers/models/*/modeling_*.py` for `from fla.` imports.

A model "wants tilelang" iff its modeling file imports an FLA op,
which is the same signal `is_flash_linear_attention_available()` is
the runtime test for. The scan happens once per worker subprocess
and is cached for the process lifetime; an empty result (eg
transformers not importable) means "no tilelang pre-install" --
the FLA runtime hook still drives the install via the gate when
the loaded model actually probes it.

Verified against the live installed transformers, the auto-derived
set is {qwen3_5, qwen3_5_moe, qwen3_next}, with `_model_wants_tilelang`
matching the HF Hub names `unsloth/Qwen3.5-2B`, `Qwen/Qwen3.5-MoE-A3B`,
`mlx-community/qwen3-next-80b`, and correctly rejecting Llama,
Mistral, Nemotron-H, Falcon-H1, etc. Future GDN models (Qwen3.7,
OLMo-Hybrid-FA, ...) are picked up automatically once they ship in
transformers; no further worker edits needed.

Also trim docstrings / comments through the FLA / tilelang / HIP /
hook block: constants get 1-line trailing comments, function
docstrings collapse to 1-3 lines, and the fast-path-hooks banner
shrinks from a 27-line block to 4 lines. The file drops from 2847
to 2630 lines without losing the load-bearing WHY notes
(--no-deps protects torch; `__dict__.get` avoids lazy-module
__getattr__; two-step tvm-ffi repair keeps torch off the dep
graph; HIP setdefault disables FLA's TileLang dispatch even with
tilelang already installed).

7 new tests (50 -> 57 total): discovery returns only FLA-using
model_types; discovery cache reuse; missing transformers handled;
OSError on a modeling file is non-fatal; `_model_wants_tilelang`
matches real HF repo names across separator variants; empty
discovery -> always False; normalization across `-`, `.`, `/`,
space.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* test: hermetize the non-allowlist hook test against transformers 5.4.0+

transformers 5.4.0 added `olmo_hybrid` as an FLA-using model_type, so
the auto-discovered allowlist now includes it -- and the test's prior
choice of `allenai/OLMo-Hybrid-1B` as a "non-Qwen FLA-only" example
became an allowlist member. CI on Python 3.11 / 3.13 caught this.

Swap to a guaranteed-not-in-allowlist fake model_name AND patch
_discover_fla_model_types to a known {qwen3_5, qwen3_5_moe, qwen3_next}
set so the test stays valid as upstream transformers adds new
FLA-using architectures.

Renames the test to reflect the actual semantic under test:
"outside-allowlist -> no tilelang".

* ci: retrigger Windows Studio API after llama.cpp prebuilt staging WinError 5 flake

* tests: move MLX smoke gate changes to dedicated PR #5537

The seven MLX smoke commits in this PR's history (_on_step grad_norm,
max_grad_value pin, loss + round-trip gates) are unrelated to the
FLA / tilelang work. They now live in #5537 so this PR's diff is
limited to the studio worker installer changes.

Net effect on tests/studio/run_real_mlx_smoke.py vs main: zero.

* studio: friendlier install banners (drop hook / gate-name jargon)

User-visible status text now reads:
  Installing flash-linear-attention==<ver> for faster training...
  Installing TileLang==<ver> for faster training...
  Installing causal-conv1d for faster training...
  Installing flash-attn for faster training...

Removed the transient "Hook fired for is_flash_linear_attention_available;
installing kernel..." banner — the install banner that immediately follows
already tells the user what is happening, in plain English.

The internal logger.info messages (server-side log) still carry the
gate names + "Hook fired ..." for debugging.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-18 03:49:06 -07:00
Etherll
11e5b58827
studio: gate image input on effective vision capability (#5492)
* Studio: gate image input on a usable mmproj for GGUF vision models

* Improve image gating and model capability sync

Tighten image-handling and model capability syncing across the chat flow. Key changes:

- chat-adapter: Replace per-message current-user image check with a simpler gate that blocks if ANY image is present in the outbound payload when the selected model cannot handle vision. Show the toast reason and flip the per-thread running flag on→off to avoid hanging wait promises before throwing.

- shared-composer: Simplify and correct image-attachment gating for single vs compare modes. Use an attach-time gate that defers to send/ensureModelLoaded in compare mode, introduce attachUnavailableReason, and only block immediately for single-mode. Remove an unused models selector.

- shared-composer: Sync the runtime models[] entry with the response from ensureModelLoaded so UI/send gates read fresh capabilities (isVision, isGguf, isAudio, audioType, hasAudioInput). This addresses catalog lag (e.g., GGUF mmproj arriving after the catalog snapshot).

- UX tweak: the file-picker button no longer outright blocks on image availability; addFiles still filters images per-file and toasts appropriately.

These changes prevent mid-stream server rejections, avoid deadlocks, and ensure model capability checks are accurate when attaching images or audio.

* studio: only pass --mmproj to llama-server when effective_is_vision

When a text-only GGUF (static is_vision=False) was paired with a
family-matching mmproj path, the launcher appended both --mmproj and
--spec-default, leaving llama-server in an inconsistent state while
Studio reported is_vision=False. Gate the --mmproj flag on
effective_is_vision so the launch command tracks the runtime
capability the rest of Studio sees.

* studio: reject image content in streaming /v1/responses for non-vision GGUF

_responses_stream forwards the OpenAI request body directly to
llama-server's /v1/chat/completions, bypassing the image-vs-vision
guard that openai_chat_completions enforces for the wrapped path.
Add the same check at the top of the streaming entry point so an
SDK client that posts an image to a non-vision GGUF receives a
typed 400 instead of an opaque downstream error.

* studio: gate external chat providers in the image input helper

External selections (cohere, deepseek, mistral, openrouter, ...) live
in externalProviders, not in runtime.models[], so activeModel is
undefined for them and the helper short-circuited to allow. Result:
images attached to a non-vision external chat model were dropped
silently downstream instead of rejected up front.

Add providerTypeSupportsVision to external-providers.ts (false for
known text-only providers, true for known vision-capable ones, null
for unknown / custom self-hosted) and thread externalSupportsVision
+ externalModelLabel through the helper. shared-composer.tsx,
runtime-provider.tsx (VisionImageAdapter.add), and chat-adapter.ts
pre-stream gate all resolve the provider type and pass it.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: Roland Tannous <115670425+rolandtannous@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-18 01:35:46 -07:00
Paul Durkin
388ade4c84
fix(studio/worker): inject --gcc-install-dir for HIP source builds on Ubuntu 24.04 (#5517)
* fix(studio/worker): inject --gcc-install-dir for HIP source builds on Ubuntu 24.04

On Ubuntu 24.04 + ROCm clang-20, the HIP source-build fallback in
`_install_package_wheel_first` (causal-conv1d, mamba-ssm source fallback,
flash-attn source fallback) dies at:

  /opt/rocm-X.Y/lib/llvm/lib/clang/20/include/__clang_hip_runtime_wrapper.h:112:10:
    fatal error: 'cstdlib' file not found

Root cause: clang-20 picks the highest-numbered /usr/lib/gcc/x86_64-linux-gnu/<N>
runtime dir by default. On 24.04 that's gcc-14, whose runtime objects ship in
the gcc-14 package but whose C++ headers (/usr/include/c++/14) come from
libstdc++-14-dev — NOT in the default apt set. libstdc++-13-dev IS in the
default set, so /usr/include/c++/13 exists. clang has no way to discover
that asymmetry and the build fails.

Fix: new `_hipcc_gcc_install_dir()` helper iterates gcc 14 → 11 and returns
the first /usr/lib/gcc/x86_64-linux-gnu/<N> dir where BOTH the runtime AND
/usr/include/c++/<N> exist. The HIP branch of `_install_package_wheel_first`
appends `--gcc-install-dir=<that path>` to HIPCC_COMPILE_FLAGS_APPEND before
invoking pip. Respects an existing `--gcc-install-dir` in the env var
(user-set takes precedence); preserves any other flags the user has set
(appends to the end rather than overwriting). No-op on non-HIP, non-Linux,
non-x86_64.

Mirrors the same fix bbf004c added to studio/setup.sh for the llama.cpp HIP
build branch (#5301), but via env var since pip-driven source builds can't
take CMake flags directly.

Verified on Ryzen AI MAX+ 395 / Radeon 8060S (gfx1151) / Ubuntu 24.04 /
ROCm 7.13 nightly: `_hipcc_gcc_install_dir()` returns
`/usr/lib/gcc/x86_64-linux-gnu/13`, which matches the manual workaround
that already lets `pip install causal-conv1d` succeed on this hardware.

Tests added (8 new in test_training_worker_flash_attn.py):
- test_hipcc_gcc_install_dir_picks_highest_with_headers
- test_hipcc_gcc_install_dir_picks_14_when_headers_exist
- test_hipcc_gcc_install_dir_returns_none_when_no_match
- test_hipcc_gcc_install_dir_returns_none_on_non_linux
- test_hipcc_gcc_install_dir_returns_none_on_non_x86_64
- test_install_injects_gcc_install_dir_on_hip_source_build
- test_install_appends_to_existing_hipcc_compile_flags
- test_install_respects_user_gcc_install_dir
- test_install_does_not_inject_env_on_cuda

Per @danielhanchen's suggestion in
https://github.com/unslothai/unsloth/pull/5434#issuecomment-4469980122

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* review: apply gemini-code-assist suggestion on _run_kwargs env handling

Use _run_kwargs.get("env", os.environ).copy() + key-mutation instead of
rebuilding env from os.environ directly. Today both forms are equivalent
(no earlier code in _install_package_wheel_first sets _run_kwargs["env"]),
but the .get().copy() pattern survives any future env modification added
upstream of this block without silently throwing it away.

No behavioural change; tests already assert the final HIPCC_COMPILE_FLAGS_APPEND
value, not the env-construction pattern.

Per https://github.com/unslothai/unsloth/pull/5517#discussion_r... (gemini-code-assist[bot])

---------

Co-authored-by: h34v3nzc0dex <h34v3nzc0dex@users.noreply.github.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2026-05-18 01:05:30 -07:00
Daniel Han
3876c87034
studio: extend offline DNS auto-detect to inference parent + training (#5512)
* studio: extend offline DNS auto-detect to inference parent + training

#5505 fixed the GGUF/llama-server load path. Studio still has two
adjacent code paths that burn ~30-60s of soft-failed timeouts before
the worker subprocess starts when DNS to huggingface.co is dead and
the model is already in the local HF cache.

Inference parent process (routes/inference.py:load_model):

* ModelConfig.from_identifier now runs inside _hf_offline_if_dns_dead
  so the LoRA-detect hf_model_info call and the urllib config probes
  in utils/transformers_version.py short-circuit when DNS is dead.
* utils/models/model_config.py: extracted the inline HF_HUB_OFFLINE/
  TRANSFORMERS_OFFLINE check used by list_gguf_variants and
  detect_gguf_model_remote into a shared _env_offline() helper, then
  reused it to gate the LoRA-detect hf_model_info call.
* utils/transformers_version.py: _check_tokenizer_config_needs_v5 and
  _check_config_needs_550 now early-return False when offline instead
  of issuing a 10s urllib.urlopen against huggingface.co/raw/main.

Training worker (core/training/worker.py:run_training_process):

* Add the same 2s DNS probe used by core/inference/worker.py at the
  top of the training subprocess. On failure, set HF_HUB_OFFLINE,
  TRANSFORMERS_OFFLINE, and HF_DATASETS_OFFLINE before the rest of
  the subprocess imports torch/transformers/unsloth, so every
  from_pretrained, snapshot_download, and load_dataset call below
  resolves from cache. Scope is per-subprocess; the orchestrator
  always spawns a fresh worker per training run.

Training trainer (core/training/trainer.py:load_model):

* Skip the proactive hf_model_info gated-repo probe when _env_offline()
  is true. The API is unreachable anyway, and a gated model that is
  already cached is exactly the scenario the user is trying to train
  against. from_pretrained surfaces the real error if access is
  actually denied.

Tests (tests/test_offline_inference_parent.py, 7 new cases):

* _env_offline truthy/falsy parsing across HF_HUB_OFFLINE and
  TRANSFORMERS_OFFLINE.
* transformers_version urllib short-circuit when offline.
* LoRA detect hf_model_info skip when offline.

Existing tests/test_offline_gguf_cache_fallback.py still passes
(26 cases) because the inline env check was extracted, not changed.

* tests: prefer real httpx over stub in offline-test files

The studio test stub convention only included the 6 httpx exception
names that existed callers needed. Newer huggingface_hub (1.15+)
imports HTTPError, Response, Request, HTTPStatusError, AsyncClient,
and more at module import time. When httpx is truly absent the stub
chase becomes a treadmill.

Use the real package when installed (the CI install list already
includes httpx, so this is the production environment). Fall back to
the stub only when httpx is genuinely missing.

No code under test changes.

* studio: detect cached LoRA adapters offline; tighten test

Two follow-ups from the review pass on #5512:

* ModelConfig.from_identifier no longer skips the remote LoRA-detect
  hf_model_info call when _env_offline() is true. huggingface_hub
  short-circuits the call via OfflineModeIsEnabled in ~0ms when
  HF_HUB_OFFLINE is set, so the original 25s concern was moot once
  routes/inference.py wrapped the call in _hf_offline_if_dns_dead.
  Skipping the API meant users with a cached LoRA adapter
  (adapter_config.json on disk) got is_lora=False and the load
  failed. After the API call (which raises fast offline) a new
  cache-fallback walks the HF cache snapshot for adapter_config.json
  via the existing _iter_hf_cache_snapshots helper.

* test_hf_model_info_not_called_when_offline replaced. The old test
  raised AssertionError inside production code that catches Exception,
  so it passed even if the call happened. New tests use MagicMock and
  assert call_count >= 1, plus a fixture that stages a fake HF cache
  with adapter_config.json to verify the offline cache detection.

Test count goes from 7 to 8 in test_offline_inference_parent.py.
Combined with test_offline_gguf_cache_fallback.py: 34 pass in 9.75s.

* Fix/adjust offline training DNS probe per PR #5505 review

Same fix as #5505's _probe_dns_dead refactor: run gethostbyname on a
daemon thread with join timeout so concurrent sockets in the parent
interpreter never inherit a process-wide socket.setdefaulttimeout
mutation. Adds a static-pin regression test that the inference parent
file does not regress on this.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Trim verbose code comments per review feedback

Shorten the longer explanatory comments added by this PR while keeping
the WHY of each non-obvious branch:

- trainer.py: collapse the 5-line proactive gated-check comment.
- training/worker.py: trim the offline auto-detect preamble and the
  "logger isn't configured" note.
- routes/inference.py: shorten the DNS-probe wrap rationale.
- transformers_version.py: collapse the two urllib short-circuit notes.
- model_config.py: shorten the LoRA detect + cache-fallback notes.
- tests/test_offline_inference_parent.py: tighter module docstring,
  trim class docstrings, drop multi-line explainer comments inside the
  tests; behaviour and coverage unchanged (9/9 tests still pass).

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-18 00:31:33 -07:00
Daniel Han
c690b28e99
Studio: warn when llama.cpp prebuilt is at least 3 days behind (#5529)
* Studio: warn when llama.cpp prebuilt is at least 3 days behind

Layered on #5528. Generalises the MTP-specific staleness warning to
every llama.cpp prebuilt update, not just the ones that add MTP. If
the installed prebuilt is at least 3 days old AND its tag differs
from the latest published tag on the helper release repo (default
unslothai/llama.cpp), Studio nudges the user to run
"unsloth studio update".

How it works

Reads the install marker UNSLOTH_PREBUILT_INFO.json that
install_llama_prebuilt.py already writes to install_dir. The marker
carries the installed tag, the helper repo, and an installed_at_utc
timestamp. Studio compares those against the latest published tag
from the GitHub releases API for the helper repo.

GitHub fetch is cached at two levels:
- Process-level memo for /status hot path.
- Disk-level cache (24h TTL) at ~/.unsloth/studio/cache/llama_cpp_freshness/
  so cold-start Studio launches do not always hit the API.

On a transient fetch failure (offline, rate-limited) we keep the
last-good disk value alive rather than poisoning the cache with None.
The check fails open: if anything is missing (marker, timestamp,
GitHub response), stale stays False so users never see a misleading
banner.

Surfaced in two places

1. Startup banner (logs + stderr) in main.py:lifespan(), alongside the
   MTP capability probe added in #5528. Single line, e.g.:
     WARNING: llama.cpp prebuilt is 5 days behind: installed b9190,
     latest b9300. Run "unsloth studio update" to refresh.

2. /api/inference/status now returns:
     llama_cpp_prebuilt_stale: bool
     llama_cpp_installed_tag:  str | None
     llama_cpp_latest_tag:     str | None
   so the frontend can render a banner / popup with the actual tag
   delta the user is missing.

3-day threshold

Mirrors the typical Unsloth llama.cpp release cadence. Anything
shorter would nag users who restart Studio at the wrong moment;
longer leaves real bugs sitting on the user's machine. Configurable
via the threshold_days kwarg if a future call site wants a different
window.

Tests

17 new cases in tests/test_llama_cpp_freshness.py cover marker
discovery in both cmake and root install layouts, missing / invalid
marker, GitHub fetch caching across process restarts (disk cache hit
after the in-memory cache is reset), the stale / not-stale decision
matrix (tag mismatch + age threshold), fail-open behaviour when
GitHub is unreachable, custom threshold, singular/plural day in the
warning string, and unparseable installed_at_utc. The broader
205-test inference regression suite still passes.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-18 00:21:50 -07:00
Daniel Han
fc04809bfe
Studio: warn when llama.cpp prebuilt is too old for MTP (#5528)
* Studio: warn when llama.cpp prebuilt is too old for MTP

Layered on #5527. Adds a one-shot llama-server --help capability probe
so users get a clear signal when their prebuilt is missing MTP support,
plus a graceful fallback if they load an MTP GGUF against an outdated
binary.

What's surfaced:

1. Startup log + stderr line in main.py:lifespan() if MTP isn't
   advertised:
     WARNING: llama.cpp prebuilt is missing MTP support
     (--spec-type mtp / draft-mtp). Run `unsloth studio update` to
     refresh it. MTP GGUFs will load without speculative decoding.
2. Load-time graceful fallback in load_model's spec block: skip the
   auto-emit and log a clear warning instead of letting llama-server
   fail with an unknown-flag error.
3. /api/inference/status now returns llama_cpp_supports_mtp: bool so
   the frontend can show a banner / popup.

Probe internals:

- Class-level cache keyed on (binary_path, mtime). One subprocess call
  the first time, instant thereafter. Touching the binary (e.g. via
  `unsloth studio update`) invalidates the cache automatically because
  the mtime changes, so the new build is picked up without restarting
  the server.
- Recognises both upstream naming forms: the original draft-mtp from
  llama.cpp PR #22673 and the renamed mtp variant in later commits.
- Spec block uses whichever token the binary accepts so we emit the
  right value regardless of which release the user has.

Tests:

- 6 new cases in test_llama_cpp_mtp_detection.py covering each probe
  variant (draft-mtp, renamed mtp, pre-MTP build, missing binary,
  mtime-based cache invalidation).
- Existing 38 MTP detection cases still pass; broader 188-test
  regression suite (server args, reload inheritance, gguf metadata,
  load progress, context fit, model validation) still green.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-18 00:19:47 -07:00
Daniel Han
5a6e94d422
Studio: auto-enable MTP speculative decoding for MTP GGUFs (#5527)
* Studio: auto-enable MTP speculative decoding for MTP GGUFs

Detect Unsloth's MTP (multi-token-prediction) GGUFs and auto-emit the
right --spec-type draft-mtp flags for llama-server (llama.cpp PR
#22673), so users get the speedup without configuration.

Detection prefers the GGUF metadata field <arch>.nextn_predict_layers
(verified on Qwen3.6-27B-MTP-GGUF / qwen35 and Qwen3.6-35B-A3B-MTP-GGUF
/ qwen35moe). Falls back to a -MTP marker in the identifier / filename
so HF-mode loads can detect MTP from the repo name before the GGUF is
downloaded.

Flag presets follow the Unsloth MTP guide:
  GPU:     --spec-type draft-mtp --spec-draft-n-max 6
  CPU/Mac: --spec-type draft-mtp --spec-draft-n-max 3 \
           --spec-type ngram-mod --spec-ngram-mod-n-match 24 \
           --spec-ngram-mod-n-min 48 --spec-ngram-mod-n-max 6

User overrides win: if the caller passes --spec-type / --spec-default
via unsloth run / unsloth studio run pass-through (or HTTP
llama_extra_args), the auto-emit steps aside so llama-server only sees
the user's flag. Scalar tuning knobs like --spec-draft-n-max compose
with the auto preset via llama-server's last-wins parsing.

_already_in_target_state mirrors the same promotion so a repeat /load
with unchanged settings against an MTP backend running draft-mtp
short-circuits cleanly instead of forcing a reload.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-18 00:15:42 -07:00
Daniel Han
a09e70e8be
tests/studio: lock in Windows GPU detection fix (#5106) with a synthetic CI test (#5376)
* tests/studio: end-to-end Windows GPU detection mock test (#5106)

Locks in the combined fix from #5322 + #5324 with a synthetic
Windows scenario that CI runners without GPUs can execute. The
test packs the real PyPI win_amd64 wheel layouts (cu12 modular and
the new unsuffixed cu13 nvidia/cu13/bin/x86_64 layout) plus the
exact filename set of the upstream b9103 cudart-llama-bin-win-cuda
bundles, then mocks nvidia-smi output and asserts that:

 * Studio's nvidia-smi probe parses the CSV and reports the GPU.
 * After PR #5322 the install_dir/build/bin/Release/ tree contains
   all three cudart bundle DLLs alongside llama-server.exe.
 * After PR #5324 the PATH built by start_llama_server's win32
   branch lists pip nvidia + torch/lib dirs in addition to the
   binary_dir.
 * cudart64_X.dll, cublas64_X.dll, and cublasLt64_X.dll are
   each reachable from at least one PATH entry, with cudart
   specifically reachable from BOTH the install dir and a pip
   nvidia dir (defence in depth).
 * Bare venvs without pip nvidia wheels still work via #5322's
   binary_dir drop; pre-#5322 installs still work via #5324's
   PATH augmentation.
 * A reconstructed pre-PR scenario (cudart absent from binary_dir
   and pip dirs not on PATH) leaves cudart unreachable, confirming
   the test would catch a future regression.

Bonus housekeeping in studio/install_llama_prebuilt.py: drop the
pointless f-prefix on the literal "llama-" in the
windows_cuda_attempts pairing guard (no behaviour change; lint
nit flagged in the post-merge review).

The mocks model real artifact contents I verified empirically:
 * pip download nvidia-cuda-runtime --platform win_amd64
   produces nvidia/cu13/bin/x86_64/cudart64_13.dll.
 * unzip on the b9103 cudart-llama-bin-win-cuda-13.1-x64.zip
   produces exactly cudart64_13.dll + cublas64_13.dll +
   cublasLt64_13.dll, no executables.
 * objdump -p on the b9103 ggml-cuda.dll shows a static PE
   import on cublas64_13.dll (the root cause of #5106 when
   cublas64_13.dll is unreachable).

Refs #5106 #5322 #5324

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* test_5106_windows_gpu_detection_mock: don't shadow real httpx

This file's name sorts before every other file in studio/backend/tests/
(starts with the digit '5'), so pytest collects it first. The previous
``sys.modules.setdefault("httpx", _httpx_stub)`` ran before any other
test imported real httpx, which meant the stub permanently shadowed
the real module for the rest of the collection. Tests that did
``from httpx import HTTPError, Response`` (test_anthropic_messages,
test_browse_folders_route, test_training_*, etc) then failed at
collection with ``ImportError: cannot import name 'HTTPError'``
because the stub did not define those names. The existing
test_llama_cpp_windows_nvidia_path.py did not trigger the same issue
because it sorts after test_a* / test_b* / etc, by which point the
real httpx has already been imported and setdefault is a no-op.

Switch the stub installation to ``importlib.util.find_spec(name) is
None`` so we only fall back to the stub when the real module truly is
not installed. Backend CI installs httpx, structlog, and the
studio/backend/loggers package is reachable via the sys.path
augmentation a few lines above, so on CI all three find_spec calls
succeed and no stubs are installed at all.

Also add HTTPError and Response to the stub module for the offline
case, so anyone running this test outside CI with httpx absent still
gets a stub that satisfies the broader test suite's imports.

Refs #5106

* test_5106 + llama_cpp: extract win32 PATH helper and harden the regression test

Follow-up to PR #5376's review feedback. Three real findings from the
bot reviewers, plus one stale one.

1. (codex P2 line 201, gemini medium line 209) The regression test's
   _build_path_dirs_like_start_llama_server hand-copied the win32
   branch of LlamaCppBackend.start_llama_server, so a future drop or
   reorder of _windows_pip_nvidia_dll_dirs(sys.prefix) in production
   would have passed the test silently.

   Extract a new staticmethod LlamaCppBackend._build_windows_path_dirs
   (binary_dir, prefix, cuda_path). Production start_llama_server now
   calls this helper. The test's wrapper is reduced to a one-line
   delegate that forwards to the staticmethod, so the regression
   asserts against the exact production logic instead of a parallel
   copy of it.

2. (codex P2 line 245) test_nvidia_smi_probe_reports_synthetic_gpu did
   not clear CUDA_VISIBLE_DEVICES. On a shared GPU runner with the
   variable set in the parent shell, _get_gpu_free_memory() filters
   the mocked CSV and returns [] or falls through to the torch
   fallback. Cleared CUDA_VISIBLE_DEVICES and NVIDIA_VISIBLE_DEVICES
   via monkeypatch.delenv(..., raising=False).

3. (codex P2 line 66) _maybe_stub gated on importlib.util.find_spec
   ("loggers"), which returns a spec because studio/backend/loggers/
   is on sys.path. But the actual import chain loads
   loggers/handlers.py which does `from fastapi import Request,
   Response` at module load. In a lightweight env without fastapi
   installed, the stub never lands and `from core.inference.llama_cpp
   import LlamaCppBackend` raises during collection. Switched
   _maybe_stub to a real import attempt under try / except ImportError
   so the stub falls into place when the package is discoverable but
   not importable. CI has fastapi so this is purely a developer-
   machine ergonomics fix.

The fourth comment (codex P1 line 85 "Keep the httpx stub from leaking
across tests") was already addressed by 7437e735, which replaced the
unconditional sys.modules.setdefault with the find_spec-gated
_maybe_stub. No code change needed.

Production behaviour is unchanged: _build_windows_path_dirs returns
exactly the same ordering start_llama_server used inline
([binary_dir, *pip_dirs, cuda_bin?, cuda_bin_x64?]).

Verification (run inside studio/backend):
  pytest tests/test_5106_windows_gpu_detection_mock.py -v
    -> 10 passed
  pytest tests/test_llama_cpp_*.py tests/test_llama_server_args.py
       tests/test_5106_windows_gpu_detection_mock.py -q
    -> 171 passed
  CUDA_VISIBLE_DEVICES=1 pytest tests/test_5106_windows_gpu_detection_mock.py::TestWindowsGpuDetectionAfter5106Fix::test_nvidia_smi_probe_reports_synthetic_gpu
    -> 1 passed

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Rename Windows GPU detection test to a generic filename and trim comments

- studio/backend/tests/test_5106_windows_gpu_detection_mock.py
  -> studio/backend/tests/test_windows_gpu_detection_mock.py
  The file is the generic regression suite for Windows GPU detection;
  encoding the issue number in the filename is noise.
- Shorten module docstring, helper docstrings, per-test docstrings and
  inline comments in the renamed test file. No behaviour change,
  all 10 cases still pass.
- Shorten the _build_windows_path_dirs docstring in
  studio/backend/core/inference/llama_cpp.py and update the test-path
  reference; trim the win32 call-site comment to one line.

Local verification:
- pytest studio/backend/tests/test_windows_gpu_detection_mock.py -- 10 passed.
- pytest studio/backend/tests/test_llama_cpp_windows_nvidia_path.py
  studio/backend/tests/test_llama_server_args.py
  studio/backend/tests/test_windows_gpu_detection_mock.py -- 110 passed.

* Studio: harden _wait_for_health against transient httpx ReadError

The probe loop in LlamaCppBackend._wait_for_health only caught
ConnectError and TimeoutException. On Windows, when llama-server.exe
accepts the TCP probe and then dies before sending HTTP headers, the
peer process RST closes the socket. httpx maps this to ReadError
("WinError 10054 -- An existing connection was forcibly closed by the
remote host"), which fell through the except clause and bubbled out of
_wait_for_health, the routes/inference.py load_model handler, and back
to /api/inference/load as an opaque 500.

The crash diagnostic Studio actually wants to surface lives on the
self._process.poll() branch at the top of the loop body: "llama-server
exited with code X. Output: ...". We never reached that branch on the
WinError 10054 path because the very first probe blew up.

Expand the except to also swallow ReadError and RemoteProtocolError so
the next 0.5-second iteration runs the poll() branch. Outcomes:
  * Process really died: structured exit-code + last-stdout log line.
  * Single transient probe blip: silently retried; load succeeds.

Adds studio/backend/tests/test_llama_cpp_wait_for_health.py with five
cases covering happy-path 200, transient ReadError + dead process,
RemoteProtocolError + dead process, ConnectError cycling until success,
and dead process before the first probe. The new cases would have
failed against the old except clause -- ReadError / RemoteProtocolError
would have propagated instead of returning False.

Found while triaging the Windows Studio GGUF CI flake on this PR's
5a6ddc34 push: llama-server.exe (b9203 prebuilt) crashed within 2.2 s of
launch on the GPU-less runner, and Studio reported "WinError 10054"
instead of an upstream-tag-attributable exit-code line.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: danielhanchen <michaelhan2050@gmail.com>
2026-05-18 00:06:01 -07:00
Daniel Han
62410bf2af
studio: proxy-aware login rate-limit; allow google favicons in CSP (#5489)
* studio: proxy-aware login rate-limit; allow google favicons in CSP

Two follow-ups to #5375's auth + headers hardening.

Login rate-limit:
The per-IP bucket keyed on request.client.host alone. Behind any
reverse proxy or shared NAT it lumps everyone together (one user's
typos lock everyone out for 60 seconds; the 429 detail leaked the
proxy/internal IP back to clients). The bucket key is now
(client-ip, username.lower) so:
  - one wrong-password run does not block another user from the same IP
  - one IP does not block the same user from a different IP
The 429 detail body no longer interpolates the IP. Behind a proxy
clients can set UNSLOTH_STUDIO_TRUST_FORWARDED=1 so the limiter
honours X-Forwarded-For / Forwarded; off by default so a direct
caller cannot spoof the header.

CSP img-src:
components/assistant-ui/sources.tsx renders citation favicons from
https://www.google.com/s2/favicons. The current img-src allows
t0..t3.gstatic.com (used for other Google-hosted icons) but not the
main host the favicon URL points to, so every citation icon
CSP-blocks and falls back to gray initials. Adding www.google.com to
img-src is the same shape as #5409's connect-src HF allowlist fix.

Tests:
  - test_login_rate_limit.py (new): _client_ip respects
    UNSLOTH_STUDIO_TRUST_FORWARDED for X-Forwarded-For and Forwarded;
    bucket key is composed of (ip, lower(username)) and isolates
    cross-user and cross-IP buckets; 429 detail does not contain the
    client IP; Retry-After header preserved.
  - test_middleware.py: new test_img_src_allows_google_favicons pins
    that www.google.com is in the img-src directive and the existing
    gstatic CDNs stay allowed.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: normalise forwarded IPs, IP-wide aggregate cap, unknown-user sentinel

Reviewer follow-ups to the proxy-aware login rate-limit PR.

Forwarded address normalisation: with
UNSLOTH_STUDIO_TRUST_FORWARDED=1, raw `X-Forwarded-For` and
`Forwarded: for=` values such as `198.51.100.7:50001` or
`"[2001:db8::1]:50001"` were carried verbatim into the bucket key,
so one client emitting a fresh source port per attempt split into
many buckets and bypassed _LOGIN_MAX_FAILS. _normalize_forwarded_addr
now strips quotes, optional `[..]:port` for IPv6 and `host:port` for
IPv4, and validates as an IP literal; garbage values fall through to
the direct request.client.host. Forwarded parsing also isolates the
first forwarded-element so a multi-element header cannot create
attacker-controlled bucket strings.

Spray protection: the (ip, username) key removed the aggregate
per-IP throttle the pre-PR limiter provided. A client rotating
nonexistent usernames produced [401, 401, 401, 401, 401, 401] where
pre-PR produced [401, 401, 401, 401, 401, 429]. Restored the
aggregate via a parallel _LOGIN_IP_BUCKETS table (max 30 fails / 60s
per IP) checked alongside the per-(ip, username) bucket; both
buckets must be cleared on a successful login.

Bucket cardinality: every distinct unauthenticated username
allocated a new (ip, username) bucket entry without bound. 1,000
random usernames from one IP produced 1,000 buckets. Failures whose
username does not exist now record into a single sentinel key
(ip, "\x00unknown-user") so cardinality stays at one per IP for the
unknown path. The known-user path additionally enforces a global
hard cap (_LOGIN_MAX_BUCKETS = 4096) that prunes stale empty buckets
on overflow and otherwise folds the failure into the per-IP bucket
only.

Test:
  - python -m pytest studio/backend/tests/test_login_rate_limit.py -q
    -> 19 passed (was 12 before this commit; +5 forwarded-address
       normalisation, +1 sentinel bucket, +1 bucket cap)

CSP comment refreshed to mention `www.google.com` alongside
*.gstatic.com so future readers see why the host is allowlisted.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: tokenise img-src assertion to silence CodeQL substring rule

The new CSP google-favicon test used 'host string in directive string'
which CodeQL flagged as py/incomplete-url-substring-sanitization
(the substring could appear at an arbitrary position in a URL).
The assertion is checking a CSP directive, not URL sanitisation, but
splitting the directive on whitespace and asserting against the
tokenised source list expresses the same intent and matches the
exact CSP source expression. CodeQL no longer treats it as a URL
substring check.

Test: python -m pytest studio/backend/tests/test_middleware.py -q
      -> 14 passed

* studio: use any(src == host) for CSP source asserts

CodeQL's py/incomplete-url-substring-sanitization still flagged the
tokenised "host in img_sources" check. Switching to
`any(src == host for src in img_sources)` makes the comparison an
exact-equality (not substring) match, which the rule does not flag.

Test: python -m pytest studio/backend/tests/test_middleware.py -q
      -> 14 passed

* studio: trim verbose rate-limit + CSP comments

Compress the 6-line constants header on _LOGIN_BUCKETS to 3 lines and
the per-helper docstrings on _trust_forwarded_for / _normalize_forwarded_addr
to one line each. Same code, fewer in-flow tutorials.

Note in the CSP comment that www.google.com is the active favicon host
(used by sources.tsx for s2/favicons citations); *.gstatic.com stays as
legacy faviconV2 coverage but the SPA no longer fetches it.

33 tests in test_login_rate_limit.py + test_middleware.py still pass.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-18 00:02:15 -07:00
Daniel Han
d79fd92798
studio: scope cancel-cleanup to in-flight tmp dirs; walk back tool_call_id (#5488)
* studio: scope cancel-cleanup to in-flight tmp dirs; walk back tool_call_id

Two follow-ups to #5375's training and chat hardening.

_cleanup_cancelled_checkpoints used to rmtree every checkpoint-N
directory on Cancel. That is the opposite of what the user expects.
A user cancelling an 8h run with save_steps=2000 loses every
completed checkpoint they could have resumed from. The 67 MB residue
the audit memo flagged is the HF Trainer atomic-rename partial
(tmp-checkpoint-N), not the completed ones. The cleanup now targets
only tmp-checkpoint subdirs; completed checkpoint-N directories are
user-owned and stay. Symlinked output_dir and symlinked children are
skipped so the realpath containment cannot be levered into deleting
arbitrary content via a symlink trick.

ChatMessage._validate_role_shape stamped a random secrets.token_hex
id on tool messages with no tool_call_id. That id is uncorrelated
with the prior assistant tool_calls id, so strict passthrough
backends (OpenAI, Anthropic) reject the request as orphaned and
llama.cpp treats the tool result as "no preceding call" and
hallucinates. The synthesis moves up to ChatCompletionRequest, where
the whole conversation is visible: for each tool message missing an
id we walk back to the most recent assistant turn with tool_calls
(stopping at user turns), prefer a function.name match, otherwise
take the first unconsumed tool_call. Synthesis is the fallback when
no candidate assistant turn exists, preserving the prior round-trip
guarantee for orphaned tool messages.

Tests:
  - test_cleanup_cancelled_checkpoints.py (new): pins that completed
    checkpoint subdirs survive, tmp-checkpoint partials are removed,
    non-int suffixes (checkpoint-final, checkpoint-best) are left
    alone, output_dir outside outputs_root is refused, symlinked
    output_dir and symlinked child are both skipped, missing dir is
    a no-op.
  - test_inference_model_validation.py: 6 new walkback cases covering
    name-match preference, first-unconsumed fallback, explicit-id
    passthrough, multi-tool-result pairing, synth-on-no-parent, and
    no-cross-user-turn invariant.
  - test_openai_tool_passthrough.py: the two ChatMessage-level
    synth-on-missing tests are rewritten to assert that the per-
    message validator now leaves tool_call_id untouched; resolution
    coverage lives in the request-level tests above.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: explicit tool_call_id reserve, numeric tmp-checkpoint suffix only

Reviewer follow-ups to the training-cleanup + tool_call_id walkback PR.

tool_call_id walkback: a mixed assistant turn with [call_a, call_b]
followed by a tool result that carried tool_call_id="call_a" and a
sibling tool result with no id resolved to ['call_a', 'call_a']
because the explicit id never reserved call_a in the consumed set.
Added a pre-pass over the message list that walks back from every
role="tool" message carrying an explicit id and marks the matching
(asst_idx, tc_idx) consumed, then the missing-id walkback runs against
that pre-populated set. The second result now resolves to call_b.

While here, also harden the function-shape check: if a provider
ships a malformed tool_call where `function` is a string rather than
a dict, the old `(tc.get("function") or {}).get("name")` raised
AttributeError on the string's .get; now isinstance-gated so the
walkback falls through to the fallback id without raising.

Cancel cleanup: `tmp-checkpoint-*` is too broad. HF Trainer's
in-flight partials are always `tmp-checkpoint-<integer-step>`, so
constrain the cleanup regex to `^tmp-checkpoint-\d+$`. A user folder
named `tmp-checkpoint-final`, `tmp-checkpoint-backup`, or
`tmp-checkpoint-user-notes` is now preserved.

ChatMessage docstring still pointed at the pre-PR contract that
required `tool_call_id` on every role="tool" message. Updated to say
missing ids are accepted at message scope and resolved at
ChatCompletionRequest scope. Inline comment above the cancel-cleanup
call now describes the actual behaviour (in-flight tmp partials,
completed checkpoints preserved).

Test:
  - python -m pytest studio/backend/tests/test_inference_model_validation.py
    studio/backend/tests/test_cleanup_cancelled_checkpoints.py
    studio/backend/tests/test_openai_tool_passthrough.py -q
    -> 76 passed (was 67 before this commit; +2 walkback regression
       tests, +1 numeric-suffix preservation test)

* studio: trim verbose comments in cleanup + tool_call_id walkback

Move the HF tmp-checkpoint regex to module scope as a named constant.
Drop the multi-paragraph docstring on _cleanup_cancelled_checkpoints
and the inline call-site rationale; the function name + the test
class already cover the why.

Compress _resolve_missing_tool_call_ids docstring from a six-line
explanation to two. Same logic, fewer in-flow tutorials.

76 tests in cleanup + inference-model-validation + tool-passthrough pass.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-18 00:01:48 -07:00
Daniel Han
aab371a068
studio: tighten sandbox blocklist precision (bash, hf upload, NOFILE) (#5487)
* studio: tighten sandbox blocklist precision (bash, hf upload, NOFILE)

Three precision fixes in core/inference/tools.py. Same security
boundary; fewer false positives that broke legitimate sandbox use.

bash blocklist:
The per-token loop introduced in #5375 fired on any blocklist word in
any token position, so the entirely benign `grep -r curl .`,
`echo source the data`, and `ls /usr/bin/curl` were rejected with
"blocked command 'curl'". The position-anchored regex already covers
real command-position invocations, including `;rm`, `&&wget`, `$(rm)`,
`<(rm)`, backticked subshells, and `/usr/bin/sudo`. The token loop is
re-scoped: it only fires when the previous shlex token is a shell
separator (or at start of line), so split-quoting obfuscations like
`r''m -rf /` are still caught (shlex collapses them to a single
command-position token) while argument-position blocklist words pass
through. Trailing meta-chars glued to a shlex token (`rm;`) are
stripped before basename matching.

hf upload AST gate:
`_method_call_is_hf_upload` previously matched any method named
`upload_file` / `upload_folder` / `upload_large_folder` / `create_commit`
on any receiver, so paramiko.SFTPClient.upload_file, boto3.create_commit,
and similar non-HF SDK methods were rejected. The fallback now requires
an `import huggingface_hub` / `import hf_api` / `from huggingface_hub
import ...` somewhere in the same module. Fully-qualified
huggingface_hub.upload_file(...) calls are unchanged.

NOFILE env knob:
`RLIMIT_NOFILE = (1024, 1024)` was the only sandbox rlimit without an
env override. 1024 is below Linux's typical soft default and below
what multi-shard safetensors mmap chains need on Llama-3 70B-class
loads. Default is now 16384 with UNSLOTH_STUDIO_SANDBOX_NOFILE, parity
with the other rlimits.

15 new bash-blocklist-position tests pin both the false-positive
fixes and the still-blocked invariants (semicolon, &&, subshell,
backtick, split-quote, /usr/bin/ prefix, nested bash -c).
4 new hf-upload-import-gate tests pin both the false-positive
allowances and that HF-imported uses are still blocked.
1 new pin asserts the NOFILE env var is wired.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: cover command wrappers, find -exec, dynamic HF imports, NOFILE clamp

Reviewer follow-ups to the sandbox blocklist precision change.

Command-position scanner missed Bash command-prefix wrappers and inline
shell assignments. shlex tokenised `env curl`, `time curl`, `nohup rm`,
`FOO=bar curl`, `sudo rm`, etc. with the prefix at command position and
the real command at argument position, so the position-anchored check
returned set() while pre-PR's per-token scan caught them. Likewise the
position-anchored regex requires `^` or a shell separator before the
command, so `env curl` slipped through.

Reworked the scanner to track an expect_command flag plus a
prefix_pending flag:
  - assignments (FOO=bar) keep expect_command=True for the next token,
  - flags ('-oL', '--') keep it intact while prefix_pending is set,
  - numeric duration args ('timeout 1 cmd') skip without breaking
    expect_command,
  - known wrappers (env, command, builtin, exec, time, nohup, nice,
    setsid, stdbuf, timeout, ionice, chroot, sudo, doas, su, xargs)
    set prefix_pending so the wrapper's command is still checked,
  - shell separators now include `{`, `}`, `)`, `then`, `do`,
    `else`, `elif` so brace groups and if/then/while/do bodies are
    recognised as command positions.

Also lex with `shlex.shlex(punctuation_chars=";&|()`")` so split-quote
forms like `echo done; r''m -rf /tmp/x` and `echo done;r''m` tokenise
as `[..., ';', 'rm', ...]` and the command position check fires.

Added a small `find -exec CMD ... ;` / `-execdir CMD ... ;` pass so
`find . -exec rm -f {} +` and friends are caught even though the
direct token is at argument position to `find`.

Dynamic Hugging Face imports were treated as no-HF-in-scope. The
upload-method gate now also resolves `__import__('huggingface_hub')`,
`importlib.import_module('huggingface_hub')`, and bare
`import_module('huggingface_hub')` (via `from importlib import
import_module`) as HF imports, so HfApi().upload_file via dynamic
import is still blocked.

RLIMIT_NOFILE: setrlimit(NOFILE, (16384, 16384)) silently failed if
the parent's hard cap is below the requested value; the broad
except swallowed the OSError and left the sandbox at the parent's
default. Clamp the requested value to the inherited hard limit
before calling setrlimit.

Test cleanup: the existing test_cat_with_word_source_allowed had
`assert ... or True` so it could not fail; rewrote it to assert the
actual return value plus the two membership checks. Added
parametrised coverage for shell prefix wrappers, find -exec / xargs,
brace groups, if/then, while/do, split-quote command-name forms, and
dynamic HF import upload patterns.

Test:
  - python -m pytest studio/backend/tests/test_sandbox_tools.py -q
    -> 90 passed (was 67 before this commit)
  - full studio/backend/tests/ minus llama_cpp_load_progress_live and
    GPU CUDA_VISIBLE_DEVICES tests (pre-existing isolation flake)
    -> 1063 passed

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: catch bare-name HF upload calls in AST gate

`from huggingface_hub import upload_file; upload_file(...)` is a
canonical HF call shape that the previous Attribute-only check missed:
the bare-name call lands as ast.Name (not ast.Attribute), so the
fuzzy gate skipped it.

Extend _method_call_is_hf_upload to also match ast.Name when HF is in
scope. Same import-gating discipline as the Attribute branch, so
paramiko/boto3 and locally-defined `def upload_file(...)` helpers
without HF imports still pass.

Pins: 4 new TestHfUploadImportGate cases (upload_file/folder/create_commit
bare-name imports blocked; local upload_file without HF import allowed).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: scope HF uploads to sandbox-local literals; block env / token leaks

The previous gate dropped every HF upload call. Two refinements make it
precise enough to allow legitimate sandbox->HF uploads while still
catching credential / file exfil:

- path_or_fileobj / folder_path / create_commit operation paths must be
  sandbox-local relative-path literals (no '/', '~', drive letter, or
  '..' segments). Variable / dynamic paths are rejected.

- Any positional or keyword argument that statically resolves to
  os.environ / os.environ.get / os.getenv / bare getenv / subprocess
  shape readers is rejected (env-var exfil).

- token / hf_token / api_token / api_key / auth_token / access_token /
  password / secret kwargs are always rejected; sandbox env strips all
  parent credentials by construction, so any value here is hard-coded
  or lifted.

Recursive subtree walk in _reads_env_or_secret catches wrapper shapes
(str(os.environ), json.dumps(os.environ.items()), etc.).

Add TestSandboxEnvIsolation: pin that _build_safe_env builds the env
from a whitelist, not by stripping. Cover Linux/macOS/WSL/Windows
secret shapes. The whitelist is PATH / HOME / TMPDIR / LANG / TERM /
PYTHONIOENCODING (+ VIRTUAL_ENV / SystemRoot when applicable); HOME
points at the sandbox workdir, so HF / wandb / aws SDKs cannot reach
the operator's ~/.cache credentials.

Test classes added:
- TestHfUploadSandboxLocalPaths (relative literals allowed; absolute,
  drive-letter, '~', '..', mid-path traversal, dynamic vars, and
  open() of unsafe paths blocked, including create_commit recursion).
- TestHfUploadEnvAndSecretLeakBlock (os.environ subscript/get/getenv,
  bare getenv, subprocess.check_output, str(os.environ), token=,
  hf_token=, api_key=, and create_commit operations referencing env).
- TestSandboxEnvIsolation (no parent secret leaks into sandbox env).

131 tests in test_sandbox_tools.py pass.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-18 00:01:17 -07:00
Daniel Han
2fbbf8ade6
studio: expose launcher capability bits on unauth /api/health (#5486)
* studio: expose launcher capability bits on unauth /api/health

PR #5375 reduced the unauthenticated /api/health response to {status,
timestamp} only, on the theory that the rest of the payload was useful
fingerprinting. That was too aggressive: the Tauri watchdog reads
`service == "Unsloth UI Backend"` and `studio_root_id` to re-adopt its
own backend across restarts (src-tauri/src/desktop_backend_owner.rs
and commands.rs), and the SPA bootstrap fetches the same payload
unauth to detect chat-only mode and native path lease support before
any token is available (frontend src/config/env.ts and
features/native-intents/use-native-readiness.ts). With the post-#5375
shape, the watchdog kills its own healthy backend, the SPA never
flips out of "full Studio" mode on chat-only Linux/Windows, and the
About tab shows "dev" in place of the real version.

The actual fingerprint-ish fields are `version` / `studio_version` /
`device_type` (and to a lesser extent the hostname inside
`device_type`). `service`, `studio_root_id` (already a hex digest of
the install path, not the raw path), `chat_only`, the desktop_*
capability flags, and `native_path_leases_supported` do not leak the
install path or version.

This patch keeps the auth gate but rebalances which fields sit on
each side of it:

  unauth      service, studio_root_id, chat_only, desktop_protocol_version,
              desktop_manageability_version, supports_desktop_auth,
              supports_desktop_backend_ownership, native_path_leases_supported,
              desktop_owner (when present)

  authed      + version, studio_version, device_type

Existing must-change-password sessions still fall through to the base
payload because get_current_subject (strict) rejects them; that
matches prior behaviour.

test_middleware.py is updated to pin the new contract: launcher bits
present unauth, fingerprint fields present only with a valid bearer.

* studio: complete launcher-bits health unauth contract on Tauri + About tab

Reviewer follow-ups to the unauth /api/health launcher bits split.

Tauri preflight:
backend_capability_stale_reason() fell through to
backend_version_stale_reason(health.version.as_deref()) when capability
bits were present but version was absent. With the unauth payload now
exposing service + studio_root_id + desktop_* bits but gating version
behind a bearer, the desktop watchdog was reading the new payload,
parsing all capability bits, then classifying the same-root backend
as desktop_backend_version_missing and refusing to adopt it.

A backend that exposes desktop_protocol_version=1,
desktop_manageability_version>=1, supports_desktop_auth=true and
supports_desktop_backend_ownership=true was introduced together with
MIN_DESKTOP_BACKEND_VERSION=2026.5.3 in #5341, so a present capability
bitset is itself a version-compatibility signal. Skip the version
sub-check when version is None/empty; keep it for non-empty values
so genuinely-too-old backends that do echo a version still get
desktop_backend_version_too_old.

About tab:
fetchStudioVersions() did a bare fetch(apiUrl("/api/health")), which
the unauth payload no longer carries version/studio_version for, so
Settings -> About kept rendering "dev"/"dev" for any logged-in user.
Attach Authorization: Bearer <token> when getAuthToken() returns one;
fall back to bare fetch (still 200, just truncated payload) for the
not-logged-in case. No new endpoint.

Comment:
studio_root_id is no longer a hex digest of the install path; it is
an opaque per-install id written by the launcher. Updated the inline
comment to match.

Test:
  - python -m pytest studio/backend/tests/test_middleware.py::TestHealthAuthGate studio/backend/tests/test_desktop_auth.py -q
    -> 29 passed
  - npm run typecheck clean, npm run build produces fresh dist

* Trigger CI rerun for flaky Mac Chat UI step
2026-05-17 21:29:24 -07:00
Michael Han
3ff6204aa7
studio: load cached GGUF models when fully offline (#5505)
* studio: load cached GGUF models when fully offline

When huggingface.co is unreachable, GGUF model loads fail in three distinct
places even though the bits are already in ~/.cache/huggingface/hub. Each
failure has a different surface symptom:

1. list_gguf_variants() raises straight through HTTPException(500), so the
   variant dropdown shows 'Failed to list GGUF variants'.

2. detect_gguf_model_remote() silently returns None after retries fail. The
   caller then treats a GGUF-only repo as non-GGUF and routes it through the
   transformers/MLX path. On Apple Silicon this surfaces as 'Unsloth currently
   only works on NVIDIA, AMD and Intel GPUs.'

3. _download_gguf() loses list_repo_files() to the network and falls back to a
   filename heuristic ('{repo}-{variant}.gguf'). When the repo name does not
   echo the filenames (e.g. repo 'Qwen3.6-27B-MTP-GGUF' contains a file
   'Qwen3.6-27B-UD-Q4_K_XL.gguf' with no MTP), hf_hub_download cannot find
   that invented filename in the cache and aborts.

Fix in three layers:

- list_gguf_variants / detect_gguf_model_remote: honor HF_HUB_OFFLINE and
  fall back to scanning the local HF cache snapshot when the API throws.
  detect_gguf_model_remote still keeps its retry loop for transient flakes;
  the cache fallback only kicks in after every attempt fails.

- _download_gguf: when list_repo_files() fails, look up variant -> real
  filename inside the cached snapshot before resorting to the heuristic.

- llama_cpp.load_model / inference worker startup: when DNS for
  huggingface.co fails (2s probe), set HF_HUB_OFFLINE=1 for the process so
  every hf_hub_download call below resolves from cache instantly instead of
  spending ~25s on five exponential retries.

Online behavior is unchanged: the API is tried first and only used to fail
over. The cache scan is a strict subset of what list_local_gguf_variants
already does today for local paths.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: tighten inline comments on offline GGUF fallback

* studio: address review feedback on offline GGUF fallback

Fixes from the review pass on #5505:

* ruff F823 (lint CI red): the late `import os` at the bottom of
  LlamaCppBackend.load_model made `os` a function-local name, so my
  new `os.environ` reference at the top of the same method was a
  use-before-bind. Surfaces at runtime as
  'cannot access local variable os where it is not associated with a value'
  and is why the Mac/Windows Studio API jobs were failing too. The
  env-var mutation has been moved into a module-level contextmanager,
  so load_model no longer touches `os` directly.

* Codex P1: cache variant match now uses the relative path, not the
  basename. Layouts like `BF16/foo.gguf` (variant token only in
  parent dir) were silently skipped, falling through to the bogus
  `{repo}-{variant}.gguf` heuristic and failing offline loads of
  models stored under quant-named subdirs.

* Codex P1: HF_HUB_OFFLINE no longer persists past one model load.
  llama_cpp.load_model now uses a contextmanager that probes DNS,
  sets HF_HUB_OFFLINE/TRANSFORMERS_OFFLINE only when DNS is dead,
  and pops them in finally (preserving any prior user setting of
  TRANSFORMERS_OFFLINE). Pre-existing user-set HF_HUB_OFFLINE is
  respected as a no-op. worker.py keeps the startup probe because the
  orchestrator spawns a fresh worker per load -- comment updated to
  make that lifecycle explicit, and a warning is now logged.

* Gemini: cache-dir lookup centralized in `_iter_hf_cache_snapshots`.
  Three near-identical copies (in list/detect helpers and the
  llama_cpp offline scan) now go through one helper.

* Gemini: `huggingface_hub.utils.is_offline_mode` does not exist in
  1.x (verified locally); `huggingface_hub.constants.HF_HUB_OFFLINE`
  is snapshot-at-import-time and does not reflect runtime mutations.
  Manual env-var parsing kept.

* socket probe now saves and restores the prior default timeout
  instead of unconditionally setting None on exit, so it composes
  with caller code that already configured a timeout.

* worker.py probe now logs a warning when offline mode is auto-enabled
  so debugging the case isn't blind.

* studio: regression tests for offline GGUF cache fallback

Lock in the offline fallback path from #5505 so future refactors can't
silently regress either bug. 26 tests, 0.55 s, no network/GPU/subprocess.

Covers:

* _iter_hf_cache_snapshots: missing cache, missing repo, missing
  snapshots/, newest-mtime ordering, case-insensitive repo match.
* _list_gguf_variants_from_hf_cache and the list_gguf_variants
  online/offline-env/API-exception/reraise paths.
* _detect_gguf_from_hf_cache and detect_gguf_model_remote 3x-fail
  fallback. Pre-existing RepositoryNotFoundError early-return preserved.
* Codex P1 #1 regression: BF16/foo.gguf (quant only in subdir name)
  must resolve via _detect_gguf_from_hf_cache, which now matches the
  snapshot-relative path rather than the basename.
* _probe_dns_dead: returns True/False, restores prior socket timeout.
* Codex P1 #2 regression: _hf_offline_if_dns_dead sets env only inside
  the block, restores on exit (including on exception), re-probes DNS
  on the next call so a transient hiccup cannot lock the long-lived
  LlamaCppBackend singleton offline. Honors a user-set HF_HUB_OFFLINE
  as a no-op. Preserves a user-set TRANSFORMERS_OFFLINE across exit.

Follows the existing studio backend test stub pattern (loggers /
structlog / httpx stubs + backend dir on sys.path).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: extend offline cache fallback to _download_mmproj and quant label

Two follow-up fixes from the review pass on #5505:

* _download_mmproj() now mirrors _download_gguf()'s offline path:
  when list_repo_files() fails, scan the local HF cache snapshot for
  any GGUF whose basename starts with mmproj-. Without this, offline
  vision GGUF loads succeed at the main weight (the existing PR fix)
  but the mmproj returns None and llama-server starts without vision
  support. Same _iter_hf_cache_snapshots helper, F16 preference and
  fallback to the first match are preserved.

* _extract_quant_label() now considers parent directory segments when
  the basename has no quant token. Layouts like BF16/foo.gguf are
  already documented in this file and are returned by the new
  snapshot-relative-path filter in _download_gguf; before this fix
  their variant label collapsed to "foo" (the last hyphen segment of
  the basename). Regex is the same; the search just walks parent
  segments innermost-first if the basename misses.

Tests (studio/backend/tests/test_offline_gguf_cache_fallback.py):

* TestExtractQuantLabelSubdir: basename quant unchanged, quant-only-
  in-parent, UD- prefix in parent, deeper nesting picks the
  innermost matching segment.
* TestDownloadMmprojOfflineCacheFallback: cache fallback returns the
  mmproj when list_repo_files fails, F16 preference holds when both
  variants are in cache, no-mmproj cache returns None.
* httpx stub now prefers the real package when installed (the CI
  install list already includes it) and falls back to the stub only
  when httpx is genuinely missing. Newer huggingface_hub imports
  HTTPError/Response/Request at module load, so the previous
  fixed-set stub broke when those names were added upstream.

26 existing cases plus 7 new = 33 pass in 0.74s.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix/adjust offline cache + DNS probe per PR #5505 review

Four review findings tightened, with regression tests:

- list_local_gguf_variants subdir collapse (P1 codex 10:08): pass the
  snapshot-relative path to _extract_quant_label so BF16/foo.gguf and
  Q4_K_M/foo.gguf produce distinct labels instead of folding to the same
  basename pseudo-quant.
- list_gguf_variants cache fallback (P2 codex 12:10): surface
  RepositoryNotFoundError / GatedRepoError / RevisionNotFoundError /
  EntryNotFoundError to the caller instead of masking with stale cache,
  matching detect_gguf_model_remote.
- _detect_gguf_from_hf_cache mmproj (P2 codex 12:10): exclude mmproj
  files from the candidate list so a partial cache with only a vision
  projector cannot route the projector as the main model.
- _probe_dns_dead global timeout (P2 codex 13:06): run the gethostbyname
  on a daemon thread with join timeout so concurrent sockets in the same
  interpreter never inherit a process-wide socket.setdefaulttimeout
  mutation. Same shape applied in worker.py's startup probe.

* Make llama-server health check tolerant of warmup races

Two layered fixes for the Windows GGUF smoke CI Tool calling Tests
flake that exit-22'd on a single httpx.ReadError during llama-server
warmup. The 'windows-latest -> windows-2025-vs2026' image rollout is
hitting main with the identical symptom.

A. _wait_for_health: catch httpx.ReadError, RemoteProtocolError,
   WriteError alongside ConnectError and TimeoutException. A TCP RST
   mid-read while llama-server is still binding the port (WinError
   10054) is a 'still warming up' signal, not fatal. The existing
   _process.poll() check still wins for real crashes.

B. _drain_stdout + spawn: tee llama-server stdout/stderr to a
   per-launch log file at ~/.unsloth/studio/logs/llama-server/
   <port>.log. Any future subprocess crash leaves a forensic trace
   on disk even when Studio's traceback only captures the symptom
   (ReadError) and not the cause. Best-effort: a logging-side OSError
   never blocks the load.

Regression coverage: TestWaitForHealthRetriesOnReadError pins the
retry behaviour for the three new exception types and verifies that a
real process exit still short-circuits the loop.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* ci(windows): retry inference/load + collect llama-server logs

Composite fix for the Tool calling Tests flake that exit-22'd on a
single httpx.ReadError during llama-server warm-up. The
windows-latest -> windows-2025-vs2026 runner image rollout has been
hitting main with the identical symptom.

- All three jobs (openai-anthropic, tool-calling, json-images) now
  retry POST /api/inference/load up to 3 times with 10s backoff and
  preserve the response body for post-mortem. One transient 500 no
  longer fails the whole job.
- A new "Collect llama-server logs" step copies the per-launch
  llama-server stdout teed by Studio under ~/.unsloth/studio/logs/
  llama-server/ into the workspace, and the upload-artifact step
  now includes logs/llama-server/*.log so any future subprocess
  crash leaves a forensic trace.

---------

Co-authored-by: shimmyshimmer <shimmyshimmer@users.noreply.github.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2026-05-17 21:25:39 -07:00
Ashwin Upadhyay
867b1e1873
studio/openai: align chat completions docstring with stream=false default (closes #5047) (#5524)
* studio/openai: align chat completions docstring with stream=false default

The schema and regression test for ChatCompletionRequest.stream were
already corrected to default `false` (matching OpenAI's spec), but the
route docstring still claimed streaming was the default -- misleading
for anyone reading the source while debugging the original report.

Updates the docstring to reflect the actual behavior, adds an explicit
note pointing to #5047, and tags the existing regression test with the
issue reference and the .NET / System.Text.Json client class so the
intent survives future cleanup.

Closes #5047

* studio/openai: address review — move #5047 ref out of OpenAPI doc, add route-level test

Two follow-ups on review feedback:

- Drop "(see #5047)" from the openai_chat_completions docstring so the
  internal issue number doesn't leak into the FastAPI-generated
  OpenAPI / Swagger schema. The reference now lives in a code comment
  next to the actual `if payload.stream:` branch, where it's most
  useful for the next person debugging the same class of report.

- Add test_post_without_stream_field_decodes_to_stream_false_over_http:
  a TestClient-based wire-level guard that POSTs a body without
  `stream` (the exact shape naive curl / .NET / System.Text.Json
  clients send) and asserts both that the request deserialises into
  stream=False *and* that the response Content-Type is
  application/json, never text/event-stream. The existing
  constructor-level test would silently miss regressions introduced
  by middleware or alias rewrites that mutate the body before the
  pydantic model is built.

Refs #5047

* studio/openai: route-level test mounts real router instead of synthetic echo app

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: Roland Tannous <115670425+rolandtannous@users.noreply.github.com>
Co-authored-by: Roland Tannous <rolandtannous@gravityq.ai>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-18 01:19:50 +04:00
DoubleMathew
e7e02230e3
Fix /recommended-folders 500 on unreadable model directories (Python 3.12+) (#5523)
* Fix /recommended-folders 500 on unreadable model dirs (Python 3.12+)

get_recommended_folders probed candidate paths with a bare
Path(p).is_dir(). On Python <= 3.11 that returned False for an
unreadable path, but on Python 3.12+ is_dir() propagates
PermissionError (EACCES) instead, so a stock root-owned ollama
install at /usr/share/ollama/.ollama/models (mode 700) made the
endpoint 500 through the entire middleware stack.

Move the directory-accessibility check into a stdlib-only helper
utils.fs_access.is_accessible_dir that swallows OSError and keeps
the existing os.access(R_OK|X_OK) filter, restoring the pre-3.12
"unreadable path is simply not a candidate" behaviour. Add a
dependency-free regression test.

* Address review: inline helper, no new file, cover all probe sites

- Drop studio/backend/utils/fs_access.py and the cross-module import;
  the guard is now a small module-level _safe_is_dir() in models.py
  (also moots the import-placement / ModuleNotFoundError feedback,
  since there is no longer an import to misplace).
- Apply _safe_is_dir to every system-location probe with the
  vulnerable bare is_dir() pattern, not just /recommended-folders:
  _build_browse_allowlist._add and the /browse-folders _add_sug
  helper, so the same Python 3.12+ PermissionError cannot 500 those
  endpoints either. Each site keeps its exact prior semantics
  (recommended-folders retains its os.access(R_OK|X_OK) filter);
  the only behavioural change is "no longer crashes".
- Rewrite the regression test to extract the real _safe_is_dir from
  source via ast, keeping it dependency-free without standing up the
  FastAPI app, and correct the mode-000 case.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Roland Tannous <115670425+rolandtannous@users.noreply.github.com>
2026-05-18 00:16:14 +04:00
Daniel Han
b59e02e977
Studio: stop hint, Uvicorn log rename, reachability check + Mac UI CI retry hardening (#5503)
* Studio: clearer stop hint, Uvicorn log rename, external reachability check

Three startup-banner UX improvements to make it obvious how to stop
Studio, what the externally reachable URL really is, and whether that
URL actually works from outside.

1. Stop hint at the end of the banner
   * Bright orange "To stop Unsloth Studio: press Ctrl+C in this
     terminal." line, with a dim "(On macOS this is Control+C, not
     Command+C.)" follow-up so the macOS Cmd-vs-Ctrl confusion is
     headed off.
   * When bound to 127.0.0.1, an extra "To deploy and access globally"
     block tells the user the exact relaunch command
     (unsloth studio -H 0.0.0.0 -p PORT) with a trusted-networks
     caveat.

2. Uvicorn startup log rewrite
   * Installs a stdlib logging.Filter on the uvicorn / uvicorn.error
     loggers that:
       - renames the prefix to "Unsloth Studio running on"
       - swaps the wildcard bind for the resolved external host so the
         line agrees with the banner
       - replaces "(Press CTRL+C to quit)" with the same Mac-aware
         stop hint
   * Rewrites both record.msg and record.color_message so it works
     under plain and colorized log formatters.

3. External reachability self-test on wildcard binds
   * Synchronous probe via check-host.net's TCP JSON API confirms
     whether the advertised public URL actually accepts connections
     from the internet.
   * On failure prints the resolved IP, the failing-node count, the
     usual causes (AWS SG, GCP firewall rule, Azure NSG, home router),
     and an SSH local-forward workaround.
   * Verifies 127.0.0.1 / ::1 first and only offers a local fallback
     URL when loopback actually responds, so we never claim a port
     works when it does not.
   * Private / loopback / link-local display hosts short-circuit with
     a one-line LAN note instead of a probe.
   * Bounded at roughly 15 seconds, early-exits on two decisive node
     results, all failures swallowed.

Banner is split into print_studio_access_banner(include_stop_hint=...)
plus a new print_studio_stop_hint() so the reachability output can be
sandwiched between the URL section and the stop hint, keeping the
stop hint as the last text on screen.

Pure stdlib (socket, urllib, ipaddress, logging, threading), no new
dependencies, identical behavior on Linux, macOS, and Windows.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* CI: harden Mac Studio UI tests against Chromium ERR_NO_BUFFER_SPACE

The Mac Studio UI workflow already retries the Playwright scripts on
the racy 'Unexpected end of JSON input' pipeTransport crash, but
falls through on ERR_NO_BUFFER_SPACE -- a separate Chromium failure
that fires when the macos-14 free-runner kernel briefly runs out of
socket buffers. Same fix shape, two layers:

* In-script: when a change-password page.goto() attempt fails with
  ERR_NO_BUFFER_SPACE, sleep 5s then 15s before the next attempt so
  the OS has time to recover socket buffers. Other failures retry
  immediately as before.
* Workflow: extend both Playwright retry blocks (chat-ui and
  extra-ui) to also trigger the full Studio kill + reset + reboot
  retry on ERR_NO_BUFFER_SPACE, not just on the pipeTransport JSON
  crash.

Real assertion / timeout failures still bypass retry and surface
immediately. Linux and Windows workflows are unchanged; the flake
is macOS-runner-specific.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-17 07:44:06 -07:00