Commit graph

5,547 commits

Author SHA1 Message Date
Daniel Han
fcfbf166ff
studio(ui): use the --primary brand token for the avatar fallback color (#5987)
* studio(ui): use the --primary brand token for the avatar fallback color

The fallback profile avatar hardcoded #14b789, a slightly different green
from the app's general brand color (--primary = #17b88b, used by the send
button and every other primary-colored control). Next to primary-colored UI
-- e.g. the artifact preview/code panel -- the avatar's off-brand shade looked
inconsistent ("changes color weirdly"). Point avatarBgStyle() at
var(--primary) so the avatar always renders the general brand green and
follows the theme token.

Verified live in Studio: the avatar was rgb(20,183,137) (#14b789) while
--primary resolves to rgb(23,184,139) (#17b88b); the fix unifies them. This
is the only hardcoded brand-green left in the frontend -- every other
brand-green element already uses --primary / bg-primary.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* studio(ui): add literal fallback to the avatar --primary token

UserAvatar is a reusable component; if it is ever rendered outside the theme
root (where --primary is undefined), var(--primary) alone would compute to
transparent. Use var(--primary, #17b88b) so the avatar stays branded in that
edge case. When --primary is defined (the normal case, app-wide) it always
wins, so this changes nothing in practice -- verified in a browser:
var(--primary)=rgb(23,184,139), and an undefined var correctly falls back to
the literal.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 23:44:24 -07:00
Matt Van Horn
03349d1e05
feat: support text-only loading of Gemma 3 27B via FastLanguageModel (skip SiglipVisionModel) (#5816)
* feat: support text-only loading of Gemma 3 27B via FastLanguageModel (skip SiglipVisionModel)

* test: instantiate text-only Gemma3 model and assert no vision tower

Existing tests were AST source-introspection plus a config-resolves-to-
text-config check; none actually instantiated a model from the
text-only config. Add a small integration test that builds a shrunken
Gemma3TextConfig (CPU-cheap), instantiates the matching CausalLM
class, and asserts the resulting model exposes the LM head and has no
vision_tower or multi_modal_projector attribute.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Deduplicate _get_text_only_config into _utils for PR #5816

* Fall back to full model when a VLM has no text-only class for PR #5816

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Preserve quantization_config and clarify warning for text-only loading for PR #5816

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Only take text-only path when the VLM has its own text decoder for PR #5816

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Convert source string-match assertions to AST checks per Gemini review

* Load real VLM text weights on transformers 5.x for text-only mode in PR #5816

transformers >=5 changed Gemma3ForCausalLM base_model_prefix from language_model to model, so a VLM checkpoint's text weights (gemma3: language_model.model.*, gemma3n: model.language_model.*) no longer auto-strip onto the text decoder and were silently initialized random. Add a version-gated key_mapping that remaps them onto the text keys, returning None on transformers <5 where the prefix still strips and a mapping would break the load.

Apply the same family-guarded remap on the load_in_fp8 offline path and for direct FastBaseModel callers, and remap quantization llm_int8_skip_modules off the wrapper prefix after stripping.

Add a regression test that loads real VLM checkpoint weights (the prior tests only instantiated a fresh model so they missed this) and drop the bitsandbytes dependency from the quantization-config test.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Separate FP8 text-only cache and hoist the text-only guard for PR #5816

Address review of the text-only changes: (1) _offline_quantize_to_fp8 produced different artifacts for text-only vs full VLM but reused the same <name>-fp8-<mode> cache dir, so one mode could load the other's saved model; decide text-only before the cache name and add a -text-only suffix. (2) FastBaseModel.from_pretrained rewrote the VLM auto class to AutoModelForCausalLM before loading auto_config and before the family check, leaving is_vlm wrong for the fast_inference/vLLM block; hoist the family-guarded text-only decision above those checks and drop the redundant later block. (3) Wire the text-only regression test into the curated CPU pytest job so it runs in CI across the transformers matrix.

* Trim text-only code comments for PR #5816

Shorten and de-duplicate the comments added for the text-only loading work; keep the non-obvious rationale (the transformers >=5 base_model_prefix change) and drop the obvious parts. Comments only, no code changes; AST-based tests still pass on transformers 4.57.6 and 5.4.0.

* Make text-only loading opt-in via a public text_only argument for PR #5816

Rename the internal _force_text_only flag to a public text_only parameter on FastLanguageModel, FastModel and FastBaseModel (and the fp8 helper), defaulting False on all three. Text-only loading is now opt-in (text_only=True) instead of forced on by FastLanguageModel; the family guard and key remap are unchanged. Updated the AST tests for the new parameter and forwarding.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Trim text-only code comments for clarity

---------

Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2026-06-09 22:52:39 -07:00
Daniel Han
8848a310df
Studio: clean-room compact RAG (knowledge bases, hybrid search, fast indexing) (#5910)
Adds a self-contained RAG stack to Studio: knowledge bases with chunked indexing, hybrid (dense + lexical) retrieval, and an automatic first-pass context inject into chat. Embeddings run through a local llama-server GGUF backend (default unsloth/bge-small-en-v1.5-GGUF) with a sentence-transformers fallback. The chat tool loop gains a search_knowledge_base tool, a per-turn re-search cap, and source citation, layered on top of the shared ToolLoopController.
2026-06-09 21:17:04 -07:00
Nilay
436525d6de
Studio: stop the providers dialog from resetting custom provider form state (#6051)
* fix custom provider state

* Address provider seeding review feedback

---------

Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com>
Co-authored-by: imagineer99 <samleejackson0@gmail.com>
2026-06-09 21:43:24 +01:00
Wasim Yousef Said
2554636ded
Studio: follow-up fix for GGUF developer prompts (#6115)
* Studio: merge developer prompts for GGUF chat

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-06-09 18:11:38 +02:00
oobabooga
57be5868f9
Studio: improve OpenAI- and Anthropic-compatible API spec compliance (#6010)
* Studio: fix OpenAI- and Anthropic-compatible API spec compliance

* Studio: fix API spec-compliance gaps on passthrough and streaming paths

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: carry context_length_exceeded through the OpenAI passthrough error path

* Studio: count tool-schema tokens in the Anthropic server-tool stream, and small stream-handling guards

* Studio: guard message_delta usage against None and normalize developer role before proxying

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: honor max_completion_tokens on the external-provider proxy path

* Studio: forward llama-server cached_tokens into OpenAI prompt_tokens_details

* Studio: sanitize messages in count_tokens to match the /v1/messages prompt

* Studio: report max_tokens for truncated tool calls and guard null usage in metadata events

* Studio: drop the request-id middleware (headers aren't declared in either spec)

* Studio: include the required request_id field in Anthropic error bodies

* Studio: honor max_completion_tokens on the audio (TTS / audio-input) paths

* Studio: add the _effective_max_tokens helper and route all max-token sites through it

* Studio: align API compatibility edge cases

* Studio: clarify multi-choice chat support

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: clarify logprobs chat support

* Studio: opt the local chat UI into the streaming usage chunk so the context bar and tok/s repopulate

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: forward seed to llama-server, and fix Anthropic server-tool stop_reason, tool_result id correlation, and parallel-tool execution cap

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: align OpenAI chat completion spec edge cases

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: align backend API compatibility tests

* Studio: honor tool caps and internal stream usage

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: coerce nullable stream usage counts

* Studio: preserve system prompts with developer messages

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: wasimysaid <wasimysdev@gmail.com>
2026-06-09 17:13:25 +02:00
pre-commit-ci[bot]
a01960f20f
[pre-commit.ci] pre-commit autoupdate (#6104)
updates:
- [github.com/astral-sh/ruff-pre-commit: v0.15.15 → v0.15.16](https://github.com/astral-sh/ruff-pre-commit/compare/v0.15.15...v0.15.16)

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-06-09 07:29:47 -07:00
Wasim Yousef Said
ccb471f5bf
Improve local chat tool call flow (#5962)
Unify the Studio local tool-call loop (GGUF + safetensors) behind a shared ToolLoopController: ordered preface-then-tool-card rendering, duplicate-call de-looping with a forced final answer, XML-leak containment, and a parser fix that accepts closed <function=...> calls followed by trailing prose. Includes backend tests for the controller, strict parser, and GGUF route cursor reset.
2026-06-09 07:28:44 -07:00
Wasim Yousef Said
0d6d7dd4b3
Studio: make Helper LLM startup pre-cache opt in (#6113)
* Studio: make Helper LLM startup pre-cache opt in

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-06-09 15:28:34 +02:00
Wasim Yousef Said
33f4397b78
Studio fix recipe dataset preview (#6031)
* Studio: fix recipe dataset preview

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-06-09 14:02:00 +02:00
Eyera
aec41d17ed
feat(studio): Hub + Download Manager (#5916)
Adds the Studio Hub and download manager: browse Hugging Face models and datasets, download GGUF and safetensors with live progress and cancellation, and manage on-device inventory. The Hub does not require a GPU, so it is available on chat-only hosts.

CI: all substantive checks pass, including the three Core jobs after unsloth-zoo#736. The two red checks are non-code flakes, a transient npm-registry DNS resolution failure in the package scan and one quantized vision-model output assertion whose sibling shards passed.
2026-06-09 04:11:24 -07:00
Daniel Han
85314ed162
Studio frontend: reduce and tighten code comments (#6099)
Trim and tighten code comments across studio/frontend TS/JS. Comment-only: every changed file verified code-identical to main via the TypeScript printer signature comparison.
2026-06-08 23:10:35 -07:00
Daniel Han
187144d4e7
Reduce and tighten code comments and docstrings repo-wide (#6095)
Trim and tighten code comments and docstrings across the repository. Comment-only: every changed file verified code-identical to main via AST/token comparison.
2026-06-08 23:09:51 -07:00
Daniel Han
8292e699e4
Studio: make code comments and docstrings more succinct (#6029)
Trim and tighten code comments and docstrings across studio/ Python. Comment-only: every changed file verified code-identical to main via AST/token comparison.
2026-06-08 23:07:28 -07:00
oobabooga
ebf28e7e07
Studio: open the MCP dialog to the server list so servers can be managed (#6100)
Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com>
2026-06-08 19:18:05 +01:00
Datta Nimmaturi
6f36e8403a
Merge nvfp4_load CI fixes
Merged latest main, resolved KTO test conflicts, fixed nvfp4 test to use synthetic configs, fixed TRL/GRPO KTO drift
2026-06-08 23:23:53 +05:30
Datta Nimmaturi
b30e2b4b15
Merge qwen35_export CI fixes
Merged latest main, resolved save.py/KTO test conflicts, fixed TRL/GRPO KTO drift
2026-06-08 22:02:05 +05:30
Datta Nimmaturi
ca476c41f8
Merge studio_gemma4_vlm CI fixes
Merged latest main, resolved model_config.py conflict, removed redundant VLM checks
2026-06-08 20:20:08 +05:30
Datta Nimmaturi
6f27ecc66e
Merge moe-lora-target-fix CI fixes
Merged latest main, resolved _utils.py and KTO test conflicts
2026-06-08 20:20:00 +05:30
Daniel Han
3ce187da02
Formatting: ruff line-length 100, kwarg-spacing passes, drop blank after short local imports (#6079)
Raise ruff line-length to 100 and extend the local pre-commit format pipeline (def-signature magic-comma normalization, short multi-line assert collapse, kwarg '=' spacing, blank-line-after-short-import removal, adjacent string-literal / f-string+plain merge, redundant-pass pruning). Every transform re-checks the file AST and is dropped if it would differ; the whole-repo reformat is verified AST-identical per file and idempotent.
2026-06-08 04:24:13 -07:00
Daniel Han
8ccdf596aa
Studio: stop leaking internal exceptions to API clients; harden sandbox path (#6072)
* Studio: stop leaking internal exceptions to API clients; harden sandbox path

Security hardening for the FastAPI backend.

Error exposure (CodeQL py/stack-trace-exposure): many route handlers returned
raw caught-exception text to clients via HTTPException detail / response bodies,
which can leak internal filesystem paths and stack detail. Add shared helpers in
utils/utils.py (safe_error_detail, log_and_http_error) that log the full
exception server-side and return a generic message, and sweep the route layer
(inference, models, export, training, datasets, chat_history, providers,
mcp_servers, settings, data_recipe/{jobs,seed,validate,mcp}) to use them.
Intentionally user-facing validation messages, the existing _friendly_error SSE
paths, and upstream-service body passthrough (llama-server / OpenAI) are kept;
absolute server paths echoed in models.py browse/read errors are redacted.

Path injection (CodeQL py/path-injection): serve_sandbox_file already does
basename + realpath containment; add a strict filename allowlist
(^[A-Za-z0-9._-]{1,255}$) before the path is built as defense-in-depth and to
give the analyzer a clear sanitizer.

No behavior change beyond error-message text; status codes preserved.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: keep curated error messages, fix remaining load leak

- inference.py /load non-native path: redact str(e) instead of leaking it
  (matched the native branch which already redacted).
- llama_extra_args validation: return the curated, path-redacted message
  instead of the generic fallback so users see the offending flag.
- sandbox file serving: allowlist now forbids only separators/control chars
  via fullmatch, so generated images like 'loss curve.png' render again
  while traversal is still blocked by basename + extension + realpath.
- Add safe_curated_detail() for domain/validation exceptions whose message
  is intentionally user-facing; apply it to data_recipe job/validate,
  chat conflict, provider test, and MCP probe paths (these were collapsing
  to 'An internal error occurred', and 'connection' even mis-mapped to an
  upstream-service message). Generic Exception paths keep safe_error_detail.
- log_and_http_error: tolerate stdlib loggers (no structlog kwargs).
- delete_openai_container: log transport errors with exc_info like list/create.
- Drop helper/HTTPException imports this change left unused.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* log_and_http_error: log original error traceback on stdlib-logger fallback

* Tidy error-helper and sandbox comments for PR #6072

* Trim redundant comments in studio error-hardening routes for PR #6072

* Re-trigger CI now that unsloth-zoo #727 is merged (Core pulls zoo main)

* Address PR #6072 review feedback

- inference.py: keep the actionable NativePathLeaseError detail (path-redacted)
  instead of collapsing it to the generic message, matching the other curated
  validation paths in this file.
- utils.py: log via a single formatted log.error(exc_info=error) call that works
  for structlog and stdlib loggers; drop the now-unneeded try/except helper.
- models.py: use Path.name instead of os.path.basename(str(current)).

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-06-08 03:40:59 -07:00
Daniel Han
e20a6c3020
Restore KTO logps truncation guard for TRL (re-apply dropped #5996) (#6086)
* Restore KTO logps truncation guard for TRL (re-apply dropped #5996)

#5996 ported the KTO truncation guard to TRL's _compute_logps refactor but was
dropped from main in the 2026-06-05 history rewrite. Re-apply the
unsloth/models/rl_replacements.py guard (kto_trainer_get_batch_logps +
kto_trainer_align_completion_logps); its regexes still match TRL main's current
compute_ref_log_probs / _compute_kl_logps shape
(per_token_logps = selective_log_softmax(shift_logits, ...)), so the guard
remains effective.

Also extend the version_compat detection to recognize that current shape: TRL
refactored KTO again (no get_batch_logps / _compute_logps), so the test was
failing on TRL main even though the rewrite still applies.

* KTO patcher: match single or double quotes in dict-key regexes

Per review: _KTO_COMPLETION_RE / _KTO_KL_RE hardcoded double quotes for the
TRL dict keys, so a formatter or TRL version using single quotes would make the
patch silently skip. Accept both quote styles. Verified the regexes still match
TRL main's current experimental/kto source.
2026-06-08 03:40:47 -07:00
Daniel Han
b2b4e4c376
CI: allowlist deepseek_ocr2 in the compiler full-model-sweep (#6085)
transformers-latest ships a new deepseek_ocr2 model whose source-rewriter
compile exceeds the 60s per-model budget on the CI runner, same as the
existing beit/sam/sam_hq entries. Add it to KNOWN_BROKEN_COMPILE Category F
so HF=latest Core stops failing on a new upstream model. The slow compile
path itself remains a follow-up for unsloth_zoo.
2026-06-07 21:43:09 -07:00
Michael Han
cf97faed9f
Studio: keep chat in place when composer attachments resize it (#6070)
* Studio: keep chat in place when composer attachments resize it

Attaching or removing a file in the chat composer could yank the whole
conversation to the bottom, and the grown composer covered the end of
the chat with no way to scroll it back into view.

Root cause: the Viewport composes refs with an identity that changes on
re-render, so React re-runs our scroll ref on unrelated renders and the
autoscroll hook treated every rebind as a fresh mount, pinning to the
bottom. On top of that the viewport reserved a fixed 160px under the
last message regardless of composer size.

- Treat same-element ref rebinds as no-ops in the autoscroll hook; only
  a genuinely new viewport element pins and resets detach state
- Size the bottom spacer from the measured composer height plus a 24px
  gap so the chat can always be scrolled above the composer
- On composer growth, detach from the bottom instead of auto-scrolling;
  the user scrolls down to reveal the covered lines
- On composer shrink, defer the spacer shrink until it cannot clamp
  scrollTop, then release it invisibly on scroll or on bottom-pinning
  moments (run start, thread switch, thread load)

* Studio: release deferred composer spacer when a run owns the bottom

Sending with attachments cleared the chips after thread.runStart had
already fired, so the spacer shrink was deferred while the user sat
pinned at the bottom, leaving a permanent extra gap above the composer.
Apply shrinks immediately while a run is active or within 1s of run
start; the run-start pin owns the bottom then, so the clamp is the
intended glide. Caught by a cross-engine Playwright pass (Chromium,
Firefox, WebKit) over the pre and post builds.

* Studio: track the viewport element in state so listeners survive remounts

The deferred-shrink scroll listener was attached once against a ref, but
the keyed overlay provider remounts the viewport subtree on thread
switches, leaving the listener bound to the unmounted element. Removing
an attachment near the bottom in the new thread then left the oversized
spacer stuck until a run started. Track the viewport element in state so
the listener and the clamp math follow the new element.

Reproduced and verified with a thread-switch scenario on Chromium,
Firefox and WebKit; full matrix re-run green.

* Studio: release deferred composer spacer shrink when at the bottom (#6070)

---------

Co-authored-by: shimmyshimmer <michael@unsloth.ai>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2026-06-07 01:58:01 -07:00
Michael Han
da1c5b4b94
Studio: remove red border on chat error messages (#6063)
Co-authored-by: shimmyshimmer <shimmyshimmer@users.noreply.github.com>
2026-06-07 01:57:58 -07:00
Michael Han
1e811acd62
Studio: tag MLX loaded models as MLX instead of Base in chat (#6067)
* Studio: tag MLX loaded models as MLX instead of Base in chat

* Studio: tag MLX named hub defaults via name heuristic
2026-06-07 01:57:55 -07:00
Michael Han
1b588cd141
Studio: emit usage and timings for MLX generation speed stats (#6068)
* Studio: emit usage and timings for MLX generation speed stats

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: make MLX generation stats request scoped

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-06-07 01:57:52 -07:00
Daniel Han
686a30f95e
Studio: stop ROCm amd-smi tests leaking a fake loggers into sys.modules (#6055)
Follow-up to #6027. The four TestAmdGpuMonitoring tests also set
sys.modules["loggers"] = MagicMock() without cleanup, leaking a mock
loggers module into later tests. Switch them to monkeypatch.setitem so
the stub is undone at teardown, matching the worker test fix in #6027.
2026-06-06 21:19:21 -07:00
Daniel Han
0003f889e6
Studio: stop ROCm worker test leaking a fake utils into sys.modules (#6027)
test_direct_wheel_url_returns_none_without_cuda_major set sys.modules
"utils"/"utils.hardware" (and structlog/loggers) to MagicMocks without
cleanup. Once run.py started importing utils.cpu_threads (#5760), the
leaked non-package utils made later tests in the same job fail with
'No module named utils.cpu_threads; utils is not a package', e.g. all of
test_selection_logic.py's TestStudioLocalhostIpv6Warning. Use
monkeypatch.setitem so the stubs are undone after the test.
2026-06-05 07:52:26 -07:00
Lee Jackson
783c9d1e83
Studio: fix chat preset persistence with fast mode (#5870)
* fix: persist chat presets with fast mode

* Add schema drift guard test for chat inference settings (#5862)

Asserts ChatInferenceSettings declares every InferenceParams field the
frontend persists (all but checkpoint). With extra="forbid", a field
present in the UI but missing here 400s PUT /api/chat/settings, which is
exactly how fastMode regressed. Catches the next occurrence at CI time.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-06-05 07:16:50 -07:00
Daniel Han
fe37921223
Studio: fix load_freeze audio-type tests for #6000's Gemma 4 <|audio|> probe (#6018)
* Studio: fix load_freeze audio-type tests for #6000 Gemma 4 <|audio|> probe

#6000 extended LlamaCppBackend._detect_audio_type_strict audio_vlm arm to
also probe Gemma 4 `<|audio|>` (alongside Gemma 3n `<audio_soft_token>`),
but did not update the load_freeze simulation suite (last touched by #5922).
Its "no-match" and "bicodec" fixtures only defeat `<audio_soft_token>`; the
unmapped `<|audio|>` probe falls through to FakeLlamaServer 1-token default,
so detect_audio_type now returns audio_vlm where these tests expect
None / bicodec:

  - test_functional_equivalence_no_match
  - test_functional_equivalence_bicodec_match
  - test_response_shape_matches_pre_fix_for_no_match

main push-CI does not run "Repo tests (CPU)" (pull_request-only), so this
surfaces in every open PR merge-ref (e.g. #5940, which is unrelated to audio).

Fix: map `<|audio|>` to a 2-token response in the three fixtures that intend
a non-audio_vlm result (restoring their original semantics), and add a
positive test_functional_equivalence_audio_vlm_match locking in #6000 new
`<|audio|>` detection.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-06-05 07:15:45 -07:00
Lee Jackson
fe604fde20
Studio: accept system-role messages in Claude Code requests (#6006)
Normalize misplaced system-role messages in /v1/messages by hoisting their content into the top-level Anthropic system field, fixing the 422 that newer Claude Code clients trigger. Null and non-text system content is ignored rather than stringified.

Fixes #6001
2026-06-05 05:02:54 -07:00
Lee Jackson
9806e36aa4
Studio: enable GGUF tools with vision inputs (#6009)
* fix: enable GGUF tools with vision inputs

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix: GGUF vision tool routing

* Dedupe system messages on GGUF vision tool path for PR #6009

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2026-06-05 03:46:04 -07:00
Matt Van Horn
f22e92c8e4
fix: persist Studio thread synchronously on first runStart so mid-stream refresh keeps the prompt (#5814)
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2026-06-05 02:46:34 -07:00
Matt Van Horn
5cdbfef390
fix: warn when localhost resolves to ::1 but Studio is bound only to 127.0.0.1 (#5994)
* fix: warn when localhost resolves to ::1 but Studio is bound only to 127.0.0.1

* studio: fix localhost/::1 warning suppression and cover _run wiring

Addresses the Codex review on #5994 plus review-team findings:

- Remove the `_local_port_open("::1", port)` early-return. Studio binds
  127.0.0.1 only, so a successful connect to ::1:<port> means a *different*
  process is there -- exactly when http://localhost opens the wrong service
  and the user most needs the warning. Dropping the probe also removes the
  ~0.25s startup latency and the probe/warn race.
- Extract the banner/warning block from `_run` into `_emit_startup_output`
  so the wiring is unit-testable, and make the mismatch vs wildcard paths
  an explicit if/elif (they are mutually exclusive by construction).
- Hoist the `_working_local_url` confirmation out of the try block and
  reorder `_stdout_color_ok` before its only caller.
- Tests: add `_emit_startup_output` integration coverage (banner
  include_stop_hint, warning emission, single stop hint), a regression test
  that ::1 being occupied does NOT suppress the warning, dual-stack and
  non-positive-port cases; drop the unreachable `None` getaddrinfo arm.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: Etherll <mrmrmidessam@gmail.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-06-05 02:40:56 -07:00
Michael Han
9f1d029c18
Studio: refine tool call and reasoning trigger UI (#5873)
* Studio: refine tool call and reasoning trigger UI

Tool call triggers:
- Chevron fades in on hover or keyboard focus and sits next to the
  label instead of being pinned to the right edge, matching the other
  collapsible triggers.
- Labels wrap instead of truncating so long tool names and search
  queries stay fully readable.
- Smaller chevron for a lighter look.

Reasoning trigger:
- Smaller chevron to match.
- Thinking box drops its bottom padding and raises the streaming max
  height so more of the thinking text is visible.

* Studio: pointer cursors and sidebar 3-dots polish

Collapsible triggers:
- Pointer cursor on the reasoning, tool call, and tool group triggers
  so they read as clickable.

Chat sidebar:
- Swap the chat row 3-dots menu to the vertical more-vertical icon.
- Pointer cursor on the chat row and its menu button.
- Chat row right padding opens up on hover (pr-4 at rest, pr-8 on
  hover) so the title keeps a comfortable gap and clears the menu.

* studio: refine tool-call spinner, chevron, and reasoning spacing

- Use the lucide arc spinner for running tool calls and the app-wide
  Spinner, so loading states match the rest of the UI.
- Collapse long tool-call labels to a single line with an ellipsis,
  reveal the full label when the row is expanded, and fix the clipped
  descenders.
- Keep the collapse chevron next to the label and add top spacing above
  the reasoning trigger.
- Remove the redundant nested spinner in the web search running state.

* Studio: drop tool call group background fill

The ghost tool call group used a translucent bg-muted/10 fill that read
as a faint lighter box around every group in dark mode. Remove the fill
and rounding so the group sits flush on the chat background.

---------

Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2026-06-05 01:55:03 -07:00
Michael Han
2b51bec946
Fix chat text cutoff at composer dock and speed up plus icon spin (#5989)
The composer dock backdrop was a solid block with a hard top edge, so
chat text scrolling underneath got visibly clipped. Replace it with a
gradient that fades the top 28px to transparent.

Also shorten the plus to x rotation in the composer from 300ms to 250ms,
including the reduced motion override.
2026-06-05 01:54:30 -07:00
Daniel Han
4c06c1dcc7
Studio: enable audio input for Gemma 4 GGUFs; default chat model to Qwen3.5-4B-MTP (#6000)
* Studio: enable audio input for Gemma 4 GGUF models

Audio file upload was disabled for Gemma 4 vision+audio GGUFs (e.g.
gemma-4-12b-it-GGUF) even though their mmproj carries an audio encoder
(clip.has_audio_encoder, gemma4ua). Two causes:

- Audio-input detection only matched Gemma 3n's <audio_soft_token>;
  Gemma 4 uses <|audio|>, so audio_vlm was never detected.
- The GGUF load/status responses hardcoded has_audio_input=False, so the
  flag was dropped even when audio_vlm was detected (affected Gemma 3n
  GGUFs too).

Changes:
- Recognize <|audio|> alongside <audio_soft_token> in the llama-server
  token probe and the tokenizer-config pattern.
- Read clip.has_audio_encoder from the mmproj as an independent,
  model-agnostic signal (read_mmproj_audio_capability).
- Emit the computed has_audio_input on the GGUF load/status responses.
- Tests for the new pattern and the mmproj reader.

* Studio: default chat model and dataset helper to Qwen3.5-4B-MTP

Switch the auto-loaded chat default and the dataset-analysis helper GGUF
from gemma-4-E2B-it to unsloth/Qwen3.5-4B-MTP-GGUF (UD-Q4_K_XL).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-06-04 00:56:53 -07:00
Daniel Han
0425a3c0a1
Normalize shell scripts to LF in .gitattributes (#5997)
Shell scripts are stored as LF in git, but without an eol rule a Windows
clone with core.autocrlf=true checks them out as CRLF. The trailing \r then
breaks them when run in WSL/Linux -- e.g. `set -e` becomes `set -e\r` and
dash/sh aborts with "set: Illegal option -". This bites developers who clone
on Windows and run the repo's *.sh directly in WSL, increasingly common with
the AMD Strix Halo ROCm-on-WSL support.

Add `*.sh text eol=lf` so every shell script always checks out with LF
regardless of the contributor's platform or core.autocrlf setting. All
tracked *.sh use Unix shebangs; none need CRLF. PowerShell/batch scripts are
left untouched -- they tolerate LF and are unaffected by this bug.

Verified with `git ls-files --eol`: every *.sh now resolves to
i/lf w/lf attr/text eol=lf.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 00:39:29 -07:00
Long Yixing
63dc27f76e
fix(studio): disable mlx gc for none (#5991) 2026-06-04 00:38:45 -07:00
Daniel Han
636455a7d6 Revert "Port KTO logps truncation guard to TRL 1.x _compute_logps refactor (#5996)"
This reverts commit 157cecb25c.
2026-06-04 07:17:55 +00:00
Daniel Han
b1ee492982 Revert "CI: mark deepseek_ocr2 as known-broken compile timeout (#5995)"
This reverts commit 4eac527247.
2026-06-04 07:17:55 +00:00
Daniel Han
4eac527247
CI: mark deepseek_ocr2 as known-broken compile timeout (#5995) 2026-06-04 00:08:19 -07:00
Daniel Han
157cecb25c
Port KTO logps truncation guard to TRL 1.x _compute_logps refactor (#5996)
* Port KTO logps truncation guard to TRL 1.x _compute_logps refactor

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-06-04 00:07:58 -07:00
Daniel Han
b0572bd233
Bump install.sh / install.ps1 pin to unsloth>=2026.6.1 (#5977) 2026-06-03 10:22:24 -07:00
Daniel Han
c7d2ed1920
Fix macOS Apple Silicon installs resolving torch against x86_64 (#5976) v0.1.44-beta
* Fix macOS Apple Silicon installs that resolve torch against x86_64

On Apple Silicon, `uv venv --python 3.13` can reuse a cached x86_64
(Rosetta) CPython, often because uv itself is an x86_64 build. The
resulting venv reports macosx_*_x86_64 to the wheel resolver, but PyTorch
has shipped no macOS x86_64 wheels since 2.2.2, so the torch install fails
with "no wheels with a matching platform tag (macosx_..._x86_64)".

Two changes, both scoped to macOS arm64 and additive (no other install
path is affected):

- Create the venv with an arch-explicit `cpython-X.Y-macos-aarch64-none`
  request on Apple Silicon (no --python override), so uv cannot fall back
  to a cached x86_64 interpreter.
- Harden the existing x86_64 venv guard: when the venv python cannot be
  executed (x86_64 binary on a Mac without Rosetta), the platform.machine()
  probe returns empty and the recreate was silently skipped. Fall back to
  reading the binary's Mach-O arch via lipo/file so migrated or
  pre-existing x86_64 venvs are still recreated as arm64.

* Harden arm64 static-arch fallback: file -L and set -e safety

Address review feedback on the lipo/file fallback:
- uv symlinks the venv's bin/python to the base interpreter; plain `file`
  reports the symlink ("symbolic link to ...") and the arch substring never
  matches. Use `file -L` to dereference (lipo already follows the link).
- Append `|| true` so the command substitution cannot abort the installer
  under set -e on a Mac that has neither lipo nor file.

---------

Co-authored-by: danielhanchen <michaelhan2050@gmail.com>
2026-06-03 07:29:18 -07:00
Daniel Han
08d02610d9 Versioning 2026-06-03 06:35:55 -07:00
Daniel Han
9a6f404837
Fix UnicodeEncodeError when printing emoji on legacy Windows consoles (#5948)
On Windows only, at unsloth import, reconfigure stdout/stderr to UTF-8 when they are not already, so emoji and box-drawing glyphs do not crash legacy code-page consoles (e.g. cp1252) at SFTTrainer init. No-op on Linux/macOS and when output is already UTF-8 (PYTHONUTF8, modern terminals), and fully guarded so it can never raise.
2026-06-03 06:15:06 -07:00
Datta Nimmaturi
3f68dd5f0e
Patch sibling config module so GRPOConfig resolves to the patched class (#5946)
Fixes #3931. After patching a TRL trainer, also patch the sibling config module (e.g. trl.trainer.grpo_config.GRPOConfig) to the Unsloth-patched config, so importing the config from its own module returns the patched class carrying unsloth_grpo_mini_batch. Defensive (try/except + hasattr) so it safely no-ops when no sibling config module exists.
2026-06-03 06:14:54 -07:00
Daniel Han
aa0db1ff5b
fix(studio): don't double-quote the reset-password hint for spaced paths (#5975)
Addresses review feedback on #5971. _reset_password_command() already
shell-quotes the launcher path on POSIX (shlex.quote), so wrapping the result in
another pair of single quotes in the error string produced a mangled hint for
installs / home dirs containing spaces, e.g.

  Run ''/tmp/Unsloth Studio/.../unsloth' studio reset-password' in your terminal

which a shell mis-parses. Drop the outer quotes and put the command at the end of
the message so it is unambiguous and copy-pasteable in every case:

  Incorrect password. To reset it, run this in your terminal: <cmd>

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-03 06:10:23 -07:00