Commit graph

5,295 commits

Author SHA1 Message Date
Daniel Han
09ed2b963d Make Gemma-4 MoE swap honor full quantization_config and cover vLLM path
- Normalize string compute_dtype values from dict-style quantization_config
  (e.g. {"bnb_4bit_compute_dtype": "bfloat16"}) into torch.dtype before
  forwarding to bitsandbytes. A raw string would propagate to
  bnb.nn.Linear4bit and crash the first forward with
  "Invalid device string: 'bfloat16'".

- Forward bnb_4bit_quant_type from the user's BitsAndBytesConfig or dict
  config into swap_gemma4_experts_to_per_expert_linear4bit so swapped
  experts match the quantization type used by the rest of the model
  (previously always nf4 even when the caller requested fp4).

- Wrap the swap and its warning into a local closure and invoke it from
  both the regular auto_model.from_pretrained branch and the
  fast_inference=True / convert_vllm_to_huggingface branch. The closure
  is idempotent on non-Gemma-4 models, so the vLLM call is free when the
  loaded model has no Gemma4TextExperts modules.

- Escalate partial-state swap failures: if the helper raises after one
  or more Gemma4TextExperts modules were already committed to 4-bit,
  the wrapper re-raises a RuntimeError instructing the caller to reload
  the model. Previously a warning implied a clean BF16 fallback, which
  is false when partial conversion has already occurred.

The closure is multi-line (long block) because it needs to capture the
already-resolved quantization parameters and be reusable across both
load paths; the alternative is duplicating the entire block.
2026-05-16 15:46:16 +00:00
Daniel Han
41f8792e7c Honor quantization_config.load_in_4bit for Gemma-4 MoE swap
The Gemma-4 MoE per-expert Linear4bit swap previously gated only on the
positional load_in_4bit argument. loader.py forwards load_in_4bit=False
to FastBaseModel.from_pretrained whenever the caller supplies a
quantization_config (BitsAndBytesConfig), so callers that opt in via
UNSLOTH_GEMMA4_MOE_4BIT=1 plus BitsAndBytesConfig(load_in_4bit=True)
silently bypassed the swap. The adjacent guardrail already normalises
load_in_4bit from quantization_config; the swap gate now does the same
and sources bnb_4bit_compute_dtype from quantization_config when no
local bnb_config is built.

The except branch around the swap also previously stated "Falling back
to BF16 experts", which misrepresents the model state when the helper
fails partway through (already-swapped Gemma4TextExperts modules stay
in 4-bit; only the remainder remain BF16). The warning now counts the
modules marked _unsloth_gemma4_moe_4bit_swapped and reports the partial
state, advising a reload to recover a uniform state.

The comment above the fused-Parameter dels in gemma4_moe_4bit.py
overstated the swap's memory bound; rephrased to describe the actual
per-module peak (fused BF16 plus accumulated per-expert nf4).
2026-05-16 15:46:16 +00:00
Daniel Han
5a09093305 Merge quantization-bypass guardrail base into feature branch 2026-05-16 15:46:16 +00:00
Daniel Han
737643b173 Scrub .github/workflows for staging push (matches staging base) 2026-05-16 14:20:53 +00:00
pre-commit-ci[bot]
610e15819a [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-16 14:13:55 +00:00
Daniel Han
531706afa2 Harden quantization-bypass guardrail and apply to all bnb load paths
- Fix NameError on full_finetuning in FastLlamaModel.from_pretrained
  causal-LM branch; pull from kwargs instead of an undefined local.
- Apply the guardrail to the sequence-classification branch so
  num_labels users get the silent-bypass warning too. Block is appended
  purely additively; the prior `# Attach dispatch hooks` comment and
  `_attach_bnb_multidevice_hooks` import (added earlier for multi-GPU
  bnb dispatch hardening) are preserved untouched so that protection is
  not regressed.
- Count torch.int8 alongside torch.uint8 as quantized payload; bnb
  Linear8bitLt / Int8Params stores 8-bit weights as int8 post-cuda, so
  the previous accumulator zeroed and false-fired the partial warning.
- Tighten partial-bypass condition to require quantized_bytes > 0, so
  the warning only fires when quantization actually produced payload.
- Drop the overbroad "mlp.gate" skip pattern; it suppressed fused/custom
  mlp.gate_proj / mlp.gate_up_proj bulk weights that are exactly the
  partial-bypass shape this guard must report. Standard Linear4bit
  gate_proj weights are uint8 and counted as quantized earlier, so no
  new false positive on healthy 4-bit loads.
- Use the real bnb class name (Linear8bitLt) in the 8-bit total-bypass
  message instead of synthesising "Linear8bit".
- Accept an optional quantization_config so callers passing
  BitsAndBytesConfig directly (loader.py sets load_in_4bit_kwargs=False
  in that path) still get the bypass check.
- Drop a duplicate transformers_version import in vision.py.
2026-05-16 14:10:15 +00:00
Daniel Han
f533c2257d
Merge branch 'fix-issue-5344-quantization-guardrail' into feat-gemma4-moe-4bit-swap 2026-05-15 21:14:28 -07:00
Daniel Han
6faebabff9
Merge branch 'main' into fix-issue-5344-quantization-guardrail 2026-05-15 21:13:24 -07:00
Daniel Han
54a86c3514
ci: route every hf download through xet-tuned stall-retry wrapper (#5476)
Root cause of the Mac json-images 30 min timeout (run 25950714888 /
PR #5430): huggingface_hub>=1.15 deprecated `hf_transfer` and routes
every transfer through `hf-xet`. The CI step's unpinned
`pip install --upgrade huggingface_hub hf_transfer` jumped to 1.15.0
+ hf-xet 1.5.0, the 940 MB mmproj finished in ~21s, then the 3 GB
gemma-4 GGUF made it to ~46% and went completely silent for the
remaining 29 minutes -- no progress bytes, no error, no exit -- until
the job timeout fired.

This wraps every CI `hf download` in a new
`.github/scripts/hf-download-with-retry.sh`:

  * Drops the no-op `HF_HUB_ENABLE_HF_TRANSFER=1` prefix and the
    `hf_transfer` install (both are deprecated on 1.15+ and only
    emit a FutureWarning now).
  * Exports the hf-xet high-performance knobs Daniel asked for:
        HF_XET_HIGH_PERFORMANCE=1
        HF_XET_CHUNK_CACHE_SIZE_BYTES=0
        HF_XET_NUM_CONCURRENT_RANGE_GETS=64
        HF_XET_RECONSTRUCT_WRITE_SEQUENTIALLY=0
        HF_XET_CLIENT_READ_TIMEOUT=500
  * Watchdogs each attempt: if `hf download` has not exited after
    HF_DOWNLOAD_STALL_SECONDS (default 180s = 3 min), SIGTERM,
    sleep 2, SIGKILL, then loop. Retries are unbounded; the
    enclosing job's `timeout-minutes` is the real cap.
  * Optional 3rd positional `LOCAL_DIR` -- omitted lets `hf` use
    the default HF_HUB_CACHE, which is what the HF_HOME-priming
    jobs need.

19 call sites migrated across mlx-ci.yml + 9 studio-*-smoke.yml
workflows. The inline `python -c "from huggingface_hub import
hf_hub_download; ..."` block in mlx-ci.yml is also routed through
the wrapper so every hf transfer in CI gets the same treatment.

Also reverts the json-images timeout 45 -> 30 from #5475: the bump
was masking this hang, not fixing it.
2026-05-15 21:11:56 -07:00
Daniel Han
a68747aaff
Merge branch 'fix-issue-5344-quantization-guardrail' into feat-gemma4-moe-4bit-swap 2026-05-15 20:54:06 -07:00
Daniel Han
cefcf8e2d9
Merge branch 'main' into fix-issue-5344-quantization-guardrail 2026-05-15 20:52:57 -07:00
Daniel Han
295844670b
ci: bump Mac json-images timeout 30 -> 45 min (cache-miss path) (#5475)
The `JSON, images` job in `studio-mac-inference-smoke.yml` (Job 3
of Mac Studio GGUF CI) downloads ~4 GB on a cache miss: 3 GB
gemma-4-E2B-it-UD-Q4_K_XL.gguf + ~1 GB mmproj-F16.gguf. The 30 min
cap was tight even with `HF_HUB_ENABLE_HF_TRANSFER=1` and parallel
downloads, and timed out the cache-miss run on PR #5430 mid-download
(run 25950714888) before Studio install or the smoke assertions ran.

Once the actions/cache restore hits, the job comes in under 10 min,
so 45 min only costs runner time on the first run after a cache
key bump (v1->v2 was just bumped in #5459, which is what produced
this failure). Jobs 1 (openai-anthropic, 270M model) and 2
(tool-calling, ~1.5 GB model) are not bumped -- their 25 min cap
has been comfortable.
2026-05-15 20:52:36 -07:00
Daniel Han
56f9d94254
Merge branch 'fix-issue-5344-quantization-guardrail' into feat-gemma4-moe-4bit-swap 2026-05-15 20:50:27 -07:00
Daniel Han
5e90beaf69
Merge branch 'main' into fix-issue-5344-quantization-guardrail 2026-05-15 20:49:23 -07:00
Daniel Han
fb4bd0b777
ci: drop cache: 'npm' from setup-node (silent abort on Windows) (#5474)
`actions/setup-node@v6.4.0` with `cache: 'npm'` silently aborts the
entire job on Windows runners when the npm cache path returned by
`npm config get cache` (`C:\npm\cache`) does not yet exist on a fresh
runner -- the step exits 24s in with no error message and every
following step gets skipped. See npm/cli#7308 for the underlying
EEXIST / missing-dir race in the npm cache directory.

This mirrors the existing precedent in
`studio-windows-ui-smoke.yml`'s `setup-python` block, which already
dropped `cache: 'pip'` for the same reason (post-step fatal error on
a missing pip cache dir). The frontend `npm ci` is fast enough
without the cache that the reliability gain is worth the ~30s.
2026-05-15 20:49:05 -07:00
Daniel Han
cab5a88e65
Merge branch 'fix-issue-5344-quantization-guardrail' into feat-gemma4-moe-4bit-swap 2026-05-15 19:42:24 -07:00
Daniel Han
642324a41d
Merge branch 'main' into fix-issue-5344-quantization-guardrail 2026-05-15 19:41:20 -07:00
Daniel Han
77e0929735
revert: stop touching DEVICE_TYPE == "cuda" branches for CPU CI (#5473)
#5429 (cb15a7a5) tightened three production-path branches to
DEVICE_TYPE == "cuda" and torch.cuda.is_available() and added a
new else: SUPPORTS_BFLOAT16 = False arm to let `import unsloth.trainer`
survive on a CPU-only CI host. We already ship the package on Intel
XPU / AMD HIP / NVIDIA CUDA and don't want any extra branching in
those hot paths.

Move the entire CPU-CI handling to one place -- the top of
unsloth/device_type.get_device_type() -- so the UNSLOTH_ALLOW_CPU=1
sentinel short-circuits detection and returns "cuda" before any
torch probe runs. Every downstream DEVICE_TYPE == "cuda" branch
then behaves identically to a real CUDA host, with no additional
checks. The two existing duplicate UNSLOTH_ALLOW_CPU returns
later in the function are dropped (the new top-of-function check
covers both).

Revert the three call-site changes:
- unsloth/_gpu_init.py:212  -> back to `if DEVICE_TYPE == "cuda":`
- unsloth/_gpu_init.py:247  -> back to `if DEVICE_TYPE == "cuda":`
- unsloth/models/_utils.py:1207 -> back to `if DEVICE_TYPE == "cuda":`
- unsloth/_gpu_init.py: drop the new `else: SUPPORTS_BFLOAT16 = False`
  branch (dead under the top-of-function short-circuit).

Keep the two env-var gates that are needed for zoo's drift detectors
to inspect pristine TRL source (no behavioural change on production
hosts that never set UNSLOTH_ALLOW_CPU):
- unsloth/_gpu_init.py: `if env != "1": _patch_trl_trainer()`
- unsloth/models/rl.py:PatchFastRL: `if env == "1": return`

Verified:
- CUDA_VISIBLE_DEVICES=5 python -c "import unsloth.trainer" produces
  UnslothSFTTrainer.__init__ (TRL still patched on real hosts).
- UNSLOTH_ALLOW_CPU=1 + aggressive cuda spoof import succeeds and
  trl.SFTTrainer.__init__.__qualname__ stays SFTTrainer.__init__.
- pytest tests/_zoo_compiler_cache_shim.py -> 5 passed, 1 skipped.
2026-05-15 19:41:09 -07:00
Daniel Han
3948cc0522
Merge branch 'fix-issue-5344-quantization-guardrail' into feat-gemma4-moe-4bit-swap 2026-05-15 15:55:12 -07:00
Daniel Han
af63ff2eed
Merge branch 'main' into fix-issue-5344-quantization-guardrail 2026-05-15 15:54:06 -07:00
Daniel Han
e775f941a4
tests/openai: patch httpx.AsyncClient ctor so delete tests hit mock (#5469)
delete_openai_container intentionally creates a fresh
httpx.AsyncClient per call (see external_provider docstring: shared
pool produced false 'deleted: true' responses while the container
survived). The existing _mock_http_client only swapped the shared
module-level _http_client, so the four delete tests bypassed the
mock entirely and hit the real OpenAI API, returning 401
Unauthorized on Python 3.10 / 3.12 / 3.13.

Extend the helper to also monkey-patch httpx.AsyncClient itself
to a factory that injects the test's MockTransport into any
freshly constructed client. List/create paths still use the
shared client and pass unchanged.

Verified locally: pytest tests/test_openai_container_crud.py
-> 8 passed.
2026-05-15 15:53:54 -07:00
Lee Jackson
ba0cae1aff
Stop: drop Ollama API key, clean up code execution UI (#5464)
* chat: drop Ollama API key, clean up code execution UI

* studio/chat: fix undefined candidateId + keyboard a11y on container list

- Auto-bind effect referenced `candidateId`, which is not declared in
  this scope (only `candidate` is) — would fail the TS/Next build.
  Use `candidate.id` to match the variable that's actually defined.
- Container list items get `role="button"` when `canActivate` is true
  but had no keyboard activation. Add `onKeyDown` for Enter/Space and
  `tabIndex={0}` so the row is focusable and activatable from the
  keyboard, matching the existing onClick behavior.

* studio/chat: restore declarations dropped by the main merge

The 75646444d auto-merge with main (#5466) silently dropped the
declarations a4f19171c added in regions #5466 also rewrote, while
leaving the usages further down in the file. No textual conflict
markers, but the result referenced undeclared names:

- REFRESH_POLL_MS constant (drives the 30s list refresh interval).
- pendingDelete / setPendingDelete / deleting / setDeleting state
  (drives the in-sheet AlertDialog delete confirm — replaces the
  window.confirm() that landed via #5466).
- Per-row locals inside the container list .map callback: running,
  isActive (recomputed with running), ttlMinutes, canActivate,
  statusLabel (drive click-to-activate, expired/active badges, and
  the muted styling for expired containers).

Also wire setDeleting(false) + setPendingDelete(null) into the
confirmDelete finally so the AlertDialog closes after the delete
call resolves; previously the busy state never cleared.

The all-containers list now iterates sortedContainers (matches the
picker above and the "newest-active first" UX) instead of the
unsorted visibleContainers.

---------

Co-authored-by: Roland Tannous <115670425+rolandtannous@users.noreply.github.com>
Co-authored-by: Roland Tannous <rolandtannous@gravityq.ai>
2026-05-16 02:17:03 +04:00
Daniel Han
457930c7db
Merge branch 'fix-issue-5344-quantization-guardrail' into feat-gemma4-moe-4bit-swap 2026-05-15 15:11:08 -07:00
Daniel Han
f4fbe4a49f
Merge branch 'main' into fix-issue-5344-quantization-guardrail 2026-05-15 15:10:03 -07:00
Daniel Han
2de99a23d8
studio/install: strip top-level dir from repaired symlink target (#5467)
The repair in 5465 returned the full archive entry name (e.g.
"llama-b9165 libggml-rpc.0.11.1.dylib") but safe_link_target joins
the return value with target.parent (which already lives under
base llama-b9165). That doubled the prefix to
base llama-b9165 llama-b9165 libggml-rpc.0.11.1.dylib, the
resolved path never existed, and extract_tar_safely still raised
'tar archive contained unresolved link entries'.

Strip the top-level dir before returning so the linkname is
relative to target.parent, mirroring how unmangled symlinks are
stored in the tar (basename-only relative to the symlink).

Verified end-to-end against the upstream b9165 tarball: extraction
succeeds and every symlink resolves to an existing file.
2026-05-15 15:09:50 -07:00
Roland Tannous
a70bf02bb8
studio/chat: OpenAI container picker delete reliability (#5466)
* studio/chat: fix OpenAI container delete UX (expired filter, TTL cap, idempotent 404, refresh-on-error)

- Filter status="expired" from /containers/list so the picker only
  shows usable containers. OpenAI keeps expired entries in the list
  indefinitely, which made delete look broken.
- Cap ttl_minutes at 20 (backend Field + frontend TTL_MAX + persistence
  clamp). OpenAI's actual hard limit is 20; the prior 10080 cap caused
  integer_above_max_value rejections on create.
- Treat 404 on delete as idempotent success in the frontend client so
  already-gone containers don't surface a scary error toast.
- Run refresh() in finally for onCreate/onDelete so the picker stays
  in sync with OpenAI even when the call errors.
- Add route-level test for the expired filter.

* studio/chat: add diagnostic logging for OpenAI /containers DELETE

Trace what arrives at /external/openai/containers/delete (subject,
container_id, base_url) and what we send to OpenAI (URL, presence
of Authorization, value of OpenAI-Beta) plus the full response
status + body (capped at 300 chars). Helps confirm whether the
beta header is on the wire and whether OpenAI's response actually
reports deleted=true, when users report the delete "not taking".

No secrets are logged — Authorization is reported as a boolean.

* studio/chat: log raw /containers list response from OpenAI

Sibling to the delete diagnostics. After a confirmed delete
(deleted=true on the wire), we want to see whether the very next
list call returns the just-deleted id — that distinguishes
"OpenAI eventually-consistent list" from "frontend stale state".
Logs each entry's id + status only; no names, no timestamps.

* studio/chat: fingerprint decrypted API key for container CRUD

Logs kind (sk-proj-/sk-/other), length, and last-4 chars only —
never the full secret. Lets us compare what the backend actually
uses against the key the user expects, since the same DELETE
request shape can produce different results across keys
(project-scoped containers: list is permissive but delete requires
the owning project's key).

* studio/chat: use fresh httpx client for /v1/containers DELETE

Same key, same headers, same URL via the shared _http_client
returned deleted=true but the container persisted in subsequent
list calls. A fresh httpx.AsyncClient with the identical request
shape (verified with a standalone reproducer) deleted the same
container cleanly. Suspect connection-pool state from earlier
chat-completion streams interferes at the edge — switching to a
per-call client side-steps it entirely. Scoped to delete only;
list/create keep using the shared pool until we can confirm the
same fix is needed there.

* studio/chat: log OpenAI response headers on container DELETE

Adds cf-ray / x-request-id / openai-organization / openai-project /
openai-processing-ms to the delete-response diagnostic line. Lets
us cross-reference a failing delete against OpenAI support (or
against a working standalone reproducer) using the unique
request-id and edge node.

* studio/chat: client-side tombstone for just-deleted OpenAI containers

OpenAI's /v1/containers DELETE returns {"deleted": true} but the
list endpoint can keep returning the same container for several
minutes (replica lag or in-use silent no-op — undocumented per
developers.openai.com/api/docs/guides/tools-shell). Our backend
sends the correct DELETE with OpenAI-Beta: containers=v1 and a
standalone reproducer shows the same behavior, so the right fix
is UI-side rather than waiting on OpenAI.

After a successful delete, the id goes into a per-component
tombstone map with a 5-minute expiry. visibleContainers (now the
single chokepoint feeding sortedContainers, auto-bind, and the
all-containers list) filters those ids out. A 30s sweep clears
expired tombstones so the picker recovers automatically if OpenAI
eventually catches up (or the container's TTL elapses).

* studio/chat: tombstones live for the page lifetime; drop API key fingerprint log

- Tombstones change from Map<id, expiry> to Set<id>: once tombstoned,
  the id stays hidden from the picker until page reload. OpenAI's list
  can keep returning a deleted id for an undocumented and variable
  amount of time; automatically un-tombstoning after a fixed window
  surfaces it again and creates more confusion than it solves. The
  container's own TTL eventually expires the entry on OpenAI's side,
  and the expired-status filter at the backend list route hides it
  anyway.
- Remove the periodic sweep effect (dead code without expiries).
- Remove the api-key fingerprint log added during debugging — it
  served its purpose (confirmed parity) and isn't needed long-term.
2026-05-16 01:53:13 +04:00
Daniel Han
96b7497a7e
Merge branch 'fix-issue-5344-quantization-guardrail' into feat-gemma4-moe-4bit-swap 2026-05-15 14:46:09 -07:00
Daniel Han
62a36d92b8
Merge branch 'main' into fix-issue-5344-quantization-guardrail 2026-05-15 14:45:03 -07:00
Daniel Han
4f59c8e539
studio/install: repair upstream llama.cpp prebuilt mangled symlinks (#5465)
The macos-arm64 prebuilt tarball for llama.cpp b9165 and b9169 ships
symlinks whose linkname is missing both the directory separator AND
the leading character of the target basename:

  llama-b9165/libggml-rpc.0.dylib -> llama-b9165ibggml-rpc.0.11.1.dylib

extract_tar_safely correctly classified those as unresolved and made
install.sh fall back to source-build, which Mac CI then fails as a
hard error (Studio must use the prebuilt llama-bNNNN-bin-macos-arm64
on Apple Silicon).

Add _try_repair_missing_slash inside safe_link_target: when a
linkname starts with the member's top-level dir but no following
slash, search the archive for an entry under that dir whose name
ends with the mangled suffix. Accept only when the suffix uniquely
identifies a real archive entry, so legitimate archives are
untouched.

Verified against /tmp/llama-b9165.tar.gz: all 18 link entries
repair to real files in the archive.
2026-05-15 14:44:52 -07:00
Daniel Han
1a89a0b616
Merge branch 'fix-issue-5344-quantization-guardrail' into feat-gemma4-moe-4bit-swap 2026-05-15 14:19:22 -07:00
Daniel Han
f6014c32ca
Merge branch 'main' into fix-issue-5344-quantization-guardrail 2026-05-15 14:18:17 -07:00
Daniel Han
4b23af48b1
tests: raise pwsh/bash subprocess timeout from 10s to 60s (#5463)
CI surfaced a flaky failure on Linux 'Repo tests (CPU)':
  TestPwshPrForcePromotion.test_baked_in_pr_force_promotes ->
  subprocess.TimeoutExpired after 10s on /usr/bin/pwsh startup.

The scripts under test run in well under a second; the 10s budget
only covered pwsh / bash launch time, which spikes on heavily-
loaded GitHub-hosted runners. Raise the default helper timeout to
60s for both run_bash and run_pwsh. Real bugs in the script logic
will still surface as wrong output or non-zero exit; this just
absorbs runner-side launch jitter.
2026-05-15 14:18:04 -07:00
Daniel Han
48291ffab4
Merge branch 'fix-issue-5344-quantization-guardrail' into feat-gemma4-moe-4bit-swap 2026-05-15 13:15:53 -07:00
Daniel Han
466d41c1b9
Merge branch 'main' into fix-issue-5344-quantization-guardrail 2026-05-15 13:14:48 -07:00
Daniel Han
85cf0a41ea
ci: switch Windows Stop Studio to a cmd no-op marker (#5462)
The prior set +e + redirect + exit 0 fix in #5460 did not stop the
Stop Studio step from exiting 143 (SIGTERM) on Git Bash; bash on
windows-latest exits with that signal before any inline guard
runs, regardless of redirection. The teardown does not gate
correctness -- the runner reclaims the Studio child process at
job end -- so swap the shell from Git Bash to cmd and just emit
a marker line.

After this, Job 3 (JSON, images) and the two other Windows GGUF
CI jobs cannot fail at the teardown step.
2026-05-15 13:14:34 -07:00
Roland Tannous
2622b79606
studio/chat: built-in code execution for OpenAI + Anthropic (#5461)
* studio/chat: built-in code execution for Anthropic Claude 4.x

Wire Anthropic's server-side code_execution_20250825 tool to the
existing Code pill in the composer. Pill lights up only for Claude
Opus/Sonnet/Haiku 4.x models that the docs list as compatible; pairs
independently with Search. Backend appends the tool entry plus the
code-execution-2025-08-25 beta header, and translates the SSE
server_tool_use / *_tool_result blocks (bash + text_editor sub-tools)
into the _toolEvent shape the frontend renderer consumes. File
uploads via the Files API are a deliberate follow-up.

* studio/chat: enable code execution pill in in-thread composer too

thread.tsx renders its own composer with a separate CodeToolsToggle
that was still gated on supportsTools only, so the pill stayed
disabled inside an active thread even after picking Anthropic 4.x.
Surface the capability through the runtime store
(supportsBuiltinCodeExecution, set from chat-page alongside
supportsBuiltinWebSearch) and read it in the toggle.

* studio/chat: built-in code execution for OpenAI cloud gpt-5.5

Extend the Code pill to OpenAI cloud's gpt-5.5 / gpt-5.5-pro via the
shell tool on /v1/responses. Per-thread container reuse: capture the
container_id from each response on a synthetic container_ready event,
persist it onto the ThreadRecord, and pass it back as
environment.type="container_reference" on follow-up turns so the
model sees filesystem state from prior turns until OpenAI's idle
expiry. Stale ids surface a container_invalidated event that clears
the thread record so the next turn falls back to container_auto.

Gated strictly on OpenAI cloud (api.openai.com base URL) — Ollama,
llama.cpp, vLLM, and custom OpenAI-compat presets won't see the
shell tool entry even when their providerType collapses to "openai".

* studio/chat: OpenAI shell-tool container management UI

Side-panel section (settings sheet → Code Execution) for managing
OpenAI's shell-tool containers per thread. Three controls:

- New-container idle timeout (provider-level default, pre-fills the
  create dialog and is used by the lazy-create path on a thread's
  first turn when set to a non-default value).
- Active container picker for the active thread — pick any existing
  container or stay on "Auto-create per thread".
- Inline create form (name + idle TTL) and per-row delete actions.

Three new backend endpoints under /api/inference/external/openai/
containers/{list,create,delete} proxy to OpenAI /v1/containers using
the encrypted API key. All three reject non-cloud base URLs up front
so the picker stays scoped to api.openai.com.

Deleting a container clears all thread bindings pointing at it; the
next turn falls back to auto-create.

* studio/chat: inherit container across threads + styled active picker

New threads on the same OpenAI provider now default to the most
recently used container instead of "Auto-create per thread" — both
in the chat-adapter (so a send works even if the side panel was
never opened) and in the side panel itself (auto-binds the active
thread when the dropdown loads on a thread that has no container).

Picker is visually emphasized with an accent panel and the
currently-active row in the list below is highlighted with the same
accent so the two views stay in sync.

* studio/chat: friendly English-word names for auto-created containers

Replaces the "chat-<thread-id-slug>" auto-name with a random
English-word + short hex suffix (e.g. "kestrel-3f9c"). Applies only
to the chat-adapter's lazy-create path; the OpenAI container_auto
path stays unnamed (only fires when no custom TTL is set).

* studio/chat: always pre-create OpenAI containers via frontend

Drops the TTL-based gate on the chat-adapter's lazy-create path so
every code-execution container the user ever sees in the picker has
a friendly English-word name. The backend's container_auto fallback
stays as a safety net (used only if the POST /v1/containers call
fails); in practice that branch should be rare.

* studio/chat: send OpenAI-Beta header for /v1/containers CRUD

Without OpenAI-Beta: containers=v1, OpenAI returns 200
{"deleted": true} for DELETE /v1/containers/{id} but does not
actually remove the container. The list call then keeps returning it,
making it look like Studio's "Delete container" button is broken.

Verified 2026-05-15 against api.openai.com: DELETE with the beta
header returns 200 and removes the container; the same DELETE without
the header returns the same 200 deleted:true body but the container
stays alive.

- Add _container_headers() that merges OpenAI-Beta on top of the
  shared auth headers; route list / create / delete through it.
- Verify the DELETE response body reports {"deleted": true}; raise
  httpx.HTTPError otherwise so the route surfaces a 5xx instead of
  silently reporting success on a silent no-op.
- Add tests covering header propagation and the deleted-flag guard
  (true, false, missing key, non-JSON body, 4xx passthrough).

* studio/chat: surface unpersisted-thread picker no-op as a toast

The "Active for this thread" container picker uses
db.threads.update(activeThreadId, ...), which silently returns 0 rows
affected when the thread record isn't yet in IndexedDB. That happens
on a brand-new thread where the user toggles code execution on and
opens settings before sending the first message — the chat adapter
only materializes the thread row on first send. The picker would
appear to ignore the user's selection and snap back to "Auto-create
per thread".

- onPick now awaits the update and toasts an actionable hint
  ("Send a message first to pin a container to this thread.") when
  the update affected zero rows.
- Auto-bind effect comment clarifies why it stays best-effort silent.

The auto-bind effect itself is unchanged: it's a heuristic that
should not nag the user when it can't apply.

* studio/chat: let user pick OpenAI container before first send

Previously the picker silently no-op'd until the user sent the first
message, because Dexie's ThreadRecord is only materialized inside the
runtime-provider's `initialize` hook (assistant-ui's first-message
callback). That kept users from binding a thread to an existing
OpenAI container up front; they had to either send a message and
risk the chat adapter auto-creating one, or accept the cross-thread
inheritance default.

- Export `ensureThreadRecord` from runtime-provider so other surfaces
  can materialize the row idempotently.
- In OpenAICodeExecSection.onPick, await ensureThreadRecord before
  the update, with modelType="base" (the settings sheet that hosts
  this section is only rendered in single-thread mode).

Behaviour after this commit:
- New thread + user picks a container in the sidebar → thread row is
  created with that container_id; first send uses it, no auto-create.
- New thread + user does nothing → row still absent; first send goes
  through the existing inherit/lazy-create path as before.
- The auto-bind effect remains silent best-effort: it does not
  eagerly create the thread row, so it cannot pre-empt the user's
  pick on a fresh thread.

* studio/chat: drop "Auto-create per thread" option, default to latest

The dropdown previously offered "Auto-create per thread" as an
explicit value (null in storage), with the chat-adapter then
inheriting from the most recent container at send-time. That made
the picker display disagree with what the backend would actually do:
the picker said "auto", but the backend was reusing an existing
container.

Behaviour after this commit, when code execution is enabled on an
OpenAI cloud provider:
- Containers list non-empty: dropdown defaults to the container with
  the latest lastActiveAt, eagerly bound via ensureThreadRecord +
  db.threads.update so the bind survives even when the thread row
  has not been materialized by the chat adapter yet. User can pick
  any other container in the list.
- Containers list empty: render a disabled placeholder "(none yet —
  will be created on first send)". The chat-adapter's lazy-create
  path (chat-adapter.ts:1040-1082) mints the first container on
  first send and writes it back to the thread; the next refresh
  surfaces it in the picker.

Expiration mid-operation is unchanged: the existing
container_invalidated _toolEvent clears the thread's stored id and
the next turn re-creates.

* studio/chat: fix picker stuck on "Selecting most recent…" + manual-create binding

Two follow-up fixes to the picker rework in d0cbeb99b.

1) The dropdown was getting stuck on the "Selecting most recent…"
   placeholder option even after the auto-bind write completed,
   because the select was controlled by `activeContainerId` (whatever
   sits in Dexie) and there's a brief window between the auto-bind
   firing and useLiveQuery propagating the new row back. Decoupled
   the rendered value from the Dexie state: compute the displayed id
   locally as `activeContainerId ?? sortedContainers[0]?.id`, so the
   most-recent container's name shows up immediately. The auto-bind
   effect still writes the bind to Dexie so the chat adapter sees it
   on send. Dropped the placeholder option entirely.

2) The manual "Create container" flow (`onCreate`) bound the new
   container to the active thread with a bare `db.threads.update`.
   On a brand-new thread that hadn't been materialized yet, the
   update affected 0 rows; the user's next send then went through
   cross-thread inheritance / lazy-create and could land on a stale
   container, surfacing as "container does not exist". Same fix as
   `onPick`: ensureThreadRecord before update so the bind lands.
2026-05-15 23:39:06 +04:00
Lee Jackson
a9b8c9a221
Studio: make API key optional for local providers (llama.cpp/vLLM/Ollama) (#5457)
* make API key optional for local providers (llama.cpp/vLLM/Ollama)D

* chore: reduce comments

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-15 23:33:22 +04:00
Daniel Han
3fa2ccb4f0
Merge branch 'fix-issue-5344-quantization-guardrail' into feat-gemma4-moe-4bit-swap 2026-05-15 11:48:11 -07:00
Daniel Han
514a17d95e
Merge branch 'main' into fix-issue-5344-quantization-guardrail 2026-05-15 11:47:06 -07:00
Daniel Han
ac3e9e98f2
ci: make Windows Stop Studio teardown tolerate Git Bash signal exit (#5460)
The Windows-runner "Stop Studio" step's kill + sleep block has
been observed to exit 143 (SIGTERM) even when the upstream test
work passed. Most recently caught on PR #5432 Job 3 "JSON, images":
all four assertions (json_object, plain inference, image/openai,
image/anthropic) printed PASS, then the kill step ran for ~2
seconds and exited 143, failing the job.

Teardown does not gate correctness. Wrap all three Stop Studio
steps with set +e + redirected error streams + explicit exit 0
so transient Git Bash signal weirdness no longer masks a green
test run.
2026-05-15 11:46:52 -07:00
Daniel Han
91d3e6925d
Merge branch 'fix-issue-5344-quantization-guardrail' into feat-gemma4-moe-4bit-swap 2026-05-15 11:03:34 -07:00
Daniel Han
71f5d7e547
Merge branch 'main' into fix-issue-5344-quantization-guardrail 2026-05-15 11:02:29 -07:00
Daniel Han
90ac4c87f7
ci: stop a partial mmproj cache from poisoning Mac Studio GGUF CI (#5459)
The "JSON, images" Mac Studio GGUF CI job hit a stale cache for
${{ runner.os }}-gguf-...-mmproj-F16.gguf-v1 that contains only the
main GGUF, not the mmproj sibling. cache-hit==true so the download
step was skipped, then the post-load \`ls\` failed:
  ls: ...gguf-cache/mmproj-F16.gguf: No such file or directory

Three guards layered:

1) Bump cache key v1 -> v2 to invalidate the poisoned entry on the
   GitHub-hosted side.
2) New verify-cache step explicitly checks BOTH files are present
   before trusting cache-hit. If not, fall through to download.
3) Save step gains a hashFiles() check on the mmproj path so a
   partial mmproj download cannot land back in the cache.

Behaviour on a clean run is unchanged; cache hit + verify ok skips
the re-download, partial-hit triggers fresh download, success
saves a complete archive.
2026-05-15 11:02:16 -07:00
DoubleMathew
3596ce12df
Restore Flash > SDPA > Flex priority for non-gemma3 models (#5455)
* update attn preferences

* address gemini review suggestion

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Datta Nimmaturi <venkatadattasainimmaturi@gmail.com>
2026-05-15 12:43:49 -05:00
Daniel Han
5753728eeb
Merge branch 'fix-issue-5344-quantization-guardrail' into feat-gemma4-moe-4bit-swap 2026-05-15 10:39:01 -07:00
Daniel Han
a187a18581
Merge branch 'main' into fix-issue-5344-quantization-guardrail 2026-05-15 10:37:51 -07:00
Daniel Han
51dd5fac79
ci: add tx >=5,<6 slow compile model_types to KNOWN_BROKEN_COMPILE (#5458)
The per-model SIGALRM cap landed on the previous fix now exposes
beit / sam / sam_hq as compile-too-slow on transformers >=5,<6 +
trl >=1,<2 -- each exceeds the 60s per-model budget. They are
real slow paths in unsloth_compile_transformers's source rewriter
when handling beit / SAM's encoder layers on the new transformers
line, not infra flakes (the prior fix logged sweep progress per
25 models so the slow ones are pinpointable in CI logs).

Bucket them into Category F (compile exceeds budget) so the sweep
stays green and each is tracked for follow-up zoo fixes in the
same shape as the existing 27 known-broken entries. Surface
behaviour stays identical: any NEW slow model_type still fails
the cell with a TimeoutError tag.
2026-05-15 10:37:37 -07:00
Daniel Han
763a6936cc
Merge branch 'fix-issue-5344-quantization-guardrail' into feat-gemma4-moe-4bit-swap 2026-05-15 09:38:50 -07:00
Daniel Han
11af7dbb31
Merge branch 'main' into fix-issue-5344-quantization-guardrail 2026-05-15 09:37:40 -07:00
Daniel Han
c7c3840b5f
ci: cap each compiler-sweep iteration with SIGALRM + log progress (#5456)
Core (HF=latest + TRL=latest) (transformers >=5,<6, trl >=1,<2) hangs
30+ minutes in the compiler-sweep test under the new shim layout,
exceeding the 35-min job timeout and showing up as cancelled with no
log of which model_type wedged. unsloth_compile_transformers does
real source rewriting + torch.compile decoration and can deadlock
inside a single problem model on a new transformers point release.

Per-model SIGALRM cap (60s) so one infinite-loop model_type cannot
wedge the whole sweep. Print sweep progress every 25 models so the
log surfaces the slow model_type the next time this regresses --
crucial for finding the upstream/transformers compile bug.

Timeout errors land in the same KNOWN / NEW_FAILURES bucket as any
other compile exception, so the matrix still surfaces real
regressions instead of silently absorbing them.
2026-05-15 09:37:26 -07:00