unsloth/studio/backend/tests
Daniel Han ab56e4a9ed
Studio: serialise GGUF reload and inherit unsloth-run extra args (#5427)
* Studio: serialise GGUF reload and inherit unsloth-run extra args

Closes #5401.

Three related GGUF reload bugs reproduced against `unsloth studio run -m unsloth/Qwen3-0.6B-GGUF --gguf-variant Q4_K_M --top-k 20 --seed 42`:

1. The `POST /api/inference/load` already-loaded short-circuit only compared `model_identifier` and `hf_variant`. A same-(model, variant) Apply that flipped `cache_type_kv` / `speculative_type` / `chat_template_override` / `max_seq_length` / `llama_extra_args` returned `status="already_loaded"` and the new setting silently never reached llama-server.

2. The frontend chat-settings Apply path POSTs `/unload` then `/load` without round-tripping `llama_extra_args`. Every reload after `unsloth run --some-flag X` quietly dropped `--some-flag X` from the spawned `llama-server` command line.

3. `LlamaCppBackend.load_model` released `_lock` between Phase 1 (kill) and Phase 3 (spawn) so two concurrent loads each passed Phase 1 with `self._process is None`. Both ran Phase 2 (download), both reached Phase 3, and the Phase 3 defensive `_kill_process()` from #5171 collapsed them to one survivor only after both `subprocess.Popen` calls had landed. For the 86 GB MoE in #5161 / the model in #5401 the overlap window was tens of seconds, long enough to OOM the host. With a 0.6B model the pgrep timeline showed two simultaneous PIDs for 3.3 s on `main`.

Fix:

`studio/backend/core/inference/llama_cpp.py`

* Add `self._serial_load_lock = threading.Lock()`. The whole body of `load_model` runs under this lock so two concurrent `/api/inference/load` requests are strictly sequential. The fine-grained `_lock` and the Phase 3 defensive `_kill_process()` from #5171 are kept as a second layer. `/unload`, `/status`, and `/load-progress` are unaffected because they only touch the fine-grained lock or read properties.
* Add `self._extra_args` plus an `extra_args` property, written inside `load_model` whenever the caller supplies a non-`None` value. `unload_model()` deliberately does not reset it so the route layer can inherit the args across the frontend's `/unload` + `/load` gap.

`studio/backend/routes/inference.py`

* Add `_request_matches_loaded_settings(request, llama_backend)` that compares `max_seq_length`, `cache_type_kv`, `speculative_type`, `chat_template_override`, and `llama_extra_args` between the incoming request and the live backend. Same-(model, variant) requests whose runtime settings differ now fall through to a real reload instead of returning `already_loaded`. A missing `llama_extra_args` field on the request is treated as "inherit current", so the short-circuit still fires when the only difference is the frontend not echoing the CLI flags back.
* GGUF load branch inherits `llama_extra_args` from `llama_backend.extra_args` when the request omits the field, re-validates through `validate_extra_args`, and forwards the result to `load_model(...)`. An explicit `[]` from the caller is still honoured as "clear".

Verified end to end against a live `unsloth studio run` instance:

| Scenario                                                        | Before    | After                                                                    |
| --------------------------------------------------------------- | --------- | ------------------------------------------------------------------------ |
| `/load` same (model, variant, settings)                         | 1 PID, `already_loaded` | unchanged                                                                |
| `/load` same model, variant, new `cache_type_kv=q8_0` ctx=8192  | `already_loaded`, settings dropped | `loaded`, `/status` reports the new settings, new server has `-c 8192 --cache-type-k q8_0 --top-k 20 --seed 42` |
| Frontend Apply `/unload` + `/load`, new settings, no `llama_extra_args` field | Drops `--top-k 20 --seed 42` | Preserves `--top-k 20 --seed 42`                                          |
| `/unload` + two parallel `/load`                                | Two PIDs for 3.3 s | Max simultaneous count = 1 across the full pgrep timeline                  |
| `/load` with `llama_extra_args=[]` (explicit clear)             | n/a       | `loaded`, new server has no `--top-k` / `--seed`                          |
| `/load` with `llama_extra_args=["--top-k","30","--seed","7"]` (override) | n/a       | `loaded`, new server has the supplied flags                                |

`pytest studio/backend/tests` is green except for one pre-existing terminal-width-sensitive assertion (`test_studio_api.py::test_help_output`) and the pre-existing `test_studio_api.py` fixture errors that fail on unmodified main too. No new regressions.

* Studio: track requested n_ctx so Auto-slider flips trigger a reload

Review feedback on PR #5427 from gemini-code-assist.

The original short-circuit compared ``request.max_seq_length`` against
``llama_backend.context_length`` (the effective context). VRAM-fit
logic can cap the running server below what the caller asked for, so
this comparison incorrectly returns ``already_loaded`` when the user
flips the slider from an explicit length (e.g. 8192) back to "Auto"
(0): the explicit request was capped to, say, 4096, and the new "Auto"
request reads ``backend.context_length == 4096`` and decides nothing
changed.

Track the originally requested ``n_ctx`` on the backend instead and
compare against that. ``requested_n_ctx == 0`` means the last load
asked for the model's native length; ``request.max_seq_length == 0``
matches it.

Verified in the sandbox suite (now 90 tests):

- ``test_explicit_to_auto_triggers_reload`` -- loaded with explicit
  8192, then Apply with ``max_seq_length=0`` falls through to a real
  reload and the new server runs at the native 40960.
- ``test_auto_to_explicit_triggers_reload`` -- inverse direction.
- ``test_explicit_to_same_explicit_short_circuits`` -- re-Apply with
  the same explicit value still short-circuits (no needless reload).
- Existing scenarios (kv change, spec change, template change, extra
  args inherit, parallel-load stress, frontend Apply flow) unchanged.

``pytest studio/backend/tests`` still green on the same set of tests;
the pre-existing ``test_help_output`` failure and ``test_studio_api``
fixture errors are unaffected.

* Studio: tighten comments in the 5401 fix

Trim the verbose explanatory comments and docstrings introduced in
f9cbec3b and dd0b1d58 down to one-line summaries. The "why" still
points at issue #5401; the multi-paragraph rationale belonged in the
PR body, not the source. No behaviour change.

* ci: retrigger after zoo drift + IPython fixes landed in main

* ci: retrigger Mac Studio UI CI after transient fetch flake

* Studio: address six P2 followups on the 5401 reload PR

Tightens the inheritance and serial-load paths to close the six P2
findings raised by codex-connector on PR #5427 against `f9cbec3b` /
`dd0b1d58`.

1. Re-check loaded state before killing queued loads. Two duplicate
   `/api/inference/load` requests both pass the route-level
   `is_loaded` gate before the first publishes `_healthy = True`. The
   second waits on `_serial_load_lock`, enters Phase 1, and tears down
   the just-spawned llama-server for a redundant full reload. Added
   `LlamaCppBackend._already_in_target_state(...)` and a short-circuit
   at the top of the serial-lock block: if the live server already
   satisfies the kwargs, return True without killing.

2. Don't inherit CLI overrides that shadow new first-class settings.
   `unsloth run -c 4096` is a permitted pass-through; the validator
   docs explicitly call out `-c`/`--ctx-size`. Stored in `_extra_args`
   and appended after Studio's own flags, the inherited `-c 4096`
   silently won the last-wins parse against a new
   `max_seq_length=8192`. Added `strip_shadowing_flags` in
   `llama_server_args.py` (covers `-c`, `--cache-type-k/v`, `--spec-*`,
   `--chat-template*`, `--jinja`/`--no-jinja`) and the route runs the
   inherited list through it before validate + forward.

3. Restrict inherited llama args to the same GGUF model. `_extra_args`
   is deliberately preserved across `unload_model()` for the chat-
   settings Apply flow (`/unload` + `/load` with no `llama_extra_args`
   field). Now also track `_extra_args_source = (model_identifier,
   hf_variant)` so the route can refuse cross-model inheritance.
   `LlamaCppBackend.extra_args_source` exposes the tuple.

4. Persist extras only after a successful load. `_extra_args` was
   written at the top of `load_model` before Popen + health check, so
   a failed startup left bad args in place to poison the next UI
   retry. The write (along with `_requested_n_ctx`) is now deferred
   until after `_healthy = True`.

5. Ignore speculative diffs for vision loads. `load_model` silently
   gates speculative decoding on `not is_vision`, so the backend's
   `_speculative_type` stays `None` for vision models. The route's
   comparator now normalises the request's value to `"off"` when
   `llama_backend.is_vision` to avoid a no-op reload of a vision
   server every time the dropdown defaults to `default`. The
   `_already_in_target_state` helper applies the same rule.

6. Wait for the replacement server before short-circuiting. `_kill_process`
   did not clear `_healthy`; the new first-class settings
   (`_cache_type_kv`, `_speculative_type`, `_chat_template_override`)
   are written under `_lock` BEFORE Popen + `_wait_for_health`. A
   duplicate `/load` arriving during the new server's warm-up window
   could short-circuit against the not-yet-healthy replacement and the
   caller would start inference against a server that was still
   loading. `_kill_process` now sets `_healthy = False` in its
   `finally` block so `is_loaded` returns False from the moment the
   old server is killed until the new one finishes warm-up.

Tests:

- Sandbox suite under `./temp/sim_5401/` extended to 136 tests (was
  90): new unit coverage for `strip_shadowing_flags` (12 cases),
  `_kill_process` clears `_healthy`, `extra_args_source` lifecycle and
  cross-model behaviour, failed-load preserving prior extras, and the
  duplicate-load short-circuit at `load_model` level. New live
  integration cases verify shadow-strip via `pgrep` on the live
  llama-server cmdline, cross-model refusal, and PID stability across
  a duplicate-load race. All 136 pass.
- `pytest studio/backend/tests --deselect test_studio_api.py`:
  1079 passed, 46 skipped, identical to the pre-change count. The
  pre-existing `test_studio_api.py` fixture errors and the
  terminal-width-sensitive `test_help_output` are unaffected.
- Ruff: clean on the three modified files.

* Studio: tighten GGUF reload inheritance and duplicate-load guard

Re-narrow llama_extra_args to None after validate_extra_args when the
incoming request omitted the field, so the backend can distinguish
"caller omitted, inherit prior load" from "caller explicitly cleared
to []". Without this a queued duplicate /load reaches the backend as
[] and fails _already_in_target_state's exact-equality check, killing
the just-started llama-server. The pass-through validate call from
the original "forward llama-server args from unsloth studio run /
unsloth run" change is preserved as-is; only the post-pass narrowing
is new. Cross-source loads now explicitly clear extras so a model
switch can't accidentally inherit via the backend's "no opinion"
semantics.

Store the caller's hf_variant kwarg (None for local GGUF files) in
_extra_args_source instead of the derived self._hf_variant
(an extracted filename quant label like "Q4_K_M"). Same-source check
in the route is now symmetric for HF and direct-file loads.

Add gguf_path to _already_in_target_state and prefer on-disk path
identity when both backend and caller have a path. This stops the
duplicate-load guard from killing a healthy server on repeat local
loads (where hf_variant is None on the caller side but extracted on
the backend side).

Split shadow-flag stripping into per-group toggles (context / cache /
spec / template). The route now opts into stripping only the groups
whose first-class field was actually set on the incoming request, so
an inherited --chat-template-file survives an Apply that omits
chat_template_override. _request_matches_loaded_settings detects
shadowing extras on the inherit path and falls through to a real
reload so the strip can run.

Mark --spec-default, --jinja, --no-jinja as boolean inside the
shadow stripper so the value-consuming heuristic no longer eats the
following positional token.

* Studio: trim comments around GGUF reload inheritance

* Studio: cover GGUF reload inheritance and shadow-flag stripping

* Studio: drop redundant issue refs from inheritance comments

* Studio: drop redundant issue refs from inheritance comments

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: key inheritance source off resolved gguf_variant

codex-connector P2 on PR #5427 cd14cae1: the inheritance gate at
``routes/inference.py:696`` compared the stored ``source[1]`` against
``request.gguf_variant``, but the HF branch loaded with
``hf_variant = config.gguf_variant`` (the *resolved* variant after
ModelConfig auto-pick). When the caller omitted ``gguf_variant`` on a
follow-up Apply, ``source[1] == "Q4_K_M"`` but
``(request.gguf_variant or "") == ""``, ``same_source`` returned False,
and the chat-settings Apply silently dropped CLI pass-through flags
for every auto-pick / local-file load.

Fix both sides of the comparison to key off ``config.gguf_variant``:

* The route compares ``source[1]`` to ``config.gguf_variant`` (the
  resolved label) rather than the request field.
* The local-mode load_model call now passes
  ``hf_variant = config.gguf_variant`` so ``_extra_args_source``
  stores the same string the route reads back. The HF branch already
  did this.

Sandbox: added test_source_records_caller_variant_not_extracted_label
to lock the storage key contract.
``pytest studio/backend/tests --deselect test_studio_api.py``:
1100 passed, identical to pre-change.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: deny upstream --ui family on llama-server pass-through

The validator's web-UI block named only ``--webui`` / ``--no-webui``,
which is llama.cpp's pre-rename spelling. Current upstream
(``tools/server/README.md``) uses ``--ui`` / ``--no-ui`` plus
``--ui-config``, ``--ui-config-file``, and ``--ui-mcp-proxy`` /
``--no-ui-mcp-proxy``. Without these in the denylist a user could
``unsloth run --ui`` and enable llama-server's built-in web UI on
the port Studio's reverse proxy targets, breaking the UI surface.

Keep the legacy ``--webui`` group so the validator still rejects
old binaries that haven't been re-spelled.

Cross-referenced against the README's full flag list; this was the
only gap for the post-#5401 inheritance / shadow-strip work. Pass-
through flags from every other README category (sampling, jinja,
ctx, cache, threads, GPU, reasoning, grammar, chat-template-kwargs)
already validate cleanly; sandbox suite exercises ~60 of them in
the new ``test_08_llama_server_pass_through.py``.

``pytest studio/backend/tests --deselect test_studio_api.py``:
1100 passed, identical to pre-change.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-17 04:16:09 -07:00
..
__init__.py Final cleanup 2026-03-12 18:28:04 +00:00
conftest.py Studio: Expose openai and anthropic compatible external API end points (#4956) 2026-04-13 21:08:11 +04:00
test_anthropic_code_execution.py studio/chat: built-in code execution for OpenAI + Anthropic (#5461) 2026-05-15 23:39:06 +04:00
test_anthropic_messages.py Studio: support images on /v1/messages (Anthropic-compat) (#5128) 2026-04-22 03:25:07 +04:00
test_anthropic_thinking_translation.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
test_browse_folders_route.py Studio: add folder browser modal for Custom Folders (#5035) 2026-04-15 08:04:33 -07:00
test_cache_case_resolution.py Add tests for cache case resolution (from PR #4822) (#4823) 2026-04-03 13:58:26 -07:00
test_cached_gguf_routes.py Studio: support GGUF variant selection for non-suffixed repos (#5023) 2026-04-15 15:32:01 +04:00
test_data_recipe_github_progress.py Studio: add github_repo seed reader and GitHub Support Bot recipe (#5169) 2026-04-24 12:02:03 -07:00
test_data_recipe_seed.py fix(seed): disable remote code execution in seed inspect dataset loads (#4275) 2026-03-13 19:37:43 +04:00
test_desktop_auth.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
test_detect_mmproj_file.py fix(studio/mmproj): block cross-family projectors in flat local GGUF dirs (#5347) (#5350) 2026-05-14 20:31:20 -07:00
test_export_log_cursor.py studio: stream export worker output into the export dialog (#4897) 2026-04-14 08:55:43 -07:00
test_gguf_metadata.py fix(studio/mmproj): block cross-family projectors in flat local GGUF dirs (#5347) (#5350) 2026-05-14 20:31:20 -07:00
test_gguf_reload_inheritance.py Studio: serialise GGUF reload and inherit unsloth-run extra args (#5427) 2026-05-17 04:16:09 -07:00
test_gpu_selection.py Update VRAM estimator to cater to broader model configs (#5175) 2026-05-05 04:12:36 -07:00
test_gpu_selection_sandbox.py [Studio] multi gpu finetuning/inference via "balanced_low0/sequential" device_map (#4602) 2026-03-30 02:33:15 -07:00
test_host_defaults.py Default Studio host to 127.0.0.1 and prompt before auto-start (#5267) 2026-05-04 13:03:16 +04:00
test_inference_model_validation.py Studio: Dark theme refactor, right sidebar redesign, and chat UI polish (#5150) 2026-05-07 14:33:31 +04:00
test_kv_cache_estimation.py fix KVCache estimates for gemma4 style sliding window models (#5225) 2026-05-05 04:06:46 -07:00
test_llama_cpp_cache_aware_disk_check.py Studio: make GGUF disk-space preflight cache-aware (#5012) 2026-04-14 08:53:37 -07:00
test_llama_cpp_context_fit.py Studio: pin GPU at 95% headroom and warn on silent CPU fallback (#5323) 2026-05-13 04:48:15 -07:00
test_llama_cpp_load_progress.py Studio: live model-load progress + rate/ETA on download and load (#5017) 2026-04-14 09:46:22 -07:00
test_llama_cpp_load_progress_live.py Studio: split model-load progress label across two rows (#5020) 2026-04-14 10:58:16 -07:00
test_llama_cpp_load_progress_matrix.py Studio: split model-load progress label across two rows (#5020) 2026-04-14 10:58:16 -07:00
test_llama_cpp_max_context_threshold.py fix KVCache estimates for gemma4 style sliding window models (#5225) 2026-05-05 04:06:46 -07:00
test_llama_cpp_no_context_shift.py Studio: hard-stop at n_ctx with a 'Context limit reached' toast (#5021) 2026-04-14 10:58:20 -07:00
test_llama_cpp_windows_nvidia_path.py Studio: add torch's pip nvidia DLL dirs to PATH on Windows (#5324) 2026-05-11 05:42:09 -07:00
test_llama_server_args.py Studio: serialise GGUF reload and inherit unsloth-run extra args (#5427) 2026-05-17 04:16:09 -07:00
test_log_filter_no_truncation.py Studio: stop truncating long log lines as suspected base64 (#5335) 2026-05-08 13:07:18 +04:00
test_middleware.py studio: security and hardening pass (auth rate-limit, sandbox, path containment, schema validation, headers) (#5375) 2026-05-13 06:12:18 -07:00
test_mlx_inference_backend.py MLX training support for Studio on Apple Silicon (#5340) 2026-05-14 05:24:20 -07:00
test_mlx_training_worker_config.py studio: skip flash-attn install on Blackwell GPUs (sm_100+) (#5420) 2026-05-14 18:13:50 +04:00
test_models_get_model_config_case_resolution.py Add tests for cache case resolution (from PR #4822) (#4823) 2026-04-03 13:58:26 -07:00
test_native_context_length.py Studio: Fix chat template disappearing after browser refresh (#5209) 2026-05-01 08:19:09 -07:00
test_openai_code_execution.py studio/chat: built-in code execution for OpenAI + Anthropic (#5461) 2026-05-15 23:39:06 +04:00
test_openai_container_crud.py tests/openai: patch httpx.AsyncClient ctor so delete tests hit mock (#5469) 2026-05-15 15:53:54 -07:00
test_openai_responses_translation.py Studio: o3 reasoning summary payload (#5426) 2026-05-15 17:13:28 +04:00
test_openai_tool_passthrough.py studio: security and hardening pass (auth rate-limit, sandbox, path containment, schema validation, headers) (#5375) 2026-05-13 06:12:18 -07:00
test_providers_api.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
test_pytorch_mirror.py Add configurable PyTorch mirror via UNSLOTH_PYTORCH_MIRROR env var (#5024) 2026-04-15 11:39:11 +04:00
test_responses_api.py Studio: Expose openai and anthropic compatible external API end points (#4956) 2026-04-13 21:08:11 +04:00
test_responses_tool_passthrough.py Studio: forward standard OpenAI tools / tool_choice on /v1/responses (Codex compat) (#5122) 2026-04-21 13:17:20 +04:00
test_sandbox_tools.py studio: security and hardening pass (auth rate-limit, sandbox, path containment, schema validation, headers) (#5375) 2026-05-13 06:12:18 -07:00
test_studio_api.py Studio: forward standard OpenAI tools / tool_choice to llama-server (#5099) 2026-04-18 12:53:23 +04:00
test_studio_train_validation.py studio: security and hardening pass (auth rate-limit, sandbox, path containment, schema validation, headers) (#5375) 2026-05-13 06:12:18 -07:00
test_tool_policy_gates.py unsloth run: add --enable-tools/--disable-tools server-side tool policy (#5277) 2026-05-05 12:45:15 +04:00
test_tool_policy_state.py unsloth run: add --enable-tools/--disable-tools server-side tool policy (#5277) 2026-05-05 12:45:15 +04:00
test_trained_model_scan.py studio: security and hardening pass (auth rate-limit, sandbox, path containment, schema validation, headers) (#5375) 2026-05-13 06:12:18 -07:00
test_training_history_update.py Studio: Dark theme refactor, right sidebar redesign, and chat UI polish (#5150) 2026-05-07 14:33:31 +04:00
test_training_raw_support.py studio: drop unused max_grad_value schema + route plumbing (#5424) 2026-05-14 05:43:58 -07:00
test_training_worker_flash_attn.py studio: skip flash-attn install on Blackwell GPUs (sm_100+) (#5420) 2026-05-14 18:13:50 +04:00
test_transformers_version.py split venv_t5 into tiered 5.3.0/5.5.0 and fix trust_remote_code (#4878) 2026-04-07 20:05:01 +04:00
test_utils.py Add AMD ROCm/HIP support across installer and hardware detection (#4720) 2026-04-10 01:56:12 -07:00
test_vision_cache.py Studio: split vision-cache exception test to match transient vs permanent (#5145) 2026-04-23 00:22:40 -07:00
test_vram_estimation.py Update VRAM estimator to cater to broader model configs (#5175) 2026-05-05 04:12:36 -07:00