unsloth/studio/backend/tests
Daniel Han d79fd92798
studio: scope cancel-cleanup to in-flight tmp dirs; walk back tool_call_id (#5488)
* studio: scope cancel-cleanup to in-flight tmp dirs; walk back tool_call_id

Two follow-ups to #5375's training and chat hardening.

_cleanup_cancelled_checkpoints used to rmtree every checkpoint-N
directory on Cancel. That is the opposite of what the user expects.
A user cancelling an 8h run with save_steps=2000 loses every
completed checkpoint they could have resumed from. The 67 MB residue
the audit memo flagged is the HF Trainer atomic-rename partial
(tmp-checkpoint-N), not the completed ones. The cleanup now targets
only tmp-checkpoint subdirs; completed checkpoint-N directories are
user-owned and stay. Symlinked output_dir and symlinked children are
skipped so the realpath containment cannot be levered into deleting
arbitrary content via a symlink trick.

ChatMessage._validate_role_shape stamped a random secrets.token_hex
id on tool messages with no tool_call_id. That id is uncorrelated
with the prior assistant tool_calls id, so strict passthrough
backends (OpenAI, Anthropic) reject the request as orphaned and
llama.cpp treats the tool result as "no preceding call" and
hallucinates. The synthesis moves up to ChatCompletionRequest, where
the whole conversation is visible: for each tool message missing an
id we walk back to the most recent assistant turn with tool_calls
(stopping at user turns), prefer a function.name match, otherwise
take the first unconsumed tool_call. Synthesis is the fallback when
no candidate assistant turn exists, preserving the prior round-trip
guarantee for orphaned tool messages.

Tests:
  - test_cleanup_cancelled_checkpoints.py (new): pins that completed
    checkpoint subdirs survive, tmp-checkpoint partials are removed,
    non-int suffixes (checkpoint-final, checkpoint-best) are left
    alone, output_dir outside outputs_root is refused, symlinked
    output_dir and symlinked child are both skipped, missing dir is
    a no-op.
  - test_inference_model_validation.py: 6 new walkback cases covering
    name-match preference, first-unconsumed fallback, explicit-id
    passthrough, multi-tool-result pairing, synth-on-no-parent, and
    no-cross-user-turn invariant.
  - test_openai_tool_passthrough.py: the two ChatMessage-level
    synth-on-missing tests are rewritten to assert that the per-
    message validator now leaves tool_call_id untouched; resolution
    coverage lives in the request-level tests above.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: explicit tool_call_id reserve, numeric tmp-checkpoint suffix only

Reviewer follow-ups to the training-cleanup + tool_call_id walkback PR.

tool_call_id walkback: a mixed assistant turn with [call_a, call_b]
followed by a tool result that carried tool_call_id="call_a" and a
sibling tool result with no id resolved to ['call_a', 'call_a']
because the explicit id never reserved call_a in the consumed set.
Added a pre-pass over the message list that walks back from every
role="tool" message carrying an explicit id and marks the matching
(asst_idx, tc_idx) consumed, then the missing-id walkback runs against
that pre-populated set. The second result now resolves to call_b.

While here, also harden the function-shape check: if a provider
ships a malformed tool_call where `function` is a string rather than
a dict, the old `(tc.get("function") or {}).get("name")` raised
AttributeError on the string's .get; now isinstance-gated so the
walkback falls through to the fallback id without raising.

Cancel cleanup: `tmp-checkpoint-*` is too broad. HF Trainer's
in-flight partials are always `tmp-checkpoint-<integer-step>`, so
constrain the cleanup regex to `^tmp-checkpoint-\d+$`. A user folder
named `tmp-checkpoint-final`, `tmp-checkpoint-backup`, or
`tmp-checkpoint-user-notes` is now preserved.

ChatMessage docstring still pointed at the pre-PR contract that
required `tool_call_id` on every role="tool" message. Updated to say
missing ids are accepted at message scope and resolved at
ChatCompletionRequest scope. Inline comment above the cancel-cleanup
call now describes the actual behaviour (in-flight tmp partials,
completed checkpoints preserved).

Test:
  - python -m pytest studio/backend/tests/test_inference_model_validation.py
    studio/backend/tests/test_cleanup_cancelled_checkpoints.py
    studio/backend/tests/test_openai_tool_passthrough.py -q
    -> 76 passed (was 67 before this commit; +2 walkback regression
       tests, +1 numeric-suffix preservation test)

* studio: trim verbose comments in cleanup + tool_call_id walkback

Move the HF tmp-checkpoint regex to module scope as a named constant.
Drop the multi-paragraph docstring on _cleanup_cancelled_checkpoints
and the inline call-site rationale; the function name + the test
class already cover the why.

Compress _resolve_missing_tool_call_ids docstring from a six-line
explanation to two. Same logic, fewer in-flow tutorials.

76 tests in cleanup + inference-model-validation + tool-passthrough pass.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-18 00:01:48 -07:00
..
__init__.py Final cleanup 2026-03-12 18:28:04 +00:00
conftest.py Studio: Expose openai and anthropic compatible external API end points (#4956) 2026-04-13 21:08:11 +04:00
test_anthropic_code_execution.py studio/chat: built-in code execution for OpenAI + Anthropic (#5461) 2026-05-15 23:39:06 +04:00
test_anthropic_messages.py Studio: support images on /v1/messages (Anthropic-compat) (#5128) 2026-04-22 03:25:07 +04:00
test_anthropic_thinking_translation.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
test_browse_folders_route.py Studio: add folder browser modal for Custom Folders (#5035) 2026-04-15 08:04:33 -07:00
test_cache_case_resolution.py Add tests for cache case resolution (from PR #4822) (#4823) 2026-04-03 13:58:26 -07:00
test_cached_gguf_routes.py Studio: support GGUF variant selection for non-suffixed repos (#5023) 2026-04-15 15:32:01 +04:00
test_cleanup_cancelled_checkpoints.py studio: scope cancel-cleanup to in-flight tmp dirs; walk back tool_call_id (#5488) 2026-05-18 00:01:48 -07:00
test_data_recipe_github_progress.py Studio: add github_repo seed reader and GitHub Support Bot recipe (#5169) 2026-04-24 12:02:03 -07:00
test_data_recipe_seed.py fix(seed): disable remote code execution in seed inspect dataset loads (#4275) 2026-03-13 19:37:43 +04:00
test_desktop_auth.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
test_detect_mmproj_file.py fix(studio/mmproj): block cross-family projectors in flat local GGUF dirs (#5347) (#5350) 2026-05-14 20:31:20 -07:00
test_export_log_cursor.py studio: stream export worker output into the export dialog (#4897) 2026-04-14 08:55:43 -07:00
test_gguf_metadata.py fix(studio/mmproj): block cross-family projectors in flat local GGUF dirs (#5347) (#5350) 2026-05-14 20:31:20 -07:00
test_gguf_reload_inheritance.py Studio: serialise GGUF reload and inherit unsloth-run extra args (#5427) 2026-05-17 04:16:09 -07:00
test_gpu_selection.py Update VRAM estimator to cater to broader model configs (#5175) 2026-05-05 04:12:36 -07:00
test_gpu_selection_sandbox.py [Studio] multi gpu finetuning/inference via "balanced_low0/sequential" device_map (#4602) 2026-03-30 02:33:15 -07:00
test_host_defaults.py Default Studio host to 127.0.0.1 and prompt before auto-start (#5267) 2026-05-04 13:03:16 +04:00
test_inference_model_validation.py studio: scope cancel-cleanup to in-flight tmp dirs; walk back tool_call_id (#5488) 2026-05-18 00:01:48 -07:00
test_kv_cache_estimation.py fix KVCache estimates for gemma4 style sliding window models (#5225) 2026-05-05 04:06:46 -07:00
test_llama_cpp_cache_aware_disk_check.py Studio: make GGUF disk-space preflight cache-aware (#5012) 2026-04-14 08:53:37 -07:00
test_llama_cpp_context_fit.py Studio: pin GPU at 95% headroom and warn on silent CPU fallback (#5323) 2026-05-13 04:48:15 -07:00
test_llama_cpp_load_progress.py Studio: live model-load progress + rate/ETA on download and load (#5017) 2026-04-14 09:46:22 -07:00
test_llama_cpp_load_progress_live.py Studio: split model-load progress label across two rows (#5020) 2026-04-14 10:58:16 -07:00
test_llama_cpp_load_progress_matrix.py Studio: split model-load progress label across two rows (#5020) 2026-04-14 10:58:16 -07:00
test_llama_cpp_max_context_threshold.py fix KVCache estimates for gemma4 style sliding window models (#5225) 2026-05-05 04:06:46 -07:00
test_llama_cpp_no_context_shift.py Studio: hard-stop at n_ctx with a 'Context limit reached' toast (#5021) 2026-04-14 10:58:20 -07:00
test_llama_cpp_windows_nvidia_path.py Studio: add torch's pip nvidia DLL dirs to PATH on Windows (#5324) 2026-05-11 05:42:09 -07:00
test_llama_server_args.py Studio: serialise GGUF reload and inherit unsloth-run extra args (#5427) 2026-05-17 04:16:09 -07:00
test_log_filter_no_truncation.py Studio: stop truncating long log lines as suspected base64 (#5335) 2026-05-08 13:07:18 +04:00
test_middleware.py studio: expose launcher capability bits on unauth /api/health (#5486) 2026-05-17 21:29:24 -07:00
test_mlx_inference_backend.py MLX training support for Studio on Apple Silicon (#5340) 2026-05-14 05:24:20 -07:00
test_mlx_training_worker_config.py studio: skip flash-attn install on Blackwell GPUs (sm_100+) (#5420) 2026-05-14 18:13:50 +04:00
test_models_get_model_config_case_resolution.py Add tests for cache case resolution (from PR #4822) (#4823) 2026-04-03 13:58:26 -07:00
test_native_context_length.py Studio: Fix chat template disappearing after browser refresh (#5209) 2026-05-01 08:19:09 -07:00
test_offline_gguf_cache_fallback.py studio: load cached GGUF models when fully offline (#5505) 2026-05-17 21:25:39 -07:00
test_openai_code_execution.py studio/chat: built-in code execution for OpenAI + Anthropic (#5461) 2026-05-15 23:39:06 +04:00
test_openai_container_crud.py tests/openai: patch httpx.AsyncClient ctor so delete tests hit mock (#5469) 2026-05-15 15:53:54 -07:00
test_openai_responses_translation.py Studio: o3 reasoning summary payload (#5426) 2026-05-15 17:13:28 +04:00
test_openai_tool_passthrough.py studio: scope cancel-cleanup to in-flight tmp dirs; walk back tool_call_id (#5488) 2026-05-18 00:01:48 -07:00
test_providers_api.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
test_pytorch_mirror.py Add configurable PyTorch mirror via UNSLOTH_PYTORCH_MIRROR env var (#5024) 2026-04-15 11:39:11 +04:00
test_recommended_folders_permission.py Fix /recommended-folders 500 on unreadable model directories (Python 3.12+) (#5523) 2026-05-18 00:16:14 +04:00
test_responses_api.py Studio: Expose openai and anthropic compatible external API end points (#4956) 2026-04-13 21:08:11 +04:00
test_responses_tool_passthrough.py Studio: forward standard OpenAI tools / tool_choice on /v1/responses (Codex compat) (#5122) 2026-04-21 13:17:20 +04:00
test_sandbox_tools.py studio: tighten sandbox blocklist precision (bash, hf upload, NOFILE) (#5487) 2026-05-18 00:01:17 -07:00
test_studio_api.py Studio: forward standard OpenAI tools / tool_choice to llama-server (#5099) 2026-04-18 12:53:23 +04:00
test_studio_train_validation.py studio: security and hardening pass (auth rate-limit, sandbox, path containment, schema validation, headers) (#5375) 2026-05-13 06:12:18 -07:00
test_tool_policy_gates.py unsloth run: add --enable-tools/--disable-tools server-side tool policy (#5277) 2026-05-05 12:45:15 +04:00
test_tool_policy_state.py unsloth run: add --enable-tools/--disable-tools server-side tool policy (#5277) 2026-05-05 12:45:15 +04:00
test_trained_model_scan.py studio: security and hardening pass (auth rate-limit, sandbox, path containment, schema validation, headers) (#5375) 2026-05-13 06:12:18 -07:00
test_training_history_update.py Studio: Dark theme refactor, right sidebar redesign, and chat UI polish (#5150) 2026-05-07 14:33:31 +04:00
test_training_raw_support.py studio: drop unused max_grad_value schema + route plumbing (#5424) 2026-05-14 05:43:58 -07:00
test_training_worker_flash_attn.py studio: skip flash-attn install on Blackwell GPUs (sm_100+) (#5420) 2026-05-14 18:13:50 +04:00
test_transformers_version.py split venv_t5 into tiered 5.3.0/5.5.0 and fix trust_remote_code (#4878) 2026-04-07 20:05:01 +04:00
test_utils.py Add AMD ROCm/HIP support across installer and hardware detection (#4720) 2026-04-10 01:56:12 -07:00
test_vision_cache.py Studio: split vision-cache exception test to match transient vs permanent (#5145) 2026-04-23 00:22:40 -07:00
test_vram_estimation.py Update VRAM estimator to cater to broader model configs (#5175) 2026-05-05 04:12:36 -07:00