unsloth/studio/backend/core/inference
Daniel Han a8cfba2eb2 Gate the unsloth retry in the diffusion patch backend
The retry added for the clean-environment patch failures is not free: importing
unsloth pulls torch in behind it, which costs ~940 MB of RSS measured in a
process that had neither, and on a host with no accelerator it fails anyway. A
cross-platform CI job that had generated fine at ~900 s later died 19 s in with
SIGTERM and every 'if: always()' step skipped, which is the runner being torn
down rather than a step failing.

Retry only when torch is already imported (true of the server and of anything
patching a real module, and the condition that stops the retry from being what
loads torch), unsloth is installed but not yet imported, and the first failure
was the ImportError the sentinel guard raises. The clean-environment case it was
added for still passes 29/29.
2026-07-27 07:15:52 +00:00
..
sandbox_site Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
__init__.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
_html_to_md.py Studio: stream live tool output with SSE heartbeats, fix web page extraction, and surface interrupted turns (#7083) 2026-07-15 08:41:00 -07:00
_vulkan_probe.py Studio: add Vulkan llama.cpp support (#5819) 2026-07-09 03:39:48 -07:00
anthropic_compat.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
api_monitor.py Studio: trim serving-log noise and surface llama-server engine stats (#6377) 2026-06-17 05:37:57 -07:00
audio_codecs.py Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
chat_eos.py Studio: stop chat generation on the assistant-turn-end token (fixes Qwen3.5 loop) (#6804) 2026-07-06 10:07:56 -07:00
chat_template_helpers.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
chat_templates.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
defaults.py Add DeepSeek-V4-Flash-GGUF to Studio with none/high/max reasoning (#6908) 2026-07-07 06:13:43 -07:00
diffusion.py Fix diffusion policy and classification issues from review 2026-07-27 06:51:41 +00:00
diffusion_arch_patches.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_attention.py Fix diffusion policy and classification issues from review 2026-07-27 06:51:41 +00:00
diffusion_auto_policy.py Fix diffusion policy and classification issues from review 2026-07-27 06:51:41 +00:00
diffusion_batched.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_cache.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_compile_cache.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_cond_cache.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_controlnet.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_device.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_eager_patches.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_engine_router.py Bound the gallery blob cache, and three interlock fixes 2026-07-27 05:08:45 +00:00
diffusion_families.py Fix diffusion policy and classification issues from review 2026-07-27 06:51:41 +00:00
diffusion_gguf_compile.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_hidream.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_ideogram4.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_inference_info.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_krea2.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_lora.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_memory.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_patch_backend.py Gate the unsloth retry in the diffusion patch backend 2026-07-27 07:15:52 +00:00
diffusion_precision.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_prequant.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_speed.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_te_prequant.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-27 04:29:44 +00:00
diffusion_transformer_quant.py Fix diffusion policy and classification issues from review 2026-07-27 06:51:41 +00:00
external_provider.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
gpu_arbiter.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
image_gallery.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
inference.py Add Intel XPU support to Unsloth Studio (#4724) 2026-07-24 02:22:07 -03:00
key_exchange.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
llama_admission.py Studio: queue local GGUF OpenAI-compatible requests before llama-server (#7047) 2026-07-10 17:05:48 -03:00
llama_cpp.py Merge branch 'main' of https://github.com/unslothai/unsloth into r6763 2026-07-27 07:15:43 +00:00
llama_http.py fix(studio/llama_cpp): disable trust_env on the loopback health probe (#6750) (#6752) 2026-06-30 19:09:26 +02:00
llama_keepwarm.py Merge origin/main into image-generation (PR #6763) 2026-07-25 00:34:38 -07:00
llama_server_args.py persist llama.cpp KV cache across idle auto-unload (slot save/restore) (#7204) 2026-07-20 00:12:42 -07:00
llama_stats.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
local_model_resolver.py Add interactive Agents command builder (#7312) 2026-07-26 17:09:19 -07:00
mcp_client.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
mcp_config_import.py studio: show MCP "Import config" on the add-server form (#6030) 2026-06-11 16:17:22 +01:00
message_content.py fix(studio): handle multimodal list content in inference text paths (#4383) (#6480) 2026-06-23 01:26:11 -07:00
mlx_inference.py Studio: reuse MLX prompt cache across turns instead of re-prefilling (#7311) 2026-07-22 02:35:33 -07:00
model_ids.py studio: list the full local model catalog from /v1/models (#6519) 2026-06-26 20:42:06 -03:00
orchestrator.py Add Intel XPU support to Unsloth Studio (#4724) 2026-07-24 02:22:07 -03:00
passthrough_healing.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
presence_penalty.py Studio: apply presence_penalty on the safetensors and MLX inference paths (#6923) 2026-07-06 22:24:47 -07:00
pricing.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
providers.py Allow API key for Ollama connections (#7173) 2026-07-18 22:47:00 -07:00
runtime_context.py Expose runtime context length for hub models (#6154) 2026-06-11 22:13:53 +03:00
safetensors_agentic.py Studio: default tool-call permission to Approve for me, prompt only on high-risk actions (#7285) 2026-07-26 17:07:31 -07:00
sd_cpp_args.py Keep the sd.cpp text encoder on CPU under Metal 2026-07-27 06:58:36 +00:00
sd_cpp_backend.py Bound the gallery blob cache, and three interlock fixes 2026-07-27 05:08:45 +00:00
sd_cpp_engine.py Bound the gallery blob cache, and three interlock fixes 2026-07-27 05:08:45 +00:00
sd_cpp_server.py Bound the gallery blob cache, and three interlock fixes 2026-07-27 05:08:45 +00:00
stt_ggml_sidecar.py Studio: add local speech-to-text dictation engine (#7095) 2026-07-23 01:39:03 -07:00
stt_sidecar.py Studio STT: only load safetensors weights for custom dictation models (RCE fix) (#7364) 2026-07-23 03:15:45 -07:00
tensor_fallback.py studio: deterministic VRAM auto-fit for GGUF (MTP reserve, compute buffer, total-based budget) (#6312) 2026-06-17 03:10:22 -07:00
tool_call_parser.py Studio: Inkling support fixes (#7153) 2026-07-15 11:22:38 -07:00
tool_loop_controller.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
tool_stream_exec.py Studio: stream live tool output with SSE heartbeats, fix web page extraction, and surface interrupted turns (#7083) 2026-07-15 08:41:00 -07:00
tools.py Studio: add Deep Research (#7219) 2026-07-26 23:36:02 -07:00
video.py Fix diffusion policy and classification issues from review 2026-07-27 06:51:41 +00:00
video_families.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
video_gallery.py Make the OpenAI image URL fetchable, keep WebM audio, stream example imports 2026-07-26 23:39:49 +00:00
video_ltx2.py Stop staging the dense text encoder for an fp8 video load 2026-07-27 04:28:12 +00:00
web_access_policy.py Studio: add Deep Research (#7219) 2026-07-26 23:36:02 -07:00
worker.py Add Intel XPU support to Unsloth Studio (#4724) 2026-07-24 02:22:07 -03:00