unsloth/studio/backend/core/inference
Daniel Han 30094835d5 Only retry the unsloth import where it can succeed
The gate still let the retry run on hosts unsloth does not support, which is
where it is most harmful: a 7 GB macOS runner lost the Studio server 26 s into a
load, and the Linux runner was torn down mid-generation. Neither MPS nor plain
CPU can complete the import, so the retry there pays the cost and fails anyway.

Require an accelerator unsloth actually supports (CUDA/ROCm via torch.cuda, or
XPU), with UNSLOTH_ALLOW_CPU as the documented override, and hoist the predicate
to module level so it is tested directly rather than through the import system.
On a CPU-only host the retry no longer fires at all; on CUDA the clean-environment
case it was added for still passes 29/29.
2026-07-27 07:23:29 +00:00
..
sandbox_site Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
__init__.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
_html_to_md.py Studio: stream live tool output with SSE heartbeats, fix web page extraction, and surface interrupted turns (#7083) 2026-07-15 08:41:00 -07:00
_vulkan_probe.py Studio: add Vulkan llama.cpp support (#5819) 2026-07-09 03:39:48 -07:00
anthropic_compat.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
api_monitor.py Studio: trim serving-log noise and surface llama-server engine stats (#6377) 2026-06-17 05:37:57 -07:00
audio_codecs.py Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
chat_eos.py Studio: stop chat generation on the assistant-turn-end token (fixes Qwen3.5 loop) (#6804) 2026-07-06 10:07:56 -07:00
chat_template_helpers.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
chat_templates.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
defaults.py Add DeepSeek-V4-Flash-GGUF to Studio with none/high/max reasoning (#6908) 2026-07-07 06:13:43 -07:00
diffusion.py Fix diffusion policy and classification issues from review 2026-07-27 06:51:41 +00:00
diffusion_arch_patches.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_attention.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-27 07:18:14 +00:00
diffusion_auto_policy.py Fix diffusion policy and classification issues from review 2026-07-27 06:51:41 +00:00
diffusion_batched.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_cache.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_compile_cache.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_cond_cache.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_controlnet.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_device.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_eager_patches.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_engine_router.py Bound the gallery blob cache, and three interlock fixes 2026-07-27 05:08:45 +00:00
diffusion_families.py Fix diffusion policy and classification issues from review 2026-07-27 06:51:41 +00:00
diffusion_gguf_compile.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_hidream.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_ideogram4.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_inference_info.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_krea2.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_lora.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_memory.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_patch_backend.py Only retry the unsloth import where it can succeed 2026-07-27 07:23:29 +00:00
diffusion_precision.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_prequant.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_speed.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_te_prequant.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-27 04:29:44 +00:00
diffusion_transformer_quant.py Fix diffusion policy and classification issues from review 2026-07-27 06:51:41 +00:00
external_provider.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
gpu_arbiter.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
image_gallery.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
inference.py Add Intel XPU support to Unsloth Studio (#4724) 2026-07-24 02:22:07 -03:00
key_exchange.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
llama_admission.py Studio: queue local GGUF OpenAI-compatible requests before llama-server (#7047) 2026-07-10 17:05:48 -03:00
llama_cpp.py Merge branch 'main' of https://github.com/unslothai/unsloth into r6763 2026-07-27 07:15:43 +00:00
llama_http.py fix(studio/llama_cpp): disable trust_env on the loopback health probe (#6750) (#6752) 2026-06-30 19:09:26 +02:00
llama_keepwarm.py Merge origin/main into image-generation (PR #6763) 2026-07-25 00:34:38 -07:00
llama_server_args.py persist llama.cpp KV cache across idle auto-unload (slot save/restore) (#7204) 2026-07-20 00:12:42 -07:00
llama_stats.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
local_model_resolver.py Add interactive Agents command builder (#7312) 2026-07-26 17:09:19 -07:00
mcp_client.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
mcp_config_import.py studio: show MCP "Import config" on the add-server form (#6030) 2026-06-11 16:17:22 +01:00
message_content.py fix(studio): handle multimodal list content in inference text paths (#4383) (#6480) 2026-06-23 01:26:11 -07:00
mlx_inference.py Studio: reuse MLX prompt cache across turns instead of re-prefilling (#7311) 2026-07-22 02:35:33 -07:00
model_ids.py studio: list the full local model catalog from /v1/models (#6519) 2026-06-26 20:42:06 -03:00
orchestrator.py Add Intel XPU support to Unsloth Studio (#4724) 2026-07-24 02:22:07 -03:00
passthrough_healing.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
presence_penalty.py Studio: apply presence_penalty on the safetensors and MLX inference paths (#6923) 2026-07-06 22:24:47 -07:00
pricing.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
providers.py Allow API key for Ollama connections (#7173) 2026-07-18 22:47:00 -07:00
runtime_context.py Expose runtime context length for hub models (#6154) 2026-06-11 22:13:53 +03:00
safetensors_agentic.py Studio: default tool-call permission to Approve for me, prompt only on high-risk actions (#7285) 2026-07-26 17:07:31 -07:00
sd_cpp_args.py Keep the sd.cpp text encoder on CPU under Metal 2026-07-27 06:58:36 +00:00
sd_cpp_backend.py Bound the gallery blob cache, and three interlock fixes 2026-07-27 05:08:45 +00:00
sd_cpp_engine.py Bound the gallery blob cache, and three interlock fixes 2026-07-27 05:08:45 +00:00
sd_cpp_server.py Bound the gallery blob cache, and three interlock fixes 2026-07-27 05:08:45 +00:00
stt_ggml_sidecar.py Studio: add local speech-to-text dictation engine (#7095) 2026-07-23 01:39:03 -07:00
stt_sidecar.py Studio STT: only load safetensors weights for custom dictation models (RCE fix) (#7364) 2026-07-23 03:15:45 -07:00
tensor_fallback.py studio: deterministic VRAM auto-fit for GGUF (MTP reserve, compute buffer, total-based budget) (#6312) 2026-06-17 03:10:22 -07:00
tool_call_parser.py Studio: Inkling support fixes (#7153) 2026-07-15 11:22:38 -07:00
tool_loop_controller.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
tool_stream_exec.py Studio: stream live tool output with SSE heartbeats, fix web page extraction, and surface interrupted turns (#7083) 2026-07-15 08:41:00 -07:00
tools.py Studio: add Deep Research (#7219) 2026-07-26 23:36:02 -07:00
video.py Fix diffusion policy and classification issues from review 2026-07-27 06:51:41 +00:00
video_families.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
video_gallery.py Make the OpenAI image URL fetchable, keep WebM audio, stream example imports 2026-07-26 23:39:49 +00:00
video_ltx2.py Stop staging the dense text encoder for an fp8 video load 2026-07-27 04:28:12 +00:00
web_access_policy.py Studio: add Deep Research (#7219) 2026-07-26 23:36:02 -07:00
worker.py Add Intel XPU support to Unsloth Studio (#4724) 2026-07-24 02:22:07 -03:00