unsloth/studio/backend/core/inference
Daniel Han bc00a8e797 Serialize the GPU handoffs, gate DiT training on a GPU, and keep 3.9 installable
Six review findings, three of them evict-then-fail orderings:

- The chat load reclaimed the GPU without telling the arbiter it existed. A
  chat load holds no llama-server process until its GGUF has downloaded,
  which is minutes, so a competing Images/Video acquire in that window
  found nothing to cancel, took the GPU, and the chat load then spawned
  onto the same device. It now registers an in-flight marker through
  acquire_for's register hook (under the arbiter lock, as the image and
  video loads do), the evictor cancels a marked load, and the route undoes
  itself if ownership moved while it loaded.
- The Hub-download conflict check ran after that handoff, so a GGUF the
  download manager already owns destroyed the resident Images/Video
  pipeline and then 409'd, having loaded nothing. It moves above the
  handoff, together with the marker it handshakes with.
- The image load released the engine router's transition lock before
  registering the load, so a second load choosing the other engine could
  unload the still-idle engine this one captured; the load then landed on a
  deactivated engine, where generate, status, unload and the arbiter's
  evictor can no longer reach it. Registration now happens under that lock
  and refuses if the engine changed.
- Training a DiT family on a host with no GPU was accepted: nf4 is not a
  CPU fallback, its 4-bit load goes through bitsandbytes, which requires
  CUDA, XPU or MPS. The start unloaded the working Images pipeline, pulled
  the text encoders, and only then died in the child. Rejected before the
  teardown now, and /info stops advertising a precision that always 400s.
  SDXL keeps its documented fp32-on-CPU path.
- Both diffusion pages kept the routed-pick marker forever, so re-picking
  the same checkpoint (after chat evicted it) neither loaded nor cleared
  the query string. The marker is released once the query is gone. The
  Images key also carried a stray NUL byte, which made the file read as
  binary to grep and other tooling.
- diffusers dropped Python 3.9 in 0.38, so the unconditional >=0.39.0 pin
  left pip no candidate at all on 3.9 and made every install that composes
  the huggingface extras unresolvable there. The floor is conditional now.

Also fixes tests that were already red on the branch: two hand-built
request fakes had gone stale against fields this branch added, and the
handoff-ordering test only failed on a host with fewer than two GPUs.
2026-07-26 14:46:02 +00:00
..
sandbox_site Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
__init__.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
_html_to_md.py Studio: stream live tool output with SSE heartbeats, fix web page extraction, and surface interrupted turns (#7083) 2026-07-15 08:41:00 -07:00
_vulkan_probe.py Studio: add Vulkan llama.cpp support (#5819) 2026-07-09 03:39:48 -07:00
anthropic_compat.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
api_monitor.py Studio: trim serving-log noise and surface llama-server engine stats (#6377) 2026-06-17 05:37:57 -07:00
audio_codecs.py Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
chat_eos.py Studio: stop chat generation on the assistant-turn-end token (fixes Qwen3.5 loop) (#6804) 2026-07-06 10:07:56 -07:00
chat_template_helpers.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
chat_templates.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
defaults.py Add DeepSeek-V4-Flash-GGUF to Studio with none/high/max reasoning (#6908) 2026-07-07 06:13:43 -07:00
diffusion.py Add the diffusion download plan endpoint 2026-07-26 04:56:16 -07:00
diffusion_arch_patches.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_attention.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_auto_policy.py Add the HiDream-I1 family to the image backend 2026-07-17 13:24:57 +00:00
diffusion_batched.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-25 09:00:00 +00:00
diffusion_cache.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_compile_cache.py Batch diffusion inference with per-image seeds, an inference conditioning cache, and GGUF loader fixes 2026-07-22 07:03:38 +00:00
diffusion_cond_cache.py Namespace the trainer conditioning cache per checkpoint, bound the learning rate 2026-07-26 11:51:13 +00:00
diffusion_controlnet.py Fix invalid-UTF-8 500s, the flat Canny map and the dropped DiT knobs 2026-07-25 19:59:44 -07:00
diffusion_device.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_eager_patches.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_engine_router.py Serialize the GPU handoffs, gate DiT training on a GPU, and keep 3.9 installable 2026-07-26 14:46:02 +00:00
diffusion_families.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-25 09:00:00 +00:00
diffusion_gguf_compile.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_hidream.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-25 09:00:00 +00:00
diffusion_ideogram4.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_inference_info.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_krea2.py Host pre-cast fp8 text encoders for four more families 2026-07-18 10:24:51 +00:00
diffusion_lora.py Support LoRA adapters on torchao int8/fp8 quantized image pipelines 2026-07-17 09:37:17 +00:00
diffusion_memory.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-25 09:00:00 +00:00
diffusion_patch_backend.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_precision.py Report the fp8-cast compute dtype without swapping the encoder class 2026-07-18 10:25:02 +00:00
diffusion_prequant.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-25 09:00:00 +00:00
diffusion_speed.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_te_prequant.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-25 09:00:00 +00:00
diffusion_transformer_quant.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-25 09:00:00 +00:00
external_provider.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
gpu_arbiter.py Serialize the GPU handoffs, gate DiT training on a GPU, and keep 3.9 installable 2026-07-26 14:46:02 +00:00
image_gallery.py Studio: gate gallery serve/export on ownership; keep image progress active until persisted; reserve diffusion training before the dataset scan 2026-07-13 14:40:18 +00:00
inference.py Add Intel XPU support to Unsloth Studio (#4724) 2026-07-24 02:22:07 -03:00
key_exchange.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
llama_admission.py Studio: queue local GGUF OpenAI-compatible requests before llama-server (#7047) 2026-07-10 17:05:48 -03:00
llama_cpp.py Serialize the GPU handoffs, gate DiT training on a GPU, and keep 3.9 installable 2026-07-26 14:46:02 +00:00
llama_http.py fix(studio/llama_cpp): disable trust_env on the loopback health probe (#6750) (#6752) 2026-06-30 19:09:26 +02:00
llama_keepwarm.py Merge origin/main into image-generation (PR #6763) 2026-07-25 00:34:38 -07:00
llama_server_args.py persist llama.cpp KV cache across idle auto-unload (slot save/restore) (#7204) 2026-07-20 00:12:42 -07:00
llama_stats.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
local_model_resolver.py Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
mcp_client.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
mcp_config_import.py studio: show MCP "Import config" on the add-server form (#6030) 2026-06-11 16:17:22 +01:00
message_content.py fix(studio): handle multimodal list content in inference text paths (#4383) (#6480) 2026-06-23 01:26:11 -07:00
mlx_inference.py Studio: reuse MLX prompt cache across turns instead of re-prefilling (#7311) 2026-07-22 02:35:33 -07:00
model_ids.py studio: list the full local model catalog from /v1/models (#6519) 2026-06-26 20:42:06 -03:00
orchestrator.py Add Intel XPU support to Unsloth Studio (#4724) 2026-07-24 02:22:07 -03:00
passthrough_healing.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
presence_penalty.py Studio: apply presence_penalty on the safetensors and MLX inference paths (#6923) 2026-07-06 22:24:47 -07:00
pricing.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
providers.py Allow API key for Ollama connections (#7173) 2026-07-18 22:47:00 -07:00
runtime_context.py Expose runtime context length for hub models (#6154) 2026-06-11 22:13:53 +03:00
safetensors_agentic.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
sd_cpp_args.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
sd_cpp_backend.py Batch diffusion inference with per-image seeds, an inference conditioning cache, and GGUF loader fixes 2026-07-22 07:03:38 +00:00
sd_cpp_engine.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
sd_cpp_server.py Tighten comments across the remaining image stack files 2026-07-12 11:40:05 +00:00
stt_ggml_sidecar.py Studio: add local speech-to-text dictation engine (#7095) 2026-07-23 01:39:03 -07:00
stt_sidecar.py Studio STT: only load safetensors weights for custom dictation models (RCE fix) (#7364) 2026-07-23 03:15:45 -07:00
tensor_fallback.py studio: deterministic VRAM auto-fit for GGUF (MTP reserve, compute buffer, total-based budget) (#6312) 2026-06-17 03:10:22 -07:00
tool_call_parser.py Studio: Inkling support fixes (#7153) 2026-07-15 11:22:38 -07:00
tool_loop_controller.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
tool_stream_exec.py Studio: stream live tool output with SSE heartbeats, fix web page extraction, and surface interrupted turns (#7083) 2026-07-15 08:41:00 -07:00
tools.py fix(studio): support hostname-based enterprise proxies (#7416) 2026-07-26 02:53:00 +01:00
video.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-26 12:40:35 +00:00
video_families.py Correct the ltx-2 resident TE estimate to the bf16 cast size 2026-07-18 08:00:25 +00:00
video_gallery.py Fix invalid-UTF-8 500s, the flat Canny map and the dropped DiT knobs 2026-07-25 19:59:44 -07:00
video_ltx2.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-25 09:00:00 +00:00
worker.py Add Intel XPU support to Unsloth Studio (#4724) 2026-07-24 02:22:07 -03:00