unsloth/studio/backend/core/inference
Daniel Han 3f6057a2b2 Bound the gallery blob cache, and three interlock fixes
Four review findings, all reproduced first:

- The gallery object-URL caches were unbounded. A clip runs from a few MB to a
  few hundred, both pages stay mounted after their first visit, and entries were
  only dropped on delete, so scrolling pinned everything for the session. Both
  pages now share a byte-budgeted LRU (512 MB video / 192 MB images) keyed off
  the visibility signal the near-viewport fetching already provides. On-screen
  media, the selected clip or image, and the item just fetched are never
  evicted, so eviction is invisible and a single item larger than the whole
  budget cannot evict itself into a refetch loop.

- The image, video and chat load guards ran two independent training probes but
  returned early when the FIRST one raised, so an unreadable LLM backend
  disabled the diffusion interlock and a load could proceed straight into an
  active diffusion trainer on the same GPU. The probes are independent now.

- An engine switch swallowed a failed teardown and published the new engine
  anyway, which is exactly the leak the unload exists to prevent: the arbiter's
  evictor, /images/unload and the next load all resolve through
  get_active_diffusion_engine(), so the still-resident pipeline (or a live
  sd-server) became unreachable and the next load allocated on top of it. The
  switch now fails and leaves the old engine published, so it stays reclaimable.

- The native generation timeout was 30 minutes while the Images page waits up to
  6 hours (SETTLE_MAX_MS), so slow-but-progressing CPU jobs died deterministically
  at the deadline. Measured on GPU-less runners, a 512x512 4-step Q2_K generation
  took 900 s on Linux and 1465 s on Windows, so larger images or step counts clear
  half an hour easily. The ceiling now matches the page's window and applies to
  the whole request: chunks of a split batch share one deadline instead of each
  getting a full budget. Cancellation is unchanged.

Declined: gating the huggingfacenotorch extra off Python 3.9 over the
conditional diffusers marker. The marker is deliberate and its comment says why:
diffusers dropped 3.9 in 0.38, so pinning >=0.39 outright leaves pip no candidate
and the whole extra unresolvable there. The pipelines it names live in
studio/backend, which cannot install on 3.9 anyway (studio.txt pins
matplotlib==3.10.9 and fastmcp>=3.0.2, both requires_python >=3.10), and the
extra is the general core one, so the alternative drops 3.9 for library users who
never touch Studio.
2026-07-27 05:08:45 +00:00
..
sandbox_site
__init__.py
_html_to_md.py
_vulkan_probe.py
anthropic_compat.py
api_monitor.py
audio_codecs.py
chat_eos.py
chat_template_helpers.py
chat_templates.py
defaults.py
diffusion.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_arch_patches.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_attention.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_auto_policy.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_batched.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_cache.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_compile_cache.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_cond_cache.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_controlnet.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_device.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_eager_patches.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_engine_router.py Bound the gallery blob cache, and three interlock fixes 2026-07-27 05:08:45 +00:00
diffusion_families.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_gguf_compile.py
diffusion_hidream.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_ideogram4.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_inference_info.py
diffusion_krea2.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_lora.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_memory.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_patch_backend.py
diffusion_precision.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_prequant.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_speed.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
diffusion_te_prequant.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-27 04:29:44 +00:00
diffusion_transformer_quant.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
external_provider.py
gpu_arbiter.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
image_gallery.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
inference.py
key_exchange.py
llama_admission.py
llama_cpp.py Merge remote-tracking branch 'origin/main' into r6763_mainmerge 2026-07-27 02:49:24 +00:00
llama_http.py
llama_keepwarm.py Merge origin/main into image-generation (PR #6763) 2026-07-25 00:34:38 -07:00
llama_server_args.py
llama_stats.py
local_model_resolver.py Add interactive Agents command builder (#7312) 2026-07-26 17:09:19 -07:00
mcp_client.py
mcp_config_import.py
message_content.py
mlx_inference.py
model_ids.py
orchestrator.py
passthrough_healing.py
presence_penalty.py
pricing.py
providers.py
runtime_context.py
safetensors_agentic.py Studio: default tool-call permission to Approve for me, prompt only on high-risk actions (#7285) 2026-07-26 17:07:31 -07:00
sd_cpp_args.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
sd_cpp_backend.py Bound the gallery blob cache, and three interlock fixes 2026-07-27 05:08:45 +00:00
sd_cpp_engine.py Bound the gallery blob cache, and three interlock fixes 2026-07-27 05:08:45 +00:00
sd_cpp_server.py Bound the gallery blob cache, and three interlock fixes 2026-07-27 05:08:45 +00:00
stt_ggml_sidecar.py
stt_sidecar.py
tensor_fallback.py
tool_call_parser.py
tool_loop_controller.py
tool_stream_exec.py
tools.py Studio: default tool-call permission to Approve for me, prompt only on high-risk actions (#7285) 2026-07-26 17:07:31 -07:00
video.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-27 04:29:44 +00:00
video_families.py Trim the comments across the diffusion backend 2026-07-26 20:31:19 +00:00
video_gallery.py Make the OpenAI image URL fetchable, keep WebM audio, stream example imports 2026-07-26 23:39:49 +00:00
video_ltx2.py Stop staging the dense text encoder for an fp8 video load 2026-07-27 04:28:12 +00:00
worker.py