unsloth/studio/backend/core/inference
Daniel Han 514f4c60fe Harden the video speed stack: cache quality pin, device identity, transactional caches, quant safety
- Wan2.2-A14B step cache: pin the balanced FBCache threshold to 0.08 even when
  quant is active (per-family override in diffusion_cache.py). Auto-fp8 made the
  generic quant promotion (0.12) the family's effective default at pairwise LPIPS
  0.128, over the 0.08 quality gate the balanced preset is held to. Measured
  operating point with fp8 actually engaged (1280x720/81f/50 steps, B200):
  fb@0.08 = 1.08x at 0.129 vs the old fb@0.12 = 2.58x at 0.181; documented in
  the preset table. Explicit thresholds and the fast preset are unaffected.

- MagCache curves: validated the shipped 33-frame calibrations at the production
  121-frame default for hunyuanvideo-1.5-720p, hunyuanvideo-1.5 (480p) and
  wan2.2-ti2v-5b. Fresh 121-frame calibrations differ by <= 0.024 max abs entry
  and produce byte-identical frames at the auto presets (hv720 quality 1.69x at
  LPIPS 0.042, hv480 quality 1.66x at 0.018, wan5b balanced 1.74x at 0.026, all
  pairwise vs the same-load uncached stack), so the curves ship unchanged with
  the frame-count transfer documented next to them.

- Dual-GPU CFG parallelism: the secondary-device pick now prefers a device whose
  name and compute capability match the primary, and the gate declines a
  mismatched pair in auto mode (eager kernel selection is arch-dependent, so the
  advertised bit-identity cannot hold across different GPU models); an explicit
  cfg_parallel=on proceeds but is downgraded to lossless=False with a warning.

- A14B expert step cache is now all-or-none, mirroring the transactional quant
  loop: a mixed outcome (cache engaged on one expert but not the other) is
  rolled back and reported uncached with the failure reason, on both the load
  path and the generation-time auto toggle.

- Partial torchao quantization is no longer reported as dense: after an
  in-place quantize_/caster failure, the DiT / text encoder / VAE is scanned
  for leftover torchao tensor-subclass parameters and the load fails with a
  clear error when any are found (a half-quantized module cannot run as dense,
  and offload's Module.to() crashes on torchao tensors). Failures that swapped
  nothing keep the best-effort dense fallback.

- Cleanup: apply_attention_backend / apply_speed_optims / the attention trim
  are called once on the pipe (they already fan out over every DiT internally),
  so the second A14B expert no longer passes through them twice; the stale
  dual-DiT helper comment is rewritten to match the two helper shapes.

Tests: device-identity picker/gate/lossy-plan coverage, per-family threshold
pin scoping, all-or-none rollback in both failure directions, and partial-quant
detection for all three quant modules.
2026-07-11 10:06:52 +00:00
..
__init__.py Studio: stop chat generation on the assistant-turn-end token (fixes Qwen3.5 loop) (#6804) 2026-07-06 10:07:56 -07:00
_html_to_md.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
_vulkan_probe.py Studio: add Vulkan llama.cpp support (#5819) 2026-07-09 03:39:48 -07:00
anthropic_compat.py Studio: stream reasoning tokens in the tool-loop generator (fixes DeepSeek thinking not streaming with a pill on) (#6947) 2026-07-07 19:50:40 -03:00
api_monitor.py Studio: trim serving-log noise and surface llama-server engine stats (#6377) 2026-06-17 05:37:57 -07:00
audio_codecs.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
chat_eos.py Studio: stop chat generation on the assistant-turn-end token (fixes Qwen3.5 loop) (#6804) 2026-07-06 10:07:56 -07:00
chat_template_helpers.py Studio: render thinking blocks for safetensors inference with prefilled <think> templates (#6816) 2026-07-08 08:14:03 -07:00
chat_templates.py Studio: bundle Gemma 4 chat templates (E2B/E4B + larger) and auto-apply to unsloth/gemma-4-*-GGUF (#6245) 2026-06-12 05:49:39 -07:00
defaults.py Add DeepSeek-V4-Flash-GGUF to Studio with none/high/max reasoning (#6908) 2026-07-07 06:13:43 -07:00
diffusion.py fix(review): portable bench scripts, accurate VAE auto docs, explicit TE deny 2026-07-10 06:13:48 +00:00
diffusion_arch_patches.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_attention.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-11 06:09:43 +00:00
diffusion_auto_policy.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-09 08:53:52 +00:00
diffusion_cache.py Harden the video speed stack: cache quality pin, device identity, transactional caches, quant safety 2026-07-11 10:06:52 +00:00
diffusion_cfg_parallel.py Harden the video speed stack: cache quality pin, device identity, transactional caches, quant safety 2026-07-11 10:06:52 +00:00
diffusion_compile_cache.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_controlnet.py diffusion: address review round (FBCache context guard, aiter/ROCm, video cleanup, prequant + ControlNet gating) 2026-07-09 08:52:16 +00:00
diffusion_device.py Trim the verbose comments added by the fixes 2026-07-04 13:25:08 -03:00
diffusion_eager_patches.py Add grad norm chart, clearer completion state, Windows caption keys, GGUF compute copy 2026-07-04 04:31:04 +00:00
diffusion_engine_router.py Validate a local video checkpoint's own suffix; unload the old engine before publishing the new one 2026-07-07 22:55:36 +00:00
diffusion_families.py Merge remote-tracking branch 'origin/diffusion-more-families' into fold-integration 2026-07-07 01:08:50 +00:00
diffusion_gguf_compile.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_ideogram4.py ideogram-4: build the FP8 text encoder at target dtype, size it as bf16-resident, optimize both DiTs 2026-07-06 10:21:10 +00:00
diffusion_inference_info.py Advertise per-family footprints and surface Auto badges for resolved controls 2026-07-04 07:47:01 +00:00
diffusion_krea2.py diffusion: add AGPL-3.0 SPDX header to Krea2 / LoRA / ControlNet files 2026-07-09 06:46:45 +00:00
diffusion_lora.py diffusion: add AGPL-3.0 SPDX header to Krea2 / LoRA / ControlNet files 2026-07-09 06:46:45 +00:00
diffusion_memory.py Merge image-generation bug fixes (#6872) 2026-07-07 16:14:06 +00:00
diffusion_patch_backend.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_precision.py Harden the video speed stack: cache quality pin, device identity, transactional caches, quant safety 2026-07-11 10:06:52 +00:00
diffusion_prequant.py Merge remote-tracking branch 'origin/image-generation' into r7021 2026-07-09 11:09:08 +00:00
diffusion_speed.py Merge branch 'image-generation' into video-diffusion-improvements 2026-07-10 17:44:25 +00:00
diffusion_transformer_quant.py Harden the video speed stack: cache quality pin, device identity, transactional caches, quant safety 2026-07-11 10:06:52 +00:00
diffusion_vae_quant.py Harden the video speed stack: cache quality pin, device identity, transactional caches, quant safety 2026-07-11 10:06:52 +00:00
external_provider.py Studio: harden background consumer loops and streaming paths against silent UI freezes (#6653) 2026-06-26 03:31:33 -07:00
gpu_arbiter.py Video inference engine: LTX-2 family registry, VideoBackend, MP4 gallery 2026-07-04 13:08:43 +00:00
image_gallery.py Studio diffusion (Phase 1): cross-platform device policy, fp16 guard, lock split, validate-before-evict (#6670) 2026-06-30 16:33:47 -03:00
inference.py Studio: render thinking blocks for safetensors inference with prefilled <think> templates (#6816) 2026-07-08 08:14:03 -07:00
key_exchange.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
llama_cpp.py Studio: harden OpenAI-compatible GGUF streaming (#6950) 2026-07-09 12:09:08 -03:00
llama_http.py fix(studio/llama_cpp): disable trust_env on the loopback health probe (#6750) (#6752) 2026-06-30 19:09:26 +02:00
llama_keepwarm.py Run video generation as a background job so secure mode's tunnel cap cannot 524 it 2026-07-10 09:19:02 +00:00
llama_server_args.py studio: return a clean model id from the OpenAI API instead of the local .gguf path (#6518) 2026-06-26 16:07:53 -03:00
llama_stats.py Studio: trim serving-log noise and surface llama-server engine stats (#6377) 2026-06-17 05:37:57 -07:00
local_model_resolver.py Studio: opt-in OpenAI /v1 model auto-switch and idle keep-warm (#6392) 2026-07-01 06:42:23 -07:00
mcp_client.py Studio: enable stdio MCP servers on a loopback bind (#6295) 2026-06-15 03:02:32 +01:00
mcp_config_import.py studio: show MCP "Import config" on the add-server form (#6030) 2026-06-11 16:17:22 +01:00
message_content.py fix(studio): handle multimodal list content in inference text paths (#4383) (#6480) 2026-06-23 01:26:11 -07:00
mlx_inference.py Studio: render thinking blocks for safetensors inference with prefilled <think> templates (#6816) 2026-07-08 08:14:03 -07:00
model_ids.py studio: list the full local model catalog from /v1/models (#6519) 2026-06-26 20:42:06 -03:00
orchestrator.py feat(cli): support MLX distributed inference (#6845) 2026-07-08 03:25:39 -07:00
passthrough_healing.py Studio: parse Mistral [TOOL_CALLS] and rehearsal tool-call shapes (#5704) 2026-07-06 18:52:13 -07:00
presence_penalty.py Studio: apply presence_penalty on the safetensors and MLX inference paths (#6923) 2026-07-06 22:24:47 -07:00
pricing.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
providers.py Studio: Add custom provider option to Connections (#6112) 2026-06-12 13:09:35 +02:00
runtime_context.py Expose runtime context length for hub models (#6154) 2026-06-11 22:13:53 +03:00
safetensors_agentic.py Studio chat: tool-call nudging on by default (API stays opt-in) (#6883) 2026-07-06 19:41:19 -07:00
sd_cpp_args.py Reject sd-cli batch runs and clear stale output targets before a run 2026-07-05 01:49:19 +00:00
sd_cpp_backend.py diffusion: accept vae_quant in the native sd.cpp backend load interface 2026-07-09 07:58:30 +00:00
sd_cpp_engine.py Merge remote-tracking branch 'origin/image-generation' into fix/imggen-review-bugs 2026-07-06 09:16:30 -03:00
sd_cpp_server.py Studio diffusion: LoRA adapters for the Images workflow (#6771) 2026-07-03 15:42:56 -03:00
tensor_fallback.py studio: deterministic VRAM auto-fit for GGUF (MTP reserve, compute buffer, total-based budget) (#6312) 2026-06-17 03:10:22 -07:00
tool_call_parser.py Studio chat: tool-call nudging on by default (API stays opt-in) (#6883) 2026-07-06 19:41:19 -07:00
tool_loop_controller.py Studio: clean-room compact RAG (knowledge bases, hybrid search, fast indexing) (#5910) 2026-06-09 21:17:04 -07:00
tools.py Whole-document context for RAG chat attachments (#6693) 2026-06-30 15:55:23 +02:00
video.py Harden the video speed stack: cache quality pin, device identity, transactional caches, quant safety 2026-07-11 10:06:52 +00:00
video_families.py Absorb the first-generation compile hitch with a post-load background prewarm 2026-07-11 06:49:07 +00:00
video_gallery.py Delete gallery videos MP4-first so a locked MP4 can't orphan the record 2026-07-07 11:57:37 +00:00
video_ltx2.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-04 14:47:26 +00:00
worker.py feat(cli): support MLX distributed inference (#6845) 2026-07-08 03:25:39 -07:00