unsloth/studio/backend/core/inference
Daniel Han be04ba00f4 video/image: honor explicit Speed=off for companions + trim, probe explicit TE kernels, bench fidelity
Address the Codex review round on the video/quant work:

- Companion auto-quant now honors an explicit Speed=off. Both loaders already pin the DiT dense
  under an explicit off (bit-exact reference), but the unset text-encoder / VAE quant still promoted
  to auto and silently fp8/int8'd the companions, breaking the bit-exact request. An UNSET speed
  still auto-quantises; an explicit companion scheme still forces it.
- The HunyuanVideo joint-attention trim is a speed lever (it swaps to the fused SDPA kernel), so gate
  it on a non-off speed tier exactly like the adjacent attention-backend selection -- the off path
  keeps the stock dense-mask attention.
- Explicit torchao text-encoder modes (int8 / fp8_dynamic / nvfp4) now run the same kernel smoke
  test the auto ladder uses. They could clear the capability gate yet fail the real GEMM on a build
  where quantize_ wraps the encoder but the kernel is broken; the caster's try/except only covers the
  cast, not the first forward, so the load would report engaged then crash at generation. Now it
  falls back to dense. Layerwise fp8 has no torchao GEMM, so the probe is a no-op for it.
- The trim pre-hook's fallback restores the caller's original kwargs (it may have emptied the image
  stream / trimmed a text stream before failing), so the stock dense-mask path runs on exactly what
  it expects, matching the empty-prompt guard.
- video_speedmem_bench mirrors the loader: installs the Hunyuan trim before the backend set (gated on
  an active tier) and skips the auto int8 quant when it is the fp8-denied memory fallback and dense
  fits resident, so the shipped/auto rows measure what the loader actually runs.

Tests: TE explicit-mode kernel probe (+ layerwise-fp8 bypass), trim mid-trim restore, and loader-level
speed=off companion suppression + trim skip for both backends. 262 backend tests pass; ruff clean.
2026-07-09 09:24:28 +00:00
..
__init__.py Studio: stop chat generation on the assistant-turn-end token (fixes Qwen3.5 loop) (#6804) 2026-07-06 10:07:56 -07:00
_html_to_md.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
anthropic_compat.py Studio: stream reasoning tokens in the tool-loop generator (fixes DeepSeek thinking not streaming with a pill on) (#6947) 2026-07-07 19:50:40 -03:00
api_monitor.py Studio: trim serving-log noise and surface llama-server engine stats (#6377) 2026-06-17 05:37:57 -07:00
audio_codecs.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
chat_eos.py Studio: stop chat generation on the assistant-turn-end token (fixes Qwen3.5 loop) (#6804) 2026-07-06 10:07:56 -07:00
chat_template_helpers.py studio: tool calling for DeepSeek (R1/V3/V3.1), GLM 4.x, Kimi K2 on safetensors + MLX (#5624) 2026-07-06 15:40:46 -07:00
chat_templates.py Studio: bundle Gemma 4 chat templates (E2B/E4B + larger) and auto-apply to unsloth/gemma-4-*-GGUF (#6245) 2026-06-12 05:49:39 -07:00
defaults.py Add DeepSeek-V4-Flash-GGUF to Studio with none/high/max reasoning (#6908) 2026-07-07 06:13:43 -07:00
diffusion.py video/image: honor explicit Speed=off for companions + trim, probe explicit TE kernels, bench fidelity 2026-07-09 09:24:28 +00:00
diffusion_arch_patches.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_attention.py video/image: honor explicit Speed=off for companions + trim, probe explicit TE kernels, bench fidelity 2026-07-09 09:24:28 +00:00
diffusion_auto_policy.py ideogram-4: build the FP8 text encoder at target dtype, size it as bf16-resident, optimize both DiTs 2026-07-06 10:21:10 +00:00
diffusion_cache.py Key the auto step-cache on the pipe's default strength when the request omits it 2026-07-06 14:49:47 +00:00
diffusion_compile_cache.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_controlnet.py Fix union ControlNet mode for bare repo-id inputs and resident single-file Reapply 2026-07-08 03:15:09 +00:00
diffusion_device.py Trim the verbose comments added by the fixes 2026-07-04 13:25:08 -03:00
diffusion_eager_patches.py Add grad norm chart, clearer completion state, Windows caption keys, GGUF compute copy 2026-07-04 04:31:04 +00:00
diffusion_engine_router.py Validate a local video checkpoint's own suffix; unload the old engine before publishing the new one 2026-07-07 22:55:36 +00:00
diffusion_families.py Merge remote-tracking branch 'origin/diffusion-more-families' into fold-integration 2026-07-07 01:08:50 +00:00
diffusion_gguf_compile.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_ideogram4.py ideogram-4: build the FP8 text encoder at target dtype, size it as bf16-resident, optimize both DiTs 2026-07-06 10:21:10 +00:00
diffusion_inference_info.py Advertise per-family footprints and surface Auto badges for resolved controls 2026-07-04 07:47:01 +00:00
diffusion_krea2.py Add missing AGPL-3.0 SPDX headers to krea2 / lora / controlnet sources 2026-07-08 10:45:27 +00:00
diffusion_lora.py Add missing AGPL-3.0 SPDX headers to krea2 / lora / controlnet sources 2026-07-08 10:45:27 +00:00
diffusion_memory.py Merge image-generation bug fixes (#6872) 2026-07-07 16:14:06 +00:00
diffusion_patch_backend.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_precision.py video/image: honor explicit Speed=off for companions + trim, probe explicit TE kernels, bench fidelity 2026-07-09 09:24:28 +00:00
diffusion_prequant.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-09 06:13:22 +00:00
diffusion_speed.py Merge remote-tracking branch 'origin/video-hunyuan-gate' into fold-integration 2026-07-07 01:15:20 +00:00
diffusion_transformer_quant.py Run video DiT dense+compile when it fits instead of a slower int8 fallback 2026-07-09 02:31:22 +00:00
diffusion_vae_quant.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-09 06:13:22 +00:00
external_provider.py Studio: harden background consumer loops and streaming paths against silent UI freezes (#6653) 2026-06-26 03:31:33 -07:00
gpu_arbiter.py Video inference engine: LTX-2 family registry, VideoBackend, MP4 gallery 2026-07-04 13:08:43 +00:00
image_gallery.py Studio diffusion (Phase 1): cross-platform device policy, fp16 guard, lock split, validate-before-evict (#6670) 2026-06-30 16:33:47 -03:00
inference.py Studio: apply presence_penalty on the safetensors and MLX inference paths (#6923) 2026-07-06 22:24:47 -07:00
key_exchange.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
llama_cpp.py Studio: stream reasoning tokens in the tool-loop generator (fixes DeepSeek thinking not streaming with a pill on) (#6947) 2026-07-07 19:50:40 -03:00
llama_http.py fix(studio/llama_cpp): disable trust_env on the loopback health probe (#6750) (#6752) 2026-06-30 19:09:26 +02:00
llama_keepwarm.py Studio: gate untrusted companion base_repo on image loads, track image/video generation for the training guard, and body-cap the whole /v1 surface 2026-07-07 05:04:49 +00:00
llama_server_args.py studio: return a clean model id from the OpenAI API instead of the local .gguf path (#6518) 2026-06-26 16:07:53 -03:00
llama_stats.py Studio: trim serving-log noise and surface llama-server engine stats (#6377) 2026-06-17 05:37:57 -07:00
local_model_resolver.py Studio: opt-in OpenAI /v1 model auto-switch and idle keep-warm (#6392) 2026-07-01 06:42:23 -07:00
mcp_client.py Studio: enable stdio MCP servers on a loopback bind (#6295) 2026-06-15 03:02:32 +01:00
mcp_config_import.py studio: show MCP "Import config" on the add-server form (#6030) 2026-06-11 16:17:22 +01:00
message_content.py fix(studio): handle multimodal list content in inference text paths (#4383) (#6480) 2026-06-23 01:26:11 -07:00
mlx_inference.py Studio: apply presence_penalty on the safetensors and MLX inference paths (#6923) 2026-07-06 22:24:47 -07:00
model_ids.py studio: list the full local model catalog from /v1/models (#6519) 2026-06-26 20:42:06 -03:00
orchestrator.py Speed up Studio startup path (#6899) 2026-07-07 18:08:07 -07:00
passthrough_healing.py Studio: parse Mistral [TOOL_CALLS] and rehearsal tool-call shapes (#5704) 2026-07-06 18:52:13 -07:00
presence_penalty.py Studio: apply presence_penalty on the safetensors and MLX inference paths (#6923) 2026-07-06 22:24:47 -07:00
pricing.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
providers.py Studio: Add custom provider option to Connections (#6112) 2026-06-12 13:09:35 +02:00
runtime_context.py Expose runtime context length for hub models (#6154) 2026-06-11 22:13:53 +03:00
safetensors_agentic.py Studio chat: tool-call nudging on by default (API stays opt-in) (#6883) 2026-07-06 19:41:19 -07:00
sd_cpp_args.py Reject sd-cli batch runs and clear stale output targets before a run 2026-07-05 01:49:19 +00:00
sd_cpp_backend.py diffusion: accept vae_quant in the native sd.cpp backend load interface 2026-07-09 07:58:30 +00:00
sd_cpp_engine.py Merge remote-tracking branch 'origin/image-generation' into fix/imggen-review-bugs 2026-07-06 09:16:30 -03:00
sd_cpp_server.py Studio diffusion: LoRA adapters for the Images workflow (#6771) 2026-07-03 15:42:56 -03:00
tensor_fallback.py studio: deterministic VRAM auto-fit for GGUF (MTP reserve, compute buffer, total-based budget) (#6312) 2026-06-17 03:10:22 -07:00
tool_call_parser.py Studio chat: tool-call nudging on by default (API stays opt-in) (#6883) 2026-07-06 19:41:19 -07:00
tool_loop_controller.py Studio: clean-room compact RAG (knowledge bases, hybrid search, fast indexing) (#5910) 2026-06-09 21:17:04 -07:00
tools.py Whole-document context for RAG chat attachments (#6693) 2026-06-30 15:55:23 +02:00
video.py video/image: honor explicit Speed=off for companions + trim, probe explicit TE kernels, bench fidelity 2026-07-09 09:24:28 +00:00
video_families.py Merge remote-tracking branch 'origin/video-hunyuan-gate' into fold-integration 2026-07-07 01:15:20 +00:00
video_gallery.py Delete gallery videos MP4-first so a locked MP4 can't orphan the record 2026-07-07 11:57:37 +00:00
video_ltx2.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-04 14:47:26 +00:00
worker.py Studio: apply presence_penalty on the safetensors and MLX inference paths (#6923) 2026-07-06 22:24:47 -07:00