unsloth/studio/backend/core/inference
Daniel Han a5928064a0 video: skip padded text tokens in HunyuanVideo-1.5 joint attention
HunyuanVideo-1.5's DiT runs a joint [video; text] self-attention and, on every
block and step, builds a dense [B,1,N,N] boolean mask so the video never attends
to the padded text. A dense bool attn_mask disables every fused SDPA kernel
(flash rejects it; cuDNN and memory-efficient fall back), so the attention runs
the slow math-style path: at the production shape (121 frames, 480p, N about 50k)
one attention call is ~421ms with the mask vs ~19ms with attn_mask=None. The text
is ~99.5% padding (a t2v prompt fills ~9 of ~1985 slots), so nearly all of that
cost is spent masking padding.

install_hunyuan_attention_trim installs an eager forward pre-hook that drops the
all-zero image stream (t2v) and trims the mllm/byt5 text streams to their
globally-valid columns, plus a null-mask attention processor that runs
attn_mask=None once no partially-padded column remains (the batch-1 /
per-guidance-branch case) and otherwise delegates to the stock dense-mask
processor. The model already zeroes and masks the padded text and discards its
attention output (only the video split feeds proj_out), so removing it is exact
for the video; the only numeric change is the SDPA kernel (masked fallback to
fused). Measured on a B200: 23.3s to 1.3s per DiT forward at 121 frames (~18x with
regional compile, 0 graph breaks); per-forward cosine 0.99998 vs stock; equal
distance to an fp32 reference (LPIPS fp32-vs-stock 0.292, fp32-vs-trim 0.307), so
it is not less accurate than the current bf16 default.

Wired auto-on for HunyuanVideo-1.5 in the video loader, before the attention
backend set so the requested kernel pins onto the new processors; a no-op for
every other family and reversible (stock dense-mask path on any anomaly). Adds
hermetic tests and the diagnostic/validation scripts.
2026-07-09 05:09:49 +00:00
..
__init__.py Studio: stop chat generation on the assistant-turn-end token (fixes Qwen3.5 loop) (#6804) 2026-07-06 10:07:56 -07:00
_html_to_md.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
anthropic_compat.py Studio: stream reasoning tokens in the tool-loop generator (fixes DeepSeek thinking not streaming with a pill on) (#6947) 2026-07-07 19:50:40 -03:00
api_monitor.py Studio: trim serving-log noise and surface llama-server engine stats (#6377) 2026-06-17 05:37:57 -07:00
audio_codecs.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
chat_eos.py Studio: stop chat generation on the assistant-turn-end token (fixes Qwen3.5 loop) (#6804) 2026-07-06 10:07:56 -07:00
chat_template_helpers.py studio: tool calling for DeepSeek (R1/V3/V3.1), GLM 4.x, Kimi K2 on safetensors + MLX (#5624) 2026-07-06 15:40:46 -07:00
chat_templates.py Studio: bundle Gemma 4 chat templates (E2B/E4B + larger) and auto-apply to unsloth/gemma-4-*-GGUF (#6245) 2026-06-12 05:49:39 -07:00
defaults.py Add DeepSeek-V4-Flash-GGUF to Studio with none/high/max reasoning (#6908) 2026-07-07 06:13:43 -07:00
diffusion.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-08 10:05:51 +00:00
diffusion_arch_patches.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_attention.py video: skip padded text tokens in HunyuanVideo-1.5 joint attention 2026-07-09 05:09:49 +00:00
diffusion_auto_policy.py ideogram-4: build the FP8 text encoder at target dtype, size it as bf16-resident, optimize both DiTs 2026-07-06 10:21:10 +00:00
diffusion_cache.py Key the auto step-cache on the pipe's default strength when the request omits it 2026-07-06 14:49:47 +00:00
diffusion_compile_cache.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_controlnet.py Fix union ControlNet mode for bare repo-id inputs and resident single-file Reapply 2026-07-08 03:15:09 +00:00
diffusion_device.py Trim the verbose comments added by the fixes 2026-07-04 13:25:08 -03:00
diffusion_eager_patches.py Add grad norm chart, clearer completion state, Windows caption keys, GGUF compute copy 2026-07-04 04:31:04 +00:00
diffusion_engine_router.py Validate a local video checkpoint's own suffix; unload the old engine before publishing the new one 2026-07-07 22:55:36 +00:00
diffusion_families.py Merge remote-tracking branch 'origin/diffusion-more-families' into fold-integration 2026-07-07 01:08:50 +00:00
diffusion_gguf_compile.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_ideogram4.py ideogram-4: build the FP8 text encoder at target dtype, size it as bf16-resident, optimize both DiTs 2026-07-06 10:21:10 +00:00
diffusion_inference_info.py Advertise per-family footprints and surface Auto badges for resolved controls 2026-07-04 07:47:01 +00:00
diffusion_krea2.py Add missing AGPL-3.0 SPDX headers to krea2 / lora / controlnet sources 2026-07-08 10:45:27 +00:00
diffusion_lora.py Add missing AGPL-3.0 SPDX headers to krea2 / lora / controlnet sources 2026-07-08 10:45:27 +00:00
diffusion_memory.py Merge image-generation bug fixes (#6872) 2026-07-07 16:14:06 +00:00
diffusion_patch_backend.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_precision.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-08 09:48:12 +00:00
diffusion_prequant.py Restore fp8 DiT quant for Wan video via a per-family embedder exclude 2026-07-08 23:22:34 +00:00
diffusion_speed.py Merge remote-tracking branch 'origin/video-hunyuan-gate' into fold-integration 2026-07-07 01:15:20 +00:00
diffusion_transformer_quant.py Run video DiT dense+compile when it fits instead of a slower int8 fallback 2026-07-09 02:31:22 +00:00
diffusion_vae_quant.py Size-gate VAE auto-quant: quantize large (video) VAEs only, skip tiny image VAEs 2026-07-08 12:25:53 +00:00
external_provider.py Studio: harden background consumer loops and streaming paths against silent UI freezes (#6653) 2026-06-26 03:31:33 -07:00
gpu_arbiter.py Video inference engine: LTX-2 family registry, VideoBackend, MP4 gallery 2026-07-04 13:08:43 +00:00
image_gallery.py Studio diffusion (Phase 1): cross-platform device policy, fp16 guard, lock split, validate-before-evict (#6670) 2026-06-30 16:33:47 -03:00
inference.py Studio: apply presence_penalty on the safetensors and MLX inference paths (#6923) 2026-07-06 22:24:47 -07:00
key_exchange.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
llama_cpp.py Studio: stream reasoning tokens in the tool-loop generator (fixes DeepSeek thinking not streaming with a pill on) (#6947) 2026-07-07 19:50:40 -03:00
llama_http.py fix(studio/llama_cpp): disable trust_env on the loopback health probe (#6750) (#6752) 2026-06-30 19:09:26 +02:00
llama_keepwarm.py Studio: gate untrusted companion base_repo on image loads, track image/video generation for the training guard, and body-cap the whole /v1 surface 2026-07-07 05:04:49 +00:00
llama_server_args.py studio: return a clean model id from the OpenAI API instead of the local .gguf path (#6518) 2026-06-26 16:07:53 -03:00
llama_stats.py Studio: trim serving-log noise and surface llama-server engine stats (#6377) 2026-06-17 05:37:57 -07:00
local_model_resolver.py Studio: opt-in OpenAI /v1 model auto-switch and idle keep-warm (#6392) 2026-07-01 06:42:23 -07:00
mcp_client.py Studio: enable stdio MCP servers on a loopback bind (#6295) 2026-06-15 03:02:32 +01:00
mcp_config_import.py studio: show MCP "Import config" on the add-server form (#6030) 2026-06-11 16:17:22 +01:00
message_content.py fix(studio): handle multimodal list content in inference text paths (#4383) (#6480) 2026-06-23 01:26:11 -07:00
mlx_inference.py Studio: apply presence_penalty on the safetensors and MLX inference paths (#6923) 2026-07-06 22:24:47 -07:00
model_ids.py studio: list the full local model catalog from /v1/models (#6519) 2026-06-26 20:42:06 -03:00
orchestrator.py Speed up Studio startup path (#6899) 2026-07-07 18:08:07 -07:00
passthrough_healing.py Studio: parse Mistral [TOOL_CALLS] and rehearsal tool-call shapes (#5704) 2026-07-06 18:52:13 -07:00
presence_penalty.py Studio: apply presence_penalty on the safetensors and MLX inference paths (#6923) 2026-07-06 22:24:47 -07:00
pricing.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
providers.py Studio: Add custom provider option to Connections (#6112) 2026-06-12 13:09:35 +02:00
runtime_context.py Expose runtime context length for hub models (#6154) 2026-06-11 22:13:53 +03:00
safetensors_agentic.py Studio chat: tool-call nudging on by default (API stays opt-in) (#6883) 2026-07-06 19:41:19 -07:00
sd_cpp_args.py Reject sd-cli batch runs and clear stale output targets before a run 2026-07-05 01:49:19 +00:00
sd_cpp_backend.py Refuse deleting an in-use native companion repo while a GGUF is loaded 2026-07-08 00:08:55 +00:00
sd_cpp_engine.py Merge remote-tracking branch 'origin/image-generation' into fix/imggen-review-bugs 2026-07-06 09:16:30 -03:00
sd_cpp_server.py Studio diffusion: LoRA adapters for the Images workflow (#6771) 2026-07-03 15:42:56 -03:00
tensor_fallback.py studio: deterministic VRAM auto-fit for GGUF (MTP reserve, compute buffer, total-based budget) (#6312) 2026-06-17 03:10:22 -07:00
tool_call_parser.py Studio chat: tool-call nudging on by default (API stays opt-in) (#6883) 2026-07-06 19:41:19 -07:00
tool_loop_controller.py Studio: clean-room compact RAG (knowledge bases, hybrid search, fast indexing) (#5910) 2026-06-09 21:17:04 -07:00
tools.py Whole-document context for RAG chat attachments (#6693) 2026-06-30 15:55:23 +02:00
video.py video: skip padded text tokens in HunyuanVideo-1.5 joint attention 2026-07-09 05:09:49 +00:00
video_families.py Merge remote-tracking branch 'origin/video-hunyuan-gate' into fold-integration 2026-07-07 01:15:20 +00:00
video_gallery.py Delete gallery videos MP4-first so a locked MP4 can't orphan the record 2026-07-07 11:57:37 +00:00
video_ltx2.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-04 14:47:26 +00:00
worker.py Studio: apply presence_penalty on the safetensors and MLX inference paths (#6923) 2026-07-06 22:24:47 -07:00