unsloth/studio/backend/core/inference
Daniel Han 2c4386ffc1 Pass the calibrated distilled sigma curve to LTX-2.3 8-step runs
The 22B distilled DiT was trained against ltx_core's fixed
DISTILLED_SIGMA_VALUES, but the diffusers scheduler derives 8-step
spacing from resolution-shifted flow matching and lands far off at
every reachable mu (second sigma 0.945-0.981 vs 0.99375, tail
0.37-0.61 -> 0.1 vs 0.725 -> 0.42 -> 0). At the distilled default step
count the backend now passes the list verbatim, neutralising the
scheduler's dynamic shift and terminal stretch for the call (they
distort even explicit sigmas) and restoring them afterwards. Other
step counts and the dev/base DiT keep the scheduler's own spacing.

Live-verified on B200: the scheduler holds the exact curve after an
8-step distilled GGUF generation, config restored, healthy clip. Also
reword the transformer_quant resolved reason to the measured reality:
quant halves resident weights and hosted checkpoints cut load time,
while per-step speed is roughly bf16 parity.
2026-07-18 10:59:08 +00:00
..
__init__.py Studio: stop chat generation on the assistant-turn-end token (fixes Qwen3.5 loop) (#6804) 2026-07-06 10:07:56 -07:00
_html_to_md.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
_vulkan_probe.py Studio: add Vulkan llama.cpp support (#5819) 2026-07-09 03:39:48 -07:00
anthropic_compat.py Studio: stream reasoning tokens in the tool-loop generator (fixes DeepSeek thinking not streaming with a pill on) (#6947) 2026-07-07 19:50:40 -03:00
api_monitor.py Studio: trim serving-log noise and surface llama-server engine stats (#6377) 2026-06-17 05:37:57 -07:00
audio_codecs.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
chat_eos.py Studio: stop chat generation on the assistant-turn-end token (fixes Qwen3.5 loop) (#6804) 2026-07-06 10:07:56 -07:00
chat_template_helpers.py Studio: render thinking blocks for safetensors inference with prefilled <think> templates (#6816) 2026-07-08 08:14:03 -07:00
chat_templates.py Studio: bundle Gemma 4 chat templates (E2B/E4B + larger) and auto-apply to unsloth/gemma-4-*-GGUF (#6245) 2026-06-12 05:49:39 -07:00
defaults.py Add DeepSeek-V4-Flash-GGUF to Studio with none/high/max reasoning (#6908) 2026-07-07 06:13:43 -07:00
diffusion.py Host pre-cast fp8 text encoders for four more families 2026-07-18 10:25:14 +00:00
diffusion_arch_patches.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_attention.py Tighten fault-path comments added by the video/diffusion hardening pass 2026-07-13 17:33:01 +00:00
diffusion_auto_policy.py Add the HiDream-I1 family to the image backend 2026-07-17 13:26:18 +00:00
diffusion_cache.py Add Wan2.2-I2V-A14B image-to-video support to the video backend 2026-07-17 11:18:01 +00:00
diffusion_cfg_parallel.py Tighten fault-path comments added by the video/diffusion hardening pass 2026-07-13 17:33:01 +00:00
diffusion_compile_cache.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_controlnet.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_device.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_eager_patches.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_engine_router.py Tighten comments and docstrings added by the image-generation fixes 2026-07-13 05:29:09 +00:00
diffusion_families.py Host pre-cast fp8 text encoders for four more families 2026-07-18 10:25:14 +00:00
diffusion_gguf_compile.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_hidream.py Extend the fp8 TE quant to HiDream's Llama text_encoder_4 2026-07-18 07:55:27 +00:00
diffusion_ideogram4.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_inference_info.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_krea2.py Host pre-cast fp8 text encoders for four more families 2026-07-18 10:25:14 +00:00
diffusion_lora.py Support LoRA adapters on torchao int8/fp8 quantized image pipelines 2026-07-17 09:37:39 +00:00
diffusion_memory.py Harden the diffusion memory plan against transient free-VRAM undercounts 2026-07-17 08:46:24 +00:00
diffusion_patch_backend.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
diffusion_precision.py Report the fp8-cast compute dtype without swapping the encoder class 2026-07-18 10:25:14 +00:00
diffusion_prequant.py Load hosted pre-quantized checkpoints on the video quant path 2026-07-18 05:58:40 +00:00
diffusion_speed.py Reflow verbose diffusion comments to fewer lines 2026-07-13 03:30:48 +00:00
diffusion_te_prequant.py Host pre-cast fp8 text encoders for four more families 2026-07-18 10:25:14 +00:00
diffusion_transformer_quant.py Add Wan2.2-I2V-A14B image-to-video support to the video backend 2026-07-17 11:18:01 +00:00
diffusion_vae_quant.py Harden video diffusion cache, CFG-parallel replica, and layerwise-fp8 rollback 2026-07-13 01:30:22 +00:00
external_provider.py Studio: harden background consumer loops and streaming paths against silent UI freezes (#6653) 2026-06-26 03:31:33 -07:00
gpu_arbiter.py Studio: close arbiter load-registration race and surface native progress + local pipeline folders 2026-07-13 06:31:00 +00:00
image_gallery.py Tighten comments and docstrings added by the image-generation fixes 2026-07-13 05:29:09 +00:00
inference.py Studio: render thinking blocks for safetensors inference with prefilled <think> templates (#6816) 2026-07-08 08:14:03 -07:00
key_exchange.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
llama_admission.py Studio: queue local GGUF OpenAI-compatible requests before llama-server (#7047) 2026-07-10 17:05:48 -03:00
llama_cpp.py Studio: queue local GGUF OpenAI-compatible requests before llama-server (#7047) 2026-07-10 17:05:48 -03:00
llama_http.py fix(studio/llama_cpp): disable trust_env on the loopback health probe (#6750) (#6752) 2026-06-30 19:09:26 +02:00
llama_keepwarm.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
llama_server_args.py studio: return a clean model id from the OpenAI API instead of the local .gguf path (#6518) 2026-06-26 16:07:53 -03:00
llama_stats.py Studio: trim serving-log noise and surface llama-server engine stats (#6377) 2026-06-17 05:37:57 -07:00
local_model_resolver.py Studio: opt-in OpenAI /v1 model auto-switch and idle keep-warm (#6392) 2026-07-01 06:42:23 -07:00
mcp_client.py Studio: enable stdio MCP servers on a loopback bind (#6295) 2026-06-15 03:02:32 +01:00
mcp_config_import.py studio: show MCP "Import config" on the add-server form (#6030) 2026-06-11 16:17:22 +01:00
message_content.py fix(studio): handle multimodal list content in inference text paths (#4383) (#6480) 2026-06-23 01:26:11 -07:00
mlx_inference.py Studio: render thinking blocks for safetensors inference with prefilled <think> templates (#6816) 2026-07-08 08:14:03 -07:00
model_ids.py studio: list the full local model catalog from /v1/models (#6519) 2026-06-26 20:42:06 -03:00
orchestrator.py feat(cli): support MLX distributed inference (#6845) 2026-07-08 03:25:39 -07:00
passthrough_healing.py Studio: parse Mistral [TOOL_CALLS] and rehearsal tool-call shapes (#5704) 2026-07-06 18:52:13 -07:00
presence_penalty.py Studio: apply presence_penalty on the safetensors and MLX inference paths (#6923) 2026-07-06 22:24:47 -07:00
pricing.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
providers.py Studio: Add custom provider option to Connections (#6112) 2026-06-12 13:09:35 +02:00
runtime_context.py Expose runtime context length for hub models (#6154) 2026-06-11 22:13:53 +03:00
safetensors_agentic.py Studio chat: tool-call nudging on by default (API stays opt-in) (#6883) 2026-07-06 19:41:19 -07:00
sd_cpp_args.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
sd_cpp_backend.py Support LoRA adapters on torchao int8/fp8 quantized image pipelines 2026-07-17 09:37:39 +00:00
sd_cpp_engine.py Tighten comments across the image generation stack 2026-07-12 10:55:39 +00:00
sd_cpp_server.py Tighten comments across the remaining image stack files 2026-07-12 11:40:05 +00:00
tensor_fallback.py studio: deterministic VRAM auto-fit for GGUF (MTP reserve, compute buffer, total-based budget) (#6312) 2026-06-17 03:10:22 -07:00
tool_call_parser.py Studio chat: tool-call nudging on by default (API stays opt-in) (#6883) 2026-07-06 19:41:19 -07:00
tool_loop_controller.py Studio: clean-room compact RAG (knowledge bases, hybrid search, fast indexing) (#5910) 2026-06-09 21:17:04 -07:00
tools.py Whole-document context for RAG chat attachments (#6693) 2026-06-30 15:55:23 +02:00
video.py Pass the calibrated distilled sigma curve to LTX-2.3 8-step runs 2026-07-18 10:59:08 +00:00
video_families.py Correct the ltx-2 resident TE estimate to the bf16 cast size 2026-07-18 08:00:34 +00:00
video_gallery.py Tighten comments and docstrings added by the image-generation fixes 2026-07-13 05:29:09 +00:00
video_ltx2.py Pass the calibrated distilled sigma curve to LTX-2.3 8-step runs 2026-07-18 10:59:08 +00:00
worker.py feat(cli): support MLX distributed inference (#6845) 2026-07-08 03:25:39 -07:00