unsloth/studio/backend/core/inference
Daniel Han a5195517cf Load Ideogram 4 fp8 repo by dequantizing and remapping its DiTs and text encoder
The ideogram-ai/ideogram-4-fp8 repo stores its two DiTs and the Qwen3-VL text
encoder in a vendor float8 layout that diffusers 0.39.0 (and diffusers main)
cannot read, so a stock Ideogram4Pipeline.from_pretrained produced a pipeline
with randomly initialized attention weights left on the meta device: the load
then died at pipe.to(device) with "Cannot copy out of meta tensor", and any load
that got past that would have generated noise.

Two things broke:

- The DiT attention is stored FUSED as attention.qkv.weight ([3*hidden, hidden],
  Q/K/V rows stacked) plus attention.o.weight, while the diffusers transformer has
  split to_q/to_k/to_v/to_out.0. from_pretrained mapped neither name and left them
  meta + random.
- Every quantized weight is float8_e4m3 with a per-output-channel weight_scale;
  the real weight is fp8.float() * weight_scale[:, None]. diffusers dropped the
  scales and loaded the raw fp8 values (range +-448) as the weights, so even the
  weights that did map were wrong.

load_ideogram4_transformer now reads the shards, dequantizes every scaled weight,
splits the fused qkv into to_q/to_k/to_v and renames o to to_out.0, then loads the
result into a config-constructed model. It fails loudly if any key stays unmatched
so a partly random model can never ship. The dequantized fp8 projections match the
byte-identical -nf4 export (already in the diffusers split layout with a bnb
quantization_config) to cosine ~0.997, so the split order and scale axis are
confirmed. The conversion is gated on the fp8 marker (a *.weight_scale key) read
from the shard header only, so the -nf4 repos skip it and load through the stock
from_pretrained path without a wasteful full-shard read.

The fp8 text encoder needed the same float8 dequant (its keys already match the
transformers Qwen3-VL module, so no rename). load_ideogram4_text_encoder handles
the fp8 repo and delegates the bnb-4bit and dense repos to the shared krea shim.

One more incompatibility was in the diffusers pipeline itself: it calls
transformers create_causal_mask(inputs_embeds = ...) with no cache_position, but
on transformers 4.57.6 the parameter is spelled input_embeds and cache_position is
required. _patch_create_causal_mask installs a signature-aware wrapper that renames
the kwarg and supplies cache_position, and is self-disabling on a matching signature.

Adds unit tests for the fp8 dequant/split conversion and the causal-mask patch.
Verified live on a B200: ideogram-4-fp8 (both CFG paths), ideogram-4-nf4-diffusers,
and krea-2 with the retroanime LoRA all load and generate coherent images.
2026-07-04 14:30:58 +00:00
..
__init__.py Studio: make code comments and docstrings more succinct (#6029) 2026-06-08 23:07:28 -07:00
_html_to_md.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
anthropic_compat.py Tool-call healing (default on) and opt-in nudging for the client-tool passthrough (#6801) 2026-07-03 08:22:42 -07:00
api_monitor.py Studio: trim serving-log noise and surface llama-server engine stats (#6377) 2026-06-17 05:37:57 -07:00
audio_codecs.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
chat_template_helpers.py Studio: make code comments and docstrings more succinct (#6029) 2026-06-08 23:07:28 -07:00
chat_templates.py Studio: bundle Gemma 4 chat templates (E2B/E4B + larger) and auto-apply to unsloth/gemma-4-*-GGUF (#6245) 2026-06-12 05:49:39 -07:00
defaults.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
diffusion.py Load Ideogram 4 fp8 repo by dequantizing and remapping its DiTs and text encoder 2026-07-04 14:30:58 +00:00
diffusion_arch_patches.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_attention.py Auto-install optional attention kernels and toggle the step cache per generation 2026-07-04 07:33:07 +00:00
diffusion_auto_policy.py Add Ideogram 4 family, structured HunyuanImage exclusion, curated Krea 2 LoRAs 2026-07-04 12:46:24 +00:00
diffusion_cache.py Auto-install optional attention kernels and toggle the step cache per generation 2026-07-04 07:33:07 +00:00
diffusion_compile_cache.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_controlnet.py Studio diffusion: ControlNet for the Images workflow (diffusers) (#6773) 2026-07-03 15:57:28 -03:00
diffusion_device.py Studio diffusion (Phase 16): route no-GPU loads to the native sd.cpp engine (#6724) 2026-07-01 15:43:56 -03:00
diffusion_eager_patches.py Add grad norm chart, clearer completion state, Windows caption keys, GGUF compute copy 2026-07-04 04:31:04 +00:00
diffusion_engine_router.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_families.py Add Ideogram 4 family, structured HunyuanImage exclusion, curated Krea 2 LoRAs 2026-07-04 12:46:24 +00:00
diffusion_gguf_compile.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_ideogram4.py Load Ideogram 4 fp8 repo by dequantizing and remapping its DiTs and text encoder 2026-07-04 14:30:58 +00:00
diffusion_inference_info.py Advertise per-family footprints and surface Auto badges for resolved controls 2026-07-04 07:47:01 +00:00
diffusion_krea2.py Fail clearly when a local Krea 2 dir lacks model_index.json 2026-07-04 04:31:09 +00:00
diffusion_lora.py Add Ideogram 4 family, structured HunyuanImage exclusion, curated Krea 2 LoRAs 2026-07-04 12:46:24 +00:00
diffusion_memory.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_patch_backend.py Studio diffusion: image workflows (safetensors, image-conditioned, editing) + Images UI redesign (#6769) 2026-07-03 13:45:14 -03:00
diffusion_precision.py Studio diffusion (Phase 4): native stable-diffusion.cpp engine for CPU/Mac (#6679) 2026-07-01 15:03:53 -03:00
diffusion_prequant.py Studio diffusion: fix FP8 transformer quant producing noise (per-row scaling) (#6772) 2026-07-03 13:51:26 -03:00
diffusion_speed.py [pre-commit.ci] auto fixes from pre-commit.com hooks 2026-07-04 07:47:51 +00:00
diffusion_transformer_quant.py Merge diffusion-train-tab-2: qwen dense-quant family deny (black frames, measured) 2026-07-04 08:52:32 +00:00
external_provider.py Studio: harden background consumer loops and streaming paths against silent UI freezes (#6653) 2026-06-26 03:31:33 -07:00
gpu_arbiter.py Studio diffusion (Phase 16): route no-GPU loads to the native sd.cpp engine (#6724) 2026-07-01 15:43:56 -03:00
image_gallery.py Studio diffusion (Phase 1): cross-platform device policy, fp16 guard, lock split, validate-before-evict (#6670) 2026-06-30 16:33:47 -03:00
inference.py fix(studio): handle multimodal list content in inference text paths (#4383) (#6480) 2026-06-23 01:26:11 -07:00
key_exchange.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
llama_cpp.py Studio: reserve CUDA context and mmproj/MTP soft overhead in the GGUF fit budget (#6718) 2026-07-03 13:07:30 -03:00
llama_http.py fix(studio/llama_cpp): disable trust_env on the loopback health probe (#6750) (#6752) 2026-06-30 19:09:26 +02:00
llama_keepwarm.py Studio: opt-in OpenAI /v1 model auto-switch and idle keep-warm (#6392) 2026-07-01 06:42:23 -07:00
llama_server_args.py studio: return a clean model id from the OpenAI API instead of the local .gguf path (#6518) 2026-06-26 16:07:53 -03:00
llama_stats.py Studio: trim serving-log noise and surface llama-server engine stats (#6377) 2026-06-17 05:37:57 -07:00
local_model_resolver.py Studio: opt-in OpenAI /v1 model auto-switch and idle keep-warm (#6392) 2026-07-01 06:42:23 -07:00
mcp_client.py Studio: enable stdio MCP servers on a loopback bind (#6295) 2026-06-15 03:02:32 +01:00
mcp_config_import.py studio: show MCP "Import config" on the add-server form (#6030) 2026-06-11 16:17:22 +01:00
message_content.py fix(studio): handle multimodal list content in inference text paths (#4383) (#6480) 2026-06-23 01:26:11 -07:00
mlx_inference.py Expose runtime context length for hub models (#6154) 2026-06-11 22:13:53 +03:00
model_ids.py studio: list the full local model catalog from /v1/models (#6519) 2026-06-26 20:42:06 -03:00
orchestrator.py Studio: harden background consumer loops and streaming paths against silent UI freezes (#6653) 2026-06-26 03:31:33 -07:00
passthrough_healing.py Tool-call healing (default on) and opt-in nudging for the client-tool passthrough (#6801) 2026-07-03 08:22:42 -07:00
pricing.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
providers.py Studio: Add custom provider option to Connections (#6112) 2026-06-12 13:09:35 +02:00
runtime_context.py Expose runtime context length for hub models (#6154) 2026-06-11 22:13:53 +03:00
safetensors_agentic.py Studio: show tool-call progress for large GGUF tool arguments (#6484) 2026-06-22 05:50:10 -07:00
sd_cpp_args.py Studio diffusion: LoRA adapters for the Images workflow (#6771) 2026-07-03 15:42:56 -03:00
sd_cpp_backend.py Pin the sd.cpp CPU backend to physical cores 2026-07-04 07:40:34 +00:00
sd_cpp_engine.py Studio diffusion: fix FP8 transformer quant producing noise (per-row scaling) (#6772) 2026-07-03 13:51:26 -03:00
sd_cpp_server.py Studio diffusion: LoRA adapters for the Images workflow (#6771) 2026-07-03 15:42:56 -03:00
tensor_fallback.py studio: deterministic VRAM auto-fit for GGUF (MTP reserve, compute buffer, total-based budget) (#6312) 2026-06-17 03:10:22 -07:00
tool_call_parser.py Fix Gemma 4 GGUF OpenAI API streams (#6476) 2026-06-23 06:13:56 -07:00
tool_loop_controller.py Studio: clean-room compact RAG (knowledge bases, hybrid search, fast indexing) (#5910) 2026-06-09 21:17:04 -07:00
tools.py Whole-document context for RAG chat attachments (#6693) 2026-06-30 15:55:23 +02:00
worker.py Generalize transformers tier selection by probing AutoConfig (#6550) 2026-06-22 08:20:06 -07:00