diffusers 0.39 ships every LTX-2.3 model class but its single-file loader
maps all LTX-2 checkpoints to the 2.0 config, so 2.3 checkpoints (9-row
modulation tables, gated attention, per-modality connectors) fail a shape
check at load. The community transformer-only GGUFs also lack the text
projections, VAEs, and vocoder that 2.3 moved out of the transformer.
New core/inference/video_ltx2.py detects a 2.3 checkpoint from its header
(6 vs 9 modulation rows, no weight data read) and assembles the full
pipeline: the DiT through from_single_file with the 2.3 config overrides
and the prompt_adaln key renames the stock converter lacks, the 8-layer
per-modality connectors from the same checkpoint plus the text projection
companion file, and the 2.3 video VAE, audio VAE, and BWE vocoder from
the companion files in unsloth/LTX-2.3-GGUF. Configs and rename tables
mirror diffusers' own scripts/convert_ltx2_to_diffusers.py, which the
library loader has not absorbed yet. Assembled through the constructor
because the base repo pins LTX2Vocoder while 2.3 needs LTX2VocoderWithBWE
and the from_pretrained type gate rejects the substitution.
Verified on a B200: distilled-1.1 Q4_K_M GGUF loads in 37s, generates a
49-frame 768x512 clip with synchronized audio in 18s (8 steps), frames
on-prompt and non-black, container decodes fully. Meta-tensor validation
confirms exact key and shape match for all five converted components.
Unit tests cover 2.3 detection (gguf + safetensors headers), combined
checkpoint partitioning, and companion-set choice.