Comment-only pass over the Python this PR touches: drop what the code already
says, collapse multi-line explanations that still read on one line, and keep
the reasoning that is not recoverable from the code. No code, docstring
semantics or behaviour changes; verified with an AST comparison against the
previous revision, and the backend suite is unchanged (same 37 environment
failures as before: the API integration tests that need a live keyed server,
the flash-attn install hooks, and the GPU memory fields).
Three from the latest review.
The video download plan always asked for the wide base file list, so an LTX-2.3
pick staged the 2.0 base's VAEs, vocoder and connectors that the checkpoint
supplies itself, while the companion files the 2.3 assembly does read were left
out of the plan and pulled inline at load, outside the panel's progress, cancel
and disk preflight. The plan now recognises a 2.3 pick by name (the load keeps
the authoritative header probe, and under-guessing only falls back to the
load-time pull), narrows the base list, and stages the extras in the same entry
as the checkpoint so one repo stays one scoped job.
A pick routed from the chat picker arrives as ?model= and ?quant= with no picker
metadata, so a bare local .gguf or .safetensors was loaded as a pipeline: an
explicit model_kind wins over the backend's filename sniffing, so it evicted the
resident model and then failed on the missing model_index.json. Both pages now
derive the load kind from the path, the same way their own picker handlers do.
A torchao int8/fp8 build takes adapters only at load time. Switching artifact
inside one family keeps the LoRA selection, since the family did not change,
but the load did not bake it, so the next generation was rejected with 'reload
the model with the adapter selection' while the picker still showed the adapter
as active. The selection is now dropped once per resident build, with a message
saying to pick and load again.
The 22B distilled DiT was trained against ltx_core's fixed
DISTILLED_SIGMA_VALUES, but the diffusers scheduler derives 8-step
spacing from resolution-shifted flow matching and lands far off at
every reachable mu (second sigma 0.945-0.981 vs 0.99375, tail
0.37-0.61 -> 0.1 vs 0.725 -> 0.42 -> 0). At the distilled default step
count the backend now passes the list verbatim, neutralising the
scheduler's dynamic shift and terminal stretch for the call (they
distort even explicit sigmas) and restoring them afterwards. Other
step counts and the dev/base DiT keep the scheduler's own spacing.
Live-verified on B200 through the video branch backend: the scheduler
holds the exact curve after an 8-step distilled GGUF generation, config
restored, healthy clip. Also reword the transformer_quant resolved
reason to the measured reality: quant halves resident weights and
hosted checkpoints cut load time, while per-step speed is roughly bf16
parity.
The Lightricks/LTX-2.3-fp8 checkpoints store float8 weights with
per-tensor weight_scale and input_scale companions (verified from the
file headers: 1496 F8_E4M3 tensors, 2924 scale tensors). A plain dtype
cast would silently corrupt every quantized layer, so the 2.3 assembly
now detects the companions and raises with a pointer to the GGUF quants,
which offer comparable fidelity through the supported path. Dequantizing
the scaled fp8 layout is a possible follow-up.
diffusers 0.39 ships every LTX-2.3 model class but its single-file loader
maps all LTX-2 checkpoints to the 2.0 config, so 2.3 checkpoints (9-row
modulation tables, gated attention, per-modality connectors) fail a shape
check at load. The community transformer-only GGUFs also lack the text
projections, VAEs, and vocoder that 2.3 moved out of the transformer.
New core/inference/video_ltx2.py detects a 2.3 checkpoint from its header
(6 vs 9 modulation rows, no weight data read) and assembles the full
pipeline: the DiT through from_single_file with the 2.3 config overrides
and the prompt_adaln key renames the stock converter lacks, the 8-layer
per-modality connectors from the same checkpoint plus the text projection
companion file, and the 2.3 video VAE, audio VAE, and BWE vocoder from
the companion files in unsloth/LTX-2.3-GGUF. Configs and rename tables
mirror diffusers' own scripts/convert_ltx2_to_diffusers.py, which the
library loader has not absorbed yet. Assembled through the constructor
because the base repo pins LTX2Vocoder while 2.3 needs LTX2VocoderWithBWE
and the from_pretrained type gate rejects the substitution.
Verified on a B200: distilled-1.1 Q4_K_M GGUF loads in 37s, generates a
49-frame 768x512 clip with synchronized audio in 18s (8 steps), frames
on-prompt and non-black, container decodes fully. Meta-tensor validation
confirms exact key and shape match for all five converted components.
Unit tests cover 2.3 detection (gguf + safetensors headers), combined
checkpoint partitioning, and companion-set choice.