Codex review: the M=1 modulation/embedder exclusion was wired only into the dense
runtime quantiser; the offline builder scripts/build_prequant_checkpoint.py called
make_filter_fn(min_features) with no exclusion. So an int8 prequant checkpoint
quantised the AdaLN modulation and conditioning-embedder linears, and loading it
via transformer_prequant_path (the load path only loads already-quantised tensors,
it can't re-skip them) reintroduced the torch._int_mm M=1 crash this phase fixes
for the runtime path.
Extracted int8_exclude_name_tokens(scheme) as the single source of truth (int8 ->
the M=1 exclusion, every other scheme -> none) and use it in both the runtime
quantiser and the builder, so a prequant artifact's quantised-layer set always
matches the runtime. fp8/fp4/mx artifacts are byte-identical (empty exclusion).
Test: int8_exclude_name_tokens returns the exclusion for int8 and () for
fp8/nvfp4/mxfp8.
The Phase 8 fast transformer_quant path materialises the dense bf16 transformer on
the GPU and torchao-quantises it in place, so its load peak is ~2x GGUF's (~21 vs
13.4 GB) plus a ~12 GB download. Add a pre-quantized branch: quantise once offline
(scripts/build_prequant_checkpoint.py) and at runtime build the transformer skeleton
on the meta device (accelerate.init_empty_weights) and load_state_dict(assign=True)
the quantized weights, so the dense bf16 never touches the GPU.
Measured (B200, Z-Image fp8): full-pipeline GPU load peak 21.2 -> 14.6 GB (matching
GGUF's 13.4), on-disk 12 -> 6.28 GB, output bit-identical (LPIPS 0.0). It is the same
torchao config + min_features filter the runtime path uses, applied ahead of time.
New core/inference/diffusion_prequant.py (resolve_prequant_source +
load_prequantized_transformer, best-effort, lazy imports). diffusion.py
_load_dense_quant_pipeline tries the pre-quant source first and falls back to the
dense materialise+quantise path, then to GGUF, so the default is unchanged.
DiffusionLoadRequest gains transformer_prequant_path; DiffusionFamily gains an empty
prequant_repos map for hosted checkpoints (hosting deferred). Hermetic CPU tests for
the resolver, the meta-init+assign loader, and the backend branch selection +
fallbacks; GPU verification via scripts/verify_prequant_backend.py.