A 28-pair accuracy gate on a B200 (same-seed vs the dense bf16 reference) found per-row fp8 dynamic quant renders EVERY qwen-image frame black (mean luma 0.0000, SSIM 0.016), reproduced identically with on-the-fly quantize_ on the dense transformer, so it is the model's activation range, not a checkpoint artifact. mxfp8 shows real semantic damage at 1024px (CLIP delta mean 0.0146, worst cases 0.064/0.102) and nvfp4 measures LPIPS mean 0.51. int8 dynamic (per-token scales) is excellent on Qwen: LPIPS mean 0.069, SSIM 0.958. The per-scheme smoke probe only proves the GEMM kernel runs, so it cannot catch model-level breakage. Add _FAMILY_SCHEME_DENY consulted by select_transformer_quant_scheme: auto skips denied schemes (Qwen lands on int8) and an explicit denied request returns None, the same GGUF-fallback contract as an unsupported scheme. Family is threaded from the three diffusion.py call sites; existing behavior is unchanged for every other family. 4 new tests; 529 diffusion tests green; CI-sim green. |
||
|---|---|---|
| .. | ||
| assets | ||
| auth | ||
| core | ||
| hub | ||
| loggers | ||
| models | ||
| plugins | ||
| requirements | ||
| routes | ||
| state | ||
| storage | ||
| tests | ||
| utils | ||
| __init__.py | ||
| _platform_compat.py | ||
| cloudflare_tunnel.py | ||
| colab.py | ||
| main.py | ||
| run.py | ||
| startup_banner.py | ||