A 28-pair accuracy gate on a B200 (same-seed vs the dense bf16 reference) found per-row fp8 dynamic quant renders EVERY qwen-image frame black (mean luma 0.0000, SSIM 0.016), reproduced identically with on-the-fly quantize_ on the dense transformer, so it is the model's activation range, not a checkpoint artifact. mxfp8 shows real semantic damage at 1024px (CLIP delta mean 0.0146, worst cases 0.064/0.102) and nvfp4 measures LPIPS mean 0.51. int8 dynamic (per-token scales) is excellent on Qwen: LPIPS mean 0.069, SSIM 0.958. The per-scheme smoke probe only proves the GEMM kernel runs, so it cannot catch model-level breakage. Add _FAMILY_SCHEME_DENY consulted by select_transformer_quant_scheme: auto skips denied schemes (Qwen lands on int8) and an explicit denied request returns None, the same GGUF-fallback contract as an unsupported scheme. Family is threaded from the three diffusion.py call sites; existing behavior is unchanged for every other family. 4 new tests; 529 diffusion tests green; CI-sim green. |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| src-tauri | ||
| __init__.py | ||
| install_llama_prebuilt.py | ||
| install_node_prebuilt.py | ||
| install_python_stack.py | ||
| install_sd_cpp_prebuilt.py | ||
| LICENSE.AGPL-3.0 | ||
| node_prebuilt_pins.json | ||
| package-lock.json | ||
| package.json | ||
| setup.bat | ||
| setup.ps1 | ||
| setup.sh | ||
| Unsloth_Studio_Colab.ipynb | ||