The A/B harness measured two regimes. bf16 loads (the Studio default on Ampere+) are bit-identical with the flag on across all six families, 36/36 same-seed cases, because the flag only changes fp16 GEMM accumulation. fp16 loads (the pre-Ampere fallback dtype) show real same-seed drift on the families that genuinely run fp16 GEMMs: SDXL up to 0.050 mean abs diff, FLUX.1 0.028, FLUX.2-klein 0.045, all finite, no new black frames. qwen-image renders black in fp16 with the flag off too and z-image fp16 fails in attention, so both are dtype limitations, not accumulation ones. So the gate now takes the compute dtype and the speed tier: bf16 engages on any active tier (provably output-neutral), fp16 engages only under max, the tier that already trades exactness for measured speed. The deny-list stays empty by measurement. |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| src-tauri | ||
| __init__.py | ||
| install_llama_prebuilt.py | ||
| install_node_prebuilt.py | ||
| install_python_stack.py | ||
| install_sd_cpp_prebuilt.py | ||
| LICENSE.AGPL-3.0 | ||
| node_prebuilt_pins.json | ||
| package-lock.json | ||
| package.json | ||
| setup.bat | ||
| setup.ps1 | ||
| setup.sh | ||
| Unsloth_Studio_Colab.ipynb | ||