The A/B harness measured two regimes. bf16 loads (the Studio default on Ampere+) are bit-identical with the flag on across all six families, 36/36 same-seed cases, because the flag only changes fp16 GEMM accumulation. fp16 loads (the pre-Ampere fallback dtype) show real same-seed drift on the families that genuinely run fp16 GEMMs: SDXL up to 0.050 mean abs diff, FLUX.1 0.028, FLUX.2-klein 0.045, all finite, no new black frames. qwen-image renders black in fp16 with the flag off too and z-image fp16 fails in attention, so both are dtype limitations, not accumulation ones. So the gate now takes the compute dtype and the speed tier: bf16 engages on any active tier (provably output-neutral), fp16 engages only under max, the tier that already trades exactness for measured speed. The deny-list stays empty by measurement. |
||
|---|---|---|
| .. | ||
| assets | ||
| auth | ||
| core | ||
| hub | ||
| loggers | ||
| models | ||
| plugins | ||
| requirements | ||
| routes | ||
| state | ||
| storage | ||
| tests | ||
| utils | ||
| __init__.py | ||
| _platform_compat.py | ||
| cloudflare_tunnel.py | ||
| colab.py | ||
| main.py | ||
| run.py | ||
| startup_banner.py | ||