Commit graph

2 commits

Author SHA1 Message Date
pre-commit-ci[bot]
c7e30e9ea8 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-06-26 12:03:36 +00:00
Daniel Han
3a21f12500 Studio diffusion (Phase 8): prefer fp8 over nvfp4 in Blackwell auto ladder
Validated NVFP4 on torch 2.11 + torchao CUTLASS FP4 in an isolated env. The FP4
tensor-core GEMM is genuinely active there (a 16384^3 GEMM hits ~3826 TFLOPS,
2.52x bf16 and 1.37x fp8), but it only beats fp8 on very large GEMMs. At the
diffusion transformer's shapes (hidden ~3072, MLP ~12288, M~4096) NVFP4 is both
slower (0.81x fp8 end to end on Z-Image 1024px) and less accurate (LPIPS 0.166
vs fp8's 0.044). Reorder the Blackwell auto ladder to fp8 before nvfp4 so auto is
correct even on a future MSLK-equipped box; nvfp4 stays an explicit opt-in. Add
scripts/nvfp4_t211_probe.py (extension diagnostics + GEMM micro + end-to-end).
2026-06-26 08:56:23 +00:00