perf(video): accuracy-first round 2 for HunyuanVideo-1.5: compile parity, cache quality presets, dual-GPU CFG
Cuts the shipped default's LPIPS vs the bit-exact reference from 0.224 to 0.139 while going faster (24.9 s to 21.2 s at 720p/33f/30 steps, 22.7x vs reference), and makes the remaining speed/accuracy trade a user knob. - inductor precision parity: set emulate_precision_casts=True for the regional compile (fused pointwise kernels kept fp32 intermediates where eager rounds to bf16 between ops); full-clip LPIPS vs bit-exact 0.221 to 0.052 at zero speed cost. Snapshot/restored with the other process-wide backend flags. - cache x compile composition fix: diffusers cache hooks are torch.compiler.disable'd, so every COMPUTED step ran eager (1.69 vs 1.09 s/step) under MagCache/FBCache in both enable orders. Re-point each hook's fn_ref.original_forward at a torch.compile'd wrapper of the same bound method (armed only where the speed layer compiled the block; restored before every disable_cache so the uncached path stays pristine). Balanced MagCache at 50 steps: 1.48x to 2.17x, identical skip counts, bit-identical uncached rerun after enable/disable cycles. - transformer_cache_quality knob (quality|balanced|fast; API + UI + bench) mapping to (threshold, max_skip_steps, retention_ratio). Auto resolves to the near-lossless quality preset (0.06, 2, 0.3; 1.63-1.64x at pairwise LPIPS 0.05-0.09) for the HunyuanVideo-1.5 families and to balanced (the pre-knob values, byte-identical behaviour) everywhere else. - TE auto-quant resolves dense for HunyuanVideo-1.5: TE fp8_dynamic alone moves the clip to LPIPS 0.236 vs bit-exact for zero speed win (the quantised encoder perturbs the conditioning and the trajectory amplifies it chaotically); VAE fp8 stays in auto (0.053, at the compile floor). Explicit schemes honored. - dual-GPU CFG branch parallelism (new diffusion_cfg_parallel.py): transformer proxy + DiT replica on the most-free second CUDA device + worker thread, branch-routed off the pipeline's own cache_context names. Auto engages only where measured bit-identical (eager tier: max abs diff 0.0, 1.66x); the compiled stack is explicit cfg_parallel=on (1.52x over the sequential default; per-device compiled artifacts differ by 1 bf16 ulp/step, documented in the resolved record). Fail-soft gates: family allowlist, guider CFG, pipeline kind, dense DiT, no offload, free-VRAM check; single-GPU loads are untouched and the memory plan stays single-device. - video API: the transformer_cache literal now accepts auto/magcache (an explicit magcache request was rejected at the pydantic layer); the mxfp8 family deny records the round-2 measurement (block-32 MX scaling fixes the zero-row collapse, no black frames, but is latency-neutral at LPIPS 0.37: fails both ship bars). Measured on B200 via the production lever path (video_speedmem_bench.py, which gained a --cache-quality lever and companion-quant isolation configs). Tests: 441 passing across the video inference suite (32 new for cfg-parallel, 20 for presets/arming, 3 for the inductor flag, 2 for TE auto-dense); ruff clean.
This commit is contained in:
parent
f7824b9db1
commit
7dbdd28161
15 changed files with 1929 additions and 22 deletions
|
|
@ -131,6 +131,14 @@ _TE_AUTO_LADDER: tuple[tuple[tuple[int, int], tuple[str, ...]], ...] = (
|
|||
# rarer case where even keep-bf16 int8 (or fp8) misses the bar for a specific encoder.
|
||||
_TE_FAMILY_SCHEME_DENY: dict[str, frozenset[str]] = {}
|
||||
|
||||
# Families whose AUTO text-encoder quant resolves dense (see select_te_quant_scheme):
|
||||
# measured out-of-bar trajectory drift for zero speed win on the video families below.
|
||||
# Unlike the deny table this only steers the AUTO default; an explicit scheme request
|
||||
# (text_encoder_quant="fp8_dynamic") is still honored verbatim.
|
||||
_TE_AUTO_DENSE_FAMILIES: frozenset[str] = frozenset(
|
||||
{"hunyuanvideo-1.5", "hunyuanvideo-1.5-720p"}
|
||||
)
|
||||
|
||||
# Map a TE torchao scheme to the transformer smoke-probe scheme (same torchao GEMM), so ``auto``
|
||||
# degrades gracefully when a build lacks a kernel. Layerwise fp8 has no torchao GEMM to probe.
|
||||
_TE_SMOKE_SCHEME = {TE_QUANT_FP8_DYNAMIC: "fp8", TE_QUANT_INT8: "int8", TE_QUANT_NVFP4: "nvfp4"}
|
||||
|
|
@ -209,6 +217,16 @@ def select_te_quant_scheme(
|
|||
requested = normalize_te_quant(requested)
|
||||
if requested is None or requested != TE_QUANT_AUTO:
|
||||
return requested
|
||||
# AUTO resolves dense for these families regardless of hardware. TE quant perturbs
|
||||
# the CONDITIONING and a multi-step video trajectory amplifies that chaotically: on
|
||||
# HunyuanVideo-1.5-720p (B200, 720p/33f/30 steps) TE fp8_dynamic ALONE moves the
|
||||
# clip to LPIPS 0.236 vs the bit-exact reference while the rest of the shipped
|
||||
# stack sits at 0.052-0.053, for ZERO speed win (35.48 vs 35.36 s e2e -- the TE
|
||||
# runs once per generation). The ~6.7 GB of weight savings is not worth being the
|
||||
# single dominant accuracy cost of the default stack. (VAE fp8 stays in auto:
|
||||
# measured 0.053, at the compile floor -- decode-only, no trajectory to amplify.)
|
||||
if (family or "").strip().lower() in _TE_AUTO_DENSE_FAMILIES:
|
||||
return None
|
||||
from .diffusion_transformer_quant import _capability, _is_consumer_gpu
|
||||
|
||||
cap = _capability()
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue