perf(video): accuracy-first round 2 for HunyuanVideo-1.5: compile parity, cache quality presets, dual-GPU CFG

Cuts the shipped default's LPIPS vs the bit-exact reference from 0.224 to 0.139
while going faster (24.9 s to 21.2 s at 720p/33f/30 steps, 22.7x vs reference),
and makes the remaining speed/accuracy trade a user knob.

- inductor precision parity: set emulate_precision_casts=True for the regional
  compile (fused pointwise kernels kept fp32 intermediates where eager rounds to
  bf16 between ops); full-clip LPIPS vs bit-exact 0.221 to 0.052 at zero speed
  cost. Snapshot/restored with the other process-wide backend flags.
- cache x compile composition fix: diffusers cache hooks are
  torch.compiler.disable'd, so every COMPUTED step ran eager (1.69 vs 1.09
  s/step) under MagCache/FBCache in both enable orders. Re-point each hook's
  fn_ref.original_forward at a torch.compile'd wrapper of the same bound method
  (armed only where the speed layer compiled the block; restored before every
  disable_cache so the uncached path stays pristine). Balanced MagCache at 50
  steps: 1.48x to 2.17x, identical skip counts, bit-identical uncached rerun
  after enable/disable cycles.
- transformer_cache_quality knob (quality|balanced|fast; API + UI + bench)
  mapping to (threshold, max_skip_steps, retention_ratio). Auto resolves to the
  near-lossless quality preset (0.06, 2, 0.3; 1.63-1.64x at pairwise LPIPS
  0.05-0.09) for the HunyuanVideo-1.5 families and to balanced (the pre-knob
  values, byte-identical behaviour) everywhere else.
- TE auto-quant resolves dense for HunyuanVideo-1.5: TE fp8_dynamic alone moves
  the clip to LPIPS 0.236 vs bit-exact for zero speed win (the quantised encoder
  perturbs the conditioning and the trajectory amplifies it chaotically); VAE
  fp8 stays in auto (0.053, at the compile floor). Explicit schemes honored.
- dual-GPU CFG branch parallelism (new diffusion_cfg_parallel.py): transformer
  proxy + DiT replica on the most-free second CUDA device + worker thread,
  branch-routed off the pipeline's own cache_context names. Auto engages only
  where measured bit-identical (eager tier: max abs diff 0.0, 1.66x); the
  compiled stack is explicit cfg_parallel=on (1.52x over the sequential
  default; per-device compiled artifacts differ by 1 bf16 ulp/step, documented
  in the resolved record). Fail-soft gates: family allowlist, guider CFG,
  pipeline kind, dense DiT, no offload, free-VRAM check; single-GPU loads are
  untouched and the memory plan stays single-device.
- video API: the transformer_cache literal now accepts auto/magcache (an
  explicit magcache request was rejected at the pydantic layer); the mxfp8
  family deny records the round-2 measurement (block-32 MX scaling fixes the
  zero-row collapse, no black frames, but is latency-neutral at LPIPS 0.37:
  fails both ship bars).

Measured on B200 via the production lever path (video_speedmem_bench.py, which
gained a --cache-quality lever and companion-quant isolation configs). Tests:
441 passing across the video inference suite (32 new for cfg-parallel, 20 for
presets/arming, 3 for the inductor flag, 2 for TE auto-dense); ruff clean.
This commit is contained in:
Daniel Han 2026-07-10 14:29:14 +00:00
commit 7dbdd28161
15 changed files with 1929 additions and 22 deletions

View file

@ -531,3 +531,81 @@ def test_fp16_accum_allowed_on_fp16_dtype_under_max(monkeypatch):
)
assert applied["fp16_accum"] is True
assert torch.backends.cuda.matmul.allow_fp16_accumulation is True
# ── inductor precision-cast emulation (compile-vs-eager numeric parity) ─────────
def _stub_inductor_config(monkeypatch, torch, *, emulate = False):
"""Attach a fake ``_inductor.config`` to the stubbed torch module (diffusion_speed
resolves it as attributes off the imported torch, never via sys.modules -- so the
real torch._inductor lingering in sys.modules cannot leak into stubbed tests)."""
cfg = types.SimpleNamespace(emulate_precision_casts = emulate)
torch._inductor = types.SimpleNamespace(config = cfg)
return cfg
def test_regional_compile_enables_emulate_precision_casts(monkeypatch):
# Inductor's fused pointwise kernels keep intermediates in fp32 where eager rounds
# to bf16 between ops; over a multi-step denoise that compounds to a visible drift
# (LPIPS 0.221 vs bit-exact on HunyuanVideo-1.5-720p). emulate_precision_casts
# restores eager's rounding at zero measured speed cost (LPIPS 0.052), so the
# regional compile path must switch it on.
torch = _stub_torch(monkeypatch)
_stub_gguf_accel(monkeypatch)
cfg = _stub_inductor_config(monkeypatch, torch, emulate = False)
pipe = _Pipe(with_compile = True)
applied = apply_speed_optims(
pipe, _target(), is_gguf = False, family = _family(), speed_mode = SPEED_DEFAULT
)
assert applied["compiled"] is True
assert cfg.emulate_precision_casts is True
def test_snapshot_restores_emulate_precision_casts(monkeypatch):
# The flag is process-global, so the unload path must restore the pre-load value
# exactly like the TF32 / cudnn.benchmark globals.
torch = _stub_torch(monkeypatch)
cfg = _stub_inductor_config(monkeypatch, torch, emulate = False)
snap = snapshot_backend_flags()
assert snap["inductor_emulate_precision_casts"] is False
cfg.emulate_precision_casts = True
restore_backend_flags(snap)
assert cfg.emulate_precision_casts is False
def test_missing_inductor_config_is_tolerated(monkeypatch):
# A build without torch._inductor (or with the flag renamed) must neither break the
# snapshot nor the compile path.
_stub_torch(monkeypatch) # the stub torch has no _inductor attribute
_stub_gguf_accel(monkeypatch)
snap = snapshot_backend_flags()
assert "inductor_emulate_precision_casts" not in snap
pipe = _Pipe(with_compile = True)
applied = apply_speed_optims(
pipe, _target(), is_gguf = False, family = _family(), speed_mode = SPEED_DEFAULT
)
assert applied["compiled"] is True
def test_regional_compile_arms_cache_hook_inners(monkeypatch):
# The production load order engages the step cache BEFORE compile, so the regional
# compile pass must re-arm the already-installed cache hooks with compiled inner
# forwards (otherwise every computed step runs eager under the hook's
# torch.compiler.disable; measured 1.69 vs 1.09 s/step on HunyuanVideo-1.5-720p).
_stub_torch(monkeypatch)
_stub_gguf_accel(monkeypatch)
from core.inference import diffusion_cache as dc_mod
armed = []
monkeypatch.setattr(
dc_mod,
"_compile_hooked_block_inners",
lambda transformer, logger = None: armed.append(transformer) or 1,
)
pipe = _Pipe(with_compile = True)
applied = apply_speed_optims(
pipe, _target(), is_gguf = False, family = _family(), speed_mode = SPEED_DEFAULT
)
assert applied["compiled"] is True
assert armed == [pipe.transformer]