- diffusion_attention: clear the HunyuanVideo-1.5 null-mask flag with an always_call post-hook so it is scoped to one hooked forward and never latches across an exception; add attention_backend_supported_on_device to arch-gate an already-resolved backend on a specific (heterogeneous) CUDA device. - video: make the explicit MagCache resize transactional via _step_cache_all_or_none (refuse to stack a fresh cache over one that could not be disabled; roll a mixed resize back and report the true state); raise on a failed all-or-none rollback instead of falsely reporting an uncached pipeline. - diffusion_cfg_parallel: re-validate the attention backend on the replica device and pin native there when unsupported; mirror the primary's max tier on the replica (max-autotune compile + direct QKV fusion) via a new speed_mode arg; prefer a viable heterogeneous secondary GPU over an unusable identical one; clear the const cache at each plan_generation. - diffusion_vae_quant / diffusion_precision: detect a partial diffusers layerwise-fp8 mutation (leftover casting hooks the torchao detector cannot see) and fail the load closed, while a clean failure still falls back to dense. - video_speedmem_bench: engage the dual-expert cache all-or-none like the loader. - frontend video api: add text_encoder_quant / vae_quant and the auto/off literals to VideoLoadRequest so typed callers match the backend contract. |
||
|---|---|---|
| .. | ||
| assets | ||
| auth | ||
| core | ||
| hub | ||
| loggers | ||
| models | ||
| plugins | ||
| requirements | ||
| routes | ||
| state | ||
| storage | ||
| tests | ||
| utils | ||
| __init__.py | ||
| _platform_compat.py | ||
| cloudflare_tunnel.py | ||
| colab.py | ||
| main.py | ||
| run.py | ||
| startup_banner.py | ||