Reject partial dual-DiT quantization in video_speedmem_bench (the loader fails that load all-or-none; a mixed quantized/dense row is unloadable), toggle the generation-time FBCache recheck on every expert view like the loader's per-view iteration, and rescore lpips_vs_reference in a post-pass so a --configs order that lists reference late no longer publishes null. In quant_speedmem_bench, track per-encoder engagement via a weight-storage fingerprint so a partial multi-encoder cast cannot certify a still-dense encoder with a ~1.0 cosine, and load vae_force_fp32 families (Wan) at fp32 with a matching latent dtype so the dense VAE row measures what production runs. Gate the attention-trim tests with pytest.importorskip so a no-torch environment keeps the backend test suite collectable. |
||
|---|---|---|
| .. | ||
| assets | ||
| auth | ||
| core | ||
| hub | ||
| loggers | ||
| models | ||
| plugins | ||
| requirements | ||
| routes | ||
| state | ||
| storage | ||
| tests | ||
| utils | ||
| __init__.py | ||
| _platform_compat.py | ||
| cloudflare_tunnel.py | ||
| colab.py | ||
| main.py | ||
| run.py | ||
| startup_banner.py | ||