Drop the now-redundant second-expert group offload on the A14B load
After the upstream cascade generalised _apply_group_offload to stream every DiT it finds on the pipe (transformer AND transformer_2), apply_memory_plan's group tier already block-streams both A14B experts, and model/sequential offload hook every top-level module. The explicit per-expert pass this branch added is therefore redundant: on a real group-tier load it re-registers the group-offload hooks transformer_2 already carries, which diffusers rejects with a duplicate-hook ValueError. That was swallowed by the helper's try/except and surfaced as a misleading "second expert stays resident" warning even though the expert was in fact streamed. Remove the pass and document why MoE needs none.
This commit is contained in:
parent
c68c9fab9e
commit
1d193f58bb
1 changed files with 6 additions and 12 deletions
|
|
@ -1004,18 +1004,12 @@ class VideoBackend:
|
|||
clear_gpu_cache()
|
||||
raise RuntimeError("Video load was cancelled or superseded.")
|
||||
offload_policy, vae_tiling = apply_memory_plan(pipe, plan, device = device, logger = logger)
|
||||
if offload_policy == "group" and len(views) > 1:
|
||||
# Group offload streams only ``pipe.transformer``; the second expert would
|
||||
# otherwise sit resident (~57 GB bf16 on the A14B) and defeat the tier.
|
||||
# model/sequential offload hook every top-level module, so only group needs
|
||||
# this. Applied through the view so the helper streams transformer_2.
|
||||
from .diffusion_memory import _apply_group_offload
|
||||
for view in views[1:]:
|
||||
if not _apply_group_offload(view, device, logger):
|
||||
logger.warning(
|
||||
"video.memory: group offload did not engage on the second "
|
||||
"expert; it stays resident"
|
||||
)
|
||||
# A dual-DiT MoE pipe (Wan2.2-A14B) needs no extra per-expert offload pass here:
|
||||
# apply_memory_plan's group tier (_apply_group_offload) already block-streams every
|
||||
# DiT it finds on the pipe -- transformer AND transformer_2 -- and model/sequential
|
||||
# offload hook every top-level module, so the second expert is covered under all tiers.
|
||||
# A second _apply_group_offload on transformer_2 would re-register the group-offload
|
||||
# hooks it already carries, which diffusers rejects with a duplicate-hook ValueError.
|
||||
if not vae_tiling:
|
||||
# Decode of a whole clip is the video memory peak; tiling is near-free
|
||||
# in quality and keeps the decode bounded, so it is always on.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue