Studio diffusion (Phase 10): reset the global attention backend on native, gate arch-specific kernels, accept sdpa
- apply_attention_backend now restores the native default when no backend is requested or a
kernel fails. diffusers keeps a process-wide active attention backend that
set_attention_backend updates, and a fresh transformer's processors follow it, so a load
that wanted native could silently inherit a backend (e.g. cuDNN) an earlier speed-profile
load pinned, breaking the bit-identical/off guarantee.
- select_attention_backend drops flash3/flash4 up front when the CUDA capability is below
Hopper/Blackwell. diffusers only checks the kernels package at set time, so an explicit
request on the wrong card set fine then crashed mid-generation; it now falls back to native.
- Add the sdpa alias to the attention_backend Literal so an API request with sdpa (already a
valid alias of native) is accepted instead of 422-rejected by Pydantic.
- Drop the dead replace('-','_') normalization (no alias uses dashes/underscores).
- perf_levers_probe.py output dir is now relative to the script, not a hardcoded path.
This commit is contained in:
parent
c7e30e9ea8
commit
cf5f23996f
4 changed files with 152 additions and 18 deletions
|
|
@ -25,7 +25,7 @@ import numpy as np
|
|||
|
||||
BASE = "Tongyi-MAI/Z-Image-Turbo"
|
||||
PROMPT = "A cinematic photograph of a red fox in a snowy forest at dawn, highly detailed"
|
||||
OUT = Path("/mnt/disks/unslothai/ubuntu/workspace_81/outputs/quant_research/perf_levers_images")
|
||||
OUT = Path(__file__).resolve().parent.parent / "outputs" / "quant_research" / "perf_levers_images"
|
||||
|
||||
|
||||
_LP = {"fn": None}
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue