Studio diffusion (Phase 16) review round 2: native CPU arbiter, status offload, load race
Codex review on the native-engine routing: - The /images/load route took the GPU arbiter (acquire_for(DIFFUSION) -> evict chat) unconditionally after engine selection. A native sd.cpp load on a pure-CPU host never touches the GPU, so that needlessly tore down the resident chat model. The handoff is now gated: diffusers always takes it, a force-native sd.cpp load on a CUDA/XPU/MPS box still takes it, but a native sd.cpp load on a CPU host skips it. - sd_cpp status() hardcoded offload_policy 'none' / cpu_offload False even when _run_load computed real offload flags (balanced/low_vram/cpu_offload off-CPU), so the setting was unverifiable. status now derives them from state.offload_flags (still 'none' on CPU, where the flags are empty). - _run_load committed the new state without cancelling/waiting on a generation that started during the (slow) asset download, so a stale sd-cli run against the OLD model could finish afterward and persist an image from the previous model once the new load reported ready. The commit now signals the in-flight cancel and waits on _generate_lock before swapping _state (taken only at commit, so the download never serialises against generation), mirroring the diffusers load path. Tests: CPU native load skips the arbiter while a GPU native load takes it; status reports offload active when flags are set; _run_load cancels and waits for an in-flight generation before committing.
This commit is contained in:
parent
656d11c731
commit
8b385c44fb
4 changed files with 156 additions and 8 deletions
|
|
@ -336,11 +336,25 @@ class SdCppDiffusionBackend:
|
|||
# than commit a "ready" state that crashes on the first generation.
|
||||
if engine.version() is None:
|
||||
raise RuntimeError("sd-cli binary is present but not runnable.")
|
||||
# A generation that started during the (slow) asset download is still running
|
||||
# against the OLD model. Abort it, then WAIT on _generate_lock for it to exit
|
||||
# before publishing the new state -- otherwise that stale sd-cli run can finish
|
||||
# afterward and persist an image from the previous model once this load reports
|
||||
# ready (mirrors the diffusers load_pipeline commit). _generate_lock is taken
|
||||
# only here, not during the download, so the long fetch never serialises against
|
||||
# generation; the inner token re-check guards an unload/newer load arriving while
|
||||
# we waited.
|
||||
with self._lock:
|
||||
if self._load_token != _load_token:
|
||||
return # superseded / cancelled
|
||||
self._state = state
|
||||
self._loading = None
|
||||
if self._active_generate_cancel is not None:
|
||||
self._active_generate_cancel.set()
|
||||
with self._generate_lock:
|
||||
with self._lock:
|
||||
if self._load_token != _load_token:
|
||||
return # superseded / cancelled while waiting
|
||||
self._state = state
|
||||
self._loading = None
|
||||
except SdCppCancelled:
|
||||
return
|
||||
except Exception as exc: # noqa: BLE001 -- surfaced via load_progress
|
||||
|
|
@ -596,8 +610,12 @@ class SdCppDiffusionBackend:
|
|||
"base_repo": state.base_repo,
|
||||
"device": state.device,
|
||||
"dtype": "gguf",
|
||||
"cpu_offload": False,
|
||||
"offload_policy": "none",
|
||||
# Reflect the offload flags actually passed to sd-cli, so a balanced/low_vram
|
||||
# (or cpu_offload) load is verifiable from status instead of always reading
|
||||
# "none". On CPU _run_load leaves offload_flags empty (the flags are no-ops),
|
||||
# so this correctly stays "none" there.
|
||||
"cpu_offload": bool(state.offload_flags),
|
||||
"offload_policy": "active" if state.offload_flags else "none",
|
||||
"vae_tiling": False,
|
||||
"memory_mode": None,
|
||||
"speed_mode": state.native_speed,
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue