Codex review on the native-engine routing:
- The /images/load route took the GPU arbiter (acquire_for(DIFFUSION) -> evict chat)
unconditionally after engine selection. A native sd.cpp load on a pure-CPU host
never touches the GPU, so that needlessly tore down the resident chat model. The
handoff is now gated: diffusers always takes it, a force-native sd.cpp load on a
CUDA/XPU/MPS box still takes it, but a native sd.cpp load on a CPU host skips it.
- sd_cpp status() hardcoded offload_policy 'none' / cpu_offload False even when
_run_load computed real offload flags (balanced/low_vram/cpu_offload off-CPU), so
the setting was unverifiable. status now derives them from state.offload_flags
(still 'none' on CPU, where the flags are empty).
- _run_load committed the new state without cancelling/waiting on a generation that
started during the (slow) asset download, so a stale sd-cli run against the OLD
model could finish afterward and persist an image from the previous model once the
new load reported ready. The commit now signals the in-flight cancel and waits on
_generate_lock before swapping _state (taken only at commit, so the download never
serialises against generation), mirroring the diffusers load path.
Tests: CPU native load skips the arbiter while a GPU native load takes it; status
reports offload active when flags are set; _run_load cancels and waits for an
in-flight generation before committing.