Round 41 review findings (2 P1, 5/12 reviewers consensus on the dominant one):
1. routes/export.py: load_checkpoint already refuses 409 when training
or another export is active, but /export/{merged,base,gguf,lora} and
/cleanup went through _export_public_window without those checks.
A user could start training, then trigger an export (or cleanup),
and both would double-own the GPU. Factor the training-active and
export-active guards into _raise_if_training_active_for_export and
_raise_if_export_active_for_export, call them inside the context
manager so all /export/* + /cleanup share the same fail-closed
semantics as load_checkpoint, and wrap /cleanup with the window.
2. core/inference/diffusion.py: DiffusionBackend.unload_model cleared
_pipe / _repo_id / _family / ... under _lock BEFORE _release(old)
and _drain_cuda_cache. Between the lock release and cache drain,
status() reported is_loaded=False / is_loading=False, so the
helper-busy check (which OR-s those two) could let an AI Assist
GGUF backend start while diffusion VRAM was still being freed.
Set _loading=True inside the lock as a busy marker before clearing
the slot, and only clear it in a finally after release + drain
complete.