Round 14 reviewer aggregate (logs/review_round14_aggregate.md):
P1 fixes:
- routes/export.py /load-checkpoint now runs the active-export 409
guard BEFORE the chat / diffusion unloads, so a rejected request
no longer tears down unrelated GPU state.
- core/inference/llama_cpp.py wraps the WHOLE load_model body in a
single try/finally that publishes loading_model_identifier across
download, metadata read, VRAM settle, process spawn, and health
check. Done via a thin load_model wrapper around the existing
body (renamed _load_model_impl) to avoid reindenting hundreds of
lines.
- routes/models.py /delete-finetuned now checks
loading_model_identifier so a pending HF GGUF download cannot
have its destination directory rmtree'd before llama-server
spawns.
- core/inference/diffusion.py stores the original caller-supplied
gguf_filename (e.g. ``BF16/model.gguf``) in a new self._gguf_filename
field and exposes it as active_gguf_filename. UI-facing
gguf_filename still collapses to basename for the panel.
- routes/models.py /delete-cached llama guard now allows safe
different-variant deletes when hf_variant differs, matching the
diffusion path's variant-aware behaviour.
- core/inference/diffusion.py tracks self._cpu_offload_enabled and
forces a CPU torch.Generator when offload is on, so seeded
generation no longer crashes on CUDA hosts with the default offload
enabled.
P2 fixes:
- core/inference/diffusion.py detect_family normalises mixed
separators (``Qwen_Image-Edit-GGUF``, ``Qwen-Image_Edit-GGUF``,
``QwenImageEdit-GGUF``) so every Qwen-Image-Edit spelling is
excluded from the base Qwen-Image family.
- core/inference/diffusion.py logger.info / logger.error in
load_model run repo_id and effective_base through _redact_hf_tokens
so URL-embedded ``hf_xxxxx`` tokens never reach structured-log
sinks.
- core/inference/diffusion.py _release_other_gpu_owners_for_diffusion
now raises RuntimeError when an export job is active instead of
logging and continuing, so direct backend callers cannot bypass
the route layer's 409 guard.
- core/inference/diffusion.py full-diffusers repo / base_repo paths
expand ``~`` via _expand_existing_local_path so
``repo_id="~/models/my-flux"`` no longer falls through to the Hub.
Tests:
- 5 new regression cases (mixed Qwen-Image-Edit separators, token
redaction, status full-filename, CPU offload generator device,
staging Windows leaf already-set sanity).
- All 68 diffusion backend + route tests pass.