Round 16 reviewer aggregate (logs/review_round16_aggregate.md):
P1 fixes:
- routes/models.py /delete-cached llama guard pairs loading_id with
loading_hf_variant so deleting a different cached quant (Q8_0)
while another variant (Q4_K_M) is loading is no longer blocked.
- core/inference/diffusion.py load_model now calls
_release_other_gpu_owners_for_diffusion BEFORE
_release_chat_backend_for_diffusion. The other-owners helper
RAISES on active training/export, so a route -> worker race or
direct backend caller no longer drops the user's chat model
before the diffusion load is refused.
- routes/models.py /delete-cached diffusion guard fails CLOSED
(503) on HF cache scan failure instead of silently falling
through to repo-id-only matching, which could miss a loaded
local snapshot path.
- routes/inference.py _release_llama_for and
_release_safetensors_chat_for now raise 503 on actual unload
failure (exception or False return), so new GPU workloads do
not start while the old chat process still owns VRAM.
- core/inference/diffusion.py status() now takes
include_internal=False by default and only exposes the
guard-facing active_*/pending_* paths when callers opt in. The
public /api/inference/images/status route gets the redacted
payload; routes/models.py delete guards pass
include_internal=True so they still see the raw paths.
- core/inference/diffusion.py generate_image_with_metadata routes
the response model through _display_repo_id so /images/generate
cannot echo back an absolute local path.
P2 fixes:
- routes/inference.py /images/load now maps backend "Could not
verify training/export status" to 503 instead of 409, matching
the route-level pre-check.
- core/inference/diffusion.py _release_other_gpu_owners_for_diffusion
raises "Could not verify export status" when the
is_export_active() probe itself raises, instead of silently
treating it as active export.
- core/inference/diffusion.py detect_family compares compact family
spellings (Flux2Klein) against per-token compact strings so
unsloth/Flux2Klein-GGUF matches the flux.2-klein family without
matching the embedded substring inside flux.20.
- main.py installs a RequestValidationError handler that scrubs
hf_xxxxx tokens out of the 422 response body so a rejected
``repo_id`` containing a URL-embedded HF token does not echo it
back to the browser.
Tests:
- 3 new regression cases (Flux2Klein compact alias, public status
redaction, generate_image_with_metadata redaction).
- All 75 diffusion backend + route tests pass.