Review follow-ups on the video inference backend:
- validate_load_request now rejects a -GGUF repo picked as a diffusers
pipeline (no gguf_filename) up front, instead of failing minutes later
in from_pretrained after the GPU owner was already evicted.
- New _detect_load_family helper shared by validate_load_request and
_run_load: when the repo id alone does not carry the family, fall back
to detecting it from the picked GGUF filename, so both paths agree.
- routes/video.py now threads base_repo into validate_load_request so an
untrusted companion repo is refused before the arbiter handoff.
- unload() now drains _generate_lock before _teardown_state so a
cancelled clip actually exits the denoise loop before the VRAM is
reported free.
- load_pipeline re-checks the load token after the generate-lock barrier
and raises if the load was superseded while waiting.
- Pre-commit global mutations (backend flags, gguf compile installs) are
registered per load token and rolled back in _run_load's error path
via _rollback_precommit_globals, so a failed load no longer leaks
process-wide state.
- fp32 memory estimates now apply a 2x dtype scale on non-CPU devices
for pipeline, single-file and companion sizes (bf16 tables assume
2 bytes/param); GGUF quant estimates stay unscaled.
Tests: GGUF-repo-as-pipeline rejection, _detect_load_family fallback and
override semantics; fake route backend accepts base_repo. 66 passed
across test_video_backend, test_video_routes, test_video_families,
test_video_gallery.
The latent cache holds two fp32 posterior tensors per crop/flip variant per
image, pinned on CUDA hosts, so datasets with thousands of images can exhaust
host or pinned memory with no fallback. Estimate the cache size from the first
real encoded latent and fall back to per-step VAE encoding when it exceeds a
4 GiB budget. UNSLOTH_DIFFUSION_FORCE_LATENT_CACHE bypasses the gate; the
existing UNSLOTH_DIFFUSION_NO_LATENT_CACHE opt-out is unchanged.
Replace the with-replacement per-batch index draw in the SDXL and DiT LoRA
trainers with a shared PermutationBatchSampler that visits every image once per
cycle before repeating, so short runs cover the whole dataset. The sampler
reshuffles from the run's rng so the index stream stays seed-deterministic.
Guard the diffusion run detail route against a valid-JSON non-object record,
which previously raised TypeError and returned a 500; it now 404s like the list
path's shape check.
Add regression tests for both.
The latest huggingface-hub release added the Sandboxes feature. Its
bootstrap (_sandbox.py) fetches the static sbx-server binary into /tmp with
an Authorization header and marks it executable, which is exactly the
staged-dropper pattern the scanner hunts, and three while True polling loops
in _sandbox.py / hf_api.py / utils/_http.py match the beaconing heuristic.
All four verified against the official huggingface/huggingface_hub
repository: the snippet is the documented sandbox server injection and the
loops are deadline-style job and sandbox polling. Entries generated with
--write-baseline and reviewed line by line; scan_packages.py huggingface-hub
now exits 0 with the four findings suppressed.