Review follow-ups on the video inference backend:
- validate_load_request now rejects a -GGUF repo picked as a diffusers
pipeline (no gguf_filename) up front, instead of failing minutes later
in from_pretrained after the GPU owner was already evicted.
- New _detect_load_family helper shared by validate_load_request and
_run_load: when the repo id alone does not carry the family, fall back
to detecting it from the picked GGUF filename, so both paths agree.
- routes/video.py now threads base_repo into validate_load_request so an
untrusted companion repo is refused before the arbiter handoff.
- unload() now drains _generate_lock before _teardown_state so a
cancelled clip actually exits the denoise loop before the VRAM is
reported free.
- load_pipeline re-checks the load token after the generate-lock barrier
and raises if the load was superseded while waiting.
- Pre-commit global mutations (backend flags, gguf compile installs) are
registered per load token and rolled back in _run_load's error path
via _rollback_precommit_globals, so a failed load no longer leaks
process-wide state.
- fp32 memory estimates now apply a 2x dtype scale on non-CPU devices
for pipeline, single-file and companion sizes (bf16 tables assume
2 bytes/param); GGUF quant estimates stay unscaled.
Tests: GGUF-repo-as-pipeline rejection, _detect_load_family fallback and
override semantics; fake route backend accepts base_repo. 66 passed
across test_video_backend, test_video_routes, test_video_families,
test_video_gallery.
The load tail re-ran the already-filtered speed_optims tuple through
.items() as if it were still the raw applied dict from apply_speed_optims.
An empty tuple short-circuited to {} so CPU test runs passed, but on a real
GPU at least channels_last engages, the tuple is truthy, and every load
failed with 'tuple' object has no attribute 'items'. Store the filtered
tuple directly and add a regression test that forces one optimisation to
engage.
tokenizer/chat_template.jinja ships as its own file in the LTX-2 and
HunyuanVideo-1.5 repos and apply_chat_template reads it at generation
time, so a scoped snapshot without it loads fine and then crashes the
first generation.
A dual-DiT pipeline (Ideogram 4's unconditional tower) placed its second
denoiser resident under the group tier, which defeats the tier since the
pair rarely fits where one alone did not. Stream transformer_2 and
unconditional_transformer alongside the transformer and keep only the
smaller companions resident.
The trusted-repo allowlist admitted the 720p t2v repack while the only
Hunyuan family entry carried 480p presets, so a 720p load silently
defaulted to 832x480. It now resolves a dedicated entry whose repo-id
alias outranks the generic token (same guider config, verified: both
repos ship guidance 6.0).
A scheduler-wrapped cancel unwinds pipe call by exception and skips the
pipeline's end-of-call maybe_free_model_hooks, leaving onloaded offload
modules on the GPU until the next request; generate() now frees them
before surfacing the cancelled sentinel.
A warm-cache predownload sweep never consults the cancel event (each cached
file returns instantly), so an unload during it was ignored until a cold
file hit the network. The 2.3 detection also only probed bare-file local
repos; resolve directory repos through the same child resolver the loader
uses so their base pull is scoped too.
- Run the trainer's caption discovery in the start route BEFORE freeing GPU
residents, so a missing or uncaptionable dataset 400s without evicting the
loaded chat/Images model.
- sd.cpp unload now waits out a cancelled one-shot generation on the generate
lock before reporting the device free, matching the diffusers backend.
- Clearing a caption that came from metadata.jsonl writes an empty sidecar
tombstone instead of unlinking (both readers treat an existing sidecar as
authoritative), so the cleared label cannot resurface.
- The ControlNet wrapper pipe is only cached while its load is still current,
closing the unload race the model cache already handled.
The prefetch already scopes the file list (no packaged root singles, no
dtype-variant twins, no ONNX/Flax exports), but from_pretrained was then
called with the hub id, and its own snapshot sweep re-downloaded the
skipped files anyway: 24 GB per FLUX.1 repo and 65 GB on FLUX.2-dev, as
found in the blob cache. Return the snapshot dir from the prefetch (keyed
on the pipeline manifest) and hand it to every pipeline-assembly
from_pretrained site; any prefetch failure keeps the hub id and the old
behavior.
A 2.3 checkpoint (GGUF or single file) carries the DiT and, with its extras
files, the connectors, both VAEs and the vocoder; only the 2.0 base repo's
scheduler, text encoder and tokenizer are read. Detect 2.3 from the
checkpoint header after the pull, re-estimate, and scope the base
pre-download accordingly (about 6 GB less per fresh install).
A bare from_pretrained snapshot of Lightricks/LTX-2 pulls the whole 314 GB
repo: 170 GB of packaged root checkpoints and a second 50 GB text-encoder
shard set, when the pipeline reads about 93 GB. Build the needed file list
once (shared with the progress estimate so the two cannot disagree),
download it per file with cancellation, and hand from_pretrained the local
snapshot dir. Clamp the progress counter to the estimate so stale cache
blobs can no longer report over 100 percent.