Review follow-ups on the video inference backend: - validate_load_request now rejects a -GGUF repo picked as a diffusers pipeline (no gguf_filename) up front, instead of failing minutes later in from_pretrained after the GPU owner was already evicted. - New _detect_load_family helper shared by validate_load_request and _run_load: when the repo id alone does not carry the family, fall back to detecting it from the picked GGUF filename, so both paths agree. - routes/video.py now threads base_repo into validate_load_request so an untrusted companion repo is refused before the arbiter handoff. - unload() now drains _generate_lock before _teardown_state so a cancelled clip actually exits the denoise loop before the VRAM is reported free. - load_pipeline re-checks the load token after the generate-lock barrier and raises if the load was superseded while waiting. - Pre-commit global mutations (backend flags, gguf compile installs) are registered per load token and rolled back in _run_load's error path via _rollback_precommit_globals, so a failed load no longer leaks process-wide state. - fp32 memory estimates now apply a 2x dtype scale on non-CPU devices for pipeline, single-file and companion sizes (bf16 tables assume 2 bytes/param); GGUF quant estimates stay unscaled. Tests: GGUF-repo-as-pipeline rejection, _detect_load_family fallback and override semantics; fake route backend accepts base_repo. 66 passed across test_video_backend, test_video_routes, test_video_families, test_video_gallery. |
||
|---|---|---|
| .. | ||
| assets | ||
| auth | ||
| core | ||
| hub | ||
| loggers | ||
| models | ||
| plugins | ||
| requirements | ||
| routes | ||
| state | ||
| storage | ||
| tests | ||
| utils | ||
| __init__.py | ||
| _platform_compat.py | ||
| cloudflare_tunnel.py | ||
| colab.py | ||
| main.py | ||
| run.py | ||
| startup_banner.py | ||