Two evict/OOM fixes on the diffusion load paths: - The video load moved a pipeline onto the GPU (apply_memory_plan) and committed it while holding no lock, so an unload / GPU-arbiter eviction -- which bumps the load token and then barriers on _generate_lock before freeing -- could hand VIDEO to chat/images and let the new owner allocate concurrently with the in-flight placement, OOMing. Hold _generate_lock across placement + the locked commit, mirroring the image backend, so an evicting owner waits until this worker's placement is torn down or committed. Lock order stays _generate_lock -> _lock (unload takes _lock then releases it before the barrier), so there is no deadlock. - resolve_local_single_file reinterpreted an On-Device folder as a base single_file load whenever it held exactly one .safetensors, so a PEFT LoRA adapter folder (adapter_config.json + adapter_model.safetensors) with a family-token name was picked as a base checkpoint, evicting the resident model before from_single_file failed on the adapter weights. Skip adapter folders (adapter_config.json) and the adapter_model basename so the pick stays a pipeline load and 400s in validation, before the GPU handoff. Adds regression tests for both. |
||
|---|---|---|
| .. | ||
| assets | ||
| auth | ||
| core | ||
| hub | ||
| loggers | ||
| models | ||
| plugins | ||
| requirements | ||
| routes | ||
| state | ||
| storage | ||
| tests | ||
| utils | ||
| __init__.py | ||
| _platform_compat.py | ||
| cloudflare_tunnel.py | ||
| colab.py | ||
| main.py | ||
| run.py | ||
| startup_banner.py | ||