- detect_family adds _FAMILY_EXCLUDE so 'stable-diffusion-3.5' no longer matches the SD3 Medium family and 'qwen-image-edit' no longer matches Qwen-Image. Both were misleading silent loads. - from_single_file now forwards config=<effective_base>, subfolder='transformer', and the HF token. Diffusers-format GGUFs (FLUX.2 klein, Qwen-Image, SD3) need the matching base config or the transformer load picks the wrong shapes; gated GGUFs need the token both for download and config read. - Move _release_chat_backend_for_diffusion + new _release_other_gpu_owners_for_diffusion to AFTER the GGUF download and pipeline class lookup so a typo or transient Hub error does not kill the user's currently-loaded chat model. Peak VRAM still stays at one model's worth because the releases run right before from_pretrained. - _release_other_gpu_owners_for_diffusion: shut down the export subprocess and any active training subprocess before a diffusion load. Symmetric with the export load path. - routes/training.py: unload diffusion before starting training so the new subprocess does not race FLUX/Qwen for VRAM. - routes/export.py: also unload the GGUF llama-server before export load (the existing inference-backend unload only covered the safetensors path). |
||
|---|---|---|
| .. | ||
| assets | ||
| auth | ||
| core | ||
| loggers | ||
| models | ||
| plugins | ||
| requirements | ||
| routes | ||
| state | ||
| storage | ||
| tests | ||
| utils | ||
| __init__.py | ||
| _platform_compat.py | ||
| colab.py | ||
| main.py | ||
| run.py | ||
| startup_banner.py | ||