1. Interruptible downloads: load_model now checks a cancel event between shard downloads. unload_model sets the event so cancel stops the download at the next shard boundary. 2. /api/models/cached-gguf endpoint: scans the HF cache for already-downloaded GGUF repos with their total size and cache path. 3. "Downloaded" section in Hub model picker: shows cached GGUF repos at the top (before Recommended) so users can quickly re-load previously downloaded models without re-downloading. |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| audio_codecs.py | ||
| inference.py | ||
| llama_cpp.py | ||
| orchestrator.py | ||
| worker.py | ||