Previously only GGUF models showed download progress in Chat. Non-GGUF models (safetensors, bnb quantized, etc.) showed a static message with no progress indication. This adds progress tracking for all model types and fixes several related issues. Backend: - Add /api/models/download-progress endpoint that checks the HF cache blobs directory for completed and .incomplete files. Uses model_info() (cached per repo) to determine expected total size for percentage. - Add /api/models/cached-models endpoint that lists non-GGUF model repos from the HF cache via scan_cache_dir(). - Fix progress stuck at 0.99: when no .incomplete files remain, report 1.0 immediately (blob deduplication can make byte totals mismatch). Frontend: - Remove the ggufVariant gate so download progress polling works for all non-cached models, not just GGUFs. - Use GGUF-specific endpoint when variant + expectedBytes available, otherwise use the general download-progress endpoint. - Fix toast stuck after load: check loadingModelRef.current before and after the async poll to prevent overwriting the success toast. - First poll at 500ms instead of waiting for the 2s interval. - Show downloaded non-GGUF models in the Hub model picker "Downloaded" section alongside GGUFs. |
||
|---|---|---|
| .. | ||
| assets | ||
| auth | ||
| core | ||
| loggers | ||
| models | ||
| plugins | ||
| requirements | ||
| routes | ||
| state | ||
| tests | ||
| utils | ||
| __init__.py | ||
| colab.py | ||
| main.py | ||
| run.py | ||