Two fixes for GGUF variant dropdown: 1. useGpuInfo now sums memory across all GPU devices instead of only reading devices[0]. This matches llama-server's multi-GPU allocation where models can be split across GPUs. 2. When the backend-recommended variant (e.g. UD-Q4_K_XL) exceeds total GPU VRAM, the frontend picks the largest variant that fits instead. If all variants are OOM, it recommends the smallest one (most likely to work with --fit). |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| __init__.py | ||
| install_python_stack.py | ||
| LICENSE.AGPL-3.0 | ||
| setup.bat | ||
| setup.ps1 | ||
| setup.sh | ||
| Unsloth_Studio_Colab.ipynb | ||