Two fixes for accurate GGUF OOM detection: 1. /api/system now uses nvidia-smi to enumerate all physical GPUs instead of torch.cuda which only sees CUDA_VISIBLE_DEVICES. This matches llama-server which can use all GPUs regardless of the env var. Falls back to torch-based detection if nvidia-smi unavailable. 2. Frontend GGUF OOM check now uses 70% of total GPU memory as the budget, matching the PR's _select_gpus logic (30% reserved for KV cache and compute buffers). Previously used checkVramFit's 100% threshold which was too generous. |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| __init__.py | ||
| install_python_stack.py | ||
| LICENSE.AGPL-3.0 | ||
| setup.bat | ||
| setup.ps1 | ||
| setup.sh | ||
| Unsloth_Studio_Colab.ipynb | ||