1. Backend: When a model fails with "No config file found" or similar
unsupported-model errors, wrap the message with "This model is not
supported yet. Try a different model." instead of showing the raw
Unsloth exception.
2. Frontend: Compute estimated download size from the HF search API's
safetensors.parameters dtype breakdown (BF16=2B/param, I32=4B/param,
F32=4B/param, etc.) and show it in the model picker instead of just
the param count. For example, Kimi-K2.5 now shows "~554 GB" instead
of "171B" (which was misleading since 171B params != 171GB download).
Updated GGUF fit classification to match llama-server's --fit behavior:
- fits: model <= 70% of total GPU memory (all GPUs)
- tight: model > 70% GPU but <= 70% GPU + 70% available system RAM
(llama-server uses --fit to offload layers to CPU)
- OOM: model exceeds both GPU and system RAM budgets
useGpuInfo now also returns systemRamAvailableGb from /api/system so the
frontend can compute the combined GPU+RAM budget.
Two fixes for GGUF variant dropdown:
1. useGpuInfo now sums memory across all GPU devices instead of only
reading devices[0]. This matches llama-server's multi-GPU allocation
where models can be split across GPUs.
2. When the backend-recommended variant (e.g. UD-Q4_K_XL) exceeds total
GPU VRAM, the frontend picks the largest variant that fits instead.
If all variants are OOM, it recommends the smallest one (most likely
to work with --fit).
GGUF was in the global EXCLUDED_TAGS set which filtered it from all
consumers of useHfModelSearch, including the chat page. Move GGUF
exclusion to an opt-in excludeGguf option so only training and
onboarding pages filter out GGUF models.