Commit graph

885 commits

Author SHA1 Message Date
Shine1i
227e6a7402 Improve model loading toast UX 2026-03-15 22:25:32 +01:00
Daniel Han
ac99649c47 studio: toast UX -- 5s auto-dismiss with close button
- Default toast duration 5s (was infinite for loading toasts)
- Add closeButton to Sonner Toaster for manual dismiss via X
- Sonner pauses dismiss timer on hover automatically
2026-03-15 12:27:51 +00:00
pre-commit-ci[bot]
050240b27a [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-03-15 05:24:06 -07:00
Daniel Han
11612f6dc9 studio: fix GGUF download UX -- progress bar, cancel, sorting, auto-scroll
- Run GGUF load_model in asyncio.to_thread so the event loop stays free
  for progress polling during download (was blocking all requests).
- Extract download phase out of the lock in LlamaCppBackend.load_model
  so unload_model/cancel can take effect immediately during download.
- Fix "downloaded" badge for split GGUFs: check total cached bytes
  across all shards vs expected size, not just first shard existence.
- Respect CUDA_VISIBLE_DEVICES in /api/system GPU reporting so the
  frontend GGUF fit estimation uses actual available VRAM.
- Sort tight variants (need CPU offload) smallest-first instead of
  largest-first -- closer to GPU budget = faster inference.
- Fix cancel: use refs instead of React state for abort controller and
  toast ID so both cancel buttons (text + toast) work reliably. Make
  cancel synchronous (fire-and-forget unload) for instant UI response.
  Check abortCtrl.signal.aborted after loadModel returns to prevent
  ghost model state. Skip rollback and suppress errors on cancel.
- Dynamic top 4 GGUF models fetched from HF API sorted by downloads,
  prepended to the default recommended list.
- Remove turnAnchor="top" for auto-scroll to bottom during generation.
- Set default toast duration to 10s (was infinite for loading toasts).
- Deduplicate cached GGUF repos using scan_cache_dir API (fixes
  Qwen/X-GGUF vs qwen/x-gguf duplicates from lowercased HF cache).
- Pre-compile repo_id validation regex to silence CodeQL ReDoS warning.
- Change welcome text and default suggestion text.
2026-03-15 05:24:06 -07:00
Daniel Han
bb57236e29 studio: revert -- always respect CUDA_VISIBLE_DEVICES in GPU memory query 2026-03-15 05:24:06 -07:00
pre-commit-ci[bot]
851cb2af68 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-03-15 05:24:06 -07:00
Daniel Han
5603ced75f studio: ignore CUDA_VISIBLE_DEVICES in GPU memory query for llama-server
_get_gpu_free_memory was filtering by CUDA_VISIBLE_DEVICES, so with
CUDA_VISIBLE_DEVICES='0' set by the training env, llama-server only
saw 1 GPU and used --fit for CPU offloading instead of spreading
across all 8 GPUs.

Since llama-server manages its own GPU allocation (the _select_gpus
method picks GPUs and sets CUDA_VISIBLE_DEVICES for the subprocess),
the query must see ALL physical GPUs to make the right decision.
2026-03-15 05:24:06 -07:00
Daniel Han
1dfba866be studio: fix download progress -- track per-variant, include incomplete blobs
1. Progress endpoint now takes a variant parameter and only counts
   .gguf files matching that variant (not all files in the repo cache,
   which would include previously downloaded variants)

2. Tracks .incomplete files in HF blobs dir for in-progress single-shard
   downloads, capping at 99% until the file is fully committed

3. Fixed loading text: "Loading model..." for cached, "Downloading
   model..." for new downloads, with appropriate descriptions

4. Wording: "Downloading and loading model. Large models can take a
   while." instead of "This may include downloading."
2026-03-15 05:24:06 -07:00
pre-commit-ci[bot]
b1dda44745 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-03-15 05:24:06 -07:00
Daniel Han
475ba417dc studio: context-aware loading text + download progress bar
1. Loading text: shows "Loading model..." for cached models,
   "Downloading model..." for new downloads. Toast description
   adapts accordingly.

2. Download progress: polls /api/models/gguf-download-progress every
   2s during downloads, updating the toast with percentage and GB
   downloaded. Progress is estimated by checking the HF cache folder
   size against the expected total bytes.

3. Passes isDownloaded and expectedBytes through the full chain from
   variant click to selectModel for accurate UI state.
2026-03-15 05:24:06 -07:00
pre-commit-ci[bot]
061de08f86 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-03-15 05:24:06 -07:00
Daniel Han
1e9d19126b studio: fix P1 issues from PR review comments
1. n_gpu_layers kwarg: accept (and ignore) in load_model signature
   so callers like llm_assist.py don't get TypeError

2. mmproj exclusion: filter out mmproj files in _find_smallest_fitting_variant
   so fallback doesn't pick a tiny vision projection as the "model"

3. Shard preservation after fallback: re-discover shards for the
   fallback variant instead of resetting to empty list, so split
   GGUFs download all shards

4. Orphan cleanup safety: only kill llama-server processes whose
   cmdline contains ".unsloth/", avoiding termination of unrelated
   llama-server instances on the same machine

5. Path expression sanitization: validate repo_id format before using
   it in cache directory lookups
2026-03-15 05:24:06 -07:00
Daniel Han
cf45ff7232 studio: fix downloaded check -- compare basename not full path
The variant filename includes a subfolder prefix (e.g.
UD-Q4_K_XL/Kimi-K2.5-UD-Q4_K_XL-00001-of-00013.gguf) but rglob
returns just the filename. Use Path.name for the comparison.
2026-03-15 05:24:06 -07:00
Daniel Han
92670a90dd studio: fix case-insensitive HF cache lookup for downloaded GGUF variants
HF cache dirs use the exact case from the repo_id at download time
(e.g. models--unsloth--kimi-k2.5-gguf) which may differ from the
canonical HF repo_id (unsloth/Kimi-K2.5-GGUF). Use case-insensitive
matching to find the cache directory.
2026-03-15 05:24:06 -07:00
Daniel Han
7b65073311 studio: show 'downloaded' badge instead of 'recommended' when variant is cached 2026-03-15 05:24:06 -07:00
Daniel Han
bcb382def9 studio: sort downloaded GGUF variants before recommended
Downloaded variants now take priority over the recommended badge in
sort order. Within the same tier (downloaded+fits, etc.), recommended
still sorts first. Order: downloaded -> recommended -> fits -> tight -> OOM
2026-03-15 05:24:06 -07:00
pre-commit-ci[bot]
64ab7554b1 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-03-15 05:24:06 -07:00
Daniel Han
4d35699c65 studio: show downloaded status in GGUF variant list, sort downloaded first
- Backend: /gguf-variants now checks HF cache for each variant's file
  and returns a downloaded flag per variant
- Frontend: downloaded variants sort before non-downloaded (after
  recommended), and show a green "downloaded" badge
- Sort order: recommended -> downloaded+fits -> downloaded+tight ->
  fits -> tight -> OOM
2026-03-15 05:24:06 -07:00
pre-commit-ci[bot]
904ac86f4a [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-03-15 05:24:06 -07:00
Daniel Han
897d8b426a studio: interruptible GGUF downloads, cached models endpoint, Downloaded section
1. Interruptible downloads: load_model now checks a cancel event
   between shard downloads. unload_model sets the event so cancel
   stops the download at the next shard boundary.

2. /api/models/cached-gguf endpoint: scans the HF cache for
   already-downloaded GGUF repos with their total size and cache path.

3. "Downloaded" section in Hub model picker: shows cached GGUF repos
   at the top (before Recommended) so users can quickly re-load
   previously downloaded models without re-downloading.
2026-03-15 05:24:06 -07:00
Daniel Han
226ece0c9e studio: fix cancel to actually kill llama-server during loading
The unload endpoint checked is_loaded (requires healthy=True), but
during initial loading the server is not yet healthy. Cancel had no
effect because the unload route fell through to the Unsloth backend.

Fix: add is_active property (process exists, loading or loaded) and
check it in the unload route so cancel kills llama-server even during
the download/loading phase.

Also: toast cancel button now properly triggers the backend unload.
2026-03-15 05:24:06 -07:00
Daniel Han
a0fdf03340 studio: add Cancel button to model loading toast popup
Replace toast.promise with a manual toast.loading that includes a
Cancel action button. Users can now cancel model downloads/loads from
the toast notification itself, not just from the header bar spinner.
2026-03-15 05:24:06 -07:00
pre-commit-ci[bot]
1c4efa6c3d [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-03-15 05:24:06 -07:00
Daniel Han
c59f028150 studio: kill orphaned llama-server processes on startup
When the studio process is killed (SIGTERM/SIGKILL), atexit handlers
may not run in the subprocess orchestrator, leaving llama-server
processes orphaned and holding GPU memory. This caused OOM errors when
trying to load a new model after a studio restart.

On init, LlamaCppBackend now runs pgrep to find and SIGKILL any stale
llama-server processes before starting fresh.
2026-03-15 05:24:06 -07:00
Daniel Han
7b19cb418e studio: sort TIGHT (CPU offload) GGUF variants after GPU-only fits
Sort order is now: recommended -> fits (largest first) -> tight/CPU
offload (largest first) -> OOM (smallest first). Previously tight
variants were mixed with fits variants.
2026-03-15 05:24:06 -07:00
Daniel Han
5bb783850a studio: GGUF OOM accounts for CPU offload via --fit (GPU + system RAM)
Updated GGUF fit classification to match llama-server's --fit behavior:

- fits:  model <= 70% of total GPU memory (all GPUs)
- tight: model > 70% GPU but <= 70% GPU + 70% available system RAM
         (llama-server uses --fit to offload layers to CPU)
- OOM:   model exceeds both GPU and system RAM budgets

useGpuInfo now also returns systemRamAvailableGb from /api/system so the
frontend can compute the combined GPU+RAM budget.
2026-03-15 05:24:06 -07:00
pre-commit-ci[bot]
1625565da2 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-03-15 05:24:06 -07:00
Daniel Han
c9c485a7b0 studio: use nvidia-smi for all GPUs + 70% VRAM threshold for GGUF OOM
Two fixes for accurate GGUF OOM detection:

1. /api/system now uses nvidia-smi to enumerate all physical GPUs
   instead of torch.cuda which only sees CUDA_VISIBLE_DEVICES. This
   matches llama-server which can use all GPUs regardless of the env
   var. Falls back to torch-based detection if nvidia-smi unavailable.

2. Frontend GGUF OOM check now uses 70% of total GPU memory as the
   budget, matching the PR's _select_gpus logic (30% reserved for KV
   cache and compute buffers). Previously used checkVramFit's 100%
   threshold which was too generous.
2026-03-15 05:24:06 -07:00
Daniel Han
f5f631e5d1 studio: add cancel button for model loading/downloading
Adds a Cancel button next to the "Downloading model..." spinner so
users can abort long downloads. Clicking it aborts the in-flight load,
calls unloadModel to kill any running llama-server process, and clears
the loading state.
2026-03-15 05:24:06 -07:00
Daniel Han
4600131fea studio: sort OOM GGUF variants smallest-to-largest
OOM variants are more useful sorted ascending by size since smaller ones
are more likely to run with --fit. Non-OOM variants remain largest-first
(best quality).
2026-03-15 05:24:06 -07:00
Daniel Han
ea45370ab8 studio: use total multi-GPU VRAM for OOM checks, recommend smallest when all OOM
Two fixes for GGUF variant dropdown:

1. useGpuInfo now sums memory across all GPU devices instead of only
   reading devices[0]. This matches llama-server's multi-GPU allocation
   where models can be split across GPUs.

2. When the backend-recommended variant (e.g. UD-Q4_K_XL) exceeds total
   GPU VRAM, the frontend picks the largest variant that fits instead.
   If all variants are OOM, it recommends the smallest one (most likely
   to work with --fit).
2026-03-15 05:24:06 -07:00
Daniel Han
10c4db04d8 studio: fix React hooks order -- move useMemo before early returns
The useMemo for sortedVariants was placed after the loading/error early
returns, which violated React's rules of hooks (hooks must be called in
the same order every render). Move it before the conditional returns.

Fixes: Minified React error #310
2026-03-15 05:24:06 -07:00
Daniel Han
3c1b8d7ab7 studio: sort GGUF dropdown client-side -- recommended first, OOM last, rest by size descending
Move the sort logic from the backend to the frontend GgufVariantExpander
component where GPU VRAM info is available. The backend now does a simple
size-descending sort. The frontend pins the recommended variant at the
top, pushes OOM variants to the bottom, and sorts the rest by file size
descending (largest/best quality first).
2026-03-15 05:24:06 -07:00
Daniel Han
dd2d979b40 studio: sort GGUF quants largest-first so best quality that fits is at the top 2026-03-15 05:24:06 -07:00
pre-commit-ci[bot]
1f861e185b [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-03-15 05:24:06 -07:00
Daniel Han
2f5347cb4d studio: sort GGUF quant variants -- recommended first, then UD by size, then standard by size
The variants list was returned in HuggingFace file listing order (alphabetical),
making the dropdown confusing (e.g. BF16 before Q4_0). Now sorted as:

1. Recommended variant (from _pick_best_gguf) pinned at top
2. Other UD (Unsloth Dynamic) variants sorted by disk size ascending
3. Non-UD variants sorted by disk size ascending
2026-03-15 05:24:06 -07:00
Daniel Han
928868f07d studio: auto-find free port if requested port is in use
If the requested port (default 8000) is already in use, auto-
increment and try the next port, up to 20 attempts. Prints a
message like "Port 8000 is in use, using port 8001 instead".

Previously, if port 8000 was busy, uvicorn would fail with
"[Errno 98] address already in use" and the studio would not
start. Now it gracefully finds the next free port.

Uses socket.bind() to check availability before starting uvicorn.
Cross-platform (Linux, macOS, Windows).
2026-03-15 05:24:06 -07:00
Daniel Han
ab6fdccfb5 studio: reorder GGUF preference -- UD-Q4_K_XL first, all UD above standard
Reorder _GGUF_QUANT_PREFERENCE so all UD (Unsloth Dynamic) variants
come before standard quants. UD-Q4_K_XL is the default (best
size/quality tradeoff), followed by other UD quants in decreasing
preference order.

For repos without UD variants (e.g., bartowski), falls through to
standard quants starting with Q4_K_M.

Verified with:
  - unsloth/Qwen3.5-35B-A3B-GGUF -> UD-Q4_K_XL
  - bartowski/Qwen_Qwen3.5-35B-A3B-GGUF -> Q4_K_M
  - unsloth/DeepSeek-V3.2-GGUF -> UD-Q4_K_XL (9 shards)
  - unsloth/Llama-3.2-1B-Instruct-GGUF -> UD-Q4_K_XL
2026-03-15 05:24:06 -07:00
pre-commit-ci[bot]
1dba26012c [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-03-15 05:24:06 -07:00
Daniel Han
8ccb461570 studio: group GGUF shards by variant in size-based fallback
The smallest-fitting-variant fallback now groups split GGUF shards
by their variant prefix and sums all shard sizes per variant.

For example, DeepSeek-V3.2 UD-Q4_K_XL has 9 shards totaling
379.8 GB. The previous code treated each shard as a separate
"variant" and would have incorrectly selected a single 50 GB shard
as fitting, ignoring the other 8 shards needed.

Tested with unsloth/DeepSeek-V3.2-GGUF (237 GGUF files, 27
variants from 150 GB to 1.25 TB). Correctly groups and sorts
all variants by total size.
2026-03-15 05:24:06 -07:00
pre-commit-ci[bot]
d5a18e5a00 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-03-15 05:24:06 -07:00
Daniel Han
93ec05ced2 studio: default to UD-Q4_K_XL for GGUFs, fall back to smallest
Two changes for GGUF variant selection:

1. Default variant preference now starts with UD-Q4_K_XL (Unsloth
   Dynamic quantization) which provides better quality per bit than
   standard Q4_K_M. Also added UD-Q2_K_XL, UD-IQ2_M, UD-IQ1_M,
   UD-IQ1_S as small fallback options.

2. If the selected variant doesn't fit on disk, automatically fall
   back to the smallest GGUF variant in the repo that does fit.
   Queries all GGUF file sizes via get_paths_info() and picks the
   smallest one under the free disk space limit. If nothing fits,
   raises a clear error.

This means users with limited disk space won't get a download
error -- they'll get a smaller quantization instead.
2026-03-15 05:24:06 -07:00
pre-commit-ci[bot]
12f3f4361d [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-03-15 05:24:06 -07:00
Daniel Han
38d700ecb0 studio: check disk space before downloading GGUF models
Query file sizes from HuggingFace via get_paths_info() before
downloading, and compare against free disk space on the cache
partition. Raises a clear error if there is not enough space,
instead of failing mid-download.

Uses get_paths_info() instead of repo_info() because xet-stored
repos return size=None from repo_info().siblings, but
get_paths_info() returns the actual file sizes.

If the size check fails for any reason (network error, API change),
it logs a warning and continues with the download anyway.
2026-03-15 05:24:06 -07:00
pre-commit-ci[bot]
f4fbbcaec8 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-03-15 05:24:06 -07:00
Daniel Han
f4f69f16a6 studio: centralize cache directory for all downloads
Set HF_HOME, HF_HUB_CACHE, HF_XET_CACHE, UV_CACHE_DIR, and
VLLM_CACHE_ROOT to a unified location under ~/.unsloth/studio/cache/
on startup. This keeps all model downloads, datasets, and caches
in one place instead of scattered across ~/.cache/huggingface,
~/.cache/uv, etc.

Layout:
  ~/.unsloth/studio/cache/
    huggingface/       (HF_HOME)
      hub/             (HF_HUB_CACHE -- model/dataset downloads)
      xet/             (HF_XET_CACHE -- xet blob store)
    uv/                (UV_CACHE_DIR -- uv package cache)
    vllm/              (VLLM_CACHE_ROOT -- vllm compiled kernels)

Only sets variables that are not already in the environment, so
user overrides (e.g. HF_HOME=/data/models) are respected.

Cross-platform: uses Path.home() which resolves correctly on
Linux (~), macOS (~), and Windows (C:\Users\<user>).
2026-03-15 05:24:06 -07:00
Daniel Han
f1293fe7d8 studio: respect existing CUDA_VISIBLE_DEVICES in GPU selection
If CUDA_VISIBLE_DEVICES is already set in the environment (e.g.,
by the user or a wrapper script), only consider those GPUs when
selecting devices for llama-server. nvidia-smi reports all physical
GPUs regardless of CUDA_VISIBLE_DEVICES, so we filter its output
to match the allowed set.

Without this, the GPU selector could pick a GPU outside the user's
allowed set, overriding their restriction.
2026-03-15 05:24:06 -07:00
pre-commit-ci[bot]
e885d7308e [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-03-15 05:24:06 -07:00
Daniel Han
12183e0656 studio: smart GPU allocation for GGUF inference
Automatically select the best GPU(s) for a GGUF model based on
file size and available VRAM, instead of relying on hardcoded
-ngl -1 or letting llama-server guess.

Logic:
1. Measure total GGUF file size (including split shards)
2. Query free memory per GPU via nvidia-smi
3. If the model fits in 70% of the most-free GPU's memory,
   pin to that single GPU (CUDA_VISIBLE_DEVICES=X, no --fit)
4. If it needs multiple GPUs, pick the N most-free GPUs
   (CUDA_VISIBLE_DEVICES=X,Y, no --fit)
5. If it's too large for all GPUs combined, omit
   CUDA_VISIBLE_DEVICES and use --fit on to let llama-server
   handle partial offloading

The 70% threshold accounts for KV cache and compute buffers
that sit on top of the model weights.

Removed the -ngl parameter (was hardcoded to -1). llama-server's
default of "auto" handles layer offloading correctly, especially
with --fit on for oversized models.

Tested on 8x B200:
  - 1B model (0.75 GB):  picks 1 GPU, no --fit
  - 27B model (17 GB):   picks 1 GPU, no --fit
  - 405B model (230 GB): picks 2 GPUs, no --fit
  - 2TB model:           all GPUs, --fit on
2026-03-15 05:24:06 -07:00
pre-commit-ci[bot]
7202f81985 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-03-15 05:24:06 -07:00