* Removing .precommit config
* edited colab comments
* studio: update Unsloth_Studio_Colab.ipynb
* studio: update Unsloth_Studio_Colab.ipynb
* studio: add Colab T4 GPU metadata to force T4 instance
* style: update colab popup to black/white theme with gem icon and play button
* feat: center landscape image in colab notebook
* style: shrink popup to fit content, truncate URL display
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* feat: center landscape image in colab notebook
* feat: use GitHub raw URL for studio landscape image in notebook
* chore: update colab notebook
---------
Co-authored-by: LeoBorcherding <LeoBorcherding@users.noreply.github.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* studio: extract param count from model name as fallback
When HuggingFace API doesn't return totalParams for a model,
extract the param count from the model name (e.g. "Qwen3-0.6B"
-> "0.6B", "Llama-3.2-1B-Instruct" -> "1B"). Applied to both
the recommended list and HF search results.
* studio: read GGUF context_length via fast header parser, set max tokens
- Fast GGUF metadata reader (~30-55ms) parses only KV header, skips
tensor data and large arrays (tokenizer vocab etc)
- Extracts context_length and chat_template from GGUF metadata
- Returns context_length in LoadResponse for frontend to use
- Frontend sets maxTokens to actual context_length for GGUFs (e.g.
262144 for Qwen3.5-9B, 131072 for Qwen2.5-7B)
- Max Tokens slider shows "Max" and is locked for GGUFs
- Auto-load path also uses actual context_length from load response
- Toast auto-dismiss (5s) and close button for auto-load toast
* studio: GGUF TTS audio support (from PR #4318)
Add GGUF TTS audio generation via llama-server. When a GGUF model
loads, the backend probes its vocabulary to detect audio codecs
(SNAC/BiCodec/DAC/CSM/Whisper). If detected, the codec is pre-loaded
and the model is reported as audio to the frontend.
During chat, TTS models route to the audio generation path which sends
a per-codec prompt to llama-server's /completion endpoint, extracts
generated tokens/text, and decodes to WAV using AudioCodecManager.
Also strips base64 audio data from prior assistant messages to prevent
context overflow.
Co-authored-by: Manan Shah <mananshah511@gmail.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Remove package-lock.json from tracking
* studio: per-model inference defaults, GGUF max tokens fix, reasoning toggle
- Add inference_defaults.json with per-model-family sampling parameters
for ~50 families (Qwen3.5, Qwen3, Gemma-3, Llama-3, DeepSeek, etc.).
Values sourced from unslothai/docs and Ollama params blobs.
- Family-based lookup in inference_config.py: extracts model family from
identifier, matches against patterns (longest match first), merges with
priority: model-specific YAML > family JSON > default.yaml.
- Fix GGUF Max Tokens slider locked at "Max": store ggufContextLength
separately from maxTokens so the slider is adjustable (step=64).
- Fix Ministral YAML: top_p was literal string "default", now 0.95.
- Add reasoning toggle for thinking models (Qwen3.5, Qwen3, DeepSeek-R1,
DeepSeek-V3.1, etc.): detect enable_thinking support from GGUF chat
template metadata, pass --jinja to llama-server, send
chat_template_kwargs per-request. Frontend shows "Reasoning is ON/OFF"
pill button next to attachment button in composer.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* studio: remove default system prompt injection
Backend was injecting "You are a helpful AI assistant." when no system
prompt was provided. Neither unslothai/docs nor Ollama specify a default
system prompt for most models. Now defaults to empty string, letting the
model's own chat template handle system behavior.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* studio: use lightbulb icons and "Think" label for reasoning toggle
Lightbulb on when thinking enabled, lightbulb-off when disabled.
Label is just "Think" in both states; grayed out styling when off.
* studio: fix HTML file upload breaking chat
Replace SimpleTextAttachmentAdapter with custom TextAttachmentAdapter
(excludes text/html) and HtmlAttachmentAdapter that strips tags via
DOMParser, removing scripts/styles and extracting readable text content
instead of dumping raw HTML markup into the conversation.
* studio: show chat template in Configuration panel
Display the model's Jinja2 chat template in a new "Chat Template"
section under Settings (now open by default). For GGUFs, reads from
GGUF metadata; for safetensors, reads from tokenizer.chat_template.
Template is editable with a "Restore default chat template" button
that appears when modified. Section only shows when a model with a
chat template is loaded.
* studio: editable chat template with Apply & Reload
Chat template section now functional:
- Editing the template shows "Apply & Reload" (reloads model with
custom template) and "Revert changes" buttons
- For GGUFs: writes template to temp .jinja file, passes
--chat-template-file to llama-server on reload
- For non-GGUF: passes chat_template_override in load request
- Settings section now open by default
- selectModel supports forceReload to reload same model
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* studio: fix DeepSeek reasoning detection and auto-load metadata
- Set _model_identifier before _read_gguf_metadata so DeepSeek
"thinking" template detection works (was always None before)
- Populate ggufContextLength, supportsReasoning, reasoningEnabled,
defaultChatTemplate in autoLoadSmallestModel GGUF path
* studio: add spacing before BETA badge in navbar
Add gap-1.5 on the logo Link container to space the BETA label
from the wordmark.
Co-authored-by: Imagineer99 <Imagineer99@users.noreply.github.com>
* studio: vertically center BETA badge with logo
---------
Co-authored-by: Manan Shah <mananshah511@gmail.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Imagineer99 <Imagineer99@users.noreply.github.com>
* Strip <think> blocks from LLM assist model output
* Add debug logging for raw LLM assist output
* Quiet llama-server logs, use structlog in llm_assist
* Fix think-tag stripping when response is inside tags
* Remove debug logging of raw model output
* Clarify GGUF download logs: show cache hit vs actual download
* Clarify heuristic-detected mapping in UI text
* Default helper model to Qwen3-4B-Instruct-2507 UD-Q4_K_XL
* Remove package-lock.json from tracking, add to .gitignore
* Auto-open mapping dialog on Start Training for custom_heuristic format
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Use last think block when extracting inner content (review feedback)
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
- GGUF: use -c 0 for model's native context size (no 4096 cap)
- GGUF: hide Max Seq Length slider (irrelevant), set Max Tokens to Max
- Non-GGUF: default Max Tokens to 4096
- Max Tokens slider shows "Max" label when at ceiling for GGUFs
- Run non-GGUF load_model in asyncio.to_thread for progress polling
- Auto-load smallest downloaded model when chatting without selection
- Wait for in-progress model load before inference (modelLoading store flag)
- Recommended list: 4 GGUFs + 4 hub models after case-insensitive dedup
- Model selector waits for cached data before rendering
- Toast close button repositioned, Sampling section open by default
- Add logging to _get_repo_size_cached exception handler
- Use -c 0 for llama-server (model's native context size, no 4096 cap)
- Run non-GGUF backend.load_model in asyncio.to_thread for progress polling
- Auto-load smallest downloaded model when user chats without selecting one
- Wait for in-progress model load before inference (no "No model loaded" error)
- Add modelLoading flag to zustand store for cross-component coordination
- Dynamic top models: send 8 GGUFs + 8 hub models, frontend caps 4+4 after dedup
- Case-insensitive dedup: downloaded models correctly hide from recommended list
- Prevent duplicate toasts: guard against double selectModel calls
- Model selector waits for cached data before rendering (no empty flash)
- Toast close button positioned at top-right with proper spacing
- Sampling section expanded by default in chat settings
- Global toast close button styling fix
Change all repetition_penalty defaults from 1.1 (or 1.05/1.2 in
presets) to 1.0 across the entire backend and frontend. Most models
handle repetition well on their own and a non-1.0 penalty can degrade
output quality, especially for code, structured output, and creative
tasks.
Files changed:
- Backend: inference.py, llama_cpp.py, orchestrator.py, worker.py,
models/inference.py (Field defaults)
- Frontend: chat-settings-sheet.tsx (Creative/Precise presets),
runtime-provider.tsx (auto-title generation)
The _VISION_CHECK_SCRIPT subprocess used logger.info() but logger was
never defined in the subprocess context. This caused a NameError on
every vision check, making all transformers 5.x models (Qwen3.5,
GLM, etc.) fall back to text-only mode even when they support vision.
Replace logger.info() with print() since the parent process reads
the subprocess stdout via result.stdout.
llama-server uses stb_image internally which does not support WebP,
TIFF, AVIF, and other formats that browsers accept for upload.
Uploading a WebP image to a vision GGUF model caused a 400 error:
"Failed to load image or audio file" / "failed to decode image bytes".
Convert all uploaded images to PNG via PIL before base64-encoding and
forwarding to llama-server. This handles WebP, TIFF, BMP, GIF, AVIF,
and any other format PIL supports. RGBA images are converted to RGB
first since PNG with alpha can cause issues in some vision pipelines.
GGUF repos with mmproj files (e.g. Qwen3.5-0.8B-GGUF) are already
detected as vision-capable by list_gguf_variants(), and is_vision is
set correctly in ModelConfig. However, the HF download path only
downloaded the main GGUF file without the mmproj projection file,
so llama-server started without --mmproj and rejected image uploads
with "text-only model" errors.
Add _download_mmproj() to LlamaCppBackend that:
- Lists repo files for mmproj*.gguf matches
- Prefers mmproj-F16.gguf (best quality), falls back to any mmproj
- Downloads via hf_hub_download (uses the same HF cache)
In load_model(), when is_vision=True and no explicit mmproj_path was
provided (HF mode), auto-download the mmproj after the main GGUF.
The downloaded path is passed to llama-server via --mmproj.
1. Backend: When a model fails with "No config file found" or similar
unsupported-model errors, wrap the message with "This model is not
supported yet. Try a different model." instead of showing the raw
Unsloth exception.
2. Frontend: Compute estimated download size from the HF search API's
safetensors.parameters dtype breakdown (BF16=2B/param, I32=4B/param,
F32=4B/param, etc.) and show it in the model picker instead of just
the param count. For example, Kimi-K2.5 now shows "~554 GB" instead
of "171B" (which was misleading since 171B params != 171GB download).
Three fixes on top of the download progress feature:
1. Backend: Replace broken "no .incomplete = done" completion check
with a 95% byte threshold. HF downloads files sequentially, so
between files there are briefly no .incomplete files even though
the download is far from done (e.g. Kimi-K2.5 reported "done"
after downloading 22KB of config files out of 595GB).
2. Frontend: Track hasShownProgress flag. Only show "Download
complete. Loading into memory..." if we actually displayed
download progress before. For already-cached models where the
first poll returns progress=1.0, this avoids the misleading
"Download complete" message.
3. Frontend: Deduplicate recommended vs downloaded -- filter out
models already in the "Downloaded" section. Cache the fetched
lists at module level so re-mounting the popover does not flash
an empty "Downloaded" section.
Previously only GGUF models showed download progress in Chat. Non-GGUF
models (safetensors, bnb quantized, etc.) showed a static message with
no progress indication. This adds progress tracking for all model types
and fixes several related issues.
Backend:
- Add /api/models/download-progress endpoint that checks the HF cache
blobs directory for completed and .incomplete files. Uses model_info()
(cached per repo) to determine expected total size for percentage.
- Add /api/models/cached-models endpoint that lists non-GGUF model repos
from the HF cache via scan_cache_dir().
- Fix progress stuck at 0.99: when no .incomplete files remain, report
1.0 immediately (blob deduplication can make byte totals mismatch).
Frontend:
- Remove the ggufVariant gate so download progress polling works for all
non-cached models, not just GGUFs.
- Use GGUF-specific endpoint when variant + expectedBytes available,
otherwise use the general download-progress endpoint.
- Fix toast stuck after load: check loadingModelRef.current before and
after the async poll to prevent overwriting the success toast.
- First poll at 500ms instead of waiting for the 2s interval.
- Show downloaded non-GGUF models in the Hub model picker "Downloaded"
section alongside GGUFs.
* fix: Ctrl+C not breaking out of backend on Linux
threading.Event.wait() without a timeout blocks at the C level on
Linux, preventing Python from delivering SIGINT. Use a 1-second
timeout loop so the interpreter can process pending signals.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* user can upload eval dataset, removed bugs
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* resolving merge conflicts
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* resolving gpt comments
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Roland Tannous <115670425+rolandtannous@users.noreply.github.com>
Two issues caused the studio frontend to render without any styling
when installed via `pip install` (non-editable):
1. `pyproject.toml` package-data only included `frontend/dist/**/*`.
The `include-package-data = true` setting relies on `git ls-files`,
which fails in isolated builds (pip/uv copy source to a temp dir
without `.git`). This meant `frontend/src/`, `package.json`,
`vite.config.ts`, and other build files were missing from the
installed package. Tailwind had no source files to scan.
2. Python venvs auto-create a `.gitignore` with a bare `*` pattern.
Tailwind v4's oxide scanner walks parent directories and respects
`.gitignore` -- so even when source files are present, the venv's
`*` pattern causes the scanner to skip all `.tsx` files. The result
is a 34KB CSS skeleton with zero utility classes instead of the
expected 265KB.
Additionally, Vite adds `crossorigin` to script/link tags by default.
This forces CORS mode on font subresource loads, which Firefox
HTTPS-Only Mode does not exempt -- causing all @font-face downloads
to fail silently when Studio is served over HTTP.
Changes:
- pyproject.toml: Expand package-data to include frontend source,
config files, setup scripts, and backend requirements using glob
patterns (no node_modules)
- studio/setup.sh: Temporarily hide parent .gitignore files containing
a bare `*` during `npm run build`, with trap-based restoration
- studio/backend/main.py: Strip `crossorigin` attributes from HTML
at serve time so fonts load correctly on any protocol
* fix: graceful shutdown on Windows (signal handlers for Ctrl+C)
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* chat only with gguf for mac devices
* resolving gpt comments
* add change-password for chat only
* hide lora adaptors dropdown
* solving gpt comments
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* addressing the comment
* fixing auth flow
---------
Co-authored-by: Datta Nimmaturi <venkatadattasainimmaturi@gmail.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
- Run GGUF load_model in asyncio.to_thread so the event loop stays free
for progress polling during download (was blocking all requests).
- Extract download phase out of the lock in LlamaCppBackend.load_model
so unload_model/cancel can take effect immediately during download.
- Fix "downloaded" badge for split GGUFs: check total cached bytes
across all shards vs expected size, not just first shard existence.
- Respect CUDA_VISIBLE_DEVICES in /api/system GPU reporting so the
frontend GGUF fit estimation uses actual available VRAM.
- Sort tight variants (need CPU offload) smallest-first instead of
largest-first -- closer to GPU budget = faster inference.
- Fix cancel: use refs instead of React state for abort controller and
toast ID so both cancel buttons (text + toast) work reliably. Make
cancel synchronous (fire-and-forget unload) for instant UI response.
Check abortCtrl.signal.aborted after loadModel returns to prevent
ghost model state. Skip rollback and suppress errors on cancel.
- Dynamic top 4 GGUF models fetched from HF API sorted by downloads,
prepended to the default recommended list.
- Remove turnAnchor="top" for auto-scroll to bottom during generation.
- Set default toast duration to 10s (was infinite for loading toasts).
- Deduplicate cached GGUF repos using scan_cache_dir API (fixes
Qwen/X-GGUF vs qwen/x-gguf duplicates from lowercased HF cache).
- Pre-compile repo_id validation regex to silence CodeQL ReDoS warning.
- Change welcome text and default suggestion text.
_get_gpu_free_memory was filtering by CUDA_VISIBLE_DEVICES, so with
CUDA_VISIBLE_DEVICES='0' set by the training env, llama-server only
saw 1 GPU and used --fit for CPU offloading instead of spreading
across all 8 GPUs.
Since llama-server manages its own GPU allocation (the _select_gpus
method picks GPUs and sets CUDA_VISIBLE_DEVICES for the subprocess),
the query must see ALL physical GPUs to make the right decision.
1. Progress endpoint now takes a variant parameter and only counts
.gguf files matching that variant (not all files in the repo cache,
which would include previously downloaded variants)
2. Tracks .incomplete files in HF blobs dir for in-progress single-shard
downloads, capping at 99% until the file is fully committed
3. Fixed loading text: "Loading model..." for cached, "Downloading
model..." for new downloads, with appropriate descriptions
4. Wording: "Downloading and loading model. Large models can take a
while." instead of "This may include downloading."
1. Loading text: shows "Loading model..." for cached models,
"Downloading model..." for new downloads. Toast description
adapts accordingly.
2. Download progress: polls /api/models/gguf-download-progress every
2s during downloads, updating the toast with percentage and GB
downloaded. Progress is estimated by checking the HF cache folder
size against the expected total bytes.
3. Passes isDownloaded and expectedBytes through the full chain from
variant click to selectModel for accurate UI state.
1. n_gpu_layers kwarg: accept (and ignore) in load_model signature
so callers like llm_assist.py don't get TypeError
2. mmproj exclusion: filter out mmproj files in _find_smallest_fitting_variant
so fallback doesn't pick a tiny vision projection as the "model"
3. Shard preservation after fallback: re-discover shards for the
fallback variant instead of resetting to empty list, so split
GGUFs download all shards
4. Orphan cleanup safety: only kill llama-server processes whose
cmdline contains ".unsloth/", avoiding termination of unrelated
llama-server instances on the same machine
5. Path expression sanitization: validate repo_id format before using
it in cache directory lookups
The variant filename includes a subfolder prefix (e.g.
UD-Q4_K_XL/Kimi-K2.5-UD-Q4_K_XL-00001-of-00013.gguf) but rglob
returns just the filename. Use Path.name for the comparison.
HF cache dirs use the exact case from the repo_id at download time
(e.g. models--unsloth--kimi-k2.5-gguf) which may differ from the
canonical HF repo_id (unsloth/Kimi-K2.5-GGUF). Use case-insensitive
matching to find the cache directory.
- Backend: /gguf-variants now checks HF cache for each variant's file
and returns a downloaded flag per variant
- Frontend: downloaded variants sort before non-downloaded (after
recommended), and show a green "downloaded" badge
- Sort order: recommended -> downloaded+fits -> downloaded+tight ->
fits -> tight -> OOM
1. Interruptible downloads: load_model now checks a cancel event
between shard downloads. unload_model sets the event so cancel
stops the download at the next shard boundary.
2. /api/models/cached-gguf endpoint: scans the HF cache for
already-downloaded GGUF repos with their total size and cache path.
3. "Downloaded" section in Hub model picker: shows cached GGUF repos
at the top (before Recommended) so users can quickly re-load
previously downloaded models without re-downloading.
The unload endpoint checked is_loaded (requires healthy=True), but
during initial loading the server is not yet healthy. Cancel had no
effect because the unload route fell through to the Unsloth backend.
Fix: add is_active property (process exists, loading or loaded) and
check it in the unload route so cancel kills llama-server even during
the download/loading phase.
Also: toast cancel button now properly triggers the backend unload.
When the studio process is killed (SIGTERM/SIGKILL), atexit handlers
may not run in the subprocess orchestrator, leaving llama-server
processes orphaned and holding GPU memory. This caused OOM errors when
trying to load a new model after a studio restart.
On init, LlamaCppBackend now runs pgrep to find and SIGKILL any stale
llama-server processes before starting fresh.
Two fixes for accurate GGUF OOM detection:
1. /api/system now uses nvidia-smi to enumerate all physical GPUs
instead of torch.cuda which only sees CUDA_VISIBLE_DEVICES. This
matches llama-server which can use all GPUs regardless of the env
var. Falls back to torch-based detection if nvidia-smi unavailable.
2. Frontend GGUF OOM check now uses 70% of total GPU memory as the
budget, matching the PR's _select_gpus logic (30% reserved for KV
cache and compute buffers). Previously used checkVramFit's 100%
threshold which was too generous.
Move the sort logic from the backend to the frontend GgufVariantExpander
component where GPU VRAM info is available. The backend now does a simple
size-descending sort. The frontend pins the recommended variant at the
top, pushes OOM variants to the bottom, and sorts the rest by file size
descending (largest/best quality first).
The variants list was returned in HuggingFace file listing order (alphabetical),
making the dropdown confusing (e.g. BF16 before Q4_0). Now sorted as:
1. Recommended variant (from _pick_best_gguf) pinned at top
2. Other UD (Unsloth Dynamic) variants sorted by disk size ascending
3. Non-UD variants sorted by disk size ascending
If the requested port (default 8000) is already in use, auto-
increment and try the next port, up to 20 attempts. Prints a
message like "Port 8000 is in use, using port 8001 instead".
Previously, if port 8000 was busy, uvicorn would fail with
"[Errno 98] address already in use" and the studio would not
start. Now it gracefully finds the next free port.
Uses socket.bind() to check availability before starting uvicorn.
Cross-platform (Linux, macOS, Windows).
Reorder _GGUF_QUANT_PREFERENCE so all UD (Unsloth Dynamic) variants
come before standard quants. UD-Q4_K_XL is the default (best
size/quality tradeoff), followed by other UD quants in decreasing
preference order.
For repos without UD variants (e.g., bartowski), falls through to
standard quants starting with Q4_K_M.
Verified with:
- unsloth/Qwen3.5-35B-A3B-GGUF -> UD-Q4_K_XL
- bartowski/Qwen_Qwen3.5-35B-A3B-GGUF -> Q4_K_M
- unsloth/DeepSeek-V3.2-GGUF -> UD-Q4_K_XL (9 shards)
- unsloth/Llama-3.2-1B-Instruct-GGUF -> UD-Q4_K_XL