- Python 71.5%
- TypeScript 22.7%
- Shell 1.9%
- PowerShell 1.6%
- Rust 1.5%
- Other 0.7%
* fix: throttle and cache HuggingFace modelInfo API calls The frontend was firing 40 to 60 parallel modelInfo requests on app startup with zero caching or deduplication, causing HF rate limits. Adds a caching layer (hf-cache.ts) with TTL cache, inflight request dedup, and a concurrency limiter. Also debounces the HF token input so typing a token no longer re-fires all model searches per keystroke. * fix: only fetch VRAM info for visible models in chat selector * Fix cache key isolation and VRAM badge stability for PR #4696 - Cache key now includes a token fingerprint (last 8 chars) instead of a boolean, so switching HF tokens gives separate cache entries instead of serving stale data from the previous token. - Extract token via credentials?.accessToken to match the @huggingface/hub API surface. - Extend CachedResult type with safetensors/tags fields so downstream consumers no longer need unsafe `as` casts. - Merge VRAM param map with previous state on scroll instead of replacing it, preventing a brief flash of missing VRAM badges when new models become visible. * Fix VRAM badges missing for search-filtered recommended models When a user types a search query, filteredRecommendedIds can include models beyond the currently visible page. These models had no VRAM data because useRecommendedModelVram only received visibleRecommendedIds. Now we pass the union of visibleRecommendedIds and filteredRecommendedIds to the VRAM hook, so recommended models surfaced by search also show their VRAM badges. The hf-cache layer ensures no duplicate network calls. * Apply biome formatting to hf-cache.ts and use-recommended-model-vram.ts Auto-formatted with biome check --write to match project lint rules: - Block statements for single-line if/for bodies - Import sorting (type imports first) - Consistent line wrapping * Fix extractToken to handle both current and deprecated HF auth forms The @huggingface/hub CredentialsParams type is a union: - { accessToken: "hf_..." } (current preferred form) - { credentials: { accessToken: "..." } } (deprecated form) Previously only checked params.credentials?.accessToken (deprecated path). Now checks both forms so the cache key is correct regardless of which calling convention is used. * Simplify extractToken, map merge, and set construction - extractToken: remove type assertions, use direct property access with truthiness checks for cleaner union type handling - VRAM map merge: use Map spread constructor instead of manual for loop - idsForVram: use Set spread construction for more concise dedup * Add rationale comment for MAX_CONCURRENT=3 in hf-cache.ts * Skip GGUF repos in VRAM fetch and pre-populate cache from listModels Two changes to reduce redundant HF API calls: 1. Filter GGUF repos from idsForVram before passing to useRecommendedModelVram. GGUF repos have no safetensors metadata and the render layer already shows a static "GGUF" badge -- fetching modelInfo for them is a no-op that wastes a semaphore slot and a network round-trip. 2. Add primeCacheFromListing() to hf-cache.ts and call it from listModels yield sites in mergedModelIterator and priorityThenListingIterator. listModels returns the same type (ModelEntry & Pick<ApiModelInfo, T>) as modelInfo with the same additionalFields, so the data is interchangeable. Priming only writes if the key is not already fresh, so it never overwrites a recent modelInfo response. This means models discovered via listModels are already in cache when useRecommendedModelVram later calls cachedModelInfo for them, eliminating duplicate network requests. * Fix cache key mismatch: prime both token and anonymous slots The VRAM hook calls cachedModelInfo without credentials (anonymous key), but listModels results were primed only under the authenticated key. For authenticated users the priming was a no-op -- cache miss every time. Fix: prime both the token-specific slot and the anonymous slot when an access token is present. Public model metadata (safetensors, tags) is identical regardless of auth so this is safe. Also add a defensive guard in primeCacheFromListing for empty name. * Auto-prime anonymous cache slot from authenticated modelInfo fetches When cachedModelInfo is called with a token, the result was only stored under the token-specific key (e.g. model::abc12345). The VRAM hook calls cachedModelInfo without credentials and reads the anonymous slot (model::anon), causing a cache miss and duplicate fetch for every priority model. Now cachedModelInfo also writes to the anonymous slot on success when a token is present. Public model metadata (safetensors, tags) is identical regardless of auth, so this is safe and eliminates ~10 duplicate API calls on first page load. * Guard anonymous cache priming against gated/private models Only prime the anonymous cache slot for non-gated, non-private models. Previously, authenticated modelInfo responses and listing results were unconditionally copied into the anonymous slot, which could briefly expose gated/private model metadata after clearing the HF token. Now checks result.gated and result.private before writing the anon slot. Public unsloth/ models (the common case) still benefit from the optimization; gated models like meta-llama/* require a fresh fetch per auth context. * Extract primeFromListing helper to deduplicate cache priming logic The cache priming pattern (prime token slot + conditionally prime anon slot for non-gated models) was duplicated in three places. Extracted into a single primeFromListing() function for maintainability. * Export CachedResult type, add isStale helper, simplify primeFromListing - Export CachedResult so consumers can use it directly instead of the indirect Parameters<typeof ...> pattern. - Extract isStale(key) helper to deduplicate the cache freshness check that was repeated in primeCacheFromListing, cachedModelInfo, and the anonymous-slot priming logic. - Simplify primeFromListing to use CachedResult directly for both the data parameter and the gated/private guard, eliminating the double cast. --------- Co-authored-by: Daniel Han <danielhanchen@gmail.com> |
||
|---|---|---|
| .github | ||
| images | ||
| scripts | ||
| studio | ||
| tests | ||
| unsloth | ||
| unsloth_cli | ||
| .gitattributes | ||
| .gitignore | ||
| .pre-commit-ci.yaml | ||
| .pre-commit-config.yaml | ||
| build.sh | ||
| cli.py | ||
| CODE_OF_CONDUCT.md | ||
| CONTRIBUTING.md | ||
| COPYING | ||
| install.ps1 | ||
| install.sh | ||
| LICENSE | ||
| pyproject.toml | ||
| README.md | ||
| unsloth-cli.py | ||
Run and train AI models with a unified local interface.
Features • Quickstart • Notebooks • Documentation • Reddit
Unsloth Studio (Beta) lets you run and train text, audio, embedding, vision models on Windows, Linux and macOS.
⭐ Features
Unsloth provides several key features for both inference and training:
Inference
- Search + download + run models including GGUF, LoRA adapters, safetensors
- Export models: Save or export models to GGUF, 16-bit safetensors and other formats.
- Tool calling: Support for self-healing tool calling and web search
- Code execution: lets LLMs test code in Claude artifacts and sandbox environments
- Auto-tune inference parameters and customize chat templates.
- We work directly with teams behind gpt-oss, Qwen3, Llama 4, Mistral, Gemma 1-3, and Phi-4, where we’ve fixed bugs that improve model accuracy.
- Upload images, audio, PDFs, code, DOCX and more file types to chat with.
Training
- Train and RL 500+ models up to 2x faster with up to 70% less VRAM, with no accuracy loss.
- Custom Triton and mathematical kernels. See some collabs we did with PyTorch and Hugging Face.
- Data Recipes: Auto-create datasets from PDF, CSV, DOCX etc. Edit data in a visual-node workflow.
- Reinforcement Learning (RL): The most efficient RL library, using 80% less VRAM for GRPO, FP8 etc.
- Supports full fine-tuning, RL, pretraining, 4-bit, 16-bit and, FP8 training.
- Observability: Monitor training live, track loss and GPU usage and customize graphs.
- Multi-GPU training is supported, with major improvements coming soon.
⚡ Quickstart
Unsloth can be used in two ways: through Unsloth Studio, the web UI, or through Unsloth Core, the code-based version. Each has different requirements.
Unsloth Studio (web UI)
Unsloth Studio (Beta) works on Windows, Linux, WSL and macOS.
- CPU: Supported for Chat and Data Recipes currently
- NVIDIA: Training works on RTX 30/40/50, Blackwell, DGX Spark, Station and more
- macOS: Currently supports chat and Data Recipes. MLX training is coming very soon
- AMD: Chat + Data works. Train with Unsloth Core. Studio support is out soon.
- Coming soon: Training support for Apple MLX, AMD, and Intel.
- Multi-GPU: Available now, with a major upgrade on the way
macOS, Linux, WSL:
curl -fsSL https://unsloth.ai/install.sh | sh
Windows:
irm https://unsloth.ai/install.ps1 | iex
Launch
unsloth studio -H 0.0.0.0 -p 8888
Update
To update, use the same install commands as above. Or run (does not work on Windows):
unsloth studio update
Docker
Use our Docker image unsloth/unsloth container. Run:
docker run -d -e JUPYTER_PASSWORD="mypassword" \
-p 8888:8888 -p 8000:8000 -p 2222:22 \
-v $(pwd)/work:/workspace/work \
--gpus all \
unsloth/unsloth
Developer, Nightly, Uninstall
To see developer, nightly and uninstallation etc. instructions, see advanced installation.
Unsloth Core (code-based)
Linux, WSL:
curl -LsSf https://astral.sh/uv/install.sh | sh
uv venv unsloth_env --python 3.13
source unsloth_env/bin/activate
uv pip install unsloth --torch-backend=auto
Windows:
winget install -e --id Python.Python.3.13
winget install --id=astral-sh.uv -e
uv venv unsloth_env --python 3.13
.\unsloth_env\Scripts\activate
uv pip install unsloth --torch-backend=auto
For Windows, pip install unsloth works only if you have PyTorch installed. Read our Windows Guide.
You can use the same Docker image as Unsloth Studio.
AMD, Intel:
For RTX 50x, B200, 6000 GPUs: uv pip install unsloth --torch-backend=auto. Read our guides for: Blackwell and DGX Spark.
To install Unsloth on AMD and Intel GPUs, follow our AMD Guide and Intel Guide.
✨ Free Notebooks
Train for free with our notebooks. Read our guide. Add dataset, run, then deploy your trained model.
| Model | Free Notebooks | Performance | Memory use |
|---|---|---|---|
| Qwen3.5 (4B) | ▶️ Start for free | 1.5x faster | 60% less |
| gpt-oss (20B) | ▶️ Start for free | 2x faster | 70% less |
| Qwen3.5 GSPO | ▶️ Start for free | 2x faster | 70% less |
| gpt-oss (20B): GRPO | ▶️ Start for free | 2x faster | 80% less |
| Qwen3: Advanced GRPO | ▶️ Start for free | 2x faster | 70% less |
| Gemma 3 (4B) Vision | ▶️ Start for free | 1.7x faster | 60% less |
| embeddinggemma (300M) | ▶️ Start for free | 2x faster | 20% less |
| Mistral Ministral 3 (3B) | ▶️ Start for free | 1.5x faster | 60% less |
| Llama 3.1 (8B) Alpaca | ▶️ Start for free | 2x faster | 70% less |
| Llama 3.2 Conversational | ▶️ Start for free | 2x faster | 70% less |
| Orpheus-TTS (3B) | ▶️ Start for free | 1.5x faster | 50% less |
- See all our notebooks for: Kaggle, GRPO, TTS, embedding & Vision
- See all our models and all our notebooks
- See detailed documentation for Unsloth here
🦥 Unsloth News
- Introducing Unsloth Studio: our new web UI for running and training LLMs. Blog
- Qwen3.5 - 0.8B, 2B, 4B, 9B, 27B, 35-A3B, 112B-A10B are now supported. Guide + notebooks
- Train MoE LLMs 12x faster with 35% less VRAM - DeepSeek, GLM, Qwen and gpt-oss. Blog
- Embedding models: Unsloth now supports ~1.8-3.3x faster embedding fine-tuning. Blog • Notebooks
- New 7x longer context RL vs. all other setups, via our new batching algorithms. Blog
- New RoPE & MLP Triton Kernels & Padding Free + Packing: 3x faster training & 30% less VRAM. Blog
- 500K Context: Training a 20B model with >500K context is now possible on an 80GB GPU. Blog
- FP8 & Vision RL: You can now do FP8 & VLM GRPO on consumer GPUs. FP8 Blog • Vision RL
- gpt-oss by OpenAI: Read our RL blog, Flex Attention blog and Guide.
📥 Advanced Installation
The below advanced instructions are for Unsloth Studio. For Unsloth Core advanced installation, view our docs.
Developer installs: macOS, Linux, WSL:
git clone https://github.com/unslothai/unsloth
cd unsloth
./install.sh --local
unsloth studio -H 0.0.0.0 -p 8888
Then to update :
unsloth studio update
Developer installs: Windows PowerShell:
git clone https://github.com/unslothai/unsloth.git
cd unsloth
Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass
.\install.ps1 --local
unsloth studio -H 0.0.0.0 -p 8888
Then to update :
unsloth studio update
Nightly: MacOS, Linux, WSL:
git clone https://github.com/unslothai/unsloth
cd unsloth
git checkout nightly
./install.sh --local
unsloth studio -H 0.0.0.0 -p 8888
Then to launch every time:
unsloth studio -H 0.0.0.0 -p 8888
Nightly: Windows:
Run in Windows Powershell:
git clone https://github.com/unslothai/unsloth.git
cd unsloth
git checkout nightly
Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass
.\install.ps1 --local
unsloth studio -H 0.0.0.0 -p 8888
Then to launch every time:
unsloth studio -H 0.0.0.0 -p 8888
Uninstall
You can uninstall Unsloth Studio by deleting its install folder usually located under $HOME/.unsloth/studio on Mac/Linux/WSL and %USERPROFILE%\.unsloth\studio on Windows. Using the rm -rf commands will delete everything, including your history, cache:
- MacOS, WSL, Linux:
rm -rf ~/.unsloth/studio - Windows (PowerShell):
Remove-Item -Recurse -Force "$HOME\.unsloth\studio"
For more info, see our docs.
Deleting model files
You can delete old model files either from the bin icon in model search or by removing the relevant cached model folder from the default Hugging Face cache directory. By default, HF uses:
- MacOS, Linux, WSL:
~/.cache/huggingface/hub/ - Windows:
%USERPROFILE%\.cache\huggingface\hub\
💚 Community and Links
| Type | Links |
|---|---|
| Join Discord server | |
| Join Reddit community | |
| 📚 Documentation & Wiki | Read Our Docs |
| Follow us on X | |
| 🔮 Our Models | Unsloth Catalog |
| ✍️ Blog | Read our Blogs |
Citation
You can cite the Unsloth repo as follows:
@software{unsloth,
author = {Daniel Han, Michael Han and Unsloth team},
title = {Unsloth},
url = {https://github.com/unslothai/unsloth},
year = {2023}
}
If you trained a model with 🦥Unsloth, you can use this cool sticker!
License
Unsloth uses a dual-licensing model of Apache 2.0 and AGPL-3.0. The core Unsloth package remains licensed under Apache 2.0, while certain optional components, such as the Unsloth Studio UI are licensed under the open-source license AGPL-3.0.
This structure helps support ongoing Unsloth development while keeping the project open source and enabling the broader ecosystem to continue growing.
Thank You to
- The llama.cpp library that lets users run and save models with Unsloth
- The Hugging Face team and their libraries: transformers and TRL
- The Pytorch and Torch AO team for their contributions
- NVIDIA for their NeMo DataDesigner library and their contributions
- And of course for every single person who has contributed or has used Unsloth!