- Python 71.5%
- TypeScript 22.7%
- Shell 1.9%
- PowerShell 1.6%
- Rust 1.5%
- Other 0.7%
* Fix inference stall during prefill by removing retry storm The _stream_with_retry method used a 0.5s read timeout and retried by sending a brand new POST request each time. During prompt prefill (which can take 5-30+ seconds for long contexts or reasoning models), this caused 10-60 duplicate requests that forced llama-server to restart processing from scratch each time, resulting in 10-20s stalls visible as "Generating" with no progress in the UI. Fix: send the request ONCE with a 120s read timeout for the initial response headers. Cancel support during the prefill wait is handled by a background thread that monitors cancel_event (checked every 0.3s) and closes the response to unblock the httpx read immediately. This preserves the ability to stop/cancel/refresh during generation. The existing 0.5s timeout on the httpx.Client is still used by _iter_text_cancellable for per-token cancel checking during streaming (after prefill), which is unaffected by this change. * Fix race in cancel watcher when response is not yet created When cancel_event fires before client.stream() returns (response is still None), the watcher would hit return and exit without closing anything. The main thread stays blocked for up to 120s. Fix: after cancel is requested, keep polling _response_ref every 0.1s until the response object appears (then close it) or _cancel_closed is set (main thread finished on its own). * Minor cleanup: remove redundant None check, add debug logging in cancel watcher Address Gemini review: cancel_event is guaranteed non-None when the watcher thread runs, and logging the close exception aids debugging. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Retry r.close() on failure instead of giving up If r.close() raises, stay in the polling loop and retry rather than returning and leaving the main thread blocked for up to 120s. * fix: keep short read timeout during token streaming The prefill_timeout (read=120s) was passed to client.stream(), which applied to ALL reads -- not just the initial response headers. This meant _iter_text_cancellable's ReadTimeout-based cancel checking was broken during token streaming: the Stop button could take up to 120s to respond instead of 0.5s. Fix: keep the client's short read timeout (0.5s) for the stream call. During prefill, catch ReadTimeout in a loop and re-check cancel_event instead of re-sending the POST (which was the original retry storm). Once the first bytes arrive, yield the response with a PrependStream wrapper so iter_text() sees the buffered first chunk. This preserves both: - Fast cancel during prefill (via cancel watcher + ReadTimeout loop) - Fast cancel during streaming (via _iter_text_cancellable's 0.5s ReadTimeout, which now fires correctly again) * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: swap to short-timeout stream after prefill completes Address two review issues: 1. _PrependStream did not inherit from httpx.SyncByteStream, so Response.iter_raw() would raise RuntimeError. Replaced with a _ShortTimeoutStream that inherits SyncByteStream properly. 2. client.stream() entry itself raises ReadTimeout during slow prefill (before headers arrive). The previous fix tried to catch this at the body-read level but missed the connection-level timeout. New approach: keep the 120s read timeout for client.stream() so the connection survives long prefills. Once headers arrive, replace the response stream with _ShortTimeoutStream -- a wrapper that uses a background reader thread and a Queue with a short get() timeout to re-raise ReadTimeout at the original 0.5s interval. This way _iter_text_cancellable's cancel-checking remains responsive during token streaming while prefill gets the long timeout it needs. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: move _ShortTimeoutStream before LlamaCppBackend class The class was placed inside LlamaCppBackend's body, splitting the class in two and making _codec_mgr and other attributes unreachable. Move it to module level before LlamaCppBackend. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix: remove _ShortTimeoutStream, use watcher for all cancel _ShortTimeoutStream had two critical issues: 1. Raising ReadTimeout from a generator kills it -- Python finalizes generators after an uncaught exception, so the next next() call hits StopIteration and streaming ends mid-response. 2. The unbounded Queue in the background reader loses backpressure, causing memory spikes with slow clients. Simpler approach: use the 120s read timeout for the entire stream and rely on the cancel watcher thread for all cancellation (both prefill and streaming). The watcher closes the response on cancel_event, which unblocks any blocking httpx read within ~0.3s. This eliminates the need for short timeout tricks entirely. Cancel latency: - Prefill: ~0.3s (watcher polls cancel_event every 0.3s) - Streaming: ~0.3s (same watcher mechanism) - Both faster than the old 0.5s ReadTimeout approach * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * docs: clarify cancel limitations in _stream_with_retry The docstrings claimed ~0.3s cancel in all cases, but httpx cannot interrupt a blocked read before the response object exists. Update the docstrings to accurately describe the behavior: - Cancel during prefill (header wait) is deferred until headers arrive - Cancel during streaming works via response.close() from the watcher - _iter_text_cancellable docstring updated to reflect the watcher-based cancel mechanism instead of the old ReadTimeout polling --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> |
||
|---|---|---|
| .github | ||
| images | ||
| scripts | ||
| studio | ||
| tests | ||
| unsloth | ||
| unsloth_cli | ||
| .gitattributes | ||
| .gitignore | ||
| .pre-commit-ci.yaml | ||
| .pre-commit-config.yaml | ||
| build.sh | ||
| cli.py | ||
| CODE_OF_CONDUCT.md | ||
| CONTRIBUTING.md | ||
| COPYING | ||
| LICENSE | ||
| pyproject.toml | ||
| README.md | ||
| unsloth-cli.py | ||
Run and train AI models with a unified local interface.
Features • Quickstart • Notebooks • Documentation • Discord
Unsloth Studio lets you run and train models for text, audio, embedding, vision and more. Available on Windows, Linux and macOS.
⭐ Features
Unsloth provides several key features for both inference and training:
Inference
- Search + download + run models including GGUF, LoRA adapters, safetensors
- Export models: Save or export models to GGUF, 16-bit safetensors and other formats.
- Tool calling: Support for self-healing tool calling and web search
- Code execution: lets LLMs run code, data and verify results so answers are more accurate.
- Auto-tune inference parameters and customize chat templates.
- Upload images, audio, PDFs, code, DOCX and more file types to chat with.
Training
- Train 500+ models up to 2x faster with up to 70% less VRAM, with no accuracy loss.
- Supports full fine-tuning, pretraining, 4-bit, 16-bit and, FP8 training.
- Observability: Monitor training live, track loss and GPU usage and customize graphs.
- Data Recipes: Auto-create datasets from PDF, CSV, DOCX etc. Edit data in a visual-node workflow.
- Reinforcement Learning: The most efficient RL library, using 80% less VRAM for GRPO, FP8 etc.
- Multi-GPU training is supported, with major improvements coming soon.
⚡ Quickstart
Unsloth can be used in two ways: through Unsloth Studio, the web UI, or through Unsloth Core, the code-based version. Each has different requirements.
Unsloth Studio (web UI)
Unsloth Studio works on Windows, Linux, WSL and macOS.
- CPU: Supported for chat inference only
- NVIDIA: Training works on RTX 30/40/50, Blackwell, DGX Spark, Station and more
- macOS: Currently supports chat only; MLX training is coming very soon
- AMD: Chat works. Train with Unsloth Core. Studio support is coming soon.
- Coming soon: Training support for Apple MLX, AMD, and Intel.
- Multi-GPU: Available now, with a major upgrade on the way
Windows, MacOS, Linux or WSL:
pip install --upgrade pip && pip install uv
uv pip install unsloth --torch-backend=auto
unsloth studio setup
unsloth studio -H 0.0.0.0 -p 8888
Use our Docker image unsloth/unsloth container. Read our Docker Guide.
You can also install directly from source:
pip install --upgrade pip && pip install uv
git clone --filter=blob:none https://github.com/unslothai/unsloth.git
cd unsloth
uv pip install -e . --torch-backend=auto
unsloth studio setup
unsloth studio -H 0.0.0.0 -p 8888
Unsloth Core (code-based)
Windows, Linux, WSL
pip install --upgrade pip && pip install uv
uv pip install unsloth --torch-backend=auto
For Windows, pip install unsloth works only if you have Pytorch installed. Read our Windows Guide.
You can use the same Docker image as Unsloth Studio.
AMD, Intel
For RTX 50x, B200, 6000 GPUs: uv pip install unsloth --torch-backend=auto. Read our guides for: Blackwell and DGX Spark.
To install Unsloth on AMD and Intel GPUs, follow our AMD Guide and Intel Guide.
✨ Free Notebooks
Train for free with our notebooks. Read our guide. Add dataset, run, then deploy your trained model.
| Model | Free Notebooks | Performance | Memory use |
|---|---|---|---|
| Qwen3.5 (4B) | ▶️ Start for free | 1.5x faster | 60% less |
| gpt-oss (20B) | ▶️ Start for free | 2x faster | 70% less |
| gpt-oss (20B): GRPO | ▶️ Start for free | 2x faster | 80% less |
| Qwen3: Advanced GRPO | ▶️ Start for free | 2x faster | 50% less |
| Gemma 3 (4B) Vision | ▶️ Start for free | 1.7x faster | 60% less |
| embeddinggemma (300M) | ▶️ Start for free | 2x faster | 20% less |
| Mistral Ministral 3 (3B) | ▶️ Start for free | 1.5x faster | 60% less |
| Llama 3.1 (8B) Alpaca | ▶️ Start for free | 2x faster | 70% less |
| Llama 3.2 Conversational | ▶️ Start for free | 2x faster | 70% less |
| Orpheus-TTS (3B) | ▶️ Start for free | 1.5x faster | 50% less |
- See all our notebooks for: Kaggle, GRPO, TTS, embedding & Vision
- See all our models and all our notebooks
- See detailed documentation for Unsloth here
🦥 Unsloth News
- Introducing Unsloth Studio: our new web UI for running and training LLMs. Blog
- Qwen3.5 - 0.8B, 2B, 4B, 9B, 27B, 35-A3B, 112B-A10B are now supported. Guide + notebooks
- Train MoE LLMs 12x faster with 35% less VRAM - DeepSeek, GLM, Qwen and gpt-oss. Blog
- Embedding models: Unsloth now supports ~1.8-3.3x faster embedding fine-tuning. Blog • Notebooks
- New 7x longer context RL vs. all other setups, via our new batching algorithms. Blog
- New RoPE & MLP Triton Kernels & Padding Free + Packing: 3x faster training & 30% less VRAM. Blog
- 500K Context: Training a 20B model with >500K context is now possible on an 80GB GPU. Blog
- FP8 & Vision RL: You can now do FP8 & VLM GRPO on consumer GPUs. FP8 Blog • Vision RL
- gpt-oss by OpenAI: Read our RL blog, Flex Attention blog and Guide.
🔗 Links and Resources
| Type | Links |
|---|---|
| Join Reddit community | |
| 📚 Documentation & Wiki | Read Our Docs |
| Follow us on X | |
| 💾 Installation | Pip & Docker Install |
| 🔮 Our Models | Unsloth Catalog |
| ✍️ Blog | Read our Blogs |
Citation
You can cite the Unsloth repo as follows:
@software{unsloth,
author = {Daniel Han, Michael Han and Unsloth team},
title = {Unsloth},
url = {https://github.com/unslothai/unsloth},
year = {2023}
}
If you trained a model with 🦥Unsloth, you can use this cool sticker!
License
Unsloth uses a dual-licensing model of Apache 2.0 and AGPL-3.0. The core Unsloth package remains licensed under Apache 2.0, while certain optional components, such as the Unsloth Studio UI are licensed under the open-source license AGPL-3.0.
This structure helps support ongoing Unsloth development while keeping the project open source and enabling the broader ecosystem to continue growing.
Thank You to
- The llama.cpp library that lets users run and save models with Unsloth
- The Hugging Face team and their libraries: transformers and TRL
- The Pytorch and Torch AO team for their contributions
- And of course for every single person who has contributed or has used Unsloth!