Commit graph

4,672 commits

Author SHA1 Message Date
Shine1i
71916f1dce feat: add ShineBorder UI component and learning recipe templates to enhance data recipes page 2026-02-23 22:38:06 +01:00
Shine1i
8739a01f56 Merge branch 'nightly' into feature/canvas-lab 2026-02-23 21:54:35 +01:00
Shine1i
b8231a6be6 refactor: keep seed block pos 2026-02-23 21:53:32 +01:00
Shine1i
1323e0af53 refactor: add batch processing support with configuration options and execution enhancements 2026-02-23 21:32:20 +01:00
Shine1i
59a15cb5bc refactor: enhance recipe validation flows with error collection, seed-specific updates, and improved UX in execution dialogs 2026-02-23 20:34:53 +01:00
Shine1i
d4655eb8bf refactor: streamline recipe execution flows with validation support and enhanced run dialog interactions 2026-02-23 20:28:41 +01:00
Shine1i
91cbb0e933 refactor: improve dialog rendering and logging setup for stability and configurability 2026-02-23 20:16:03 +01:00
Leo Borcherding
cdeed53a97 fix: disable eval by default, set eval_steps to 0.0
- Changed default eval_steps from 0.01 to 0.0 across backend and frontend
- Fixed UI to allow eval_steps=0 (removed min=0.001 constraint)
- Added conditional eval logic with helpful console messages
- Updated tooltip to explain how to disable evaluation
- Tested: confirmed eval disabled by default with eval_steps=0.0
2026-02-23 13:07:47 -06:00
imagineer99
6cedc339c6 feat: add Upload / Save / Reset training config from local YAML 2026-02-23 18:55:08 +00:00
Shine1i
71ab9ff4b4 refactor: enhance seed configuration handling with added fields, dynamic chunking logic, and streamlined interactions 2026-02-23 19:40:13 +01:00
Shine1i
424b00b701 refactor: improve seed source handling with additional type support, enhanced parsing logic, and text chunking optimization 2026-02-23 19:29:54 +01:00
Shine1i
3e17e2b0f6 refactor: enhance seed source handling with new source types and streamlined inspection flows 2026-02-23 18:46:02 +01:00
imagineer99
71d698d182 feat: sort and filter dataset search results by model type relevance 2026-02-23 16:22:45 +00:00
Roland Tannous
77b0978d5f Merge pull request #228 from unslothai/fix/cap-num-proc-multigpu-deadlock
Cap dataset.map num_proc on multi-GPU machines to prevent fork deadlocks
2026-02-23 19:04:22 +04:00
Roland Tannous
d74174f7f5 Cap dataset.map num_proc on multi-GPU machines to prevent fork deadlocks 2026-02-23 14:25:31 +00:00
Roland Tannous
6acc2dbf8f Merge branch 'nightly' into feature/transformers-v5-support 2026-02-23 13:40:16 +00:00
Roland Tannous
4cb0cfdaf5 Remove firebase-debug.log and setup_leo.sh from tracking and add to .gitignore 2026-02-23 17:38:22 +04:00
Roland Tannous
c017ce802f Remove firebase-debug.log and setup_leo.sh from tracking and add to .gitignore 2026-02-23 17:37:46 +04:00
Roland Tannous
313e77c5fd Merge branch 'nightly' into feature/transformers-v5-support 2026-02-23 13:32:51 +00:00
Roland Tannous
2e2aa54ad2 Merge pull request #225 from unslothai/fix/fix-response-on-completion-truncation
fix: error on >30% sample drop after `train_on_responses_only` instead of silent DataLoader crash
2026-02-23 16:28:32 +04:00
Roland Tannous
3015916d26 fix: error on >30% sample drop after train_on_responses_only instead of silent DataLoader crash 2026-02-23 12:21:06 +00:00
Roland Tannous
b03938f6ad Merge nightly into main Brings main up to date with nightly, including chat attachments, model-per-thread persistence, speech recognition, VRAM recommendations, MoE model configs, VLM fixes, and compile cache cleanup. Conflicts resolved by taking nightly's version for all diverged files (main-only changes were a feature add + immediate revert with net zero effect). 2026-02-23 14:52:57 +04:00
Daniel Han
2ed86865fb Suppress FBGEMM CUTLASS stdout spam on Blackwell GPUs (#4092)
* Suppress FBGEMM CUTLASS "Arch conditional MMA" stdout spam on Blackwell GPUs

On Blackwell GPUs (B200/B100, SM100), FBGEMM's f8f8bf16_blockwise kernel
is hardcoded to cutlass::arch::Sm90 with no SM100 code path. When
test_has_fbgemm() probes this kernel, it fires 2304 "ERROR : Arch
conditional MMA instruction used without targeting appropriate compute
capability" lines before aborting and returning zeros.

The existing HidePrintMessage filter on sys.stderr (line 109) does not
catch these because CUDA device-side printf writes to stdout fd 1 at the
C level, bypassing Python's sys.stdout/sys.stderr entirely.

Fix: add suppress_cuda_printf() context manager in import_fixes.py that
redirects fd 1 and fd 2 to /dev/null at the OS level, with
torch.cuda.synchronize() and libc fflush before restoring. Wrap the
test_has_fbgemm() call in fp8.py with this context manager.

Tested on B200 with fbgemm-gpu-genai 1.4.0+cu130 and 1.5.0+cu130:
- Before: 2304 warning lines on every import
- After: 0 warning lines
- UNSLOTH_HAS_FBGEMM correctly set to 0 (Triton fallback works)
- Works with both UNSLOTH_ENABLE_LOGGING=0 and =1

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Guard _libc init and fflush to prevent fd leak on failure

---------

Co-authored-by: Ubuntu <ubuntu@ip-172-31-16-253.us-east-2.compute.internal>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-02-23 01:27:10 -08:00
Daniel Han
fec06247c9 Fix VLM processor load degradation and vLLM CUDA version detection (#4091)
* Fix VLM processor load degradation and vLLM CUDA version detection

vision.py - Fix VLM processor load for issue #4085:
- Before loading the processor, scan local config files and strip the
  _Unsloth_Patched_ prefix. AutoProcessor.from_pretrained silently
  degrades to a text-only tokenizer instead of raising an exception
  when it encounters the unrecognized class name, so the existing
  get_auto_processor fallback never triggers. Sanitizing the configs
  before loading fixes backwards compat for old corrupted saves.
- After loading, detect when AutoProcessor returned a text-only
  tokenizer for a VLM model (has no image_processor attribute) and
  trigger the manual fallback constructor.

import_fixes.py - Fix vLLM CUDA version mismatch detection:
- _is_broken_vllm_error now also matches CUDA shared library errors
  (libcudart, libcublas, libnvrtc) with "cannot open shared object
  file". Previously it only matched errors containing "vllm._c" in
  the message text, which missed cases where the error message was
  about the missing CUDA library itself (e.g. vllm built for CUDA 12
  on a CUDA 13 system).
- New _get_vllm_cuda_mismatch_message function extracts the CUDA
  version from the error, compares to the system CUDA version via
  torch.version.cuda, and returns a targeted install command using
  the correct GitHub releases wheel URL.
- disable_broken_vllm uses the targeted message when a CUDA mismatch
  is detected, falling back to the existing generic message otherwise.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: Ubuntu <ubuntu@ip-172-31-16-253.us-east-2.compute.internal>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-02-23 01:06:53 -08:00
Roland Tannous
5416bdd4e6 added shutil import to main.py 2026-02-23 07:44:59 +00:00
Roland Tannous
f16a7f2d17 Merge nightly into feature/transformers-v5-support 2026-02-23 07:40:28 +00:00
Roland Tannous
bfb84221fe Merge pull request #221 from unslothai/feature/clear-unsloth-compile-cache
feat: clear unsloth_compiled_cache on startup, shutdown, and between …
2026-02-23 11:30:46 +04:00
Roland Tannous
dbbcdb4f09 feat: clear unsloth_compiled_cache on startup, shutdown, and between model loads 2026-02-23 07:26:22 +00:00
Roland Tannous
2a117f57a9 Merge pull request #220 from unslothai/feature/moe-training-models-configs
Add model defaults for MoE models (Qwen3 MoE, GLM Flash) and GLM response mapping
2026-02-23 10:09:27 +04:00
Roland Tannous
fb1c321ad3 Add GLM, Qwen3 MoE, TinyQwen3 MoE, and Ministral 3 VL model defaults and GLM train_on_responses_only mapping 2026-02-23 05:51:43 +00:00
Roland Tannous
3fe85d36cc Remove stale .venv_overlay on server startup to prevent transformers version conflicts 2026-02-23 05:08:27 +00:00
Roland Tannous
6995aaf077 Clean up stale .venv_overlay directory during setup 2026-02-22 20:30:00 +00:00
Roland Tannous
e3fb4f53df Patch adapter_config.json with unsloth_training_method and auto-detect load_in_4bit for LoRA inference 2026-02-22 20:27:52 +00:00
Roland Tannous
5de6246142 Purge own utils/core modules and use lazy imports so is_vision_model picks up fresh AutoConfig after version switch 2026-02-22 20:04:35 +00:00
samit
7dcaa52083 added cancel training button on the overlay 2026-02-22 12:03:26 -08:00
Roland Tannous
c12d75c472 Add transformers version switch to model config and vision check endpoints for dropdown selection 2026-02-22 19:56:12 +00:00
Roland Tannous
60997a75eb Install transformers into both site-packages and overlay to fix sub-package resolution during version switch 2026-02-22 19:44:12 +00:00
Roland Tannous
7cde520176 Move transformers overlay to local .venv_overlay/, add huggingface-hub to overlay install 2026-02-22 19:34:46 +00:00
Roland Tannous
15cb9b0f37 Use sys.path overlay to switch transformers versions in-process instead of modifying site-packages 2026-02-22 19:19:12 +00:00
Roland Tannous
0050e78aa3 Fix in-memory transformers version detection and aggressive module purge for 5.1.0/4.57.1 switching 2026-02-22 19:08:18 +00:00
Roland Tannous
1c2653fcc2 aggressive reload_transformers 2026-02-22 18:52:06 +00:00
Roland Tannous
4d06258e93 Auto-switch transformers version (5.1.0/4.57.1) for Ministral-3, GLM-4.7-Flash, Qwen3-30B-A3B models with LoRA adapter resolution 2026-02-22 18:29:40 +00:00
Roland Tannous
bb1bd49a68 Merge pull request #217 from unslothai/fix/update-config-yamls
Fix vision LoRA defaults for VLMs and clean up text-only model configs
2026-02-22 19:14:54 +04:00
Roland Tannous
132cdb0547 fix: correct vision LoRA defaults for VLMs and remove vision fields from text-only model configs 2026-02-22 15:09:10 +00:00
Roland Tannous
7ac391d1e0 Merge pull request #215 from unslothai/fix/vlm-processing-class
fix: pass full Processor as processing_class for VLM SFTTrainer
2026-02-22 18:14:14 +04:00
Roland Tannous
202b7cdfa7 fix: pass full Processor as processing_class for VLM SFTTrainer 2026-02-22 14:11:12 +00:00
Roland Tannous
a4346b954e Merge pull request #213 from unslothai/fix/fix-clear-chat-new-model
Fix: Clear Chat and Manage Model Lifecycle on Model Switch
2026-02-22 17:54:27 +04:00
Roland Tannous
761953b50e feat(chat): persist model per thread and auto-load on thread switch 2026-02-22 13:35:45 +00:00
Roland Tannous
536a735acc feat(chat): eject current model and start fresh thread on model switch 2026-02-22 13:17:03 +00:00
Roland Tannous
a490dca8a0 Merge pull request #211 from unslothai/fix/param-count-display
Fix: Remove download count fallback when model param count is unavailable
2026-02-22 16:52:25 +04:00