Commit graph

46 commits

Author SHA1 Message Date
Roland Tannous
9ca45826d4 feat: parallel URL image probe with time estimate and progress reporting
- Add 200-sample parallel probe using ThreadPoolExecutor + safe_num_proc
  to estimate download speed and failure rate before full conversion
- Abort with clear error if >=30% of probe images fail to download
- Show estimated download time in the training overlay modal
- Parallel batch conversion for URL-based datasets (vs sequential for local)
- Add warning field to /check-format response for URL-based image datasets
- Display URL warning in dataset preview dialog (amber banner)
- Thread progress_callback from trainer through format_and_template_dataset
  to convert_to_vlm_format for real-time status updates
2026-03-04 23:40:38 +00:00
Roland Tannous
2b704221f7 fix: abort training pipeline on dataset conversion failure 2026-03-04 23:29:43 +00:00
Roland Tannous
a80188848d feat: add index range dataset slicing to studio training page
Add Start/End index inputs under Advanced in the dataset card,
allowing users to slice a dataset by row range before training.
Wired end-to-end: frontend store, API payload, backend Pydantic
model, and trainer dataset loading (inclusive on both ends).
2026-03-04 23:24:09 +00:00
Roland Tannous
91783c0fb2 Revert "Add index range dataset slicing to Studio training page" 2026-03-05 03:21:07 +04:00
Roland Tannous
11ebea6a4b feat: add index range dataset slicing to studio training page
Add Start/End index inputs under Advanced in the dataset card,
allowing users to slice a dataset by row range before training.
Wired end-to-end: frontend store, API payload, backend Pydantic
model, and trainer dataset loading (inclusive on both ends).
2026-03-04 21:48:40 +00:00
Roland Tannous
645d7d357a fix: abort training pipeline on dataset conversion failure 2026-03-04 06:42:48 +00:00
Roland Tannous
01082b84e5 Merge branch 'nightly' into feat/gguf-llama-cpp-inference 2026-02-25 16:06:03 +04:00
Roland Tannous
6f0b7bc38a fix: use raw github URL for vision.py patch + add VLM processor diagnostic logging 2026-02-25 10:29:05 +00:00
Roland Tannous
2149bc74ee Merge pull request #232 from unslothai/fix/disable-eval-by-default
# fix/disable eval by default
2026-02-24 13:35:11 +04:00
Roland Tannous
2be2933846 skip eval split and HF split detection when eval_steps is disabled 2026-02-24 09:26:54 +00:00
Leo Borcherding
cdeed53a97 fix: disable eval by default, set eval_steps to 0.0
- Changed default eval_steps from 0.01 to 0.0 across backend and frontend
- Fixed UI to allow eval_steps=0 (removed min=0.001 constraint)
- Added conditional eval logic with helpful console messages
- Updated tooltip to explain how to disable evaluation
- Tested: confirmed eval disabled by default with eval_steps=0.0
2026-02-23 13:07:47 -06:00
Roland Tannous
d74174f7f5 Cap dataset.map num_proc on multi-GPU machines to prevent fork deadlocks 2026-02-23 14:25:31 +00:00
Roland Tannous
3015916d26 fix: error on >30% sample drop after train_on_responses_only instead of silent DataLoader crash 2026-02-23 12:21:06 +00:00
Roland Tannous
dbbcdb4f09 feat: clear unsloth_compiled_cache on startup, shutdown, and between model loads 2026-02-23 07:26:22 +00:00
Roland Tannous
202b7cdfa7 fix: pass full Processor as processing_class for VLM SFTTrainer 2026-02-22 14:11:12 +00:00
Manan17
798bfb8f6f Setting it to total cpu_count // 4 2026-02-20 06:32:01 +00:00
Roland Tannous
adc0c78dbc reduce dataset_num_proc to 1/4 of cpu_count 2026-02-18 20:53:32 +00:00
Roland Tannous
e33920974b fix: unwrap ProcessorMixin to raw tokenizer for text-only SFTTrainer on VLM-architecture models 2026-02-18 19:13:20 +00:00
Roland Tannous
a6e2fa5b3a Merge remote-tracking branch 'origin/nightly' into fix/dataset-mapping-vlm-text-datasets 2026-02-18 18:11:45 +04:00
Manan17
76cd1dc24c fixing the hangup of training after multiple back to back training processes 2026-02-18 08:18:13 +00:00
Manan17
c37bf686a6 Dividing the total cpu_count // 3 2026-02-18 07:59:57 +00:00
Manan17
58116e7e7a fix the linear path on backend 2026-02-18 07:08:32 +00:00
Roland Tannous
d7853efd21 debug statements 2026-02-18 00:37:09 +00:00
Roland Tannous
0d8b67b706 fix: normalize target_modules [all-linear] list to string for Unsloth/PEFT compatibility 2026-02-18 00:32:21 +00:00
Roland Tannous
d2332622d1 fix: defensively rename VLM chat column to match model's forward() signature 2026-02-17 23:49:13 +00:00
Shine1i
dc0cec772d feat: enhance training stop and reset flow with detailed checks 2026-02-17 23:32:22 +01:00
Shine1i
0be3e6f525 feat: integrate gradient norm tracking in training runtime and metrics
- Enhanced chart logic to filter and visualize finite gradient norm values.
2026-02-17 18:26:59 +01:00
Roland Tannous
ff0aec180a Merge branch 'nightly' into feature/eval-split-auto-detection 2026-02-17 01:11:30 +04:00
Roland Tannous
fa0ca59215 feat: auto-detect model+dataset compatibility to select VLM vs LLM training path 2026-02-16 19:18:49 +00:00
Roland Tannous
5df3a0b250 feat: add eval_enabled flag and format-first-then-split for eval dataset 2026-02-16 14:13:55 +00:00
Roland Tannous
0aea3f149d feat: add eval split auto-detection, eval_steps hyperparam, and eval_loss chart integration 2026-02-16 13:51:10 +00:00
Roland Tannous
37452d56cf feat: add eval split auto-detection, eval_steps hyperparam, and eval_loss chart integration 2026-02-16 13:38:54 +00:00
Roland Tannous
6ecc03485d Merge pull request #97 from unslothai/fix/progress-metics
Resolved the progress metrics
2026-02-16 11:55:01 +04:00
sshah229
63b34660ed modified the num_tokens logic 2026-02-16 00:38:40 -07:00
Roland Tannous
d0964652af feat: thread dataset subset/split params from API routes through to load_dataset calls 2026-02-16 03:56:22 +00:00
sshah229
0b1c635b43 resolved the prgress metrics 2026-02-15 05:35:32 -07:00
Manan17
6e4cde3bf8 Adding save-steps to the SFTConfig 2026-02-15 09:37:54 +00:00
Manan17
6ccbc4edce Fixing stuck training processes 2026-02-15 05:38:06 +00:00
Manan17
97c6a09b84 feat: add cancel or save and stop training 2026-02-15 00:00:22 +00:00
Roland Tannous
4f0fad2156 fix: increase SSE progress timeout to 30min and allow step-0 updates 2026-02-14 05:47:22 +00:00
Roland Tannous
67edebfeb3 feat: wire custom_format_mapping through training pipeline to format_and_template_dataset 2026-02-13 21:07:36 +00:00
Roland Tannous
f52bddc23f refactor: remove gradio dependency from training backend 2026-02-13 09:25:49 +00:00
Roland Tannous
75f775d088 fix: change epoch type from int to float to match TrainerState 2026-02-13 06:51:55 +00:00
Roland Tannous
da1cde971c use get_device() for device selection and clear_gpu_cache() for GPU memory cleanup in inference, trainer, and export 2026-02-11 16:56:52 +00:00
Roland Tannous
62ddcfa019 Refactor [dataset_utils.py](cci:7://file:///home/support/new-ui-prototype/studio/backend/utils/datasets/dataset_utils.py:0:0-0:0) into focused modules 2026-02-03 14:38:02 +00:00
Roland Tannous
544d6944d1 root studio folder 2026-02-02 09:13:49 +00:00