Commit graph

51 commits

Author SHA1 Message Date
Roland Tannous
adc0c78dbc reduce dataset_num_proc to 1/4 of cpu_count 2026-02-18 20:53:32 +00:00
Roland Tannous
e33920974b fix: unwrap ProcessorMixin to raw tokenizer for text-only SFTTrainer on VLM-architecture models 2026-02-18 19:13:20 +00:00
Roland Tannous
a6e2fa5b3a Merge remote-tracking branch 'origin/nightly' into fix/dataset-mapping-vlm-text-datasets 2026-02-18 18:11:45 +04:00
Manan17
76cd1dc24c fixing the hangup of training after multiple back to back training processes 2026-02-18 08:18:13 +00:00
Manan17
c37bf686a6 Dividing the total cpu_count // 3 2026-02-18 07:59:57 +00:00
Manan17
58116e7e7a fix the linear path on backend 2026-02-18 07:08:32 +00:00
Roland Tannous
d7853efd21 debug statements 2026-02-18 00:37:09 +00:00
Roland Tannous
0d8b67b706 fix: normalize target_modules [all-linear] list to string for Unsloth/PEFT compatibility 2026-02-18 00:32:21 +00:00
Roland Tannous
d2332622d1 fix: defensively rename VLM chat column to match model's forward() signature 2026-02-17 23:49:13 +00:00
Shine1i
dc0cec772d feat: enhance training stop and reset flow with detailed checks 2026-02-17 23:32:22 +01:00
Roland Tannous
c9fdce63e7 fix: skip sudo check on WSL during GGUF export to prevent password prompt hang 2026-02-17 19:30:02 +00:00
Shine1i
0be3e6f525 feat: integrate gradient norm tracking in training runtime and metrics
- Enhanced chart logic to filter and visualize finite gradient norm values.
2026-02-17 18:26:59 +01:00
Manan17
c7b7ecab4f Adding metadata for checkpoints 2026-02-16 23:46:17 +00:00
Wasim Yousef Said
7b8220598e Merge pull request #124 from unslothai/feature/bug-fixes
feat: support disabling top-k sampling with -1 and standardize normalization
2026-02-16 13:21:12 -08:00
Roland Tannous
ff0aec180a Merge branch 'nightly' into feature/eval-split-auto-detection 2026-02-17 01:11:30 +04:00
Shine1i
0db7da96cc feat: support disabling top-k sampling with -1 and standardize normalization logic
- Updated top-k parameter range to accept -1 in models and frontend.
- Added utility to normalize top-k for backend compatibility.
2026-02-16 21:33:24 +01:00
Roland Tannous
fa0ca59215 feat: auto-detect model+dataset compatibility to select VLM vs LLM training path 2026-02-16 19:18:49 +00:00
Roland Tannous
5df3a0b250 feat: add eval_enabled flag and format-first-then-split for eval dataset 2026-02-16 14:13:55 +00:00
Roland Tannous
0aea3f149d feat: add eval split auto-detection, eval_steps hyperparam, and eval_loss chart integration 2026-02-16 13:51:10 +00:00
Roland Tannous
37452d56cf feat: add eval split auto-detection, eval_steps hyperparam, and eval_loss chart integration 2026-02-16 13:38:54 +00:00
Roland Tannous
f0298edeb8 refactor: move checkpoint scanning to utils/models and /checkpoints endpoint to models router 2026-02-16 09:32:11 +00:00
Roland Tannous
6ecc03485d Merge pull request #97 from unslothai/fix/progress-metics
Resolved the progress metrics
2026-02-16 11:55:01 +04:00
sshah229
63b34660ed modified the num_tokens logic 2026-02-16 00:38:40 -07:00
Roland Tannous
909955767b feat: add min_p sampling parameter to /chat/completions generation pipeline 2026-02-16 06:33:17 +00:00
Manan17
19276ae60b Fixing the get checkpoint api 2026-02-16 04:47:28 +00:00
Roland Tannous
d0964652af feat: thread dataset subset/split params from API routes through to load_dataset calls 2026-02-16 03:56:22 +00:00
Shine1i
571959e383 feat: add cancelation support for chat generation and streaming tasks 2026-02-15 18:23:27 +01:00
sshah229
0b1c635b43 resolved the prgress metrics 2026-02-15 05:35:32 -07:00
Manan17
6e4cde3bf8 Adding save-steps to the SFTConfig 2026-02-15 09:37:54 +00:00
Manan17
6ccbc4edce Fixing stuck training processes 2026-02-15 05:38:06 +00:00
Manan17
97c6a09b84 feat: add cancel or save and stop training 2026-02-15 00:00:22 +00:00
Roland Tannous
be3934860f strip extra debug statements 2026-02-14 19:23:51 +00:00
Roland Tannous
3ff3def555 replace model unloading and peft loading mechanism for compare feature 2026-02-14 19:18:49 +00:00
Roland Tannous
e7ae901737 del model.peft_config instead of using model.delete_adapter 2026-02-14 17:32:15 +00:00
Roland Tannous
7d8e991c1f added print statements for activate_lora_adapter 2026-02-14 17:25:37 +00:00
Roland Tannous
6fefbe9f0b swipped logger for print statements as logger isn't propagating 2026-02-14 17:21:26 +00:00
Roland Tannous
d0b94eae75 added logging 2026-02-14 17:09:07 +00:00
Roland Tannous
b5c8136957 exclude default from model.delete_adapter 2026-02-14 17:03:52 +00:00
Roland Tannous
35a6e40268 _apply_adapter_state now calls revert_to_base_model and activate_lora_adapter properly 2026-02-14 16:57:24 +00:00
Roland Tannous
f67ee58347 feat(inference): add use_adapter field for per-request adapter toggling in compare mode 2026-02-14 14:52:13 +00:00
Roland Tannous
418a374125 migrate _generate_vision_response to use TextIteratorStreamer + background thread 2026-02-14 09:30:32 +00:00
Roland Tannous
4f0fad2156 fix: increase SSE progress timeout to 30min and allow step-0 updates 2026-02-14 05:47:22 +00:00
Roland Tannous
67edebfeb3 feat: wire custom_format_mapping through training pipeline to format_and_template_dataset 2026-02-13 21:07:36 +00:00
Roland Tannous
f52bddc23f refactor: remove gradio dependency from training backend 2026-02-13 09:25:49 +00:00
Roland Tannous
75f775d088 fix: change epoch type from int to float to match TrainerState 2026-02-13 06:51:55 +00:00
Roland Tannous
da1cde971c use get_device() for device selection and clear_gpu_cache() for GPU memory cleanup in inference, trainer, and export 2026-02-11 16:56:52 +00:00
Roland Tannous
59d5f24eb5 integrate global hardware detection at lifespan entrypoint 2026-02-11 15:34:26 +00:00
Roland Tannous
107bd2be4c feat: add Apple Silicon (MPS) compatibility to backend utils + tests 2026-02-11 14:00:39 +00:00
Roland Tannous
b4ec0389f0 refactor/inference-api-routes-part-1 2026-02-03 16:57:57 +00:00
Roland Tannous
62ddcfa019 Refactor [dataset_utils.py](cci:7://file:///home/support/new-ui-prototype/studio/backend/utils/datasets/dataset_utils.py:0:0-0:0) into focused modules 2026-02-03 14:38:02 +00:00