Commit graph

42 commits

Author SHA1 Message Date
Shine1i
b08b606b21 feat(studio): studio storage roots path utilities 2026-03-09 23:48:31 +00:00
Roland Tannous
a0f03d3080 Add AGPL-3.0 SPDX headers to all source files 2026-03-09 20:17:45 +00:00
Roland Tannous
dcedc4df56 merge nightly, resolve conflict in use-chat-model-runtime 2026-03-09 13:19:17 +00:00
Roland Tannous
f416b7aa3d training: restore YAML fallback for trust_remote_code (no UI toggle) 2026-03-09 13:10:24 +00:00
Roland Tannous
83b1ff05ef respect trust_remote_code toggle, return helpful error when required 2026-03-09 13:06:55 +00:00
Roland Tannous
5ebd4de2ef backend: resolve trust_remote_code from YAML when not set by frontend 2026-03-09 11:58:23 +00:00
Shine1i
4aa171b079 feat(recipe-studio, datasets): improve dataset handling and update metadata logic 2026-03-09 02:47:32 +01:00
samit
6aa50d353f exposed trust_remote_code through the UI 2026-03-08 16:28:56 -07:00
Roland Tannous
7db2c90cc6 merge nightly into audio branch (mock test) 2026-03-08 10:23:44 +00:00
Roland Tannous
ad6739be7a fix: wait for training shutdown before export load, clear stop flag on reset
1. Export route: stop_training() only signals the subprocess — wait up to
   30s for it to actually exit before loading the export checkpoint, avoiding
   a GPU memory race.

2. Training reset: clear _should_stop so /api/train/status returns phase=idle
   instead of staying stuck on phase=stopped after a user-triggered stop.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 04:16:10 +00:00
Roland Tannous
c82e3d86bb fix: unload competing subprocesses before load across all routes 2026-03-06 06:05:31 +00:00
Roland Tannous
ce9bfd7476 fix: unload inference model before training to free GPU memory
When starting training, shut down the inference subprocess first
so the training subprocess has full GPU memory available.
2026-03-05 22:28:11 +00:00
Roland Tannous
021c3aafdd fix: handle None job_id before first training run 2026-03-05 16:59:37 +00:00
Roland Tannous
f90af41c5f feat: subprocess-based training for transformers version switching 2026-03-05 15:40:32 +00:00
Manan17
8203637d89 resolved merge conflicts 2026-03-05 07:59:43 +00:00
Roland Tannous
f57664e268 Merge nightly into feature/transformers-v5-support 2026-03-05 06:49:44 +00:00
Roland Tannous
64889cd5fc feat: add index range dataset slicing to studio training page
Add Start/End index inputs under Advanced in the dataset card,
allowing users to slice a dataset by row range before training.
Wired end-to-end: frontend store, API payload, backend Pydantic
model, and trainer dataset loading (inclusive on both ends).
2026-03-04 23:24:09 +00:00
Roland Tannous
9333f99dd3 Revert "Add index range dataset slicing to Studio training page" 2026-03-05 03:21:07 +04:00
Roland Tannous
02b17ec6d9 feat: add index range dataset slicing to studio training page
Add Start/End index inputs under Advanced in the dataset card,
allowing users to slice a dataset by row range before training.
Wired end-to-end: frontend store, API payload, backend Pydantic
model, and trainer dataset loading (inclusive on both ends).
2026-03-04 21:48:40 +00:00
Manan17
6fd1dd2c0a variable changes and some cleanup 2026-03-03 09:35:11 +00:00
Manan17
2c5621dd8c merging with nightly 2026-03-01 02:27:45 +00:00
Roland Tannous
f5b30448e8 Auto-switch transformers version (5.1.0/4.57.1) for Ministral-3, GLM-4.7-Flash, Qwen3-30B-A3B models with LoRA adapter resolution 2026-02-22 18:29:40 +00:00
Shine1i
b31461790f feat: enhance training stop and reset flow with detailed checks 2026-02-17 23:32:22 +01:00
Shine1i
f47c424be3 feat: integrate gradient norm tracking in training runtime and metrics
- Enhanced chart logic to filter and visualize finite gradient norm values.
2026-02-17 18:26:59 +01:00
Roland Tannous
108ec254cb Merge branch 'nightly' into feature/eval-split-auto-detection 2026-02-17 01:11:30 +04:00
Roland Tannous
18879a521b feat: auto-detect model+dataset compatibility to select VLM vs LLM training path 2026-02-16 19:18:49 +00:00
Roland Tannous
3b117189c5 feat: add eval_enabled flag and format-first-then-split for eval dataset 2026-02-16 14:13:55 +00:00
Roland Tannous
90c3561adb feat: add eval split auto-detection, eval_steps hyperparam, and eval_loss chart integration 2026-02-16 13:38:54 +00:00
Roland Tannous
d49506b7b1 feat: add live GPU monitor with nvidia-smi polling during training 2026-02-16 11:47:43 +00:00
Roland Tannous
be584ccfa7 Merge pull request #97 from unslothai/fix/progress-metics
Resolved the progress metrics
2026-02-16 11:55:01 +04:00
Roland Tannous
38cb5c9496 feat: thread dataset subset/split params from API routes through to load_dataset calls 2026-02-16 03:56:22 +00:00
sshah229
7fd55ce14f resolved the prgress metrics 2026-02-15 05:35:32 -07:00
Manan17
354b7d0aca feat: add cancel or save and stop training 2026-02-15 00:00:22 +00:00
Roland Tannous
8304060b9f fix: increase SSE progress timeout to 30min and allow step-0 updates 2026-02-14 05:47:22 +00:00
Roland Tannous
286d5ff0a7 feat: wire custom_format_mapping through training pipeline to format_and_template_dataset 2026-02-13 21:07:36 +00:00
Shine1i
0ebfb5be76 feat: add support for serialized previews in dataset API and improve training initialization logging 2026-02-13 13:47:17 +01:00
Roland Tannous
6beddf9f9e fix: change epoch type from int to float to match TrainerState 2026-02-13 06:51:55 +00:00
Roland Tannous
d17c1b99d8 feat: add SSE reconnection resilience with spec-compliant event fields, Last-Event-ID resume, and metric_history fallback in /status 2026-02-12 17:58:48 +00:00
Roland Tannous
5ae20f6099 move inline pydantic models - fix existing models routes integration 2026-02-11 12:39:58 +00:00
sshah229
6e5cd50c34 fixed the errors- renamed jwt to authentication, used raw jwt, and removed search route 2026-02-07 03:14:30 -07:00
sshah229
fc673227b7 Refactored the training and model routes and added the jwt authentication 2026-02-06 03:15:30 -07:00
Roland Tannous
75d8dcc824 root studio folder 2026-02-02 09:13:49 +00:00
Renamed from backend/routes/training.py (Browse further)