Roland Tannous
12b53ca260
fix: download all GGUF shards for split models (e.g. 7B Q8_0)
...
LlamaCppBackend.load_model() and precache_helper_gguf() only downloaded
the first matching GGUF file. For split models (e.g. 7B Q8_0 with 3
shards), llama-server needs all shards present. Now collects and
downloads all matching files.
2026-03-10 15:08:20 +00:00
Roland Tannous
20264e973e
debug: decode first sample after train_on_completions masking
2026-03-10 14:08:14 +00:00
Roland Tannous
9fd08a3f25
debug: fix dataset access - result is a dict, use dataset['dataset']
2026-03-10 13:19:31 +00:00
Roland Tannous
a5d9f611e6
debug: improve sample preview with type info and traceback
2026-03-10 12:56:24 +00:00
Roland Tannous
10ba97c34c
debug: switch to print() for subprocess visibility
2026-03-10 12:49:01 +00:00
Roland Tannous
19cbf575ac
debug: add temporary log statements for dataset preview and VLM instruction
2026-03-10 12:35:55 +00:00
Roland Tannous
afad614bfa
feat: add LLM-assisted dataset detection using ephemeral GGUF helper
...
Uses Qwen2.5-3B-Instruct Q8_0 via LlamaCppBackend to complement
heuristic-based dataset detection when heuristics are uncertain.
- New llm_assist.py: VLM instruction generation, column classification,
and user-friendly warning generation for dataset issues
- Pre-cache helper GGUF on FastAPI startup (background thread)
- Reorder training pipeline: dataset processing runs BEFORE model load
to avoid VRAM contention (detect → dataset → model → train)
- Add pre_detect_and_load_tokenizer() for lightweight detection
- LLM warnings on VLM conversion failures (broken URLs, missing images)
- LLM column classification fallback when heuristics return unknown
- Graceful degradation: all paths unchanged when helper unavailable
2026-03-10 09:20:45 +00:00
Roland Tannous
22eb0eea29
Revert "Merge pull request #347 from unslothai/feature/studio-storage-roots"
...
This reverts commit e9c7b97d23 , reversing
changes made to b75cc9b959 .
2026-03-10 01:52:47 +00:00
Shine1i
b08b606b21
feat(studio): studio storage roots path utilities
2026-03-09 23:48:31 +00:00
Roland Tannous
a0f03d3080
Add AGPL-3.0 SPDX headers to all source files
2026-03-09 20:17:45 +00:00
Shine1i
992e07495f
Merge remote-tracking branch 'origin/nightly' into feature/fixes-client
2026-03-09 19:07:42 +01:00
Roland Tannous
65d3539bac
Merge pull request #342 from unslothai/local-dataset
...
dataset upload
2026-03-09 21:22:23 +04:00
Roland Tannous
a8992279b6
fix: split dataset 80/20 when eval split matches train split
2026-03-09 16:36:44 +00:00
Shine1i
b9f2820cd6
chore(data-recipe): bump data-designer to 0.5.2 and pin duckdb<1.5
2026-03-09 17:27:02 +01:00
Roland Tannous
a7d78d16be
fix: restore eval_enabled early signal for subprocess training
2026-03-09 15:35:49 +00:00
Roland Tannous
d85176ba1a
fix: allow eval-only progress events through worker callback filter
2026-03-09 14:39:49 +00:00
Roland Tannous
07ba02d610
include all candidate files when scanning a directory, not just the first
2026-03-09 13:52:45 +00:00
Roland Tannous
dcedc4df56
merge nightly, resolve conflict in use-chat-model-runtime
2026-03-09 13:19:17 +00:00
Roland Tannous
f416b7aa3d
training: restore YAML fallback for trust_remote_code (no UI toggle)
2026-03-09 13:10:24 +00:00
Roland Tannous
83b1ff05ef
respect trust_remote_code toggle, return helpful error when required
2026-03-09 13:06:55 +00:00
Manan17
b430e23c0e
dataset upload
2026-03-09 05:50:18 +00:00
Shine1i
4aa171b079
feat(recipe-studio, datasets): improve dataset handling and update metadata logic
2026-03-09 02:47:32 +01:00
samit
433220d338
Adding trust_remote_code to the orchestrator and worker
2026-03-08 16:44:41 -07:00
Shine1i
d951e5aef0
merge nightly
2026-03-09 00:32:33 +01:00
samit
6aa50d353f
exposed trust_remote_code through the UI
2026-03-08 16:28:56 -07:00
Roland Tannous
c38d24b01d
fix: replace is_dataset_multimodal with is_dataset_image/is_dataset_audio in training orchestrator
2026-03-08 19:40:00 +00:00
Manan17
111caf636f
Audio_VLM bug fix
2026-03-08 19:14:07 +00:00
Roland Tannous
08a6cf0c87
feat: route audio inference (TTS, ASR, Whisper) through orchestrator/worker subprocess
2026-03-08 18:25:27 +00:00
Roland Tannous
7db2c90cc6
merge nightly into audio branch (mock test)
2026-03-08 10:23:44 +00:00
Manan17
ae828f6142
adding export support
2026-03-08 04:18:20 +00:00
Roland Tannous
ff85a45050
fix: clear stale model state on failed inference subprocess reload
2026-03-07 23:32:53 +00:00
Roland Tannous
cfad8ec36d
fix: reset checkpoint metadata on failed export checkpoint reload
2026-03-07 23:29:34 +00:00
Manan17
2ce36df03c
check fir gated repo
2026-03-07 21:32:50 +00:00
Roland Tannous
9543fa0d1c
fix: scope dataloader_num_workers=0 to Windows + transformers 5.x only
2026-03-07 17:55:59 +00:00
Roland Tannous
dcebfe718a
fix: prevent training hang on Windows by adding triton-windows support
2026-03-07 17:53:36 +00:00
Roland Tannous
c882a3d2f7
fix: propagate PYTHONPATH to child subprocesses, revert tokenizer patching
2026-03-07 11:28:24 +00:00
Roland Tannous
ac608be800
fix: patch TokenizersBackend in export output after save_pretrained
2026-03-07 10:57:51 +00:00
Roland Tannous
42bd976a2f
fix: patch TokenizersBackend by model name - Qwen3.5→Qwen2Tokenizer, GLM→PreTrainedTokenizer
2026-03-07 10:29:59 +00:00
Roland Tannous
44e9b838ae
fix: patch Qwen3.5 broken tokenizer_class TokenizersBackend across all backends
2026-03-07 09:43:25 +00:00
Roland Tannous
f101befca7
fix: bump transformers 5.x pin from 5.1.0 to 5.2.0 for Qwen3.5 support
2026-03-07 09:10:09 +00:00
Roland Tannous
2c4c598832
fix: fail fast if runtime pip install of transformers 5.x fails
2026-03-07 08:40:25 +00:00
Roland Tannous
082a6f876e
fix: join prior pump thread before starting new training job
2026-03-07 08:37:03 +00:00
Roland Tannous
717d5d621d
fix: drain stale events from resp_queue after generation cancel
2026-03-07 08:12:16 +00:00
Roland Tannous
2558125a14
fix: serialize generation with _gen_lock to prevent concurrent queue readers
...
Two overlapping /chat/completions requests could both read from the shared
resp_queue, consuming and dropping each other's token events. Replace the
request_id filtering (which silently dropped non-matching messages) with a
threading.Lock that serializes generation — correct for single-GPU inference.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 04:06:51 +00:00
Roland Tannous
cac409198f
fix: log final GGUF file locations after relocation
2026-03-06 18:04:27 +00:00
Roland Tannous
e72293a033
fix: increase export timeout to 1 hour for large model GGUF conversion
2026-03-06 17:59:42 +00:00
Shine1i
72468a3292
feat(recipe-studio, validators): tweak OXC validator with lint suppression support and improve error normalization logic
2026-03-06 09:40:53 +01:00
Roland Tannous
7e59440029
fix: pin huggingface_hub==1.3.0 in .venv_t5 (satisfies transformers 5.x)
2026-03-06 06:19:28 +00:00
Roland Tannous
661ac4be96
feat: subprocess-based export, pin huggingface_hub==0.36.0
2026-03-06 06:03:09 +00:00
Shine1i
4bd7eca3f3
feat(recipe-studio, validators): extend OXC validator with code shape support and integrate into recipe studio
2026-03-06 02:04:05 +01:00