Roland Tannous
1d06e2f54c
switch dataset upload from base64 JSON to multipart/form-data with streamed writes
2026-03-09 13:55:45 +00:00
Roland Tannous
56412f2362
include all candidate files when scanning a directory, not just the first
2026-03-09 13:52:45 +00:00
Roland Tannous
4c5ded4c52
normalize uploaded filename extension to lowercase for consistent downstream checks
2026-03-09 13:35:55 +00:00
Manan17
a08b73e385
remove file size limit
2026-03-09 07:04:02 +00:00
Manan17
a49638c504
dataset upload
2026-03-09 05:50:18 +00:00
Wasim Yousef Said
91e81227bd
Merge pull request #273 from unslothai/feature/data-reciper-enchansments
...
UX + layout polish & WIP data-reciper client & backend finalization p2
2026-03-09 02:57:32 +01:00
Shine1i
3b1663b1e9
feat(recipe-studio, datasets): improve dataset handling and update metadata logic
2026-03-09 02:47:32 +01:00
Roland Tannous
254f10e37a
Merge pull request #328 from unslothai/fix/chat-unloading-model
...
fixed model unload before load without validation
2026-03-09 04:40:05 +04:00
Shine1i
a2dde15367
merge nightly
2026-03-09 00:32:33 +01:00
Roland Tannous
7b76fccb9b
fix: loosen executorch pin for python 3.13 compat
2026-03-08 19:56:09 +00:00
Roland Tannous
a1778d6655
fix: replace is_dataset_multimodal with is_dataset_image/is_dataset_audio in training orchestrator
2026-03-08 19:40:00 +00:00
Manan17
80b704d7b7
Audio_VLM bug fix
2026-03-08 19:14:07 +00:00
Roland Tannous
7ee81dd7df
feat: route audio inference (TTS, ASR, Whisper) through orchestrator/worker subprocess
2026-03-08 18:25:27 +00:00
Roland Tannous
1435dbaf59
merge nightly into audio branch (mock test)
2026-03-08 10:23:44 +00:00
Manan17
ef714f010e
adding export support
2026-03-08 04:18:20 +00:00
Roland Tannous
a7c34b42be
fix: clear stale model state on failed inference subprocess reload
2026-03-07 23:32:53 +00:00
Roland Tannous
4f766bbe25
fix: reset checkpoint metadata on failed export checkpoint reload
2026-03-07 23:29:34 +00:00
Manan17
6487f81113
check fir gated repo
2026-03-07 21:32:50 +00:00
Manan17
905f989521
fixing sesame model
2026-03-07 19:06:00 +00:00
Roland Tannous
8454e6dd2b
fix: scope dataloader_num_workers=0 to Windows + transformers 5.x only
2026-03-07 17:55:59 +00:00
Roland Tannous
ef9184c731
fix: prevent training hang on Windows by adding triton-windows support
2026-03-07 17:53:36 +00:00
Roland Tannous
e25705a211
fix: propagate PYTHONPATH to child subprocesses, revert tokenizer patching
2026-03-07 11:28:24 +00:00
Roland Tannous
9330588015
fix: patch TokenizersBackend in export output after save_pretrained
2026-03-07 10:57:51 +00:00
Roland Tannous
76c78afb8f
fix: patch TokenizersBackend by model name - Qwen3.5→Qwen2Tokenizer, GLM→PreTrainedTokenizer
2026-03-07 10:29:59 +00:00
Roland Tannous
d60cd2843f
fix: patch Qwen3.5 broken tokenizer_class TokenizersBackend across all backends
2026-03-07 09:43:25 +00:00
Roland Tannous
bd60562145
fix: bump transformers 5.x pin from 5.1.0 to 5.2.0 for Qwen3.5 support
2026-03-07 09:10:09 +00:00
Roland Tannous
0b3397cc3a
fix: fail fast if runtime pip install of transformers 5.x fails
2026-03-07 08:40:25 +00:00
Roland Tannous
f3aeceeb24
fix: join prior pump thread before starting new training job
2026-03-07 08:37:03 +00:00
Roland Tannous
f7a3092cbd
fix: correct project root depth in model_config.py vision check
2026-03-07 08:15:29 +00:00
Roland Tannous
728420b290
fix: drain stale events from resp_queue after generation cancel
2026-03-07 08:12:16 +00:00
Samit
5f902af456
fixed model unload before load
2026-03-06 22:01:27 -08:00
Roland Tannous
25b51fad3b
fix: wait for training shutdown before export load, clear stop flag on reset
...
1. Export route: stop_training() only signals the subprocess — wait up to
30s for it to actually exit before loading the export checkpoint, avoiding
a GPU memory race.
2. Training reset: clear _should_stop so /api/train/status returns phase=idle
instead of staying stuck on phase=stopped after a user-triggered stop.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 04:16:10 +00:00
Roland Tannous
609e3168a1
fix: serialize generation with _gen_lock to prevent concurrent queue readers
...
Two overlapping /chat/completions requests could both read from the shared
resp_queue, consuming and dropping each other's token events. Replace the
request_id filtering (which silently dropped non-matching messages) with a
threading.Lock that serializes generation — correct for single-GPU inference.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 04:06:51 +00:00
Roland Tannous
9470957bb9
fix: log final GGUF file locations after relocation
2026-03-06 18:04:27 +00:00
Roland Tannous
5a828ebd43
fix: increase export timeout to 1 hour for large model GGUF conversion
2026-03-06 17:59:42 +00:00
Roland Tannous
4b7ad23b3a
feat: broaden Qwen3.5 matching to cover entire family
2026-03-06 16:48:28 +00:00
Roland Tannous
ed1e63c814
feat: add Qwen3.5-35B-A3B and Qwen3-Next to transformers 5.x model list
2026-03-06 10:54:48 +00:00
Shine1i
5e5feb5c00
feat(recipe-studio, validators): tweak OXC validator with lint suppression support and improve error normalization logic
2026-03-06 09:40:53 +01:00
Roland Tannous
d910759121
feat: add OpenAI-compatible /v1/chat/completions endpoint
2026-03-06 07:48:09 +00:00
Roland Tannous
c3bc19494f
fix: pin huggingface_hub==1.3.0 in .venv_t5 (satisfies transformers 5.x)
2026-03-06 06:19:28 +00:00
Roland Tannous
c5f4503b9e
fix: unload competing subprocesses before load across all routes
2026-03-06 06:05:31 +00:00
Roland Tannous
6b32af0bdc
feat: subprocess-based export, pin huggingface_hub==0.36.0
2026-03-06 06:03:09 +00:00
Roland Tannous
b5cfd0952c
fix: use subprocess with transformers 5.x for vision detection
...
Models like GLM-4.7-Flash have architectures (glm4_moe_lite) that
AutoConfig in the main process (transformers 4.57.x) can't recognize.
Instead of a raw config.json workaround, run the AutoConfig check in
a subprocess with .venv_t5/ activated — same pattern as training and
inference workers. This is more robust and consistent.
2026-03-06 04:51:23 +00:00
Roland Tannous
e5c7a18f72
fix: handle unrecognized model architectures in vision detection
...
AutoConfig.from_pretrained() fails for models needing transformers 5.x
(e.g. glm4_moe_lite) when running with 4.57.x. Add a raw config.json
fallback that bypasses AutoConfig's architecture registry — fetches
config.json directly from local path or HuggingFace Hub and checks
for vision indicators without needing the architecture to be registered.
2026-03-06 04:46:51 +00:00
Roland Tannous
1167be2798
refactor: consolidate version switching to .venv_t5, remove .venv_overlay
...
All version switching now uses .venv_t5/ (pre-installed by setup.sh).
The old .venv_overlay/ with runtime pip installs is removed.
ensure_transformers_version() (used only by export) now does a
lightweight sys.path swap instead of pip installing at runtime.
2026-03-06 04:37:06 +00:00
Shine1i
93063c3212
feat(recipe-studio, validators): extend OXC validator with code shape support and integrate into recipe studio
2026-03-06 02:04:05 +01:00
Shine1i
cf8cb9109b
feat(recipe-studio): add support for inference_extra_body configuration with collapsible UI and enhanced validation logic
2026-03-05 23:33:48 +01:00
Roland Tannous
cbe2896705
fix: unload inference model before training to free GPU memory
...
When starting training, shut down the inference subprocess first
so the training subprocess has full GPU memory available.
2026-03-05 22:28:11 +00:00
Roland Tannous
31334cece1
fix: indentation error in orchestrator load_model
2026-03-05 19:43:30 +00:00
Shine1i
3cafc0506e
feat(data-recipes, validators): extend OXC validator with linting mode support and integrate new modes into recipe studio
2026-03-05 20:19:31 +01:00