Commit graph

986 commits

Author SHA1 Message Date
Roland Tannous
58467aac3d fix: validate pip exit codes for .venv_t5 installs in setup.ps1 2026-03-07 23:25:13 +00:00
Roland Tannous
b8577565cd add .venv_t5/ to .gitignore 2026-03-07 23:21:48 +00:00
Manan17
2ce36df03c check fir gated repo 2026-03-07 21:32:50 +00:00
Manan17
4faa6e9152 fixing sesame model 2026-03-07 19:06:00 +00:00
Roland Tannous
9543fa0d1c fix: scope dataloader_num_workers=0 to Windows + transformers 5.x only 2026-03-07 17:55:59 +00:00
Roland Tannous
dcebfe718a fix: prevent training hang on Windows by adding triton-windows support 2026-03-07 17:53:36 +00:00
Roland Tannous
c882a3d2f7 fix: propagate PYTHONPATH to child subprocesses, revert tokenizer patching 2026-03-07 11:28:24 +00:00
Roland Tannous
ac608be800 fix: patch TokenizersBackend in export output after save_pretrained 2026-03-07 10:57:51 +00:00
Roland Tannous
42bd976a2f fix: patch TokenizersBackend by model name - Qwen3.5→Qwen2Tokenizer, GLM→PreTrainedTokenizer 2026-03-07 10:29:59 +00:00
Roland Tannous
44e9b838ae fix: patch Qwen3.5 broken tokenizer_class TokenizersBackend across all backends 2026-03-07 09:43:25 +00:00
Roland Tannous
d5138c529f fix: bump transformers to 5.2.0 and pin huggingface_hub in setup.ps1 2026-03-07 09:12:12 +00:00
Roland Tannous
f101befca7 fix: bump transformers 5.x pin from 5.1.0 to 5.2.0 for Qwen3.5 support 2026-03-07 09:10:09 +00:00
Roland Tannous
2c4c598832 fix: fail fast if runtime pip install of transformers 5.x fails 2026-03-07 08:40:25 +00:00
Roland Tannous
082a6f876e fix: join prior pump thread before starting new training job 2026-03-07 08:37:03 +00:00
Roland Tannous
48456070a7 fix: correct project root depth in model_config.py vision check 2026-03-07 08:15:29 +00:00
Roland Tannous
717d5d621d fix: drain stale events from resp_queue after generation cancel 2026-03-07 08:12:16 +00:00
samit
51f5b56f1c modified the title 2026-03-06 22:33:30 -08:00
Samit
7ec41afaf0 fixed model unload before load 2026-03-06 22:01:27 -08:00
Roland Tannous
ad6739be7a fix: wait for training shutdown before export load, clear stop flag on reset
1. Export route: stop_training() only signals the subprocess — wait up to
   30s for it to actually exit before loading the export checkpoint, avoiding
   a GPU memory race.

2. Training reset: clear _should_stop so /api/train/status returns phase=idle
   instead of staying stuck on phase=stopped after a user-triggered stop.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 04:16:10 +00:00
Roland Tannous
08032e35a7 fix: add /v1 proxy entry to vite dev server config
Without this, /v1/chat/completions requests in local dev are served by
Vite instead of being proxied to the FastAPI backend.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 04:09:28 +00:00
Roland Tannous
2558125a14 fix: serialize generation with _gen_lock to prevent concurrent queue readers
Two overlapping /chat/completions requests could both read from the shared
resp_queue, consuming and dropping each other's token events. Replace the
request_id filtering (which silently dropped non-matching messages) with a
threading.Lock that serializes generation — correct for single-GPU inference.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 04:06:51 +00:00
Roland Tannous
cac409198f fix: log final GGUF file locations after relocation 2026-03-06 18:04:27 +00:00
Roland Tannous
e72293a033 fix: increase export timeout to 1 hour for large model GGUF conversion 2026-03-06 17:59:42 +00:00
Roland Tannous
2cfcaa8b61 feat: broaden Qwen3.5 matching to cover entire family 2026-03-06 16:48:28 +00:00
Roland Tannous
c31e8a6ed7 feat: add Qwen3.5-35B-A3B and Qwen3-Next to transformers 5.x model list 2026-03-06 10:54:48 +00:00
Manan17
8afaf55eb5 derive effective model type from isVisionModel for dataset search filterin 2026-03-06 08:47:56 +00:00
Shine1i
72468a3292 feat(recipe-studio, validators): tweak OXC validator with lint suppression support and improve error normalization logic 2026-03-06 09:40:53 +01:00
Roland Tannous
527db4ffff feat: add OpenAI-compatible /v1/chat/completions endpoint 2026-03-06 07:48:09 +00:00
Manan17
fcd52d9bd0 fixing update model type 2026-03-06 07:44:57 +00:00
imagineer99
e0e4619ee2 fix: harden chart container sizing with legacy event rechecks 2026-03-06 07:16:36 +00:00
Roland Tannous
7e59440029 fix: pin huggingface_hub==1.3.0 in .venv_t5 (satisfies transformers 5.x) 2026-03-06 06:19:28 +00:00
imagineer99
e231f366ea fix: preserve chart sizing updates without ResizeObserver 2026-03-06 06:06:50 +00:00
Roland Tannous
c82e3d86bb fix: unload competing subprocesses before load across all routes 2026-03-06 06:05:31 +00:00
Roland Tannous
661ac4be96 feat: subprocess-based export, pin huggingface_hub==0.36.0 2026-03-06 06:03:09 +00:00
imagineer99
6f13185379 fix: align slider fill bar with thumb across value range 2026-03-06 05:21:44 +00:00
Roland Tannous
f15970c02a fix: use subprocess with transformers 5.x for vision detection
Models like GLM-4.7-Flash have architectures (glm4_moe_lite) that
AutoConfig in the main process (transformers 4.57.x) can't recognize.
Instead of a raw config.json workaround, run the AutoConfig check in
a subprocess with .venv_t5/ activated — same pattern as training and
inference workers. This is more robust and consistent.
2026-03-06 04:51:23 +00:00
Roland Tannous
67121ce427 fix: handle unrecognized model architectures in vision detection
AutoConfig.from_pretrained() fails for models needing transformers 5.x
(e.g. glm4_moe_lite) when running with 4.57.x. Add a raw config.json
fallback that bypasses AutoConfig's architecture registry — fetches
config.json directly from local path or HuggingFace Hub and checks
for vision indicators without needing the architecture to be registered.
2026-03-06 04:46:51 +00:00
Roland Tannous
d55e9abcca refactor: consolidate version switching to .venv_t5, remove .venv_overlay
All version switching now uses .venv_t5/ (pre-installed by setup.sh).
The old .venv_overlay/ with runtime pip installs is removed.
ensure_transformers_version() (used only by export) now does a
lightweight sys.path swap instead of pip installing at runtime.
2026-03-06 04:37:06 +00:00
Manan17
cf9eaa2add Fixing dataset split issues 2026-03-06 01:13:55 +00:00
Shine1i
4bd7eca3f3 feat(recipe-studio, validators): extend OXC validator with code shape support and integrate into recipe studio 2026-03-06 02:04:05 +01:00
Shine1i
769515cb5f feat(recipe-studio): add inference_timeout configuration and validation logic 2026-03-06 00:51:28 +01:00
Shine1i
5c33886cce feat(recipe-studio): add support for inference_extra_body configuration with collapsible UI and enhanced validation logic 2026-03-05 23:33:48 +01:00
Roland Tannous
ce9bfd7476 fix: unload inference model before training to free GPU memory
When starting training, shut down the inference subprocess first
so the training subprocess has full GPU memory available.
2026-03-05 22:28:11 +00:00
imagineer99
4c634c4d2a fix: guard recharts ResponsiveContainer behind measured container dimensions 2026-03-05 21:54:27 +00:00
Roland Tannous
e866d16787 fix: indentation error in orchestrator load_model 2026-03-05 19:43:30 +00:00
Shine1i
d687ca6437 feat(data-recipes, validators): extend OXC validator with linting mode support and integrate new modes into recipe studio 2026-03-05 20:19:31 +01:00
Roland Tannous
b598cf0c14 fix: always spawn fresh subprocess per model load
Reusing a subprocess after unsloth patches torch internals causes
inspect.getsource() failures when loading a different model type.
Each load now gets a clean Python interpreter.
2026-03-05 19:15:37 +00:00
Roland Tannous
5c3efd610f fix: use mp.Event for instant cross-process generation cancel
Replaces cmd_queue-based cancel polling with a shared mp.Event.
Fixes two issues:
- Loading a new model while generating no longer hangs (cancel is instant)
- Subprocess shuts down cleanly after explicit stop generation
2026-03-05 18:54:17 +00:00
Shine1i
f9eabf285d feat(data-recipes, validators): add OXC validator runtime and integration with recipe studio 2026-03-05 19:48:26 +01:00
Roland Tannous
b70faf8cb7 feat: subprocess-based inference for transformers version switching
Inference now runs in a persistent subprocess, solving the same
transformers version-switching problem that was fixed for training.
The subprocess stays alive between requests (model in GPU memory)
and is only restarted when switching transformers versions.

New files:
- core/inference/worker.py: subprocess entry point with command loop
- core/inference/orchestrator.py: parent-side proxy with same API

Modified:
- core/inference/__init__.py: exports orchestrator as default backend
- routes/inference.py: removed in-process ensure_transformers_version()
2026-03-05 17:47:57 +00:00