Commit graph

708 commits

Author SHA1 Message Date
Roland Tannous
f15970c02a fix: use subprocess with transformers 5.x for vision detection
Models like GLM-4.7-Flash have architectures (glm4_moe_lite) that
AutoConfig in the main process (transformers 4.57.x) can't recognize.
Instead of a raw config.json workaround, run the AutoConfig check in
a subprocess with .venv_t5/ activated — same pattern as training and
inference workers. This is more robust and consistent.
2026-03-06 04:51:23 +00:00
Roland Tannous
67121ce427 fix: handle unrecognized model architectures in vision detection
AutoConfig.from_pretrained() fails for models needing transformers 5.x
(e.g. glm4_moe_lite) when running with 4.57.x. Add a raw config.json
fallback that bypasses AutoConfig's architecture registry — fetches
config.json directly from local path or HuggingFace Hub and checks
for vision indicators without needing the architecture to be registered.
2026-03-06 04:46:51 +00:00
Roland Tannous
d55e9abcca refactor: consolidate version switching to .venv_t5, remove .venv_overlay
All version switching now uses .venv_t5/ (pre-installed by setup.sh).
The old .venv_overlay/ with runtime pip installs is removed.
ensure_transformers_version() (used only by export) now does a
lightweight sys.path swap instead of pip installing at runtime.
2026-03-06 04:37:06 +00:00
Manan17
cf9eaa2add Fixing dataset split issues 2026-03-06 01:13:55 +00:00
Shine1i
4bd7eca3f3 feat(recipe-studio, validators): extend OXC validator with code shape support and integrate into recipe studio 2026-03-06 02:04:05 +01:00
Shine1i
769515cb5f feat(recipe-studio): add inference_timeout configuration and validation logic 2026-03-06 00:51:28 +01:00
Shine1i
5c33886cce feat(recipe-studio): add support for inference_extra_body configuration with collapsible UI and enhanced validation logic 2026-03-05 23:33:48 +01:00
Roland Tannous
ce9bfd7476 fix: unload inference model before training to free GPU memory
When starting training, shut down the inference subprocess first
so the training subprocess has full GPU memory available.
2026-03-05 22:28:11 +00:00
imagineer99
4c634c4d2a fix: guard recharts ResponsiveContainer behind measured container dimensions 2026-03-05 21:54:27 +00:00
Roland Tannous
e866d16787 fix: indentation error in orchestrator load_model 2026-03-05 19:43:30 +00:00
Shine1i
d687ca6437 feat(data-recipes, validators): extend OXC validator with linting mode support and integrate new modes into recipe studio 2026-03-05 20:19:31 +01:00
Roland Tannous
b598cf0c14 fix: always spawn fresh subprocess per model load
Reusing a subprocess after unsloth patches torch internals causes
inspect.getsource() failures when loading a different model type.
Each load now gets a clean Python interpreter.
2026-03-05 19:15:37 +00:00
Roland Tannous
5c3efd610f fix: use mp.Event for instant cross-process generation cancel
Replaces cmd_queue-based cancel polling with a shared mp.Event.
Fixes two issues:
- Loading a new model while generating no longer hangs (cancel is instant)
- Subprocess shuts down cleanly after explicit stop generation
2026-03-05 18:54:17 +00:00
Shine1i
f9eabf285d feat(data-recipes, validators): add OXC validator runtime and integration with recipe studio 2026-03-05 19:48:26 +01:00
Roland Tannous
b70faf8cb7 feat: subprocess-based inference for transformers version switching
Inference now runs in a persistent subprocess, solving the same
transformers version-switching problem that was fixed for training.
The subprocess stays alive between requests (model in GPU memory)
and is only restarted when switching transformers versions.

New files:
- core/inference/worker.py: subprocess entry point with command loop
- core/inference/orchestrator.py: parent-side proxy with same API

Modified:
- core/inference/__init__.py: exports orchestrator as default backend
- routes/inference.py: removed in-process ensure_transformers_version()
2026-03-05 17:47:57 +00:00
Roland Tannous
021c3aafdd fix: handle None job_id before first training run 2026-03-05 16:59:37 +00:00
Roland Tannous
36b5c6af88 fix: lazy imports in core/__init__ to prevent subprocess importing ML libs early 2026-03-05 16:56:45 +00:00
Roland Tannous
794b8fe866 fix: exclude bitsandbytes from module purge to prevent duplicate operator registration 2026-03-05 16:40:20 +00:00
Roland Tannous
06fbeb22cf fix: remove in-process version switching from models routes 2026-03-05 16:22:32 +00:00
Roland Tannous
fb29d7f999 fix: remove UnslothTrainer/get_trainer from core __init__ exports 2026-03-05 15:57:07 +00:00
Roland Tannous
f90af41c5f feat: subprocess-based training for transformers version switching 2026-03-05 15:40:32 +00:00
Shine1i
6fc829e3ef merge: nightly into feature/data-reciper-enchansments 2026-03-05 14:51:08 +01:00
Shine1i
3ff0e74216 feat(recipe-studio): runtime edge handling with template refs and reversed edge support 2026-03-05 14:46:48 +01:00
Shine1i
b54238bbf5 refactor(studio): replace inputValue with searchQuery for improved clarity, add input reason tracking, and streamline dataset filtering logic 2026-03-05 14:06:17 +01:00
Shine1i
5674c6a815 feat(recipe-studio): improve tab switch fit logic with animation and delay support 2026-03-05 13:29:38 +01:00
Shine1i
2ff1d4de36 feat(recipe-studio): normalize and slugify run_name, update job naming logic 2026-03-05 12:25:51 +01:00
Shine1i
f988202290 refactor(studio): add local data-recipe dataset selection + training wiring 2026-03-05 12:25:51 +01:00
Shine1i
59995d8447 feat(data-recipes, recipe-studio): refactor and enhance recipe templates with updated model configurations, structure changes, and added validation logic 2026-03-05 12:14:01 +01:00
Shine1i
015e93491f feat(recipe-studio): persist advanced collapsible states across components and sessions 2026-03-05 11:56:40 +01:00
Shine1i
e50b2345fd feat(data-recipes, recipe-studio): recipies changes, image context selector 2026-03-05 11:46:42 +01:00
Shine1i
b43e9033fd feat(data-recipes): update recipe templates 2026-03-05 11:29:20 +01:00
Manan17
5d1a162ddd remove tracked OuteTTS embedded repo reference 2026-03-05 08:44:23 +00:00
Manan17
8203637d89 resolved merge conflicts 2026-03-05 07:59:43 +00:00
Manan17
2fa933640e fix SNAC training crash on variable-length sequences with DataCollatorForSeq2Seq 2026-03-05 07:04:53 +00:00
Roland Tannous
f57664e268 Merge nightly into feature/transformers-v5-support 2026-03-05 06:49:44 +00:00
Roland Tannous
37bff450d5 Merge pull request #314 from unslothai/fix/vlm-dataset-conversion-error-handling-local
Fix VLM training abort on URL-based dataset conversion failure
2026-03-05 10:10:58 +04:00
Roland Tannous
a1706c894f fix: check for http(s) prefix instead of bare string type for URL detection 2026-03-05 06:10:10 +00:00
Roland Tannous
2116cc5cca fix: remove benchmark scripts from git tracking
These are standalone benchmark scripts that were force-added despite being
gitignored. They have no test functions and run network calls at module
level, which breaks pytest collection in CI.
2026-03-05 06:06:47 +00:00
imagineer99
ebb57c765d feat(data-recipes): add OCR learning recipe template 2026-03-05 00:58:29 +00:00
Roland Tannous
b38df21bb5 fix: clear dataset slice state when switching to uploaded file
Prevents stale slice values from silently truncating uploaded datasets.
2026-03-04 23:42:23 +00:00
Roland Tannous
e04b9d53d6 feat: parallel URL image probe with time estimate and progress reporting
- Add 200-sample parallel probe using ThreadPoolExecutor + safe_num_proc
  to estimate download speed and failure rate before full conversion
- Abort with clear error if >=30% of probe images fail to download
- Show estimated download time in the training overlay modal
- Parallel batch conversion for URL-based datasets (vs sequential for local)
- Add warning field to /check-format response for URL-based image datasets
- Display URL warning in dataset preview dialog (amber banner)
- Thread progress_callback from trainer through format_and_template_dataset
  to convert_to_vlm_format for real-time status updates
2026-03-04 23:40:38 +00:00
Roland Tannous
db30b4105f test: add parallel download benchmark with ThreadPoolExecutor 2026-03-04 23:29:43 +00:00
Roland Tannous
1f03754c95 feat: add tqdm progress bar to VLM conversion and download benchmark test 2026-03-04 23:29:43 +00:00
Roland Tannous
63f723cc36 fix: add early probe to fail fast on datasets with too many broken image URLs 2026-03-04 23:29:43 +00:00
Roland Tannous
9487d17b94 fix: use fsspec for URL image downloads with per-sample error handling 2026-03-04 23:29:43 +00:00
Roland Tannous
e6eea64df5 test: add URL image loading comparison script 2026-03-04 23:29:43 +00:00
Roland Tannous
b376c54c21 fix: abort training pipeline on dataset conversion failure 2026-03-04 23:29:43 +00:00
Roland Tannous
bc244aeb23 fix: cast URL image columns to HF Image() type in VLM conversion 2026-03-04 23:29:43 +00:00
Roland Tannous
945cfa9460 fix: remove unnecessary tooltip copy from train split start 2026-03-04 23:24:09 +00:00
Roland Tannous
263cd19c44 refactor: move train split slice controls back to Advanced section
Place Train Split Start / End inputs inside the Advanced collapsible
with descriptive tooltips clarifying they slice the training split.
Revert the selectors component to its original eval-split-only layout.
2026-03-04 23:24:09 +00:00