unsloth/studio/backend
Roland Tannous 4eabc74f34 feat: subprocess-based inference for transformers version switching
Inference now runs in a persistent subprocess, solving the same
transformers version-switching problem that was fixed for training.
The subprocess stays alive between requests (model in GPU memory)
and is only restarted when switching transformers versions.

New files:
- core/inference/worker.py: subprocess entry point with command loop
- core/inference/orchestrator.py: parent-side proxy with same API

Modified:
- core/inference/__init__.py: exports orchestrator as default backend
- routes/inference.py: removed in-process ensure_transformers_version()
2026-03-05 17:47:57 +00:00
..
assets Add GLM, Qwen3 MoE, TinyQwen3 MoE, and Ministral 3 VL model defaults and GLM train_on_responses_only mapping 2026-02-23 05:51:43 +00:00
auth fix: replace datetime.UTC with timezone.utc for Python 3.9+ compatibility 2026-02-24 14:37:00 -06:00
core feat: subprocess-based inference for transformers version switching 2026-03-05 17:47:57 +00:00
loggers root studio folder 2026-02-02 09:13:49 +00:00
models feat: parallel URL image probe with time estimate and progress reporting 2026-03-04 23:40:38 +00:00
requirements feat: introduce single-env Python dependency management for streamlined compatibility 2026-02-24 07:45:40 +01:00
routes feat: subprocess-based inference for transformers version switching 2026-03-05 17:47:57 +00:00
state root studio folder 2026-02-02 09:13:49 +00:00
tests added @needs_torch to test_cuda_oom 2026-02-11 16:12:37 +00:00
utils fix: exclude bitsandbytes from module purge to prevent duplicate operator registration 2026-03-05 16:40:20 +00:00
colab.py feat: Add simple 2-cell Colab notebook (no tunnel needed) 2026-02-17 04:57:30 -06:00
main.py Merge nightly into feature/transformers-v5-support 2026-03-05 06:49:44 +00:00
run.py fix path in run_server 2026-02-14 05:06:17 +00:00