Commit graph

37 commits

Author SHA1 Message Date
Shine1i
b08b606b21 feat(studio): studio storage roots path utilities 2026-03-09 23:48:31 +00:00
Roland Tannous
a0f03d3080 Add AGPL-3.0 SPDX headers to all source files 2026-03-09 20:17:45 +00:00
Roland Tannous
dd24009f06 Merge pull request #340 from unslothai/fix/auth-audio
Added auth to audio generate endpoint
2026-03-09 17:25:49 +04:00
Roland Tannous
dcedc4df56 merge nightly, resolve conflict in use-chat-model-runtime 2026-03-09 13:19:17 +00:00
Roland Tannous
83b1ff05ef respect trust_remote_code toggle, return helpful error when required 2026-03-09 13:06:55 +00:00
Roland Tannous
5ebd4de2ef backend: resolve trust_remote_code from YAML when not set by frontend 2026-03-09 11:58:23 +00:00
samit
1eb68678a0 added auth to audio generate endpoint 2026-03-08 17:46:54 -07:00
Roland Tannous
1b04a40bc8 Merge pull request #328 from unslothai/fix/chat-unloading-model
fixed model unload before load without validation
2026-03-09 04:40:05 +04:00
samit
6aa50d353f exposed trust_remote_code through the UI 2026-03-08 16:28:56 -07:00
Roland Tannous
7db2c90cc6 merge nightly into audio branch (mock test) 2026-03-08 10:23:44 +00:00
Samit
7ec41afaf0 fixed model unload before load 2026-03-06 22:01:27 -08:00
Roland Tannous
527db4ffff feat: add OpenAI-compatible /v1/chat/completions endpoint 2026-03-06 07:48:09 +00:00
Roland Tannous
c82e3d86bb fix: unload competing subprocesses before load across all routes 2026-03-06 06:05:31 +00:00
Roland Tannous
b70faf8cb7 feat: subprocess-based inference for transformers version switching
Inference now runs in a persistent subprocess, solving the same
transformers version-switching problem that was fixed for training.
The subprocess stays alive between requests (model in GPU memory)
and is only restarted when switching transformers versions.

New files:
- core/inference/worker.py: subprocess entry point with command loop
- core/inference/orchestrator.py: parent-side proxy with same API

Modified:
- core/inference/__init__.py: exports orchestrator as default backend
- routes/inference.py: removed in-process ensure_transformers_version()
2026-03-05 17:47:57 +00:00
Roland Tannous
f57664e268 Merge nightly into feature/transformers-v5-support 2026-03-05 06:49:44 +00:00
Roland Tannous
da00f5ed1d Merge branch 'nightly' into feature/support-for-audio-models 2026-03-02 15:55:25 +04:00
Roland Tannous
986bef4f99 fix: support mmproj for local vision GGUF models + fix Windows pipe deadlock 2026-03-01 12:58:38 +00:00
Manan17
8cdeb006b6 code cleanup 2026-03-01 08:04:38 +00:00
Manan17
9e89f31bc7 revamping up the code and adding inference 2026-03-01 02:30:31 +00:00
Roland Tannous
d6922f5e83 Merge pull request #282 from unslothai/fix/inference-auth
Added auth to inference endpoints
2026-02-27 13:18:38 +04:00
Manan17
b4311cca82 Aggregating sharded models, showing fit/oom for quantizations 2026-02-27 08:23:15 +00:00
samit
6a9969d67b added auth to inference endpoints 2026-02-27 00:20:36 -08:00
Roland Tannous
2ebeba8588 Switch GGUF backend from /v1/completions to /v1/chat/completions
Fixes two bugs:
1. Chat template tags (<|im_start|>, <|im_end|>) leaking into output
   because /v1/completions treated them as literal text
2. Image hallucination because image_b64 was never passed to llama-server

Now llama-server handles chat templates natively and receives images
as OpenAI-format multimodal content parts for vision models.
2026-02-24 19:21:01 +04:00
Roland Tannous
3ee4f1359a Use llama-server -hf mode, add GGUF variant selector, fix vision detection
Replace Python-side GGUF download with llama-server's native -hf flag for
HuggingFace repos. Add frontend variant picker so users can choose
quantization (Q4_K_M, Q8_0, BF16, etc.) with file sizes. Fix vision
detection via mmproj files instead of hardcoding is_vision=False.
2026-02-24 19:03:06 +04:00
Roland Tannous
2f985ccbb5 Add GGUF model inference via llama-server backend 2026-02-24 17:40:05 +04:00
Roland Tannous
778762eb28 Patch adapter_config.json with unsloth_training_method and auto-detect load_in_4bit for LoRA inference 2026-02-22 20:27:52 +00:00
Roland Tannous
f5b30448e8 Auto-switch transformers version (5.1.0/4.57.1) for Ministral-3, GLM-4.7-Flash, Qwen3-30B-A3B models with LoRA adapter resolution 2026-02-22 18:29:40 +00:00
Roland Tannous
6b839a1481 feat: add min_p sampling parameter to /chat/completions generation pipeline 2026-02-16 06:33:17 +00:00
Shine1i
f6397bf1ac feat: add cancelation support for chat generation and streaming tasks 2026-02-15 18:23:27 +01:00
sshah229
2483b98985 added the inference fetching from model mappers 2026-02-15 02:48:53 -07:00
Roland Tannous
ac8128519d decouple reliance of backend on frontend for is_lora 2026-02-14 20:13:50 +00:00
Roland Tannous
8403bac48d feat(inference): add use_adapter field for per-request adapter toggling in compare mode 2026-02-14 14:52:13 +00:00
Roland Tannous
480418b595 feat(inference): accept OpenAI multimodal content parts (image_url) in /chat/completions 2026-02-14 09:06:25 +00:00
Roland Tannous
c78cb11f81 feat: add OpenAI-compatible POST /chat/completions endpoint with streaming and non-streaming support 2026-02-12 19:00:05 +00:00
Roland Tannous
5ae20f6099 move inline pydantic models - fix existing models routes integration 2026-02-11 12:39:58 +00:00
Roland Tannous
c17ba10f96 refactor/inference-api-routes-part-1 2026-02-03 16:57:57 +00:00
Roland Tannous
75d8dcc824 root studio folder 2026-02-02 09:13:49 +00:00
Renamed from backend/routes/inference.py (Browse further)