unsloth/studio/backend
Roland Tannous 0e7c8a2e5e Switch GGUF backend from /v1/completions to /v1/chat/completions
Fixes two bugs:
1. Chat template tags (<|im_start|>, <|im_end|>) leaking into output
   because /v1/completions treated them as literal text
2. Image hallucination because image_b64 was never passed to llama-server

Now llama-server handles chat templates natively and receives images
as OpenAI-format multimodal content parts for vision models.
2026-02-24 19:21:01 +04:00
..
assets Add GLM, Qwen3 MoE, TinyQwen3 MoE, and Ministral 3 VL model defaults and GLM train_on_responses_only mapping 2026-02-23 05:51:43 +00:00
auth fix: rollback auth.db user row if token generation fails during setup 2026-02-13 10:11:07 +00:00
core Switch GGUF backend from /v1/completions to /v1/chat/completions 2026-02-24 19:21:01 +04:00
loggers root studio folder 2026-02-02 09:13:49 +00:00
models Use llama-server -hf mode, add GGUF variant selector, fix vision detection 2026-02-24 19:03:06 +04:00
requirements add huggingface-hub==0.36.0 due to colab error 2026-02-20 18:26:20 +00:00
routes Switch GGUF backend from /v1/completions to /v1/chat/completions 2026-02-24 19:21:01 +04:00
state root studio folder 2026-02-02 09:13:49 +00:00
tests added @needs_torch to test_cuda_oom 2026-02-11 16:12:37 +00:00
utils Use llama-server -hf mode, add GGUF variant selector, fix vision detection 2026-02-24 19:03:06 +04:00
colab.py feat: Add simple 2-cell Colab notebook (no tunnel needed) 2026-02-17 04:57:30 -06:00
main.py feat: clear unsloth_compiled_cache on startup, shutdown, and between model loads 2026-02-23 07:26:22 +00:00
run.py fix path in run_server 2026-02-14 05:06:17 +00:00