Commit graph

36 commits

Author SHA1 Message Date
Roland Tannous
5f98d232d0 fix: align llama-server binary discovery with upstream unsloth-zoo paths 2026-03-03 17:03:01 +00:00
Roland Tannous
e7619a1291 Move llama.cpp clone/build from in-tree to ~/.unsloth/llama.cpp
- setup.sh: builds at ~/.unsloth/llama.cpp instead of ./llama.cpp
- setup.ps1: builds at %USERPROFILE%/.unsloth/llama.cpp
- inference llama_cpp.py: searches ~/.unsloth/ first, in-tree as legacy
- export.py: updated comments (unsloth-zoo handles path natively)
2026-03-02 04:04:41 +00:00
Roland Tannous
70d1567fe3 Download GGUF via huggingface_hub instead of llama-server -hf (fixes HTTPS not supported on Windows) 2026-03-01 13:05:10 +00:00
Roland Tannous
0430d22cc2 Auto-add CUDA DLLs to PATH when launching llama-server on Windows 2026-03-01 13:05:10 +00:00
Roland Tannous
afa1344452 Build llama.cpp in-tree, auto-detect driver CUDA version for compatible toolkit 2026-03-01 13:05:10 +00:00
Roland Tannous
af90c9c3d2 Fix llama-server binary lookup for Windows (.exe, Release dir, ~/.unsloth) 2026-03-01 13:05:10 +00:00
Roland Tannous
986bef4f99 fix: support mmproj for local vision GGUF models + fix Windows pipe deadlock 2026-03-01 12:58:38 +00:00
Manan17
b4311cca82 Aggregating sharded models, showing fit/oom for quantizations 2026-02-27 08:23:15 +00:00
Roland Tannous
efaa0bacfb Merge branch 'nightly' into feat/gguf-llama-cpp-inference 2026-02-25 16:06:03 +04:00
Roland Tannous
2ebeba8588 Switch GGUF backend from /v1/completions to /v1/chat/completions
Fixes two bugs:
1. Chat template tags (<|im_start|>, <|im_end|>) leaking into output
   because /v1/completions treated them as literal text
2. Image hallucination because image_b64 was never passed to llama-server

Now llama-server handles chat templates natively and receives images
as OpenAI-format multimodal content parts for vision models.
2026-02-24 19:21:01 +04:00
Roland Tannous
3ee4f1359a Use llama-server -hf mode, add GGUF variant selector, fix vision detection
Replace Python-side GGUF download with llama-server's native -hf flag for
HuggingFace repos. Add frontend variant picker so users can choose
quantization (Q4_K_M, Q8_0, BF16, etc.) with file sizes. Fix vision
detection via mmproj files instead of hardcoding is_vision=False.
2026-02-24 19:03:06 +04:00
Roland Tannous
5b7555cd3f Fix llama-server: build in-tree, fix path resolution, add LD_LIBRARY_PATH 2026-02-24 18:19:29 +04:00
Roland Tannous
2f985ccbb5 Add GGUF model inference via llama-server backend 2026-02-24 17:40:05 +04:00
Manan17
6fa6b0bf28 Fixing base model export issue for vlms 2026-02-24 01:34:11 +00:00
Roland Tannous
198433363a feat: clear unsloth_compiled_cache on startup, shutdown, and between model loads 2026-02-23 07:26:22 +00:00
Roland Tannous
ef118d0d05 fix: load proper vision processor from base model when FastVisionModel returns raw tokenizer, add tokenize=False to vision chat template 2026-02-21 04:40:29 +00:00
Manan17
e9710874e1 Mapping proper tokenizer for VLMs 2026-02-21 01:57:05 +00:00
Manan17
756aa56cd2 fixed the vlm's text only errors 2026-02-20 22:23:26 +00:00
Manan17
444ece6b07 Fixing compare feature 2026-02-19 20:15:44 +00:00
Shine1i
4be6eefed3 feat: support disabling top-k sampling with -1 and standardize normalization logic
- Updated top-k parameter range to accept -1 in models and frontend.
- Added utility to normalize top-k for backend compatibility.
2026-02-16 21:33:24 +01:00
Roland Tannous
6b839a1481 feat: add min_p sampling parameter to /chat/completions generation pipeline 2026-02-16 06:33:17 +00:00
Shine1i
f6397bf1ac feat: add cancelation support for chat generation and streaming tasks 2026-02-15 18:23:27 +01:00
Roland Tannous
4399687f93 strip extra debug statements 2026-02-14 19:23:51 +00:00
Roland Tannous
4d868e8d2b replace model unloading and peft loading mechanism for compare feature 2026-02-14 19:18:49 +00:00
Roland Tannous
225b3f1750 del model.peft_config instead of using model.delete_adapter 2026-02-14 17:32:15 +00:00
Roland Tannous
f122154cf3 added print statements for activate_lora_adapter 2026-02-14 17:25:37 +00:00
Roland Tannous
754ccf1a67 swipped logger for print statements as logger isn't propagating 2026-02-14 17:21:26 +00:00
Roland Tannous
b930a17b1d added logging 2026-02-14 17:09:07 +00:00
Roland Tannous
9bee0a3f63 exclude default from model.delete_adapter 2026-02-14 17:03:52 +00:00
Roland Tannous
0b305fd822 _apply_adapter_state now calls revert_to_base_model and activate_lora_adapter properly 2026-02-14 16:57:24 +00:00
Roland Tannous
8403bac48d feat(inference): add use_adapter field for per-request adapter toggling in compare mode 2026-02-14 14:52:13 +00:00
Roland Tannous
4ab8f81780 migrate _generate_vision_response to use TextIteratorStreamer + background thread 2026-02-14 09:30:32 +00:00
Roland Tannous
a6ee9ee957 use get_device() for device selection and clear_gpu_cache() for GPU memory cleanup in inference, trainer, and export 2026-02-11 16:56:52 +00:00
Roland Tannous
c17ba10f96 refactor/inference-api-routes-part-1 2026-02-03 16:57:57 +00:00
Roland Tannous
47ead076cf Refactor [dataset_utils.py](cci:7://file:///home/support/new-ui-prototype/studio/backend/utils/datasets/dataset_utils.py:0:0-0:0) into focused modules 2026-02-03 14:38:02 +00:00
Roland Tannous
75d8dcc824 root studio folder 2026-02-02 09:13:49 +00:00