llama-server sends thinking/reasoning tokens as "reasoning_content" in the SSE delta (separate from "content"). The studio was only reading delta.content, so all reasoning tokens from models like Qwen3.5, Qwen3-Thinking, DeepSeek-R1, etc. were silently dropped. This caused "replies with nothing" for thinking models: the model would spend its entire token budget on reasoning, produce zero content tokens, and the user would see an empty response. Fix: read reasoning_content from the delta and wrap it in <think>...</think> tags. The frontend already has full support for these tags (parse-assistant-content.ts splits them into reasoning parts, reasoning.tsx renders a collapsible "Thinking..." indicator). Verified with Qwen3.5-27B-GGUF (UD-Q4_K_XL): - Before: "What is 2+2?" -> empty response (all tokens in reasoning) - After: shows collapsible thinking + answer "4" |
||
|---|---|---|
| .. | ||
| data_recipe | ||
| export | ||
| inference | ||
| training | ||
| __init__.py | ||