unsloth/studio/backend/core/inference
Daniel Han 62827a9f3f Studio: soften tool-use nudge for small models and add synthesise directive
The existing _TOOL_ACTION_NUDGE tells the model "For any factual question,
call web_search" and "Never describe what you plan to do -- just call the
tool immediately". On small GGUF models (<9B) this causes two failure modes
we can measure:

1. First turn: the model calls web_search on questions it could answer from
   training data, often queuing several parallel searches in one assistant
   message. Measured on Qwen3.5-4B UD-Q4_K_XL, n=30: 28/30 tool_call, 2/30
   answer. On Qwen3.5-4B Q4_K_M: 30/30 tool_call, 0/30 answer.

2. Subsequent turns: even with tool results already in context, the
   "prefer tools" directive keeps dominating and the model searches again
   instead of synthesising. Measured on UD-Q4_K_XL with one tool_result
   present: 23/30 tool_call.

Two changes:

_TOOL_ACTION_NUDGE_SMALL: a softer nudge used for models under 9B. Asks
for tool use only when current information or a calculation is actually
needed, and explicitly discourages queuing multiple parallel tool calls.
Larger models keep the original aggressive nudge -- they weren't the ones
over-triggering.

_TOOL_SYNTHESISE_NUDGE: appended whenever the conversation already has
a tool result, and injected into the system message inside the internal
tool-call loop once the first tool result has been added. Phrasing
matters here -- framing this as a concrete action ("write the final
answer to the user's original question using what you have") works.
Framing it as an opt-out clause ("do not call more tools unless...") is
actually worse than no nudge, measured 33% vs 57% synthesis rate on
UD-Q4_K_XL.

Measured end-to-end on the same "How do you fine-tune an audio model
with Unsloth?" query, n=30 per cell:

Qwen3.5-4B UD-Q4_K_XL
                            OLD              NEW
  first turn tool_call      28/30 (93%)      2/30  (7%)
  first turn answer         2/30             28/30
  second turn tool_call     23/30 (77%)      6/30  (20%)
  second turn answer        7/30             24/30

Qwen3.5-4B Q4_K_M (LM Studio)
                            OLD              NEW
  first turn tool_call      30/30 (100%)     0/30  (0%)
  first turn answer         0/30             30/30
  second turn tool_call     2/30             0/30
  second turn answer        28/30            30/30

Test scripts under tests/test_ab_nudge.py and tests/test_stronger_synth.py.
2026-04-16 11:03:17 +00:00
..
__init__.py Final cleanup 2026-03-12 18:28:04 +00:00
_html_to_md.py fix: studio web search SSL failures and empty page content (#4754) 2026-04-01 06:12:02 -07:00
anthropic_compat.py Studio: Expose openai and anthropic compatible external API end points (#4956) 2026-04-13 21:08:11 +04:00
audio_codecs.py studio: per-model inference defaults, GGUF slider fix, reasoning toggle (#4325) 2026-03-16 06:37:55 -07:00
defaults.py UI Changes (#4782) 2026-04-02 08:05:55 -07:00
inference.py Pin bitsandbytes to continuous-release_main on ROCm (4-bit decode fix) (#4954) 2026-04-10 06:25:39 -07:00
llama_cpp.py Studio: soften tool-use nudge for small models and add synthesise directive 2026-04-16 11:03:17 +00:00
orchestrator.py fix(studio): prioritize curated defaults over HF download ranking in Recommended (#4792) 2026-04-02 10:46:53 -07:00
tools.py fix(studio): harden sandbox security for terminal and python tools (#4827) 2026-04-03 13:33:42 -07:00
worker.py split venv_t5 into tiered 5.3.0/5.5.0 and fix trust_remote_code (#4878) 2026-04-07 20:05:01 +04:00