unsloth/studio/backend/core/inference
Daniel Han 3a118e5df0 studio: auto-continue when model stops mid-plan
The existing intent-signal re-prompt fires only when tools are armed
and the response is short. Models often stop mid-plan in other shapes
too: a trailing "Let me clone the repo.", a "Let me ...:" header
followed by a numbered list, or a bare trailing colon. When this
happens the turn ends with the structured workload only partially
delivered.

Add a neutral "Continue." nudge that runs alongside the tool-coercive
re-prompt:

- _TRAILING_PLAN_INTENT, _TRAILING_PLAN_LIST, _TRAILING_PLAN_COLON
  cover the three observed shapes, scanned over the last 600 chars.
- _trailing_plan_hit() returns True if any of them match.
- _MAX_CONTINUES (3) is independent of _MAX_REPROMPTS so the two
  paths cannot starve each other.
- _continue_count threads through the agentic loop; auto-continue
  fires with "Continue." regardless of tool armament.

Regex tested against the patterns above plus negative controls
(complete sentences, benign "let me" earlier in the buffer) before
landing.
2026-05-19 04:35:39 +00:00
..
__init__.py Final cleanup 2026-03-12 18:28:04 +00:00
_html_to_md.py fix: studio web search SSL failures and empty page content (#4754) 2026-04-01 06:12:02 -07:00
anthropic_compat.py Studio: support images on /v1/messages (Anthropic-compat) (#5128) 2026-04-22 03:25:07 +04:00
audio_codecs.py Add native GGUF intake to Studio (#5246) 2026-05-04 11:46:18 +02:00
defaults.py studio: engage draft-mtp on vision MTP GGUFs (drop incorrect vision gate) (#5560) 2026-05-18 08:42:55 -07:00
external_provider.py fix(studio): handle expired OpenAI shell-tool containers without surfacing error in chat (#5547) 2026-05-18 05:47:57 -07:00
inference.py Pin bitsandbytes to continuous-release_main on ROCm (4-bit decode fix) (#4954) 2026-04-10 06:25:39 -07:00
key_exchange.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
llama_cpp.py studio: auto-continue when model stops mid-plan 2026-05-19 04:35:39 +00:00
llama_server_args.py Studio: auto-enable MTP speculative decoding for MTP GGUFs (#5527) 2026-05-18 00:15:42 -07:00
mlx_inference.py MLX training support for Studio on Apple Silicon (#5340) 2026-05-14 05:24:20 -07:00
orchestrator.py Add native GGUF intake to Studio (#5246) 2026-05-04 11:46:18 +02:00
providers.py Polish/cloud to providers (#5450) 2026-05-15 19:29:21 +04:00
tools.py studio: tighten sandbox blocklist precision (bash, hf upload, NOFILE) (#5487) 2026-05-18 00:01:17 -07:00
worker.py studio: load cached GGUF models when fully offline (#5505) 2026-05-17 21:25:39 -07:00