unsloth/studio/backend/core/inference
danielhanchen fd53dab931 studio: tighten trailing-plan list anchor + grow tool-iter cap on demand
Codex P2 review on #5549 surfaced two related risks in the auto-continue
plumbing:

1. `_TRAILING_PLAN_LIST` was compiled with `(?ims)`. The `m` flag makes
   the terminal `\s*$` match end-of-line, so a complete answer like
   "Here's my plan:\n- a\n- b\n\nDone, that should work." still matched
   the list-block sub-pattern and tripped a spurious `Continue.` retry.
   Drop the `m` (and the unused `s`) flag and re-anchor with `\Z` so the
   list pattern only fires when the list is genuinely the last thing in
   the buffer.

2. The agent loop pre-reserved `_MAX_REPROMPTS + _MAX_CONTINUES` (= 6)
   extra iterations on top of the caller's `max_tool_iterations`
   unconditionally. That weakens the caller-provided budget: a turn
   that never trips the reprompt or continue path could still run up
   to N+6 full iterations and execute their tool calls.

   Switch the bound to a dynamic cap that grows only as reprompts /
   continues are actually consumed: `iteration < max_tool_iterations +
   _reprompt_count + _continue_count`. With both counters at zero the
   loop honors the caller cap exactly; once a continue or reprompt
   fires it earns its own slot back.

   Implemented with `itertools.count()` so the existing `continue`
   statements in the loop body keep their semantics.

Regex behaviour pinned by `scripts/r6_trailing_plan_regex_test.py`
(updated separately for the new list-tail case).
2026-05-19 04:35:39 +00:00
..
__init__.py Final cleanup 2026-03-12 18:28:04 +00:00
_html_to_md.py fix: studio web search SSL failures and empty page content (#4754) 2026-04-01 06:12:02 -07:00
anthropic_compat.py studio: reset Anthropic adapter cursor on auto-continue boundary 2026-05-19 04:35:39 +00:00
audio_codecs.py Add native GGUF intake to Studio (#5246) 2026-05-04 11:46:18 +02:00
defaults.py studio: engage draft-mtp on vision MTP GGUFs (drop incorrect vision gate) (#5560) 2026-05-18 08:42:55 -07:00
external_provider.py fix(studio): handle expired OpenAI shell-tool containers without surfacing error in chat (#5547) 2026-05-18 05:47:57 -07:00
inference.py Pin bitsandbytes to continuous-release_main on ROCm (4-bit decode fix) (#4954) 2026-04-10 06:25:39 -07:00
key_exchange.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
llama_cpp.py studio: tighten trailing-plan list anchor + grow tool-iter cap on demand 2026-05-19 04:35:39 +00:00
llama_server_args.py Studio: auto-enable MTP speculative decoding for MTP GGUFs (#5527) 2026-05-18 00:15:42 -07:00
mlx_inference.py MLX training support for Studio on Apple Silicon (#5340) 2026-05-14 05:24:20 -07:00
orchestrator.py Add native GGUF intake to Studio (#5246) 2026-05-04 11:46:18 +02:00
providers.py Polish/cloud to providers (#5450) 2026-05-15 19:29:21 +04:00
tools.py studio: tighten sandbox blocklist precision (bash, hf upload, NOFILE) (#5487) 2026-05-18 00:01:17 -07:00
worker.py studio: load cached GGUF models when fully offline (#5505) 2026-05-17 21:25:39 -07:00