unsloth/studio/backend/core/inference
alkinun f747108212
studio: extract tool-call XML parser into a reusable helper module (#5583)
Move the inline tool-call XML parser and stripper out of
studio/backend/core/inference/llama_cpp.py into a new
studio/backend/core/tool_healing.py so external inference servers
(llama-server wrappers, llama-swap, custom shims) can reuse the same
logic without importing the inference orchestrator, structlog, httpx,
or anything from torch / transformers / unsloth.

Closes #5502.

What this PR does:

- New file studio/backend/core/tool_healing.py contains the regex
  constants (_TOOL_CLOSED_PATS, _TOOL_ALL_PATS, _TC_JSON_START_RE,
  _TC_FUNC_START_RE, _TC_END_TAG_RE, _TC_FUNC_CLOSE_RE,
  _TC_PARAM_START_RE, _TC_PARAM_CLOSE_RE), parse_tool_calls_from_text,
  and strip_tool_call_markup. The regexes and function bodies are
  byte-for-byte the same as the previous inline implementation in
  llama_cpp.py; only the @staticmethod decorator and the closure-only
  `if not auto_heal_tool_calls: return text` short-circuit are dropped
  (the latter stays in the caller as a fast path when healing is off).
- studio/backend/core/inference/llama_cpp.py now imports the regexes
  and helpers from .tool_healing. LlamaCppBackend._parse_tool_calls_from_text
  becomes a one-line delegate; the _strip_tool_markup closure keeps the
  auto_heal_tool_calls fast path and delegates the work.
- Helper module imports cleanly without torch, transformers, structlog,
  httpx, or numpy. studio.backend.core itself is already stdlib-only
  at import time (lazy __getattr__), so `from
  studio.backend.core.tool_healing import parse_tool_calls_from_text,
  strip_tool_call_markup` is the lightweight import path issue #5502
  asked for.

No behaviour change for existing Studio paths. parse_tool_calls_from_text
and strip_tool_call_markup produce the same OpenAI-shape output the
old inline code produced for every input.

Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2026-05-19 05:06:17 -07:00
..
__init__.py Final cleanup 2026-03-12 18:28:04 +00:00
_html_to_md.py fix: studio web search SSL failures and empty page content (#4754) 2026-04-01 06:12:02 -07:00
anthropic_compat.py Studio: support images on /v1/messages (Anthropic-compat) (#5128) 2026-04-22 03:25:07 +04:00
audio_codecs.py Add native GGUF intake to Studio (#5246) 2026-05-04 11:46:18 +02:00
defaults.py studio: engage draft-mtp on vision MTP GGUFs (drop incorrect vision gate) (#5560) 2026-05-18 08:42:55 -07:00
external_provider.py fix(studio): handle expired OpenAI shell-tool containers without surfacing error in chat (#5547) 2026-05-18 05:47:57 -07:00
inference.py Pin bitsandbytes to continuous-release_main on ROCm (4-bit decode fix) (#4954) 2026-04-10 06:25:39 -07:00
key_exchange.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
llama_cpp.py studio: extract tool-call XML parser into a reusable helper module (#5583) 2026-05-19 05:06:17 -07:00
llama_server_args.py Studio: auto-enable MTP speculative decoding for MTP GGUFs (#5527) 2026-05-18 00:15:42 -07:00
mlx_inference.py MLX training support for Studio on Apple Silicon (#5340) 2026-05-14 05:24:20 -07:00
orchestrator.py Add native GGUF intake to Studio (#5246) 2026-05-04 11:46:18 +02:00
providers.py Polish/cloud to providers (#5450) 2026-05-15 19:29:21 +04:00
tools.py studio: tighten sandbox blocklist precision (bash, hf upload, NOFILE) (#5487) 2026-05-18 00:01:17 -07:00
worker.py studio: load cached GGUF models when fully offline (#5505) 2026-05-17 21:25:39 -07:00