unsloth/studio/backend/core
Daniel Han e9ea45b6a5
Studio: coerce tool_call arguments to dict before chat templating (fixes MLX tool follow-up error) (#6807)
* Studio: coerce tool_call arguments to dict before chat templating

Strict tool chat templates (e.g. mlx-community Qwen3.5 checkpoints) iterate
arguments.items() and raise "TypeError: Can only get item pairs from a mapping"
when a prior assistant tool call is re-rendered on the next turn. The agentic
loop stores arguments in the OpenAI JSON-string form (as_assistant_tool_call),
which is correct on the wire and for llama-server, but the transformers / MLX
paths apply_chat_template directly and hit the strict Jinja templates.

Normalize each assistant tool_call's function.arguments from a JSON string to a
dict inside apply_chat_template_for_generation (shared by both the MLX and
safetensors paths). A dict renders on strict and lenient templates alike;
non-JSON / non-dict values are left untouched, and the OpenAI-format
as_assistant_tool_call (used by the GGUF path + API responses) is unchanged.

Verified against the real mlx-community/Qwen3.5-2B-8bit template: string args
raised the tester's error, the fix renders cleanly, and the lenient
unsloth/Qwen3.5-0.8B template still works.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: make tool-arg coercion a string-first fallback (non-regressive)

Render the original OpenAI string-arg form first and only coerce arguments to a
dict when the template raises the mapping TypeError, instead of always coercing.
Any template that already renders is now byte-identical (a template that emits
arguments verbatim keeps the JSON string, not a Python dict repr).

Verified across Llama-3, Qwen2.5, Qwen3, Qwen3.5, Phi-3.5 (byte-identical) and
mlx-community/Qwen3.5-2B-8bit (strict -> fixed). Gemma-3 / Mistral tool-template
errors are unrelated (role alternation / tool-id length) and identical with or
without the change.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Make core.inference package init lazy so dependency-light helpers import standalone

Importing any core.inference submodule ran the package __init__, which
eagerly imported orchestrator and llama_cpp; both pull loggers ->
structlog (and httpx), so a dependency-light helper like
chat_template_helpers dragged in the full heavy stack and its unit test
failed to collect in a backend env without structlog. Defer those
imports to attribute access via PEP 562 __getattr__, mirroring the lazy
pattern already in core/__init__.py. The re-exports resolve unchanged on
first access.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Retry dict-coercion for strict templates that raise non-TypeError

apply_chat_template_for_generation only retried the OpenAI JSON-string arguments
coercion when the first render raised TypeError (the arguments.items() form). The
bundled gemma-4.jinja instead rejects string arguments with raise_exception, which
surfaces as a Jinja error, so a second tool turn with string function.arguments
propagated and failed rather than retrying with the parsed dict.

Broaden the outer catch to Exception, still gated on there being a string arg to
normalize (normalized is messages -> re-raise), so unrelated template errors and
templates that already render are unaffected.

* Tighten comments in tool-call argument coercion helper and tests

* Tighten tool-call argument coercion comments

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-07-06 10:12:22 -07:00
..
data_recipe Studio: harden background consumer loops and streaming paths against silent UI freezes (#6653) 2026-06-26 03:31:33 -07:00
export Studio: multi-select export formats, portable FP8/INT8, GGUF LoRA, and source parity (#6767) 2026-07-03 08:25:10 -07:00
inference Studio: coerce tool_call arguments to dict before chat templating (fixes MLX tool follow-up error) (#6807) 2026-07-06 10:12:22 -07:00
rag Studio: customizable RAG embedding model with HF search, settings tab reorganization (#6800) 2026-07-02 05:26:33 -07:00
training Add MLX-aware public Unsloth trainer API (#6462) 2026-07-02 23:02:26 +01:00
__init__.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
_torchao_stub.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
import_guards.py Studio: self-heal unsloth namespace shadows; clearer failed-load messages (#6532) 2026-06-21 22:43:31 -07:00
tool_healing.py studio: tool calling + healing parity for Llama-3, Mistral, Gemma 4 on safetensors + MLX (#5620) 2026-07-06 10:06:06 -07:00