unsloth/studio/backend/tests/test_plan_classifier_accuracy.py
Nilay 22493242a3
Studio: Don't re-prompt finished answers in the tool loop (#7505)
* don't re-prompt finished answers in the tool loop

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* keep a separate post-tool reprompt budget and tighten the intent regexes

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Reset the repeat guard after a tool runs and suppress 'I should call ...' forced stalls

* Cover 'must' in forced-retry suppression, keep appended answers, and count RAG autoinject as a prior tool run

* Anchor obligation suppression to sentence starts and wire the repeat guard into the safetensors loop

* Keep deletions out of restatement and nudge pronoun-free first-step plans

* Tighten repeat similarity, anchor subjectless plans, and restore first-step plan forms

* Keep first-person plan framing and punctuation-bearing terms out of repeat detection

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Keep leading term punctuation, accept colon-delimited first steps, and drop invoke/query from suppression

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Tighten comments on the plan-without-action re-prompt guards

* Compare plans by token sequence, suppress subjectless modals, and accept dash-delimited first steps

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: narrow the first-step plan match and make repeat detection content-based

Restrict the bare "First, <word>" intent alternative to a pronoun, an explicit
plan, or an investigative verb, so ordinal prose ("First place went to Alice")
and user-facing advice ("First, install the package") no longer count as a plan
without action.

Keep punctuation-only tokens in the repeat comparison, so "the value is 5" and
"the value is < 5" stay distinct, and compare content-word sequences instead of
a similarity ratio: any ratio is length-dependent, so one corrected token in a
54-token plan still scored 0.98 and cost the model its remaining nudge.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: tighten comments in the plan-without-action re-prompt path

* studio: keep a forced retry that pivots from a plan to an answer

The obligation-plan branch discarded the whole turn, so a retry such as
"I should call web_search, but the answer is Tokyo." reached the user as
nothing at all. Suppress the plan only when nothing follows it: a pivot
after the match keeps the output, and _FINAL_ANSWER_SIGNAL now recognises
"the answer is" and "to summarise" alongside "answer:".

Leaking a plan sentence is cosmetic, dropping an answer is not, so the
doubtful case now resolves towards shipping the turn.

* studio: keep articles in repeat comparison and exclude missing-answer phrasing

Articles are not filler: dropping them made "search for The Who" and
"search for Who" compare equal, so a corrected target ended the nudge.

_FINAL_ANSWER_SIGNAL matched "the answer is not in the provided context",
which announces a missing answer, so the plan behind it shipped as the final
response instead of being suppressed. Negated forms are now excluded.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: tighten the pivot and final-answer signals, drop filler-insensitive repeats

The purpose clause in "call web_search to summarize the results" matched the
final-answer signal, so the plan shipped instead of being suppressed; that
alternative is gone. A pivot word now has to carry text of its own, since
"I should call web_search, though." answers nothing.

Repeat detection no longer ignores filler words. No word is reliably filler:
dropping them to absorb rewording also absorbed the target ("OK Go" became
"Go"). A missed repeat costs one nudge out of the cap; a false one strands the
plan unexecuted.

* studio: exempt offers of help, and add a measured accuracy floor

Offering to help hands control back exactly like the existing "let me know"
exemption. On a corpus of real model turns, "I'll do my best to help" and
"allow me to assist" close a clarification request and never precede a tool
call, but they were read as intent and re-prompted. "help you" keeps its plan
reading when an action verb follows it.

The new test scores the classifier against 300 turns captured from three local
GGUF models, each one a finished answer: the turn called no tool, and three
regenerations behind the production nudge produced no tool call either. Over
those turns, wasted nudges go from 36 (12.0%) on main to 5 (1.7%), and retries
whose text would be discarded from 60 (20.2%) to 1 (0.3%).

Until now these patterns were tuned on hand-written example sentences, which
cannot show how often the classifier is right on real output.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2026-07-29 02:38:17 -07:00

96 lines
3.9 KiB
Python

# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""An accuracy floor for the plan-without-action classifier, on real model output.
The rest of the tool-loop suites pin behaviour on hand-written example sentences,
which is how the patterns here were tuned. That says nothing about how often the
classifier is right on what models actually emit, so this file scores it against a
corpus captured from local models (``tests/data/plan_vs_answer.jsonl``).
How the corpus was built: three GGUF models (Qwen3-0.6B, Qwen3-1.7B,
Llama-3.2-1B-Instruct) were driven through llama-server with the real Studio tool
schemas over prompts spanning tool-requiring questions, questions needing no tool,
list-formatted answers, ambiguous requests, non-English, and follow-ups issued after
a tool had already run. Turns cut off by the token cap were dropped, since a
truncation is not a stall.
Every turn here is a *finished answer*: the turn called no tool, and when the
production nudge was appended and the turn regenerated three times, not one retry
produced a tool call. A forceful re-prompt could not extract an action, so there was
no action left to take. Nudging these is wasted work, and in the GGUF loop the
retry's text can then be discarded, which costs the user a visible answer.
Measured when this landed, over the 300 turns:
tree nudged retry discarded
origin/main (pre-PR) 36 (12.0%) 60 (20.2%)
this PR 5 ( 1.7%) 1 ( 0.3%)
The budgets below sit above the measured counts so that innocuous wording changes
do not fail the build, and far below the pre-PR counts so a real regression does.
A failure prints the offending turns: fix the pattern, or if the turn really is a
stall, correct its label here.
"""
import json
from pathlib import Path
from core.inference.llama_cpp import _should_suppress_forced_no_tool_output
from core.inference.tool_call_parser import is_short_intent_without_action
DATA = Path(__file__).parent / "data" / "plan_vs_answer.jsonl"
# Measured 5 of 300; pre-PR was 36.
NUDGE_BUDGET = 9
# Measured 1 of 300; pre-PR was 60. Tighter, because this one destroys output.
DISCARD_BUDGET = 4
def _corpus():
with open(DATA, encoding = "utf-8") as fh:
return [json.loads(line) for line in fh if line.strip()]
def _report(rows, limit = 10):
lines = []
for row in rows[:limit]:
text = " ".join(row["text"].split())
lines.append(
f" [{row['model']}/{row['prompt_class']}] {row['prompt']!r}\n {text[:200]!r}"
)
if len(rows) > limit:
lines.append(f" ... and {len(rows) - limit} more")
return "\n".join(lines)
def test_corpus_is_intact():
"""Guards the budgets: they mean nothing if the corpus silently shrinks."""
corpus = _corpus()
assert len(corpus) == 300
assert all(row["text"].strip() for row in corpus)
# Every row is a finished answer by construction.
assert all(row["retry_tool_calls"] == 0 for row in corpus)
def test_finished_answers_are_rarely_nudged():
"""A finished answer costs a whole extra generation when it is nudged."""
nudged = [row for row in _corpus() if is_short_intent_without_action(row["text"])]
assert len(nudged) <= NUDGE_BUDGET, (
f"{len(nudged)}/300 finished answers classified as plans "
f"(budget {NUDGE_BUDGET}):\n{_report(nudged)}"
)
def test_finished_answers_are_not_discarded():
"""The retry's text is all the user gets, so discarding it is the worst case."""
discarded = [
row
for row in _corpus()
if row["retry_text"].strip()
and _should_suppress_forced_no_tool_output(row["retry_text"], row["text"])
]
assert len(discarded) <= DISCARD_BUDGET, (
f"{len(discarded)}/300 finished retries would be discarded "
f"(budget {DISCARD_BUDGET}):\n{_report(discarded)}"
)