* don't re-prompt finished answers in the tool loop * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * keep a separate post-tool reprompt budget and tighten the intent regexes * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Reset the repeat guard after a tool runs and suppress 'I should call ...' forced stalls * Cover 'must' in forced-retry suppression, keep appended answers, and count RAG autoinject as a prior tool run * Anchor obligation suppression to sentence starts and wire the repeat guard into the safetensors loop * Keep deletions out of restatement and nudge pronoun-free first-step plans * Tighten repeat similarity, anchor subjectless plans, and restore first-step plan forms * Keep first-person plan framing and punctuation-bearing terms out of repeat detection * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Keep leading term punctuation, accept colon-delimited first steps, and drop invoke/query from suppression * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Tighten comments on the plan-without-action re-prompt guards * Compare plans by token sequence, suppress subjectless modals, and accept dash-delimited first steps * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * studio: narrow the first-step plan match and make repeat detection content-based Restrict the bare "First, <word>" intent alternative to a pronoun, an explicit plan, or an investigative verb, so ordinal prose ("First place went to Alice") and user-facing advice ("First, install the package") no longer count as a plan without action. Keep punctuation-only tokens in the repeat comparison, so "the value is 5" and "the value is < 5" stay distinct, and compare content-word sequences instead of a similarity ratio: any ratio is length-dependent, so one corrected token in a 54-token plan still scored 0.98 and cost the model its remaining nudge. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * studio: tighten comments in the plan-without-action re-prompt path * studio: keep a forced retry that pivots from a plan to an answer The obligation-plan branch discarded the whole turn, so a retry such as "I should call web_search, but the answer is Tokyo." reached the user as nothing at all. Suppress the plan only when nothing follows it: a pivot after the match keeps the output, and _FINAL_ANSWER_SIGNAL now recognises "the answer is" and "to summarise" alongside "answer:". Leaking a plan sentence is cosmetic, dropping an answer is not, so the doubtful case now resolves towards shipping the turn. * studio: keep articles in repeat comparison and exclude missing-answer phrasing Articles are not filler: dropping them made "search for The Who" and "search for Who" compare equal, so a corrected target ended the nudge. _FINAL_ANSWER_SIGNAL matched "the answer is not in the provided context", which announces a missing answer, so the plan behind it shipped as the final response instead of being suppressed. Negated forms are now excluded. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * studio: tighten the pivot and final-answer signals, drop filler-insensitive repeats The purpose clause in "call web_search to summarize the results" matched the final-answer signal, so the plan shipped instead of being suppressed; that alternative is gone. A pivot word now has to carry text of its own, since "I should call web_search, though." answers nothing. Repeat detection no longer ignores filler words. No word is reliably filler: dropping them to absorb rewording also absorbed the target ("OK Go" became "Go"). A missed repeat costs one nudge out of the cap; a false one strands the plan unexecuted. * studio: exempt offers of help, and add a measured accuracy floor Offering to help hands control back exactly like the existing "let me know" exemption. On a corpus of real model turns, "I'll do my best to help" and "allow me to assist" close a clarification request and never precede a tool call, but they were read as intent and re-prompted. "help you" keeps its plan reading when an action verb follows it. The new test scores the classifier against 300 turns captured from three local GGUF models, each one a finished answer: the turn called no tool, and three regenerations behind the production nudge produced no tool call either. Over those turns, wasted nudges go from 36 (12.0%) on main to 5 (1.7%), and retries whose text would be discarded from 60 (20.2%) to 1 (0.3%). Until now these patterns were tuned on hand-written example sentences, which cannot show how often the classifier is right on real output. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com>
96 lines
3.9 KiB
Python
96 lines
3.9 KiB
Python
# SPDX-License-Identifier: AGPL-3.0-only
|
|
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
|
|
|
"""An accuracy floor for the plan-without-action classifier, on real model output.
|
|
|
|
The rest of the tool-loop suites pin behaviour on hand-written example sentences,
|
|
which is how the patterns here were tuned. That says nothing about how often the
|
|
classifier is right on what models actually emit, so this file scores it against a
|
|
corpus captured from local models (``tests/data/plan_vs_answer.jsonl``).
|
|
|
|
How the corpus was built: three GGUF models (Qwen3-0.6B, Qwen3-1.7B,
|
|
Llama-3.2-1B-Instruct) were driven through llama-server with the real Studio tool
|
|
schemas over prompts spanning tool-requiring questions, questions needing no tool,
|
|
list-formatted answers, ambiguous requests, non-English, and follow-ups issued after
|
|
a tool had already run. Turns cut off by the token cap were dropped, since a
|
|
truncation is not a stall.
|
|
|
|
Every turn here is a *finished answer*: the turn called no tool, and when the
|
|
production nudge was appended and the turn regenerated three times, not one retry
|
|
produced a tool call. A forceful re-prompt could not extract an action, so there was
|
|
no action left to take. Nudging these is wasted work, and in the GGUF loop the
|
|
retry's text can then be discarded, which costs the user a visible answer.
|
|
|
|
Measured when this landed, over the 300 turns:
|
|
|
|
tree nudged retry discarded
|
|
origin/main (pre-PR) 36 (12.0%) 60 (20.2%)
|
|
this PR 5 ( 1.7%) 1 ( 0.3%)
|
|
|
|
The budgets below sit above the measured counts so that innocuous wording changes
|
|
do not fail the build, and far below the pre-PR counts so a real regression does.
|
|
A failure prints the offending turns: fix the pattern, or if the turn really is a
|
|
stall, correct its label here.
|
|
"""
|
|
|
|
import json
|
|
from pathlib import Path
|
|
|
|
from core.inference.llama_cpp import _should_suppress_forced_no_tool_output
|
|
from core.inference.tool_call_parser import is_short_intent_without_action
|
|
|
|
DATA = Path(__file__).parent / "data" / "plan_vs_answer.jsonl"
|
|
|
|
# Measured 5 of 300; pre-PR was 36.
|
|
NUDGE_BUDGET = 9
|
|
# Measured 1 of 300; pre-PR was 60. Tighter, because this one destroys output.
|
|
DISCARD_BUDGET = 4
|
|
|
|
|
|
def _corpus():
|
|
with open(DATA, encoding = "utf-8") as fh:
|
|
return [json.loads(line) for line in fh if line.strip()]
|
|
|
|
|
|
def _report(rows, limit = 10):
|
|
lines = []
|
|
for row in rows[:limit]:
|
|
text = " ".join(row["text"].split())
|
|
lines.append(
|
|
f" [{row['model']}/{row['prompt_class']}] {row['prompt']!r}\n {text[:200]!r}"
|
|
)
|
|
if len(rows) > limit:
|
|
lines.append(f" ... and {len(rows) - limit} more")
|
|
return "\n".join(lines)
|
|
|
|
|
|
def test_corpus_is_intact():
|
|
"""Guards the budgets: they mean nothing if the corpus silently shrinks."""
|
|
corpus = _corpus()
|
|
assert len(corpus) == 300
|
|
assert all(row["text"].strip() for row in corpus)
|
|
# Every row is a finished answer by construction.
|
|
assert all(row["retry_tool_calls"] == 0 for row in corpus)
|
|
|
|
|
|
def test_finished_answers_are_rarely_nudged():
|
|
"""A finished answer costs a whole extra generation when it is nudged."""
|
|
nudged = [row for row in _corpus() if is_short_intent_without_action(row["text"])]
|
|
assert len(nudged) <= NUDGE_BUDGET, (
|
|
f"{len(nudged)}/300 finished answers classified as plans "
|
|
f"(budget {NUDGE_BUDGET}):\n{_report(nudged)}"
|
|
)
|
|
|
|
|
|
def test_finished_answers_are_not_discarded():
|
|
"""The retry's text is all the user gets, so discarding it is the worst case."""
|
|
discarded = [
|
|
row
|
|
for row in _corpus()
|
|
if row["retry_text"].strip()
|
|
and _should_suppress_forced_no_tool_output(row["retry_text"], row["text"])
|
|
]
|
|
assert len(discarded) <= DISCARD_BUDGET, (
|
|
f"{len(discarded)}/300 finished retries would be discarded "
|
|
f"(budget {DISCARD_BUDGET}):\n{_report(discarded)}"
|
|
)
|