Commit graph

765 commits

Author SHA1 Message Date
pre-commit-ci[bot]
9f3c49a733 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 22:02:12 +00:00
Daniel Han
32b64aef97 Studio: r22 fixes - prose backtick guard, take/follow steps verbs
- _has_unclosed_code_fence() ignores a fence run when the trailing
  text on the same line starts with a space (typical English prose
  like "Use \`\`\` to start a markdown fence."). Real fence openers
  either end the line right after the delimiters or carry an info
  string with no leading space (\`\`\`python, \`\`\`bash-session).

- _DIRECT_NUMBERED_PLAN_FRAMING accepts "take these steps",
  "follow these steps", and "perform these actions" as first-person
  intent verbs. Plans like "I'll take these steps:\n1. Open URL\n
  2. Read" still re-prompt instead of being read as final answers.
2026-05-24 22:01:57 +00:00
pre-commit-ci[bot]
394cdf48ea [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 21:51:15 +00:00
Daniel Han
834b34c68d Studio: r21 fixes - strip orphan tool-call XML before artifact, broaden first-person plan verbs
- Re-prompt path calls _strip_tool_markup(final=True) on content_accum
  before measuring intent / artifact / length. An orphan
  ``<tool_call>...</tool_call>`` block containing a code fence no
  longer hides the intent-only visible answer from the artifact check.

- _DIRECT_NUMBERED_PLAN_FRAMING splits into two branches:
  * First-person intent ("I'll", "Let me", "I will", etc.) accepts a
    broader work-verb set (open, read, search, check, review, inspect,
    examine, etc.). Direct first-person announcements are strong
    plan-like signals.
  * Bare "First, ..." / "Step N: ..." keeps the narrow verb set so
    algorithmic answers ("First, use binary search:") stay valid.
  Catches stalls like "I will check the docs:\n1. Gather..." and
  "Let me read the uploaded file:\n1. Identify the columns..." that
  previously slipped past the freshness-gated lookup verbs.
2026-05-24 21:50:39 +00:00
pre-commit-ci[bot]
64a400be2d [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 21:35:08 +00:00
Daniel Han
6ea922a58d Studio: r20 fixes - skip markup-count when real artifact exists
When a real complete artifact is already in the response, prose
mentions of bare <html> / <svg> tags in explanatory text are common
(for example "Use the <html> tag for the root"). The unbalanced-
open/close count would falsely classify the response as mid-stream
and wipe the valid answer. The artifact-counting cross-check now
only runs when NO real artifact has been emitted yet; once a real
artifact exists, mid-stream second markup is rare enough that the
count-based detector is not worth the false-positive cost.

This also unblocks complete <html> answers that nest <svg> children
or contain JS string literals like "<svg width=10>", since those
unmatched markup tokens were being flagged as unclosed.
2026-05-24 21:34:52 +00:00
pre-commit-ci[bot]
ae8f70f05e [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 21:22:40 +00:00
Daniel Han
2ea2f3519a Studio: r19 fixes - cross-strip closed artifacts, iterate skeleton matches, bare-intent colon plan
- _has_answer_artifact() now strips closed code fences before checking
  for unclosed markup, and strips closed markup before checking for
  unclosed code fences. A Python / JS snippet containing literal
  "<html>" / "<svg>" strings no longer trips the unclosed-markup
  cross-check, and complete HTML containing a JS string with literal
  backticks no longer trips the unclosed-fence cross-check.

- _looks_like_real_artifact() iterates every artifact match. An empty
  <html></html> / <svg></svg> skeleton followed by a real complete
  page no longer hides the real artifact.

- _is_empty_markup_skeleton() strips an optional <!doctype ...> prefix
  before testing the empty-skeleton pattern, so
  "<!doctype html><html></html>" plan-only mentions also re-prompt.

- _BARE_INTENT_NUMBERED_PLAN catches the tight "I'll:\n1. Open ..." /
  "Let me:\n1. Parse ..." shape where bare first-person intent +
  colon + newline is immediately followed by numbered action items.
  No work verb is required between the intent and the list.
2026-05-24 21:22:26 +00:00
pre-commit-ci[bot]
7c08ade3a4 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 21:11:13 +00:00
Daniel Han
4809503cce Studio: r18 fixes - unbalanced markup detection, empty-skeleton reject, list-item verb cross-check
- _has_unclosed_markup_block() now compares open / close tag counts.
  A response with one closed <html> followed by a second still-open
  <html> (multi-page mid-stream) or <svg></svg><svg> is unbalanced,
  so the artifact path returns False and the re-prompt fires. The
  helper now runs BEFORE _HAS_ANSWER_ARTIFACT so an earlier complete
  artifact cannot mask a later open block.

- _looks_like_real_artifact() rejects empty <html></html> /
  <svg></svg> skeletons. Plan-only mentions ("First, I'll create an
  <html></html> skeleton, then add CSS.") no longer suppress the
  re-prompt.

- _NUMBERED_ACTION_ITEM + _STRONG_INTENT_BEFORE_LIST catches plans
  where the work verbs sit in the list ITEMS rather than before the
  list (e.g. "First, I'll:\n1. Load the CSV.\n2. Compute total").
  The verb whitelist is intentionally narrow (load, parse, calculate,
  compute, analyze, run, execute, fetch, download, query, inspect,
  extract) so ordinary algorithm answers ("First, use binary search:
  1. Search the left half") stay valid. The intent gate excludes
  bare "First" / "Step N:" for the same reason - direct first-person
  pronoun is required.
2026-05-24 21:10:59 +00:00
pre-commit-ci[bot]
aba4824de2 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 20:59:40 +00:00
Daniel Han
e078605936 Studio: r17 fixes - unclosed-markup cross-check, compare/review lookup verbs, defer artifact scan
- _has_unclosed_markup_block() short-circuits the numbered-list fallback
  when the response contains an open <html> or <svg> with no matching
  close. A partial markup body that happens to contain two numbered
  lines no longer reads as a final answer.

- _TOOL_ACTION_VERBS adds freshness-gated "compare" and "review" so
  plans phrased as "Compare the latest release sources" or "Review
  the current documentation" still re-prompt.

- Re-prompt call site defers the visible-artifact regex scan until
  the cheap gates (tools enabled, _reprompt_count, length window,
  intent regex) have all passed. Long final answers that can never
  re-prompt no longer pay the artifact-scan cost.
2026-05-24 20:59:26 +00:00
pre-commit-ci[bot]
2d4d7136ed [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 20:47:47 +00:00
Daniel Han
7b0ed8333a Studio: r16 fixes - inline fence tracking, broaden direct-intent plan with first/step prefixes
- _has_unclosed_code_fence() now scans every line with re.search and a
  shared FENCE_RUN regex, so an inline opening fence such as
  "First, let me write it. \`\`\`python" is tracked alongside the
  column-0 openers. A numbered list emitted INSIDE an inline-open
  fence no longer reads as a final answer.

- _DIRECT_NUMBERED_PLAN_FRAMING adds "first" and "step N(:?)" to its
  intent prefixes and "look up" to its verb whitelist. Plans like
  "First, analyze the uploaded CSV:\n1. Load rows\n2. Compute total"
  or "I'll look that up:\n1. Search the docs" now re-prompt instead
  of being mis-classified as final answers. The verb whitelist still
  excludes bare search/find/check/verify so "First, use binary
  search:\n1. Search the left half" stays an answer.
2026-05-24 20:47:34 +00:00
Daniel Han
9fa736bef1 Studio: r15 fixes - direct-intent numbered plan stalls, unclosed-fence short-circuit
- _DIRECT_NUMBERED_PLAN_FRAMING matches first-person intent ("I'll",
  "Let me", etc.) plus a narrow follow-up verb ("do this", "do these",
  "create", "build", "set up", "calculate", "parse", "run", etc.)
  followed by a numbered list. This catches stalls like "First, I'll
  do this:\n1. Search for X." or "Let me do this:\n1. Parse the
  JSON.\n2. Calculate the average." where the model announces actions
  but never invokes a tool. The verb whitelist stays narrow so
  "Let me explain" / "Let me show" / "Let me draft a poem" answers
  are NOT misclassified.

- _has_answer_artifact() now checks for an unclosed code fence BEFORE
  consulting _HAS_ANSWER_ARTIFACT. A response with one complete fence
  followed by a second, still-open fence (mid-stream multi-file
  answers) no longer suppresses the re-prompt; the unclosed second
  fence wins.
2026-05-24 20:33:15 +00:00
Daniel Han
4652a4b03c Studio: r14 fixes - longer CommonMark closing fence, explicit-plan header standalone, use-python tool wording
- _HAS_ANSWER_ARTIFACT closing fence now accepts strictly more delimiters
  than the opener (CommonMark rule). The opener stays anchored on both
  sides so a 4-open / 3-close payload still does not match, but a
  legitimate 3-open / 4-close (and 3-tilde / 4-tilde) answer is now
  recognised as a completed artifact.

- _EXPLICIT_PLAN_HEADER triggers the plan classification by itself when
  the response contains \"Here's my plan\" / \"Here's my approach\" /
  \"Here's the plan\". Numbered stalls like \"Here's my plan:\n1. Analyze\n
  2. Draft\" re-prompt again without needing a freshness-gated verb.
  Plain \"Plan:\" / \"My weekly plan:\" stay valid answers because they
  lack the possessive first-person header.

- _TOOL_ACTION_VERBS adds \"use python (tool) to ...\", \"use the python
  tool\", \"invoke the python tool\", and \"use the search tool\" so
  numbered plans that route through these phrasings still re-prompt.
2026-05-24 20:20:38 +00:00
Daniel Han
b7c7427eb8 Studio: add ReDoS regression test for full-window plan-framing scan 2026-05-24 16:17:24 +00:00
pre-commit-ci[bot]
cff8a6a1bc [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 16:00:05 +00:00
Daniel Han
e16f898c26 Studio: r13 fixes - freshness-gated lookup verbs, anchored CommonMark fences, open-fence cross-check
- _TOOL_ACTION_VERBS gates the lookup verbs (search / look up /
  browse / google / fetch / research / investigate / find / check /
  verify) on a freshness or web/internet/online target. Plain answer
  prose like \"binary search: 1. Search the left half\" or \"1. Find
  the bug\" stays a valid answer, while \"1. Search the web for X\"
  / \"1. Google the current chart\" / \"1. Research the latest docs\"
  still re-prompts. Strong unambiguous patterns (web search, query
  the web, call a tool, run python) remain bare.

- _HAS_ANSWER_ARTIFACT anchors the fence opener and closer with
  (?<!\\`) / (?!\\`) lookarounds so a 4-backtick opener cannot
  backtrack to a 3-backtick fence and treat the surplus delimiter as
  info-string text. Same rule for tildes.

- _has_answer_artifact now consults a small _has_unclosed_code_fence
  helper before the numbered-list fallback. A numbered list embedded
  INSIDE an open fence no longer masquerades as a final answer.

- Existing plan-framing tests updated to use freshness-gated lookup
  phrasing so they continue to assert the intended invariants.
2026-05-24 15:58:19 +00:00
pre-commit-ci[bot]
693bd89a79 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 15:36:23 +00:00
Daniel Han
ef3ee3d8bd Studio: r12 fixes - CommonMark 3+ fences, query/consult synonyms, full-window plan scan, visible-reasoning artifact
- _HAS_ANSWER_ARTIFACT now matches fences with three OR MORE backticks
  / tildes using a named-group backreference (CommonMark rule). Models
  routinely emit \`\`\`\` / \`\`\`\`\` when the body itself contains a triple
  fence. The previous regex only matched exactly three.

- _TOOL_ACTION_VERBS adds \"query / consult the web / internet / online
  sources\" so numbered plan stalls phrased with these synonyms still
  re-prompt instead of being read as final answers.

- _PLAN_LIST_FRAMING widens the intent-to-action scan from 80 chars to
  the full short candidate (caller already gates at _REPROMPT_MAX_CHARS
  = 2000). Realistic plans where item 1 is preamble and item 2 is the
  explicit tool action no longer slip through.

- Re-prompt call site separates VISIBLE-content artifact check from
  hidden reasoning. When content_accum is empty AND has_content_tokens
  is False, reasoning_accum is the user-visible text and counts for
  the artifact check. Otherwise reasoning stays hidden and an artifact
  inside it must not suppress the re-prompt.
2026-05-24 15:36:00 +00:00
pre-commit-ci[bot]
562af754c8 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 15:22:59 +00:00
Daniel Han
197b7755ea Merge remote-tracking branch 'origin/main' into fix-tool-reprompt-overfire 2026-05-24 15:22:38 +00:00
Daniel Han
9e79a3e2c6 Studio: harden re-prompt guard - visible-only artifact check, closing-fence end-of-line, freshness-gated find/check/verify
- Re-prompt path now treats a closed artifact in hidden reasoning as
  no artifact for the user; only visible content_accum counts. Stops
  hidden chain-of-thought from suppressing the tool-forcing nudge
  when content_accum is empty.

- Closed backtick / tilde fences must end the line cleanly. Trailing
  prose after the closing fence (```not actually closed) no longer
  reads as a complete artifact.

- _TOOL_ACTION_VERBS admits find / check / verify only when paired
  with a freshness signal (current / latest / today / up-to-date /
  live / online / web). Numbered plan stalls like \"1. Find the
  current Billboard chart\" re-prompt again, while \"1. Find the
  bug\" / \"2. Check the answer\" stay valid answer text.
2026-05-24 15:22:21 +00:00
pre-commit-ci[bot]
94e7e12aaf [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 15:08:32 +00:00
Daniel Han
d38bb3f077 Studio: revert Plan: / "Here is the plan" intent and narrow plan verbs
Reviewer round 9 (5 of 10 reviewers) flagged that the new bare
``Plan:`` / ``Approach:`` / ``Here is the plan`` intent branches
reintroduced the original "wipe a complete answer" failure for
realistic final answers whose topic happens to contain a tool-action
word. Triggers for prompts like "Create a lesson plan for teaching
search skills" when the model answers:

  Plan:
  1. Search skills: students learn query keywords.
  2. Source evaluation: compare domains.
  3. Reflection: write what worked.

``_INTENT_SIGNAL`` matched the new ``Plan:`` lookahead because
``search`` appears within 120 chars, then ``_PLAN_LIST_FRAMING``
disqualified the numbered list, and the synthetic STOP turn wiped a
valid answer.

Revert the additions in ``_INTENT_SIGNAL``:
  * Drop ``Plan:`` / ``Approach:`` (newline + action-verb lookahead).
  * Drop ``Here is the plan`` / ``Here are my steps`` (action-verb
    lookahead).

Plan stalls phrased with explicit first-person intent ("I'll search...",
"First, I'll fetch...", "Let me look up...") are still caught by the
existing intent patterns and ``_PLAN_LIST_FRAMING``.

Also narrow the plan-list action-verb whitelist to tool-specific verbs
(``search`` / ``look up`` / ``fetch`` / ``browse`` / ``web search`` /
``call (a) tool`` / ``run python`` / ``execute python``). Broad verbs
like ``use`` / ``compare`` / ``check`` / ``find`` / ``think`` /
``respond`` / ``answer`` / ``analyse`` / ``explore`` / ``outline`` /
``reason`` are removed because real answer lists use them ("1. Use
BFS", "1. Compare versions").

Finally, fix the test module's ``loggers`` / ``structlog`` stub
injection to only fire when the real module is missing AND to set
``__path__ = []`` on the stub. Previously the bare ``ModuleType`` could
poison ``sys.modules`` for any later test that imports a real
submodule (``from loggers.handlers import ...``).

Net behavioural change vs the previous commit: stricter on what
counts as a plan stall, never wipes a final answer titled
``Plan:`` / ``My plan:`` / ``Here is the plan you asked for``.
2026-05-24 15:08:18 +00:00
Daniel Han
a6f6022bd0 Studio: require nearby action verb for Plan: / "Here is the plan" intents
Reviewer round 8 surfaced a real false positive in the previous commit:
a final answer naturally titled "Plan:" / "My plan:" / "Approach:" with
numbered content items now slipped through _INTENT_SIGNAL and got
wiped by the synthetic STOP turn. Examples:

  Plan:
  1. Warm-up: Students review fractions.
  2. Group practice.
  3. Assessment.

  My plan:
  1. Breakfast: oatmeal and fruit.
  2. Lunch: rice bowl.
  3. Dinner: lentil soup.

  Here is the plan you asked for. It is two pages long.

Add a lookahead requiring one of the conservative re-prompt action
verbs (search / fetch / verify / look up / call / compare / think /
respond / etc.) to appear within 120 chars after the "Plan:" /
"Approach:" / "Here is the plan" / "Here are my steps" marker. Plan
stalls whose items are tool actions ("Plan:\n1. search the docs\n2.
summarise the result") still match and re-prompt; prose plans whose
items are content do not.

Also mirror "first" in _PLAN_LIST_FRAMING so numbered action plans
that start with "First" stay disqualified even after the helper enters
the numbered-list branch.

Factor the action-verb set out as _REPROMPT_ACTION_VERBS so both
regexes share one source of truth.

Six new regression samples: three lesson / meal / weather plans that
must NOT wipe, three action-plan headers that must re-prompt, three
prose "Here is the plan" answers that must not wipe.
2026-05-24 14:56:04 +00:00
Daniel Han
3dc26e7acf Studio: require newline after Plan: / Approach: header so inline product text does not re-prompt
After narrowing the colon marker to lines starting with a generic
determiner ("My plan:" / "The approach:" / ...), inline product or
pricing answers like "Your current Plan: Pro includes local chats",
"The plan: Basic is free, Pro is $10/month", or
"My plan: use dynamic programming" still slipped into the re-prompt
path and could wipe a valid answer.

Add a lookahead requiring a newline (with optional trailing horizontal
whitespace) after the colon, so only header-style framings like
"Plan:\n1. search\n2. summarise" or "My approach:\n1. fetch" count.
Inline "Plan: <text>" is now treated as ordinary prose.

Add eight regression samples (lesson plan, meal plan, marketing plan,
pricing plan, recommended approach, migration plan, dynamic-programming
plan, currently active plan) all of which previously re-prompted under
the unanchored matcher and now correctly do not.
2026-05-24 14:41:44 +00:00
Daniel Han
64ae2ac4c5 Studio: sync re-prompt guard intent forms and tighten Plan: anchor
Two more gaps surfaced by another reviewer sweep on the previous commit:

1. _PLAN_LIST_FRAMING was missing several intent forms that
   _INTENT_SIGNAL accepts, so numbered tool-action plans phrased with
   "Allow me", "I'm going to", "I'm gonna", "I am gonna", or "I shall"
   were silently classified as completed answers and skipped the
   tool-call re-prompt. Mirror the full intent set from _INTENT_SIGNAL
   so the two regexes stay in lock-step.

2. Bare \b(?:plan|approach): in _INTENT_SIGNAL / _PLAN_LIST_FRAMING
   matched any in-text occurrence of "plan:" / "approach:", including
   "lesson plan:" / "meal plan:" / "migration plan:". A direct answer
   like "Here is a lesson plan:\n1. Warm-up\n2. Group practice" would
   trip _INTENT_SIGNAL and risk wiping the response. Anchor the colon
   marker to start of line and only allow generic determiners (my, the,
   our, a, this, that) between the line start and the keyword.

3. Add "Here is the plan" / "Here are my steps" to both _INTENT_SIGNAL
   and _PLAN_LIST_FRAMING so non-apostrophe phrasings of the same
   framing pattern are caught.

Added regression tests covering every intent form against a numbered
action plan, and a line-anchor test that distinguishes generic plan
framings ("My plan:", "The approach:") from content noun phrases
("lesson plan:", "meal plan:").
2026-05-24 14:36:04 +00:00
Daniel Han
a2ab9895b1 Studio: require apostrophe in _PLAN_LIST_FRAMING i'll alternative
The follow-up commit used ``i['’]?ll`` (apostrophe optional) in
``_PLAN_LIST_FRAMING``. With the apostrophe optional the alternative
also matches the word "ill" (sick), so a response like
"She is ill. Here is the list:\n1. ...\n2. ..." plus an unrelated
action verb within 80 chars was misclassified as a plan and re-prompted.

Make the apostrophe required (``i['’]ll``) to mirror the original
_INTENT_SIGNAL definition. Add a regression test that pins the
distinction: "ill" as adjective does not trigger plan framing, but
"I'll" / "I will" do.
2026-05-24 14:29:46 +00:00
Daniel Han
4165878734 Studio: extend re-prompt guard for tilde fences, Plan: intent, and contemplative verbs
Three follow-up gaps surfaced by another reviewer sweep on the previous commit:

1. Tilde-fenced code (~~~lang ... ~~~) was not detected. CommonMark allows
   it and several models emit it when the body itself contains backticks.
   Add a tilde alternative to _HAS_ANSWER_ARTIFACT mirroring the backtick
   form (any info string, optional indent on close, length-bounded body).

2. Bare "Plan:" / "Approach:" lines did not match _INTENT_SIGNAL, so a
   "Plan:\n1. search\n2. summarise" stall slipped past the entry gate
   entirely. Add the colon form to the step / plan framing alternative.

3. The plan-framing verb whitelist missed common contemplative verbs
   (think / respond / answer / analy[sz]e / explore / outline / gather /
   query / reason) so plan stalls phrased without explicit "Here's my
   plan" framing were misclassified as completed answers. Keep the
   whitelist conservative: write / create / make / build / read / list /
   try are intentionally out because real answer lists use them
   ("1. Write a poem", "1. Read War and Peace").

Added regression tests for each fix plus an extra ReDoS budget test for
the doctype/<html alternation worst case (about 7 ms today; assert < 50
ms so a future quantifier change that drops the inner {0,4000} bound
fails loudly).
2026-05-24 14:25:35 +00:00
pre-commit-ci[bot]
c2c448684c [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-24 14:06:25 +00:00
Daniel Han
69ef56edb3 Studio: tighten re-prompt artifact guard for non-alpha fences, planning lists, incomplete HTML
Addresses three follow-ups flagged on the first cut of this PR by static
reviewers and parallel reviewer runs:

1. Numbered plan-only stalls were treated as completed answers. A
   response like `Here's my plan:\n1. Search the web\n2. Summarise`
   matched both `_INTENT_SIGNAL` and the numbered-list branch of
   `_HAS_ANSWER_ARTIFACT`, so the tool-forcing re-prompt was skipped.
   That contradicted the PR's stated invariant that plan-only stalls
   still re-prompt. The list now has to be paired with no plan framing
   (no `Here's my plan` / `plan:` / `approach:`, no intent phrase
   followed by a tool-action verb) to count as an artifact.

2. Closed code fences with non-alpha info strings (`python3`, `c++`,
   `c#`, `objective-c`, `ts-node`, `bash-session`, `python linenums="1"`)
   were not recognised by the `[a-zA-Z]*` info-string class. Complete
   answers in those languages still re-prompted and could be wiped.
   The info-string class is now `[^\r\n]{0,200}` and the closing fence
   may be indented.

3. Bare `<!doctype` or `<html` text was treated as an artifact. A
   plan-only response that mentions `<html>` in prose now no longer
   bypasses the re-prompt; the HTML branch requires a closing
   `</html>` (doctype prefix optional).

All `[\s\S]{...}?` runs are length-bounded so ReDoS-style adversarial
input stays linear. ReDoS guard tests cover CRLF spam and repeated
`<html ` openings without close.
2026-05-24 14:04:49 +00:00
Daniel Han
f7f540a58b
Studio: strip orphan tool_call XML leaking into visible content (#5735)
* Studio: strip orphan tool_call XML from streamed visible content

The speculative-buffer state machine in
`studio/backend/core/inference/llama_cpp.py` can slice a tool_call XML
block between the silent DRAINING path and the user-visible
content_accum, depending on when in the model's emission the BUFFERING
-> STREAMING -> DRAINING transitions fire. Three leak shapes were
observed in a 2026-05-22 sweep of 900 Qwen3.5 / Qwen3.6 GGUF runs:

  Pre-fix XML leak rate: 20/900 (2.22%), concentrated 6.7% on the
  larger Q8 / MTP configs:

    Qwen3.6-35B-A3B Q8_0         4/60  (6.7%)
    Qwen3.6-35B-A3B-MTP Q4       4/60  (6.7%)
    Qwen3.5-35B-A3B Q8_0         3/60  (5.0%)
    Qwen3.6-27B Q8_0             3/60  (5.0%)

The existing `_TOOL_XML_RE` only matched well-formed
`<tool_call>...</tool_call>` and `<function=...></function>` pairs, so
unterminated openings (close was DRAINED) and orphan closes (opening
was DRAINED) survived the strip and reached the user.

Fix relaxes the regex to also strip:
  1. Orphan opening up to end-of-string: `(?:</tool_call>|\Z)`
  2. Orphan closing tag: bare `</tool_call>` / `</function>`

Verified on the full sweep: 20/900 -> 0/900 (100% of detected leaks
eliminated). 16 unit tests in `test_tool_xml_strip.py` pin all three
leak shapes plus the well-formed cases, plus parametrised checks on
the 5 actual real-world leak samples from the sweep data.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: strip tail-only </parameter> orphan + tighten regex

The 2026-05-22 gdpval sweep surfaced a 4th XML-leak shape not caught
by the earlier regex: a bare `</parameter>\n\n` at end-of-buffer (7
of 192 trials, all Qwen3.5-27B + a few Qwen3.6-27B). The model emits
the full `<tool_call><function=...><parameter=...>...content...
</parameter></function></tool_call>` envelope, the speculative buffer
DRAINS the opening tags as intended, but EOS (max_tokens cutoff)
truncates the outer `</function></tool_call>` close, leaving just
`</parameter>` as the visible tail.

We strip this ONLY when end-anchored (`\s*\Z`) so legitimate
mid-text uses (user code samples, documentation discussing the
Qwen tool-call XML shape) survive. Verified on the 192-trial
gdpval corpus: before=7, after=0.

While at it, fold the five top-level alternations into three by
sharing tag-name and prefix subgroups:

  <tool_call>...    + <function=\w+>...    +    -->  <(?:tool_call|function=\w+)>...
  </tool_call>      | </function>                  -->  </(?:tool_call|function)>

Semantically identical (verified by replay over the 192-trial
corpus + adversarial inputs, 0 diffs) and 1.34x faster on real
workloads. Backtracking-safety pinned by two new perf guards
(256KB '<' spam, 1000x orphan opens).

Tests: 16 -> 28 (6 new functional + 4 well-formed-vs-orphan +
2 perf guards).

* Tighten comments in XML-strip regex and tests

Code says what it does; comments were repeating it. Strip the verbose
explanations down to the WHY-only bits (engine quirk, tail-anchor
rationale, real-world source of each test sample). No code changes.

inference.py:  21 -> 12 lines around _TOOL_XML_RE
test_tool_xml_strip.py: 343 -> 259 lines (-84)
Tests: 28/28 still pass.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-24 05:00:08 -07:00
Daniel Han
dfb3eedf77
ci: broaden Linux + narrow Windows llama.cpp runtime patterns + trim #5741 comments (#5746)
* ci: broaden Linux llama.cpp runtime pattern to lib*.so*

#5741 patched the explicit Linux pattern list to add
``libllama-*-impl.so*`` after ggml-org/llama.cpp#23462 (between
b9279 and b9283) split each binary's entry code into a paired
``lib<binary>-impl.so`` shared library. Same class of upstream
repackaging will hit us again whenever a new shared lib is added.

Mirror what macOS already does and replace the per-lib list with a
single ``lib*.so*`` glob. ``copy_globs`` (line 3614) unions
patterns, so the per-variant ``libggml-cuda.so*`` / ``libggml-hip.so*``
entries were never filtering anything; the spec lives in
``runtime_payload_health_groups`` (line 5209) which keeps the
explicit minimum-required list per variant.

Dry-run against b9296-bin-ubuntu-x64.tar.gz: 40 files copied (all
ggml, llama, mtmd, impl variants + the two binaries we ship), 22
skipped (other CLIs, rpc-server, LICENSE). Functionally equal to
the post-#5741 set.

* cleanup: trim #5741 comments on the pydantic split

Comments added in #5741 explained the original bug in full each
time. They are mostly redundant with the commit message and the PR.
Trim them to one short paragraph per site.

No behavior change.

* ci: narrow Windows runtime pattern to llama-server.exe + llama-quantize.exe

Studio only invokes llama-server and llama-quantize. Mac and Linux
already filter to those two binaries; Windows was the odd one out
with ``*.exe`` copying every CLI upstream ships (llama-cli,
llama-bench, llama-mtmd-cli, ...).

Dry-run on b9296 (win cpu-x64, cpu-arm64, cuda-13.1, hip-radeon):
20 unused EXEs skipped per variant, all DLLs (incl. the new
llama-*-impl.dll family) still copied via ``*.dll``.

``existing_install_matches_choice`` already checks llama-server.exe
exists explicitly (line 5297), so the health gate is unchanged.
2026-05-23 21:48:12 -07:00
pre-commit-ci[bot]
6639a3b31a [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-23 14:00:39 +00:00
Daniel Han
2db8b81854 Studio: harden re-prompt artifact regex for CRLF + catastrophic backtracking
Two robustness fixes for the `_HAS_ANSWER_ARTIFACT` regex from the
parent commit, both caught by a thorough simulation suite covering
Linux/Mac/Windows line-ending portability and adversarial inputs.

1. **CRLF line endings.** The original `\n` literals missed Windows-
   authored or CRLF-converted content (model echoing a pasted prompt,
   etc.). Replaced with `\r?\n` everywhere a newline is required, so
   closed code fences, numbered lists, and end-to-end re-prompt
   decisions all work on `\r\n` as well as `\n`.

2. **Catastrophic backtracking on whitespace spam.** The numbered-list
   alternative `(?:^|\r?\n)\s*\d+\.\s+\S.*?\r?\n\s*\d+\.` was
   O(n^2) on long whitespace runs: `\s*` greedy + `\d+` failing +
   `\s` matching `\r\n` led to repeated backtracking through the
   newline characters. Measured at ~630ms for 10KB of `\r\n` repeats.
   Fix: restrict the post-newline indent to `[ \t]*` (spaces / tabs
   only). After `\r?\n` we are at column 0 and only spaces / tabs
   are a sensible leading indent for a list item; greedy whitespace
   was never needed. New worst case on the same input: <1ms (1000x
   speedup).

Added 5 in-tree tests:
  - test_artifact_regex_handles_crlf_code_fence
  - test_artifact_regex_handles_crlf_numbered_list
  - test_artifact_regex_handles_mixed_lf_crlf
  - test_no_backtrack_on_crlf_spam (asserts <50ms on 10KB \r\n)
  - test_no_reprompt_on_crlf_complete_python_game

All 18 reprompt-guard tests pass. All 253 llama_cpp-related tests pass.
Out-of-tree simulation suite (84 tests) passes on both Python 3.12 and
Python 3.13 inside isolated uv venvs.
2026-05-23 14:00:39 +00:00
pre-commit-ci[bot]
cb6ebc032a [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-05-23 14:00:39 +00:00
Daniel Han
078ae64cdf Studio: don't re-prompt after model already produced a complete answer
The plan-without-action re-prompt at
`studio/backend/core/inference/llama_cpp.py` fires when the model
emits intent-only language ("first I'll ...", "let me ...") without
calling a tool. Previously the heuristic only checked an intent regex
and a 2000-char length cap. The same intent words occur in long
explanations that accompany REAL code or markup, so a complete reply
like "First, let me set up pygame. ```python ... ```" still tripped
the re-prompt, and the synthetic follow-up ("STOP. Do NOT write code
or explain.") wiped the user-visible answer.

Reproduced at scale in a 900-run sweep across 15 Qwen3.5/3.6 GGUF
configs: prompts that emit code or markup (Create a Python game,
Create a Flappy Bird game, weather dashboard HTML, sloth SVG)
landed empty `final_text` for the majority of seeds even on the
strongest configs.

Fix adds a `_HAS_ANSWER_ARTIFACT` regex covering:
  - closed code fences (```...```)
  - HTML pages (<!doctype, <html)
  - complete SVG (<svg...</svg>)
  - 2+ item numbered lists

and a `and not _HAS_ANSWER_ARTIFACT.search(_stripped)` guard on the
re-prompt condition. Plan-only stalls still re-prompt; complete
responses no longer do.

13 new unit tests in `test_llama_cpp_reprompt_guard.py` pin both
directions (artifact present -> no re-prompt; plan-only -> still
re-prompts).
2026-05-23 14:00:39 +00:00
Daniel Han
83b20976f7
ci: unblock Studio Windows + Linux + Mac smoke (#5741)
Bundles three independent CI regressions hitting the maintainer PR
backlog. Each one is verified end-to-end on a staging fork against
real Ubuntu / macOS / Windows GitHub-hosted runners before this
lands.

1. Windows --no-torch install: pydantic + pydantic-core drift to
   incompatible versions under `uv pip install --no-deps -r
   no-torch-runtime.txt` because pip resolves each independently
   from latest. pydantic.VERSION 2.13.4 pins pydantic-core==2.46.4
   but pydantic-core 2.47.0 was the freshest published wheel, so
   `import pydantic` raised
   `SystemError: pydantic-core 2.47.0 is incompatible with the
   current pydantic version`. Resolve pydantic WITH deps in a
   focused pip call (install.sh, install.ps1,
   install_python_stack.py) before the --no-deps no-torch-runtime
   pass so pip pins pydantic-core to the version pydantic declares.
   pydantic's transitive deps (annotated-types, pydantic-core,
   typing-extensions, typing-inspection) are torch-free. Drop the
   redundant `Patch Studio venv with full typer / pydantic dep
   trees` workaround from the four Windows smoke YAMLs.
   Supersedes #5733 + #5734.

2. Linux Studio Update CI: upstream llama.cpp b9261+ split each
   binary's entry code into a paired `libllama-<binary>-impl.so`
   shared library. `llama-server` and `llama-quantize` NEEDED-link
   against `libllama-server-impl.so` / `libllama-quantize-impl.so`
   with RUNPATH `$ORIGIN`, so the prebuilt overlay must copy those
   alongside the binaries. Without that, ldd reports them missing,
   preflight rejects, the installer falls back to source build, and
   studio-update-smoke annotates `setup.sh idempotency regressed`.
   Add `libllama-*-impl.so*` to the Linux runtime patterns and lock
   the pattern in test_rocm_support.TestRuntimePatterns.

3. Mac Studio UI Chat: change-password submit clicked while
   disabled. The disable gate only checked new + confirm password
   length, but Playwright's first click landed before the
   current-password field's React state had committed, so the form
   was simultaneously logically-invalid (current_password empty) and
   the button was disabled. Tighten the gate to require
   `currentPassword.length >= 8` and mirror the same check in the
   submit handler so Enter / autofill cannot bypass.
   Supersedes #5738.
2026-05-23 06:59:16 -07:00
Daniel Han
ebe504b558
Studio: PDF / document attachments for Anthropic + OpenAI (#5689)
* Studio: PDF / document attachments for Anthropic + OpenAI

Studio's local-GGUF chat already supports image attachments via the
`image_url` content part shape. PDFs and other documents had no
plumbing for the external-provider path: there was no normalised
content type the frontend could send that translated to Anthropic's
native `document` block or OpenAI's `input_file`.

Add a Studio-side `input_document` content part on assistant /
user messages with three shapes:

  {type: "input_document",
   file_data: "data:application/pdf;base64,<DATA>",
   filename?: "name.pdf",
   media_type?: "application/pdf"}

  {type: "input_document",
   file_url: "https://example.com/doc.pdf",
   filename?: "doc.pdf"}

Translation:

- Anthropic Messages API: emits a `document` block with
  `{source: {type:"base64", media_type, data}}` or
  `{source: {type:"url", url}}`, plus an optional `title` from
  `filename`. PDFs are extracted server-side by Anthropic per their
  vision/document docs and counted toward input tokens.
- OpenAI Responses API: emits `{type:"input_file", file_data |
  file_url, filename?}`. PDFs are extracted server-side.

Empty / unparseable `input_document` parts are silently dropped so
a malformed frontend payload can't blow up the request.

Tests:

- New `test_multimodal_document.py` with 6 cases pinning the
  outbound body shape for base64 + URL inputs on both providers,
  and the empty-part drop behavior on both.
- The Anthropic assertions strip the prompt-cache wrapper
  (`cache_control:{type:ephemeral}` that the tail-message caching
  layer adds) before comparing the document core fields, so this
  test stays focused on the translation, not the caching layer.

Live verified end-to-end against both providers: a 363-byte
single-page "HELLO" PDF, base64-encoded, attached as a `document`
block to Opus 4.7 and as an `input_file` to gpt-5.5. Both models
correctly extracted the word "HELLO" from the PDF.

Follow-up (out of scope):

- Pydantic schema entry on ChatMessage.content for `input_document`
  (today it rides through because ChatCompletionRequest uses
  extra=allow). Will tighten when the frontend attach button lands.
- Frontend file-picker UX for non-image attachments on the external
  provider path.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: gate empty-content msg + skip empty data-URI payload

Gemini High + Codex P2 on PR #5689:

1. Anthropic translation appended an empty `anthropic_parts` array
   when every part was dropped (e.g. user sent only an unparseable
   input_document). Anthropic 400s on "messages.N.content: at least
   one block is required". Skip the whole-message append when no
   parts survived. The OpenAI Responses path already had the
   equivalent guard, so this brings the two providers into parity.

2. `data:application/pdf;base64,` with no payload (or whitespace-only)
   parses to an empty `source.data` string. Anthropic rejects that
   with 400 as well. Skip the document block before constructing it.

Plus 2 new test cases pinning both behaviors:

- `test_anthropic_empty_only_document_drops_whole_message`: confirms
  a turn whose only content is an unparseable input_document does
  NOT make it onto the outbound `messages` array.
- `test_anthropic_empty_data_uri_payload_is_dropped`: confirms an
  empty-payload data-URI is filtered out at translation time.

(Note re: gemini's other High note about adding `input_document` to
the Pydantic ContentPart union -- ChatCompletionRequest is configured
with `extra=allow` so the part rides through today. Tightening the
union belongs with the frontend attach-button PR that surfaces the
field; called out as follow-up in the PR description.)

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: register input_document in ContentPart + builder

Reviewer caught that the translation code on the external_provider
side was unreachable from a real ChatCompletionRequest:

- ContentPart is a discriminated Union of (text, image_url) only, so
  any `{"type": "input_document", ...}` part was rejected by Pydantic
  at request parsing with a discriminator error before the helper
  could see it.
- _build_external_messages in routes/inference.py only walked text
  and image_url parts, so even with a permissive schema the document
  parts would have been silently dropped instead of forwarded to
  the per-provider translator.

Fixes:

- Add InputDocumentContentPart with optional file_data / file_url /
  filename / media_type and Tag("input_document") on the Union.
- Extend _build_external_messages to pass input_document through as
  a plain dict for vision-capable providers (so external_provider's
  existing Anthropic `document` and OpenAI Responses `input_file`
  mappers actually run) and strip them on non-vision providers.

Tests added: schema accepts input_document, builder passes it to
vision providers, builder strips it on non-vision providers.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: validate file_data before preferring over file_url

Codex P2 caught that the OpenAI input_document translator treats any
truthy file_data as valid and never falls back to file_url. That
means a malformed `data:application/pdf;base64,` (empty payload) or
a whitespace-only data URI gets forwarded as `file_data=""` and
400s the whole turn, AND silently discards a perfectly recoverable
file_url on the same part.

Mirror the Anthropic-side guard onto the OpenAI Responses path:
treat any "data:" URI with no actual base64 payload as missing and
fall through to file_url. Standalone-empty data URIs (no fallback)
are dropped entirely instead of being sent to the wire.

Tests added: empty data URI + valid file_url -> file_url wins,
whitespace-only data URI + valid file_url -> file_url wins,
empty data URI without fallback -> part is dropped.

* Address review: Anthropic side also falls back to file_url on empty data URI

Codex P2 follow-up to my earlier fix: I added the empty-data-URI ->
file_url fallback to the OpenAI Responses translator but missed
the Anthropic translator, which still `continue`d on empty payloads
and discarded an otherwise valid file_url on the same part. Result:
when the frontend supplied both file_data (placeholder / broken)
AND a working file_url, Anthropic silently lost the attachment;
when the message contained only that part, the whole message could
be dropped before reaching the wire.

Mirrored the OpenAI guard: any "data:" URI with no actual base64
payload (`data:application/pdf;base64,` or whitespace-only) is
treated as missing, and the file_url branch takes over. The
all-parts-dropped guard further down already handles the
no-fallback case.

Tests added: empty data URI + valid file_url -> URL source on the
wire with the filename preserved; whitespace-only data URI + valid
file_url -> URL source on the wire.

* Address review: gate input_document passthrough to anthropic + openai

Codex P1: only `_stream_anthropic` and `_stream_openai_responses`
have explicit translation logic for input_document parts (the former
maps to {type:"document", source:...}, the latter to
{type:"input_file", file_data|file_url}). Every other provider
(gemini / mistral / kimi / openrouter / deepseek / qwen / custom)
goes through the generic /chat/completions passthrough that forwards
`messages` verbatim, so any input_document part on a non-vision
route on those providers would 400 with an unknown content_part
type.

Added `_INPUT_DOCUMENT_PROVIDERS = frozenset({"anthropic", "openai"})`
constant and gated the pass-through branch on `provider_type in
_INPUT_DOCUMENT_PROVIDERS`. Every other provider strips the part
(text content survives). Threaded provider_type through from
_proxy_to_external_provider's call site.

Tests updated: vision + provider in {anthropic, openai} still
forwards; six unmapped providers (gemini/mistral/kimi/openrouter/
deepseek/qwen) strip the part; missing provider_type strips
defensively. The existing non-vision drop test still passes.

* Fix stale web_fetch tool-version assertion after merging main

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:22:57 -07:00
Daniel Han
e86f3c5dc7
Studio: wire OpenAI Responses server-side context compaction (#5687)
* Studio: wire OpenAI Responses server-side context compaction

The OpenAI Responses API accepts a `context_management` field that
enables server-side compaction. When the rendered prompt crosses the
configured threshold, the API runs a server-side compaction step and
the request continues against the compacted prefix. No beta header
and no dated version pin are required, per the docs.

Changes:

- Add `compaction_threshold: Optional[int]` (ge=1_000, le=2_000_000)
  to ChatCompletionRequest. Thread through `routes/inference.py` ->
  `stream_chat_completion` -> `_stream_openai_responses`.
- In `_stream_openai_responses`, when threshold is set AND the base
  URL points at cloud OpenAI (api.openai.com), attach
  `context_management: [{type:"compaction", compact_threshold:N}]`
  to the outbound body. Non-cloud bases (ollama, llama.cpp, "custom"
  presets) silently drop the field so we don't 400 those servers.
- Add `test_openai_compaction.py` with 4 cases: cloud OpenAI sets
  the field verbatim, low-threshold probe passes through (we don't
  clamp on the OpenAI side because the API accepts whatever),
  non-cloud base drops the field, omitted threshold leaves body
  untouched.

Live verified against the real OpenAI API on gpt-5.5:
`context_management:[{type:"compaction", compact_threshold:200000}]`
returns 200 with no error.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: accept Azure OpenAI base URLs + raise compaction floor

Two reviewer follow-ups on the OpenAI compaction PR:

1. The `is_openai_cloud = "api.openai.com" in self.base_url` check
   excluded Azure OpenAI Foundry, even though Azure exposes the
   same /v1/responses extensions (context_management,
   prompt_cache_retention, container shell). Users on Azure saw
   their compaction toggle silently no-op. Broadened the check to
   also match `*.openai.azure.com` and made it case-insensitive so
   URLs copy-pasted from the Azure portal still resolve. Non-cloud
   OpenAI-compatible servers (ollama / llama.cpp / vLLM / "custom"
   preset) still fall outside the gate.

2. The schema floor on compaction_threshold was ge=1_000, which is
   well below the upstream Responses API's effective minimum
   (vercel/ai#12486, langchain-ai/langchain#35464 report
   `compact_threshold is not enabled` 400s on Azure at 100k; cloud
   uses 200k as the canonical example). Raised the floor to 10k
   so obvious typos surface as a clean 422 from FastAPI rather than
   an opaque upstream 400 the user has to debug from the SSE
   stream.

Tests added: Azure base URL carries both context_management and
prompt_cache_retention; mixed-case Azure URLs match; schema rejects
9_999 and accepts 10_000.

* Address review: drop schema-level compaction floor (cross-provider regression)

Codex P2 follow-up on the previous floor bump: ge=10_000 was
enforced globally at the ChatCompletionRequest layer, but the field
is documented as a no-op on every non-cloud OpenAI base and every
non-OpenAI provider. With the global floor, an Anthropic / ollama
/ llama.cpp / custom request that happens to carry compaction_threshold
below 10k was rejected with 422 at request validation time instead
of being silently ignored as the description promised.

Reverted the schema floor to ge=1 (any positive int) and rewrote
the description to call out per-provider routing: OpenAI cloud's
effective floor is around 200k and surfaces upstream 400s below
that; _stream_anthropic clamps sub-50k values up. Per-provider
helpers stay the single source of truth on the floor.

Test updated to pin: zero is still rejected, but every positive
value (1, 5_000, 9_999, 10_000, 200_000) passes schema validation.

* Address CodeQL: hostname-anchored OpenAI cloud detection

CodeQL py/incomplete-url-substring-sanitization fired on
`".openai.azure.com" in _base`. An attacker who controls the
configured base_url could slip cloud-only request body fields
(prompt_cache_retention, context_management compaction, container
shell) to an arbitrary server with:

  https://evil.com/api.openai.com/v1
  https://api.openai.com.attacker.com/v1
  https://attacker.com/.openai.azure.com/v1
  https://my-resource.openai.azure.com.attacker.com/openai/v1

Replaced the substring check with a `_is_openai_family_cloud`
helper that runs urllib.parse.urlparse on the URL and matches the
lowercased hostname exactly (`api.openai.com`) or via `endswith`
on the leading-dot suffix (`.openai.azure.com`). Both halves are
host-anchored so path / fake-subdomain bypasses fail.

Test added: every attacker-controlled bypass shape above must NOT
carry context_management OR prompt_cache_retention on the wire.
Existing Azure and openai.com tests still pass.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: scope compaction_threshold description to OpenAI on this branch

Codex P2: the field description on this PR mentioned Anthropic
compaction behavior, but the Anthropic wiring lives on PR 5686
(separate branch). On feat/openai-compaction alone, _stream_anthropic
has no compaction_threshold parameter, so the field is silently
ignored for Anthropic requests and the doc claim was misleading.

Trimmed the description to OpenAI cloud + Azure Foundry only on
this branch. PR 5686 already re-adds the Anthropic clause via its
own change, so the rebase / merge order on main will land the
combined description naturally once both PRs ship.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:20:45 -07:00
Daniel Han
9a737facaf
Studio: wire Anthropic server-side context compaction (#5686)
* Studio: wire Anthropic server-side context compaction

Anthropic ships server-side context compaction as a beta
(`compact-2026-01-12`). When the rendered prompt crosses the
configured input-token threshold, Anthropic runs an extra LLM pass
that summarises older turns and the request continues against the
compacted prefix. The response carries the original top-level fields
plus a new `context_management` block (with `applied_edits`) and
`usage.iterations[]` accounting per pass.

Per the docs the feature is currently supported on Opus 4.6, Opus 4.7,
Sonnet 4.6, and Mythos preview. The minimum threshold is 50k tokens;
under-50k requests 400.

Changes:

- Add prefix gate + helper `_anthropic_supports_compaction` plus
  constants `_ANTHROPIC_COMPACTION_PREFIXES`, `_ANTHROPIC_COMPACTION_BETA`,
  `_ANTHROPIC_COMPACTION_TYPE`, `_ANTHROPIC_COMPACTION_MIN`.
- Add `compaction_threshold: Optional[int]` to ChatCompletionRequest
  (50k ge bound, 2M le bound). Thread through `routes/inference.py`
  -> `stream_chat_completion` -> `_stream_anthropic`.
- In `_stream_anthropic`, when threshold is set AND the model
  accepts compaction, attach `context_management.edits[{type:
  "compact_20260112", trigger:{type:"input_tokens", value:N}}]` to
  the outbound body. Sub-50k values are clamped up to 50k to keep
  the request well-formed.
- Refactor the anthropic-beta header builder to merge any combination
  of `code-execution-2025-08-25` + `compact-2026-01-12` flags into
  one header value. Unrelated betas added at the registry level still
  pass through.
- Add `test_anthropic_compaction.py` with 16 cases: gate matrix
  (every doc-listed model), correct body shape, threshold clamping,
  beta header merge with code execution, silent no-op on unsupported
  models, omitted-threshold pass-through.

Live verified end-to-end against the real Anthropic API:
`compact_20260112` accepted on Opus 4.7, response carries
`context_management.applied_edits` + `usage.iterations[]` as
documented. (The first WebFetch-summarised version of these docs
suggested `compact_20260120`; the actual API only accepts
`compact_20260112`, matching the beta-header date. Worth pinning
behind a test so a future doc update can't drift back.)

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: drop ge=50_000 clamp + parse usage.iterations[]

Two reviewer follow-ups on the compaction PR:

1. Pydantic ge=50_000 on compaction_threshold was dead code.
   FastAPI rejected sub-50k threshold values with a 422 before the
   `max(int(...), _ANTHROPIC_COMPACTION_MIN)` clamp in
   _stream_anthropic could ever fire. Relaxed the floor to ge=1 so
   the in-helper clamp actually does its job; the schema comment
   now explains why this is intentional. Added a regression test
   that posts a value of 1 and 49_999 through the real request
   schema.

2. Anthropic publishes per-iteration token counts in
   `usage.iterations[]` whenever a fresh compaction has run, and
   the top-level input_tokens / output_tokens cover only the
   `message` iteration -- billing must add the compaction
   iterations on top. Aggregate compaction iteration tokens into
   `last_usage["compaction_input_tokens" / "compaction_output_tokens"]`
   so the cost surface (PR 5690) can read them without re-walking
   the array, and surface both figures in the closing stream
   summary log. Added two tests: one that pins the aggregation on a
   compacted turn and one that pins `None` when no fresh
   iterations land (so re-applied compaction blocks don't double-bill).

Sourcing: https://platform.claude.com/docs/en/build-with-claude/compaction

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: round-trip Anthropic compaction blocks across turns

Codex P1: once context_management is enabled and Anthropic runs
server-side compaction mid-stream, the response carries a
`{type:"compaction", content:"<summary>"}` content block on the
assistant message. The translator only handled text_delta and
input_json_delta on content_block_delta, so the compaction block
was silently dropped. Worse, the request schema's ContentPart
discriminated Union didn't accept `type:"compaction"`, and
_build_external_messages didn't pass it through, so even a
hand-crafted assistant message carrying the block would 422 at
parse time. Net result: Anthropic re-compacted from scratch on
every subsequent turn, wasting input tokens and reasoning budget.

End-to-end backend wiring of the round-trip:

1. SSE translator. _stream_anthropic now tracks a `current_compaction`
   state slot. content_block_start with type=="compaction" seeds it
   (Anthropic may include the summary on the start event AND/OR
   stream it via text_delta events on the same block index --
   handle both). text_delta inside a compaction block routes into
   the compaction buffer instead of the user-visible content
   stream, since the summary is opaque internal state, not
   assistant prose. content_block_stop emits a `compaction_block`
   tool_event carrying the full summary so the chat-adapter can
   persist it. compaction_blocks_seen is surfaced in the closing
   summary log.

2. Pydantic schema. Added CompactionContentPart with Tag("compaction")
   on the ContentPart Union so requests carrying the block parse
   cleanly. Required `content` field with a docstring pointing at
   the Anthropic docs.

3. Message builder. _build_external_messages forwards compaction
   parts on both vision and non-vision paths; the per-provider
   stream helper decides whether to forward to the wire (Anthropic
   does; other providers ignore the part). When a non-vision route
   ends up with a single text part, collapse back to a string
   so providers that don't accept content arrays still get the
   expected shape.

4. _stream_anthropic outbound translator. {type:"compaction"} parts
   on an assistant message land on the wire verbatim. Empty/missing
   `content` is skipped so a malformed stored block can't 400
   Anthropic.

Tests added (5): stream emits compaction_block tool event with the
summary intact; user-visible content stream does NOT carry the
summary text; outbound body forwards compaction parts verbatim on
the next turn; Pydantic schema accepts the part; builder passes
it through on both vision and non-vision provider routes.

Frontend follow-up: the chat-adapter needs to persist the
compaction_block tool_event onto the stored assistant message so
turn N+1 includes it in payload.messages. Pinned in the PR
description.

Sourcing: https://platform.claude.com/docs/en/build-with-claude/compaction

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: gate compaction-part passthrough to Anthropic only

Codex P1: my previous round-trip change preserved {type:"compaction"}
parts on every provider route in _build_external_messages. That
meant a chat history with prior compaction state silently leaked
the Anthropic-specific block to OpenAI/DeepSeek/Mistral/Gemini/
Kimi/OpenRouter on a provider switch, where generic
/chat/completions passthrough hands the unknown content type to
the upstream API and 400s the whole turn.

Added a `provider_type` kwarg to _build_external_messages and
gated the compaction forwarder on `provider_type == "anthropic"`.
Every other value (including the legacy None for callers that
don't pass it yet) strips the part. The Anthropic stream helper
still maps it to a native `compaction` block on the wire.

Threaded provider_type through from _proxy_to_external_provider's
call site.

Tests updated: vision + provider="anthropic" still forwards; six
non-anthropic providers strip the part; missing provider_type
strips defensively; non-vision + anthropic still forwards; non-vision
+ non-anthropic collapses back to a text string.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:19:09 -07:00
Lee Jackson
61ed4cac51
Studio: persist chat history in backend storage (#5272)
* feat: Persist chat history in backend storage

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address chat tombstone batching review

* fix: update desktop auth routes stub

* chat db settings storage

* chat db settings routes

* chat db settings client

* chat db settings store

* chat db settings wiring

* chat db history storage

* chat db settings migration

* chat db settings fallback

* chat db container metadata

* chat db legacy migration fixes

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* chat ci auth background reads

* chat auth storage fixes

* chat migration final fixes

* chat export batch message lookup

* chat history review fixes

* chat prune sync fix

* chat settings hydration retry

* gate settings persistence

* Scope chat-history rows by subject; fix hijack, clear-confirm, hydrate race

Backend storage and routes:
- chat_threads / chat_messages / chat_settings carry a NOT NULL subject
  column with composite PRIMARY KEY (id, subject). Two authenticated
  identities can no longer see or wipe each other's data.
- Pre-existing rows on an existing studio.db migrate under sentinel
  subject __legacy_unscoped__ via rename + rebuild + copy; single-user
  installs see no behavior change.
- ON CONFLICT(id, subject) DO UPDATE ... WHERE chat_messages.thread_id =
  excluded.thread_id refuses cross-thread re-parenting via upsert.
  upsert_chat_message + sync_chat_messages now raise
  ChatMessageThreadMismatch which the routes map to HTTP 409.
- replace_thread_messages rejects body messages whose threadId does not
  match the URL thread (HTTP 400) instead of silently rewriting them.
- DELETE /api/chat requires ?confirm=true, returns row count, logs the
  subject and count.
- upsert_chat_settings_merge does read + deep-merge + write inside a
  single BEGIN IMMEDIATE so concurrent writers no longer drop each
  other's updates. The route delegates to this helper.
- New POST /api/chat/messages:batch returns {thread_id -> messages[]}
  for many threads in one HTTP call. Subject-scoped. Unknown ids return
  empty lists instead of 404 so the sidebar/search caller can rebuild
  atomically.

Frontend:
- chat-runtime-store: hydrate-failure catch sets settingsHydrated:true
  so a transient backend blip no longer permanently disables
  persistence. setParams bumps inferenceParamMutationVersions
  unconditionally so a slow hydration response cannot clobber a
  pre-hydrate user edit. saveSettingsPatch replaces the serial chain
  with a debounced pendingPatch + deep merge; flush on beforeunload.
- chat-history-storage: clearStoredChats returns ClearStoredChatsResult
  distinguishing backend / legacy / both outcomes.
  listStoredChatThreadsWithMessages uses the batched fetch (one HTTP
  call) instead of Promise.all per-thread; legacy Dexie fallback only
  fires when the batch result is empty.
- chat-api: batchListChatMessages with graceful 404 / 405 fallback to
  per-thread listChatMessages for older servers.
- chat-thread-tombstones: store {id, deletedAt} tuples with 90-day GC
  and a 5000-entry cap so localStorage stays bounded. Back-compat reads
  pre-fix plain strings. Adds removeChatThreadTombstones (rollback) and
  clearAllChatThreadTombstones (post-legacy-purge clean-up).
- use-chat-sidebar-items: deleteChatItem tombstones synchronously
  BEFORE the backend round-trip and rolls back on failure (restores
  pre-PR optimistic UX). 300 ms trailing debounce on
  CHAT_HISTORY_UPDATED_EVENT plus requestSeq guard so stream-time event
  bursts produce at most one fetch per quiet window.

Tests:
- studio/backend/tests/pr5272_sim/ adds 64 regression tests covering
  schema migration from pre-fix shape, subject scoping, cross-thread
  hijack, bulk-replace mismatch, clear-confirm, concurrent settings,
  unicode + 2MB content + SQL-injection-safe binding, chunking
  boundary at 900 and 901 ids, batched endpoint (multi-subject + 1200
  ids + per-thread order), and grep contracts for the frontend patches.
  test_chat_history_storage.py updated to pass subject.

Verified locally on Linux + macOS + Windows GitHub Actions runners
(staging fork): 64 pass + 2 from the PR's own backend test on all
three OSes.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Drop subject scoping and clear-confirm gate (Studio is single-user)

Per maintainer feedback: subject scoping, cross-thread message hijack
guard, and DELETE /api/chat ?confirm=true gate are unnecessary because
Studio is intentionally single-user (the client already shows a confirm
dialog before clear-all).

This commit reverts those backend changes and keeps only the
non-multi-user pieces from the earlier fix commit:

- studio_db.py: restored to pre-fix shape; adds upsert_chat_settings_merge
  which does atomic read + deep-merge + write under BEGIN IMMEDIATE so
  two concurrent slider drags cannot drop one another's updates.
- routes/chat_history.py: restored; put_settings now calls the atomic
  merge instead of doing the read-merge-write across three separate
  connections. Adds POST /api/chat/messages:batch to collapse the
  sidebar/search rebuild from N round-trips to 1.
- frontend/api/chat-api.ts: align batchListChatMessages request and
  response keys with the backend (threadIds / messagesByThreadId).
- tests/test_chat_history_storage.py: add atomic-merge concurrency test,
  deep-merge nested-key test, and 901-id chunking-boundary test.
- Drop the pr5272_sim test directory (those tests covered the reverted
  subject-scoping/hijack/confirm behavior).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix sidebar delete crash, keepalive on settings beforeunload flush, search rebuild race

Two correctness bugs and one perf race surfaced by a fresh code review of
the prior fix commit:

- chat-api.ts: notifyChatHistoryUpdated was declared as a non-exported
  function, but use-chat-sidebar-items.ts imports it. The import would
  fail tsc with TS2305 and at runtime the optimistic-delete and
  delete-failure rollback paths would both throw.
- chat-runtime-store.ts + chat-settings-api.ts + chat-settings-storage.ts:
  the beforeunload settings flush is now actually keepalive. Without it
  the browser cancels the in-flight PUT on tab close, so the last slider
  drag is silently dropped (which is exactly the case the
  debounce+beforeunload combination was meant to protect against).
- use-chat-search-index.ts: rebuilds now coalesce with a 300ms trailing
  debounce and discard out-of-order responses via a requestSeq guard.
  Matches the sibling pattern in use-chat-sidebar-items.ts so two rapid
  CHAT_HISTORY_UPDATED_EVENTs (run-start + run-end save during a turn)
  cannot land with stale data winning.
- chat-thread-tombstones.ts: drop dead clearAllChatThreadTombstones with
  no call sites; Dexie is never wiped so the function has no use.

* fix(studio): protect chat persistence writes

* fix(studio): align chat history clear semantics

* fix(studio): show partial chat clear feedback

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix(studio): preserve chat persistence fallbacks

* fix(studio): harden chat thread persistence checks

* Preserve chat message timestamps

* Gate chat stream on history save

* Make chat thread backfill best effort

* Avoid chat message 404 probe

* Tighten chat legacy fallbacks

* chat: server-side ledger so legacy Dexie import is recoverable

The boolean localStorage sentinel
(unsloth_chat_legacy_imported_to_studio_db) made importLegacyChatsIfNeeded
non-recoverable: deleting studio.db while the browser keeps the flag
silently hides every legacy Dexie thread from the sidebar (verified by
the 3-GPU validation probe; matches the third review comment on PR
#5272). Same trap fires for browser-profile sync to a fresh machine
and any other path that wipes studio.db while keeping IndexedDB.

Source of truth moves into studio.db itself via a new
chat_legacy_import_log table keyed by legacy thread id. The ledger
disappears together with studio.db, so the next launch re-runs the
import from whatever Dexie still holds. localStorage stays as a
per-session perf hint only.

Performance, all bounded by the three new fast-paths before any
backend work:

  A) localStorage hint says "imported earlier in this session" -- 0
     network, ~0 ms. Covers the warm sidebar mount.

  B) indexedDB.databases() reports no "unsloth-chat" DB -- 0 network,
     ~1 ms. Covers every new user who never had the old browser-only
     Studio (the common case after launch).

  C) db.threads.count() + db.messages.count() are both 0 -- 0 network,
     ~5 ms. Covers returning users who migrated long ago and Dexie was
     never repopulated.

Only when all three miss does the code talk to the backend
(GET /api/chat/import-ledger -> diff vs Dexie -> existing import path
-> POST /api/chat/import-ledger to record what was just imported).
Per-thread tracking is enough because Dexie is read-only after this
PR; a thread's message set does not grow.

Backend deployments that predate the import-ledger routes are
handled transparently: the client treats 404/405 as an empty ledger
and re-runs the (idempotent via UPSERT) import on next launch.

Changes:
- storage/studio_db.py: new chat_legacy_import_log table (WITHOUT
  ROWID, PK on legacy_thread_id) + list_chat_legacy_import_log() +
  record_chat_legacy_import_log() (idempotent batch UPSERT).
- routes/chat_history.py: GET + POST /api/chat/import-ledger with the
  obvious request/response models.
- frontend api/chat-api.ts: listChatImportLedger() (returns a Set for
  O(1) diff) + recordChatImportLedger(), both with 404/405 fallback.
- frontend utils/chat-history-storage.ts: importLegacyChatsIfNeeded
  gains three fast-paths, ledger fetch on the slow path, and writes
  the ledger after a successful import. The localStorage helper is
  unchanged on the surface; it just stops being authoritative.
- tests: 5 new test_legacy_import_log_* cases (empty default, record
  + list round-trip, idempotency, input dedup, empty/null ignore).
  All 9 pre-existing tests still pass.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Make the legacy-import recovery actually recoverable

The previous commit added a server-side ledger to make Dexie -> studio.db
import recoverable after a studio.db wipe, but the localStorage perf hint
still short-circuited the import gate before the ledger was ever consulted.
After a wipe, the hint stayed "true" and the bulk re-import never ran -- the
ledger sat empty and only the per-thread lazy materialize-on-continue path
restored data.

Changes:

- Remove the localStorage short-circuit from importLegacyChatsIfNeeded so
  the ledger is checked on every fresh tab. legacyChatImportPromise keeps
  the per-session cache; the hint now only matters for the listing paths.
- Batch the slow path: one db.messages.where().anyOf().toArray() and one
  batchListChatMessages() instead of 2N round-trips. At 1k threads this
  drops a multi-second blocking import to a single request pair.
- recordChatImportLedger returns {accepted, inserted, supported}. The
  localStorage hint is only flipped when supported is true, so old
  backends (404 / 405 / 501) no longer permanently poison recovery.
- Ledger backfill: threads already present in chat_threads but missing
  from the ledger now get added too, so old-FE-then-new-FE deployments
  don't redo the diff every launch.
- Backend response field renamed recorded -> {accepted, inserted}.
  accepted is the deduped non-empty input count; inserted is the rows
  actually new (via INSERT ... RETURNING). Bounded by Field(max_length=
  10_000) on the request payload.
- Storage helpers renamed: chat_legacy_import_log -> chat_legacy_imports,
  record_* -> upsert_* to match the existing noun/verb conventions.
- DEXIE_DB_NAME exported from db.ts; duplicate constant in
  chat-history-storage.ts removed.
- 3 new route-level tests for /api/chat/import-ledger covering the
  round-trip, the (accepted, inserted) split, and the 10k payload cap.

All 18 chat-history tests pass.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: shine1i <wasimysdev@gmail.com>
Co-authored-by: danielhanchen <michaelhan2050@gmail.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2026-05-22 06:18:05 -07:00
Daniel Han
a2d2b7866f
Studio: wire Anthropic web_fetch server-side tool (#5671)
* Studio: wire Anthropic web_fetch server-side tool

Studio's Anthropic passthrough only forwarded web_search and
code_execution when enabled_tools was set. Asking Claude through Studio
to fetch a URL produced no fetch (the tool was not in the outbound
tools array), so users had to fall back to web_search even when they
already had the exact URL they wanted.

This change opts in web_fetch_20250910 when enabled_tools contains
"web_fetch". The new tool entry is appended alongside any existing
web_search / code_execution entries:

  {"type": "web_fetch_20250910", "name": "web_fetch", "max_uses": 5}

No anthropic-beta header is required (web_fetch is GA); the existing
code-execution-2025-08-25 flag continues to merge cleanly when both
tools are enabled in the same turn.

SSE translation mirrors the web_search path. A `server_tool_use` block
with name="web_fetch" emits a `tool_start` _toolEvent carrying the
URL the model asked to fetch; the matching `web_fetch_tool_result`
block emits a `tool_end` _toolEvent whose result string follows the
Title / URL / Snippet shape parseSourcesFromResult on the frontend
already expects, so the source pill renders identically. Error blocks
(`web_fetch_tool_error`) are surfaced as "Error: <error_code>" matching
the code_execution error path.

The final "Anthropic stream complete" log line picks up web_fetch_
requested / web_fetch_invocations / web_fetch_urls so support reports
of "the model did not fetch anything" can be triaged from the log.

Verified end to end against claude-haiku-4-5 with
`enabled_tools=["web_fetch"]`: the model emitted tool_start with
url=https://example.com and tool_end with the page Title + URL +
Snippet, plus the assistant message correctly read back "Example
Domain" as the title.

Tests:
- 5 new unit tests in test_anthropic_web_fetch.py covering tool
  registration, the combined web_search + web_fetch + code_execution
  request body, the pill-off case, and SSE translation for both
  success and error paths.
- All 242 existing Anthropic + OpenAI provider tests still pass.

The enabled_tools field description in models/inference.py is updated
so OpenAPI consumers see the new option.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* web_fetch: title fallback to URL, log parse failures, drop dead checks

Three review nits on the previous commit:

1. `_format_web_fetch_result` left `title` empty when Anthropic omitted
   `document.title`. The frontend `parseSourcesFromResult` only emits
   a source pill when both `Title:` and `URL:` lines are present, so
   fetches against pages without an HTML title tag silently lost
   their citation in the UI. Fall back to `title = title or url`,
   matching the web_search formatter.

2. The broad `except Exception` around `json.loads(buffer)` for the
   web_fetch input swallowed the failure with no trace. Log at debug
   so a malformed partial_json buffer can be triaged from the server
   log without changing behavior.

3. `inner` was already sanitised to a dict at the matching
   content_block_start and `_format_web_fetch_result` always returns
   a non-empty string (defaulting to "(fetch complete)"), so the
   `isinstance(inner, dict) else {}` guard and the
   `result_text or "(fetch complete)"` fallback at the emit site
   were dead code. Removed.

Added a test exercising the titleless path so the fallback stays
covered.

* chat-adapter: emit source pills for web_fetch tool calls

`parseSourcesFromResult` was only wired up for tool calls where
`toolName === "web_search"`, so the Title / URL / Snippet block the
backend formatter emits for `web_fetch_tool_result` never reached the
source-pill renderer. Users saw the raw tool result in the tool card
but the dedicated source-pill row at the message tail stayed empty.

Both web_search and web_fetch ship the same text shape today, so the
fix is to broaden the gate.

* Address review: wire web_fetch from Search pill + fix pause_turn truncation

Two reviewer follow-ups on the Anthropic web_fetch PR:

1. The backend tool wiring landed but the frontend chat-adapter
   never put `web_fetch` in `enabled_tools`, so toggling the Search
   pill only ever attached `web_search` -- web_fetch was unreachable
   from the UI. Added providerSupportsBuiltinWebFetch() (Anthropic
   today) and paired the entry with the existing Search pill, since
   the canonical workflow is "search returns URLs, fetch reads
   them" and there is no separate UI toggle yet.

2. `pause_turn` from Anthropic's stop_reason vocabulary fell through
   the finish_reason map's "stop" default, which the OpenAI-format
   client renders as end-of-message and truncates the answer. Per
   the docs pause_turn means "Claude paused a long server-tool
   turn (web_search / web_fetch) and will resume". Mapped to None
   and skipped the chunk emission so the SSE stream still ends with
   [DONE] on message_stop but no terminal finish_reason lands on
   the client. While there: added explicit mappings for `tool_use`
   (-> tool_calls) and `refusal` (-> content_filter) which were
   also falling through to "stop".

Tests added: pause_turn emits no finish_reason, end_turn still
emits "stop", refusal maps to "content_filter".

Sourcing: https://platform.claude.com/docs/en/api/messages#response-stop-reason

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:03:48 -07:00
Daniel Han
2201fd687b
Studio: per-session cost calculator + /api/providers/pricing endpoint (#5690)
* Studio: per-session cost calculator + /api/providers/pricing endpoint

Neither the Anthropic Messages API nor the OpenAI Responses API
reports a `cost` field on the response. Both expose detailed token
counts (input, output, cache hits, server-tool invocations); pricing
multipliers live in the provider docs. The frontend's "cost so far"
display was impossible without scraping the server log.

Land the math + a snapshot endpoint so the cost calculator can run
client-side from the existing usage chunk plumbing. The actual UI
hookup belongs in a frontend follow-up (and is gated on PR #5670's
usage-chunk emission landing so the frontend sees the usage block
in the first place).

Changes:

- New `core/inference/pricing.py` with:
  - Per-MTok base pricing tables for every active Anthropic and
    gpt-5.x family member. Dated snapshots inherit the canonical-id
    price via prefix match so future snapshots cost the same as the
    canonical id until pricing changes.
  - Shared multipliers for Anthropic cache writes (5m: 1.25x, 1h: 2x)
    and reads (0.1x); OpenAI cache reads (0.1x); Anthropic server
    tool surcharges ($10 / 1k web_search, $0.05 / hour code_exec
    beyond the 50-hour daily free tier).
  - `calculate_cost(provider, model, usage)` returns a per-turn USD
    breakdown plus billable token counts, with priced=False for
    unknown models so the UI can still render token counts.
  - `pricing_snapshot()` returns the whole table for the frontend
    so it doesn't re-implement the multipliers.
- New `GET /api/providers/pricing` returning the snapshot, scoped
  behind the existing auth dependency.
- New `backend/tests/test_pricing.py` with 12 cases pinning the
  math against documented values: base input/output multiplication,
  5m / 1h / read multipliers, default-to-5m fallback when the
  breakdown is absent, web_search per-1k pricing, code_execution
  per-hour pricing, dated-snapshot fallback, OpenAI cache-read
  discount accounting (cached tokens subtracted from full-price
  bucket and re-billed at 0.1x), unknown model graceful-degrade,
  and the snapshot endpoint shape.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: verified OpenAI pricing + fix billable input double-count

Address the cost-calculator review:

- OpenAI prices were 2-6x under the actual published rates.
  Cross-checked the live developers.openai.com/api/docs/pricing page
  and replaced every entry. gpt-5.5 is 5/30, gpt-5.5-pro is 30/180,
  gpt-5.4 is 2.5/15, gpt-5.4-mini 0.75/4.5, gpt-5.4-nano 0.20/1.25,
  gpt-5.3-codex 1.75/14. Added chat-latest alias to the canonical
  chat-snapshot rate. Dropped o3 / o4 / gpt-4.5 rows that are no
  longer listed on the page; calculator returns priced=False instead
  of silently billing at zero.

- billable_input_tokens was double-counting cached tokens for
  OpenAI. Anthropic excludes cache_* buckets from input_tokens so
  we add them; OpenAI folds cache_read_input_tokens into
  input_tokens already, so the tooltip read 1.8M for a 1.0M bill.
  Branched the math by provider and added a regression test.

Sourcing notes in the module docstring updated.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: canonical 4.5 ids, long-context tier, OpenAI tool fees

Three Codex P1 follow-ups on the cost calculator:

1. Canonical Anthropic 4.5 ids missing from ANTHROPIC_PRICING.
   claude-opus-4-5 / claude-sonnet-4-5 / claude-haiku-4-5 (no date
   suffix) are the ids used by backend defaults
   (PROVIDER_REGISTRY['anthropic'].default_models), but the table
   only had the dated forms. _lookup's prefix fallback doesn't help
   because the canonical id is SHORTER than the dated key, so
   str.startswith goes the wrong way and the calculator returned
   priced=False + zero cost. Added the canonical aliases for
   opus-4-5, sonnet-4-5, haiku-4-5, and opus-4-1.

2. OpenAI long-context tier. gpt-5.5 and gpt-5.4 cross over at
   272k input tokens to a 2x input / 1.5x output rate (gpt-5.5:
   $5/$30 -> $10/$45; gpt-5.4: $2.50/$15 -> $5/$22.50). Turns past
   the threshold were systematically undercounted at headline
   rates. Added long_context_threshold / long_context_input_per_mtok /
   long_context_output_per_mtok columns and a tier-selection step
   in calculate_cost; model_priced gains a "(long-context >272000)"
   suffix when the higher tier applies so the tooltip can show
   which rate was used. gpt-5.5-pro / gpt-5.4-pro / mini / nano /
   codex have no published long-context tier today, so they keep a
   single rate.

3. OpenAI server-tool surcharges. web_search is $10/1000 calls and
   the hosted shell container is $0.03 per 20-minute session on the
   default 1g tier (~$0.09/hr). server_tools_usd was previously
   stuck at 0.0 for OpenAI even when web_search and shell tools
   fired, so sessions with tool use understated cost. Added
   OPENAI_WEB_SEARCH_USD_PER_1K and OPENAI_CONTAINER_USD_PER_HOUR
   constants plus a parallel of the Anthropic surcharge block that
   reads counts from usage["openai_tool_use"]. The SSE translator
   wires the counts in a follow-up commit; the calculator is now
   ready for them. pricing_snapshot also exposes both constants so
   the frontend tooltip can render the per-call rate.

Existing tests updated to stay in the short-context tier where they
were testing base rates; new tests pin canonical 4.5 lookups,
long-context crossover on gpt-5.5/gpt-5.4, the absence of crossover
on mini/nano/codex, and OpenAI tool surcharges (web_search,
container hours, combined total).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:03:43 -07:00
Daniel Han
5b41872e8b
Studio: wire OpenAI image_generation tool (#5688)
* Studio: wire OpenAI image_generation tool

OpenAI's Responses API exposes server-side image generation as a
tool entry (`{type: "image_generation"}`); the result comes back as
an `image_generation_call` output item with the base64 image on
`result`, the actual prompt used on `revised_prompt`, plus `size`,
`quality`, `output_format`, `background`. The model decides when to
call the tool based on the user's request; rendering uses one of
the gpt-image-* backbones server-side.

Available on every gpt-5.x family member plus gpt-4.1, gpt-4o, o3,
o4-mini per the docs.

Changes:

- Append `{type:"image_generation"}` to the Responses request tools
  array when `enabled_tools` carries `image_generation` AND the base
  URL points at cloud OpenAI. Non-cloud bases (ollama, llama.cpp,
  "custom" presets that collapse to provider="openai") silently drop
  the tool to avoid 400s.
- Mirror the same logic in `_build_body` (the post-expiry retry
  builder) so retries carry the same tool set as the original
  attempt.
- Handle `image_generation_call` items in
  `response.output_item.done`: emit `tool_start` with
  `arguments:{kind:"image", prompt:<revised_prompt>}` and `tool_end`
  with `image_b64`, `image_mime`, `size`, `quality`, `background`
  so the chat adapter can render an inline preview. Image bytes go
  on the tool_end chunk; no extra fields on the chat-completions
  envelope so the OpenAI SDK shape stays clean.
- Add `import time` (used for synthesised tool_call_id fallback).
- Add `test_openai_image_generation.py` with 5 cases: tool entry on
  cloud OpenAI, combined with web_search + code_execution
  (verifies all three coexist), non-cloud drop, omitted pill leaves
  body untouched, output item translation produces the expected
  tool_start + tool_end chunks.

Live verified end-to-end: `gpt-5.4-mini` with `image_generation`
tool returned an `image_generation_call` carrying ~1MB of base64
PNG plus the gpt-image backbone's revised prompt.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Use time.time_ns() for synthesised image_generation tool_call_id

Gemini medium on PR #5688: `int(time.time() * 1000)` has 1ms
resolution; two image generations resolving in the same millisecond
would collide on the synthesised id. Bump to nanoseconds.

(In practice the upstream `image_generation_call` item always carries
its own `id`; the synthesised fallback only fires when OpenAI omits
it -- rare, but cheap to harden.)

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:03:38 -07:00
Daniel Han
b8dde0a835
Studio: support Anthropic 1h cache TTL via prompt_cache_ttl (#5685)
* Studio: support Anthropic 1h cache TTL via prompt_cache_ttl field

Anthropic exposes two ephemeral cache pools per request: the default
5-minute pool, and a 1-hour pool selected by attaching `ttl:"1h"` to
the `cache_control` marker. 1h writes are billed at 2x base input vs
1.25x for 5m, but reads stay at 0.1x for both, so a single extra read
landing more than 5 minutes after the write pays off the premium.

Studio hardcoded the 5m pool via `cache_control: {type:"ephemeral"}`
on both breakpoints. For chats with multi-minute idle gaps (people
juggling tabs, long-running tool calls between turns), the cache
expires before the next turn and every read becomes a cache_creation,
not a cache_read -- exactly the case where the 1h pool wins.

Changes:

- Add `prompt_cache_ttl: Optional[Literal["5m", "1h"]]` to
  ChatCompletionRequest. Default (None) preserves today's 5m behavior.
- Thread through `routes/inference.py` ->
  `stream_chat_completion` -> `_stream_anthropic`.
- Build a shared `cache_marker` dict in `_stream_anthropic`; attach
  `ttl` only when the request asks for one of the two valid values.
  Unknown TTL strings are silently dropped to avoid sending malformed
  markers (the upstream API would 400).
- Apply the same marker to both existing breakpoints (system block at
  line 1175 and the latest-message tail at line 1198 / 1213) so the
  pool selection is consistent across the whole prefix.
- Add `test_anthropic_cache_ttl.py` with 11 parametrized cases
  pinning the outbound body shape: omitted -> default marker;
  explicit `5m`/`1h` -> ttl field set; unknown values dropped;
  caching off -> no markers at all.

Verified upstream that `cache_control: {type:"ephemeral", ttl:"1h"}`
is accepted by the Anthropic API today; no beta header required.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Relax prompt_cache_ttl to Optional[str] (Codex P1)

Declaring `prompt_cache_ttl` as `Optional[Literal["5m", "1h"]]` made
FastAPI/Pydantic 422 the request before _stream_anthropic could even
see the field. The whole point of the downstream drop-unknown-values
behaviour was to keep a stale frontend from crashing the request;
the strict Literal at the request layer defeated that.

Loosen the schema to Optional[str]; the existing in-helper guard
already restricts forwarded values to {"5m", "1h"} (everything else
is silently dropped). Test suite stays unchanged -- the bogus-value
cases in test_anthropic_cache_ttl.py already pass arbitrary strings
through and assert they are dropped before the wire.

* Address review: confirm extended-cache-ttl beta header is GA

Reviewer asked whether the 1h cache TTL still requires the
`extended-cache-ttl-2025-04-11` anthropic-beta header. Investigated:

- Live-tested api.anthropic.com on claude-opus-4-7 (2026-05-22)
  with cache_control={type:"ephemeral", ttl:"1h"} and NO beta
  header. Got status 200 and ephemeral_1h_input_tokens populated
  on the create turn, plus cache_read_input_tokens populated on
  the reuse turn.
- Cross-checked the current prompt-caching docs: no mention of
  any beta header on the 1h TTL path.

Conclusion: the gate has been promoted to GA. The code already
does not send the beta header (the cache_marker dict only carries
`type`/`ttl`), so no wire change is needed. Pinned the contract
with two regression tests that assert the header is NOT on the
outbound request, and added a docstring note explaining the
investigation outcome so a future reader does not re-add it.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:03:32 -07:00
Daniel Han
f399e3b9d1
Studio: per-model Anthropic server-side tool versions (#5679)
* Studio: per-model Anthropic server-side tool versions

Anthropic ships date-pinned tool versions per model family. Studio
currently hard-codes `web_search_20250305`, `web_fetch_20250910`, and
`code_execution_20250825` for every model, which means Opus 4.6/4.7,
Sonnet 4.6 and the Opus/Sonnet 4.5 family never get the newer
`_20260209` / `_20260120` variants. Those newer variants add dynamic
filtering (Claude writes code to rank/filter web results before they
enter context) and REPL state persistence + programmatic tool calling
inside the sandbox, which is what the user-facing pills are supposed
to expose.

Hardcoding the legacy versions also breaks if a future model family
drops the legacy types: the request 400s instead of falling back.

Changes:

- Add `_anthropic_web_search_version`, `_anthropic_web_fetch_version`,
  `_anthropic_code_execution_version` helpers that pick the newest
  variant the model accepts and fall back to the GA versions for
  everything else.
- Add `_ANTHROPIC_CODE_EXECUTION_BETA` constant since the beta header
  (`code-execution-2025-08-25`) is shared across both code-execution
  date variants per the upstream docs.
- Wire the helpers into `_stream_anthropic` so the outbound body
  carries the right pinned version per request.
- Add parametrized dispatch tests in
  `test_anthropic_tool_versions.py` covering Opus 4.7/4.6/4.5,
  Sonnet 4.6/4.5, Haiku 4.5, Opus 4.1/4.0, Sonnet 4.0, 3.5 Sonnet,
  plus streaming integration tests that verify the outbound body
  uses the right versions on Opus 4.7 (new web_search + new
  code_execution), Haiku 4.5 (legacy both), and Sonnet 4.5 (legacy
  web_search + new code_execution).
- Update existing `test_anthropic_code_execution.py` cases that
  pinned the old version on Opus 4.7 to expect the new ones.

Verified end-to-end against the live Anthropic API: Opus 4.7 with
both pills enabled accepts the newer-pinned tools without a 400, and
Haiku 4.5 still works on the legacy fallback path.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:03:27 -07:00
Daniel Han
ba9405b908
Studio: surface prompt-cache token counts in /v1/chat/completions usage chunk (#5670)
* Studio: surface prompt-cache token counts in /v1/chat/completions usage chunk

Studio's Anthropic and OpenAI Responses proxies already capture
cache_creation_input_tokens, cache_read_input_tokens (Anthropic) and
input_tokens_details.cached_tokens (OpenAI), but they were only written
to the structlog stream. Browser and SDK clients had no way to compute
"how many tokens hit the prompt cache" without scraping the server log,
so the chat UI could not show users how much money the cache was
saving on each turn.

This change emits one extra OpenAI include_usage-style chunk
(choices: [] with a populated usage block) just before the existing
[DONE] for Anthropic and after the final finish_reason chunk for
OpenAI Responses (both response.completed and response.incomplete).
The chunk shape:

  usage.prompt_tokens_details.cached_tokens
      normalised cache-read count, present for both providers.
  usage.cache_creation_input_tokens
      Anthropic-only; tokens billed at the cache-write premium.
  usage.cache_read_input_tokens
      Anthropic-only; same value as cached_tokens, kept for callers
      that already key off the native Anthropic name.

Smoke verified end to end against a live Studio (claude-haiku-4-5
and gpt-4o-mini) plus 7 new unit tests on the helper and the two
streaming paths.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Anthropic: include cache buckets in prompt_tokens / total_tokens

Anthropic's `input_tokens` field excludes the cache buckets -- the
real prompt size is `input_tokens + cache_creation_input_tokens +
cache_read_input_tokens`. Previously the new usage chunk reported
only `input_tokens` as `prompt_tokens`, which heavily undercounted
cache-hit turns (e.g. an 18.9k-token cache_read turn looked like an
8-token prompt) and broke any downstream context / cost display fed
by `prompt_tokens` or `total_tokens`.

Fix `_build_usage_chunk` to sum all three input buckets for the
Anthropic provider while keeping the OpenAI Responses path unchanged
(OpenAI already folds cached tokens into `input_tokens`). The native
`cache_creation_input_tokens` / `cache_read_input_tokens` keys and
`prompt_tokens_details.cached_tokens` mirror are still emitted, so
clients keep full visibility of the cache split.

Tests updated to assert the summed shape.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:02:52 -07:00