Codex 02:42Z on #5549 caught that flagging the post-tool empty-status
event with boundary=True (cycle-12 commit 610c387) double-handles the
cursor reset in the Anthropic streaming path.
AnthropicStreamEmitter._handle_tool_end already:
- closes the open tool_use block,
- emits tool_result,
- increments block_index,
- opens a fresh text block, and
- resets _prev_text = "".
When llama_cpp.py then yielded boundary=True on the very next event,
_handle_boundary fired _close_block + _open_text_block on that freshly
opened (still-empty) text block. Result: every tool call produced a
spurious content_block_stop + content_block_start pair before the
post-tool model text streamed.
Fix:
- llama_cpp.py: drop boundary=True from the post-tool status emit
(line ~4768). Keep boundary=True only at the auto-continue site
(line ~4500), which has no preceding tool_end to do the cursor
work.
- routes/inference.py OpenAI-compat tool stream: mirror the
Anthropic semantics by resetting prev_text on BOTH tool_start AND
tool_end, so the post-tool empty-status no longer needs to do it.
Add backend/tests/test_anthropic_messages.py::TestAnthropicStreamEmitter::
test_post_tool_empty_status_does_not_double_close as a regression
test: content -> tool_start -> tool_end -> empty status -> content
must not bump block_index past tool_end's increment, and the post-tool
content must land in the text block tool_end opened.
95 tests pass across the anthropic + trailing-plan suites.