* studio: report the true reasoning duration and fix the Stop button for thinking models For a local GGUF the "Thought for N" label was timed entirely on the client by a brittle edge-detector, so an always-think model (Qwen3 MTP) that buffers its whole reasoning and flushes it in one chunk showed "1 second" instead of the real minute-plus. The client cannot time reasoning it receives atomically, so make the timing backend-authoritative. Backend: generate_chat_completion_with_tools measures wall-clock reasoning and emits a Studio reasoning_summary event (duration_ms) at the moment reasoning ends -- the first answer token, or end-of-stream for a reasoning-only reply -- for both the tool-detection pass and the final-answer pass. Timing resets per tool iteration so the final answer's thinking time wins on the client (which takes the latest reasoning_summary). routes/inference.py forwards the event in the GGUF tool stream. Frontend: parse the reasoning_summary SSE into a _reasoningDurationMs chunk and use it as the authoritative reasoning duration (last write wins), clamped to >= 0 and guarded to a finite number so a malformed or proxied chunk cannot produce a NaN label; the persisted value wins for the final "Thought for N" label, with the previous live timer kept only as a fallback when no metadata arrives. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| public | ||
| src | ||
| .gitignore | ||
| .gitkeep | ||
| .npmrc | ||
| biome.json | ||
| components.json | ||
| data-designer.openapi (1).yaml | ||
| eslint.config.js | ||
| index.html | ||
| package-lock.json | ||
| package.json | ||
| tsconfig.app.json | ||
| tsconfig.json | ||
| tsconfig.node.json | ||
| vite.config.ts | ||