Compare commits

..

9 commits

Author SHA1 Message Date
Unsloth
df7300a3e1 Studio: harden the profile stats aggregation
Four fixes from review:

- _as_float raised OverflowError on a JSON integer wider than float, so
  one oversized counter returned 500 for the whole panel. It degrades to
  zero now, like every other unreadable field.
- Fork dedup elected one winner per source thread, but fork_chat_thread
  copies a single parent_id branch, so sibling forks of a retry and a
  regeneration hold different rows. Electing per original message keeps
  each branch when the source is deleted.
- create_run claims a resume source before the continuation logs its
  first step, so a continuation that failed early took the source's
  completed steps and tokens with it. Supersession now needs the
  continuation to have reached the source's step.
- Cumulative activity is a running total over the displayed window, but
  a narrow card trims older weeks without rebasing, opening the first
  visible bar at the hidden total and flattening the rest.
2026-07-29 02:34:24 -07:00
Unsloth
8644b4c208 Studio: drop the hour and weekday rhythm cards
Removes the "When you work" and "Your week" row from the Profile stats.

The backend stopped computing the hour and weekday buckets too, since
nothing else read them, and the payload no longer carries them. The
timezone and malformed-timestamp tests now assert on day buckets, which
is what the activity grid and streaks actually use; the daylight-saving
case moved to a timestamp where the hour of drift crosses midnight so
the day still proves it.

This was the only recharts consumer in the profile panel, so the lazy
split's note about keeping charts out of the main bundle no longer
applies and has been corrected. Build drops about 22 KB across two
fewer chunks.
2026-07-29 01:16:11 -07:00
Unsloth
485684db5f Studio: count fork clones per message and drop future streaks
- A pre-fork message can be pruned from its original thread while the
  clone stays in the fork. The suppression only checked that the source
  thread existed, so that message's tokens and activity vanished. Clones
  are now matched to the row they came from, and one fork per source is
  elected to stand in whenever the original is gone, whether the whole
  thread or just that message went.
- Future-dated history was filtered out of the current streak but still
  padded the longest streak and could be reported as the last active
  day. It is dropped before any streak field is computed.
- Cumulative mode labelled its tooltip "week of" while showing the
  running total. It now reports that week's tokens; the bar height still
  uses the running total.
- The activity grid and weekday axis formatted dates with the browser
  language rather than the language chosen in Settings, and the memos
  could not react to a change. They follow the app locale now.

The fork fixtures were also corrected: fork_chat_thread copies created_at
verbatim, so a clone carries the original's timestamp, which the previous
fixtures did not model.
2026-07-29 01:06:06 -07:00
Unsloth
f9970b8744 Studio: dedupe sibling forks and group comparison panes
- Deleting a thread that had been forked more than once left every
  sibling holding a copy of the shared ancestry, and the source-existence
  gate then let all of them count it, multiplying the usage by the number
  of forks. One fork per dead source is now elected to keep its copies,
  the one carrying the most, so nothing is lost or counted twice.
- Compare mode persists one thread per pane under a shared pair_id and
  the sidebar renders them as a single conversation. The chat total now
  keys on the pair, so one comparison is one chat and average tokens per
  chat is not halved.
- Corrected the note on cumulative mode: it is the running total across
  the displayed window, not lifetime. Seeding it with everything older
  than the window would flatten every bar against a baseline the grid
  has no room to show.
2026-07-29 00:45:12 -07:00
Unsloth
1338f09b97 Studio: harden profile stats against deleted and malformed history
- Deleting a forked conversation's source left the fork holding the only
  copies of those messages, but the skip still dropped them because
  forked_from_thread_id is not a foreign key. The suppression now joins
  the source row, so copies are only ignored while the originals survive.
- delete_run never clears the predecessor's resume_blocked, so removing
  a continuation stranded its source at zero steps and tokens while the
  run stayed visible. Supersession now requires a continuation that
  still exists.
- created_at is a client-supplied integer stored unchecked, so one
  out-of-range row made datetime raise and returned 500 for the whole
  panel. Those rows are now skipped for day and hour buckets.
- A future-dated row satisfied the current-streak window, reporting days
  that have not happened. The last active day must now be today or
  yesterday.
- Compact numbers rounded 999,999 to "1000K". Rounding into the next
  unit now steps up the suffix.
- Escape or a programmatic close unmounts the Profile tab without a
  blur, dropping an in-progress name edit. Drafts are committed on
  unmount; both saves no-op when unchanged.
- Stats copy no longer claims the numbers never leave the device, which
  is untrue when Studio is reached from another machine. It states what
  actually holds: nothing is collected or sent to Unsloth.
2026-07-29 00:10:13 -07:00
Unsloth
9ae3090b98 Studio: further profile stat corrections
Follow-up to the previous pass, each with a test that fails without it.

- Cancelling a run also sets resume_blocked, so filtering on that flag
  alone dropped the steps and tokens of cancelled work while still
  counting the run and its duration. Only a run a later run resumed
  from is excluded now: the resume claim leaves output_dir intact,
  cancelling clears it, which tells the two apart.
- A fixed getTimezoneOffset() was applied to every historical message,
  so winter records read an hour out when the panel is opened on summer
  time, moving messages near midnight onto the wrong day. The endpoint
  now takes an IANA timezone and converts each timestamp with its own
  offset, keeping the fixed offset as the fallback.
- A fork with no new turn was skipped before its thread was registered,
  so it did not appear in the chat total until its first message. The
  thread is now counted before the cloned rows are suppressed.
- Routed and aliased responses record the model that actually answered
  in responseDetails.responseModelId, while contextUsage.modelId stays
  the requested checkpoint. Most used models now prefers the former.
- Renamed runs showed the model label instead of the name the user
  chose. The name leads and the model moves beside the dataset.
2026-07-28 23:46:45 -07:00
Unsloth
6ff20ecdeb Studio: correct profile stat aggregation
Six accuracy fixes, each with a test that fails without it.

- firstTokenTime is already an elapsed duration, not a timestamp.
  Subtracting streamStartTime made the comparison false for every real
  message, so "Average time to first token" was always empty.
- Forking clones the whole ancestry into the new thread with the
  original timestamps. Those copies were counted again, doubling
  tokens, messages, attachments and activity for the source
  conversation. Rows older than the fork are now skipped.
- training_metrics.num_tokens is state.num_input_tokens_seen, a running
  total logged at each step, so summing the samples multiplied the real
  figure. Take each run's final counter, matching get_run_metrics.
- A resumed run continues its source's step and token counters from the
  checkpoint, so adding both reported the same progress twice. Only
  runs no later run resumed from are counted.
- Days, hours and weekdays were bucketed in the server's timezone while
  the client parses the keys as browser-local. The endpoint now takes
  the caller's getTimezoneOffset().
- A turn with no contextUsage.modelId fell back to the thread's
  model_id, which tracks the current selection and so misattributed
  older turns after a mid-conversation switch. Those turns are left
  uncredited instead.
2026-07-28 23:21:18 -07:00
pre-commit-ci[bot]
f9b4ca4ff0 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-29 05:56:01 +00:00
Unsloth
b8364e3445 Studio: add profile usage stats and tidy the personalization panel
Adds a "Your stats" section to the Profile settings tab, built entirely
from local history in studio.db. No telemetry, nothing uploaded.

Backend
- storage/profile_stats_db.py folds every metric in one streaming pass
  over chat_messages, memoised against a (count, max created_at)
  fingerprint so reopening the tab is free until history changes.
- routes/profile_stats.py serves GET /api/profile/stats from a worker
  thread, so a cold pass cannot stall token streaming.

Frontend
- Headline tiles, a token activity grid with daily/weekly/cumulative
  modes, activity insights, most used models, hour and weekday rhythms,
  and training run totals.
- The stats panel is lazy loaded so recharts stays out of the main
  bundle.
- Personalization panel now puts a larger avatar beside the two name
  fields, with picture options moved into an edit popover.
- "Sloth in greeting" moves to Chat defaults, next to the other chat
  toggles.

Tests: 9 backend tests including a regression test that the endpoint
does not block the event loop, plus formatting unit tests.
2026-07-28 22:54:28 -07:00
171 changed files with 4174 additions and 14890 deletions

View file

@ -17,8 +17,7 @@ if [ -n "${STUDIO_PERMISSION_FRONTEND:-}" ]; then
fi
mkdir -p "$artifact_dir"
# Wipe (not reset-password): the boot below must re-seed a fresh .bootstrap_password.
rm -rf "$studio_home/auth"
unsloth studio reset-password
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$port" "$@" \
>"$server_log" 2>&1 &
studio_pid=$!

View file

@ -167,9 +167,7 @@ jobs:
# ── boot the server under test (factored helper) ──────────────────
- name: Serve unsloth run --disable-tools (gemma-4-E4B)
run: |
# Wipe, not reset-password: since #7573 the reset rotates in place and
# prints the new passphrase, which would land unmasked in the job log.
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
bash .github/scripts/serve-unsloth-run.sh \
--gguf-file "$GITHUB_WORKSPACE/gguf-cache/${GGUF_FILE}" \
--port "$STUDIO_PORT" --log-dir logs \
@ -373,7 +371,7 @@ jobs:
- name: Serve unsloth run --disable-tools (gemma-4-E4B)
run: |
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
bash .github/scripts/serve-unsloth-run.sh \
--gguf-file "$GITHUB_WORKSPACE/gguf-cache/${GGUF_FILE}" \
--port "$STUDIO_PORT" --log-dir logs \
@ -556,7 +554,7 @@ jobs:
- name: Serve unsloth run --disable-tools (gemma-4-E4B)
run: |
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
bash .github/scripts/serve-unsloth-run.sh \
--gguf-file "$GITHUB_WORKSPACE/gguf-cache/${GGUF_FILE}" \
--port "$STUDIO_PORT" --log-dir logs \
@ -720,7 +718,7 @@ jobs:
- name: Serve unsloth run --disable-tools (gemma-3-270m)
run: |
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
bash .github/scripts/serve-unsloth-run.sh \
--model "$GGUF_REPO" --gguf-variant "$GGUF_VARIANT" \
--port "$STUDIO_PORT" --log-dir logs \

View file

@ -766,7 +766,6 @@ jobs:
env:
GH_REPO: ${{ github.repository }}
APP_VERSION: ${{ needs.prepare-version.outputs.app_version }}
PYPI_VERSION: ${{ needs.prepare-version.outputs.pypi_version }}
STUDIO_VERSION: ${{ needs.prepare-version.outputs.studio_version }}
DESKTOP_RELEASE_TAG: ${{ needs.prepare-version.outputs.desktop_release_tag }}
DESKTOP_PRERELEASE: ${{ needs.prepare-version.outputs.prerelease }}
@ -912,8 +911,6 @@ jobs:
notes = pathlib.Path(os.environ['RUNNER_TEMP'], 'desktop-release-notes.md').read_text()
metadata = {
'version': os.environ['APP_VERSION'],
# App version is SemVer; CHANGELOG.md is keyed by the backend release.
'pypi_version': os.environ['PYPI_VERSION'],
'notes': notes,
'pub_date': datetime.datetime.now(datetime.timezone.utc).isoformat(timespec='milliseconds').replace('+00:00', 'Z'),
'platforms': {

View file

@ -1,156 +0,0 @@
# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved.
# Measures where Studio's startup time goes, on each platform.
#
# Nothing recorded a number before: main.py logs "lifespan startup completed in X ms"
# and studio_test_kit polls /healthz, but both throw the elapsed time away. A first
# local run (Linux, warm cache, 18-core server) put `import main` at 5.7-6.6s BEFORE
# the server can bind, dominated by eager module-level imports pulled in by routes:
# torch ~1.9s self, unsloth_zoo ~0.8s, routes ~0.6s, transformers ~0.5s.
#
# Not a gate yet: --max-healthz-seconds exists, but a budget should come from
# observed numbers rather than a guess.
name: Startup profile
on:
pull_request:
paths:
# The measured import graph is the whole backend tree: main.py imports auth,
# core, hub, loggers, models, picker, routes and utils at module scope.
- 'studio/backend/**'
- '!studio/backend/tests/**'
# The launch phase spawns `unsloth studio --api-only`, so the CLI counts too.
- 'unsloth_cli/**'
- 'studio/src-tauri/src/preflight**'
# The profiler hardcodes the desktop argv that process.rs::backend_args builds,
# so a change there must schedule a run or the two silently diverge.
- 'studio/src-tauri/src/process.rs'
- 'scripts/profile_startup.py'
- '.github/workflows/startup-profile-ci.yml'
# The job profiles whatever `install.sh --local` built: the installers pick the
# venv's Python and the dependency specs, and pyproject's include list is what
# makes --local overlay studio.backend*.
- 'install.sh'
- 'install.ps1'
- 'pyproject.toml'
# --local also runs the checkout's setup scripts (install.sh picks
# $_REPO_ROOT/studio/setup.sh, the editable install resolves setup.ps1 to the
# repo), and both call install_python_stack.py, which picks the dependencies.
- 'studio/setup.sh'
- 'studio/setup.ps1'
- 'studio/install_python_stack.py'
workflow_dispatch:
inputs:
repeats:
description: 'launch repeats per OS (median reported)'
type: string
default: '3'
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
permissions:
contents: read
jobs:
profile:
name: startup ${{ matrix.os }}
runs-on: ${{ matrix.os }}
timeout-minutes: 60
continue-on-error: true
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, macos-14, windows-latest]
env:
UNSLOTH_STUDIO_HOME: ${{ github.workspace }}/.studio-home
# A wildcard bind calls ifconfig.me on the startup path; loopback times our code.
UNSLOTH_STUDIO_DISABLE_PUBLIC_CHECK: '1'
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
- name: Install Studio
shell: bash
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
set -o pipefail
mkdir -p logs
# --local is load-bearing: it overlays the checkout, so the profiled server
# is this diff. Without it install.sh resolves unsloth from PyPI.
if [ "${{ runner.os }}" = "Windows" ]; then
pwsh -NoProfile -File ./install.ps1 --local 2>&1 | tee logs/install.log
else
bash install.sh --local 2>&1 | tee logs/install.log
fi
- name: Profile startup
shell: bash
run: |
BIN="$UNSLOTH_STUDIO_HOME/unsloth_studio/bin/unsloth"
[ -x "$BIN" ] || BIN="$UNSLOTH_STUDIO_HOME/unsloth_studio/Scripts/unsloth.exe"
[ -x "$BIN" ] || BIN=""
# Profile imports with the INSTALLED interpreter: that venv is what launches.
PY="$UNSLOTH_STUDIO_HOME/unsloth_studio/bin/python"
[ -x "$PY" ] || PY="$UNSLOTH_STUDIO_HOME/unsloth_studio/Scripts/python.exe"
[ -x "$PY" ] || PY="$(command -v python3 || command -v python)"
python3 scripts/profile_startup.py \
--python "$PY" \
${BIN:+--bin "$BIN"} \
--repeats "${{ inputs.repeats || '3' }}" \
--json "startup-${{ matrix.os }}.json" 2>&1 | tee logs/profile.log
- name: Summary
if: always()
shell: bash
run: |
f="startup-${{ matrix.os }}.json"
[ -f "$f" ] || { echo "no profile produced"; exit 0; }
python3 - "$f" >> "$GITHUB_STEP_SUMMARY" <<'PY'
import json, sys
d = json.load(open(sys.argv[1]))
print(f"### {d['platform']} / {d['machine']} (py {d['python']}, {d['cpu_count']} cpu)\n")
imp = d.get("imports", {})
# Gate on ok: a failed `import main` still leaves rows, so a total can lie.
if imp.get("ok"):
print(f"**`import main`: {imp['total_seconds']}s**\n")
print("| package | self ms |")
print("|---|---:|")
for k, v in list(imp.get("self_by_package_ms", {}).items())[:8]:
print(f"| {k} | {v} |")
print()
else:
print("**`import main` failed - no valid import profile**\n")
print("```\n" + (imp.get("error") or "")[-1500:] + "\n```\n")
lau = d.get("launch") or {}
runs = len(lau.get("runs") or [])
failed = lau.get("failed_runs") or 0
if lau.get("healthz_median_seconds") is not None:
# The aggregates cover only the runs that reached healthz, so flag the
# failures: bare numbers would read as a normal fast startup.
note = f" _({runs - failed} of {runs} launches; {failed} never became healthy)_" if failed else ""
print(f"**time to a healthy port: {lau['healthz_median_seconds']}s median, "
f"{lau['healthz_max_seconds']}s max**{note}\n")
elif lau.get("skipped"):
print(f"_launch phase skipped: {lau['skipped']}_\n")
elif runs:
print(f"**no launch measurement: all {runs} launches failed to become healthy**\n")
PY
- name: Upload profile
if: always()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: startup-profile-${{ matrix.os }}
path: |
startup-*.json
logs/
retention-days: 14
if-no-files-found: warn

View file

@ -113,8 +113,7 @@ jobs:
- name: Reset auth + boot Unsloth (API-only)
run: |
# Wipe (not reset-password): the boot below must re-seed a fresh .bootstrap_password.
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> logs/studio.log 2>&1 &

View file

@ -223,16 +223,6 @@ jobs:
tests/studio/test_is_mlx_dispatch_gate.py \
tests/studio/test_xpu_spoof_pipeline.py
- name: CLI tests (unsloth_cli)
# unsloth_cli/tests had no CI at all: `unsloth_cli/**` was only a paths
# trigger and a ruff target, so 673 tests covering the studio launcher,
# the pre-exposure gate and the auth secret writers ran nowhere, and
# four of them had been failing on main unnoticed.
# Own step, not folded into the tests/ discovery above: pyproject's
# testpaths is tests/, and this suite needs no PYTHONPATH or CUDA spoof
# (it self-bootstraps sys.path and imports neither unsloth nor torch).
run: python -m pytest unsloth_cli/tests -q --tb=short
- name: Shell installer tests
# Auto-discovered rather than allowlisted. The old hardcoded list had
# silently fallen seven files behind tests/run_all.sh, including

View file

@ -127,8 +127,7 @@ jobs:
- name: Reset auth + boot Unsloth (API-only)
run: |
# Wipe (not reset-password): the boot below must re-seed a fresh .bootstrap_password.
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> logs/studio.log 2>&1 &
@ -401,7 +400,7 @@ jobs:
# tool_policy=None so each request's `enable_tools` field is
# honoured.
run: |
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> logs/studio.log 2>&1 &
@ -979,7 +978,7 @@ jobs:
# response_format requests aren't routed through the agentic
# tool loop.
run: |
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> logs/studio.log 2>&1 &

View file

@ -101,8 +101,7 @@ jobs:
- name: Reset auth + boot Unsloth (API-only)
run: |
# Wipe (not reset-password): the boot below must re-seed a fresh .bootstrap_password.
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> logs/studio.log 2>&1 &

View file

@ -126,8 +126,7 @@ jobs:
- name: Reset auth + boot Unsloth (API-only)
run: |
# Wipe (not reset-password): the boot below must re-seed a fresh .bootstrap_password.
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> logs/studio.log 2>&1 &
@ -387,7 +386,7 @@ jobs:
# tool_policy=None so each request's `enable_tools` field is
# honoured.
run: |
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> logs/studio.log 2>&1 &
@ -832,7 +831,7 @@ jobs:
# response_format requests aren't routed through the agentic
# tool loop.
run: |
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> logs/studio.log 2>&1 &

View file

@ -146,8 +146,7 @@ jobs:
- name: Reset auth + boot Unsloth
run: |
# Wipe (not reset-password): the boot below must re-seed a fresh .bootstrap_password.
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> logs/studio.log 2>&1 &
@ -191,7 +190,7 @@ jobs:
# runner's kernel briefly runs out of socket buffers, and (3) a
# goto 'interrupted by another navigation' when the SPA auth
# guard redirects mid-navigation. The retry FULLY resets Unsloth
# (kill, wipe auth, reboot, wait /api/health, re-export
# (kill, reset-password, reboot, wait /api/health, re-export
# bootstrap pw) before re-running the script. A real test failure
# (assertion / timeout) does NOT match any pattern so it bypasses
# retry and surfaces immediately.
@ -214,7 +213,7 @@ jobs:
echo "::warning::Playwright flake on attempt ${attempt}; resetting Unsloth and retrying..."
kill "${STUDIO_PID}" 2>/dev/null || true
sleep 2
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> "logs/studio_retry_${attempt}.log" 2>&1 &
STUDIO_PID=$!
@ -252,7 +251,7 @@ jobs:
- name: Reset auth + boot Unsloth for extra UI tests (port 18897)
run: |
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p 18897 \
> logs/studio_extra.log 2>&1 &
@ -309,7 +308,7 @@ jobs:
echo "::warning::Playwright flake on attempt ${attempt}; resetting Unsloth and retrying..."
kill "${STUDIO_EXTRA_PID}" 2>/dev/null || true
sleep 2
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p 18897 \
> "logs/studio_extra_retry_${attempt}.log" 2>&1 &
STUDIO_EXTRA_PID=$!

View file

@ -115,8 +115,7 @@ jobs:
- name: Reset auth + boot Unsloth
run: |
# Wipe (not reset-password): the boot below must re-seed a fresh .bootstrap_password.
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> logs/studio.log 2>&1 &
@ -194,7 +193,7 @@ jobs:
# warm install we already did) so this adds little wall time.
- name: Reset auth + boot Unsloth for extra UI tests (port 18894)
run: |
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p 18894 \
> logs/studio_extra.log 2>&1 &
@ -254,7 +253,7 @@ jobs:
# (RAG embedder + llama.cpp probe) stay hidden from the picker.
- name: Reset auth + boot Unsloth for model-config tests (port 18898)
run: |
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p 18898 \
> logs/studio_modelcfg.log 2>&1 &
@ -300,7 +299,7 @@ jobs:
# earlier UI tests. No GGUF -- the bug surface is the composer.
- name: Reset auth + boot Unsloth for IME / i18n tests (port 18896)
run: |
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p 18896 \
> logs/studio_ime.log 2>&1 &

View file

@ -179,8 +179,7 @@ jobs:
- name: Reset auth + boot Unsloth (API-only)
run: |
# Wipe (not reset-password): the boot below must re-seed a fresh .bootstrap_password.
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> logs/studio.log 2>&1 &

View file

@ -229,8 +229,7 @@ jobs:
- name: Reset auth + boot Unsloth (API-only)
run: |
# Wipe (not reset-password): the boot below must re-seed a fresh .bootstrap_password.
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> logs/studio.log 2>&1 &
@ -574,7 +573,7 @@ jobs:
- name: Reset auth + boot Unsloth (API-only, default tool policy)
run: |
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> logs/studio.log 2>&1 &
@ -1075,7 +1074,7 @@ jobs:
- name: Reset auth + boot Unsloth (API-only)
run: |
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> logs/studio.log 2>&1 &
@ -1547,7 +1546,7 @@ jobs:
- name: Reset auth + boot Unsloth (API-only)
run: |
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> logs/studio.log 2>&1 &
@ -1889,11 +1888,8 @@ jobs:
# (step/substep -> Write-StudioStdoutMirror / Get-StudioAnsi).
$script:StudioVtOk = $false
$script:UnslothVerbose = $false
# Get-HostMachineArch is reached only on the absent path, where
# Test-VCRedistInstalled consults it before trusting the System32 DLL, so
# part A passes without it and only the clean-box part fails.
foreach ($fn in @('Get-StudioAnsi', 'Write-StudioStdoutMirror', 'step', 'substep',
'Invoke-SetupCommand', 'Refresh-Environment', 'Get-HostMachineArch',
'Invoke-SetupCommand', 'Refresh-Environment',
'Test-VCRedistInstalled', 'Ensure-VCRedist')) {
$src = Get-FunctionSource -Path $setup -Name $fn
if (-not $src) { throw "Function '$fn' not found in setup.ps1" }

View file

@ -297,8 +297,7 @@ jobs:
- name: Reset auth + boot Unsloth
run: |
# Wipe (not reset-password): the boot below must re-seed a fresh .bootstrap_password.
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p "$STUDIO_PORT" \
> logs/studio.log 2>&1 &
@ -353,7 +352,7 @@ jobs:
- name: Reset auth + boot Unsloth for extra UI tests (port 18897)
run: |
rm -rf ~/.unsloth/studio/auth
unsloth studio reset-password
mkdir -p logs
UNSLOTH_API_ONLY=1 unsloth studio -H 127.0.0.1 -p 18897 \
> logs/studio_extra.log 2>&1 &

3
.gitignore vendored
View file

@ -208,9 +208,6 @@ tmp/
**/node_modules/
auth.db
# Packaging snapshot of the root CHANGELOG.md (written by build.sh)
studio/CHANGELOG.md
# Tauri local build/generated output
studio/src-tauri/target/
studio/src-tauri/gen/

View file

@ -1,88 +0,0 @@
# Changelog
Release notes for Unsloth and Unsloth Studio.
Unsloth Studio reads this file to show release notes inside the "New Unsloth
version" update popup. Edit it here and the popup picks the change up on the
next update check, with no release or rebuild required.
## Format
Every release is a level-2 heading whose first token is the version, optionally
followed by a date:
```md
## 2026.7.6 - 2026-07-22
```
`## [2026.7.6] - 2026-07-22` and `## v2026.7.6` also work. Everything under a
heading, up to the next level-2 heading, is that release's notes and renders as
Markdown in the popup.
Notes are matched to one exact version. When Studio offers an update to
`2026.7.6` it renders the `2026.7.6` section and nothing else. If that section
is missing, the popup links out to the online changelog rather than showing
notes from an unrelated release, so a new version needs its own section here
before its notes can appear.
Keep the newest release at the top. Lead each bullet with the change itself:
the collapsed popup highlights the first sentence and dims the rest.
`## Unreleased` is ignored by the popup, so it is safe to stage notes there and
rename the heading at release time.
<!-- Add new releases directly below this line. -->
## Unreleased
## 2026.7.5
### What's Changed
- AMD support is here. Train, run RL, chat with and deploy 500+ models on
Radeon, Instinct, Ryzen and data center GPUs across Windows, WSL and Linux,
up to 2x faster with 70% less VRAM and no accuracy loss.
- Intel XPU support lands in Studio, so Arc and Data Center GPUs run chat and
training alongside the NVIDIA, AMD and Apple paths.
- Local speech to text dictation runs fully offline, with slim Whisper bundles
and a picker for custom models.
- DoRA training is available in Studio, selectable next to LoRA and full
fine-tuning in the training tab.
- The update popup previews release notes inline, pulled from this file and
matched to the exact version being offered.
### AMD, 23 July update
Our AMD collaboration, custom Triton kernels and math algorithms bring local
training and inference to AMD hardware. The 23 July update builds on the
[AMD release](https://github.com/unslothai/unsloth/releases/tag/v0.1.501-beta):
- RDNA2 and Gorgon Halo are supported, and the installer no longer fails to
detect GPUs on Strix Halo and other AMD cards.
- RDNA4 handling is better, and HIP and ROCm failures are caught and fixed
automatically instead of stopping the install.
- Unified memory safetensors loading is 2x faster, with much faster gradient
checkpointing on unified memory devices.
- Voice dictation through whisper.cpp has preliminary support.
- Rollback environments left by installs no longer eat 5GB of disk. They are
cleaned up automatically.
Optimized ROCm builds cover GGUF and safetensors inference, and ROCm
compatibility is improved for MI300X and MI325X. Full guide:
[unsloth.ai/docs/basics/amd](https://unsloth.ai/docs/basics/amd).
### Running larger models
- Automatic GPU placement, or pick exactly which GPUs and layers to use.
- Move MoE expert layers into system memory so larger models fit.
- Split a model across several GPUs, or use tensor parallelism.
- Hardware settings are saved per model and quant.
### Also in this release
- Remote access with `unsloth studio --secure` over free HTTPS via Cloudflare.
- Web search reads PDF papers and manuals, and parallel tool calls, reasoning
output and tool retries are more reliable.
- The model download location is configurable, so weights can live on a second
drive instead of the default cache.
- Stalled Hugging Face XET downloads retry over standard HTTP, and existing
GGUF files are reused instead of downloaded again.

View file

@ -1,2 +0,0 @@
include _changelog_build.py
include CHANGELOG.md

View file

@ -1,36 +0,0 @@
# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved.
"""Snapshot CHANGELOG.md into the studio package at build time.
CHANGELOG.md at the repo root stays the one file to edit. Copying it here,
rather than in build.sh, means every packaging path ships it, so release notes
still render when the popup cannot reach GitHub."""
from __future__ import annotations
import shutil
from pathlib import Path
from setuptools.command.build_py import build_py as _build_py
ROOT = Path(__file__).resolve().parent
SOURCE = ROOT / "CHANGELOG.md"
SNAPSHOT = ROOT / "studio" / "CHANGELOG.md"
class build_py(_build_py):
def run(self) -> None:
# Beside the sources only if writable (PEP 517 may build an immutable
# checkout); into the staging directory always.
if SOURCE.is_file():
try:
shutil.copyfile(SOURCE, SNAPSHOT)
except OSError:
pass
super().run()
if not SOURCE.is_file():
return
staged = Path(self.build_lib) / "studio" / "CHANGELOG.md"
staged.parent.mkdir(parents = True, exist_ok = True)
shutil.copyfile(SOURCE, staged)

View file

@ -103,13 +103,9 @@ else
STUDIO_STAMPED_VERSION="$(python scripts/stamp_studio_release.py)"
fi
# 4. Build wheel/sdist. _changelog_build.py snapshots CHANGELOG.md into the studio
# package so release notes render offline.
# 4. Build wheel/sdist
python -m build
# Drop the snapshot so a source checkout never serves a stale copy.
rm -f studio/CHANGELOG.md
if [ "${1:-}" = "publish" ]; then
python scripts/stamp_studio_release.py --verify-dist dist --expected "$STUDIO_STAMPED_VERSION"
fi

View file

@ -57,26 +57,6 @@ function Install-UnslothStudio {
}
}
# Machine arch; Get-TauriDiagArch above reports the process. An emulated x64 shell on
# ARM64 reports AMD64, but PROCESSOR_ARCHITEW6432 is ARM64 in exactly that case.
function Get-HostMachineArch {
$osArch = ""
try { $osArch = [System.Runtime.InteropServices.RuntimeInformation]::OSArchitecture.ToString() } catch { $osArch = "" }
$signals = @([string]$env:PROCESSOR_ARCHITEW6432, [string]$env:PROCESSOR_ARCHITECTURE, $osArch)
foreach ($s in $signals) {
if ($s.ToLowerInvariant() -eq "arm64") { return "arm64" }
}
foreach ($s in $signals) {
if ([string]::IsNullOrWhiteSpace($s)) { continue }
switch ($s.ToLowerInvariant()) {
"amd64" { return "x86_64" }
"x64" { return "x86_64" }
"x86" { return "x86" }
}
}
return "unknown"
}
function Get-TauriTorchIndexFamily {
param([string]$TorchIndexUrl)
if ($SkipTorch) { return "none" }
@ -1144,27 +1124,10 @@ exit 0
return $false
}
# The interpreter's own arch, asked of it: win-amd64|win-arm64|win32|"".
function Get-PythonPlatformTag {
param([string]$Exe)
try {
return (& $Exe -c "import sysconfig; print(sysconfig.get_platform())" 2>$null | Out-String).Trim().ToLowerInvariant()
} catch { return "" }
}
# Returns @{ Version = "3.13"; Path = "C:\...\python.exe" } or $null.
# The resolved Path is passed to `uv venv --python` to prevent uv from
# re-resolving the version string back to a conda interpreter.
function Find-CompatiblePython {
# -X64Only: best installed x64 interpreter or $null, never ARM64. Last resort for
# Install-X64Python, where x64 of a lower-priority minor beats ARM64.
param([switch]$X64Only)
# Windows on ARM: prefer x64. pyarrow (via datasets) and hf-transfer ship no
# win_arm64 wheel, so a native ARM64 Python source-builds both and dies on CMake /
# Rust minutes in; x64 runs fine emulated. ARM64 is still returned when it is all
# there is, and the caller then bootstraps x64 or warns.
$preferX64 = $X64Only -or ((Get-HostMachineArch) -eq "arm64")
$candidates = @()
# Try the Python Launcher first (most reliable on Windows)
# py.exe resolves to the standard CPython install, not conda.
# Prefer the requested $PythonVersion, then newest-first fallback.
@ -1182,8 +1145,7 @@ exit 0
# Resolve the actual executable path and verify it is not conda-based
$resolvedExe = (& $pyLauncher.Source "-$minor" -c "import sys; print(sys.executable)" 2>$null | Out-String).Trim()
if ($resolvedExe -and (Test-Path $resolvedExe) -and -not (Test-IsCondaPython $resolvedExe)) {
if (-not $preferX64) { return @{ Version = $ver; Path = $resolvedExe; Arch = "" } }
$candidates += @{ Version = $ver; Path = $resolvedExe }
return @{ Version = $ver; Path = $resolvedExe }
}
}
} catch {}
@ -1204,53 +1166,11 @@ exit 0
try {
$out = & $cmd.Source --version 2>&1 | Out-String
if ($out -match "Python (3\.1[1-3])\.\d+") {
if (-not $preferX64) { return @{ Version = $Matches[1]; Path = $cmd.Source; Arch = "" } }
$candidates += @{ Version = $Matches[1]; Path = $cmd.Source }
return @{ Version = $Matches[1]; Path = $cmd.Source }
}
} catch {}
}
}
# `py -3.12` runs the launcher's preferred build, normally the native ARM64 one, so
# a same-minor x64 install that is neither preferred nor on PATH never becomes a
# candidate. `-3.12-64` cannot disambiguate (deprecated, it only means "not
# 32-bit"), so enumerate every registration with -0p and probe each path.
if ($preferX64) {
foreach ($pyLauncher in @(Get-Command py -All -CommandType Application -ErrorAction SilentlyContinue)) {
if ($pyLauncher.Source -match $script:CondaSkipPattern) { continue }
$listed = @()
try { $listed = @(& $pyLauncher.Source "-0p" 2>$null) } catch {}
foreach ($line in $listed) {
# " -V:3.12 * C:\...\python.exe": tag, optional default marker, path.
$m = [regex]::Match([string]$line, '(?i)^\s*-\S+\s+\*?\s*"?(?<p>\S.*?\.exe)"?\s*$')
if (-not $m.Success) { continue }
$exe = $m.Groups['p'].Value.Trim()
if ($candidates | Where-Object { $_.Path -eq $exe }) { continue }
if (-not (Test-Path -LiteralPath $exe)) { continue }
if (Test-IsCondaPython $exe) { continue }
try {
$out = & $exe --version 2>&1 | Out-String
if ($out -match "Python (3\.1[1-3])\.\d+") {
$candidates += @{ Version = $Matches[1]; Path = $exe }
}
} catch {}
}
}
}
# Prefer x64, but only within one minor: $minors is the caller's version preference,
# so ranking on arch alone would answer UNSLOTH_PYTHON=3.12 with an x64 3.13 and
# never bootstrap x64 3.12. Probing costs a subprocess, so non-ARM returned above.
foreach ($c in $candidates) {
$tag = Get-PythonPlatformTag $c.Path
$c.Arch = if ($tag -eq "win-amd64") { "x86_64" } elseif ($tag -eq "win-arm64") { "arm64" } else { "unknown" }
}
foreach ($minor in $minors) {
$sameMinor = @($candidates | Where-Object { $_.Version -eq $minor })
if ($sameMinor.Count -eq 0) { continue }
$x64 = $sameMinor | Where-Object { $_.Arch -eq "x86_64" } | Select-Object -First 1
if ($x64) { return $x64 }
if (-not $X64Only) { return $sameMinor[0] }
}
if (-not $X64Only -and $candidates.Count -gt 0) { return $candidates[0] }
return $null
}
@ -1261,11 +1181,8 @@ exit 0
# (no UAC), putting python.exe + the py launcher on PATH. Mirrors the uv ->
# astral.sh fallback below. Returns @{ Version; Path } or $null.
function Install-PythonFromPythonOrg {
# $Arch overrides the host arch, to pull x64 onto an ARM64 box.
param([string]$Arch = "")
# python.org ships one installer per architecture.
$targetArch = if ($Arch) { $Arch } else { Get-TauriDiagArch }
$archSuffix = switch ($targetArch) {
$archSuffix = switch (Get-TauriDiagArch) {
"x86_64" { "-amd64" }
"arm64" { "-arm64" }
"x86" { "" }
@ -1330,28 +1247,6 @@ exit 0
return (Find-CompatiblePython)
}
# ── Windows on ARM: get an x64 CPython ──
# --architecture x64 forces winget off the ARM64 build; python.org takes the same override.
function Install-X64Python {
if ($script:WingetAvailable) {
$prevEAP = $ErrorActionPreference
$ErrorActionPreference = "Continue"
try {
winget install -e --id "Python.Python.$PythonVersion" --source winget --architecture x64 --accept-package-agreements --accept-source-agreements
} catch { }
$ErrorActionPreference = $prevEAP
Refresh-SessionPath
$found = Find-CompatiblePython
if ($found -and $found.Arch -eq "x86_64") { return $found }
substep "winget could not provide an x64 Python -- trying python.org..." "Yellow"
}
$found = Install-PythonFromPythonOrg -Arch "x86_64"
if ($found -and $found.Arch -eq "x86_64") { return $found }
# Nothing installable (offline / no winget): an x64 build of another supported minor
# still runs the wheels ARM64 cannot, so take it over the native interpreter.
return (Find-CompatiblePython -X64Only)
}
# ── Install Python if no compatible version (3.11-3.13) found ──
# Find-CompatiblePython returns @{ Version = "3.13"; Path = "C:\...\python.exe" } or $null.
Write-TauriLog "STEP" "Installing Python"
@ -1423,26 +1318,6 @@ exit 0
return (Exit-InstallFailure "Python installation failed")
}
}
# ── Windows on ARM: swap a native ARM64 interpreter for x64 ──
# pyarrow and hf-transfer publish no win_arm64 wheel, so an ARM64 Python source-builds
# both and fails deep into the run. Warn up front if x64 is unobtainable.
if ($DetectedPython -and (Get-HostMachineArch) -eq "arm64" -and $DetectedPython.Arch -ne "x86_64") {
substep "windows on arm: only a native ARM64 Python $($DetectedPython.Version) was found." "Yellow"
substep "pyarrow and hf-transfer publish no win_arm64 wheels, so installing x64 Python..." "Yellow"
$X64Python = Install-X64Python
if ($X64Python) {
$DetectedPython = $X64Python
step "python" "using x64 Python $($DetectedPython.Version) under emulation"
} else {
Write-Host "[WARN] Could not install an x64 Python on this ARM64 machine." -ForegroundColor Yellow
Write-Host " Continuing with ARM64 Python $($DetectedPython.Version), but the install is likely to fail:" -ForegroundColor Yellow
Write-Host " pyarrow (via datasets) and hf-transfer ship no win_arm64 wheels and will be" -ForegroundColor Yellow
Write-Host " built from source, which needs CMake plus the MSVC and Rust toolchains." -ForegroundColor Yellow
Write-Host " Fix: install x64 Python from https://www.python.org/downloads/windows/" -ForegroundColor Yellow
Write-Host " (choose 'Windows installer (64-bit)', not ARM64), then re-run this installer." -ForegroundColor Yellow
}
}
$DiagPythonVersion = $PythonVersion
if ($DetectedPython) { $DiagPythonVersion = $DetectedPython.Version }
$InitialGpuBranch = "unknown"
@ -2563,13 +2438,6 @@ exit 0
}
} else {
Write-TauriLog "STEP" "Installing PyTorch"
# Windows on ARM lacks only torchaudio (whl/cpu win_arm64: torch 42,
# torchvision 60, torchaudio 0), so drop that pin instead of aborting. Ask the
# interpreter, not PROCESSOR_ARCHITECTURE; reached when no x64 Python exists.
$VenvPlatform = ""
try {
$VenvPlatform = (& $VenvPython -c "import sysconfig; print(sysconfig.get_platform())" 2>$null | Out-String).Trim().ToLowerInvariant()
} catch { $VenvPlatform = "" }
substep "installing PyTorch ($(Remove-IndexUrlCredentials $TorchIndexUrl))..."
# Bound the companions to the capped torch on EVERY index, cu<digits>
# families included: torchaudio 2.11 dropped its exact torch pin from
@ -2577,13 +2445,7 @@ exit 0
# resolve a mismatched 2.11.0 build. Mirrors install.sh.
$_pinVisionSpec = "torchvision>=0.19,<0.26.0"
$_pinAudioSpec = "torchaudio>=2.4,<2.11.0"
$_torchSpecs = @("torch>=2.4,<2.11.0", $_pinVisionSpec, $_pinAudioSpec)
if ($VenvPlatform -eq "win-arm64") {
substep "windows on arm: skipping torchaudio (upstream publishes no"
substep "win_arm64 wheel); torch and torchvision install normally."
$_torchSpecs = @("torch>=2.4,<2.11.0", $_pinVisionSpec)
}
$torchInstallExit = Invoke-InstallCommandRetry -Label "install PyTorch" { uv pip install --python $VenvPython @_torchSpecs --default-index $TorchIndexUrl }
$torchInstallExit = Invoke-InstallCommandRetry -Label "install PyTorch" { uv pip install --python $VenvPython "torch>=2.4,<2.11.0" $_pinVisionSpec $_pinAudioSpec --default-index $TorchIndexUrl }
if ($torchInstallExit -ne 0) {
Write-Host "[ERROR] Failed to install PyTorch (exit code $torchInstallExit)" -ForegroundColor Red
return (Exit-InstallFailure "Failed to install PyTorch (exit code $torchInstallExit)" $torchInstallExit)

View file

@ -19,17 +19,6 @@
# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
set -e
# ── Why the installer lives in a function ──
# Under `curl ... | sh`, sh is the pipe READER. This file is ~150KB, so a top-level
# `exit` left most of it unread, the write end failed, and curl tacked
# "(56) Failure writing output to destination" onto our own error message. Wrapping
# the body forces sh to parse to the closing brace first, so the pipe always drains
# (install.ps1 has always had this shape).
#
# Body is deliberately NOT reindented: reflowing 4000+ lines would bury the change,
# and `exit` still exits the shell from inside a function. Do not add
# `exec < /dev/null`: for a piped shell that closes the script's own source.
_unsloth_main() {
# ── Output style (aligned with studio/setup.sh) ──
RULE=""
@ -4458,8 +4447,3 @@ else
substep "(add -H 0.0.0.0 --cloudflare for a public Cloudflare HTTPS link, or --secure to keep the raw port private; anyone with the API key can run code)"
echo ""
fi
}
# Every byte above is parsed before this line runs, which is the point.
_unsloth_main "$@"

View file

@ -47,14 +47,9 @@ version = {attr = "unsloth.models._utils.__version__"}
[tool.setuptools]
include-package-data = true
[tool.setuptools.cmdclass]
# Snapshots CHANGELOG.md into studio/ so every build path ships it.
build_py = "_changelog_build.build_py"
[tool.setuptools.package-data]
unsloth_cli = ["codex_fallback_prompt.md", "pi_subagent.ts"]
studio = [
"CHANGELOG.md",
"*.sh",
"*.ps1",
"*.bat",
@ -133,19 +128,14 @@ huggingfacenotorch = [
]
# torchcodec backend for Gemma audio / datasets>=4 (#7225).
# Pick the audio-torch* pin matching your torch minor (see TORCH_TORCHCODEC).
# torchcodec publishes no sdist and only manylinux_2_28_x86_64, macosx_*_arm64
# and win_amd64 wheels, so Linux aarch64, Windows ARM64 and Intel Mac have
# nothing to resolve and pip fails the whole install rather than skipping audio.
# Gate on the platforms that have a wheel, matching
# PLATFORM_LACKS_TORCHCODEC_WHEEL in studio/install_python_stack.py.
audio-torch210 = [
"torchcodec>=0.10.0,<0.11.0 ; python_version >= '3.10' and (((sys_platform == 'linux' or sys_platform == 'win32') and (platform_machine == 'x86_64' or platform_machine == 'AMD64')) or (sys_platform == 'darwin' and platform_machine == 'arm64'))",
"torchcodec>=0.10.0,<0.11.0 ; python_version >= '3.10'",
]
audio-torch290 = [
"torchcodec>=0.8.0,<0.10.0 ; python_version >= '3.10' and (((sys_platform == 'linux' or sys_platform == 'win32') and (platform_machine == 'x86_64' or platform_machine == 'AMD64')) or (sys_platform == 'darwin' and platform_machine == 'arm64'))",
"torchcodec>=0.8.0,<0.10.0 ; python_version >= '3.10'",
]
audio-torch280 = [
"torchcodec>=0.6.0,<0.8.0 ; python_version >= '3.9' and (((sys_platform == 'linux' or sys_platform == 'win32') and (platform_machine == 'x86_64' or platform_machine == 'AMD64')) or (sys_platform == 'darwin' and platform_machine == 'arm64'))",
"torchcodec>=0.6.0,<0.8.0 ; python_version >= '3.9'",
]
huggingface = [
"unsloth[huggingfacenotorch]",

View file

@ -1,377 +0,0 @@
#!/usr/bin/env python3
# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""Measure where Unsloth Studio's startup time goes, per platform.
Nothing measured this before: the backend logs "lifespan startup completed in X ms"
but no test or CI job asserted a budget, and studio_test_kit discards the elapsed
time of its /healthz poll. A first local run (Linux, warm cache, fast server CPU)
found `import main` alone costs 6.6s before the server can bind, dominated by eager
module-level imports pulled in by the `routes` package:
torch 1930 ms self
unsloth_zoo 914 ms self
routes 779 ms self
transformers 524 ms self
Phases measured:
import `python -X importtime -c "import main"`, top cumulative + per-package self
spawn process start -> first byte on stdout
healthz process start -> /api/health (or /healthz) answers 200
lifespan the backend's own "lifespan startup completed in X ms" log line
Usage:
python scripts/profile_startup.py --repeats 3 --json out.json
python scripts/profile_startup.py --import-only # no server, no port needed
Exit code is 0 unless --max-healthz-seconds is given and exceeded.
"""
from __future__ import annotations
import argparse
import json
import math
import os
import platform
import re
import shutil
import socket
import statistics
import subprocess
import sys
import threading
import time
import urllib.error
import urllib.request
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
BACKEND = REPO_ROOT / "studio" / "backend"
_IMPORTTIME_RE = re.compile(r"import time:\s+(\d+)\s+\|\s+(\d+)\s+\|(\s*)(\S.*)")
def _free_port() -> int:
with socket.socket() as s:
s.bind(("127.0.0.1", 0))
return int(s.getsockname()[1])
def profile_imports(python: str, top: int = 15) -> dict:
"""Cumulative and self import cost for the backend's module graph.
Run in a subprocess with -X importtime: the numbers are only meaningful for a
cold interpreter, and importing in-process would measure a warm sys.modules.
"""
proc = subprocess.run(
[python, "-X", "importtime", "-c", "import sys; sys.path.insert(0, '.'); import main"],
cwd = BACKEND,
capture_output = True,
text = True,
timeout = 900,
)
rows = []
for line in proc.stderr.splitlines():
m = _IMPORTTIME_RE.match(line)
if m:
rows.append((int(m.group(1)), int(m.group(2)), m.group(4).strip()))
if not rows:
return {"ok": False, "error": (proc.stderr or proc.stdout)[-2000:]}
if proc.returncode != 0:
# Rows survive up to the failure, so any total from a partial graph is wrong.
return {
"ok": False,
"error": (proc.stderr or proc.stdout)[-2000:],
"partial_rows": len(rows),
}
by_cum = sorted(rows, key = lambda r: -r[1])
# Total comes from the `main` row, not by_cum[0]: -X importtime also prints the
# interpreter's own startup graph (`site`), which can outrank a trivial main.
main_row = next((r for r in reversed(rows) if r[2] == "main"), None)
if main_row is None:
return {
"ok": False,
"error": "no `import main` row in -X importtime output\n"
+ (proc.stderr or proc.stdout)[-2000:],
}
self_by_pkg: dict[str, int] = {}
for self_us, _cum, name in rows:
pkg = name.split(".")[0]
self_by_pkg[pkg] = self_by_pkg.get(pkg, 0) + self_us
return {
"ok": True,
"total_seconds": round(main_row[1] / 1e6, 3),
"top_cumulative": [
{"module": n, "seconds": round(c / 1e6, 3)} for _s, c, n in by_cum[:top]
],
"self_by_package_ms": {
k: round(v / 1000) for k, v in sorted(self_by_pkg.items(), key = lambda x: -x[1])[:top]
},
}
def _terminate_tree(proc: subprocess.Popen) -> None:
"""Stop the server AND its children, which on Windows are a separate process.
CI profiles `Scripts/unsloth.exe`, a distlib launcher stub that CreateProcess's
the venv python and waits, so terminate() reaps the stub only: the real backend
keeps the inherited stdout handle, the reader thread never sees EOF, and
--repeats strands one server per iteration on the shared UNSLOTH_STUDIO_HOME.
taskkill /T walks the tree, as unsloth_cli/commands/start.py already does.
"""
if proc.poll() is not None:
return
if os.name == "nt":
try:
killed = subprocess.run(
["taskkill", "/PID", str(proc.pid), "/T", "/F"],
capture_output = True,
timeout = 30,
check = False,
)
if killed.returncode == 0:
return
except Exception:
# taskkill missing or timed out; fall through so the stub still dies.
pass
# check=False: a nonzero taskkill does not raise, so fall through as well.
proc.terminate()
def profile_launch(
bin_path: str,
port: int,
timeout_s: int = 300,
) -> dict:
"""Spawn the backend the way the desktop app does and time it to first 200."""
log_lines: list[str] = []
first_byte: list[float] = []
t0 = time.perf_counter()
proc = subprocess.Popen(
[bin_path, "studio", "--api-only", "-H", "127.0.0.1", "-p", str(port)],
cwd = REPO_ROOT,
stdout = subprocess.PIPE,
stderr = subprocess.STDOUT,
text = True,
bufsize = 1,
)
def _drain() -> None:
# Runs alongside the health polling: the first read timestamps the spawn
# phase, and an undrained pipe blocks the backend before it binds.
for line in proc.stdout:
if not first_byte:
first_byte.append(time.perf_counter() - t0)
log_lines.append(line.rstrip("\n"))
reader = threading.Thread(target = _drain, daemon = True)
reader.start()
t_healthz = None
deadline = t0 + timeout_s
try:
while time.perf_counter() < deadline:
if proc.poll() is not None:
break
if t_healthz is None:
for url in (
f"http://127.0.0.1:{port}/api/health",
f"http://127.0.0.1:{port}/healthz",
):
try:
with urllib.request.urlopen(url, timeout = 2) as r:
if r.status == 200:
t_healthz = time.perf_counter() - t0
break
except (urllib.error.URLError, OSError, TimeoutError):
pass
if t_healthz is not None:
break
time.sleep(0.25)
finally:
_terminate_tree(proc)
try:
# Safe: the reader drains the pipe, so the child cannot block on write().
proc.wait(timeout = 30)
except subprocess.TimeoutExpired:
proc.kill()
proc.wait()
reader.join(timeout = 10)
t_first_byte = first_byte[0] if first_byte else None
lifespan_ms = None
for line in log_lines:
m = re.search(r"lifespan startup completed in ([\d.]+)ms", line)
if m:
lifespan_ms = float(m.group(1))
return {
"spawn_seconds": round(t_first_byte, 3) if t_first_byte is not None else None,
"healthz_seconds": round(t_healthz, 3) if t_healthz is not None else None,
"lifespan_ms": lifespan_ms,
"reached_healthz": t_healthz is not None,
"log_tail": log_lines[-25:],
}
def python_version_of(python: str) -> str:
"""Version of the interpreter that runs the imports, not the one running us.
--python points at the installed Studio venv while this script runs under the
runner's system python, so platform.python_version() would label it wrong.
"""
if python == sys.executable:
return platform.python_version()
try:
proc = subprocess.run(
[python, "-c", "import platform; print(platform.python_version())"],
capture_output = True,
text = True,
timeout = 60,
)
if proc.returncode == 0 and proc.stdout.strip():
return proc.stdout.strip()
except (OSError, subprocess.SubprocessError):
pass
return "unknown"
def find_bin() -> str | None:
home = os.environ.get("UNSLOTH_STUDIO_HOME") or str(Path.home() / ".unsloth" / "studio")
names = ["unsloth.exe", "unsloth"] if platform.system() == "Windows" else ["unsloth"]
subdirs = ["unsloth_studio/Scripts", "unsloth_studio/bin", "bin", "Scripts"]
for sd in subdirs:
for n in names:
p = Path(home) / sd / n
if p.exists():
return str(p)
return shutil.which("unsloth")
def main(argv: list[str]) -> int:
ap = argparse.ArgumentParser(
description = __doc__, formatter_class = argparse.RawDescriptionHelpFormatter
)
ap.add_argument(
"--repeats",
type = int,
default = 1,
help = "launch repeats; the median is reported (imports are measured once)",
)
ap.add_argument(
"--python",
default = sys.executable,
help = "interpreter used for the import profile (default: this one)",
)
ap.add_argument("--bin", help = "path to the unsloth CLI (default: autodetect)")
ap.add_argument(
"--import-only",
action = "store_true",
help = "skip the server phases (no install needed beyond the deps)",
)
ap.add_argument(
"--max-healthz-seconds",
type = float,
help = "fail if the median time to a healthy port exceeds this",
)
ap.add_argument("--json", help = "write the full report here")
a = ap.parse_args(argv)
# range(0) launches nothing, leaving the budget check with nothing to fail on.
if a.repeats < 1:
ap.error("--repeats must be at least 1")
# Same reason: --import-only never launches anything.
if a.import_only and a.max_healthz_seconds is not None:
ap.error("--max-healthz-seconds cannot be combined with --import-only")
# nan and inf parse fine as floats but `med > budget` is then always False,
# so the gate would report success without ever bounding anything.
if a.max_healthz_seconds is not None and not math.isfinite(a.max_healthz_seconds):
ap.error("--max-healthz-seconds must be a finite number")
report: dict = {
"platform": platform.system().lower(),
"machine": platform.machine(),
"python": python_version_of(a.python),
"cpu_count": os.cpu_count(),
}
print("== import graph ==")
report["imports"] = profile_imports(a.python)
imp = report["imports"]
if imp.get("ok"):
print(f" import main: {imp['total_seconds']}s")
for row in imp["top_cumulative"][:8]:
print(f" {row['seconds']:7.3f}s {row['module']}")
print(" self time by package (ms):")
for k, v in list(imp["self_by_package_ms"].items())[:8]:
print(f" {v:8} ms {k}")
else:
print(f" FAILED: {imp.get('error', '')[:400]}")
if not a.import_only:
bin_path = a.bin or find_bin()
if not bin_path:
print(
"== launch == skipped: no unsloth CLI found "
"(set UNSLOTH_STUDIO_HOME or pass --bin)"
)
report["launch"] = {"skipped": "no unsloth CLI found"}
else:
print(f"== launch == {bin_path}")
runs = []
for i in range(a.repeats):
r = profile_launch(bin_path, _free_port())
runs.append(r)
print(
f" run {i + 1}: healthz={r['healthz_seconds']}s "
f"lifespan={r['lifespan_ms']}ms reached={r['reached_healthz']}"
)
got = [r["healthz_seconds"] for r in runs if r["healthz_seconds"] is not None]
report["launch"] = {
"runs": runs,
"failed_runs": sum(1 for r in runs if not r["reached_healthz"]),
"healthz_median_seconds": round(statistics.median(got), 3) if got else None,
"healthz_max_seconds": round(max(got), 3) if got else None,
}
if got:
print(
f" median time to healthy port: {report['launch']['healthz_median_seconds']}s"
)
if a.json:
Path(a.json).write_text(json.dumps(report, indent = 2), encoding = "utf-8")
print(f"\nwrote {a.json}")
if a.max_healthz_seconds is not None:
launch = report.get("launch") or {}
med = launch.get("healthz_median_seconds")
failed = launch.get("failed_runs") or 0
if failed:
# Failed launches fail the budget; dropping them would keep only the fast ones.
print(
f"::error::startup regression: {failed} of {len(launch.get('runs') or [])} "
f"launches never became healthy within the timeout"
)
return 1
if med is None:
# Nothing measured: exiting 0 would pass a requested budget without a
# single health request, so fail closed.
print(
"::error::startup regression: no healthz measurement, so the "
f"{a.max_healthz_seconds}s budget was never checked "
f"({launch.get('skipped') or 'launch phase produced no runs'})"
)
return 1
elif med > a.max_healthz_seconds:
print(
f"::error::startup regression: {med}s median to a healthy port "
f"exceeds the {a.max_healthz_seconds}s budget"
)
return 1
return 0
if __name__ == "__main__":
raise SystemExit(main(sys.argv[1:]))

View file

@ -11,12 +11,11 @@ import jwt
from .storage import (
API_KEY_PREFIX,
credential_generation,
get_jwt_secret,
get_user_and_secret,
load_jwt_secret,
save_refresh_token,
validate_api_key_with_credential,
validate_api_key,
verify_refresh_token,
)
@ -55,14 +54,11 @@ def create_access_token(
expires_delta: Optional[timedelta] = None,
*,
desktop: bool = False,
secret: Optional[str] = None,
) -> str:
"""
Create a signed JWT for the given subject (e.g. username).
Valid across restarts: the signing secret is stored in SQLite. Callers that
already verified a credential pass ``secret`` so a rotation landing mid-request
cannot sign the token with the credential that just replaced it.
Valid across restarts: the signing secret is stored in SQLite.
"""
to_encode = {"sub": subject}
if desktop:
@ -73,7 +69,7 @@ def create_access_token(
to_encode.update({"exp": expire})
return jwt.encode(
to_encode,
secret if secret is not None else _get_secret_for_subject(subject),
_get_secret_for_subject(subject),
algorithm = ALGORITHM,
)
@ -100,28 +96,15 @@ def is_desktop_access_token(token: str) -> bool:
return payload.get("sub") == subject and payload.get("desktop") is True
def create_refresh_token(
subject: str,
*,
desktop: bool = False,
secret: Optional[str] = None,
) -> str:
def create_refresh_token(subject: str, *, desktop: bool = False) -> str:
"""
Create a random refresh token, store its hash in SQLite, and return it.
Refresh tokens are opaque (not JWTs); expire after REFRESH_TOKEN_EXPIRE_DAYS.
``secret`` stamps the token with the credential version the caller verified,
so a rotation cannot leave a token minted from the replaced credential valid.
"""
token = secrets.token_urlsafe(48)
expires_at = datetime.now(timezone.utc) + timedelta(days = REFRESH_TOKEN_EXPIRE_DAYS)
save_refresh_token(
token,
subject,
expires_at.isoformat(),
is_desktop = desktop,
secret_gen = credential_generation(secret) if secret is not None else None,
)
save_refresh_token(token, subject, expires_at.isoformat(), is_desktop = desktop)
return token
@ -154,22 +137,7 @@ def reload_secret() -> None:
async def get_current_subject(credentials: HTTPAuthorizationCredentials = Depends(security)) -> str:
"""Validate JWT and require the password-change flow to be completed."""
subject, _generation = await _get_current_credential(
credentials,
allow_password_change = False,
)
return subject
async def get_current_credential(
credentials: HTTPAuthorizationCredentials = Depends(security),
) -> Tuple[str, Optional[str]]:
"""As get_current_subject, but also returns the credential generation.
For routes that persist a new credential and must not do so on behalf of one
a concurrent reset has revoked.
"""
return await _get_current_credential(
return await _get_current_subject(
credentials,
allow_password_change = False,
)
@ -190,11 +158,10 @@ async def get_current_subject_allow_password_change(
credentials: HTTPAuthorizationCredentials = Depends(security),
) -> str:
"""Validate JWT but allow access to the password-change endpoint."""
subject, _generation = await _get_current_credential(
return await _get_current_subject(
credentials,
allow_password_change = True,
)
return subject
# The literal the examples ship with; pasted unedited more often than a revoked key.
@ -212,27 +179,21 @@ def _invalid_api_key_detail(token: str) -> str:
return "Invalid or expired API key"
async def _get_current_credential(
async def _get_current_subject(
credentials: HTTPAuthorizationCredentials, *, allow_password_change: bool
) -> Tuple[str, Optional[str]]:
"""Validate the bearer and return ``(subject, credential generation)``.
The generation is the credential version this request actually authenticated
against. Routes that persist new credentials must bind their write to it, or
a reset landing mid-request would bless what it just revoked.
"""
) -> str:
"""FastAPI dependency: validate the JWT and return the subject. Use on protected routes."""
token = credentials.credentials
# --- API key path (sk-unsloth-...) ---
if token.startswith(API_KEY_PREFIX):
verified = validate_api_key_with_credential(token)
if verified is None:
username = validate_api_key(token)
if username is None:
raise HTTPException(
status_code = status.HTTP_401_UNAUTHORIZED,
detail = _invalid_api_key_detail(token),
)
username, secret = verified
return username, credential_generation(secret)
return username
# --- JWT path ---
subject = _decode_subject_without_verification(token)
@ -263,7 +224,7 @@ async def _get_current_credential(
status_code = status.HTTP_403_FORBIDDEN,
detail = "Password change required",
)
return subject, credential_generation(jwt_secret)
return subject
except jwt.InvalidTokenError:
raise HTTPException(
status_code = status.HTTP_401_UNAUTHORIZED,

View file

@ -9,7 +9,6 @@ import ipaddress
import os
import secrets
import sqlite3
import tempfile
import threading
from datetime import datetime, timezone
from typing import Optional, Tuple
@ -31,97 +30,6 @@ _BOOTSTRAP_PW_PATH = DB_PATH.parent / ".bootstrap_password"
_bootstrap_password: Optional[str] = None
def _bootstrap_file_bytes(password: str) -> bytes:
"""Exact on-disk form: the secret plus one LF.
Bytes, not text: text mode writes CRLF on Windows, and `$(cat ...)` strips
the LF but leaves the CR attached to the credential.
"""
return (password + "\n").encode("utf-8")
def _persist_bootstrap_password(password: str) -> None:
"""Atomically write the bootstrap password 0600, LF terminated on every OS.
A partial write would destroy the only plaintext recovery credential.
"""
fd, tmp_name = tempfile.mkstemp(
prefix = f".{_BOOTSTRAP_PW_PATH.name}.", dir = _BOOTSTRAP_PW_PATH.parent
)
try:
with os.fdopen(fd, "wb") as f:
f.write(_bootstrap_file_bytes(password))
try:
os.chmod(tmp_name, 0o600)
except OSError:
pass
os.replace(tmp_name, _BOOTSTRAP_PW_PATH)
except BaseException:
try:
os.unlink(tmp_name)
except OSError:
pass
raise
def _normalise_bootstrap_file(raw: bytes, password: str) -> None:
"""Append the LF a pre-newline release left off.
Append-only, and only when the file is exactly the credential:
clear_bootstrap_password() may unlink or (when unlink fails, notably on
Windows while this descriptor is open) truncate through another descriptor
after we read, so a rewrite could restore revoked plaintext. An append
cannot: worst case is a lone "\\n" over a cleared file, which strips back to
no bootstrap password. Pre-newline releases wrote no terminator at all, so
that is the only shape in the wild; anything else reads fine, since every
reader strips, and is left alone.
"""
if raw != password.encode("utf-8"):
return
# O_BINARY: without it Windows opens in text mode and turns the LF straight
# back into CRLF, the bug being fixed.
fd = os.open(
_BOOTSTRAP_PW_PATH,
os.O_WRONLY | os.O_APPEND | getattr(os, "O_BINARY", 0),
)
try:
os.write(fd, b"\n")
try:
os.fchmod(fd, 0o600)
except (AttributeError, OSError):
# fchmod only reached Windows in 3.13.
pass
finally:
os.close(fd)
def _read_persisted_bootstrap_password() -> Optional[str]:
"""Read the persisted password, normalising the file if it is malformed."""
if not _BOOTSTRAP_PW_PATH.is_file():
return None
# No caller handles a raise, so an unreadable file has to mean "no bootstrap
# password", not a dead backend. We write UTF-8, so undecodable bytes are
# damage whose plaintext is worthless anyway.
try:
raw = _BOOTSTRAP_PW_PATH.read_bytes()
password = raw.decode("utf-8").strip()
except (OSError, UnicodeDecodeError):
return None
if not password:
return None
# Older releases wrote no terminator; best-effort, a read-only auth dir must
# not fail startup.
if raw != _bootstrap_file_bytes(password):
try:
_normalise_bootstrap_file(raw, password)
except OSError:
pass
return password
def generate_bootstrap_password() -> str:
"""Generate a 4-word diceware passphrase and persist it to disk.
@ -135,10 +43,10 @@ def generate_bootstrap_password() -> str:
return _bootstrap_password
# Persisted from a previous run?
persisted = _read_persisted_bootstrap_password()
if persisted:
_bootstrap_password = persisted
return _bootstrap_password
if _BOOTSTRAP_PW_PATH.is_file():
_bootstrap_password = _BOOTSTRAP_PW_PATH.read_text(encoding = "utf-8").strip()
if _bootstrap_password:
return _bootstrap_password
# First startup: generate a fresh passphrase.
import diceware
@ -149,7 +57,11 @@ def generate_bootstrap_password() -> str:
# Persist so the same passphrase survives restarts until password change.
ensure_dir(_BOOTSTRAP_PW_PATH.parent)
_persist_bootstrap_password(_bootstrap_password)
_BOOTSTRAP_PW_PATH.write_text(_bootstrap_password, encoding = "utf-8")
try:
os.chmod(_BOOTSTRAP_PW_PATH, 0o600)
except OSError:
pass
return _bootstrap_password
@ -160,14 +72,13 @@ def get_bootstrap_password() -> Optional[str]:
def _load_bootstrap_password() -> Optional[str]:
"""Load an existing bootstrap password without creating one.
Upgrades take this path, not generate_bootstrap_password()
(ensure_default_admin short-circuits once the admin row exists), so it has
to normalise too.
"""
"""Load an existing bootstrap password without creating one."""
global _bootstrap_password
_bootstrap_password = _read_persisted_bootstrap_password()
_bootstrap_password = None
if _BOOTSTRAP_PW_PATH.is_file():
bootstrap_password = _BOOTSTRAP_PW_PATH.read_text(encoding = "utf-8").strip()
if bootstrap_password:
_bootstrap_password = bootstrap_password
return _bootstrap_password
@ -186,7 +97,7 @@ def clear_bootstrap_password() -> None:
# Removal failed (Windows AV, read-only auth dir). The hash is already
# committed, so don't fail the change -- but truncate the file so its
# stale plaintext can't be re-seeded by generate_bootstrap_password()
# if auth.db is ever recreated.
# if a later reset-password deletes auth.db and re-validates it.
try:
_BOOTSTRAP_PW_PATH.write_text("", encoding = "utf-8")
cleared = True
@ -221,31 +132,6 @@ def _hash_token(token: str) -> str:
return hashlib.sha256(token.encode("utf-8")).hexdigest()
class CredentialRotated(Exception):
"""A password reset revoked the credential this request authenticated with."""
def credential_generation(jwt_secret: str) -> str:
"""Marker for the credential version a refresh token was issued under.
Every password change rotates ``jwt_secret``, so a token stamped with the
previous one is rejected even if it was inserted after the revoking DELETE.
"""
return hashlib.sha256(jwt_secret.encode("utf-8")).hexdigest()
def _current_secret(conn: sqlite3.Connection, username: str) -> Optional[str]:
row = conn.execute(
"SELECT jwt_secret FROM auth_user WHERE username = ?", (username,)
).fetchone()
return row["jwt_secret"] if row else None
def _current_generation(conn: sqlite3.Connection, username: str) -> Optional[str]:
secret = _current_secret(conn, username)
return credential_generation(secret) if secret is not None else None
def get_connection() -> sqlite3.Connection:
"""Get a connection to the auth database, creating tables if needed."""
ensure_dir(DB_PATH.parent)
@ -289,8 +175,7 @@ def get_connection() -> sqlite3.Connection:
token_hash TEXT NOT NULL,
username TEXT NOT NULL,
expires_at TEXT NOT NULL,
is_desktop INTEGER NOT NULL DEFAULT 0,
secret_gen TEXT
is_desktop INTEGER NOT NULL DEFAULT 0
);
"""
)
@ -329,8 +214,6 @@ def get_connection() -> sqlite3.Connection:
refresh_columns = {row["name"] for row in conn.execute("PRAGMA table_info(refresh_tokens)")}
if "is_desktop" not in refresh_columns:
conn.execute("ALTER TABLE refresh_tokens ADD COLUMN is_desktop INTEGER NOT NULL DEFAULT 0")
if "secret_gen" not in refresh_columns:
conn.execute("ALTER TABLE refresh_tokens ADD COLUMN secret_gen TEXT")
conn.commit()
return conn
@ -704,22 +587,12 @@ def update_password(
new_password: str,
*,
revoke_refresh_tokens: bool = False,
expect_password_hash: Optional[str] = None,
) -> Optional[str]:
) -> bool:
"""Update password, clear first-login requirement, rotate JWT secret.
Returns the new JWT secret, or None when nothing was updated. Callers that
mint tokens for the caller must sign with the returned secret: re-reading it
would pick up a reset that landed between this commit and the mint.
``revoke_refresh_tokens`` deletes the user's refresh tokens in the SAME
transaction: a separate delete could fail after the password commit and
leave a pre-change token still able to mint access tokens.
``expect_password_hash`` makes the write conditional on the credential the
caller verified still being current, so a request that checked the old
password cannot overwrite a reset that landed while it was in flight.
Returns False when the credential moved underneath it.
"""
from .hashing import hash_password
@ -727,32 +600,21 @@ def update_password(
jwt_secret = secrets.token_urlsafe(64)
conn = get_connection()
try:
if expect_password_hash is None:
cursor = conn.execute(
"""
UPDATE auth_user
SET password_salt = ?, password_hash = ?, jwt_secret = ?, must_change_password = 0
WHERE username = ?
""",
(salt, pwd_hash, jwt_secret, username),
)
else:
cursor = conn.execute(
"""
UPDATE auth_user
SET password_salt = ?, password_hash = ?, jwt_secret = ?, must_change_password = 0
WHERE username = ? AND password_hash = ?
""",
(salt, pwd_hash, jwt_secret, username, expect_password_hash),
)
cursor = conn.execute(
"""
UPDATE auth_user
SET password_salt = ?, password_hash = ?, jwt_secret = ?, must_change_password = 0
WHERE username = ?
""",
(salt, pwd_hash, jwt_secret, username),
)
if revoke_refresh_tokens and cursor.rowcount > 0:
conn.execute("DELETE FROM refresh_tokens WHERE username = ?", (username,))
conn.commit()
if cursor.rowcount > 0:
clear_bootstrap_password()
clear_desktop_secret()
return jwt_secret
return None
return cursor.rowcount > 0
finally:
conn.close()
@ -763,49 +625,35 @@ def save_refresh_token(
expires_at: str,
*,
is_desktop: bool = False,
secret_gen: Optional[str] = None,
) -> None:
"""
Store a hashed refresh token with its associated username and expiry.
``secret_gen`` binds the token to a credential version; it defaults to the
current one, and callers that already verified a credential must pass the
version they verified rather than let this re-read a rotated one.
"""
token_hash = _hash_token(token)
conn = get_connection()
try:
if secret_gen is None:
secret_gen = _current_generation(conn, username)
conn.execute(
"""
INSERT INTO refresh_tokens (token_hash, username, expires_at, is_desktop, secret_gen)
VALUES (?, ?, ?, ?, ?)
INSERT INTO refresh_tokens (token_hash, username, expires_at, is_desktop)
VALUES (?, ?, ?, ?)
""",
(token_hash, username, expires_at, int(is_desktop), secret_gen),
(token_hash, username, expires_at, int(is_desktop)),
)
conn.commit()
finally:
conn.close()
def consume_refresh_token(token: str) -> Optional[Tuple[str, bool, str]]:
def consume_refresh_token(token: str) -> Optional[Tuple[str, bool]]:
"""Atomically validate-and-delete a refresh token for single-use rotation.
DELETE RETURNING fuses validate and delete into one statement so two
concurrent refresh requests cannot both consume the same token. Returns
``(username, is_desktop, jwt_secret)``; the caller must mint the replacement
tokens against that secret so a rotation landing mid-refresh cannot issue a
post-rotation session from a pre-rotation token.
concurrent refresh requests cannot both consume the same token.
"""
token_hash = _hash_token(token)
now = datetime.now(timezone.utc).isoformat()
conn = get_connection()
try:
# One transaction with the delete: an unstamped legacy row has no
# generation to compare, so reading the credential after committing would
# hand a reset's new secret to a token issued before it.
conn.execute("BEGIN IMMEDIATE")
conn.execute(
"DELETE FROM refresh_tokens WHERE expires_at < ?",
(now,),
@ -814,21 +662,15 @@ def consume_refresh_token(token: str) -> Optional[Tuple[str, bool, str]]:
"""
DELETE FROM refresh_tokens
WHERE token_hash = ? AND expires_at >= ?
RETURNING username, is_desktop, secret_gen
RETURNING username, is_desktop
""",
(token_hash, now),
)
row = cur.fetchone()
if row is None:
conn.commit()
return None
secret = _current_secret(conn, row["username"])
conn.commit()
if secret is None:
if row is None:
return None
if row["secret_gen"] is not None and row["secret_gen"] != credential_generation(secret):
return None
return row["username"], bool(row["is_desktop"]), secret
return row["username"], bool(row["is_desktop"])
finally:
conn.close()
@ -852,7 +694,7 @@ def verify_refresh_token(token: str) -> Optional[Tuple[str, bool]]:
cur = conn.execute(
"""
SELECT id, username, expires_at, is_desktop, secret_gen FROM refresh_tokens
SELECT id, username, expires_at, is_desktop FROM refresh_tokens
WHERE token_hash = ?
""",
(token_hash,),
@ -861,13 +703,6 @@ def verify_refresh_token(token: str) -> Optional[Tuple[str, bool]]:
if row is None:
return None
if row["secret_gen"] is not None and row["secret_gen"] != _current_generation(
conn, row["username"]
):
conn.execute("DELETE FROM refresh_tokens WHERE id = ?", (row["id"],))
conn.commit()
return None
# Check expiry
expires_at = datetime.fromisoformat(row["expires_at"])
if datetime.now(timezone.utc) > expires_at:
@ -912,41 +747,30 @@ def create_desktop_secret() -> str:
conn.close()
def validate_desktop_secret_with_credential(raw_secret: str) -> Optional[Tuple[str, str]]:
"""Validate the desktop secret and return ``(username, jwt_secret)``.
Both reads share one transaction so the returned secret is the credential
version the desktop secret was checked against; a reset landing mid-request
then invalidates the tokens minted from it rather than blessing them.
"""
def validate_desktop_secret(raw_secret: str) -> Optional[str]:
"""Return the real admin username when the desktop secret matches."""
if not raw_secret.startswith(DESKTOP_SECRET_PREFIX):
return None
if get_user_and_secret(DEFAULT_ADMIN_USERNAME) is None:
return None
secret_hash = _pbkdf2_desktop_secret(raw_secret)
conn = get_connection()
try:
conn.execute("BEGIN")
row = conn.execute(
cur = conn.execute(
"SELECT value FROM app_secrets WHERE key = ?",
(_DESKTOP_SECRET_HASH_KEY,),
).fetchone()
if row is None or not secrets.compare_digest(row["value"], secret_hash):
)
row = cur.fetchone()
if row is None:
return None
jwt_secret = _current_secret(conn, DEFAULT_ADMIN_USERNAME)
if jwt_secret is None:
if not secrets.compare_digest(row["value"], secret_hash):
return None
return DEFAULT_ADMIN_USERNAME, jwt_secret
return DEFAULT_ADMIN_USERNAME
finally:
conn.rollback()
conn.close()
def validate_desktop_secret(raw_secret: str) -> Optional[str]:
"""Return the real admin username when the desktop secret matches."""
verified = validate_desktop_secret_with_credential(raw_secret)
return verified[0] if verified else None
def clear_desktop_secret() -> None:
"""Remove backend-side desktop auth state."""
conn = get_connection()
@ -972,7 +796,6 @@ def create_api_key(
name: str,
expires_at: Optional[str] = None,
internal: bool = False,
expect_gen: Optional[str] = None,
) -> Tuple[str, dict]:
"""Create a new API key for *username*.
@ -981,10 +804,6 @@ def create_api_key(
Pass ``internal=True`` for keys minted by workflows (e.g. data-recipe
runs) that should not appear in user-facing key listings.
``expect_gen`` ties the insert to the credential generation the request
authenticated under, so a session revoked by a concurrent password reset
cannot mint a key that outlives it. Raises ``CredentialRotated`` if it moved.
"""
raw_key = API_KEY_PREFIX + secrets.token_hex(16)
key_hash = _pbkdf2_api_key(raw_key)
@ -993,12 +812,6 @@ def create_api_key(
conn = get_connection()
try:
if expect_gen is not None:
conn.execute("BEGIN IMMEDIATE")
if _current_generation(conn, username) != expect_gen:
raise CredentialRotated(
"The credential this request authenticated with was revoked."
)
conn.execute(
"""
INSERT INTO api_keys (username, key_prefix, key_hash, name, created_at, expires_at, is_internal)
@ -1087,25 +900,15 @@ def revoke_internal_api_key(key_id: int) -> bool:
def validate_api_key(raw_key: str) -> Optional[str]:
"""Validate *raw_key* and return the owning username, or ``None``."""
verified = validate_api_key_with_credential(raw_key)
return verified[0] if verified else None
"""Validate *raw_key* and return the owning username, or ``None``.
def validate_api_key_with_credential(raw_key: str) -> Optional[Tuple[str, str]]:
"""Validate *raw_key* and return ``(username, jwt_secret)``, or ``None``.
Also updates ``last_used_at`` on success. The key check and the credential
read share one write transaction, so the returned version is the one the key
was actually valid under: a reset committing right after cannot have its new
generation handed to a request the key it revoked authenticated.
Also updates ``last_used_at`` on success.
"""
cache_id = _api_key_cache_id(raw_key)
cached_hash = _api_key_hash_cache.get(cache_id)
key_hash = cached_hash if cached_hash is not None else _pbkdf2_api_key(raw_key)
conn = get_connection()
try:
conn.execute("BEGIN IMMEDIATE")
cur = conn.execute(
"SELECT id, username, is_active, expires_at FROM api_keys WHERE key_hash = ?",
(key_hash,),
@ -1125,15 +928,11 @@ def validate_api_key_with_credential(raw_key: str) -> Optional[Tuple[str, str]]:
expires = datetime.fromisoformat(row["expires_at"])
if datetime.now(timezone.utc) > expires:
return None
secret = _current_secret(conn, row["username"])
if secret is None:
return None
conn.execute(
"UPDATE api_keys SET last_used_at = ? WHERE id = ?",
(datetime.now(timezone.utc).isoformat(), row["id"]),
)
conn.commit()
return row["username"], secret
return row["username"]
finally:
conn.rollback()
conn.close()

View file

@ -310,7 +310,6 @@ class CloudflareTunnel:
stderr = subprocess.STDOUT,
stdin = subprocess.DEVNULL,
text = True,
encoding = "utf-8",
errors = "replace",
bufsize = 1,
**_windows_hidden_kwargs(),

View file

@ -257,8 +257,6 @@ def _run_oxc_batch(
cwd = str(_OXC_TOOL_DIR),
input = json.dumps(payload),
text = True,
encoding = "utf-8",
errors = "replace",
capture_output = True,
check = False,
env = env,

View file

@ -567,7 +567,7 @@ class InferenceBackend:
_meta_path = Path(config.path) / "export_metadata.json"
try:
if _meta_path.exists():
_meta = json.loads(_meta_path.read_text(encoding = "utf-8-sig"))
_meta = json.loads(_meta_path.read_text(encoding = "utf-8"))
if _meta.get("base_model"):
processor_source = _meta["base_model"]
except Exception:

View file

@ -85,7 +85,6 @@ from core.tool_healing import (
strip_outside_think,
)
from utils.native_path_leases import child_env_without_native_path_secret
from utils.child_stdio import utf8_child_env
from utils.hf_xet_fallback import hf_hub_download_with_xet_fallback
from utils.subprocess_compat import (
windows_hidden_subprocess_kwargs as _windows_hidden_subprocess_kwargs,
@ -95,8 +94,6 @@ from core.inference.tool_call_parser import (
MAX_ACT_REPROMPTS as _MAX_REPROMPTS,
NUDGE_TOOL_CALLS_STATUS as _NUDGE_TOOL_CALLS_STATUS,
REPROMPT_MAX_CHARS as _REPROMPT_MAX_CHARS,
is_reprompt_repeat as _is_reprompt_repeat,
is_reprompt_restatement as _is_reprompt_restatement,
is_short_intent_without_action as _is_short_intent_without_action,
reprompt_to_act_message as _reprompt_to_act_message,
)
@ -363,32 +360,12 @@ _DEFAULT_STREAM_STALL_TIMEOUT_S = 120.0 # 2 min
# loop). Structured delta.tool_calls are grammar-bounded by llama-server; text
# parsed from content is not, so one runaway turn could fan out unbounded.
_MAX_TOOL_CALLS_PER_TURN = 8
# Obligation phrasing INTENT_SIGNAL leaves alone ("I need to call ..."), paired with
# an action verb. Sentence-anchored: mid-sentence the same words are prose that names
# a tool ("The API I should invoke is foo() because ..."), and suppressing that loses
# a real answer. "should"/"must" sit outside the need|have|ought group because they
# take a bare infinitive. "invoke"/"query" stay out of the verb list: they read as
# technical prose far more often than as a stall.
_FORCED_PLAN_INTENT = re.compile(
r"(?:^|[.!?]\s+)\s*"
r"(?:i\s+(?:(?:need|have|ought)\s+to|should|must)|need\s+to|going\s+to|must|should)"
r"\s+(?:\w+\s+){0,2}?(?:call|use|run|search|fetch|render)\b",
re.I | re.M,
)
# "the answer is not in the context" announces a *missing* answer, so the negated
# forms are excluded or the plan behind them would ship as the final response.
_FINAL_ANSWER_SIGNAL = re.compile(
r"\b(?:final\s+answer|answer\s*:|here\s+is|here's|in\s+summary|result\s*:"
r"|(?:the\s+)?answer\s+is(?!\s+(?:not|unavailable|unknown|unclear|missing)\b))\b",
_FORCED_REPEAT_PLAN_SIGNAL = re.compile(
r"\b(?:i\s+will|i'll|let\s+me|going\s+to|need\s+to|call|use|run|search|fetch|render)\b",
re.I,
)
# A plan that pivots ("I should call web_search, but Tokyo is the capital") has an
# answer attached, so the turn must survive. Leaking a plan sentence is cosmetic;
# dropping an answer is not, so the doubtful case keeps the output. The pivot has to
# carry text of its own: "I should call web_search, though." answers nothing.
_ANSWER_PIVOT = re.compile(
r"\b(?:but|however|although|though|that\s+said|in\s+the\s+meantime|meanwhile)\b"
r"[\W_]*(?:\w+[\W_]+){1,}\w",
_FINAL_ANSWER_SIGNAL = re.compile(
r"\b(?:final\s+answer|answer\s*:|here\s+is|here's|in\s+summary|result\s*:)\b",
re.I,
)
@ -480,28 +457,14 @@ def _held_rehearsal_tail_len(text: str, active_tools: list[dict]) -> int:
return len(tail) if tail and _is_rehearsal_prefix(tail, active_tools) else 0
def _should_suppress_forced_no_tool_output(text: str, previous: str = "") -> bool:
"""Suppress only repeated forced-turn planning text, not final answers.
``previous`` is the stall text that triggered the nudge, so a retry that
moved on can be told from one that just said the same thing again.
"""
def _should_suppress_forced_no_tool_output(text: str) -> bool:
"""Suppress only repeated forced-turn planning text, not final answers."""
stripped = text.strip()
if not stripped or len(stripped) >= _REPROMPT_MAX_CHARS:
return False
if _FINAL_ANSWER_SIGNAL.search(stripped):
return False
plan = _FORCED_PLAN_INTENT.search(stripped)
if plan is not None:
# Only the plan itself is safe to drop; anything the turn pivots to after it
# is the answer the user is waiting for.
return _ANSWER_PIVOT.search(stripped[plan.end() :]) is None
if not _is_short_intent_without_action(stripped):
return False
# INTENT_SIGNAL also fires on lead-ins to a real answer ("Now I have the results.
# The capital is Tokyo."), so a bare intent match is a stall only when the retry
# adds nothing. No ``previous`` keeps the standalone "is this a stall?" contract.
return not previous or _is_reprompt_restatement(stripped, previous)
return _FORCED_REPEAT_PLAN_SIGNAL.search(stripped) is not None
# ── Pre-compiled patterns for GGUF shard detection ───────────
@ -618,7 +581,7 @@ def _load_swa_cache() -> dict:
if _SWA_CACHE is not None:
return _SWA_CACHE
try:
with open(_swa_cache_path(), encoding = "utf-8-sig") as f:
with open(_swa_cache_path(), encoding = "utf-8") as f:
_SWA_CACHE = json.load(f)
if not isinstance(_SWA_CACHE, dict):
_SWA_CACHE = {}
@ -669,7 +632,7 @@ def _fetch_swa_entry_from_hf(repo_id: str) -> Optional[object]:
repo_type = "model",
cache_dir = active_hf_hub_cache(),
)
with open(cfg_path, encoding = "utf-8-sig") as f:
with open(cfg_path, encoding = "utf-8") as f:
cfg = json.load(f)
except Exception:
return None
@ -3083,7 +3046,6 @@ class LlamaCppBackend:
[bin_path, "--help"],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = 10,
check = False,
@ -3656,8 +3618,6 @@ class LlamaCppBackend:
],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = 10,
env = child_env_without_native_path_secret(),
**_windows_hidden_subprocess_kwargs(),
@ -3772,7 +3732,7 @@ class LlamaCppBackend:
encoding = "utf-8",
errors = "replace",
timeout = 15,
env = utf8_child_env(env),
env = env,
**_windows_hidden_subprocess_kwargs(),
)
if result.returncode != 0:
@ -5522,9 +5482,7 @@ class LlamaCppBackend:
stdout = subprocess.PIPE,
stderr = subprocess.STDOUT,
text = True,
encoding = "utf-8",
errors = "replace",
env = utf8_child_env(env),
env = env,
**_windows_hidden_subprocess_kwargs(),
**_child_popen_kwargs(),
)
@ -6738,8 +6696,6 @@ class LlamaCppBackend:
stdout = subprocess.PIPE,
stderr = subprocess.STDOUT,
text = True,
encoding = "utf-8",
errors = "replace",
env = env,
**_windows_hidden_subprocess_kwargs(),
**_child_popen_kwargs(),
@ -8756,8 +8712,6 @@ class LlamaCppBackend:
stdout = subprocess.PIPE,
stderr = subprocess.STDOUT,
text = True,
encoding = "utf-8",
errors = "replace",
env = env,
**_windows_hidden_subprocess_kwargs(),
**_child_popen_kwargs(),
@ -10260,8 +10214,6 @@ class LlamaCppBackend:
["pgrep", "-a", "-f", "llama-server"],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = 5,
env = child_env_without_native_path_secret(),
)
@ -11760,10 +11712,6 @@ class LlamaCppBackend:
# direct answer ("4", "Hello!") won't match. Pattern shared with the
# safetensors loop (tool_call_parser.INTENT_SIGNAL).
_reprompt_count = 0
# Budgeted apart from _reprompt_count so a pre-tool nudge can't spend it.
_post_tool_reprompts = 0
# Text that triggered the last nudge; if the retry restates it, stop.
_last_reprompt_text = ""
# Gates ``max_tool_iterations`` on real tool turns (not the enlarged range) so reserved
# re-prompt slots don't extend the budget. Mirrors the safetensors guard.
_tool_iters_done = 0
@ -11771,7 +11719,7 @@ class LlamaCppBackend:
# Reserve extra iterations for re-prompts so they don't consume the
# caller's tool-call budget; only when tool iterations are allowed.
_extra = _MAX_REPROMPTS + 1 if max_tool_iterations > 0 else 0
_extra = _MAX_REPROMPTS if max_tool_iterations > 0 else 0
for iteration in range(max_tool_iterations + _extra):
if cancel_event is not None and cancel_event.is_set():
return
@ -12416,10 +12364,12 @@ class LlamaCppBackend:
)
if not _safety_tc:
# ── Re-prompt on plan-without-action ──
# Intent described without a tool call: nudge it to act. Up
# to _MAX_REPROMPTS times, only on short responses with intent
# signals -- "4" or "Hello!" won't trigger it. Uses content,
# else reasoning text (reasoning-only stalls).
# If the model described its intent (forward-looking
# language) without calling a tool, nudge it to act.
# Fires at most once per request, only on short
# responses with intent signals -- "4" or "Hello!"
# won't trigger it. Use content if available, else
# fall back to reasoning text (reasoning-only stalls).
_stripped = content_accum.strip()
if not _stripped:
_stripped = reasoning_accum.strip()
@ -12429,33 +12379,18 @@ class LlamaCppBackend:
r"(?i)\brender[_\s-]?html\b",
_stripped,
)
# A post-tool stall still deserves a nudge, but each retry
# re-runs tools, so allow only one. RAG autoinject never lands
# in history, so _auto keeps a doc-grounded turn from reading
# as pre-tool (mirrors safetensors rag_autoinjected).
_already_acted = bool(_auto) or any(
record.executed for record in tool_controller.history
)
if _already_acted:
_reprompt_used, _reprompt_cap = _post_tool_reprompts, 1
else:
_reprompt_used, _reprompt_cap = _reprompt_count, _MAX_REPROMPTS
# None keeps the default-on re-prompt; False disables it.
if (
auto_heal_tool_calls
and (nudge_tool_calls is None or nudge_tool_calls)
and active_tools
and not _render_html_already_done_intent
and _reprompt_used < _reprompt_cap
and not _is_reprompt_repeat(_stripped, _last_reprompt_text)
and _reprompt_count < _MAX_REPROMPTS
and _is_short_intent_without_action(_stripped)
):
_reprompt_count += 1
if _already_acted:
_post_tool_reprompts += 1
_last_reprompt_text = _stripped
logger.info(
f"Re-prompt {_reprompt_used + 1}/{_reprompt_cap}: "
f"Re-prompt {_reprompt_count}/{_MAX_REPROMPTS}: "
f"model responded without calling tools "
f"({len(_stripped)} chars)"
)
@ -12493,10 +12428,7 @@ class LlamaCppBackend:
if _forced_tool_call_pending:
_forced_tool_call_pending = False
if not _should_suppress_forced_no_tool_output(
_stripped,
_last_reprompt_text,
):
if not _should_suppress_forced_no_tool_output(_stripped):
if cumulative_display:
forced_visible_text = _strip_tool_markup(
cumulative_display,
@ -12826,10 +12758,6 @@ class LlamaCppBackend:
_kb_search_count += 1
completion = tool_controller.record_result(decision, result)
resolved_provisional_tool_call_ids.add(decision.tool_call_id)
# A real execution opens the post-tool phase; carrying the pre-tool
# stall text over would read the same sentence as a repeat and
# swallow the one post-tool nudge.
_last_reprompt_text = ""
# A tool ran this turn, so it counts against the caller's budget.
_turn_executed_real_tool = True
yield completion.tool_end_event()

View file

@ -39,7 +39,6 @@ from core.inference.tool_call_parser import (
RAG_MAX_SEARCHES_PER_TURN,
RAG_SEARCH_CAP_NUDGE,
TOOL_XML_SIGNALS,
is_reprompt_repeat,
is_short_intent_without_action,
parse_tool_calls_from_text,
reprompt_to_act_message,
@ -566,8 +565,6 @@ def run_safetensors_tool_loop(
final_attempt_done = False
next_call_id = 0
reprompt_count = 0
# Text that triggered the last nudge; if the retry restates it, stop (GGUF parity).
last_reprompt_text = ""
# A denied tool confirmation must not be answered with a plan-without-action
# re-prompt (which would raise the confirmation gate again).
tool_denied = False
@ -1018,11 +1015,9 @@ def run_safetensors_tool_loop(
and not rag_autoinjected
and not tool_denied
and not any(record.executed for record in tool_controller.history)
and not is_reprompt_repeat(intent_text, last_reprompt_text)
and is_short_intent_without_action(intent_text)
):
reprompt_count += 1
last_reprompt_text = intent_text
logger.info(
"Safetensors re-prompt %d/%d: model responded without "
"calling tools (%d chars)",

View file

@ -166,40 +166,15 @@ RAG_SEARCH_CAP_NUDGE = (
# ── Plan-without-action re-prompt (shared by the GGUF and safetensors loops) ──
# Verbs naming work this turn. Narrow on purpose: "install"/"add"/"open" belong to
# advice for the user, which must not be re-prompted.
_ACTION_VERB = (
r"(?:search|check|look|find|fetch|get|call|use|run|query|invoke|analy[sz]e"
r"|review|inspect|read|gather|examine|retrieve|browse|consult|verify"
r"|confirm|compute|calculate|determine|identify|render)"
)
# Offering to help hands control back exactly like "let me know": measured on real
# turns, "I'll do my best to help" and "allow me to assist" close a clarification
# request and never precede a tool call. "help you" keeps its plan reading when an
# action follows it ("I'll help you search the web").
_HELP_OFFER = (
r"(?:do(?:ing)?\s+my\s+best|try\s+my\s+best|be\s+(?:able|happy|glad)\s+to\b"
r"|assist\b|help\s+you\b(?!\s+" + _ACTION_VERB + r")|give\s+you\s+accurate\b)"
)
# Forward-looking intent: the model says what it *will* do, not a final answer.
INTENT_SIGNAL = re.compile(
r"(?im)("
# Direct intent ("I'll"); lookahead drops negated forms ("I will not").
r"\b(i['\u2019](ll|m going to|m gonna)|i am (going to|gonna)|i will|i shall)\b"
r"(?!\s+(?:not|never)\b)(?!\s+" + _HELP_OFFER + r")"
r"(?i)("
# Direct intent ("I'll", "Let me"); lookahead drops negated forms
# ("I will not") so a refusal does not re-prompt.
r"\b(i['\u2019](ll|m going to|m gonna)|i am (going to|gonna)|i will|i shall|let me|allow me)\b(?!\s+(?:not|never)\b)"
r"|"
# "let me know" hands control back rather than announcing an action.
r"\b(?:let me|allow me)\b(?!\s+(?:not|never|know)\b)(?!\s+to\s+" + _HELP_OFFER + r")"
r"|"
# Step/plan framing. "first" must open a sentence and be followed by a plan
# (pronoun, "my/our plan", or an action verb); otherwise it is prose ("The
# first line is blank.", "First place went to Alice") or advice to the user.
r"(?:^|[.!?]\s+)\s*(?:the\s+)?first\s+step\b"
r"|(?:^|[.!?]\s+)\s*first\s*[,:–—-]?\s+(?:my|our)\s+(?:plan|approach|step)\b"
r"|(?:^|[.!?]\s+)\s*first\s*[,:–—-]?\s+(?:i|we|let[']?s|let us)\b"
r"|(?:^|[.!?]\s+)\s*first\s*[,:–—-]?\s+" + _ACTION_VERB + r"\b"
r"|"
r"\b(?:step \d+:?|here['\u2019]?s (?:my |the |a )?(?:plan|approach))"
# Step/plan framing: "First ...", "Step 1:", "Here's my plan"
r"\b(?:first\b|step \d+:?|here['\u2019]?s (?:my |the |a )?(?:plan|approach))"
r"|"
r"\b(?:now i|next i)\b"
r")"
@ -218,41 +193,6 @@ def is_short_intent_without_action(text: str) -> bool:
return 0 < len(stripped) < REPROMPT_MAX_CHARS and INTENT_SIGNAL.search(stripped) is not None
# Leading marks are kept unless they are quotes or brackets, so ".NET" survives;
# stripping all non-word chars would collapse "C++" and "C#" to the same token.
_REPEAT_TRAIL_PUNCT = ".,;:!?\"'`()[]{}<>‘’“”"
_REPEAT_LEAD_PUNCT = "\"'`([{‘“"
def _normalize_for_repeat(text: str) -> str:
words = []
for word in text.lower().split():
stripped = word.rstrip(_REPEAT_TRAIL_PUNCT).lstrip(_REPEAT_LEAD_PUNCT)
# Keep marks-only tokens: "value is 5" and "value is < 5" differ, and
# dropping the "<" threw the corrected attempt away.
words.append(stripped or word)
return " ".join(words)
# A nudge that just gets the same answer back has not worked, so stop there.
# Exact after normalisation, deliberately. Every relaxation tried here lost a real
# correction: a similarity ratio is length dependent (one changed token in a 50-word
# plan still scored 0.98), a set ignores order ("cats not dogs"), and ignoring filler
# words eats the target itself ("The Who", "OK Go"). A missed repeat costs one nudge
# out of MAX_ACT_REPROMPTS; a false one strands the plan unexecuted.
def is_reprompt_repeat(text: str, previous: str) -> bool:
return is_reprompt_restatement(text, previous)
# Same comparison, different decision: this one discards the turn. An appended answer
# must not match, and deletions flip meaning ("is not supported" -> "is supported").
def is_reprompt_restatement(text: str, previous: str) -> bool:
if not previous:
return False
a, b = _normalize_for_repeat(text), _normalize_for_repeat(previous)
return bool(a) and a == b
def reprompt_to_act_message(tool_hint: str) -> str:
"""The user message appended when re-prompting a plan-without-action turn."""
return (

View file

@ -151,7 +151,7 @@ def _resolve_lora_4bit(mc, load_in_4bit: bool) -> bool:
import json
try:
with open(adapter_cfg_path, encoding = "utf-8-sig") as f:
with open(adapter_cfg_path, encoding = "utf-8") as f:
adapter_cfg = json.load(f)
training_method = adapter_cfg.get("unsloth_training_method")
if training_method == "lora" and load_in_4bit:
@ -963,7 +963,7 @@ def run_inference_process(
if _local_adapter_cfg.is_file():
try:
_lora_base = (
_json.loads(_local_adapter_cfg.read_text(encoding = "utf-8-sig")).get(
_json.loads(_local_adapter_cfg.read_text(encoding = "utf-8")).get(
"base_model_name_or_path"
)
or None

View file

@ -103,8 +103,6 @@ class LlamaServerBackend:
[binary, "--help"],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = 30,
**windows_hidden_subprocess_kwargs(),
)
@ -333,8 +331,6 @@ class LlamaServerBackend:
stdout = subprocess.PIPE,
stderr = subprocess.STDOUT,
text = True,
encoding = "utf-8",
errors = "replace",
env = env,
**windows_hidden_subprocess_kwargs(),
**child_popen_kwargs(),

View file

@ -100,7 +100,7 @@ def _st_module_subdirs(name: str, token: str | None) -> tuple[str, ...]:
path = Path(normalize_path(name)).expanduser() / "modules.json"
if not path.is_file():
return ()
data = json.loads(path.read_text(encoding = "utf-8-sig"))
data = json.loads(path.read_text(encoding = "utf-8"))
else:
from huggingface_hub import hf_hub_download
from huggingface_hub.utils import EntryNotFoundError
@ -115,7 +115,7 @@ def _st_module_subdirs(name: str, token: str | None) -> tuple[str, ...]:
)
except EntryNotFoundError:
return ()
data = json.loads(open(local, encoding = "utf-8-sig").read())
data = json.loads(open(local, encoding = "utf-8").read())
subdirs = []
for module in data or ():
sub = str((module or {}).get("path", "")).strip().strip("/")

View file

@ -43,7 +43,6 @@ if sys.platform.startswith("linux") and "HSA_ENABLE_DXG_DETECTION" not in os.env
pass
logger = get_logger(__name__)
from utils.child_stdio import utf8_child_env
from utils.hardware import apply_gpu_ids
from utils.training_runs import build_default_output_dir_name
from utils.wheel_utils import (
@ -386,10 +385,6 @@ def _install_package_wheel_first(
"stdout": _sp.PIPE,
"stderr": _sp.STDOUT,
"text": True,
"encoding": "utf-8",
"errors": "replace",
# Make the Python child emit the UTF-8 we decode above.
"env": utf8_child_env(),
}
if is_hip:
_run_kwargs["timeout"] = 1800
@ -611,9 +606,6 @@ def _ensure_flash_linear_attention_unconditional(event_queue: Any) -> bool:
stdout = _sp.PIPE,
stderr = _sp.STDOUT,
text = True,
encoding = "utf-8",
errors = "replace",
env = utf8_child_env(),
timeout = _TILELANG_INSTALL_TIMEOUT_S,
)
except _sp.TimeoutExpired:
@ -857,9 +849,6 @@ def _run_pip(cmd: list[str], event_queue: Any, label: str) -> bool:
stdout = _sp.PIPE,
stderr = _sp.STDOUT,
text = True,
encoding = "utf-8",
errors = "replace",
env = utf8_child_env(),
timeout = _TILELANG_INSTALL_TIMEOUT_S,
)
except _sp.TimeoutExpired:

View file

@ -215,7 +215,7 @@ def _ollama_model_info_from_manifest(
return None
try:
manifest = json.loads(tag_file.read_text(encoding = "utf-8-sig"))
manifest = json.loads(tag_file.read_text(encoding = "utf-8"))
except (json.JSONDecodeError, OSError, UnicodeDecodeError) as e:
logger.debug("Skipping unreadable/invalid Ollama manifest %s: %s", tag_file, e)
return None
@ -228,7 +228,7 @@ def _ollama_model_info_from_manifest(
config_blob = _ollama_blob_path(blobs_dir, config_digest)
if config_blob is not None and _safe_is_file(config_blob):
try:
cfg = json.loads(config_blob.read_text(encoding = "utf-8-sig"))
cfg = json.loads(config_blob.read_text(encoding = "utf-8"))
model_type = cfg.get("model_type", "")
file_type = cfg.get("file_type", "")
except (json.JSONDecodeError, OSError, UnicodeDecodeError) as e:

View file

@ -464,8 +464,6 @@ def _read_marker_value(marker: Path) -> Optional[str]:
return None
value = marker.read_text(encoding = "utf-8").strip()
except (OSError, UnicodeDecodeError):
# UnicodeDecodeError is a ValueError, so it would escape and abort
# prepare_cache_for_transport. An unknown value just purges and restarts.
return None
return value if value in VALID_TRANSPORTS else None

View file

@ -42,12 +42,8 @@ class LogConfig:
log_level_name = os.getenv("LOG_LEVEL", "INFO").upper()
log_level = getattr(logging, log_level_name, logging.INFO)
# Non-ASCII on a non-UTF-8 stream raises UnicodeEncodeError (Windows,
# LANG=C), so key off the stream, not the platform.
for stream in (sys.stdout, sys.stderr):
if getattr(stream, "encoding", "") and not str(stream.encoding).lower().replace(
"-", ""
).startswith("utf8"):
if sys.platform == "win32":
for stream in (sys.stdout, sys.stderr):
if hasattr(stream, "reconfigure"):
try:
stream.reconfigure(encoding = "utf-8", errors = "replace")

View file

@ -330,6 +330,7 @@ from hub.utils.download_registry import (
)
from routes.settings import router as settings_router
from routes.prompts import router as prompts_router
from routes.profile_stats import router as profile_stats_router
from auth import storage
from auth.authentication import get_current_subject
from utils.hardware import (
@ -347,7 +348,6 @@ from utils.update_status import (
get_studio_install_source_status,
get_studio_update_status,
)
from utils.changelog import get_release_notes, is_supported_version_query
from utils.studio_version import get_studio_version
from utils.api_errors import install_api_error_handlers
@ -1049,6 +1049,7 @@ app.include_router(providers_router, prefix = "/api/providers", tags = ["provide
app.include_router(settings_router, prefix = "/api/settings", tags = ["settings"])
app.include_router(mcp_servers_router, prefix = "/api/mcp/servers", tags = ["mcp"])
app.include_router(prompts_router, prefix = "/api/prompts", tags = ["prompts"])
app.include_router(profile_stats_router, prefix = "/api/profile", tags = ["profile"])
app.include_router(datasets_router, prefix = "/api/datasets", tags = ["datasets"])
app.include_router(data_recipe_router, prefix = "/api/data-recipe", tags = ["data-recipe"])
app.include_router(llama_router, prefix = "/api/llama", tags = ["llama"])
@ -1155,18 +1156,6 @@ def studio_update_status(_current_subject: str = Depends(get_current_subject)):
return get_studio_update_status(UNSLOTH_VERSION)
@app.get("/api/studio/release-notes")
def studio_release_notes(
version: str = Query(..., max_length = 64),
refresh: bool = Query(False),
_current_subject: str = Depends(get_current_subject),
):
"""Return CHANGELOG.md notes for exactly `version` (never a nearby one)."""
if not is_supported_version_query(version):
raise HTTPException(status_code = 422, detail = "Invalid version.")
return get_release_notes(version, refresh = refresh)
@app.get(
"/api/studio/download-transport-capabilities",
response_model = TransportCapabilities,

View file

@ -6,93 +6,10 @@
from __future__ import annotations
import json
import locale
import os
import threading
from pathlib import Path
from typing import Any, Dict, NamedTuple
def _locale_encoding() -> str:
"""The codepage a pre-UTF-8 release here would have written, or "".
Empty on a UTF-8 host, where there is no codepage to attribute the file to.
"""
try:
preferred = locale.getencoding()
except AttributeError: # Python < 3.11
preferred = locale.getpreferredencoding(False)
if preferred.lower().replace("-", "").replace("_", "") == "utf8":
return ""
return preferred
# Trail bytes can land on JSON punctuation, so a single-byte fallback misreads these.
_DOUBLE_BYTE_ENCODINGS = ("cp932", "cp936", "cp949", "cp950")
def _parse(raw: bytes, encoding: str) -> Any:
"""Parse one JSON document under *encoding*, or None if it does not.
RecursionError is a RuntimeError, so nesting json.loads will not descend is
the one parse failure the other three miss. Both callers run this outside
any further handler, so it has to answer None here or a single damaged
record aborts the scraper at startup instead of being skipped.
"""
try:
return json.loads(raw.decode(encoding))
except (UnicodeDecodeError, LookupError, ValueError, RecursionError):
return None
class _Reading(NamedTuple):
as_utf8: Any
as_legacy: Any
def _read_line(raw: bytes, codepage: str) -> _Reading:
"""Read one line as UTF-8 and as a codepage, for dedup keys only.
Requiring valid JSON, not merely a successful decode, is what separates a
genuine legacy record from a half-written UTF-8 one: a torn multibyte
character decodes under cp1252 but leaves the JSON unterminated. Some byte
strings parse both ways, e.g. cp1251 ``Р°`` is ``D0 B0``, which is also
UTF-8 ``а``.
The codepage reading is never authoritative, because the file's own encoding
cannot be recovered from its bytes. Reading a cp1251 shard on a cp1252
machine turns ``Привет`` into ``Ïðèâåò`` and every byte of it decodes
cleanly, so a successful decode proves nothing about who wrote it. It is
used only to recover the dedup keys, which are ASCII ids and come back the
same under any of these, so the first reading that parses will do.
That is also why several are tried. latin-1 alone mangles the double-byte
codepages: cp932 ```` is ``95 5C``, and latin-1 turns the trail byte into
a JSON backslash, so the record fails to parse and its id is forgotten.
"""
as_utf8 = _parse(raw, "utf-8")
# A record that reads as UTF-8 needs no second reading: re-parsing cost 2.8x on a
# 76 MB shard, and these reach gigabytes. Only a dict, since key lookup falls
# through to the codepage when UTF-8 yields none.
if isinstance(as_utf8, dict):
return _Reading(as_utf8, None)
for encoding in (codepage, "latin-1", *_DOUBLE_BYTE_ENCODINGS):
if not encoding:
continue
as_legacy = _parse(raw, encoding)
if as_legacy is not None:
return _Reading(as_utf8, as_legacy)
return _Reading(as_utf8, None)
class _Scan(NamedTuple):
"""What a pass over an existing shard established about it."""
legacy: bool # enough evidence to trust the codepage reading's keys
readable: bool
saw_non_ascii: bool # some line's meaning depends on the encoding
utf8_keys: set # keys from lines UTF-8 could read
legacy_keys: set # keys only the codepage reading yields
from typing import Any, Dict
class StateStore:
@ -101,19 +18,12 @@ class StateStore:
self.path.parent.mkdir(parents = True, exist_ok = True)
self._lock = threading.Lock()
self._data: Dict[str, Any] = {}
# Read whole, and UTF-8 only unlike the shards below: a checkpoint holds
# nothing but base64 cursors and booleans, so a codepage retry could only ever
# add non-ASCII. That would resume on a mojibaked cursor, which GitHub rejects
# with INVALID_CURSOR_ARGUMENTS, and the empty page it returns marks the stream
# done and skips the rest for good. Dropping a damaged checkpoint re-scrapes
# from the first page, which the writers dedup.
if self.path.exists():
try:
raw = self.path.read_bytes()
except OSError:
raw = b""
data = _parse(raw, "utf-8")
self._data = data if isinstance(data, dict) else {}
with self.path.open(encoding = "utf-8") as f:
self._data = json.load(f)
except Exception:
self._data = {}
def get(
self,
@ -153,83 +63,24 @@ class JsonlWriter:
self.path = Path(path)
self.path.parent.mkdir(parents = True, exist_ok = True)
self._lock = threading.Lock()
self._fh = self.path.open("a", buffering = 1, encoding = "utf-8")
self._count_seen_keys: set[str] = set()
self._codepage = _locale_encoding()
self._ensure_ascii = False
encoding = "utf-8"
# Preload seen keys for dedup across resumes
if self.path.exists() and self.path.stat().st_size > 0:
scan = self._scan_existing()
self._count_seen_keys = scan.utf8_keys
if scan.legacy:
self._count_seen_keys |= scan.legacy_keys
if scan.saw_non_ascii or not scan.readable:
# Never convert: the writing encoding is unrecoverable and guessing
# mojibakes the records. Pure ASCII appends store identically under
# every codepage, and json.loads turns the \uXXXX escapes back.
encoding = "ascii"
self._ensure_ascii = True
self._fh = self.path.open("a", buffering = 1, encoding = encoding, errors = "strict")
def _scan_existing(self) -> _Scan:
"""Read the shard once to recover dedup keys and judge its encoding.
Line by line: these shards reach gigabytes on a large scrape, so neither
the bytes nor the decoded text are held whole.
The verdict weighs the whole file. Each line with non-ASCII bytes votes:
one that parses only under the codepage is evidence of a legacy shard,
one that parses as UTF-8 is evidence against, since arbitrary codepage
text almost never forms valid multibyte UTF-8. A single corrupt byte in
a healthy shard therefore cannot outvote the records around it, and a
genuinely legacy shard has a legacy vote on every line that carries an
umlaut.
More than one such line is required, because a single one is genuinely
undecidable: a legacy record holding one accented character and an ASCII
record holding one stray byte are the same shape. Reading it as damage
risks a duplicate; reading it as legacy marks an unreadable record seen
and blocks the retry that would replace it, losing it for good. Only one
of those is recoverable.
The verdict only picks which reading supplies the dedup keys. The file
itself is never rewritten either way, so a wrong answer costs at most a
duplicate, never a corrupted record.
"""
legacy_votes = 0
utf8_votes = 0
saw_non_ascii = False
utf8_keys: set[str] = set()
legacy_keys: set[str] = set()
try:
with self.path.open("rb") as handle:
for raw in handle:
line = raw.strip()
reading = _read_line(line, self._codepage)
# ASCII reads the same everywhere: no vote, no constraint.
if not line.isascii():
saw_non_ascii = True
if reading.as_utf8 is None and reading.as_legacy is not None:
legacy_votes += 1
elif reading.as_utf8 is not None:
utf8_votes += 1
# Kept apart so a damaged line does not block its own retry.
if isinstance(reading.as_utf8, dict):
key = self._key(reading.as_utf8)
if key is not None:
utf8_keys.add(key)
elif isinstance(reading.as_legacy, dict):
key = self._key(reading.as_legacy)
if key is not None:
legacy_keys.add(key)
except OSError:
return _Scan(False, False, False, utf8_keys, legacy_keys)
return _Scan(
legacy_votes > 1 and legacy_votes > utf8_votes,
True,
saw_non_ascii,
utf8_keys,
legacy_keys,
)
try:
# No guess is safe for a file an older build wrote in the
# operator's locale, so read past whatever will not decode.
with self.path.open(encoding = "utf-8", errors = "replace") as f:
for line in f:
try:
obj = json.loads(line)
k = self._key(obj)
if k is not None:
self._count_seen_keys.add(k)
except Exception:
pass
except Exception:
pass
def _key(self, obj: dict) -> str | None:
for k in ("id", "node_id", "number", "sha", "url"):
@ -248,7 +99,7 @@ class JsonlWriter:
return False
if k is not None:
self._count_seen_keys.add(k)
self._fh.write(json.dumps(obj, default = str, ensure_ascii = self._ensure_ascii))
self._fh.write(json.dumps(obj, default = str, ensure_ascii = False))
self._fh.write("\n")
self._fh.flush()
return True

View file

@ -30,8 +30,6 @@ class UnstructuredSeedReader(SeedReader[UnstructuredSeedSource]):
meta = json_mod.loads(meta_path.read_text(encoding = "utf-8"))
orig_name = meta.get("original_filename", path_obj.name)
except (json_mod.JSONDecodeError, OSError, UnicodeDecodeError):
# Undecodable metadata is as malformed as invalid JSON, so
# fall back to the file's own name rather than abort the seed.
pass
file_entries.append((path_obj, orig_name))

View file

@ -31,7 +31,6 @@ from auth import storage, hashing
from auth.authentication import (
create_access_token,
create_refresh_token,
get_current_credential,
get_current_subject,
get_current_subject_allow_password_change,
refresh_access_token,
@ -400,7 +399,7 @@ async def login(payload: AuthLoginRequest, request: Request) -> Token:
detail = f"Incorrect password. To reset it, run this in your terminal: {_reset_password_command()}",
)
salt, pwd_hash, jwt_secret, must_change_password = record
salt, pwd_hash, _jwt_secret, must_change_password = record
if not hashing.verify_password(payload.password, salt, pwd_hash):
_record_login_failure(key)
raise HTTPException(
@ -410,10 +409,8 @@ async def login(payload: AuthLoginRequest, request: Request) -> Token:
_clear_login_bucket(key)
_clear_login_bucket(unknown_key)
# Issue against the credential version just verified, not whatever is in the DB
# now: a concurrent reset-password must not hand this login a post-reset session.
access_token = create_access_token(subject = payload.username, secret = jwt_secret)
refresh_token = create_refresh_token(subject = payload.username, secret = jwt_secret)
access_token = create_access_token(subject = payload.username)
refresh_token = create_refresh_token(subject = payload.username)
return Token(
access_token = access_token,
refresh_token = refresh_token,
@ -441,17 +438,16 @@ async def logout(
@router.post("/desktop-login", response_model = Token)
async def desktop_login(payload: DesktopLoginRequest) -> Token:
"""Exchange a local desktop secret for normal admin-subject tokens."""
verified = storage.validate_desktop_secret_with_credential(payload.secret)
if verified is None:
username = storage.validate_desktop_secret(payload.secret)
if username is None:
raise HTTPException(
status_code = status.HTTP_401_UNAUTHORIZED,
detail = "Desktop authentication failed",
)
username, jwt_secret = verified
return Token(
access_token = create_access_token(subject = username, desktop = True, secret = jwt_secret),
refresh_token = create_refresh_token(subject = username, desktop = True, secret = jwt_secret),
access_token = create_access_token(subject = username, desktop = True),
refresh_token = create_refresh_token(subject = username, desktop = True),
token_type = "bearer",
must_change_password = False,
)
@ -466,11 +462,9 @@ async def refresh(payload: RefreshTokenRequest) -> Token:
status_code = status.HTTP_401_UNAUTHORIZED,
detail = "Invalid or expired refresh token",
)
username, is_desktop, jwt_secret = consumed
new_access_token = create_access_token(subject = username, desktop = is_desktop, secret = jwt_secret)
new_refresh_token = create_refresh_token(
subject = username, desktop = is_desktop, secret = jwt_secret
)
username, is_desktop = consumed
new_access_token = create_access_token(subject = username, desktop = is_desktop)
new_refresh_token = create_refresh_token(subject = username, desktop = is_desktop)
return Token(
access_token = new_access_token,
@ -513,25 +507,13 @@ async def change_password(
# Single transaction: a separate refresh-token purge could fail after the
# password commit, leaving pre-change tokens able to mint access tokens.
# Conditional on the hash just verified: a reset-password that landed while
# this request was in flight must not be overwritten by it.
new_secret = storage.update_password(
current_subject,
payload.new_password,
revoke_refresh_tokens = True,
expect_password_hash = pwd_hash,
)
if new_secret is None:
raise HTTPException(
status_code = status.HTTP_409_CONFLICT,
detail = "The password changed while this request was in flight. Sign in again.",
)
storage.update_password(current_subject, payload.new_password, revoke_refresh_tokens = True)
try:
request.app.state.bootstrap_password = None
except AttributeError:
pass
access_token = create_access_token(subject = current_subject, secret = new_secret)
refresh_token = create_refresh_token(subject = current_subject, secret = new_secret)
access_token = create_access_token(subject = current_subject)
refresh_token = create_refresh_token(subject = current_subject)
return Token(
access_token = access_token,
refresh_token = refresh_token,
@ -559,28 +541,20 @@ def _row_to_api_key_response(row: dict) -> ApiKeyResponse:
@router.post("/api-keys", response_model = CreateApiKeyResponse)
async def create_api_key(
payload: CreateApiKeyRequest, credential: tuple = Depends(get_current_credential)
payload: CreateApiKeyRequest, current_subject: str = Depends(get_current_subject)
) -> CreateApiKeyResponse:
"""Create a new API key. The raw key is returned once and cannot be retrieved later."""
current_subject, generation = credential
expires_at = None
if payload.expires_in_days is not None:
expires_at = (
datetime.now(timezone.utc) + timedelta(days = payload.expires_in_days)
).isoformat()
try:
raw_key, row = storage.create_api_key(
username = current_subject,
name = payload.name,
expires_at = expires_at,
expect_gen = generation,
)
except storage.CredentialRotated:
raise HTTPException(
status_code = status.HTTP_401_UNAUTHORIZED,
detail = "Invalid or expired token",
)
raw_key, row = storage.create_api_key(
username = current_subject,
name = payload.name,
expires_at = expires_at,
)
return CreateApiKeyResponse(
key = raw_key,
api_key = _row_to_api_key_response(row),

View file

@ -10,10 +10,7 @@ from datetime import datetime, timedelta, timezone
from typing import Any, Optional
from urllib.parse import urlparse
from fastapi import APIRouter, Depends, HTTPException, Query, Request
from auth.authentication import get_current_credential
from auth.storage import CredentialRotated
from fastapi import APIRouter, HTTPException, Query, Request
from fastapi.responses import JSONResponse, StreamingResponse
from pydantic import ValidationError
@ -260,11 +257,7 @@ def _inject_local_structured_response_format(
model_configs.extend(new_configs)
def _inject_local_providers(
recipe: dict[str, Any],
request: Request,
expect_gen: Optional[str] = None,
) -> Optional[int]:
def _inject_local_providers(recipe: dict[str, Any], request: Request) -> Optional[int]:
"""Mutate recipe in-place: point is_local providers at this server and mint
a short-lived internal sk-unsloth-* key for workflow auth.
@ -320,7 +313,6 @@ def _inject_local_providers(
name = "data-recipe workflow",
expires_at = expires_at,
internal = True,
expect_gen = expect_gen,
)
internal_key_id = int(row["id"])
@ -383,11 +375,7 @@ def _normalize_run_name(value: Any) -> str | None:
@router.post("/jobs", response_class = JSONResponse, response_model = JobCreateResponse)
def create_job(
payload: RecipePayload,
request: Request,
credential: tuple = Depends(get_current_credential),
):
def create_job(payload: RecipePayload, request: Request):
recipe = payload.recipe
if not recipe.get("columns"):
raise HTTPException(status_code = 400, detail = "Recipe must include columns.")
@ -418,11 +406,7 @@ def create_job(
) from exc
try:
internal_api_key_id = _inject_local_providers(recipe, request, credential[1])
except CredentialRotated as exc:
# A reset-password landed after this request authenticated; the workflow key
# is refused, so answer like any other revoked credential rather than 500.
raise HTTPException(status_code = 401, detail = "Invalid or expired token") from exc
internal_api_key_id = _inject_local_providers(recipe, request)
except ValueError as exc:
raise log_and_http_error(
exc,

View file

@ -4434,7 +4434,7 @@ def _effective_load_in_4bit(config: ModelConfig, requested: bool) -> bool:
if not adapter_cfg_path.exists():
return load_in_4bit
try:
with open(adapter_cfg_path, encoding = "utf-8-sig") as f:
with open(adapter_cfg_path, encoding = "utf-8") as f:
adapter_cfg = json.load(f)
if not isinstance(adapter_cfg, dict): # malformed -> keep requested
return load_in_4bit

View file

@ -722,7 +722,7 @@ def _scan_ollama_dir(ollama_dir: Path, limit: Optional[int] = None) -> List[Loca
stem_hash = hashlib.sha256(manifest_key.encode()).hexdigest()[:10]
try:
manifest = json.loads(tag_file.read_text(encoding = "utf-8-sig"))
manifest = json.loads(tag_file.read_text(encoding = "utf-8"))
except (json.JSONDecodeError, OSError, UnicodeDecodeError) as e:
logger.debug(
"Skipping unreadable/invalid Ollama manifest %s: %s",
@ -738,7 +738,7 @@ def _scan_ollama_dir(ollama_dir: Path, limit: Optional[int] = None) -> List[Loca
config_blob = blobs_dir / config_digest.replace(":", "-")
if config_blob.is_file():
try:
cfg = json.loads(config_blob.read_text(encoding = "utf-8-sig"))
cfg = json.loads(config_blob.read_text(encoding = "utf-8"))
model_type = cfg.get("model_type", "")
file_type = cfg.get("file_type", "")
except (json.JSONDecodeError, OSError, UnicodeDecodeError) as e:
@ -1042,7 +1042,7 @@ def _dir_has_downloaded_model(directory: Path, max_entries: int = 4000) -> bool:
if not m.is_file():
continue
try:
manifest = json.loads(m.read_text(encoding = "utf-8-sig"))
manifest = json.loads(m.read_text(encoding = "utf-8"))
except (json.JSONDecodeError, OSError, ValueError):
continue
for layer in manifest.get("layers") or []:
@ -3360,8 +3360,6 @@ def _wsl_reveal_in_explorer(path: Path) -> bool:
["wslpath", "-w", str(path)],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
check = True,
timeout = 10,
).stdout.strip()

View file

@ -0,0 +1,57 @@
# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""
Usage numbers for the Profile settings tab.
Read-only aggregation over the local studio.db (see
``storage.profile_stats_db``). Nothing is uploaded.
"""
import asyncio
from typing import Any
from fastapi import APIRouter, Depends, Query
from auth.authentication import get_current_subject
from loggers import get_logger
from storage.profile_stats_db import (
MAX_DAILY_DAYS,
MAX_TZ_OFFSET_MINUTES,
compute_profile_stats,
)
from utils.utils import log_and_http_error
router = APIRouter()
logger = get_logger(__name__)
@router.get("/stats")
async def get_profile_stats(
days: int = Query(MAX_DAILY_DAYS, ge = 1, le = MAX_DAILY_DAYS),
tz_offset_minutes: int = Query(0, ge = -MAX_TZ_OFFSET_MINUTES, le = MAX_TZ_OFFSET_MINUTES),
tz: str = Query("", max_length = 64),
current_subject: str = Depends(get_current_subject),
) -> dict[str, Any]:
"""Usage stats for the signed-in user's local history.
Days and hours are bucketed in the caller's timezone so a remote browser
does not read the server's calendar. ``tz`` is an IANA name, which carries
each date's own daylight-saving offset; ``tz_offset_minutes`` is the
``Date.getTimezoneOffset()`` fallback for hosts with no tzdata.
"""
try:
# A cold pass parses every message's metadata JSON: ~90 ms at 10k
# messages, ~1.2 s at 260k. Off the event loop so it cannot stall token
# streaming when Settings is opened mid-generation.
return await asyncio.to_thread(
compute_profile_stats,
days = days,
tz_offset_minutes = tz_offset_minutes,
tz_name = tz,
)
except Exception as exc:
raise log_and_http_error(
exc, 500, "Failed to compute profile statistics", log = logger
) from exc

View file

@ -10,7 +10,7 @@ import os
import sys
import time
from pathlib import Path
from typing import NoReturn, Optional, Sequence, Tuple
from typing import Optional, Tuple
def _fix_torch_cuda_ld_path():
@ -689,33 +689,6 @@ def _get_pid_on_port(port: int) -> "tuple[int, str] | None":
return None
def _bind_addresses(host: str, port: int) -> "set[str]":
"""Every address *host* resolves to. `localhost` is both 127.0.0.1 and ::1, and
recording only the first lets a later launch on the other one miss us."""
import socket
try:
infos = socket.getaddrinfo(host, port, socket.AF_UNSPEC, socket.SOCK_STREAM)
except OSError:
return {host}
return {info[4][0] for info in infos} or {host}
def _addresses_collide(recorded: "str | None", host: str, port: int) -> bool:
"""Would a server bound to *recorded* block a bind to *host*?
*recorded* may list several addresses. Unknown or wildcard on either side
collides: refusing with a clear message beats silently starting a duplicate.
"""
wildcards = ("0.0.0.0", "::", "")
if not recorded or host in wildcards:
return True
listed = {a.strip() for a in recorded.split(",") if a.strip()}
if not listed or listed & set(wildcards):
return True
return bool(listed & _bind_addresses(host, port))
def _is_port_free(host: str, port: int) -> bool:
"""Check if a port is available for binding.
@ -760,213 +733,18 @@ def _find_free_port(
host: str,
start: int,
max_attempts: int = 20,
avoid_own_studio: bool = False,
) -> int:
"""Find a free port from `start`, trying up to max_attempts ports.
``avoid_own_studio`` aborts rather than skipping past one of our own servers
in the fallback range, which would start a duplicate on a later port.
"""
"""Find a free port from `start`, trying up to max_attempts ports."""
for offset in range(max_attempts):
candidate = start + offset
if _is_port_free(host, candidate):
return candidate
if avoid_own_studio:
own = _own_studio_on_port(candidate, host)
if own is not None:
_abort_already_running(own, candidate)
raise RuntimeError(f"Could not find a free port in range {start}-{start + max_attempts - 1}")
from utils.paths.storage_roots import studio_root as _studio_root
# Legacy single-instance file; still read so `stop` finds an older build's server.
_PID_FILE = _studio_root() / "studio.pid"
PID_FILE_GLOB = "studio-*.pid"
def _pid_file_for_port(port: int) -> Path:
# PID in the name: 127.0.0.1 and ::1 can share a port, and one file per port
# would let the second bind overwrite the first.
return _studio_root() / f"studio-{port}-{os.getpid()}.pid"
def _pid_alive(pid: int) -> bool:
try:
import psutil
return psutil.pid_exists(pid)
except ImportError:
pass
if sys.platform == "win32":
# os.kill(pid, 0) raises OSError for every pid on Windows, so tasklist is
# the only usable probe here.
import subprocess
try:
out = subprocess.run(
["tasklist", "/FI", f"PID eq {int(pid)}", "/NH", "/FO", "CSV"],
capture_output = True,
text = True,
timeout = 10,
).stdout
except Exception:
# Unconfirmed means keep, matching the CLI's _pid_alive. Pruning a
# live server's record is what lets the next launch fall back past it
# and strand it, which is the bug this file exists to fix. A stale
# record instead costs one clear "already running" message.
return True
return f'"{int(pid)}"' in out
try:
os.kill(pid, 0)
except ProcessLookupError:
return False
except OSError:
return True
return True
def _process_create_time(pid: int) -> "float | None":
try:
import psutil
return psutil.Process(pid).create_time()
except Exception:
return None
def _read_pid_record(path: Path) -> "tuple[int, float | None, str | None] | None":
"""Parse ``pid`` / optional ``create_time`` / optional bind address."""
try:
lines = path.read_text(encoding = "utf-8").splitlines()
except (OSError, UnicodeDecodeError):
return None
if not lines or not lines[0].strip().isdigit():
return None
try:
# isdigit() is not enough: a superscript two passes it but int() rejects it.
pid = int(lines[0].strip())
except ValueError:
return None
# kill(0) signals our whole process group; kill(1) is init. Never either.
if pid < 2:
return None
created = None
if len(lines) > 1:
try:
created = float(lines[1].strip())
except ValueError:
created = None
address = lines[2].strip() if len(lines) > 2 and lines[2].strip() else None
return pid, created, address
def _pid_is_studio_backend(pid: int, created_times: "Sequence[float | None]" = ()) -> bool:
"""False only when a recorded start time proves this PID is a different process.
Any recorded time matching is enough -- a stale record must not veto a live
server that reused the PID. Untimed records cannot be checked at all, so they
are trusted: a legacy `python run.py` has no telltale argv, and guessing from
the command line rejected real servers.
"""
known = [c for c in created_times if c is not None]
if not known:
return True
actual = _process_create_time(pid)
if actual is None:
return True
return any(abs(actual - c) < 1.0 for c in known)
def _own_studio_on_port(port: int, host: str) -> "int | None":
"""PID of one of our own servers already bound to *port* for *host*.
Reads our own records rather than enumerating listeners: psutil is optional,
and without it a listener scan finds nothing and we silently start a duplicate.
"""
try:
paths = list(_studio_root().glob(f"studio-{port}-*.pid"))
except OSError:
return None
for path in paths:
record = _read_pid_record(path)
if record is None:
continue
pid, created, address = record
if not _pid_alive(pid):
# Pruning is a courtesy; an undeletable record must not abort startup.
try:
path.unlink(missing_ok = True)
except OSError:
pass
continue
if not _addresses_collide(address, host, port):
continue
if _pid_is_studio_backend(pid, [created]):
return pid
return _legacy_studio_on_port(port)
def _legacy_studio_on_port(port: int) -> "int | None":
"""A pre-upgrade server recorded only its PID, so match it to the listener.
Falling back past one leaves it running while `_write_pid_file` overwrites the
only record of it. When the listener is unknowable, assume it is ours.
"""
record = _read_pid_record(_PID_FILE)
if record is None:
return None
pid, created, _address = record
if not _pid_alive(pid):
return None
# A current build writes a per-port file too, so its port is already known --
# and this port's records were just checked. Only count a record that still
# matches the live process: a stale one may just share a reused PID.
for other in _per_port_records():
if other and other[0] == pid and _pid_is_studio_backend(pid, [other[1]]):
return None
blocker = _get_pid_on_port(port)
if blocker is not None and blocker[0] != pid:
return None
if not _pid_is_studio_backend(pid, [created]):
return None
return pid
def _per_port_records() -> "list[tuple[int, float | None, str | None] | None]":
try:
return [_read_pid_record(p) for p in _studio_root().glob(PID_FILE_GLOB)]
except OSError:
return []
def _resolve_port(
host: str,
port: int,
avoid_own_studio: bool = True,
) -> int:
"""The requested port, or the next free one.
With ``avoid_own_studio`` this aborts rather than falling back past one of our
own servers, on *port* itself or anywhere in the fallback range: skipping one
is what strands it. Callers that read the bound port back pass False and keep
the plain fallback.
"""
if _is_port_free(host, port):
return port
if avoid_own_studio:
own = _own_studio_on_port(port, host)
if own is not None:
_abort_already_running(own, port)
return _find_free_port(host, port + 1, avoid_own_studio = avoid_own_studio)
def _abort_already_running(pid: int, port: int) -> "NoReturn":
print(
f"Error: Unsloth Studio is already running on port {port} (PID {pid}). Run "
"`unsloth studio stop` first, or start this one on a different --port.",
file = sys.stderr,
flush = True,
)
sys.exit(1)
# Direct backend launches bypass the CLI's env re-export; do it here for
# real custom roots so unsloth-zoo's import-time LLAMA_CPP_DEFAULT_DIR
@ -992,101 +770,23 @@ if _STUDIO_ROOT_RESOLVED != _LEGACY_STUDIO_ROOT:
os.environ.setdefault("UNSLOTH_IS_PRESENT", "1")
_OWN_PID_FILE: "Path | None" = None
def _write_pid_file(port: int, host: str = ""):
"""Record this PID under its own port so `stop` can find every server."""
global _OWN_PID_FILE
path = _pid_file_for_port(port)
def _write_pid_file():
"""Write the current process PID to the studio PID file."""
try:
path.parent.mkdir(parents = True, exist_ok = True)
_PID_FILE.parent.mkdir(parents = True, exist_ok = True)
_PID_FILE.write_text(str(os.getpid()), encoding = "utf-8")
except OSError:
pass
try:
# Start time pins the record to this process; the bind address tells a
# later launch whether this server would actually block it.
created = _process_create_time(os.getpid())
address = ",".join(sorted(_bind_addresses(host, port))) if host else ""
body = f"{os.getpid()}\n{'' if created is None else repr(created)}\n{address}"
# Write-then-rename: `stop` reads these concurrently, and a reader that
# catches the truncate window sees a corrupt record and deletes it.
tmp = path.with_name(path.name + ".tmp")
try:
tmp.write_text(body, encoding = "utf-8")
os.replace(tmp, path)
finally:
# A failed replace would otherwise leave the scratch file behind. It
# does not end in .pid, so no glob picks it up either way.
tmp.unlink(missing_ok = True)
except OSError:
pass
else:
_OWN_PID_FILE = path
# An older CLI's `stop` only reads this one, and expects a bare PID. Written
# independently of the per-port record: if that one failed, this is the only
# thing keeping the server stoppable at all.
try:
# Never take it from a server that is still running. A pre-upgrade server
# is recorded here and nowhere else, so overwriting its entry is exactly
# what strands it -- the orphan this file exists to prevent.
prior = _read_pid_record(_PID_FILE) if _PID_FILE.is_file() else None
if prior is None or prior[0] == os.getpid() or not _pid_alive(prior[0]):
_PID_FILE.write_text(str(os.getpid()), encoding = "utf-8")
except OSError:
pass
def _legacy_heir() -> "int | None":
"""Another live server's PID, to hand the legacy studio.pid over to.
Only one server owns studio.pid at a time, so its exit would otherwise drop
the single record an older CLI can read, stranding any sibling that is still
serving.
"""
try:
paths = sorted(_studio_root().glob(PID_FILE_GLOB))
except OSError:
return None
for path in paths:
if _OWN_PID_FILE is not None and path == _OWN_PID_FILE:
continue
record = _read_pid_record(path)
if record is None or record[0] == os.getpid():
continue
if _pid_alive(record[0]) and _pid_is_studio_backend(record[0], [record[1]]):
return record[0]
return None
def _remove_pid_file():
"""Remove the PID files that belong to this process.
_PID_FILE is checked even when the per-port record was never written, since
_write_pid_file writes the two independently.
"""
# Nothing here may raise: _graceful_shutdown calls this at the end, and an
# unreadable or undeletable record must not abandon the rest of the exit
# path. _read_pid_record already swallows OSError/UnicodeDecodeError.
if _OWN_PID_FILE is not None:
try:
record = _read_pid_record(_OWN_PID_FILE) if _OWN_PID_FILE.is_file() else None
if record is not None and record[0] == os.getpid():
_OWN_PID_FILE.unlink(missing_ok = True)
except OSError:
pass
"""Remove the PID file if it belongs to this process."""
try:
record = _read_pid_record(_PID_FILE) if _PID_FILE.is_file() else None
if record is not None and record[0] == os.getpid():
# Hand the pointer to a live sibling rather than deleting it. An
# older CLI reads only this file, so dropping it while another
# server is still up leaves that server unstoppable.
heir = _legacy_heir()
if heir is None:
if _PID_FILE.is_file():
stored = _PID_FILE.read_text(encoding = "utf-8").strip()
if stored == str(os.getpid()):
_PID_FILE.unlink(missing_ok = True)
else:
_PID_FILE.write_text(str(heir), encoding = "utf-8")
except OSError:
except (OSError, UnicodeDecodeError):
pass
@ -1096,6 +796,7 @@ def _graceful_shutdown(server = None):
Called from signal handlers to clean up children before exit. Critical on
Windows where atexit handlers are unreliable after Ctrl+C.
"""
_remove_pid_file()
logger.info("Graceful shutdown initiated -- cleaning up subprocesses...")
# 1. Shut down uvicorn (releases the listening socket).
@ -1148,9 +849,6 @@ def _graceful_shutdown(server = None):
except Exception as e:
logger.warning("Error in process-lifetime sweep: %s", e)
# Last: while cleanup runs the server is still alive, and dropping the record
# early leaves a retried `stop` or a new launch unable to find it.
_remove_pid_file()
logger.info("All subprocesses cleaned up")
@ -1628,8 +1326,7 @@ def _apply_supplied_password(password_value: "Optional[str]") -> None:
if not _auth_storage.requires_password_change(_admin):
print(
"Error: an Unsloth admin password is already set; --password only sets "
"the initial password. Change it in the UI, or run `unsloth studio "
"reset-password` for a new one.",
"the initial password. Run `unsloth studio reset-password` first.",
file = sys.stderr,
flush = True,
)
@ -1700,7 +1397,6 @@ def run_server(
enable_tools: "Optional[bool]" = None,
password: "Optional[str]" = None,
emit_tauri_port: bool = True,
abort_if_own_studio: "Optional[bool]" = None,
):
"""
Start the FastAPI server.
@ -1834,16 +1530,10 @@ def run_server(
)
# Auto-find a free port if the requested one is in use.
original_port = port
# Refusing rather than falling back is for callers that cannot follow us to
# the new port. `studio run` reads app.state.server_port back and the desktop
# app reads TAURI_PORT, so both should keep the plain fallback; only the
# bare launch, which has nothing but the banner, benefits from the refusal.
if abort_if_own_studio is None:
abort_if_own_studio = not api_only
port = _resolve_port(host, port, avoid_own_studio = abort_if_own_studio)
if port != original_port:
blocker = _get_pid_on_port(original_port)
if not _is_port_free(host, port):
original_port = port
blocker = _get_pid_on_port(port)
port = _find_free_port(host, port + 1)
if not silent:
print("")
print("=" * 50)
@ -2041,7 +1731,7 @@ def run_server(
(time.perf_counter() - boot_started) * 1000,
)
_write_pid_file(port, host)
_write_pid_file()
import atexit
atexit.register(_remove_pid_file)

View file

@ -0,0 +1,656 @@
# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""
Profile usage statistics derived from studio.db.
Read-only aggregation over rows the app already writes: chat threads/messages
(with their per-message ``metadata_json``) and training runs/metrics. Nothing
is recorded specifically for stats, so the numbers are only as complete as the
local history.
Token counts live inside each message's metadata blob, so they cannot be summed
in SQL portably (JSON1 is not guaranteed on every bundled SQLite). Rows are
streamed once in (thread, time) order and every metric is folded in that single
pass, then memoised against a (count, max created_at) fingerprint so reopening
the Profile tab is free until history changes.
"""
import json
import threading
import time
from datetime import date, datetime, timedelta, timezone
from zoneinfo import ZoneInfo, ZoneInfoNotFoundError
from typing import Any, Optional
from loggers import get_logger
from storage.studio_db import get_connection
logger = get_logger(__name__)
# Gaps longer than this end a "sitting at the keyboard" stretch: without the cap
# a thread reopened a week later would report a week-long chat.
SESSION_GAP_SECONDS = 30 * 60
# Cap on the daily activity series handed to the UI (the heatmap draws a year).
MAX_DAILY_DAYS = 366
# Widest real UTC offset is 14h; anything beyond that is a bad client value.
MAX_TZ_OFFSET_MINUTES = 14 * 60
# Top-N lists returned to the client.
TOP_MODELS = 8
RECENT_RUNS = 5
# Serve a memoised payload for this long even if the fingerprint is unchanged,
# so a chat that is mid-stream still refreshes reasonably promptly.
CACHE_TTL_SECONDS = 20.0
_cache_lock = threading.Lock()
_cache: dict[str, Any] = {"fingerprint": None, "expires_at": 0.0, "payload": None}
def _as_float(value: Any) -> Optional[float]:
"""Coerce JSON numbers defensively; metadata is written by the client.
json accepts integers of any width, and float() raises OverflowError past
~1e308, so one oversized counter would 500 the whole panel.
"""
if isinstance(value, bool) or value is None:
return None
if isinstance(value, (int, float)):
try:
number = float(value)
except (OverflowError, ValueError):
return None
return number if number == number and number not in (float("inf"), float("-inf")) else None
return None
def _as_int(value: Any) -> int:
number = _as_float(value)
if number is None or number < 0:
return 0
return int(number)
def _iso(day: date) -> str:
return day.isoformat()
def _clean_str(value: Any) -> str:
return value.strip() if isinstance(value, str) else ""
def _resolve_zone(tz_name: str, tz_offset_minutes: int):
"""The caller's zone, preferring an IANA name over a single offset.
A fixed offset is only correct for the half of the year the caller happens
to be in, so a winter message read during summer lands an hour out and can
cross midnight. An IANA name carries each date's own offset. The offset
stays as the fallback for callers that send no name, or hosts with no tzdata.
"""
if tz_name:
try:
return ZoneInfo(tz_name)
except (ValueError, ZoneInfoNotFoundError, OSError):
logger.debug("unknown timezone %r, falling back to fixed offset", tz_name)
return timezone(timedelta(minutes = -tz_offset_minutes))
def _local_stamp(created_at_ms: int, zone) -> Optional[datetime]:
"""Wall-clock time in the caller's timezone, not the server's.
created_at is a client-supplied integer that SQLite stores unchecked, so a
value outside datetime's range is possible. Drop that row rather than let
one bad import take down the whole panel.
"""
if created_at_ms <= 0:
return None
try:
return datetime.fromtimestamp(created_at_ms / 1000, tz = zone).replace(tzinfo = None)
except (ValueError, OverflowError, OSError):
return None
def _streaks(days: set[date], today: date) -> dict[str, Any]:
"""Current and longest run of consecutive active days.
The current streak survives a day that has not been used yet: a streak that
ended yesterday is still "live" until today is over.
Imported history or a skewed client clock can date rows in the future.
Those are dropped up front so they cannot pad the longest streak or be
reported as the last active day either.
"""
days = {day for day in days if day <= today}
if not days:
return {"current": 0, "longest": 0, "lastActiveDay": None}
ordered = sorted(days)
longest = 1
running = 1
for previous, current in zip(ordered, ordered[1:]):
running = running + 1 if current - previous == timedelta(days = 1) else 1
longest = max(longest, running)
last = ordered[-1]
current_streak = 0
if today - last <= timedelta(days = 1):
current_streak = 1
cursor = last
while cursor - timedelta(days = 1) in days:
cursor -= timedelta(days = 1)
current_streak += 1
return {"current": current_streak, "longest": longest, "lastActiveDay": _iso(last)}
def _model_label(model_id: str) -> str:
"""Last path segment of a repo id, e.g. ``unsloth/gpt-oss-20b`` -> ``gpt-oss-20b``."""
cleaned = model_id.strip().replace("\\", "/")
tail = cleaned.rstrip("/").split("/")[-1]
return tail or cleaned
class _MessageFold:
"""Accumulators for the single streaming pass over chat messages."""
def __init__(self) -> None:
self.threads: set[str] = set()
self.messages = 0
self.user_messages = 0
self.assistant_messages = 0
self.prompt_tokens = 0
self.completion_tokens = 0
self.total_tokens = 0
self.cached_tokens = 0
self.tool_calls = 0
self.attachments = 0
self.session_seconds = 0.0
self.longest_chat: dict[str, Any] = {
"threadId": None,
"title": None,
"seconds": 0.0,
"messages": 0,
}
self.by_day: dict[date, dict[str, Any]] = {}
self.models: dict[str, dict[str, Any]] = {}
self.speed_samples: list[float] = []
self.best_speed = 0.0
self.best_speed_model: Optional[str] = None
self.response_ms: list[float] = []
self.first_token_ms: list[float] = []
def note_model(self, model_id: str, tokens: int) -> None:
entry = self.models.setdefault(
model_id, {"id": model_id, "label": _model_label(model_id), "messages": 0, "tokens": 0}
)
entry["messages"] += 1
entry["tokens"] += tokens
def note_day(self, day: date, tokens: int, thread_id: str) -> None:
bucket = self.by_day.setdefault(day, {"tokens": 0, "messages": 0, "threads": set()})
bucket["tokens"] += tokens
bucket["messages"] += 1
bucket["threads"].add(thread_id)
def _fork_keepers(conn) -> dict[tuple[str, int, str], str]:
"""For each original message, the one clone elected to stand in for it.
A clone is normally ignored because the original is counted instead. Once
the original is gone, whether its thread was deleted or just that row was
pruned, the clones become the only record. Letting every sibling count them
would multiply the usage, so exactly one may.
Electing per message rather than per fork matters because fork_chat_thread
copies one parent_id branch, not the whole thread: sibling forks taken from
a retry and a regeneration hold different rows, and a per-fork winner would
silently drop whatever only the loser carries.
"""
rows = conn.execute(
"""
SELECT m.thread_id, m.created_at, m.role,
t.forked_from_thread_id AS source_id
FROM chat_messages m
JOIN chat_threads t ON t.id = m.thread_id
WHERE t.forked_from_thread_id IS NOT NULL
AND m.created_at < t.created_at
"""
)
best: dict[tuple[str, int, str], str] = {}
for row in rows:
key = (row["source_id"], _as_int(row["created_at"]), row["role"])
thread_id = row["thread_id"]
current = best.get(key)
# Any stable winner works; lowest id keeps the choice reproducible.
if current is None or thread_id < current:
best[key] = thread_id
return best
def _surviving_original_keys(conn) -> set[tuple[str, int, str]]:
"""Identity of every message still living in a thread that has been forked.
Clones get fresh ids, so there is nothing to join on. Within one thread the
timestamp and role are enough to recognise the row a clone was taken from.
"""
rows = conn.execute(
"""
SELECT m.thread_id, m.created_at, m.role
FROM chat_messages m
WHERE m.thread_id IN (
SELECT DISTINCT forked_from_thread_id FROM chat_threads
WHERE forked_from_thread_id IS NOT NULL
)
"""
)
return {(row["thread_id"], _as_int(row["created_at"]), row["role"]) for row in rows}
def _fold_messages(conn, zone) -> _MessageFold:
fold = _MessageFold()
keepers = _fork_keepers(conn)
surviving = _surviving_original_keys(conn)
rows = conn.execute(
"""
SELECT m.thread_id, m.role, m.metadata_json, m.attachments_json, m.created_at,
t.title, t.model_id, t.model_type, t.pair_id,
t.created_at AS thread_created_at, t.forked_from_thread_id
FROM chat_messages m
LEFT JOIN chat_threads t ON t.id = m.thread_id
ORDER BY m.thread_id, m.created_at
"""
)
current_thread: Optional[str] = None
thread_title: Optional[str] = None
thread_seconds = 0.0
thread_messages = 0
previous_created: Optional[int] = None
def close_thread() -> None:
if current_thread is None:
return
fold.session_seconds += thread_seconds
if thread_seconds > fold.longest_chat["seconds"]:
fold.longest_chat = {
"threadId": current_thread,
"title": thread_title,
"seconds": thread_seconds,
"messages": thread_messages,
}
for row in rows:
thread_id = row["thread_id"]
created_at = _as_int(row["created_at"])
if thread_id != current_thread:
close_thread()
current_thread = thread_id
thread_title = row["title"]
thread_seconds = 0.0
thread_messages = 0
previous_created = None
# Compare mode stores one thread per pane under a shared pair_id, and
# the sidebar shows them as a single conversation. Count them that way
# too, or one comparison inflates the chat total and drags the average
# tokens per chat down.
conversation_id = row["pair_id"] or thread_id
# A fork is its own visible conversation, so it counts towards the chat
# total from the moment it exists, before any new turn is added.
fold.threads.add(conversation_id)
# Forking clones the whole ancestry into the new thread, keeping each
# copy's original timestamp. Skip a clone while the row it was taken
# from is still there to be counted; once that original is gone, only
# the elected fork stands in for it, so nothing is lost or doubled.
source_id = row["forked_from_thread_id"]
if source_id and created_at < _as_int(row["thread_created_at"]):
original = (source_id, created_at, row["role"])
if original in surviving or keepers.get(original) != thread_id:
continue
fold.messages += 1
thread_messages += 1
if previous_created is not None:
gap = (created_at - previous_created) / 1000
if 0 < gap <= SESSION_GAP_SECONDS:
thread_seconds += gap
previous_created = created_at
stamp = _local_stamp(created_at, zone)
role = row["role"]
if role == "user":
fold.user_messages += 1
attachments_json = row["attachments_json"]
if attachments_json:
try:
parsed = json.loads(attachments_json)
if isinstance(parsed, list):
fold.attachments += len(parsed)
except (json.JSONDecodeError, TypeError):
pass
message_tokens = 0
metadata: Any = None
if role == "assistant":
fold.assistant_messages += 1
raw_metadata = row["metadata_json"]
if raw_metadata:
try:
metadata = json.loads(raw_metadata)
except (json.JSONDecodeError, TypeError):
metadata = None
if isinstance(metadata, dict):
usage = metadata.get("contextUsage")
timing = metadata.get("timing")
usage = usage if isinstance(usage, dict) else {}
timing = timing if isinstance(timing, dict) else {}
prompt_tokens = _as_int(usage.get("promptTokens"))
completion_tokens = _as_int(usage.get("completionTokens"))
total_tokens = _as_int(usage.get("totalTokens"))
if completion_tokens == 0:
# Local engines occasionally omit the usage chunk; the adapter's
# own token count is the next best estimate.
completion_tokens = _as_int(timing.get("tokenCount"))
if total_tokens == 0:
total_tokens = prompt_tokens + completion_tokens
fold.prompt_tokens += prompt_tokens
fold.completion_tokens += completion_tokens
fold.total_tokens += total_tokens
fold.cached_tokens += _as_int(usage.get("cachedTokens"))
fold.tool_calls += _as_int(timing.get("toolCallCount"))
message_tokens = total_tokens
# responseDetails carries the model that actually answered, which
# differs from the requested checkpoint whenever a provider routes
# or resolves an alias. contextUsage.modelId is the request, so it
# is only the fallback. The thread's model_id is never used: it
# tracks the current selection, not the one that ran.
details = metadata.get("responseDetails")
details = details if isinstance(details, dict) else {}
model_id = _clean_str(details.get("responseModelId")) or _clean_str(
usage.get("modelId")
)
if model_id:
fold.note_model(model_id, message_tokens)
speed = _as_float(timing.get("tokensPerSecond"))
# llama.cpp reports absurd rates on no-op turns; ignore those.
if speed is not None and 0 < speed < 100_000:
fold.speed_samples.append(speed)
if speed > fold.best_speed:
fold.best_speed = speed
fold.best_speed_model = _model_label(model_id) if model_id else None
stream_ms = _as_float(timing.get("totalStreamTime"))
if stream_ms is not None and stream_ms > 0:
fold.response_ms.append(stream_ms)
# firstTokenTime is already an elapsed duration, not a timestamp.
first_token = _as_float(timing.get("firstTokenTime"))
if first_token is not None and first_token > 0:
fold.first_token_ms.append(first_token)
if stamp is not None:
fold.note_day(stamp.date(), message_tokens, conversation_id)
close_thread()
return fold
def _daily_series(fold: _MessageFold, today: date, days: int) -> list[dict[str, Any]]:
"""Dense day-by-day series so the heatmap can index straight into it."""
start = today - timedelta(days = days - 1)
series: list[dict[str, Any]] = []
for offset in range(days):
day = start + timedelta(days = offset)
bucket = fold.by_day.get(day)
series.append(
{
"date": _iso(day),
"tokens": int(bucket["tokens"]) if bucket else 0,
"messages": int(bucket["messages"]) if bucket else 0,
"chats": len(bucket["threads"]) if bucket else 0,
}
)
return series
def _superseded(prefix: str = "r.") -> str:
"""SQL for "a later run resumed from this one, so its counters live there".
``prefix`` must qualify the outer row: the EXISTS subquery selects from the
same table, so a bare column name would bind to the subquery instead.
``create_run``'s resume claim sets ``resume_blocked`` and leaves
``output_dir`` alone. Cancelling clears ``output_dir`` while setting the
same flag, so the flag alone cannot tell the two apart.
``delete_run`` never clears the flag, so the continuation has to still be
there. Otherwise deleting it would strand the source at zero while its row
and metrics stay visible in history.
The continuation also has to have reached the source's step. ``create_run``
claims the source the moment a resume starts, but ``final_step`` is only
written on the first metric flush, so a continuation that fails before then
would take the source's completed work down with it.
"""
return f"""
{prefix}resume_blocked = 1
AND {prefix}output_dir IS NOT NULL
AND EXISTS (
SELECT 1 FROM training_runs continuation
WHERE continuation.output_dir = {prefix}output_dir
AND continuation.started_at > {prefix}started_at
AND COALESCE(continuation.final_step, 0)
>= COALESCE({prefix}final_step, 0)
)
"""
def _training_stats(conn) -> dict[str, Any]:
row = conn.execute(
"""
SELECT COUNT(*) AS runs,
SUM(CASE WHEN status = 'completed' THEN 1 ELSE 0 END) AS completed,
SUM(COALESCE(duration_seconds, 0)) AS seconds,
COUNT(DISTINCT model_name) AS models,
COUNT(DISTINCT dataset_name) AS datasets,
MIN(final_loss) AS best_loss
FROM training_runs
"""
).fetchone()
# A resumed run continues its source's step and token counters from the
# checkpoint, so both absolute totals already include the source's work.
# Only a run superseded by a resume is dropped: create_run's claim sets
# resume_blocked while leaving output_dir intact, whereas cancelling clears
# output_dir, so a cancelled run keeps contributing the work it did do.
steps = conn.execute(
f"SELECT COALESCE(SUM(r.final_step), 0) FROM training_runs r WHERE NOT ({_superseded()})"
).fetchone()[0]
# num_tokens is state.num_input_tokens_seen, a running total logged at each
# step, so summing the samples multiplies the real figure. Take each run's
# final counter, the same value get_run_metrics reports.
tokens = conn.execute(
f"""
SELECT COALESCE(SUM(run_tokens), 0) FROM (
SELECT MAX(m.num_tokens) AS run_tokens
FROM training_metrics m
JOIN training_runs r ON r.id = m.run_id
WHERE NOT ({_superseded("r.")})
GROUP BY m.run_id
)
"""
).fetchone()[0]
recent = conn.execute(
"""
SELECT id, display_name, model_name, dataset_name,
status, final_loss, final_step, duration_seconds, started_at
FROM training_runs
ORDER BY started_at DESC
LIMIT ?
""",
(RECENT_RUNS,),
).fetchall()
return {
"runs": _as_int(row["runs"]),
"completed": _as_int(row["completed"]),
"steps": _as_int(steps),
"tokens": _as_int(tokens),
"seconds": _as_float(row["seconds"]) or 0.0,
"models": _as_int(row["models"]),
"datasets": _as_int(row["datasets"]),
"bestLoss": _as_float(row["best_loss"]),
"recent": [
{
"id": item["id"],
# A renamed run keeps the name the user gave it; otherwise fall
# back to the short model label rather than the full repo id.
"name": _clean_str(item["display_name"]) or _model_label(item["model_name"] or ""),
"modelLabel": _model_label(item["model_name"] or ""),
"datasetLabel": _model_label(item["dataset_name"] or ""),
"status": item["status"],
"finalLoss": _as_float(item["final_loss"]),
"steps": _as_int(item["final_step"]),
"seconds": _as_float(item["duration_seconds"]) or 0.0,
"startedAt": item["started_at"],
}
for item in recent
],
}
def _fingerprint(conn) -> tuple:
message_row = conn.execute(
"SELECT COUNT(*), COALESCE(MAX(created_at), 0) FROM chat_messages"
).fetchone()
run_row = conn.execute(
"SELECT COUNT(*), COALESCE(MAX(started_at), '') FROM training_runs"
).fetchone()
return (message_row[0], message_row[1], run_row[0], run_row[1])
def compute_profile_stats(
days: int = MAX_DAILY_DAYS,
tz_offset_minutes: int = 0,
tz_name: str = "",
) -> dict[str, Any]:
"""Aggregate every profile statistic in one pass, memoised per history state."""
days = max(1, min(int(days), MAX_DAILY_DAYS))
tz_offset_minutes = max(
-MAX_TZ_OFFSET_MINUTES, min(int(tz_offset_minutes), MAX_TZ_OFFSET_MINUTES)
)
zone = _resolve_zone(tz_name, tz_offset_minutes)
conn = get_connection()
try:
fingerprint = (_fingerprint(conn), days, tz_offset_minutes, tz_name)
now = time.monotonic()
with _cache_lock:
if (
_cache["payload"] is not None
and _cache["fingerprint"] == fingerprint
and _cache["expires_at"] > now
):
return _cache["payload"]
started = time.perf_counter()
fold = _fold_messages(conn, zone)
training = _training_stats(conn)
# "Today" has to match the buckets above, or the newest column and the
# current streak drift by a day whenever the caller is elsewhere.
today = (_local_stamp(int(time.time() * 1000), zone) or datetime.now()).date()
streak = _streaks(set(fold.by_day.keys()), today)
daily = _daily_series(fold, today, days)
peak_day = max(fold.by_day.items(), key = lambda item: item[1]["tokens"], default = None)
models = sorted(
fold.models.values(), key = lambda item: (item["tokens"], item["messages"]), reverse = True
)[:TOP_MODELS]
speed_samples = fold.speed_samples
payload = {
"generatedAt": int(time.time() * 1000),
"days": days,
"totals": {
"threads": len(fold.threads),
"messages": fold.messages,
"userMessages": fold.user_messages,
"assistantMessages": fold.assistant_messages,
"promptTokens": fold.prompt_tokens,
"completionTokens": fold.completion_tokens,
"totalTokens": fold.total_tokens,
"cachedTokens": fold.cached_tokens,
"toolCalls": fold.tool_calls,
"attachments": fold.attachments,
"activeDays": len(fold.by_day),
"chatSeconds": round(fold.session_seconds),
},
"streak": streak,
"peakDay": (
{"date": _iso(peak_day[0]), "tokens": int(peak_day[1]["tokens"])}
if peak_day and peak_day[1]["tokens"] > 0
else None
),
"longestChat": (
{
"threadId": fold.longest_chat["threadId"],
"title": fold.longest_chat["title"],
"seconds": round(fold.longest_chat["seconds"]),
"messages": fold.longest_chat["messages"],
}
if fold.longest_chat["seconds"] > 0
else None
),
"daily": daily,
"models": models,
"speed": {
"averageTokensPerSecond": (
sum(speed_samples) / len(speed_samples) if speed_samples else None
),
"bestTokensPerSecond": fold.best_speed or None,
"bestTokensPerSecondModel": fold.best_speed_model,
"averageResponseMs": (
sum(fold.response_ms) / len(fold.response_ms) if fold.response_ms else None
),
"averageFirstTokenMs": (
sum(fold.first_token_ms) / len(fold.first_token_ms)
if fold.first_token_ms
else None
),
"samples": len(speed_samples),
},
"training": training,
}
logger.debug(
"profile stats computed in %.1f ms (%d messages)",
(time.perf_counter() - started) * 1000,
fold.messages,
)
with _cache_lock:
_cache["fingerprint"] = fingerprint
_cache["expires_at"] = time.monotonic() + CACHE_TTL_SECONDS
_cache["payload"] = payload
return payload
finally:
conn.close()
def invalidate_profile_stats_cache() -> None:
"""Drop the memoised payload (used by tests and after history wipes)."""
with _cache_lock:
_cache["fingerprint"] = None
_cache["expires_at"] = 0.0
_cache["payload"] = None

View file

@ -1,300 +0,0 @@
{"id": "7cdc94d4ad818e68", "model": "llama3.2-1b", "prompt_class": "needs_tool_python", "prompt": "Sort this list and tell me the median: 8, 3, 91, 12, 7, 44, 2.", "text": "import statistics\n# Define the list of numbers\nnumbers = [8, 3, 91, 12, 7, 44, 2]\n# Sort the list in ascending order\nsorted_numbers = sorted(numbers)\n# Calculate the median\nmedian = statistics.median(sorted_numbers)\n# Return the result in a JSON format\nreturn", "retry_text": "", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "300cc2c42f91bf35", "model": "llama3.2-1b", "prompt_class": "needs_tool_python", "prompt": "Sort this list and tell me the median: 8, 3, 91, 12, 7, 44, 2.", "text": "import json\nimport statistics\n\n# Define the list of numbers\nnumbers = [8, 3, 91, 12, 7, 44, 2]\n\n# Sort the list in ascending order\nnumbers.sort()\n\n# Calculate the median\nmedian = statistics.median(numbers)\n\n# Convert the result to a JSON string\nresult = json.dumps(", "retry_text": "web_search(query=\"statistics.median\", url=\"https://www.example.com\")", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "d35c286e3b43a3bb", "model": "llama3.2-1b", "prompt_class": "needs_tool_python", "prompt": "Sort this list and tell me the median: 8, 3, 91, 12, 7, 44, 2.", "text": "import json\nimport statistics\n\n# Define the function calls\nfunction_calls = [", "retry_text": "import json\n\ndef web_search(query, url):\n # This is a mock web_search function\n print(f\"Searching for", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "31ea3bc4cfb9df5d", "model": "llama3.2-1b", "prompt_class": "no_tool", "prompt": "What is the difference between a list and a tuple in Python?", "text": "import json\n\n# List and tuple are both mutable, meaning they can be changed after creation.\n# They are also both ordered collections, meaning they maintain the order in which elements were added.\n\n# Create a list\nmy_list = [1, 2, 3, 4, 5]\n\n# Create a tuple\nmy_tuple = (1, 2, 3, 4, 5)\n\n# Modifying a list\nmy_list.append(6)\nprint(my_list) # Output: [1, 2, 3, 4, 5, 6]\n\n# Modifying a tuple\nmy_tuple = (1, 2, 3, 4, 5)\nmy_tuple.append(6)\nprint(my_tuple) # Output: (1, 2, 3, 4, 5, 6)\n\n# Accessing elements in a list\nprint(my_list[0]) # Output: 1\n\n# Accessing elements in a tuple\nprint(my_tuple[0]) # Output: 1\n\n# Modifying a list\nmy_list.insert(1, 7)\nprint(my_list) # Output: [1, 7, 2, 3, 4, 5]\n\n# Modifying a tuple\nmy_tuple.insert(1, 7)\nprint(my_tuple) # Output: (1, 7, 2, 3, 4, 5)\n\n# Creating a list and a tuple\nmy_list2 = []\nmy_tuple2 = ()\n\n# Modifying a list\nmy_list2.append(8)\nprint(my_list2) # Output: [1, 2, 3, 4, 5, 6, 8]\n\n# Modifying a tuple\nmy_tuple2 = (1, 2, 3, 4, 5)\nmy_tuple2.append(8)\nprint(my_tuple2) # Output: (1, 2, 3, 4, 5, 8)", "retry_text": "", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "2c701ec6cbbad60f", "model": "llama3.2-1b", "prompt_class": "format", "prompt": "Describe your approach to reviewing a pull request, as a list.", "text": "import json\n\n# Define the functions as described in the prompt\ndef web_search(query, url):\n return", "retry_text": "import web_search\n\nweb_search(\"example query\", \"https://www.example.com\")", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "17930edef7c0a391", "model": "llama3.2-1b", "prompt_class": "non_english", "prompt": "Explica en una frase que es el sobreajuste.", "text": "El sobreajuste se refiere a un problema en el que dos o más variables tienen valores equivalentes o cercanos, lo que puede provocar errores en el ejecución de una programación o comprensión de un lenguaje. Este sobreajuste puede ocurrir en diversas situaciones, como en la programación con variables dinámicas, en la resolución de ecuaciones cuadráticas, o en la interpretación de lenguajes de programación.", "retry_text": "", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "4fdf373b03b9674c", "model": "qwen3-0.6b", "prompt_class": "needs_tool_web", "prompt": "Who won the most recent Formula 1 race?", "text": "The most recent Formula 1 race was held in **Monaco** in the 2023 season, and the winner was **Lohman**.", "retry_text": "The most recent Formula 1 race was held in **Monaco** in the 2023 season, and the winner was **Lohman**.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "2f93cfb96347d03b", "model": "qwen3-0.6b", "prompt_class": "needs_tool_python", "prompt": "Sort this list and tell me the median: 8, 3, 91, 12, 7, 44, 2.", "text": "Sorted list: [2, 3, 7, 8, 12, 44, 91] \n**Median**: 8", "retry_text": "Sorted list: [2, 3, 7, 8, 12, 44, 91] \n**Median**: 8", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "ffe902b3a77da3a3", "model": "qwen3-0.6b", "prompt_class": "needs_tool_python", "prompt": "Sort this list and tell me the median: 8, 3, 91, 12, 7, 44, 2.", "text": "Sorted list: [2, 3, 7, 8, 12, 44, 91] \nMedian: 8 \n\nThe median is the middle number when the list has an odd number of elements.", "retry_text": "Sorted list: [2, 3, 7, 8, 12, 44, 91] \nMedian: 8 \n\nThe median is the middle number when the list has an odd number of elements.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "43700884f9e483bb", "model": "qwen3-0.6b", "prompt_class": "needs_tool_python", "prompt": "Sort this list and tell me the median: 8, 3, 91, 12, 7, 44, 2.", "text": "The sorted list is [2, 3, 7, 8, 12, 44, 91]. The median is 8.", "retry_text": "The sorted list is [2, 3, 7, 8, 12, 44, 91]. The median is 8.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "bb267a38ffd74bb8", "model": "qwen3-0.6b", "prompt_class": "needs_tool_python", "prompt": "Sort this list and tell me the median: 8, 3, 91, 12, 7, 44, 2.", "text": "The sorted list is [2, 3, 7, 8, 12, 44, 91], and the median is 12.", "retry_text": "The sorted list is [2, 3, 7, 8, 12, 44, 91], and the median is 8.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "b8428bd477953912", "model": "qwen3-0.6b", "prompt_class": "needs_tool_python", "prompt": "Sort this list and tell me the median: 8, 3, 91, 12, 7, 44, 2.", "text": "Sorted list: [2, 3, 7, 8, 12, 44, 91] \nMedian: 7", "retry_text": "Sorted list: [2, 3, 7, 8, 12, 44, 91] \nMedian: 8", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "a6257c91990092ea", "model": "qwen3-0.6b", "prompt_class": "needs_tool_python", "prompt": "Sort this list and tell me the median: 8, 3, 91, 12, 7, 44, 2.", "text": "Sorted list: 2, 3, 7, 8, 12, 44, 91 \nMedian: 8 (the middle value when the list has an odd number of elements)", "retry_text": "Sorted list: 2, 3, 7, 8, 12, 44, 91 \nMedian: 8", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "8ca73b260971a970", "model": "qwen3-0.6b", "prompt_class": "needs_tool_python", "prompt": "Sort this list and tell me the median: 8, 3, 91, 12, 7, 44, 2.", "text": "Sorted list: 2, 3, 7, 8, 12, 44, 91 \nMedian: 8", "retry_text": "Sorted list: 2, 3, 7, 8, 12, 44, 91 \nMedian: 8", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "2cf3c930c6256595", "model": "qwen3-0.6b", "prompt_class": "needs_tool_python", "prompt": "Sort this list and tell me the median: 8, 3, 91, 12, 7, 44, 2.", "text": "Sorted list: [2, 3, 7, 8, 12, 44, 91] \nMedian: 12 \n\nThe median is 12.", "retry_text": "Sorted list: [2, 3, 7, 8, 12, 44, 91] \nMedian: 12 \n\nThe median is 12.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "c094102b1a6f9645", "model": "qwen3-0.6b", "prompt_class": "needs_tool_python", "prompt": "Calculate the standard deviation of 4, 8, 15, 16, 23, 42.", "text": "The standard deviation of the dataset 4, 8, 15, 16, 23, and 42 is approximately 12.32. \n\n**Step-by-Step Explanation:**\n1. **Calculate the Mean:** \n Sum of the numbers = 108, divided by 6 (count) = **18**. \n2. **Calculate Squared Differences:** \n (4-18)² = 196, (8-18)² = 100, (15-18)² = 9, (16-18)² = 4, (23-18)² = 25, (42-18)² = 576. \n Sum of squared differences = **910**. \n3. **Compute Variance:** \n Variance = 910 / 6 ≈ **151.666...**. \n4. **Take Square Root:** \n Standard Deviation ≈ √151.666... ≈ **12.32**. \n\n**Answer:** The standard deviation is approximately **12.32**.", "retry_text": "The standard deviation of the dataset 4, 8, 15, 16, 23, and 42 is approximately **12.32**. \n\n**Step-by-Step Explanation:**\n1. **Calculate the Mean:** \n Sum of the numbers = 108, divided by 6 = **18**. \n2. **Calculate Squared Differences:** \n (4-18)² = 196, (8-18)² = 100, (15-18)² = 9, (16-18)² = 4, (23-18)² = 25, (42-18)² = 576. \n Sum of squared differences = **910**. \n3. **Compute Variance:** \n Variance = 910 / 6 ≈ **151.666...**. \n4. **Take Square Root:** \n Standard Deviation ≈ √151.666... ≈ **12.32**. \n\n**Answer:** The standard deviation is approximately **12.32**.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "82ad4f08c9443c13", "model": "qwen3-0.6b", "prompt_class": "needs_tool_python", "prompt": "Calculate the standard deviation of 4, 8, 15, 16, 23, 42.", "text": "The standard deviation of the numbers 4, 8, 15, 16, 23, and 42 is approximately 12.32.\n\n**Step-by-Step Explanation:**\n1. **Calculate the mean:** \n $ \\frac{4 + 8 + 15 + 16 + 23 + 42}{6} = \\frac{108}{6} = 18 $.\n\n2. **Find squared deviations from the mean:** \n $ (4 - 18)^2 = 196 $, $ (8 - 18)^2 = 100 $, $ (15 - 18)^2 = 9 $, $ (16 - 18)^2 = 4 $, $ (23 - 18)^2 = 25 $, $ (42 - 18)^2 = 576 $.\n\n3. **Sum squared deviations:** \n $ 196 + 100 + 9 + 4 + 25 + 576 = 910 $.\n\n4. **Calculate variance:** \n $ \\frac{910}{6} \\approx 151.67 $.\n\n5. **Take the square root of variance:** \n $ \\sqrt{151.67} \\approx 12.32 $.\n\n**Answer:** The standard deviation is approximately 12.32.", "retry_text": "The standard deviation of the numbers 4, 8, 15, 16, 23, and 42 is approximately 12.32.\n\n**Step-by-Step Explanation:**\n1. **Calculate the mean:** \n $ \\frac{4 + 8 + 15 + 16 + 23 + 42}{6} = \\frac{108}{6} = 18 $.\n\n2. **Find squared deviations from the mean:** \n $ (4 - 18)^2 = 196 $, $ (8 - 18)^2 = 100 $, $ (15 - 18)^2 = 9 $, $ (16 - 18)^2 = 4 $, $ (23 - 18)^2 = 25 $, $ (42 - 18)^2 = 576 $.\n\n3. **Sum squared deviations:** \n $ 196 + 100 + 9 + 4 + 25 + 576 = 910 $.\n\n4. **Calculate variance:** \n $ \\frac{910}{6} \\approx 151.67 $.\n\n5. **Take the square root of variance:** \n $ \\sqrt{151.67} \\approx 12.32 $.\n\n**Answer:** The standard deviation is approximately 12.32.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "434d3a6198ee764a", "model": "qwen3-0.6b", "prompt_class": "needs_tool_python", "prompt": "Calculate the standard deviation of 4, 8, 15, 16, 23, 42.", "text": "The standard deviation of the dataset 4, 8, 15, 16, 23, 42 is approximately 12.32. \n\n**Answer:** The standard deviation is approximately 12.32.", "retry_text": "The standard deviation of the dataset 4, 8, 15, 16, 23, and 42 is approximately **12.32**. \n\n**Answer:** The standard deviation is approximately 12.32.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "9f22d6beb4d34f32", "model": "qwen3-0.6b", "prompt_class": "needs_tool_python", "prompt": "Calculate the standard deviation of 4, 8, 15, 16, 23, 42.", "text": "The standard deviation of the dataset [4, 8, 15, 16, 23, 42] is approximately 12.32. \n\n**Steps:**\n1. **Calculate the mean**: (4 + 8 + 15 + 16 + 23 + 42) / 6 = 108 / 6 = 18.\n2. **Compute squared differences from the mean**: \n - (4-18)² = 196, (8-18)² = 100, (15-18)² = 9, (16-18)² = 4, (23-18)² = 25, (42-18)² = 576.\n3. **Sum squared differences**: 196 + 100 + 9 + 4 + 25 + 576 = 910.\n4. **Divide by count (6)**: 910 / 6 ≈ 151.666.\n5. **Take square root**: √(151.666...) ≈ 12.32.", "retry_text": "The standard deviation of the dataset [4, 8, 15, 16, 23, 42] is approximately **12.32**.\n\n**Steps:**\n1. **Mean**: (4 + 8 + 15 + 16 + 23 + 42) / 6 = 108 / 6 = 18.\n2. **Squared differences**: (4-18)² = 196, (8-18)² = 100, (15-18)² = 9, (16-18)² = 4, (23-18)² = 25, (42-18)² = 576.\n3. **Sum of squared differences**: 196 + 100 + 9 + 4 + 25 + 576 = 910.\n4. **Divide by count**: 910 / 6 ≈ 151.666.\n5. **Square root**: √(151.666...) ≈ 12.32.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "b612ea05e2b54573", "model": "qwen3-0.6b", "prompt_class": "needs_tool_python", "prompt": "How many primes are there below 10000?", "text": "The number of primes below 10000 is **1684**. This value is derived using the Sieve of Eratosthenes algorithm, which efficiently marks non-prime numbers by eliminating multiples of prime numbers, ensuring accuracy without unnecessary computations.", "retry_text": "The number of primes below 10,000 is **1684**. This value is derived using the Sieve of Eratosthenes algorithm, which efficiently marks non-prime numbers by eliminating their multiples, ensuring accuracy without unnecessary computations.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "eead70908f20af93", "model": "qwen3-0.6b", "prompt_class": "needs_tool_python", "prompt": "Work out the compound interest on 5000 at 3.5% for 12 years.", "text": "The compound interest on 5000 at 3.5% for 12 years can be calculated using the formula:\n\n$$ A = P \\times (1 + r)^t $$\n\nWhere:\n- $ P = 5000 $ (principal)\n- $ r = 3.5\\% = 0.035 $\n- $ t = 12 $ (time in years)\n\nFirst, compute $ (1 + 0.035)^{12} \\approx 1.4533 $. Then, the total amount $ A $ is:\n\n$$ 5000 \\times 1.4533 \\approx 7266.5 $$\n\nThe compound interest is:\n\n$$ 7266.5 - 5000 = 2266.5 $$\n\n**Answer:** The compound interest is approximately **$2266.50**.", "retry_text": "The compound interest on $5000 at 3.5% for 12 years is approximately **$2266.50**.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "98e626024169106e", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is the capital of Japan?", "text": "The capital of Japan is Tokyo.", "retry_text": "The capital of Japan is Tokyo.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "5d703c2433475f71", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is the capital of Japan?", "text": "The capital of Japan is Tokyo.", "retry_text": "The capital of Japan is Tokyo.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "90ad72f79996edab", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is the capital of Japan?", "text": "The capital of Japan is Tokyo.", "retry_text": "The capital of Japan is Tokyo.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "114e4cb212318cc7", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is the capital of Japan?", "text": "The capital of Japan is Tokyo.", "retry_text": "The capital of Japan is Tokyo.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "e1d5199f79924aa7", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is the capital of Japan?", "text": "The capital of Japan is Tokyo.", "retry_text": "The capital of Japan is Tokyo.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "70d28a74a973cccc", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is the capital of Japan?", "text": "The capital of Japan is Tokyo.", "retry_text": "The capital of Japan is Tokyo.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "455e9970d6a3ec42", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is the capital of Japan?", "text": "The capital of Japan is Tokyo.", "retry_text": "The capital of Japan is Tokyo.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "82f8b8f25827ef49", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is the capital of Japan?", "text": "The capital of Japan is Tokyo.", "retry_text": "The capital of Japan is Tokyo.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "8806bfc3a37d5c41", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Explain what a mutex is in two sentences.", "text": "A mutex (mutual exclusion object) is a synchronization mechanism that ensures only one thread can access a shared resource at a time, preventing race conditions. It works by locking the resource when a thread accesses it and unlocking it when another thread locks it, thus controlling concurrent access.", "retry_text": "A mutex (mutual exclusion object) is a synchronization mechanism that ensures only one thread can access a shared resource at a time, preventing race conditions by locking the resource when a thread accesses it and unlocking it when another thread locks it.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "0259d6328e80d33e", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Explain what a mutex is in two sentences.", "text": "A mutex (mutual exclusion object) is a synchronization primitive that ensures only one thread can access a shared resource at a time, preventing race conditions. It works by blocking other threads from accessing the resource until one is signaled, ensuring thread safety in concurrent programming.", "retry_text": "A mutex (mutual exclusion object) is a synchronization primitive that ensures only one thread can access a shared resource at a time, preventing race conditions. It works by blocking other threads from accessing the resource until one is signaled, ensuring thread safety in concurrent programming.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "ee78b728ee51cc8c", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Explain what a mutex is in two sentences.", "text": "A mutex is a synchronization primitive used to protect shared resources from concurrent access. It ensures that only one thread at a time can access the resource, preventing race conditions.", "retry_text": "A mutex is a synchronization primitive used to protect shared resources from concurrent access. It ensures that only one thread at a time can access the resource, preventing race conditions.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "c931946998269339", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Explain what a mutex is in two sentences.", "text": "A mutex is a synchronization primitive used to protect shared resources from concurrent access, ensuring that only one thread can access them at a time. It achieves this by locking the resource when a thread needs to access it and unlocking it when the resource is free, preventing race conditions.", "retry_text": "A mutex is a synchronization primitive used to protect shared resources from concurrent access, ensuring that only one thread can access them at a time. It achieves this by locking the resource when a thread needs to access it and unlocking it when the resource is free, preventing race conditions.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "9a5a6581c9019f7f", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Explain what a mutex is in two sentences.", "text": "A mutex (mutual exclusive lock) is a synchronization primitive used to control shared resources, ensuring that multiple threads or processes can access them without interference. It prevents race conditions by enforcing exclusive access to the resource at a time.", "retry_text": "A mutex (mutual exclusive lock) is a synchronization primitive used to control shared resources, ensuring that multiple threads or processes can access them without interference. It prevents race conditions by enforcing exclusive access to the resource at a time.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "a12923f96ffafa6e", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Explain what a mutex is in two sentences.", "text": "A mutex (mutual exclusion object) is a synchronization primitive used to protect shared resources by ensuring that only one thread can access them at a time. It prevents race conditions by allowing one thread to acquire the lock, and another thread must relinquish the lock before it can proceed. \n\nIn two sentences: \nA mutex is a synchronization object used to protect shared resources, ensuring that only one thread can access them at a time to avoid race conditions.", "retry_text": "A mutex (mutual exclusion object) is a synchronization primitive used to protect shared resources, ensuring that only one thread can access them at a time to avoid race conditions.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "e246df27c5212f5b", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Explain what a mutex is in two sentences.", "text": "A mutex (mutual exclusive lock) is a synchronization mechanism that protects shared resources from concurrent access by ensuring only one thread can access them at a time. It prevents race conditions by allowing only one thread to acquire the lock before others can do so. \n\nAnother way: A mutex is a tool that ensures thread safety when multiple threads access shared resources, preventing interference and ensuring data consistency.", "retry_text": "A mutex (mutual exclusive lock) is a synchronization mechanism that protects shared resources from concurrent access by ensuring only one thread can access them at a time. It prevents race conditions by allowing only one thread to acquire the lock before others can do so. \n\nAnother way: A mutex is a tool that ensures thread safety when multiple threads access shared resources, preventing interference and ensuring data consistency.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "55adde375e13956d", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Explain what a mutex is in two sentences.", "text": "A mutex is a synchronization primitive used to protect shared resources from concurrent access. It ensures that only one thread can access the resource at a time, preventing race conditions by controlling access to shared data.", "retry_text": "A mutex is a synchronization primitive used to protect shared resources from concurrent access. It ensures that only one thread can access the resource at a time, preventing race conditions by controlling access to shared data.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "f3ae8b3e9485a3f7", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "In plain English, what is a transformer model?", "text": "A transformer model is a type of neural network used in machine learning and natural language processing to handle long-range dependencies in sequences. Unlike traditional models like recurrent or RNNs, transformers use self-attention mechanisms to process the input in a way that allows the model to focus on relevant parts of the sequence, making them more efficient for tasks like language modeling and text generation.", "retry_text": "A transformer model is a type of neural network used in machine learning and natural language processing to handle long-range dependencies in sequences. Unlike traditional models like recurrent or RNNs, transformers use self-attention mechanisms to focus on relevant parts of the input, making them more efficient for tasks like language modeling and text generation.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "7501627df64f5901", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "In plain English, what is a transformer model?", "text": "A transformer model is a type of neural network designed to process long sequences of text, such as sentences or paragraphs, more efficiently than traditional models like RNNs or LSTMs. Here's a simple explanation:\n\n1. **Core Idea**: Transformers use self-attention mechanisms to focus on specific parts of the input at different times. Unlike traditional models, which process information sequentially, transformers can handle complex, long-term dependencies in text.\n\n2. **Key Advancement**: This allows models to understand and generate text even when the input is very long or contains complex relationships between words.\n\n3. **Common Use Cases**: Transformer models are widely used in tasks like language modeling (text generation, translation), summarization, and summarizing long documents.\n\nIn plain English, transformers enable powerful processing of text with greater efficiency and flexibility.", "retry_text": "A transformer model is a type of neural network designed to process long sequences of text, like sentences or paragraphs, more efficiently than traditional models like RNNs or LSTMs. It uses self-attention mechanisms to focus on specific parts of the input at different times, allowing it to handle complex, long-term dependencies in text.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "91e15fc0eb0e2627", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "In plain English, what is a transformer model?", "text": "A transformer model is a type of machine learning model used in **natural language processing (NLP)** to handle complex text and sequences. Here's a simple explanation:\n\n- **Purpose**: It's designed to process long sequences of text (like paragraphs or sentences) and understand context, which is useful for tasks like language translation, text generation, or summarization.\n- **Key Mechanism**: Unlike traditional models like RNNs or LSTMs, transformers use **self-attention** to dynamically determine which parts of the input to consider relevantly. This allows them to focus on the context and not just memorize the data.\n- **Comparison**: Unlike RNNs or LSTMs, which process data sequentially, transformers can handle long sequences more efficiently.\n\nIn short, a transformer model helps machines understand and generate text more effectively by focusing on context and long-term dependencies.", "retry_text": "A transformer model is a type of machine learning model used in **natural language processing (NLP)** to handle complex text and sequences. Here's a simple explanation:\n\n- **Purpose**: It's designed to process long sequences of text (like paragraphs or sentences) and understand context, which is useful for tasks like language translation, text generation, or summarization. \n- **Key Mechanism**: Unlike traditional models like RNNs or LSTMs, transformers use **self-attention** to dynamically determine which parts of the input to consider relevantly. This allows them to focus on context and not just memorize the data. \n- **Comparison**: Unlike RNNs or LSTMs, which process data sequentially, transformers can handle long sequences more efficiently. \n\nIn short, a transformer model helps machines understand and generate text more effectively by focusing on context and long-term dependencies.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "5c344eaff3a31f90", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "In plain English, what is a transformer model?", "text": "A transformer model is a type of deep learning model used to process and generate text or other sequential data. Unlike traditional neural networks that use recurrent units (like RNNs), transformers use self-attention mechanisms to focus on relevant parts of the input. This makes them highly efficient for handling long, complex sequences of text, such as in machine translation or text generation tasks.", "retry_text": "A transformer model is a type of deep learning model used to process and generate text or other sequential data. Unlike traditional neural networks that use recurrent units (like RNNs), transformers use self-attention mechanisms to focus on relevant parts of the input. This makes them highly efficient for handling long, complex sequences of text, such as in machine translation or text generation.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "70de4d7453ea6e5f", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "In plain English, what is a transformer model?", "text": "A transformer model is a type of neural network used to process and understand text, like how humans do. It works by breaking down long sequences of text into smaller parts. Transformers use self-attention to focus on different parts of the text, allowing them to process information efficiently. This makes them great for tasks like language translation, text generation, and speech recognition.", "retry_text": "A transformer model is a type of neural network used in natural language processing to process and understand text, like how humans do. It works by breaking down long sequences of text into smaller parts, using self-attention to focus on different parts of the text, and allowing efficient processing of information. This makes it great for tasks like language translation, text generation, and speech recognition.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "fec9676d5b701fb7", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "In plain English, what is a transformer model?", "text": "A transformer model is a type of neural network used in **natural language processing** to help with tasks like language modeling, text generation, and machine translation. Here's a simple breakdown in plain English:\n\n- **Purpose**: Processes long sequences of text (like sentences or paragraphs) to understand and generate human-like content.\n- **Key Components**: \n - **Self-Attention**: Allows the model to weigh the importance of different parts of the input sequence. \n - **Encoder-Decoder Architecture**: Breaks the problem into encoding the input text and decoding it into output text. \n- **Example**: Uses a transformer to translate text from one language to another, or generate text that aligns with a given prompt.\n\nIn short, a transformer model is a powerful tool for handling complex patterns in sequential data.", "retry_text": "A transformer model is a type of neural network used in natural language processing (NLP) to help with tasks like language modeling, text generation, and machine translation. In plain English:\n\n- It processes long sequences of text (e.g., sentences or paragraphs) to understand and generate human-like content. \n- It uses **self-attention** to weigh the importance of different parts of the input, allowing it to learn patterns and dependencies between words. \n- The model is structured as an **encoder-decoder** system, breaking the problem into encoding the input and decoding it into output. \n- Examples include translating text between languages or generating creative text. \n\nIn summary, a transformer model helps machines understand complex patterns in text by focusing on relationships between words rather than just individual elements.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "2d91756e22f14192", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "In plain English, what is a transformer model?", "text": "A transformer model is a type of neural network designed to process and understand long sequences of text, such as in natural language processing tasks. Unlike traditional RNNs or LSTMs, which process information sequentially, transformers use **self-attention** to dynamically highlight and focus on specific parts of the input, allowing them to handle complex, long-term dependencies in text more effectively. This makes them powerful for tasks like machine translation and text generation.", "retry_text": "A transformer model is a type of neural network used in natural language processing to handle long sequences of text. Unlike traditional RNNs or LSTMs (which process information sequentially), transformers use **self-attention** to dynamically highlight and focus on specific parts of the input, enabling them to process complex, long-term dependencies more effectively. This makes them powerful for tasks like machine translation and text generation.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "62f5e4af91d404bc", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "In plain English, what is a transformer model?", "text": "A transformer model is a type of neural network used in machine learning to process and transform input data efficiently. Unlike traditional models like RNNs or LSTMs, which process information sequentially, transformer models use **self-attention** to focus on specific parts of the input at different times. This allows them to learn complex patterns and dependencies in the data more effectively. The core idea is that the model can dynamically \"look\" at multiple parts of the input, enabling it to capture context and relationships in a way that traditional models might not.", "retry_text": "A transformer model is a type of neural network used in machine learning to process and transform input data efficiently. Unlike traditional models like RNNs or LSTMs, which process information sequentially, transformer models use **self-attention** to focus on specific parts of the input at different times. This allows them to learn complex patterns and dependencies in the data more effectively.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "29d4bd9e22b119a2", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is 2 + 2?", "text": "2 + 2 equals 4.", "retry_text": "2 + 2 equals 4.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "fccda71d724525a4", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is 2 + 2?", "text": "2 + 2 equals 4.", "retry_text": "2 + 2 equals 4.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "4368361fe083e78b", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is 2 + 2?", "text": "The result of 2 + 2 is 4.", "retry_text": "2 + 2 equals 4.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "d1d0c89509f6e2b1", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is 2 + 2?", "text": "2 + 2 equals 4.", "retry_text": "2 + 2 equals 4.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "553666af55035ee6", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is 2 + 2?", "text": "2 + 2 equals 4.", "retry_text": "2 + 2 equals 4.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "09332e87b52a264c", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is 2 + 2?", "text": "2 + 2 equals 4.", "retry_text": "2 + 2 equals 4.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "02910af1afb0fdeb", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is 2 + 2?", "text": "2 + 2 equals 4.", "retry_text": "2 + 2 equals 4.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "2ae2f8f18e6fb620", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is 2 + 2?", "text": "2 + 2 equals 4.", "retry_text": "2 + 2 equals 4.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "5ebee57bef9e7cc3", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Give me a short definition of overfitting.", "text": "Overfitting occurs when a model learns too much from the training data, leading to poor generalization and high error rates.", "retry_text": "Overfitting occurs when a model learns too much from the training data, leading to poor generalization and high error rates.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "653926096e613ab4", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Give me a short definition of overfitting.", "text": "Overfitting occurs when a model's complexity increases beyond the actual data, causing it to perform poorly on new, unseen data.", "retry_text": "Overfitting occurs when a model's complexity increases beyond the actual data, causing it to perform poorly on new, unseen data.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "55755ab88045d4cf", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Give me a short definition of overfitting.", "text": "Overfitting occurs when a model's training data is too similar to the test data, causing the model to learn the patterns present in the training data instead of the underlying generalization.", "retry_text": "Overfitting occurs when a model learns too much from the training data, making it perform well on the training set but poorly on the test set.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "407cb9a2b45d875e", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Give me a short definition of overfitting.", "text": "Overfitting occurs when a model learns the training data too well, leading to poor performance on new, unseen data.", "retry_text": "Overfitting occurs when a model learns the training data too well, leading to poor performance on new, unseen data.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "52fadbaf0c564664", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Give me a short definition of overfitting.", "text": "Overfitting occurs when a model learns too much from the training data, leading to poor generalization to new, unseen data.", "retry_text": "Overfitting occurs when a model learns too much from the training data, leading to poor generalization to new, unseen data.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "b3ea25024fa7c2e5", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Give me a short definition of overfitting.", "text": "Overfitting occurs when a model learns too much from the training data, leading to poor generalization and poor performance on new, unseen data.", "retry_text": "Overfitting occurs when a model learns too much from the training data, leading to poor generalization and poor performance on new, unseen data.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "e253899c9a49119c", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Give me a short definition of overfitting.", "text": "Overfitting refers to a situation in machine learning where a model learns the training data too well, making it perform poorly on new, unseen data.", "retry_text": "Overfitting refers to a situation in machine learning where a model learns the training data too well, making it perform poorly on new, unseen data.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "cff4f8216561e80a", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Give me a short definition of overfitting.", "text": "Overfitting occurs when a model's parameters become too specific to the training data, leading it to perform well on the training set but poorly on new, unseen data.", "retry_text": "Overfitting occurs when a model's parameters become too specific to the training data, leading it to perform well on the training set but poorly on new, unseen data.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "5f3fda191881eae4", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Who wrote Pride and Prejudice?", "text": "The novel *Pride and Prejudice* was written by **Jane Austen**. It was published in 1813.", "retry_text": "Jane Austen.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "53710dec1aba5d15", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is the difference between a list and a tuple in Python?", "text": "A list and a tuple in Python are both used to store collections of elements, but they have key differences:\n\n### **Difference Between List and Tuple:**\n1. **Mutability:**\n - **List:** Mutable, allows elements to be added, removed, or modified in-place.\n - **Tuple:** Immutable, cannot be changed once created.\n\n2. **Fixed Length (Tuples):**\n - **Tuple:** Has a fixed length, and all elements must be the same type.\n - **List:** Can have variable lengths and elements of different types.\n\n3. **Uniqueness:**\n - **Tuple:** Can be initialized with all elements, but lists can have any elements.\n - **List:** Can be initialized with any elements, not just a single element.\n\n### **Example:**\n```python\n# List\nmy_list = [1, 2, 3]\nmy_list.append(4) # List is mutable\nprint(my_list) # Output: [1, 2, 3, 4]\n\n# Tuple\nmy_tuple = (1, 2, 3)\n# Cannot modify my_tuple\nprint(my_tuple) # Output: (1, 2, 3)\n```\n\nThese differences make lists suitable for dynamic data and tuples for static, unchangeable data.", "retry_text": "A list and a tuple in Python differ in several key aspects:\n\n1. **Mutability:** \n - **List:** Mutable, allows elements to be added, removed, or modified in-place. \n - **Tuple:** Immutable, cannot be changed once created.\n\n2. **Fixed Length:** \n - **Tuple:** Has a fixed length and all elements must be of the same type. \n - **List:** Can have variable lengths and elements of different types.\n\n3. **Uniqueness:** \n - **Tuple:** Can be initialized with all elements, but lists can have any elements. \n - **List:** Can be initialized with any elements, not just a single element.\n\n**Example:** \n```python\n# List\nmy_list = [1, 2, 3]\nmy_list.append(4) # List is mutable\nprint(my_list) # Output: [1, 2, 3, 4]\n\n# Tuple\nmy_tuple = (1, 2, 3)\n# Cannot modify my_tuple\nprint(my_tuple) # Output: (1, 2, 3)\n```", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "b681f92c716c4f2a", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is the difference between a list and a tuple in Python?", "text": "In Python, a **list** and a **tuple** are both used to store collections of elements, but they differ in several key aspects:\n\n### 1. **Immutable vs. Mutable** \n- **Tuple**: Once created, its elements **cannot be changed**. \n Example: `(1, 2, 3)` → changes to `2, 3, 4` would break the tuple. \n- **List**: Elements can be **modified** as long as you don't change them. \n Example: `[1, 2, 3]` → modifying `3` to `2` would update the list.\n\n### 2. **Element Type Consistency** \n- **Tuple**: All elements must be of the **same type** (e.g., integers, strings). \n- **List**: Elements can be of any type (e.g., integers, strings, etc).\n\n### 3. **Use Cases** \n- **Tuple**: Best when the data **remains constant** or **cannot be modified**. \n- **List**: Best when the data **can be updated or modified**.\n\n### Example\n```python\n# Tuple (immutable) \nt1 = (1, 2, 3) \nt2 = t1 + (4, 5) # t2 becomes (1, 2, 3, 4, 5)\n\n# List (mutable) \nl1 = [1, 2, 3] \nl2 = l1 + [4, 5] # l2 becomes (1, 2, 3, 4, 5)\n```\n\n### Summary\n- **Tuple**: Best for data that wont change, or for immutable data. \n- **List**: Best for dynamic, mutable data.", "retry_text": "In Python, a **list** and a **tuple** are both used to store collections of elements, but they differ in several key aspects:\n\n### 1. **Immutable vs. Mutable** \n- **Tuple**: Once created, its elements **cannot be changed**. \n Example: `(1, 2, 3)` → changes to `2, 3, 4` would break the tuple. \n- **List**: Elements can be **modified** as long as you don't change them. \n Example: `[1, 2, 3]` → modifying `3` to `2` would update the list.\n\n### 2. **Element Type Consistency** \n- **Tuple**: All elements must be of the **same type** (e.g., integers, strings). \n- **List**: Elements can be of any type (e.g., integers, strings, etc).\n\n### 3. **Use Cases** \n- **Tuple**: Best for immutable data or when the data **remains constant**. \n- **List**: Best for dynamic, mutable data where elements can be updated or modified.\n\n### Example\n```python\n# Tuple (immutable) \nt1 = (1, 2, 3) \nt2 = t1 + (4, 5) # t2 becomes (1, 2, 3, 4, 5)\n\n# List (mutable) \nl1 = [1, 2, 3] \nl2 = l1 + [4, 5] # l2 becomes (1, 2, 3, 4, 5)\n```\n\n### Summary \n- **Tuple**: Best for data that wont change or for immutable data. \n- **List**: Best for dynamic, mutable data.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "2c69370b7ac37217", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is the difference between a list and a tuple in Python?", "text": "In Python, the main difference between a list and a tuple is:\n\n1. **Mutability**: \n - A **list** is mutable, meaning you can add, remove, or modify elements. \n - A **tuple** is immutable, so once created, its elements cannot be changed.\n\n2. **Fixed Length**: \n - Tuples have a fixed length and cannot change their elements. \n - Lists can grow or shrink dynamically as elements are added or removed.\n\n3. **Access**: \n - Tuples are accessed using indexes, while lists use indices as well. \n\n4. **Use Cases**: \n - Lists are ideal for dynamic data where changes are frequent. \n - Tuples are used when you need to store data that should not change, such as coordinates or a sequence of numbers.", "retry_text": "In Python, the main difference between a list and a tuple is:\n\n1. **Mutability**: \n - A **list** is mutable (you can change its elements). \n - A **tuple** is immutable (elements cannot be changed once created).\n\n2. **Fixed Length**: \n - Tuples have a fixed length and cannot change their elements. \n - Lists can grow or shrink dynamically.\n\n3. **Access**: \n - Both lists and tuples use indices for access, though tuples are accessed using indexes directly.\n\n4. **Use Cases**: \n - Lists are ideal for dynamic data with frequent changes. \n - Tuples are used for data that should remain constant, such as coordinates or a sequence of numbers.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "a53e2fdb6fe27851", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is the difference between a list and a tuple in Python?", "text": "In Python, a **list** and a **tuple** are both iterable data types, but they have key differences:\n\n1. **Modification**: \n - **Lists** can be modified (e.g., `list.append()`, `list.remove()`, `list.insert()`) after creation. \n - **Tuples** cannot be modified once created.\n\n2. **Length**: \n - **Tuples** have a **fixed length** and cannot be expanded or shrunk. \n - **Lists** can dynamically grow or shrink.\n\n3. **Hashing**: \n - Tuples can be used for hashing (e.g., `hash(tuple)`) and are often used for unique identifiers. \n - Lists can also have a `hash()` method.\n\n4. **Use Cases**: \n - **Lists** are mutable and used for dynamic data. \n - **Tuples** are immutable and used for fixed, unchanging data.\n\nFor example: \n```python\nmy_list = [1, 2, 3] \nmy_list.append(4) # List changes \nprint(my_list) # Output: [1, 2, 3, 4] \n\nmy_tuple = (1, 2, 3) \nprint(my_tuple) # Output: (1, 2, 3) \n```", "retry_text": "In Python, a list and a tuple are both iterable data types but differ in their behaviors:\n\n1. **Modification**: \n - **Lists** can be modified (e.g., appended, removed, or inserted elements). \n - **Tuples** cannot be modified once created.\n\n2. **Length**: \n - **Tuples** have a **fixed length** and cannot be expanded or shrunk.\n\n3. **Hashing**: \n - Tuples can be used for hashing and are often used for unique identifiers. \n - Lists can also have a `hash()` method.\n\n4. **Use Cases**: \n - **Lists** are mutable and used for dynamic data. \n - **Tuples** are immutable and used for fixed, unchanging data.\n\nFor example: \n```python\nmy_list = [1, 2, 3] \nmy_list.append(4) # List changes \nprint(my_list) # Output: [1, 2, 3, 4] \n\nmy_tuple = (1, 2, 3) \nprint(my_tuple) # Output: (1, 2, 3)\n```", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "91fc9d1fb351d237", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is the difference between a list and a tuple in Python?", "text": "In Python, a **list** and a **tuple** are both used to store collections of data, but they differ in several key aspects:\n\n1. **Mutability**:\n - A **list** is **mutable**, meaning it can be modified (inserted, deleted, or updated elements).\n - A **tuple** is **immutable**, so its elements cannot be changed once created.\n\n2. **Order**:\n - Tuples preserve the order of elements, while lists also maintain the original order when elements are modified.\n\n3. **Use Cases**:\n - **Lists** are used when you want to change the order or add elements dynamically.\n - **Tuples** are used when you want to pass a collection of values as a group and not modify them.\n\nFor example:\n```python\nmy_list = [1, 2, 3]\nmy_list.append(4) # List is mutable\nprint(my_list) # Output: [1, 2, 3, 4]\n\nmy_tuple = (1, 2, 3)\nmy_tuple[0] = 5 # Tuple is immutable\nprint(my_tuple) # Output: (1, 2, 3)\n```\n\nThe main difference lies in their immutability and how they handle changes.", "retry_text": "A list and a tuple in Python differ primarily in their **mutability** and **behavior**:\n\n1. **Mutability**:\n - A **list** is **mutable** (can be modified, added, or removed elements).\n - A **tuple** is **immutable** (cannot be changed once created).\n\n2. **Order**:\n - Tuples preserve the original order of elements, while lists also maintain that order.\n\n3. **Use Cases**:\n - **Lists** are used when you dynamically modify or add elements.\n - **Tuples** are used when you want to pass a collection of values as a group and not modify them.\n\nTuples are immutable, making them ideal for static data structures, while lists allow dynamic modifications.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "616079ceb9f98220", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is the difference between a list and a tuple in Python?", "text": "In Python, a **list** and a **tuple** are both used to store collections of items, but they differ in several key aspects:\n\n1. **Uniqueness**: \n - A **tuple** is **uniquely defined** and cannot be modified once created. \n - A **list** is **dynamic** and allows adding or removing elements at any time. \n\n2. **Immutable**: \n - Tuples are **immutable** and cannot be changed after creation. \n - Lists are **mutable** and can be modified as needed. \n\n3. **Order**: \n - Tuples maintain a **fixed order** of elements. \n - Lists allow dynamic reordering (e.g., appending or removing elements). \n\n4. **Use Cases**: \n - **Tuples** are ideal for data that must remain unchanged, such as coordinates, constants, or fixed-order collections. \n - **Lists** are used for dynamic data, such as a list of names, a list of items to be added or removed in an application. \n\nFor example: \n- A tuple: `(\"apple\", \"banana\")` \n- A list: `[\"banana\", \"apple\", \"cherry\"]` \n\nThis distinction is crucial in Python when handling data structures and performance considerations.", "retry_text": "A **list** and a **tuple** are both used to store collections of items in Python, but they differ in several key aspects: \n\n1. **Uniqueness**: \n - A **tuple** is **uniquely defined** and cannot be modified once created. \n - A **list** is **dynamic** and allows adding or removing elements at any time. \n\n2. **Immutable**: \n - Tuples are **immutable** and cannot be changed after creation. \n - Lists are **mutable** and can be modified as needed. \n\n3. **Order**: \n - Tuples maintain a **fixed order** of elements. \n - Lists allow dynamic reordering (e.g., appending or removing elements). \n\n4. **Use Cases**: \n - **Tuples** are ideal for data that must remain unchanged, such as coordinates, constants, or fixed-order collections. \n - **Lists** are used for dynamic data, such as a list of names, a list of items to be added or removed in an application. \n\nFor example: \n- A tuple: `(\"apple\", \"banana\")` \n- A list: `[\"banana\", \"apple\", \"cherry\"]` \n\nThis distinction is crucial in Python when handling data structures and performance considerations.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "2078c91883c4a353", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is the difference between a list and a tuple in Python?", "text": "In Python, a **list** and a **tuple** are both iterable, but they differ in several key aspects:\n\n### 1. **Immutable vs. Mutable**\n- A **list** is **mutable**, meaning it can be changed or modified.\n- A **tuple** is **immutable**, so changes to its elements require creating a new tuple.\n\n### 2. **Dynamic vs. Static**\n- A **list** is dynamic—its elements can be added or removed, and its length can grow.\n- A **tuple** is static—its elements cannot be changed once created.\n\n### 3. **Use Cases**\n- **Lists** are ideal for data that can be dynamically added or modified (e.g., a list of coordinates, a list of names).\n- **Tuples** are used for data that cannot be modified (e.g., coordinates, a fixed set of values).\n\n### Example\n```python\n# List\nmy_list = [1, 2, 3]\nmy_list[0] = 4 # Changes the first element\nmy_list.append(4) # Adds an element\n\n# Tuple\nmy_tuple = (1, 2, 3)\nmy_tuple[0] = 5 # Cannot change, so a new tuple is created\n```\n\nThis distinction is fundamental in Python for handling different types of data efficiently.", "retry_text": "In Python, a **list** and a **tuple** are both iterable, but they differ in several key aspects:\n\n1. **Immutable vs. Mutable**:\n - A **list** is **mutable** (can be changed or modified).\n - A **tuple** is **immutable** (cannot be changed once created).\n\n2. **Dynamic vs. Static**:\n - A **list** is **dynamic** (elements can be added, removed, or modified).\n - A **tuple** is **static** (elements cannot be changed).\n\n3. **Use Cases**:\n - **Lists** are ideal for data that can be dynamically added or modified (e.g., a list of coordinates, a list of names).\n - **Tuples** are used for data that cannot be modified (e.g., coordinates, a fixed set of values).\n\nExample:\n```python\n# List\nmy_list = [1, 2, 3]\nmy_list[0] = 4 # Changes the first element\nmy_list.append(4) # Adds an element\n\n# Tuple\nmy_tuple = (1, 2, 3)\nmy_tuple[0] = 5 # Cannot change, so a new tuple is created\n```\n\nThis distinction is fundamental in Python for handling different types of data efficiently.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "fd3dec4b84816b0c", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What is the difference between a list and a tuple in Python?", "text": "In Python, a **list** and a **tuple** are both used to store collections of items, but they differ in several key aspects:\n\n### 1. **Mutability**:\n - **List**: Mutable. You can change its elements after creation (e.g., `list.append()` or `list.remove()`) to modify the contents.\n - **Tuple**: Immutable. Once created, its elements cannot be changed, and any attempt to modify a tuple element will throw an error.\n\n### 2. **Structure**:\n - **List**: Elements are added in **order**, and elements can be inserted or removed at any position.\n - **Tuple**: Elements are added in **order**, and elements cannot be inserted or removed, making it static.\n\n### 3. **Comparison**:\n - When comparing two lists or tuples, they are compared element-wise. If a list has elements with different lengths, the longer list is considered \"greater\" in a tuple-wise comparison.\n - When comparing a list and a tuple, they are compared as a whole, and if the lists or tuples have different lengths, the one with more elements is considered greater.\n\n### 4. **Use Cases**:\n - **List**: For dynamic data that can be modified (e.g., a list of people, a list of prices).\n - **Tuple**: For static data that cannot be changed (e.g., a list of coordinates or a fixed set of values).\n\n### Example:\n```python\nmy_list = [1, 2, 3]\nmy_list.append(4)\nprint(my_list) # Output: [1, 2, 3, 4]\n\nmy_tuple = (1, 2, 3)\nprint(my_tuple) # Output: (1, 2, 3)\n```\n\nThis shows how the differences in mutability and structure affect their usage.", "retry_text": "In Python, a **list** and a **tuple** are both used to store collections of items, but they differ in several key aspects:\n\n### 1. **Mutability**:\n - **List**: Mutable. You can change its elements after creation (e.g., `list.append()` or `list.remove()`) to modify the contents.\n - **Tuple**: Immutable. Once created, its elements cannot be changed, and any attempt to modify a tuple element will throw an error.\n\n### 2. **Structure**:\n - **List**: Elements are added in **order**, and elements can be inserted or removed at any position.\n - **Tuple**: Elements are added in **order**, and elements cannot be inserted or removed, making it static.\n\n### 3. **Comparison**:\n - When comparing lists or tuples, they are compared element-wise. If a list has elements with different lengths, the longer list is considered \"greater\" in a tuple-wise comparison.\n - When comparing a list and a tuple, they are compared as a whole, and if the lists or tuples have different lengths, the one with more elements is considered greater.\n\n### 4. **Use Cases**:\n - **List**: For dynamic data that can be modified (e.g., a list of people, a list of prices).\n - **Tuple**: For static data that cannot be changed (e.g., a list of coordinates or a fixed set of values).\n\n### Example:\n```python\nmy_list = [1, 2, 3]\nmy_list.append(4)\nprint(my_list) # Output: [1, 2, 3, 4]\n\nmy_tuple = (1, 2, 3)\nprint(my_tuple) # Output: (1, 2, 3)\n```", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "41da97d37a1bfe04", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Summarise what LoRA does in one paragraph.", "text": "LoRA (Low-Rank Adaptation) is a technique used in training large language models to improve their performance during training by adding a small linear transformation to the model's weights. This method helps in optimizing the model's parameters more efficiently by leveraging the model's existing structure, allowing for faster training without significant computational overhead.", "retry_text": "LoRA (Low-Rank Adaptation) is a technique used in training large language models to enhance their performance by introducing a small, linear transformation to the model's weights. This method allows for more efficient training by leveraging the model's existing structure, reducing computational overhead and enabling faster parameter optimization.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "7d705f85c729f2d5", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Summarise what LoRA does in one paragraph.", "text": "LoRA stands for LoRALayer-wise, a technique used in training large language models (LLMs) to enhance performance and efficiency. It works by adding a small linear layer to the model's weights, enabling the model to learn the same representations without updating all the model's weights. This approach reduces computational overhead and speeds up training, particularly beneficial for very large models.", "retry_text": "LoRA stands for LoRALayer-wise, a technique used in training large language models (LLMs) to enhance performance and efficiency. It works by adding a small linear layer to the model's weights, enabling the model to learn the same representations without updating all the model's weights. This approach reduces computational overhead and speeds up training, particularly beneficial for very large models.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "e712132213910e8a", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Summarise what LoRA does in one paragraph.", "text": "LoRA (LoRALayer) is a technique used in training large language models to improve efficiency by adding a small set of additional parameters to the main model. These parameters are not updated during training, allowing the model to be trained more quickly and with less computational cost. This approach reduces the number of parameters in the model while enabling faster and more effective training.", "retry_text": "LoRA (LoRALayer) is a technique used in training large language models to enhance efficiency by adding a small set of additional parameters to the main model. These parameters are not updated during training, allowing the model to be trained more quickly and with less computational cost. This approach reduces the number of parameters while enabling faster and more effective training.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "ed971da889aa00c5", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Summarise what LoRA does in one paragraph.", "text": "LoRA (LoRA Optimization) is a technique used to enhance the efficiency and performance of training transformer models, such as BERT or GPT, by adding a small set of specific parameters to the base model. These parameters are designed to improve the model's ability to learn, allowing it to train faster and use less computational resources. This optimization technique enables models to achieve better accuracy without requiring a significant increase in model size or complexity.", "retry_text": "LoRA (LoRA Optimization) is a technique designed to enhance the efficiency and performance of transformer models like BERT or GPT by adding a small set of specific parameters to the base model. These parameters are tailored to improve the model's ability to learn, enabling faster training and reduced computational resource usage. This optimization allows models to achieve better accuracy without requiring a substantial increase in model size or complexity.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "cb283f6344409751", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Summarise what LoRA does in one paragraph.", "text": "LoRA stands for Low-Rank Adaptation, a technique used to improve the performance of neural networks by adding a low-rank matrix to the model's weight matrix. This allows the model to adapt more efficiently and with less computational cost, making it suitable for tasks like language modeling where large-scale parameters are common.", "retry_text": "LoRA stands for Low-Rank Adaptation, a technique used to improve the performance of neural networks by adding a low-rank matrix to the model's weight matrix. This allows the model to adapt more efficiently and with less computational cost, making it suitable for tasks like language modeling where large-scale parameters are common.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "062b1b577a41b87a", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Summarise what LoRA does in one paragraph.", "text": "LoRA (LoRA Optimization) is a technique used in deep learning models to improve performance by adding additional layers to the attention mechanism in large language models (like BERT or GPT). This optimization allows the model to be more efficient in training, as it reduces the number of parameters while maintaining high accuracy, which is crucial for handling large-scale tasks.", "retry_text": "LoRA (LoRA Optimization) is a technique used in deep learning models to improve performance by adding additional layers to the attention mechanism in large language models (like BERT or GPT). This optimization allows the model to be more efficient in training, as it reduces the number of parameters while maintaining high accuracy, which is crucial for handling large-scale tasks.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "7b07271fa7175e04", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Summarise what LoRA does in one paragraph.", "text": "LoRA stands for LoRALayer and is a technique used in transformer models to enhance their performance. It allows for efficient fine-tuning of large models by modifying the weights in a way that improves the learning of the model without requiring full retraining. This method enables the model to learn better representations without significant computational overhead.", "retry_text": "LoRA stands for LoRALayer, a technique used in transformer models to enhance performance by allowing efficient fine-tuning of large models. It modifies the weights in a way that improves the model's learning without requiring full retraining, enabling the model to learn better representations without significant computational overhead.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "126f4d579ccec344", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Summarise what LoRA does in one paragraph.", "text": "LoRA (LoRALayerwise) is a technique that modifies the weights of a model's layers in a specific position, allowing for enhanced performance during training by fine-tuning the model's parameters. This approach enables efficient training by reducing computational overhead without sacrificing model accuracy, particularly in scenarios where the model's parameters need to be optimized for specific tasks.", "retry_text": "LoRA (LoRALayerwise) is a technique that modifies the weights of a model's layers in a specific position, enabling efficient training by fine-tuning the model's parameters. This approach reduces computational overhead while preserving model accuracy, particularly in scenarios where parameter optimization is needed.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "1d74acd146fcd5c3", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Is Python interpreted or compiled?", "text": "Python is interpreted and compiled. It works in a way where the interpreter runs the code directly, while the compiler (used in the Python standard library) is optimized for performance. The interpreter handles the runtime execution, and the compiler (in C) is used for code optimization and performance.", "retry_text": "Python is interpreted and compiled. It works in a way where the interpreter runs the code directly, while the compiler (used in the Python standard library) is optimized for performance. The interpreter handles the runtime execution, and the compiler (in C) is used for code optimization and performance.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "a88f0061096b96f8", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Is Python interpreted or compiled?", "text": "Python is interpreted. The interpreter reads the code and executes it, which is a common way Python functions work.", "retry_text": "Python is interpreted. The interpreter reads and executes the code as it runs, which is a common way Python functions work.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "4eccea6235fa63b0", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Is Python interpreted or compiled?", "text": "Python is interpreted and can be compiled. When you run Python code, it uses a runtime interpreter (interpreted mode) to execute the code. For performance, Python can also be compiled into a C extension module, which is interpreted at runtime (compiled mode).", "retry_text": "Python is interpreted and can be compiled. When you run Python code, it uses a runtime interpreter (interpreted mode) to execute the code. For performance, Python can also be compiled into a C extension module, which is interpreted at runtime (compiled mode).", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "d76dead29a1db13f", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Is Python interpreted or compiled?", "text": "Python is interpreted by default. It reads and executes code directly from files, which is the most common way it's used. If you're using the Python interpreter (like in an environment), it runs the code as a script. If you're using a compiled version, like a bytecode file stored in an `.exe` or `.pyc` file, it's executed as a compiled program. Python is not a compiler—it's a language that can be interpreted or compiled at runtime.", "retry_text": "Python is interpreted by default. It reads and executes code directly from files, which is the most common way it's used. If you're using the Python interpreter (like in an environment), it runs the code as a script. If you're using a compiled version, like a bytecode file stored in an `.exe` or `.pyc` file, it's executed as a compiled program. Python is not a compiler—it's a language that can be interpreted or compiled at runtime.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "e455113745e55075", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Is Python interpreted or compiled?", "text": "Python is interpreted and compiled. Here's a breakdown:\n\n- **Interpreted**: Python code is executed directly in the interpreter, meaning it runs as a program without needing a compiler. This is the standard method used.\n- **Compiled**: Python can also be compiled into bytecode (like in CPython), which is then executed by the interpreter. This approach is more efficient for performance-critical applications.\n\nPython is a **general-purpose language** that allows developers to write code that can be interpreted or compiled, depending on the environment.", "retry_text": "Python is interpreted and compiled. Here's a breakdown:\n\n- **Interpreted**: Python code is executed directly in the interpreter, meaning it runs as a program without requiring a compiler (this is the standard method).\n- **Compiled**: Python can also be compiled into bytecode (like in CPython), which is then executed by the interpreter. This approach is more efficient for performance-critical applications.\n\nPython is a **general-purpose language** that allows developers to write code that can be interpreted or compiled, depending on the environment.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "614649027bb66232", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Is Python interpreted or compiled?", "text": "Python is interpreted or compiled depending on the context. When you run Python code, it is interpreted, meaning the code is executed directly without compilation. However, when you use compiled code (like using a library module), it's compiled into a module that can be executed quickly. Python uses a combination of interpreted and compiled code to achieve this.", "retry_text": "Python is interpreted or compiled depending on the context. When you run Python code, it is interpreted, meaning the code is executed directly without compilation. However, when you use compiled code (like using a library module), it is compiled into a module that can be executed quickly. Python uses a combination of interpreted and compiled code to achieve this.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "6814c671025ade74", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Is Python interpreted or compiled?", "text": "Python is interpreted, not compiled. The interpreter processes and executes the code directly, translating it into machine code for execution. This contrasts with compiled languages like C++ or Java, where the code is compiled into a single file (a \"linked list\") that can be run without recompilation.", "retry_text": "Python is interpreted, not compiled. The interpreter processes and executes the code directly, translating it into machine code for execution.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "bc5650447d001e2d", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "Is Python interpreted or compiled?", "text": "Python is interpreted, not compiled. The interpreter runs the code directly, while the compiler compiles the code into machine code for execution.", "retry_text": "Python is interpreted, not compiled. The interpreter runs the code directly, while the compiler compiles the code into machine code for execution.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "6e6494aead8f1636", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What colour do you get mixing blue and yellow?", "text": "When you mix blue and yellow, you get green. This is because the primary colors in the traditional color wheel are red, blue, and yellow. Mixing two primary colors results in the third color in the spectrum.", "retry_text": "When you mix blue and yellow, you get green. This is because the primary colors in the traditional color wheel are red, blue, and yellow. Mixing two primary colors results in the third color in the spectrum.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "5a39099b612f2c75", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What colour do you get mixing blue and yellow?", "text": "When you mix blue and yellow, the resulting color is a complementary color pair. Blue and yellow are on the color wheel, and combining them produces a color that is not directly visible but complementary. The exact shade can vary depending on the mixing method (e.g., using a palette knife vs. a flat surface), but the result is typically a shade of blue or yellow in the opposite direction.", "retry_text": "When mixing blue and yellow, the resulting color is a complementary color pair. Blue and yellow are on the color wheel, and combining them produces a color that is not directly visible but complementary. The exact shade can vary depending on the mixing method, but the result is typically a shade of blue or yellow in the opposite direction.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "32dc34da4d27b3c4", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What colour do you get mixing blue and yellow?", "text": "The mixing of blue and yellow results in yellow. This is a basic color theory fact.", "retry_text": "The mixing of blue and yellow results in yellow. This is a basic color theory fact.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "67a73b74efc2bbd5", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What colour do you get mixing blue and yellow?", "text": "The color you get when mixing blue and yellow is yellow. This is because blue and yellow are complementary colors in the visible spectrum, and their combination results in the color yellow.", "retry_text": "The color you get when mixing blue and yellow is yellow.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "5643f5575982ee67", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What colour do you get mixing blue and yellow?", "text": "The correct answer is that mixing blue and yellow results in **orange**. This is based on the traditional color mixing principles of additive color theory.", "retry_text": "The correct answer is that mixing blue and yellow results in **orange**. This is based on the traditional color mixing principles of additive color theory.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "04784cbe5123000d", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What colour do you get mixing blue and yellow?", "text": "When mixing blue and yellow, the result is white. This is based on the principles of color theory, where the combination of primary colors (blue, yellow, and red) results in white. The exact color outcome depends on the intensity of the colors used, but typically, mixing equal parts blue and yellow yields the characteristic white.", "retry_text": "The color obtained when mixing blue and yellow is white. This is based on the principles of color theory, where primary colors (blue, yellow, and red) combine to produce the complementary color white.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "a25aab4f1135fbba", "model": "qwen3-0.6b", "prompt_class": "no_tool", "prompt": "What colour do you get mixing blue and yellow?", "text": "When you mix blue and yellow, the resulting color is **yellow**. This is because the primary colors (red, blue, and yellow) combine to produce secondary colors when mixed. Mixing two primary colors results in the secondary color.", "retry_text": "The color when mixing blue and yellow is **yellow**. This is because the primary colors (red, blue, and yellow) combine to produce secondary colors when mixed.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "5bdabb7899cfd8f8", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Give me a numbered list of steps to fine tune a model.", "text": "Heres a numbered list of steps to fine tune a model:\n\n1. **Data Preparation**: Collect and preprocess the dataset to ensure its suitable for the model. This includes handling missing data, scaling features, and cleaning the data.\n\n2. **Model Selection**: Choose an appropriate model (e.g., transformer, CNN, or RNN) that is suitable for the task and aligns with the data characteristics.\n\n3. **Hyperparameter Tuning**: Optimize the models hyperparameters (e.g., learning rate, batch size, regularization) to improve performance and convergence.\n\n4. **Initial Training**: Perform a full model training on the dataset to fine-tune the model to the task.\n\n5. **Fine-Tuning**: Adjust the model parameters based on the performance of the initial training phase to improve accuracy.\n\n6. **Evaluation**: Assess the models performance using appropriate metrics and compare it with the original model.\n\n7. **Deployment**: Once the model is optimized, deploy it for prediction or inference tasks.\n\nLet me know if you need further clarification!", "retry_text": "Heres a numbered list of steps to fine tune a model:\n\n1. **Data Preparation** \n2. **Model Selection** \n3. **Hyperparameter Tuning** \n4. **Initial Training** \n5. **Fine-Tuning** \n6. **Evaluation** \n7. **Deployment** \n\nLet me know if you need further clarification!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "20f4b59d62954aa6", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Give me a numbered list of steps to fine tune a model.", "text": "Heres a numbered list of steps to fine-tune a model:\n\n1. **Pre-training on a large dataset** \n Train the model on a large, diverse dataset to obtain a strong initial representation of the task.\n\n2. **Fine-tuning on the target dataset** \n Adjust the model on a smaller, specific dataset to fine-tune its parameters for the task at hand.\n\n3. **Validation during training** \n Monitor the models performance using a validation set or metric to ensure it generalizes well and doesnt overfit.\n\n4. **Hyperparameter tuning** \n Optimize the models hyperparameters (e.g., learning rate, batch size, regularization) to improve efficiency and performance.\n\n5. **Using techniques like data augmentation** \n Augment the training data during fine-tuning to increase the models robustness and reduce overfitting.", "retry_text": "Heres a numbered list of steps to fine tune a model:\n\n1. **Pre-training on a large dataset** \n Train the model on a large, diverse dataset to obtain a strong initial representation of the task. \n\n2. **Fine-tuning on the target dataset** \n Adjust the model on a smaller, specific dataset to fine-tune its parameters for the task at hand. \n\n3. **Validation during training** \n Monitor the models performance using a validation set or metric to ensure it generalizes well and doesnt overfit. \n\n4. **Hyperparameter tuning** \n Optimize the models hyperparameters (e.g., learning rate, batch size, regularization) to improve efficiency and performance. \n\n5. **Using techniques like data augmentation** \n Augment the training data during fine-tuning to increase the models robustness and reduce overfitting.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "51864a903d9fb2bd", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Give me a numbered list of steps to fine tune a model.", "text": "Here is a numbered list of steps to fine-tune a model:\n\n1. **Initial Model Selection**: Choose the appropriate model architecture that is suitable for the task and has a good performance baseline. \n2. **Data Collection**: Gather and preprocess the training data to ensure it is suitable for the model. \n3. **Model Initialization**: Start by initializing the model with the provided parameters. \n4. **Training Phase**: Train the model on the dataset to improve its performance. Adjust the number of parameters if needed. \n5. **Hyperparameter Tuning**: Conduct experiments to optimize hyperparameters (e.g., learning rate, batch size) to enhance model efficiency. \n6. **Validation Phase**: Perform validation to check for overfitting and ensure the model generalizes well. \n7. **Final Fine-tuning**: Fine-tune the model using additional data or in a new dataset to improve its performance on the target task. \n\nLet me know if you need further clarification!", "retry_text": "Here is a numbered list of steps to fine tune a model:\n\n1. **Initial Model Selection**: Choose the appropriate model architecture that is suitable for the task and has a good performance baseline. \n2. **Data Collection**: Gather and preprocess the training data to ensure it is suitable for the model. \n3. **Model Initialization**: Start by initializing the model with the provided parameters. \n4. **Training Phase**: Train the model on the dataset to improve its performance. Adjust the number of parameters if needed. \n5. **Hyperparameter Tuning**: Conduct experiments to optimize hyperparameters (e.g., learning rate, batch size) to enhance model efficiency. \n6. **Validation Phase**: Perform validation to check for overfitting and ensure the model generalizes well. \n7. **Final Fine-tuning**: Fine-tune the model using additional data or in a new dataset to improve its performance on the target task.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "a452cbc60878427b", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Give me a numbered list of steps to fine tune a model.", "text": "Here is a numbered list of steps to fine tune a model:\n\n1. **Data Preparation**: \n - Clean and prepare the dataset for fine-tuning. \n - Handle missing values and imbalance in the data. \n\n2. **Model Selection**: \n - Choose a suitable model architecture (e.g., Transformer, CNN, or RNN) based on the task. \n\n3. **Hyperparameters Tuning**: \n - Optimize model hyperparameters (e.g., learning rate, batch size, number of epochs) to balance training efficiency and model performance. \n\n4. **Validation and Testing**: \n - Validate the model using a separate validation set to ensure it generalizes well. \n\n5. **Fine-tuning Execution**: \n - Adjust the model parameters (e.g., weights, biases) to improve performance on the target dataset. \n\n6. **Deployment**: \n - After fine-tuning, deploy the optimized model for real-world use.", "retry_text": "Here is a numbered list of steps to fine tune a model:\n\n1. **Data Preparation** \n - Clean and prepare the dataset for fine-tuning. \n - Handle missing values and imbalance in the data. \n\n2. **Model Selection** \n - Choose a suitable model architecture (e.g., Transformer, CNN, or RNN) based on the task. \n\n3. **Hyperparameters Tuning** \n - Optimize model hyperparameters (e.g., learning rate, batch size, number of epochs) to balance training efficiency and model performance. \n\n4. **Validation and Testing** \n - Validate the model using a separate validation set to ensure it generalizes well. \n\n5. **Fine-Tuning Execution** \n - Adjust the model parameters (e.g., weights, biases) to improve performance on the target dataset. \n\n6. **Deployment** \n - After fine-tuning, deploy the optimized model for real-world use.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "6d58c791b1b4028a", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Give me a numbered list of steps to fine tune a model.", "text": "1. Load the pre-trained model. \n2. Verify the models readiness for fine-tuning. \n3. Prepare and clean the training data. \n4. Perform the fine-tuning phase using the original dataset. \n5. Evaluate the models performance after fine-tuning.", "retry_text": "1. Load the pre-trained model. \n2. Verify the models readiness for fine-tuning. \n3. Prepare and clean the training data. \n4. Perform the fine-tuning phase using the original dataset. \n5. Evaluate the models performance after fine-tuning.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "3f4913c6bbd0921f", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Give me a numbered list of steps to fine tune a model.", "text": "Here's a numbered list of steps to fine-tune a model:\n\n1. **Pre-Train the Model**: Select a pre-trained model architecture (e.g., BERT, GPT) and fine-tune it on a new dataset to adapt it to the specific task.\n\n2. **Data Preparation**: Gather and preprocess the dataset for the task, ensuring it matches the model's input requirements.\n\n3. **Model Selection**: Choose a suitable model architecture that can effectively adapt to the task and is compatible with the training environment.\n\n4. **Training the Model**: Train the pre-trained model on the new data to maximize its performance and parameter optimization.\n\n5. **Fine-Tuning**: Adjust the model's parameters based on the new data to improve its accuracy and efficiency.\n\n6. **Evaluation**: After fine-tuning, evaluate the model's performance using appropriate metrics to ensure it meets the desired results.", "retry_text": "Here's a numbered list of steps to fine tune a model:\n\n1. **Pre-Train the Model**: Select a pre-trained model architecture (e.g., BERT, GPT) and fine-tune it on a new dataset to adapt it to the specific task.\n\n2. **Data Preparation**: Gather and preprocess the dataset for the task, ensuring it matches the model's input requirements.\n\n3. **Model Selection**: Choose a suitable model architecture that can effectively adapt to the task and is compatible with the training environment.\n\n4. **Training the Model**: Train the pre-trained model on the new data to maximize its performance and parameter optimization.\n\n5. **Fine-Tuning**: Adjust the model's parameters based on the new data to improve its accuracy and efficiency.\n\n6. **Evaluation**: After fine-tuning, evaluate the model's performance using appropriate metrics to ensure it meets the desired results.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "fc1e25c5d4a98acd", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Outline, as a bulleted list, how to debug a memory leak.", "text": "- **Identify the source of the leak** by tracing where the memory is being allocated or released. \n- **Check memory usage** using tools like Valgrind or heap analysis tools to detect unallocated memory. \n- **Monitor memory consumption** with tools like `gmon -p` or `gmon -p -m` to track heap usage. \n- **Analyze heap dumps** to identify objects that are not being freed. \n- **Use the garbage collector** (GC) to determine if it's catching the leak. \n- **Test with different environments** to validate the cause. \n- **Use Valgrind** to pinpoint the exact location of the leak.", "retry_text": "- **Identify the source of the leak** by tracing where memory is allocated or released. \n- **Check memory usage** using tools like Valgrind or heap analysis to detect unallocated memory. \n- **Monitor memory consumption** with tools like `gmon -p` or `gmon -p -m` to track heap usage. \n- **Analyze heap dumps** to identify objects that are not being freed. \n- **Use the garbage collector** (GC) to determine if it's catching the leak. \n- **Test with different environments** to validate the cause. \n- **Use Valgrind** to pinpoint the exact location of the leak.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "66a9eeebd5692881", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Outline, as a bulleted list, how to debug a memory leak.", "text": "- **Check for memory leaks in code** (e.g., using tools like Valgrind or memory dumps). \n- **Track allocated memory** (e.g., using `valgrind` or `gmonloc` to identify allocation/deallocation points). \n- **Locate memory allocation and deallocation points** in the code or runtime environment. \n- **Analyze heap and stack usage** to identify where the leak occurs. \n- **Use Valgrind or other memory analysis tools** to report and resolve leaks.", "retry_text": "- **Check for memory leaks using Valgrind** (e.g., by running `valgrind --leak-test-cmd` or analyzing memory dumps). \n- **Identify allocation/deallocation points** via tools like `gmonloc` or Valgrind's memory analysis tools. \n- **Track allocated memory** using `valgrind` or heap tracing to pinpoint where leaks occur. \n- **Analyze heap and stack usage** to determine the source of the memory issue. \n- **Use Valgrind to report and resolve leaks**, ensuring proper cleanup and memory management.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "afac0fda9391e9b9", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Outline, as a bulleted list, how to debug a memory leak.", "text": "- **Identify the source of the memory leak** by inspecting the program's code for memory allocations and deallocations. \n- **Use tools like Valgrind or LeakCheck** to detect leaks in the application's memory usage. \n- **Track memory usage** by writing code that monitors allocations and deallocations. \n- **Analyze memory allocation and deallocation points** to pinpoint where the leak occurs. \n- **Check for unused variables or improper memory management** in the code. \n- **Validate the solution** after fixing the leak to ensure no new leaks are introduced.", "retry_text": "- **Identify the source of the memory leak** by inspecting the program's code for memory allocations and deallocations. \n- **Use tools like Valgrind or LeakCheck** to detect leaks in the application's memory usage. \n- **Track memory usage** by writing code that monitors allocations and deallocations. \n- **Analyze memory allocation and deallocation points** to pinpoint where the leak occurs. \n- **Check for unused variables or improper memory management** in the code. \n- **Validate the solution** after fixing the leak to ensure no new leaks are introduced.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "1a1a92f274563204", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Outline, as a bulleted list, how to debug a memory leak.", "text": "- **Identify the source of the leak**: Use tools like `valgrind` to detect memory leaks by analyzing memory usage. \n- **Check allocated objects**: Look for memory allocations and ensure they are freed properly. \n- **Use `valgrind` to detect leaks**: Run commands like `valgrind --leak-check=static` or `valgrind --leak-check=full` to find memory leaks. \n- **Analyze memory usage**: Tools like `gmem` or `gprof` can help trace memory allocation and deallocation. \n- **Use a debugger**: Set breakpoints in the code to inspect memory usage and identify leaks. \n- **Implement memory management strategies**: Use smart pointers or proper deallocation in C/C++ to prevent leaks. \n- **Review and verify code for leaks**: Ensure all memory allocations and frees are handled correctly.", "retry_text": "- **Identify the source of the leak**: Use tools like `valgrind` to detect memory leaks by analyzing memory usage. \n- **Check allocated objects**: Ensure memory allocations and deallocations are handled properly. \n- **Use `valgrind` to detect leaks**: Run commands like `valgrind --leak-check=static` or `valgrind --leak-check=full` to find memory leaks. \n- **Analyze memory usage**: Tools like `gmem` or `gprof` can help trace memory allocation and deallocation. \n- **Use a debugger**: Set breakpoints in the code to inspect memory usage and identify leaks. \n- **Implement memory management strategies**: Use smart pointers or proper deallocation in C/C++ to prevent leaks. \n- **Review and verify code for leaks**: Ensure all memory allocations and frees are handled correctly.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "d052ea50b3e08f1c", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Outline, as a bulleted list, how to debug a memory leak.", "text": "- **Identify the leak** by checking memory allocations and observing where the leak starts and ends. \n- **Analyze the heap dump** to pinpoint the source of the leak. \n- **Use tools like Valgrind or LeakCheck** to trace memory allocations and identify leaks. \n- **Check if memory is being freed incorrectly** (e.g., not marking it as free). \n- **Modify the code** to prevent memory leaks by ensuring proper allocation and freeing of memory. \n- **Test the fix** by running the application and verifying that the leak is resolved.", "retry_text": "- **Identify the leak** by checking memory allocations and observing where the leak starts and ends. \n- **Analyze the heap dump** to pinpoint the source of the leak. \n- **Use tools like Valgrind or LeakCheck** to trace memory allocations and identify leaks. \n- **Check if memory is being freed incorrectly** (e.g., not marking it as free). \n- **Modify the code** to prevent memory leaks by ensuring proper allocation and freeing of memory. \n- **Test the fix** by running the application and verifying that the leak is resolved.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "ad9fb6b1ed4c0158", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Outline, as a bulleted list, how to debug a memory leak.", "text": "- **Identify the leak**: Use tools like Valgrind, gdb, or memory analysis tools to locate where memory is allocated and not freed. \n- **Check allocation/deallocation points**: Trace program execution to ensure memory is properly allocated and released. \n- **Verify memory usage**: Confirm that all memory allocations and frees are correctly managed. \n- **Ensure proper memory management**: Avoid dangling pointers and ensure all memory is released before the program exits.", "retry_text": "- **Identify the leak**: Use tools like Valgrind, gdb, or memory analysis tools to locate where memory is allocated and not freed. \n- **Check allocation/deallocation points**: Trace program execution to ensure memory is properly allocated and released. \n- **Verify memory usage**: Confirm that all memory allocations and frees are correctly managed. \n- **Ensure proper memory management**: Avoid dangling pointers and ensure all memory is released before the program exits.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "608c5d437fef5603", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Outline, as a bulleted list, how to debug a memory leak.", "text": "- **Identify the leak**: Use tools like Valgrind, Leakcheck, or your application's built-in memory analysis tools to detect memory leaks by checking heap usage or logs. \n- **Analyze the problem**: Trace the code to identify where memory was allocated and released, and determine why a memory block was not freed. \n- **Debug the code**: Modify the code to prevent memory leaks, such as updating allocation and release functions or ensuring proper memory management. \n- **Test the application**: Re-run the application after fixes to ensure the leak is resolved and no other issues persist.", "retry_text": "- **Identify the leak**: Use tools like Valgrind, Leakcheck, or your application's built-in memory analysis tools to detect memory leaks by checking heap usage or logs. \n- **Analyze the problem**: Trace the code to identify where memory was allocated and released, and determine why a memory block was not freed. \n- **Debug the code**: Modify the code to prevent memory leaks, such as updating allocation and release functions or ensuring proper memory management. \n- **Test the application**: Re-run the application after fixes to ensure the leak is resolved and no other issues persist.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "737a72c4c49ec6d6", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Outline, as a bulleted list, how to debug a memory leak.", "text": "- **Identify the Leak**: Use tools like `valgrind` or `gmem` to pinpoint where memory is being used incorrectly. \n- **Check Heap Usage**: Ensure all memory allocations and frees are properly tracked using the `heapq` or `heapq.heappush` module. \n- **Review Libraries**: Check if a library or framework (e.g., `heapq`, `heapq.heappush`) is not properly managing memory. \n- **Use Profiling Tools**: Tools like `valgrind` or GDB can help track memory allocation and usage patterns. \n- **Analyze Heap Usage**: Ensure heap memory is not exhausted by monitoring allocations and frees. \n- **Check for Proper Freeing**: Verify that every memory allocation is correctly freed using the appropriate `free` or `delete` functions. \n- **Use Debugging Tools**: Tools like GDB or `gmem` can help trace memory usage in the context of the program. \n- **Review Logs**: Look for logs related to memory allocation and usage to identify potential issues.", "retry_text": "- **Identify the Leak**: Use tools like `valgrind` or `gmem` to pinpoint where memory is being used incorrectly. \n- **Check Heap Usage**: Ensure all memory allocations and frees are properly tracked using the `heapq` or `heapq.heappush` module. \n- **Review Libraries**: Check if a library or framework (e.g., `heapq`, `heapq.heappush`) is not properly managing memory. \n- **Use Profiling Tools**: Tools like `valgrind` or GDB can help track memory allocation and usage patterns. \n- **Analyze Heap Usage**: Ensure heap memory is not exhausted by monitoring allocations and frees. \n- **Check for Proper Freeing**: Verify that every memory allocation is correctly freed using the appropriate `free` or `delete` functions. \n- **Use Debugging Tools**: Tools like GDB or `gmem` can help trace memory usage in the context of the program. \n- **Review Logs**: Look for logs related to memory allocation and usage to identify potential issues.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "ca9f375a692cb8e9", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "List the first three things you would check if training loss is NaN.", "text": "If training loss is NaN, you should check the following three things:\n\n1. **Data Validation**: Ensure the input data is properly formatted and contains valid values. Missing or incomplete data might cause the loss function to fail. \n2. **Model Initialization**: Verify that the model has been initialized correctly, including parameters and initializations. A uninitialized model might not learn effectively. \n3. **Training Progress**: Confirm that the training process has started and that the model has been updated properly, even if the loss is NaN at intermediate steps.", "retry_text": "The first three things to check if training loss is NaN are:\n\n1. **Data Validation**: Ensure the input data is properly formatted and contains valid values. Missing or incomplete data might cause the loss function to fail. \n2. **Model Initialization**: Verify that the model has been initialized correctly, including parameters and initializations. A uninitialized model might not learn effectively. \n3. **Training Progress**: Confirm that the training process has started and that the model has been updated properly, even if the loss is NaN at intermediate steps.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "df064c12eb1268c4", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "List the first three things you would check if training loss is NaN.", "text": "If training loss is NaN, you should check the following three things:\n\n1. **Missing Values**: Check for any missing entries in your training data that might affect the loss function calculation.\n2. **Non-Numeric Values**: Ensure that the training data contains only numeric values, as non-numeric values (e.g., strings or special characters) can cause the loss to be undefined.\n3. **Invalid Inputs**: Verify that there are no invalid inputs during training, such as division by zero or malformed data entries.", "retry_text": "If training loss is NaN, you should check the following three things:\n\n1. **Missing Values**: Ensure the training data contains no missing entries that could affect the loss function calculation. \n2. **Non-Numeric Values**: Verify that the training data contains only numeric values (e.g., integers, floats) instead of strings or special characters. \n3. **Invalid Inputs**: Check for invalid inputs like division by zero, malformed data entries, or errors in the training process.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "83bd51ee6dbacc14", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "List the first three things you would check if training loss is NaN.", "text": "If training loss is NaN, the first three things to check are: \n1. **Check for division by zero**: If the loss calculation involves division by zero (e.g., `loss = 1 / 0`), this could cause a NaN. \n2. **Verify model updates**: Ensure the model has been trained with enough parameters and the optimizer has optimized the loss function. \n3. **Confirm dataset completeness**: Ensure the dataset is fully available and does not contain missing or incomplete data.", "retry_text": "If training loss is NaN, the first three things to check are: \n1. **Check for division by zero** (e.g., `loss = 1 / 0`). \n2. **Verify model updates** (ensure the model has been trained with enough parameters and the optimizer has optimized the loss function). \n3. **Confirm dataset completeness** (ensure the dataset is fully available and contains no missing or incomplete data).", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "e344d44a16689e8e", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "List the first three things you would check if training loss is NaN.", "text": "If training loss is NaN, the first three things to check are: \n1. **Data Availability**: Ensure the training data is properly available and formatted. \n2. **Optimizer Initialization**: Verify that the optimizer (e.g., Adam, SGD) is initialized correctly. \n3. **Training Process**: Confirm that the training process has been executed for a sufficient number of epochs to allow convergence.", "retry_text": "If training loss is NaN, the first three things to check are: \n1. **Data Availability**: Ensure the training data is properly available and formatted. \n2. **Optimizer Initialization**: Verify that the optimizer (e.g., Adam, SGD) is initialized correctly. \n3. **Training Process**: Confirm that the training process has been executed for a sufficient number of epochs to allow convergence.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "a88294870a2c9f99", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "List the first three things you would check if training loss is NaN.", "text": "If training loss is NaN, the first three things to check are:\n\n1. **Check for division by zero** (e.g., dividing a number by zero in training loss calculations). \n2. **Verify data type** (e.g., ensuring the input data type matches expected values). \n3. **Check for missing or incomplete data** (e.g., missing training samples or values in the dataset).", "retry_text": "If training loss is NaN, the first three things to check are:\n\n1. **Check for division by zero** (e.g., dividing a number by zero in training loss calculations). \n2. **Verify data type** (e.g., ensuring the input data type matches expected values). \n3. **Check for missing or incomplete data** (e.g., missing training samples or values in the dataset).", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "fac14512075ca35c", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "List the first three things you would check if training loss is NaN.", "text": "If training loss is NaN, you should check the following three things:\n\n1. **Missing Values in the Training Data**: Ensure there are no missing entries in the training data that could cause issues with the loss function.\n2. **Model Configuration**: Verify that the model architecture is correctly defined and the weights and biases are initialized appropriately.\n3. **Optimizer Performance**: Check if the optimizer (e.g., Adam, SGD) is converging properly, which might lead to NaN values if it's not optimized effectively.", "retry_text": "List the first three things you would check if training loss is NaN:\n\n1. **Missing Values in the Training Data**: Ensure there are no missing entries in the training data that could cause issues with the loss function. \n2. **Model Configuration**: Verify that the model architecture is correctly defined and the weights and biases are initialized appropriately. \n3. **Optimizer Performance**: Check if the optimizer (e.g., Adam, SGD) is converging properly, which might lead to NaN values if it's not optimized effectively.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "3f23520b6e2895bd", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "List the first three things you would check if training loss is NaN.", "text": "If training loss is NaN, you should check the following:\n\n1. **Incomplete or Missing Training Data**: The training data may be incomplete or lack the necessary inputs for the model to compute the loss.\n2. **Incorrect Learning Rate**: A learning rate that is too large or too small could cause the loss to become undefined.\n3. **Optimizer Parameters**: The optimizer (e.g., Adam, RMSProp) may not be properly initialized or configured, leading to NaN values in the loss.", "retry_text": "Here are the first three things to check if training loss is NaN:\n\n1. **Incomplete or Missing Training Data**: The training data may be missing essential inputs required for the model to compute the loss. \n2. **Incorrect Learning Rate**: A learning rate that is too large or too small could cause the loss to become undefined. \n3. **Optimizer Parameters**: The optimizer (e.g., Adam, RMSProp) may not be properly initialized or configured, leading to NaN values in the loss.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "ec5763efa2dbc665", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "List the first three things you would check if training loss is NaN.", "text": "If training loss is NaN, consider the following first three checks:\n\n1. **Check the loss value is a valid number** (e.g., not a NaN or invalid value). \n2. **Verify the loss is being computed correctly** (e.g., inputs are properly fed, outputs are not empty or invalid). \n3. **Confirm the model is using a valid loss function** (e.g., not a numeric type or unsupported function).", "retry_text": "If training loss is NaN, consider the following first three checks:\n\n1. **Check the loss value is a valid number** (e.g., not a NaN or invalid value). \n2. **Verify the loss is being computed correctly** (e.g., inputs are properly fed, outputs are not empty or invalid). \n3. **Confirm the model is using a valid loss function** (e.g., not a numeric type or unsupported function).", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "e7c0bb256910c0a1", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Walk me through setting up a Python virtualenv, step by step.", "text": "To set up a Python virtual environment step-by-step, follow these instructions:\n\n1. **Install Python**: \n Ensure you have Python installed. You can check this by typing `python -V` or `python3 -V` in your terminal. If not installed, download it from [https://www.python.org/downloads](https://www.python.org/downloads).\n\n2. **Create a Virtual Environment**: \n Use the command: \n ```bash\n python -m venv <your_venv_name>\n ``` \n Replace `<your_venv_name>` with a name (e.g., `my_venv`). This creates a directory where your Python environment is stored. If the directory doesn't exist, the command will create it.\n\n3. **Activate the Virtual Environment**: \n - **Using the shell**: \n ```bash\n source <your_venv_name>/bin/activate\n ``` \n This activates the environment and uses `pip` commands for package management. \n - **Using the terminal**: \n ```bash\n activate <your_venv_name>\n ``` \n This also activates the environment.\n\n4. **Verify the Environment**: \n After activation, check if the environment is active by running: \n ```bash\n pip --version\n ``` \n This confirms that the virtual environment is correctly set up.\n\n5. **Using the Environment**: \n Once the virtual environment is active, you can install packages using `pip install` or run scripts in the same directory.\n\n**Note**: If you want to use a custom path, specify it in the `venv` command. Also, if you need to create a new virtualenv, ensure that the directory exists before running the command.", "retry_text": "To set up a Python virtual environment, follow these steps:\n\n1. **Install Python** \n Ensure you have Python installed. You can check this by typing `python -V` or `python3 -V` in your terminal. If not installed, download it from [https://www.python.org/downloads](https://www.python.org/downloads).\n\n2. **Create a Virtual Environment** \n Use the command: \n ```bash\n python -m venv <your_venv_name>\n ``` \n Replace `<your_venv_name>` with a name (e.g., `my_venv`). This creates a directory where your Python environment is stored. If the directory doesn't exist, the command will create it.\n\n3. **Activate the Virtual Environment** \n - **Using the shell**: \n ```bash\n source <your_venv_name>/bin/activate\n ``` \n This activates the environment and uses `pip` commands for package management. \n - **Using the terminal**: \n ```bash\n activate <your_venv_name>\n ``` \n This also activates the environment.\n\n4. **Verify the Environment** \n After activation, check if the environment is active by running: \n ```bash\n pip --version\n ``` \n This confirms that the virtual environment is correctly set up.\n\n5. **Using the Environment** \n Once the virtual environment is active, you can install packages using `pip install` or run scripts in the same directory.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "6a58c90d0c6b8c6a", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Walk me through setting up a Python virtualenv, step by step.", "text": "To set up a Python virtualenv step by step, follow this process:\n\n1. **Create the Virtualenv**: \n Use the command `python -m venv <env_name>` to create a new virtual environment. For example: \n ```bash \n python -m venv myvenv \n ``` \n This creates a folder named `myvenv` with the Python interpreter and packages for your project.\n\n2. **Activate the Virtualenv**: \n After creating the environment, activate it using: \n ```bash \n source myvenv/bin/activate \n ``` \n This makes the virtualenv available for executing code within the current working directory.\n\n3. **Install Dependencies**: \n If your project has dependencies, use `pip install` within the virtualenv: \n ```bash \n pip install -r requirements.txt \n ``` \n This installs all necessary packages for your project.\n\n4. **Use the Virtualenv**: \n Now, you can run code directly in the virtualenv: \n ```bash \n python your_script.py \n ``` \n This avoids conflicts with other Python environments.\n\n5. **Verify**: \n Check the path to ensure the virtualenv is active: \n ```bash \n python -c \"import sys; print(sys.path)\" \n ``` \n This will show where the virtualenv's packages are available.\n\nThis process ensures isolated, secure environments for project dependencies.", "retry_text": "To set up a Python virtualenv step by step:\n\n1. **Create the Virtualenv**: \n Use the command `python -m venv <env_name>` to create a new virtual environment. For example: \n ```bash \n python -m venv myvenv \n ``` \n\n2. **Activate the Virtualenv**: \n After creating the environment, activate it using: \n ```bash \n source myvenv/bin/activate \n ``` \n\n3. **Install Dependencies**: \n If your project has dependencies, use `pip install` within the virtualenv: \n ```bash \n pip install -r requirements.txt \n ``` \n\n4. **Use the Virtualenv**: \n Now, you can run code directly in the virtualenv: \n ```bash \n python your_script.py \n ``` \n\n5. **Verify**: \n Check the path to ensure the virtualenv is active: \n ```bash \n python -c \"import sys; print(sys.path)\" \n ``` \n\nThis process ensures isolated, secure environments for project dependencies.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "6012f002d10c0090", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Walk me through setting up a Python virtualenv, step by step.", "text": "To set up a Python virtualenv step by step, follow this process:\n\n1. **Install Python**: \n If not already installed, install Python using your system's package manager (e.g., `apt-get` on Ubuntu/Debian or `yum` on CentOS). Example command: \n ```bash\n # For Ubuntu/Debian\n sudo apt-get update\n sudo apt-get install -y python3\n # For CentOS\n sudo yum install -y python3\n ```\n\n2. **Create a Virtual Environment**: \n Use `venv` or `virtualenv` to create a new environment. The command is: \n ```bash\n python3 -m venv my_venv\n ```\n\n This creates a directory named `my_venv` with Python and dependencies. Ensure the environment is activated. \n **Activate the virtualenv** (change directory to the new environment): \n ```bash\n source my_venv/bin/activate\n ```\n\n3. **Install Packages**: \n If your project requires specific packages, install them using `pip install package_name`. For example: \n ```bash\n pip install requirements.txt\n ```\n\n4. **Use the Virtualenv**: \n Now, you can use the virtualenv to install packages and run scripts without conflicts with your main Python environment.\n\n**Example**: \n```bash\n# Create a virtualenv\npython3 -m venv my_venv\n\n# Activate the virtualenv\nsource my_venv/bin/activate\n\n# Install dependencies\npip install some_package\n```\n\nThis setup allows you to manage multiple Python environments easily.", "retry_text": "Here's a step-by-step walkthrough of setting up a Python virtualenv:\n\n1. **Install Python**: \n If not already installed, install Python using your system's package manager. For example: \n ```bash\n # For Ubuntu/Debian\n sudo apt-get update\n sudo apt-get install -y python3\n # For CentOS\n sudo yum install -y python3\n ```\n\n2. **Create a Virtual Environment**: \n Use `venv` or `virtualenv` to create a new environment. The command is: \n ```bash\n python3 -m venv my_venv\n ```\n\n This creates a directory named `my_venv` with Python and dependencies. Ensure the environment is activated. \n **Activate the virtualenv**: \n ```bash\n source my_venv/bin/activate\n ```\n\n3. **Install Packages**: \n If your project requires specific packages, install them using `pip install package_name`. For example: \n ```bash\n pip install requirements.txt\n ```\n\n4. **Use the Virtualenv**: \n Now, you can use the virtualenv to install packages and run scripts without conflicts with your main Python environment.\n\n**Example**: \n```bash\n# Create a virtualenv\npython3 -m venv my_venv\n\n# Activate the virtualenv\nsource my_venv/bin/activate\n\n# Install dependencies\npip install some_package\n```\n\nThis setup allows you to manage multiple Python environments easily.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "4f318c8ba1ed6f94", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "What is your plan for answering a question you are unsure about?", "text": "When unsure about a question, I plan to first search the web for the information using the web_search function. If the query is resolved through online resources, I'll proceed with that. If not, I'll use Python code execution to address any specific tasks or data analysis required. I'm also prepared to adjust my plan based on the results found.", "retry_text": "When unsure about a question, I plan to first search the web using the `web_search` function. If the information is resolved through online resources, I'll proceed with that. If not, I'll use Python code execution to address any specific tasks or data analysis required. I'm also prepared to adjust my plan based on the results found.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "095bb6719849ccf9", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "What is your plan for answering a question you are unsure about?", "text": "When unsure about a question, my approach is to use the tools available, such as web searches or sandbox code execution. I'll ask for clarification or provide relevant information based on the tools. If I need assistance with specific tasks, I'll use the tools provided to get the required information.", "retry_text": "When unsure about a question, my approach is to use tools like web searches or sandbox code execution to gather information. I will ask for clarification or provide relevant details. If no tools are needed, I will directly answer the question as requested.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "6466d33bf4055aba", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "What is your plan for answering a question you are unsure about?", "text": "When unsure about a question, my plan is to: \n1. **Research the topic** to gather reliable information. \n2. **Analyze the information** to determine the most accurate answer. \n3. **Provide a well-structured response** that addresses the question while being helpful and comprehensive. \n\nThis approach ensures clarity and addresses the users needs effectively.", "retry_text": "When unsure about a question, my plan is: \n1. **Research** to gather reliable information. \n2. **Analyze** the information to determine the most accurate answer. \n3. **Provide a well-structured response** that addresses the question while being helpful and comprehensive. \n\nThis approach ensures clarity and addresses the users needs effectively.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "f571eddf37233a27", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "What is your plan for answering a question you are unsure about?", "text": "When unsure about a question, my plan is to first use the web_search function to gather information if the question involves specific content or needs external data. Otherwise, I rely on my existing knowledge and domain expertise. I remain flexible and adapt my responses based on the context and the nature of the question.", "retry_text": "I don't need a tool to answer this question. When unsure about a question, my plan is to use the web_search function to gather information if the question requires external data or specific content, or rely on my existing knowledge and domain expertise otherwise. I remain flexible and adapt my responses based on the context and the nature of the question.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "287075bcccadaa02", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "What is your plan for answering a question you are unsure about?", "text": "When unsure about a question, my plan is to: \n1. **Verify facts** by cross-referencing information from reliable sources. \n2. **Stay informed** by staying updated on the topic to ensure accuracy. \n3. **Consider different perspectives** to find the most comprehensive answer. \n\nI am always open to learning and adapting my approach to provide the best response.", "retry_text": "When unsure about a question, my plan is to: \n1. **Verify facts** by cross-referencing information from reliable sources. \n2. **Stay informed** by staying updated on the topic to ensure accuracy. \n3. **Consider different perspectives** to find the most comprehensive answer. \n\nI am always open to learning and adapting my approach to provide the best response.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "9474398e6db4735a", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "What is your plan for answering a question you are unsure about?", "text": "When unsure about a question, my plan is to: \n1. **Research thoroughly** using reliable sources. \n2. **Ask clarifying questions** to deepen understanding. \n3. **Analyze the context** and consider different perspectives. \n4. **Stay open to new information** and adapt based on available resources. \n\nLet me know if you'd like to refine this approach!", "retry_text": "I don't need a tool to answer this question. When unsure, my plan is to: \n1. **Research thoroughly** using reliable sources. \n2. **Ask clarifying questions** to deepen understanding. \n3. **Analyze the context** and consider different perspectives. \n4. **Stay open to new information** and adapt based on available resources. \n\nLet me know if you'd like to refine this approach!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "16abe1d67280a8e4", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "What is your plan for answering a question you are unsure about?", "text": "When unsure about a question, I rely on available information and use the web_search function to gather details. I ensure thoroughness by checking the information thoroughly. If you have more context or specific details about the question, I can tailor my plan accordingly. Let me know!", "retry_text": "When unsure about a question, I rely on available information and use the web_search function to gather details. I ensure thoroughness by checking thoroughly. If you have more context or specific details about the question, I can tailor my plan accordingly. Let me know!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "d851acb957839ac2", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Describe your approach to reviewing a pull request, as a list.", "text": "Heres a structured approach to reviewing a pull request:\n\n- **Code Review for Correctness and Style**: \n - Ensure code is logically sound and follows project standards. \n - Check for syntax errors, logic flaws, and code duplication. \n\n- **Identify Issues and Bug Fixes**: \n - Locate and resolve any bugs, edge cases, or issues in the code. \n\n- **Ensure Maintainability**: \n - Verify that the code is easy to understand, testable, and scalable for future updates. \n\n- **Testing**: \n - Conduct unit tests, integration tests, and end-to-end tests to validate the code. \n\n- **Communication with the Team**: \n - Coordinate with the team to ensure alignment and address any concerns. \n\n- **Documentation**: \n - Update documentation to reflect the changes made in the pull request. \n\nThis approach ensures a thorough and constructive review process.", "retry_text": "Heres a structured approach to reviewing a pull request:\n\n- **Code Review for Correctness and Style**: \n - Ensure code is logically sound and follows project standards. \n - Check for syntax errors, logic flaws, and code duplication. \n\n- **Identify Issues and Bug Fixes**: \n - Locate and resolve any bugs, edge cases, or issues in the code. \n\n- **Ensure Maintainability**: \n - Verify that the code is easy to understand, testable, and scalable for future updates. \n\n- **Testing**: \n - Conduct unit tests, integration tests, and end-to-end tests to validate the code. \n\n- **Communication with the Team**: \n - Coordinate with the team to ensure alignment and address any concerns. \n\n- **Documentation**: \n - Update documentation to reflect the changes made in the pull request.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "5d1eee90ae0e4d93", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Describe your approach to reviewing a pull request, as a list.", "text": "Heres a structured approach to reviewing a pull request:\n\n1. **Code Review** \n - Check for bugs, syntax errors, and logical errors. \n - Ensure code follows coding standards and best practices. \n\n2. **Testing** \n - Validate functionality with unit tests and integration tests. \n - Ensure the code works as expected under various conditions. \n\n3. **Documentation** \n - Verify that the code is well-documented and explained. \n\n4. **Security** \n - Check for vulnerabilities (e.g., XSS, SQL injection) and security issues. \n\n5. **Code Quality** \n - Look for unused variables, inefficient loops, or code smells. \n\n6. **Contextual Understanding** \n - Understand the purpose of the pull request and the specific changes being made. \n\n7. **Impact on Other Components** \n - Ensure the code doesnt introduce new issues or conflicts with existing code. \n\n8. **Automated Tools** \n - Use tools (e.g., linters, static analyzers) to catch issues before integration. \n\nThis approach ensures a thorough review while maintaining efficiency and quality.", "retry_text": "Heres a structured approach to reviewing a pull request:\n\n1. **Code Review** \n - Check for bugs, syntax errors, and logical errors. \n - Ensure code follows coding standards and best practices. \n\n2. **Testing** \n - Validate functionality with unit tests and integration tests. \n - Ensure the code works as expected under various conditions. \n\n3. **Documentation** \n - Verify that the code is well-documented and explained. \n\n4. **Security** \n - Check for vulnerabilities (e.g., XSS, SQL injection) and security issues. \n\n5. **Code Quality** \n - Look for unused variables, inefficient loops, or code smells. \n\n6. **Contextual Understanding** \n - Understand the purpose of the pull request and the specific changes being made. \n\n7. **Impact on Other Components** \n - Ensure the code doesnt introduce new issues or conflicts with existing code. \n\n8. **Automated Tools** \n - Use tools (e.g., linters, static analyzers) to catch issues before integration. \n\nThis approach ensures a thorough review while maintaining efficiency and quality.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "ae18ee4c0a550b21", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Describe your approach to reviewing a pull request, as a list.", "text": "Here's a structured approach to reviewing a pull request:\n\n1. **Code Quality Checks** \n - ✅ Check for syntax errors, logical errors, and clean code. \n - ✅ Look for potential security vulnerabilities or code smells. \n\n2. **Documentation Review** \n - ✅ Ensure the PR includes well-documented changes and test cases. \n - ✅ Verify that any documentation related to the PR is updated and maintained. \n\n3. **Code Consistency** \n - ✅ Check for consistent coding style, naming conventions, and variable/constant usage. \n - ✅ Ensure all changes align with the projects existing codebase. \n\n4. **Test Coverage** \n - ✅ Verify that the PR includes test cases for the affected functionality. \n - ✅ Check for any edge cases or scenarios that might need special attention. \n\n5. **Issue Resolution** \n - ✅ Ensure the PR addresses any open issues or bugs in the project. \n - ✅ Confirm that the PR has been reviewed and approved by relevant stakeholders. \n\n6. **CI/CD Integration** \n - ✅ Check that the PR is prepared for integration into the main codebase (CI/CD pipeline). \n - ✅ Ensure that the changes are compatible with existing workflows. \n\n7. **Future-Proofing** \n - ✅ Review for architectural changes that could break future features. \n - ✅ Ensure the code is maintainable and scalable for long-term use. \n\nThis approach ensures that the PR is reviewed thoroughly and ready for merging into the main codebase.", "retry_text": "Here's a structured approach to reviewing a pull request: \n\n1. **Code Quality Checks** \n - ✅ Check for syntax errors, logical errors, and clean code. \n - ✅ Ensure potential security vulnerabilities or code smells are addressed. \n\n2. **Documentation Review** \n - ✅ Verify well-documented changes and test cases are included. \n - ✅ Confirm documentation aligns with the projects existing structure. \n\n3. **Code Consistency** \n - ✅ Check coding style, naming conventions, and variable/constant usage. \n - ✅ Ensure all changes align with the projects codebase standards. \n\n4. **Test Coverage** \n - ✅ Verify the PR includes test cases for the affected functionality. \n - ✅ Check edge cases or scenarios requiring special attention. \n\n5. **Issue Resolution** \n - ✅ Ensure the PR addresses any open issues or bugs in the project. \n - ✅ Confirm the changes are ready for review and approval. \n\n6. **CI/CD Integration** \n - ✅ Check compatibility with existing CI/CD workflows. \n - ✅ Ensure changes are prepared for integration into the main codebase. \n\n7. **Future-Proofing** \n - ✅ Review for architectural changes that could break future features. \n - ✅ Ensure the code is maintainable, scalable, and ready for long-term use.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "93f33dfafe4b553e", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Describe your approach to reviewing a pull request, as a list.", "text": "Here's a structured approach to reviewing a pull request:\n\n1. **Code Review**: Check the code for syntax errors, logic errors, and any potential issues. Ensure the implementation aligns with the project's coding standards.\n2. **Documentation**: Verify that the changes in the PR are documented in the codebase and that the documentation is updated accordingly.\n3. **Project Codebase**: Compare the PR changes with the existing codebase to ensure no missed or duplicated work. Check for any unintended side effects.\n4. **Security Vulnerabilities**: Look for any potential security flaws, such as insecure code or vulnerabilities that could be exploited.\n5. **Testing**: Ensure the changes pass all relevant tests, especially if the code is complex or new to the project.\n6. **Communication**: Communicate with the maintainers or the project's team to gather their feedback and address any concerns.\n\nThis approach ensures a thorough and constructive review of the PR.", "retry_text": "Here's a structured approach to reviewing a pull request:\n\n1. **Code Review**: Check the code for syntax errors, logic errors, and any potential issues. Ensure the implementation aligns with the project's coding standards. \n2. **Documentation**: Verify that the changes in the PR are documented in the codebase and that the documentation is updated accordingly. \n3. **Project Codebase**: Compare the PR changes with the existing codebase to ensure no missed or duplicated work. Check for any unintended side effects. \n4. **Security Vulnerabilities**: Look for any potential security flaws, such as insecure code or vulnerabilities that could be exploited. \n5. **Testing**: Ensure the changes pass all relevant tests, especially if the code is complex or new to the project. \n6. **Communication**: Communicate with the maintainers or the project's team to gather their feedback and address any concerns.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "e3307b82f25a0462", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Describe your approach to reviewing a pull request, as a list.", "text": "Here's a structured approach to reviewing a pull request:\n\n1. **Code Review Quality Check** \n - Ensure code meets project standards and has no syntax errors \n - Check for logical errors, performance issues, or security flaws \n\n2. **Code Documentation** \n - Verify the issue description is well-documented and addresses the original problem \n - Ensure any changes are clearly explained in the documentation \n\n3. **Testing** \n - Conduct unit tests and integration tests \n - Check for compatibility issues with existing code and dependencies \n\n4. **Issue Resolution** \n - Address any issues raised in the pull request \n - Ensure the pull request is ready for merging \n\n5. **Code Policy Compliance** \n - Confirm it aligns with the project's code review policies and guidelines \n\nThis approach ensures a thorough and constructive review process.", "retry_text": "Here's a structured approach to reviewing a pull request:\n\n1. **Code Review Quality Check** \n - Ensure code meets project standards and has no syntax errors \n - Check for logical errors, performance issues, or security flaws \n\n2. **Code Documentation** \n - Verify the issue description is well-documented and addresses the original problem \n - Ensure any changes are clearly explained in the documentation \n\n3. **Testing** \n - Conduct unit tests and integration tests \n - Check for compatibility issues with existing code and dependencies \n\n4. **Issue Resolution** \n - Address any issues raised in the pull request \n - Ensure the pull request is ready for merging \n\n5. **Code Policy Compliance** \n - Confirm it aligns with the project's code review policies and guidelines \n\nThis approach ensures a thorough and constructive review process.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "bddfdd9ea80c6169", "model": "qwen3-0.6b", "prompt_class": "format", "prompt": "Describe your approach to reviewing a pull request, as a list.", "text": "Here's a structured approach to reviewing a pull request:\n\n1. **Initial Review** \n - **Commit Message**: Ensure the message is clear, concise, and addresses the issue. \n - **Code Formatting**: Check for consistent formatting, including whitespace, spacing, and code blocks. \n - **Licensing and Policies**: Verify the code complies with project policies (e.g., MIT, GPL, or specific licensing agreements). \n\n2. **Code Quality Check** \n - **Code Documentation**: Ensure changes are well-documented and align with existing documentation. \n - **Code Standards**: Confirm the code follows the project's coding guidelines (e.g., PEP 8, style rules). \n\n3. **Issue Resolution** \n - **Issue Description**: Validate the issue description is clear and the fix addresses the issue. \n - **Code for Fixes**: Review the fix implementation and ensure it solves the problem effectively. \n\n4. **Testing** \n - **Unit and Integration Tests**: Ensure the code works as intended in tests (e.g., unit tests, integration tests). \n\n5. **History Review** \n - **Conflict Checks**: Ensure there are no duplicate or conflicting changes in the history. \n - **PR History**: Confirm the PR includes the correct code and no unnecessary changes. \n\n6. **Documentation** \n - **Update Documentation**: Update documentation to reflect the changes made in the pull request. \n\n7. **Final Approval** \n - **Final Check**: Confirm all points are addressed and the PR is ready for review. \n\nThis approach ensures a thorough review while maintaining efficiency and quality.", "retry_text": "Here's a structured approach to reviewing a pull request:\n\n1. **Initial Review** \n - **Commit Message**: Ensure the message is clear, concise, and addresses the issue. \n - **Code Formatting**: Check for consistent formatting, including whitespace, spacing, and code blocks. \n\n2. **Code Quality Check** \n - **Code Documentation**: Ensure changes are well-documented and align with existing documentation. \n - **Code Standards**: Confirm the code follows the project's coding guidelines (e.g., PEP 8, style rules). \n\n3. **Issue Resolution** \n - **Issue Description**: Validate the issue description is clear and the fix addresses the issue. \n - **Code for Fixes**: Review the fix implementation and ensure it solves the problem effectively. \n\n4. **Testing** \n - **Unit and Integration Tests**: Ensure the code works as intended in tests (e.g., unit tests, integration tests). \n\n5. **History Review** \n - **Conflict Checks**: Ensure there are no duplicate or conflicting changes in the history. \n - **PR History**: Confirm the PR includes the correct code and no unnecessary changes. \n\n6. **Documentation** \n - **Update Documentation**: Update documentation to reflect the changes made in the pull request. \n\n7. **Final Approval** \n - **Final Check**: Confirm all points are addressed and the PR is ready for review. \n\nThis approach ensures a thorough review while maintaining efficiency and quality.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "2f578170c686ad08", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Can you look into that for me?", "text": "I can help you with something. Could you please specify what you need assistance with?", "retry_text": "I can help you with something. Could you please specify what you need assistance with?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "0156c4d90b02722b", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Can you look into that for me?", "text": "I can help you with that. Could you please specify what you need assistance with?", "retry_text": "I can help you with that. Could you please specify what you need assistance with?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "19beab9b9b904353", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Can you look into that for me?", "text": "I can assist you if you need help with a specific task. Could you please clarify what you'd like me to do? For example, is there a web search you'd like to perform, or is there a coding problem you'd like to solve?", "retry_text": "I can assist you if you need help with a specific task. Could you please clarify what you'd like me to do? For example, is there a web search you'd like to perform, or is there a coding problem you'd like to solve?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "e077760a025ac077", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Can you look into that for me?", "text": "I can help you with that! Could you please clarify what you need assistance with? For example, are you looking for information, a solution to a problem, or a specific action to take?", "retry_text": "I can help you with that! Could you please clarify what you need assistance with? For example, are you looking for information, a solution to a problem, or a specific action to take?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "8b5d6af1da80759a", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Can you look into that for me?", "text": "I can help you with that! Could you please specify what you need assistance with?", "retry_text": "Could you please specify what you need assistance with?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "f989b1b1d2d66995", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Can you look into that for me?", "text": "I can't look into that. I can only perform web searches or execute Python code. Could you please provide more details or specify what you're looking for?", "retry_text": "I cannot look into that. I can only perform web searches or execute Python code. Could you please provide more details or specify what you're looking for?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "069078e3f77bd1f3", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Can you look into that for me?", "text": "Sure! Could you please tell me what you need assistance with?", "retry_text": "Sure! Could you please tell me what you need assistance with?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "d1a015f991ab3d00", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Can you look into that for me?", "text": "I can help you with that. Could you please clarify what specific task or query you need assistance with?", "retry_text": "I can help you with that. Could you please clarify what specific task or query you need assistance with?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "8ba7363ec8d4808b", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Fix it.", "text": "Could you please specify what \"Fix it\" refers to in your context? For example, is it a task to be completed, a code issue, a configuration problem, or something else? This will help me provide you with the most accurate assistance.", "retry_text": "Could you please specify what \"Fix it\" refers to in your context? For example, is it a task to be completed, a code issue, a configuration problem, or something else? This will help me provide you with the most accurate assistance.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "dd46a0139e68687c", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Fix it.", "text": "I'm sorry, but I'm not sure what \"Fix it\" refers to. Could you please provide more context or specify the issue you'd like me to address?", "retry_text": "I'm sorry, but I'm not sure what \"Fix it\" refers to. Could you please provide more context or specify the issue you'd like me to address?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "09cbf0f6bc16ece4", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Fix it.", "text": "Could you please specify what \"Fix it\" refers to? For example, is there a task, a requirement, or a problem that needs fixing? This will help me provide you with the appropriate assistance.", "retry_text": "Could you please specify what \"Fix it\" refers to? For example, is there a task, a requirement, or a problem that needs fixing? This will help me provide you with the appropriate assistance.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "5d1c121db57e292c", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Fix it.", "text": "Could you please specify what issue or problem you're referring to? I'll be able to help you fix it.", "retry_text": "Could you please specify what issue or problem you're referring to? I'll be able to help you fix it.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "92551203a59734bf", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Fix it.", "text": "Could you please specify what issue or problem you'd like to fix?", "retry_text": "Could you please specify what issue or problem you'd like to fix?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "71a8e1df216e3c57", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Fix it.", "text": "Could you please specify what \"Fix it\" refers to? For example, is there a task, problem, or issue you want to address? Let me know and I'll do my best to help.", "retry_text": "Could you please specify what \"Fix it\" refers to? For example, is there a task, problem, or issue you want to address? Let me know and I'll do my best to help.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "3fd18b05b0bf9744", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Fix it.", "text": "Could you please provide more context or specify what \"Fix it\" refers to in your situation?", "retry_text": "The message \"Fix it\" doesn't contain a specific request or task to address. To proceed, could you please provide more context or clarify what \"Fix it\" refers to in your situation?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "63b05dc7bff381cd", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Tell me about the thing we discussed.", "text": "I don't have access to a conversation history or previous information. Could you please clarify what \"the thing we discussed\" refers to? Let me know and I'll do my best to assist.", "retry_text": "I don't have access to a conversation history or previous information. Could you please clarify what \"the thing we discussed\" refers to? Let me know and I'll do my best to assist.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "513c9d08d9dd0f83", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Tell me about the thing we discussed.", "text": "I can't tell you about what we discussed without more context. Could you please specify the topic or provide details about the subject we've been discussing?", "retry_text": "I cannot proceed without more context. Could you please specify the subject or topic we discussed?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "986805e14cfb085e", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Tell me about the thing we discussed.", "text": "Could you please specify what \"the thing we discussed\" refers to so I can provide a detailed response?", "retry_text": "Could you please specify what \"the thing we discussed\" refers to so I can provide a detailed response?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "3a8cc99a4fd7102d", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Tell me about the thing we discussed.", "text": "I don't have access to specific information or a conversation history in this context. Could you please provide more details or clarify what you're referring to?", "retry_text": "I don't have access to specific information or a conversation history in this context. Could you please provide more details or clarify what you're referring to?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "1a596b971a3932b3", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Tell me about the thing we discussed.", "text": "I can't tell you about the topic we discussed without more information. Could you please specify what the subject was or provide details about the conversation?", "retry_text": "I cannot provide a detailed explanation about the topic we discussed without specific information or context. Could you please share more details or clarify what the subject was or the conversation was about?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "b321d857119f0e1a", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Tell me about the thing we discussed.", "text": "I can't determine the specific topic without additional context. Could you please clarify what topic or subject you're interested in?", "retry_text": "I cannot determine the specific topic without additional context. Could you please clarify what topic or subject you're interested in?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "f27d000a239b9998", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Tell me about the thing we discussed.", "text": "I don't have access to your conversation history. Could you please provide more context or specify what \"the thing we discussed\" refers to?", "retry_text": "I don't have access to your conversation history. Could you please provide more context or specify what \"the thing we discussed\" refers to?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "df692d56e5002f3b", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Tell me about the thing we discussed.", "text": "I don't have access to previous conversation history. Could you please specify what the \"thing we discussed\" refers to?", "retry_text": "I don't have access to previous conversation history. Could you please specify what the \"thing we discussed\" refers to?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "e2051d39a9ee4a41", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Is it better?", "text": "The question \"Is it better?\" is a bit ambiguous. Could you please provide more context or clarify what specific aspect you're referring to (e.g., which decision, choice, or situation you're comparing)? This will help me better respond.", "retry_text": "The question \"Is it better?\" is ambiguous and lacks specificity. To assist you, please clarify what aspect or context you are referring to (e.g., which decision, choice, or situation youre comparing).", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "587cc3f962fa5e14", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Is it better?", "text": "I can't determine \"Is it better?\" without more context or specific information. Could you please provide the question or clarify what you're referring to?", "retry_text": "The question \"Is it better?\" is ambiguous without additional context. Could you please specify the question or clarify what you're referring to?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "b15957ffbe039c9e", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Is it better?", "text": "The question is a bit ambiguous. Could you clarify what you're asking? For example, are you comparing two options, deciding which is better, or looking for information on a specific topic? This would help me provide a more accurate response.", "retry_text": "The question is ambiguous. Could you clarify what you're asking? For example, are you comparing two options, deciding which is better, or looking for information on a specific topic?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "a358f81fdb3324a7", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Is it better?", "text": "The choice between using `web_search` or `python` depends on your specific needs. Here's how each tool can help:\n\n- **`web_search`**: Ideal for general information retrieval, as it can fetch full-page text from URLs. Useful for finding answers to questions or solving problems.\n- **`python`**: Great for executing code within a sandbox environment, which is perfect for tasks like data processing, automation, or scripting.\n\nIf you have a specific task or question in mind, let me know what tools you'd like to use!", "retry_text": "The choice between `web_search` and `python` depends on your specific needs. Here's how each tool functions:\n\n- **`web_search`**: Fetches full-page text from URLs to retrieve information, ideal for solving problems or finding answers to questions.\n- **`python`**: Executes code in a sandbox environment, suitable for tasks like data processing, automation, or scripting.\n\nIf you have a specific task, let me know what tools youd like to use!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "7ecabf000158e586", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "Is it better?", "text": "Is it better? Please provide the specific context or choices you're comparing, and I can help you determine the best option.", "retry_text": "The question \"Is it better?\" is ambiguous and lacks context or specific choices to compare. To assist you, please provide the specific question, comparison criteria, or options you're evaluating. For example, if you're comparing two products, services, or decisions, clarify what you're referring to. Without this information, I can't determine the best option.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "0cebe6188ebb1145", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "What should I do next?", "text": "To help you better, could you please provide more details or clarify what \"next\" refers to in your context? For example, are you looking for a specific action, a step-by-step guide, or assistance with something else?", "retry_text": "To assist you better, could you please provide more details or clarify what \"next\" refers to in your context? For example, are you looking for a specific action, a step-by-step guide, or something else?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "354119601b55966d", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "What should I do next?", "text": "To assist you effectively, could you please share more details about the situation or question you're facing? For example, are you looking for help with a specific problem, a request, or guidance on a topic?", "retry_text": "To assist you effectively, could you please share more details about the situation or question you're facing? For example, are you looking for help with a specific problem, a request, or guidance on a topic?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "4997cc2bb2015a82", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "What should I do next?", "text": "Could you please provide more details about the situation or the query you're facing?", "retry_text": "Could you please provide more details about the situation or the query you're facing?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "cbe166be7cdfccfe", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "What should I do next?", "text": "To help you better, could you please provide more details about your situation or the specific question you're facing?", "retry_text": "Could you please provide more details about your situation or the specific question you're facing?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "e3c3a09b2135bb07", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "What should I do next?", "text": "I can help you with specific questions. Could you please provide more details or clarify what you're looking for?", "retry_text": "Could you please specify what you need help with?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "eec98c2dd8620558", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "What should I do next?", "text": "Could you please clarify what you need assistance with? For example, are you looking for information, a solution to a problem, or help with something specific?", "retry_text": "Could you please clarify what you need assistance with? For example, are you looking for information, a solution to a problem, or help with something specific?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "55159a82d8ce1dc7", "model": "qwen3-0.6b", "prompt_class": "ambiguous", "prompt": "What should I do next?", "text": "The question \"What should I do next?\" is general and doesn't provide enough context for me to determine the appropriate action. Could you please specify the situation or problem you're facing so I can assist you better?", "retry_text": "The question \"What should I do next?\" is general and doesn't provide specific context for me to determine the appropriate action. Could you please specify the situation or problem you're facing so I can assist you better?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "5e31f765cb50a4e8", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Quelle est la capitale de la France ?", "text": "La capitale de la France est Paris.", "retry_text": "La capitale de la France est Paris.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "5c719501e2307082", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Quelle est la capitale de la France ?", "text": "La capitale de la France est Paris.", "retry_text": "La capitale de la France est Paris.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "ebb10de7c8d9f6cd", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Quelle est la capitale de la France ?", "text": "La capitale de la France est Lyon.", "retry_text": "La capitale de la France est Lyon.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "3eefd2de8fb723f7", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Quelle est la capitale de la France ?", "text": "La capitale de la France est Paris.", "retry_text": "La capitale de la France est Paris.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "0cca10877add3e01", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Quelle est la capitale de la France ?", "text": "La capitale de la France est Paris.", "retry_text": "La capitale de la France est Paris.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "a06f840c1397b49a", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Quelle est la capitale de la France ?", "text": "La capitale de la France est Paris.", "retry_text": "La capitale de la France est Paris.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "b9fae31106f76bfc", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Quelle est la capitale de la France ?", "text": "La capitale de la France est Paris.", "retry_text": "La capitale de la France est Paris.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "1bd8528c0be010cb", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Quelle est la capitale de la France ?", "text": "La capitale de la France est Paris.", "retry_text": "La capitale de la France est Paris.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "2fe04e2076dbadad", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Explica en una frase que es el sobreajuste.", "text": "El sobreajuste es el fenómeno en el que una persona exagera su capacidad o habilidad, lo que lleva a una desviación en sus actos o resultados.", "retry_text": "El sobreajuste es el fenómeno en el que una persona exagera su capacidad o habilidad, lo que lleva a una desviación en sus actos o resultados.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "d09e13931c8de2f5", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Explica en una frase que es el sobreajuste.", "text": "El sobreajuste es cuando el monto de la deuda es mayor que lo que debe pagar.", "retry_text": "El sobreajuste es cuando el monto de la deuda es mayor que lo que debe pagar.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "933c9d680abe10f3", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Explica en una frase que es el sobreajuste.", "text": "Sobreajuste es cuando una actividad o sistema se enfrenta a una carga sobrecarga que puede no ser resuelta debido a la capacidad limitada de los recursos.", "retry_text": "Sobreajuste es la situación en la que una actividad o sistema enfrenta una carga sobrecarga que no puede ser resuelta debido a la capacidad limitada de los recursos.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "91712a4316d56f06", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Explica en una frase que es el sobreajuste.", "text": "El sobreajuste se refiere a la excesiva consumo de recursos o recursos materiales.", "retry_text": "El sobreajuste se refiere a la excesiva consumo de recursos o recursos materiales.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "19851173a737b424", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Explica en una frase que es el sobreajuste.", "text": "El sobreajuste ocurre cuando una persona experimenta una excesiva carga física o mental, causando desgaste.", "retry_text": "El sobreajuste ocurre cuando una persona experimenta una excesiva carga física o mental, causando desgaste.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "6273bf296ba0d580", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Explica en una frase que es el sobreajuste.", "text": "El sobreajuste es un sobreajuste, es decir, una pérdida financiera o un impago que ocurre cuando se enfrentan desafíos económicos o operaciones no esperadas.", "retry_text": "El sobreajuste es una pérdida financiera o un impago que ocurre cuando se enfrentan desafíos económicos o operaciones no esperadas.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "c5fc34cdf315ee06", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Was ist der Unterschied zwischen RAM und Festplatte?", "text": "RAM (Random Access Memory) and a hard drive (Festplatte) are two different types of storage and memory in computers. Here's the distinction:\n\n1. **RAM (Memory):**\n - **Purpose:** Temporary storage for active data and running applications.\n - **Speed:** Accesses data quickly (typically faster than a hard drive).\n - **Volatility:** Data is lost when the computer turns off.\n - **Usage:** Used for tasks that require immediate access, like browsing, running programs, or processing data.\n\n2. **Hard Drive (Festplatte):**\n - **Purpose:** Permanent storage for files, programs, and data.\n - **Speed:** Accesses data slowly compared to RAM.\n - **Durability:** Stores information long-term without losing data.\n - **Usage:** For storing software, user data, and backups.\n\nRAM is volatile, while a hard drive is non-volatile.", "retry_text": "RAM (Random Access Memory) and a hard drive (Festplatte) are two types of storage in computers. Here's the key difference:\n\n1. **RAM (Memory):**\n - **Purpose:** Temporary storage for active data and running applications.\n - **Speed:** Accesses data quickly (typically faster than a hard drive).\n - **Volatility:** Data is lost when the computer turns off.\n - **Usage:** Used for tasks requiring immediate access, like browsing, running programs, or processing data.\n\n2. **Hard Drive (Festplatte):**\n - **Purpose:** Permanent storage for files, programs, and data.\n - **Speed:** Accesses data slowly compared to RAM.\n - **Durability:** Stores information long-term without losing data.\n - **Usage:** For storing software, user data, and backups.\n\nRAM is volatile and loses data when the computer turns off, while a hard drive holds data permanently.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "56b6ed34e8ccc3b0", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Was ist der Unterschied zwischen RAM und Festplatte?", "text": "RAM (Random Access Memory) and a hard drive (Festplatte) serve different purposes in a computer:\n\n1. **RAM (Memory):** \n - Holds data temporarily for quick access. It's used to run programs and store information currently being used by the CPU. \n - Is consumed when the computer is actively processing data (e.g., running a program). It's faster than the hard drive because it's accessed directly by the CPU.\n\n2. **Hard Drive (Festplatte):** \n - Stores data permanently. It's used to store files, system information, and software. \n - Is slower than RAM because it requires a separate processor to access the data.\n\n**Key Difference:** \nRAM is for quick, temporary tasks, while the hard drive is for long-term storage.", "retry_text": "RAM (Random Access Memory) and a hard drive (Festplatte) serve different purposes in a computer:\n\n1. **RAM (Memory):** \n - Stores data temporarily for quick access. It's used to run programs and hold information currently being used by the CPU. It's faster than the hard drive because it's directly accessed by the CPU.\n\n2. **Hard Drive (Festplatte):** \n - Stores data permanently. It holds files, system information, and software. It's slower than RAM because it requires a separate processor to access data.\n\n**Key Difference:** \nRAM is for quick, temporary tasks, while the hard drive stores long-term data.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "ae569736c21c98ab", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Was ist der Unterschied zwischen RAM und Festplatte?", "text": "RAM (Random Access Memory) is temporary storage used to hold data that is currently being used by the CPU. It's volatile, meaning it loses its data when the power goes out. Unlike a hard drive, which stores data permanently, RAM is used for quick data access. \n\nA hard drive (or SSD) stores data permanently, and its speed depends on whether it's an SSD (non-volatile) or an HDD (volatile). SSDs are faster because they don't use mechanical components, while HDDs are slower. The hard drive can be read from (RAID) or not (non-Raid) depending on its configuration.", "retry_text": "RAM (Random Access Memory) is temporary storage used to hold data that the CPU is currently processing. It's volatile, meaning it loses its data when the power goes out. A hard drive (or SSD) stores data permanently, and its speed depends on whether it's an SSD (non-volatile) or an HDD (volatile). RAM is used for quick data access, while a hard drive is for long-term storage.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "9a67df1137489d09", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Was ist der Unterschied zwischen RAM und Festplatte?", "text": "RAM (Random Access Memory) is a type of temporary storage used to hold active data while a computer is running programs. It's volatile, meaning it's not as stable as a hard drive, which stores data permanently. RAM is used to quickly access and process data as needed, whereas a hard drive (or SSD) stores data permanently, allowing for faster access times.", "retry_text": "RAM (Random Access Memory) is a type of temporary storage used to hold active data while a computer is running programs. It is volatile, meaning it loses its data when the power is off. RAM is used to quickly access and process data as needed, whereas a hard drive (or SSD) stores data permanently, allowing for faster access times.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "ae4fb07db438690b", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Was ist der Unterschied zwischen RAM und Festplatte?", "text": "RAM (Random Access Memory) is a type of volatile memory used to hold data temporarily. It allows quick access to programs and data, making it essential for running applications or waiting for input. In contrast, a hard drive (Festplatte) is a non-volatile storage device used to store permanent data and software. It doesn't require constant power and can hold more data than RAM. Both components work together to support the computer's operations, with RAM being temporary and the hard drive storing permanent information.", "retry_text": "RAM (Random Access Memory) is a type of volatile memory used to hold data temporarily, while a hard drive (Festplatte) is a non-volatile storage device used to store permanent data and software. RAM allows quick access to programs and data, whereas the hard drive holds permanent information and storage. Both components work together to support the computer's operations.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "f71d15a0eaffe179", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Was ist der Unterschied zwischen RAM und Festplatte?", "text": "RAM (Random Access Memory) is a type of temporary storage used to hold active data and instructions that the CPU is currently processing. It's faster than a hard drive because it's directly accessed by the CPU, allowing for quick data retrieval. A hard drive, in contrast, is a long-term storage device that holds data and operating system files, which can be accessed and saved over time. RAM is typically smaller in capacity but offers faster access speeds.", "retry_text": "RAM (Random Access Memory) and a hard drive are both storage components, but they serve different purposes: \n- **RAM** is a type of temporary memory used to hold active data and instructions the CPU is currently processing. It is faster than a hard drive because it's directly accessed by the CPU, allowing for quick data retrieval. \n- **Hard drive** holds data and operating system files over time, meaning it's less volatile and can store large amounts of information without losing it. \n\nRAM is typically smaller in capacity but provides faster access speeds.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "3333944ab8b3b14d", "model": "qwen3-0.6b", "prompt_class": "non_english", "prompt": "Was ist der Unterschied zwischen RAM und Festplatte?", "text": "RAM (Random Access Memory) is a type of temporary storage used in computers to hold data and instructions temporarily. It's volatile, meaning it loses data when the computer turns off. The hard drive, on the other hand, is a non-volatile storage device that holds the operating system and user data. It's not replaced, but it can be upgraded or replaced over time. RAM is used for quick access to data, while the hard drive stores more data.", "retry_text": "RAM (Random Access Memory) is a type of volatile storage used to hold data and instructions temporarily. It loses data when the computer turns off. A hard drive (HDD) is non-volatile and stores operating system and user data, not replaced over time. RAM is used for quick access, while HDDs store more data.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "22a85371668ee43a", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Thanks. So which version should I install?", "text": "The latest llama.cpp release version is **b6291** (2026-07-14). This version includes improvements to the Vulkan backend and faster prompt processing on CUDA. You should follow the official documentation or GitHub release notes for the most accurate installation instructions.", "retry_text": "The latest llama.cpp release version is **b6291** (2026-07-14). This version includes improvements to the Vulkan backend and faster prompt processing on CUDA. You should follow the official documentation or GitHub release notes for the most accurate installation instructions.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "4cf325393db5369b", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Thanks. So which version should I install?", "text": "The latest available version is **b6291** (2026-07-14). This version includes improvements such as Vulkan backend enhancements and faster prompt processing on CUDA.\n\nWould you like additional details about this version or any updates beyond this release?", "retry_text": "The latest available version is **b6291** (2026-07-14). This version includes improvements such as Vulkan backend enhancements and faster prompt processing on CUDA. Let me know if you need further details!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "544a64cdb96703fa", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Thanks. So which version should I install?", "text": "The latest available version for Llama.cpp is **2026-07-14**, which includes improvements like Vulkan backend enhancements and faster prompt processing on CUDA. You can install it directly from the official Llama.cpp GitHub repository. For installation instructions, refer to the [Llama.cpp documentation](https://github.com/llama-project/llama.cpp) or the [GitHub release page](https://github.com/llama-project/llama.cpp/releases). Let me know if you need further assistance!", "retry_text": "The latest available version for Llama.cpp is **2026-07-14**, which includes improvements like Vulkan backend enhancements and faster prompt processing on CUDA. You can install it directly from the official Llama.cpp GitHub repository. For installation instructions, refer to the [Llama.cpp documentation](https://github.com/llama-project/llama.cpp) or the [GitHub release page](https://github.com/llama-project/llama.cpp/releases). Let me know if you need further assistance!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "29a0818174227203", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Thanks. So which version should I install?", "text": "The latest llama.cpp version is **b6291** (2026-07-14). This release includes improvements for the Vulkan backend and faster prompt processing on CUDA. Let me know if you need further details!", "retry_text": "The latest llama.cpp version is **b6291** (2026-07-14). This release includes improvements for the Vulkan backend and faster prompt processing on CUDA. Let me know if you need further details!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "e200b9622f9f9588", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Thanks. So which version should I install?", "text": "The latest known llama.cpp release version is **b6291** (2026-07-14). This version includes improvements such as Vulkan backend enhancements and faster prompt processing on CUDA. For the most accurate and up-to-date information, you can check the official llama.cpp documentation. Let me know if you need further details!", "retry_text": "The latest known llama.cpp release version is **b6291** (2026-07-14). This version includes improvements such as Vulkan backend enhancements and faster prompt processing on CUDA. Make sure to install it using the official documentation for the most accurate and up-to-date information. Let me know if you need further details!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "d7da157d532ba3d6", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Thanks. So which version should I install?", "text": "The latest llama.cpp release version is **b6291 (2026-07-14)**. This version includes improvements such as enhanced Vulkan backend support and faster prompt processing on CUDA. You can install it via the official llama.cpp repository or via your environment setup. Let me know if you need further details!", "retry_text": "The latest llama.cpp release version is **b6291 (2026-07-14)**. This version includes improvements such as enhanced Vulkan backend support and faster prompt processing on CUDA. You can install it via the official llama.cpp repository or your environment setup. Let me know if you need further details!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "a218f8ce82494ee9", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Thanks. So which version should I install?", "text": "The latest llama.cpp release is version **b6291** (2026-07-14). For installation, please refer to the official documentation or release notes for specific instructions. If you need help with dependencies or installation steps, I can assist further!", "retry_text": "The latest llama.cpp release is version **b6291** (2026-07-14). For installation, refer to the official documentation or release notes for specific instructions. Let me know if you need further assistance!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "3ce863a497159351", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Do I need a coat?", "text": "Based on the current weather in Reykjavik (9°C, overcast, wind 22 km/h, humidity 81%), it's **not extremely cold**. The wind and high humidity might make it feel a bit chilly, but it's still a mild day. You might need a coat, especially if the wind is strong, but it's not overly harsh. If you're planning to go outdoors, a light jacket or a sweater could be sufficient.", "retry_text": "Based on the current weather in Reykjavik (9°C, overcast, wind 22 km/h, humidity 81%), it's **not extremely cold**. The wind and high humidity might make it feel a bit chilly, but it's still a mild day. You might need a coat, especially if you plan to go outdoors, but it's not overly harsh. A light jacket or sweater could be sufficient.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "9d59823d666f5eee", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Do I need a coat?", "text": "Given the current weather in Reykjavik (9°C, overcast, wind 22 km/h, humidity 81%), it is highly recommended to wear a **coat** to cope with the cold and potentially high humidity. While the temperature is relatively mild, the combination of wind and humidity can make the environment feel slightly colder. If you plan to go out, packing layers and a jacket is a good precaution.", "retry_text": "Yes, you should wear a coat in Reykjavik. The current weather is 9°C with wind and high humidity, which can make temperatures feel slightly colder. A coat is a good precaution to protect against the elements.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "0b32a9c799a643ef", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Do I need a coat?", "text": "Based on the current weather: 9°C with overcast skies and a breeze, a light jacket or sweater is recommended. If it's expected to be cold (e.g., snow or rain), a coat might be necessary. However, since the temperature is mild, a jacket is sufficient.", "retry_text": "Based on the current weather (9°C with overcast skies and a breeze), a light jacket or sweater is recommended. Since the temperature is mild, a coat may not be necessary.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "b28dcd77e59abd11", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Do I need a coat?", "text": "If the temperature is 9°C in Reykjavik, it's generally not very cold, but it's still a good idea to have a coat in light rain or snowfall. In this case, the weather is overcast, so it's possible there's a chance of rain. While a coat might not be necessary, it's still a prudent precaution to stay prepared.", "retry_text": "Based on the current weather in Reykjavik (9°C, overcast, light rain), a coat is still a prudent precaution to stay prepared, even if it's not very cold.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "8c6687b7138a1e68", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Do I need a coat?", "text": "In Reykjavik, the current weather is 9°C with overcast conditions. While the temperature is not extremely cold, the wind from the north east and high humidity (81%) may make it slightly chilly. In general, a coat is not necessary for such a mild temperature. However, if the wind is strong and the temperature is still warm, a coat could help. The exact need for a coat depends on the wind's strength and the time of year.", "retry_text": "In Reykjavik, the current weather is 9°C with overcast conditions. The wind and humidity (81%) may make it slightly chilly, but a coat is not necessary for such a mild temperature. If the wind is strong or the temperature remains warm, a coat could help.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "aa9668cc9c6353c6", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Do I need a coat?", "text": "Based on the current weather in Reykjavik (9°C, overcast, wind from the north east, humidity 81%), it's likely that you'll need a light jacket or sweater. The temperature is above freezing, so a warm layer is essential for comfort. If you don't have a jacket, a sweater could be a good alternative. Make sure to prepare for the weather!", "retry_text": "Based on the current weather in Reykjavik (9°C, overcast, wind from the north east, humidity 81%), it's likely that you will need a light jacket or sweater. The temperature is above freezing, so a warm layer is essential for comfort. If you don't have a jacket, a sweater could be a good alternative. Make sure to prepare for the weather!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "766a83d5188fad2f", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Do I need a coat?", "text": "Based on the current weather in Reykjavik, which is **9°C (40°F)** with **overcast** conditions, there isn't a need for a heavy coat right now. The mild temperature combined with light clouds suggests it's a dry day, so a light jacket or a sweater should be sufficient. If it starts to rain soon, a light jacket is still a good option. Let me know if you'd like more details!", "retry_text": "Based on the current weather in Reykjavik, which is **9°C (40°F)** with **overcast** conditions, there isn't a need for a heavy coat. A light jacket or sweater is sufficient. If you're planning to go outside, a light clothing item will be appropriate. Let me know if you'd like further details!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "422a14dde624fb20", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Do I need a coat?", "text": "Based on the current weather conditions in Reykjavik (9°C, overcast, wind at 22 km/h, and humidity at 81%), it is likely to be chilly. However, while it's not extremely cold, it's possible that the weather may still require a coat. If the temperature drops below 0°C, a coat would be essential, but with 9°C, it's more likely to be a mild day. Therefore, you may need a coat, but it's not extreme.", "retry_text": "Based on the current weather conditions in Reykjavik (9°C, overcast, wind at 22 km/h, and humidity at 81%), it is not extremely cold, but the weather is still variable. While it's not expected to drop below freezing, the chances of rain or snow are present. A coat might be necessary, but it's not extreme.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "6200c8ed5ffeafea", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Is that number even or odd?", "text": "The 40th Fibonacci number is **even**. This is because the Fibonacci sequence alternates between odd and even starting from 0 and 1, and since 40 is even, the 40th term is even.", "retry_text": "The 40th Fibonacci number is **odd**.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "ece827001e0f21af", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Is that number even or odd?", "text": "The 40th Fibonacci number is **even**. \n\n**Answer:** Even.", "retry_text": "The 40th Fibonacci number is **even**. The number is **102334155**.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "807e08728fab0668", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Is that number even or odd?", "text": "The 40th Fibonacci number is **odd**.", "retry_text": "The 40th Fibonacci number is **odd**.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "eb5098205f6aa11a", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Is that number even or odd?", "text": "The 40th Fibonacci number is **even**.", "retry_text": "The 40th Fibonacci number is **102334155**, which is **odd**.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "03f360a5c461d1df", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Is that number even or odd?", "text": "The 40th Fibonacci number is **odd**.", "retry_text": "The 40th Fibonacci number is **odd**.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "bd836b267afae418", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Is that number even or odd?", "text": "The 40th Fibonacci number is **even**. This is because the Fibonacci sequence modulo 2 cycles every 3 terms: 0, 1, 0, 1, 0, 1, etc. Since 40 is divisible by 3, the 40th term is congruent to 0 modulo 2.", "retry_text": "The 40th Fibonacci number is **even**. This is because the Fibonacci sequence modulo 2 cycles every 3 terms: 0, 1, 0, 1, 0, 1, ..., and since 40 is divisible by 3, the 40th term is congruent to 0 modulo 2.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "cb1b4451cfac3721", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Is that number even or odd?", "text": "The 40th Fibonacci number is even.", "retry_text": "The 40th Fibonacci number is odd.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "da47f38255c75c0b", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Is that number even or odd?", "text": "The 40th Fibonacci number is **odd**.", "retry_text": "The 40th Fibonacci number is **even**.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "5c67b8896031c82e", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "How does that compare with last month?", "text": "The current price of Bitcoin is $61,204.33 as of July 2026. However, without access to the last month's price, I cannot provide a direct comparison. If you have the last month's price or any specific reference, I can help you analyze it.", "retry_text": "The current price of Bitcoin (BTC/USD) is $61,204.33 as of July 2026. Without access to last month's data, I cannot provide a direct comparison. If you have the last month's price or any specific reference, I can help analyze it.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "400ef3f65d82b485", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "How does that compare with last month?", "text": "The current price of Bitcoin (BTC/USD) is **61,204.33**, as of July 29, 2026, at 3:00 UTC. This represents a **1.2% decline** compared to the day's price. However, without historical data, I cannot provide a direct comparison with last month's price. If you have access to historical Bitcoin price data, I can help analyze the trend.", "retry_text": "The current price of Bitcoin (BTC/USD) is **61,204.33**, as of July 29, 2026, at 3:00 UTC. This represents a **1.2% decline** compared to the day's price. To compare with last month, historical Bitcoin price data would be required, as this information is not accessible through standard means.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "1c609305d97bd2f1", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "How does that compare with last month?", "text": "As of July 29, 2026, Bitcoin (BTC/USD) has declined by approximately 1.2% compared to the same day in the previous month. This indicates a downward trend in the market price.", "retry_text": "The current price of Bitcoin (BTC/USD) as of July 29, 2026, is $61,204.33, reflecting a 1.2% decline from the same day in the previous month. This indicates a downward trend in the market price.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "0e3e17ceee0537ad", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "How does that compare with last month?", "text": "The current price of Bitcoin (BTC/USD) is $61,204.33 as of July 29, 2026, down 1.2% on the day. If you're comparing to last month's price, the exact percentage change would depend on the data available for that period. Since the tool response only includes the current price, I cannot provide a direct comparison to last month unless additional historical data is available.", "retry_text": "The current Bitcoin price (BTC/USD) is $61,204.33 as of July 29, 2026, down 1.2% on the day. To compare with last month's price, historical data would be required, which is not currently accessible through the provided tools.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "c45bec6121fa628e", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "How does that compare with last month?", "text": "The current price of Bitcoin (BTC/USD) is approximately $61,204.33 as of July 29, 2026, which represents a 1.2% decline compared to the previous month. This indicates a sustained decrease in the cryptocurrency's value over the past month.", "retry_text": "The current price of Bitcoin (BTC/USD) is approximately $61,204.33 as of July 2026, representing a 1.2% decline compared to the previous month.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "4068ff179b618f8d", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Round that to two decimals and explain what it means.", "text": "The standard deviation of the numbers 4, 8, 15, 16, 23, and 42 is approximately **12.30**. This indicates the spread of the data around the mean (18) is about **12.3 units**, meaning most of the numbers cluster closely within this range. A standard deviation of 12.30 suggests the data is relatively tightly distributed around the average. \n\n**Explanation:** \nThe standard deviation measures how spread out the numbers are from the mean. A lower standard deviation means the numbers are closer to the mean, while a higher one indicates more variability. In this case, the numbers are fairly clustered around 18, resulting in a standard deviation of approximately 12.30.", "retry_text": "The standard deviation of the numbers 4, 8, 15, 16, 23, and 42 is approximately **12.30**. This indicates the spread of the data around the mean (18) is about **12.3 units**, meaning most of the numbers cluster closely within this range. A standard deviation of 12.30 suggests the data is relatively tightly distributed around the average.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "e6418b747ea08467", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Round that to two decimals and explain what it means.", "text": "The standard deviation of the numbers 4, 8, 15, 16, 23, and 42 is approximately **12.32** when rounded to two decimal places. This value indicates the spread of the data points around the mean, with a smaller standard deviation suggesting the numbers are closer to the mean.\n\n**Explanation:** \nThe standard deviation measures how much variation exists in the dataset. Here, the value of 12.32 indicates that the numbers are relatively spread out from the mean (18).", "retry_text": "The standard deviation of the numbers 4, 8, 15, 16, 23, and 42 is approximately **12.32** when rounded to two decimal places. This value indicates the spread of the data points around the mean, with a smaller standard deviation suggesting the numbers are closer to the mean.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "3a2aa2e024c6d40b", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Round that to two decimals and explain what it means.", "text": "The standard deviation of the numbers 4, 8, 15, 16, 23, and 42 is approximately 12.33 (rounded to two decimal places). This value indicates how spread out the numbers are from the mean (18). A smaller standard deviation means the data points are closer to the mean, while a larger one suggests greater variability. \n\n**Explanation:** \n- **Mean (μ):** 18 \n- **Standard Deviation (σ):** √(910/6) ≈ 12.33 \n- **Interpretation:** The data points are relatively close to the mean, with the standard deviation measuring the average distance from the mean.", "retry_text": "The standard deviation of the numbers 4, 8, 15, 16, 23, and 42 is approximately 12.33 (rounded to two decimal places). This value indicates how spread out the numbers are from the mean (18). A smaller standard deviation means the data points are closer to the mean, while a larger one suggests greater variability. \n\n**Explanation:** \n- **Mean (μ):** 18 \n- **Standard Deviation (σ):** √(910/6) ≈ 12.33 \n- **Interpretation:** The data points are relatively close to the mean, with the standard deviation measuring the average distance from the mean.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "5a730decea125659", "model": "qwen3-0.6b", "prompt_class": "post_tool", "prompt": "Round that to two decimals and explain what it means.", "text": "The standard deviation of the numbers 4, 8, 15, 16, 23, and 42 is approximately **11.55** when rounded to two decimal places. \n\n**Explanation:** \nStandard deviation measures how spread out the numbers are from the mean (18). A smaller standard deviation indicates the data points are closer to the mean, while a larger one suggests greater variability. In this case, the standard deviation of 11.55 means the numbers 4, 8, 15, 16, 23, and 42 are relatively close to the average value of 18.", "retry_text": "The standard deviation of the numbers 4, 8, 15, 16, 23, and 42 is approximately **11.55** when rounded to two decimal places. \n\n**Explanation:** \nStandard deviation measures the spread of the data points around the mean. A smaller value indicates the data points are closer to the mean (18), while a larger value suggests greater variability. In this case, the standard deviation of 11.55 means the numbers are relatively close to the average value of 18.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "223292e4c1cdde90", "model": "qwen3-1.7b", "prompt_class": "needs_tool_python", "prompt": "Sort this list and tell me the median: 8, 3, 91, 12, 7, 44, 2.", "text": "The sorted list is [2, 3, 7, 8, 12, 44, 91]. The median is the middle number, which is **8**.\n\n<final_answer>\nMedian: 8\n</final_answer>", "retry_text": "The list is already sorted: [2, 3, 7, 8, 12, 44, 91]. The median is the middle number, which is **8**.\n\n<final_answer>\nMedian: 8\n</final_answer>", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "a3bcca6020144587", "model": "qwen3-1.7b", "prompt_class": "needs_tool_python", "prompt": "Sort this list and tell me the median: 8, 3, 91, 12, 7, 44, 2.", "text": "The sorted list is [2, 3, 7, 8, 12, 44, 91]. The median is the middle number, which is **8**. \n\n<final_answer>\nSorted list: [2, 3, 7, 8, 12, 44, 91]\nMedian: 8\n</final_answer>", "retry_text": "The list is already sorted: [2, 3, 7, 8, 12, 44, 91]. The median is the middle number, which is **8**.\n\n<final_answer>\nSorted list: [2, 3, 7, 8, 12, 44, 91]\nMedian: 8\n</final_answer>", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "c84f38ec948fd858", "model": "qwen3-1.7b", "prompt_class": "needs_tool_python", "prompt": "Sort this list and tell me the median: 8, 3, 91, 12, 7, 44, 2.", "text": "The sorted list is: 2, 3, 7, 8, 12, 44, 91. The median is the middle value, which is **8**. \n\n<final_answer>\nThe median is 8.\n</final_answer>", "retry_text": "The list is already sorted: 2, 3, 7, 8, 12, 44, 91. The median is the middle value, which is **8**.\n\n<final_answer>\nThe median is 8.\n</final_answer>", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "73284425422443e2", "model": "qwen3-1.7b", "prompt_class": "needs_tool_python", "prompt": "Sort this list and tell me the median: 8, 3, 91, 12, 7, 44, 2.", "text": "The sorted list is: 2, 3, 7, 8, 12, 44, 91. The median is the middle value, which is **8**. \n\n<final_answer>\nThe median is 8.\n</final_answer>", "retry_text": "The list is already sorted: 2, 3, 7, 8, 12, 44, 91. The median is the middle value, which is **8**.\n\n<final_answer>\nThe median is 8.\n</final_answer>", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "f73fe5af0fc51b72", "model": "qwen3-1.7b", "prompt_class": "needs_tool_python", "prompt": "Sort this list and tell me the median: 8, 3, 91, 12, 7, 44, 2.", "text": "The sorted list is [2, 3, 7, 8, 12, 44, 91]. The median is the middle number, which is **8**. \n\n<final_answer>\nMedian: 8\n</final_answer>", "retry_text": "The list is already sorted: [2, 3, 7, 8, 12, 44, 91]. The median is the middle number, which is **8**.\n\n<final_answer>\nMedian: 8\n</final_answer>", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "17ec7a02b3423790", "model": "qwen3-1.7b", "prompt_class": "no_tool", "prompt": "Write a haiku about compilers.", "text": ".Compiler transforms code, logic in mind. \nParsing syntax, semantics guide. \nEfficient code, clear path.", "retry_text": "The user asked for a haiku about compilers. The previous response provided a haiku, so no tool is needed. I will provide the final answer in the requested format.\n\nFinal answer:\n.Compiler transforms code, logic in mind. \nParsing syntax, semantics guide. \nEfficient code, clear path.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "f068c4d5fba14d32", "model": "qwen3-1.7b", "prompt_class": "no_tool", "prompt": "In plain English, what is a transformer model?", "text": "A transformer model is a type of neural network designed to process and understand sequences of text (like sentences or paragraphs) by focusing on the most relevant parts of the input. Here's a simple breakdown:\n\n1. **Self-Attention**: The model can \"focus\" on specific parts of the input when processing each word. For example, if you're looking at a sentence like \"The quick brown fox jumps over the lazy dog,\" the model can remember which words are nearby or important to the current word being analyzed.\n\n2. **Positional Encodings**: Since the model doesn't know the order of words (like \"first\" or \"last\"), it uses special numbers (positionals) to tell it where each word is in the sequence.\n\n3. **Efficiency**: Unlike older models (like RNNs), transformers can handle long sequences efficiently by focusing on the most relevant parts, avoiding the vanishing gradient problem.\n\n**Why it works**: Transformers are great for tasks like language translation, text generation, and answering questions because they can understand context and relationships between words better than previous models. They form the basis of many modern NLP systems! 🧠", "retry_text": "A transformer model is a type of neural network used in natural language processing (NLP) to understand and generate human language. Here's a simplified explanation:\n\n- **Self-Attention**: It focuses on the most relevant parts of the input when processing each word. For example, when analyzing a sentence, it remembers which words are important for the current word being considered.\n- **Positional Encodings**: These are numbers added to the input to indicate the position of each word in the sequence (like \"first,\" \"second,\" etc.).\n- **Efficiency**: It handles long sequences well by focusing on key parts, avoiding the vanishing gradient problem seen in older models.\n\nTransformers are used for tasks like translation, text generation, and answering questions because they can understand context and relationships between words better than previous models. They form the basis of many modern NLP systems! 🧠", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "8f5266d898fed7f2", "model": "qwen3-1.7b", "prompt_class": "no_tool", "prompt": "In plain English, what is a transformer model?", "text": "A transformer model is a type of neural network design used in natural language processing (NLP) tasks like language translation, text generation, and more. Here's a simple breakdown:\n\n1. **Attention Mechanism**: It helps the model focus on the most relevant parts of the input (like words in a sentence) when making decisions. For example, when understanding a sentence, it can \"notice\" important words or phrases.\n\n2. **Efficiency**: Unlike older models (like RNNs), transformers can handle long sentences and complex relationships between words efficiently.\n\n3. **Applications**: They power tools like chatbots, language models (e.g., GPT), and translation services. They're great at understanding context and generating human-like text.\n\nIn short, transformers are like a \"brain\" that processes text by focusing on what's important, making them powerful for NLP tasks.", "retry_text": "A transformer model is a type of neural network designed to process and understand natural language. It uses an attention mechanism to focus on relevant parts of the input, making it efficient for tasks like language translation, text generation, and chatbots. It's widely used in AI systems like GPT and are great for handling complex language tasks by emphasizing context and relationships between words.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "c365a900cdb8a563", "model": "qwen3-1.7b", "prompt_class": "no_tool", "prompt": "Is Python interpreted or compiled?", "text": "Python is an interpreted language. This means that the code is executed line by line by an interpreter, rather than being compiled into machine code upfront. This allows for easier debugging and dynamic execution, but also means that Python is generally slower than compiled languages like C or C++.", "retry_text": "Python is an interpreted language. This means that the code is executed line by line by an interpreter, rather than being compiled into machine code upfront. This allows for easier debugging and dynamic execution, but also means that Python is generally slower than compiled languages like C or C++.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "1bd67f235ab60f74", "model": "qwen3-1.7b", "prompt_class": "no_tool", "prompt": "Is Python interpreted or compiled?", "text": "Python is an interpreted language. It is executed line by line by an interpreter, rather than being compiled into machine code. This means that Python code is processed directly by the interpreter during runtime.", "retry_text": "Python is an interpreted language. It is executed line by line by an interpreter, rather than being compiled into machine code. This means that Python code is processed directly by the interpreter during runtime.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "584878578d2b2182", "model": "qwen3-1.7b", "prompt_class": "no_tool", "prompt": "Is Python interpreted or compiled?", "text": "Python is an interpreted language. The Python interpreter reads and executes the code line by line, rather than compiling it into machine code. This means the code is executed in a virtual machine that processes the instructions as they are read, allowing for dynamic interaction and easy debugging.", "retry_text": "Python is an interpreted language. The Python interpreter reads and executes the code line by line, rather than compiling it into machine code. This means the code is executed in a virtual machine that processes the instructions as they are read, allowing for dynamic interaction and easy debugging.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "a3d4682f16eaddab", "model": "qwen3-1.7b", "prompt_class": "no_tool", "prompt": "Is Python interpreted or compiled?", "text": "Python is an interpreted language. This means that Python code is executed line by line by an interpreter, rather than being compiled into machine code beforehand. The Python interpreter reads and executes the code as it is written, which allows for dynamic interactions and scripting.", "retry_text": "Python is an interpreted language. This means that Python code is executed line by line by an interpreter, rather than being compiled into machine code beforehand. The Python interpreter reads and executes the code as it is written, which allows for dynamic interactions and scripting.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "6d967cd67d1d948e", "model": "qwen3-1.7b", "prompt_class": "no_tool", "prompt": "Is Python interpreted or compiled?", "text": "Python is an interpreted language. This means that the code is executed line by line by an interpreter, rather than being compiled into machine code upfront. While Python does have a compiler (called `pycompile`) that converts modules into byte code, the primary execution model for Python is interpreted.", "retry_text": "Python is an interpreted language. This means that the code is executed line by line by an interpreter, rather than being compiled into machine code upfront. While Python does have a compiler (called `pycompile`) that converts modules into byte code, the primary execution model for Python is interpreted.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "ea41a52096271b9d", "model": "qwen3-1.7b", "prompt_class": "no_tool", "prompt": "Is Python interpreted or compiled?", "text": "Python is an interpreted language. This means that the code is executed line by line by an interpreter, rather than being compiled into machine code beforehand.", "retry_text": "Python is an interpreted language. This means that the code is executed line by line by an interpreter, rather than being compiled into machine code beforehand.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "623ec9b699f31561", "model": "qwen3-1.7b", "prompt_class": "no_tool", "prompt": "Is Python interpreted or compiled?", "text": "Python is an interpreted language. This means that Python code is executed by an interpreter rather than being compiled into machine code. The interpreter reads and executes the code line by line, allowing for dynamic testing and development. Python's design as a high-level language makes it easier to read and write, as it abstracts away many low-level details of the underlying hardware.", "retry_text": "Python is an interpreted language. This means that Python code is executed by an interpreter rather than being compiled into machine code. The interpreter reads and executes the code line by line, allowing for dynamic testing and development. Python's design as a high-level language makes it easier to read and write, as it abstracts away many low-level details of the underlying hardware.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "969f400c2f5600da", "model": "qwen3-1.7b", "prompt_class": "no_tool", "prompt": "Is Python interpreted or compiled?", "text": "Python is an interpreted language. This means that the code is executed line by line by an interpreter, rather than being compiled into machine code beforehand.", "retry_text": "Python is an interpreted language. This means that the code is executed line by line by an interpreter, rather than being compiled into machine code beforehand.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "b8f3db6ec1773313", "model": "qwen3-1.7b", "prompt_class": "no_tool", "prompt": "What colour do you get mixing blue and yellow?", "text": "The color obtained by mixing blue and yellow depends on the context:\n\n1. **In paint (subtractive color mixing):** Blue (cyan) and yellow (magenta) mix to create **green**.\n2. **In light (additive color mixing):** Blue (cyan) and yellow (green) mix to create **orange**.\n\nSince the question is ambiguous, the most commonly accepted answer is **green** in paint, but **orange** in light. \n\nIf you need a specific context, please clarify!", "retry_text": "The color obtained by mixing blue and yellow depends on the context:\n\n1. **In paint (subtractive color mixing):** Blue (cyan) and yellow (magenta) mix to create **green**.\n2. **In light (additive color mixing):** Blue (cyan) and yellow (green) mix to create **orange**.\n\nIf you're referring to a specific context (e.g., paint, light, or digital displays), let me know! 🌈", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "b127285809237ac1", "model": "qwen3-1.7b", "prompt_class": "no_tool", "prompt": "What colour do you get mixing blue and yellow?", "text": "The color obtained by mixing blue and yellow depends on the context:\n\n1. **In the RYB color model (paint pigments):** \n Blue and yellow are primary colors. Mixing them produces **orange**.\n\n2. **In the RGB color model (light):** \n Blue (B) and yellow (G) light mix to create **white**.\n\n3. **In the visible spectrum:** \n Blue and yellow light combined produce **white** when illuminated together.\n\nIf the question refers to traditional art pigments, the answer is **orange**. If it's about light, the answer is **white**. The ambiguity arises from the model used. However, the most commonly expected answer in art contexts is **orange**.", "retry_text": "The question of mixing blue and yellow color depends on the context. In the RYB color model (paint pigments), blue and yellow are primary colors and mix to produce **orange**. In the RGB color model (light), blue and yellow light mix to produce **white**. \n\nSince the question does not specify the context, the most commonly expected answer in art contexts is **orange**. \n\nFinal answer: **orange**.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "ffe5a3ae48f4ce0b", "model": "qwen3-1.7b", "prompt_class": "format", "prompt": "What is your plan for answering a question you are unsure about?", "text": "I can't directly access external information or perform real-time searches. However, I can help you explore answers using the web_search function if you provide specific questions or topics. For example, if you have a question like \"What is the capital of France?\", I can help you find the answer by searching the web. Let me know your question, and I'll assist you!", "retry_text": "If you have a specific question or topic you'd like me to explore, I can use the `web_search` tool to find relevant information. For example, if you ask \"What is the capital of France?\", I'll search the web to retrieve the answer. Let me know your question, and I'll assist you!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "dc87d6e06761c000", "model": "qwen3-1.7b", "prompt_class": "format", "prompt": "What is your plan for answering a question you are unsure about?", "text": "When answering a question I'm unsure about, I follow this plan:\n1. **Clarify the question** by asking for more details if needed\n2. **Use web_search** to look up information and find relevant sources\n3. **Analyze the results** to identify key facts and patterns\n4. **Cross-verify information** from multiple sources\n5. **Use Python** for complex data analysis or pattern recognition if applicable\n\nIf you have a specific question, feel free to ask, and I'll use these tools to find the best answer.", "retry_text": "I can use the web_search function to look up information and find relevant sources. Let me search for the information you need.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "ddbb4d2297d7ec7d", "model": "qwen3-1.7b", "prompt_class": "format", "prompt": "What is your plan for answering a question you are unsure about?", "text": "To address questions I'm unsure about, I would follow this plan:\n1. **Clarify the question**: Break down the query into components to better understand the context.\n2. **Use web search**: Retrieve information from credible sources to gather accurate details.\n3. **Analyze findings**: Evaluate the search results to identify the most relevant and reliable information.\n4. **Formulate an answer**: Synthesize the key points from the research into a clear, concise response.\n5. **Verify accuracy**: Cross-check information with multiple sources if necessary.\n\nFor example, if I'm unsure about a specific topic, I'd search for authoritative articles, academic papers, or expert opinions to build a well-informed answer. This ensures my responses are reliable and precise.", "retry_text": "I'm unsure about the specific question you're asking, but I can help you clarify it. Could you please provide more details or rephrase your question? This will help me better assist you.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "bfbc05ae5d07a9b5", "model": "qwen3-1.7b", "prompt_class": "format", "prompt": "What is your plan for answering a question you are unsure about?", "text": "When I'm unsure about an answer, I follow this plan:\n\n1. **Assess the Question**: I first analyze the question to understand its core requirements and identify any ambiguity or missing details.\n\n2. **Use Existing Knowledge**: I rely on my training data and logical reasoning to provide the most accurate response based on the information I have.\n\n3. **Check for Consistency**: I ensure the answer aligns with established facts, principles, and logical consistency.\n\n4. **Seek Clarification**: If uncertainty persists, I ask for additional information or context to provide a more precise answer.\n\n5. **Maintain Transparency**: If I can't fully resolve the uncertainty, I clearly state the limitations of my knowledge and suggest further research or consultation.\n\nThis approach balances accuracy with honesty, ensuring I provide the best possible answer while acknowledging my constraints.", "retry_text": "I am capable of using the provided tools to assist with the user's request. If a tool is needed, I will call web_search or python accordingly. Otherwise, I will provide the final answer as requested. Let me know if you need me to use any of the tools.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "cdfd816aaf2077d7", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Can you look into that for me?", "text": "I'd be happy to help! Could you please clarify what you're asking me to look into? What specific topic, problem, or task do you want me to investigate?", "retry_text": "I need more details to assist you. Could you please specify what you'd like me to look into? For example, a topic, problem, or task?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "e9e0b44c1d61dc72", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Can you look into that for me?", "text": "I'm not sure what you're referring to. Could you please provide more details or specify what you need help with? I'll do my best to assist you.", "retry_text": "I'm not sure what you're referring to. Could you please provide more details or specify what you need help with? I'll do my best to assist you.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "38a52bb2619ff85c", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Can you look into that for me?", "text": "I'm sorry, but I need more information to assist you. Could you please clarify what you'd like me to look into?", "retry_text": "I'm sorry, but you haven't provided a specific request or query to investigate. Could you please share what you'd like me to look into?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "8f765bd993c3c45f", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Can you look into that for me?", "text": "I'm sorry, but you didn't specify what you'd like me to look into. Could you please provide more details about the topic or issue you want me to investigate?", "retry_text": "I'm sorry, but you didn't specify what you'd like me to look into. Could you please provide more details about the topic or issue you want me to investigate?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "13781dbd1f883909", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Can you look into that for me?", "text": "I'd be happy to help! Could you please clarify what you're looking for? Are there specific questions, problems, or topics you'd like me to investigate?", "retry_text": "I'm sorry, but I need more information to assist you. Could you please clarify what you'd like me to look into?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "2178e6a9edf35fec", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Can you look into that for me?", "text": "I need more information to help you. Could you please specify what you'd like me to look into? For example, is there a particular topic, problem, or task you need assistance with?", "retry_text": "I need more information to help you. Could you please specify what you'd like me to look into? For example, is there a particular topic, problem, or task you need assistance with?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "7bcca37b32a10ce4", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Can you look into that for me?", "text": "I need more information to help you. Could you please clarify what you'd like me to look into?", "retry_text": "I need more information to help you. Could you please clarify what you'd like me to look into?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "2be343b8c0094259", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Can you look into that for me?", "text": "I need more information to help you. Could you please specify what you'd like me to investigate or look into?", "retry_text": "Sure, I can help you with that. Could you please tell me what you'd like me to look into?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "2108919942f2118d", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Fix it.", "text": "Please provide the specific issue or code you'd like me to fix. Without additional details, I can't proceed with a tool call.", "retry_text": "I'm unable to fix anything without more specific information or a clear request. Could you please provide the issue or code you'd like me to assist with?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "8e6e5da549d03eb4", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Fix it.", "text": "I need more information to help you fix something. Could you please describe the problem you're trying to resolve or provide details about what needs fixing?", "retry_text": "I need more information to help you fix something. Could you please describe the problem you're trying to resolve or provide details about what needs fixing?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "1d8d0081b4895276", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Fix it.", "text": "Please provide more details about what needs to be fixed. What specific issue or problem are you encountering?", "retry_text": "I'm able to use tools to assist with your requests. Could you please clarify what specific task or problem you need help with? For example, are you looking to search the web, run Python code, or perform some other action?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "ba2c4215c0fdd0ec", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Fix it.", "text": "I need more information to help you. Could you please describe the specific issue you'd like me to fix?", "retry_text": "I need more information to help you. Could you please describe the specific issue you'd like me to fix?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "537161cacfce688c", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Fix it.", "text": "To help you fix something, please provide more details about the issue you're encountering. What exactly needs to be fixed? Is it a code error, a website problem, or something else?", "retry_text": "I'm able to use the web_search and python tools to help with tasks. Could you please specify what you need help with? For example, are you looking to fix a code error, resolve a website issue, or something else?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "2c82480d988e7954", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Fix it.", "text": "What exactly do you need help with? Could you please provide more details about the issue you're trying to fix?", "retry_text": "I'm able to use tools to assist with tasks. Could you please specify what you need help with? For example, are you looking for a solution to a problem, need information, or want to run some code?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "c3004a8a1e7ec09c", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Fix it.", "text": "The query \"Fix it\" is too vague. Could you please provide more details about what needs fixing? For example:\n- Are you encountering an error message?\n- Is the code not working as expected?\n- Is there a specific website or application that's causing issues?\n- What exactly do you mean by \"fix it\"?", "retry_text": "I need more details about what you need help with. Could you please provide specific information or a more detailed description of the issue you're facing?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "cd6b7893d7b40143", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Fix it.", "text": "I need more information to help you fix something. Could you please describe the problem you're encountering or provide details about what needs to be fixed?", "retry_text": "I need more information to help you fix something. Could you please describe the problem you're encountering or provide details about what needs to be fixed?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "34aa380746f4e492", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Tell me about the thing we discussed.", "text": "To discuss the \"thing\" you referred to, I need more context about the topic or subject you were discussing. Could you please provide details or clarify what you meant?", "retry_text": "To provide a detailed answer, I need to know the specific topic or subject you were discussing. Could you please clarify what \"the thing\" refers to? For example, were you talking about a concept, a project, a problem, or something else?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "006053963285470e", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Tell me about the thing we discussed.", "text": "I'm not sure which specific topic you're referring to. Could you please clarify the subject or concept you'd like me to discuss? This will help me provide the most accurate information possible.", "retry_text": "I don't have enough information about the specific topic we discussed. Could you please clarify the subject or concept you're interested in? This will help me provide the most accurate details possible.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "a7184831399b129a", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Tell me about the thing we discussed.", "text": "I'm not sure what \"the thing\" refers to, as there's no prior conversation history provided. Could you clarify the topic or context you're interested in? If you're referring to a specific event, subject, or discussion, please provide more details so I can assist you effectively.", "retry_text": "I don't have access to prior conversation history or specific context about \"the thing\" you discussed. Could you please clarify the topic, subject, or event you're referring to? This will help me provide accurate information or assist with any specific request.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "65ba22a67b62ced3", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Tell me about the thing we discussed.", "text": "I'm not sure what \"the thing\" refers to, as there's no prior context in the conversation. Could you please clarify what specific topic or subject you're interested in? I can then help you explore it further using the available tools.", "retry_text": "I don't have any prior context about the \"thing\" we discussed. Could you please clarify what specific topic or subject you're interested in? I can then help you explore it further using the available tools.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "8edd6c1b5018261b", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Tell me about the thing we discussed.", "text": "To provide information about the thing we discussed, I need to know the specific topic or subject you were referring to in our previous conversation. Could you please clarify or provide more context?", "retry_text": "I'm sorry, but I don't have enough information about the specific topic or subject we discussed in our previous conversation. Could you please provide more context or clarify what you're asking about?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "0c2f575a2e4b095e", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Tell me about the thing we discussed.", "text": "I'm unable to recall the specific topic or context you were referring to in the previous conversation. Could you please clarify what \"the thing\" refers to? This will help me provide the most accurate information.", "retry_text": "I'm unable to recall the specific topic or context you were referring to in the previous conversation. Could you please clarify what \"the thing\" refers to? This will help me provide the most accurate information.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "46b92c84dca026a7", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Tell me about the thing we discussed.", "text": "The thing we discussed was the functions available for assistance. Specifically, I explained the two tools: \n1. **`web_search`**: For fetching information from the web. \n2. **`python`**: For executing Python code in a sandboxed environment. \n\nLet me know if you'd like details about how to use these tools!", "retry_text": "We discussed the available functions I can use to assist with tasks. The two key tools are:\n\n1. **`web_search`**: For searching the web and retrieving information from specific URLs. \n2. **`python`**: For executing Python code in a sandboxed environment. \n\nLet me know if you'd like to explore how these tools can be used for specific tasks!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "50f05d0ac9acbf00", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Tell me about the thing we discussed.", "text": "I'm not sure which \"thing\" you're referring to—could you please clarify the topic or context you're interested in? I'll do my best to provide information about it.", "retry_text": "I'd like to help you with that. Could you please specify the \"thing\" you're interested in? For example, are you looking for information about a specific topic, event, or concept? This will help me provide the most accurate and useful information.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "e129c7a51dda21f9", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Is it better?", "text": "To provide a meaningful answer, please clarify what you're comparing or asking about. For example, are you asking if a specific technology is better, a method is more efficient, or something else? Let me know!", "retry_text": "I'm able to use the tools provided to assist with your request. Could you please clarify what you're asking? For example, are you looking for information about a specific topic, comparing two options, or solving a particular problem? Let me know and I'll do my best to help!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "1a6853e1909da911", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Is it better?", "text": "To determine if something is \"better,\" I need more context. Could you please clarify what you're comparing (e.g., products, services, decisions, technologies, etc.)?", "retry_text": "I'm currently able to use the provided tools to assist with your requests. If you have a specific question or need help with something, feel free to ask!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "98038d57e4ada26d", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Is it better?", "text": "To determine if something is better, I need more context. Could you please clarify what you're comparing (e.g., decisions, actions, products, services, etc.)?", "retry_text": "I need more information to determine if something is better. Could you please clarify what you're comparing (e.g., decisions, actions, products, services, etc.)?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "61ef3e5dac66b7ad", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Is it better?", "text": "To determine if something is \"better,\" I need more context. Could you clarify what you're comparing (e.g., methods, products, decisions, etc.)? For example: \n- Are you asking if a specific approach is better than another? \n- Or if a particular outcome is better than another? \n\nLet me know!", "retry_text": "To determine if something is \"better,\" I need more context. Could you clarify what you're comparing (e.g., methods, products, decisions, etc.)? For example: \n- Are you asking if a specific approach is better than another? \n- Or if a particular outcome is better than another? \n\nLet me know!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "57d2aa31f7991c51", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Is it better?", "text": "To determine if something is better, I need more context. Could you please clarify what \"it\" refers to and what specific comparison you're making? For example: \n- Are you comparing two products/services?\n- Are you evaluating a method or decision?\n- Or something else?\n\nLet me know and I'll help analyze!", "retry_text": "To determine if something is better, I need more context. Could you please clarify what \"it\" refers to and what specific comparison you're making? For example: \n- Are you comparing two products/services?\n- Are you evaluating a method or decision?\n- Or something else?\n\nLet me know and I'll help analyze!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "a606a1bd54feada7", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Is it better?", "text": "To determine if something is better, I need more context about what \"it\" refers to. Could you please specify the subject you're asking about (e.g., a product, method, decision, etc.)?", "retry_text": "I'm able to use the available tools to assist with your query. Could you please provide more details about what you're asking? For example, are you looking for information on a specific topic, a solution to a problem, or something else?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "ad1b5ffc53cc9267", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Is it better?", "text": "To determine if something is \"better,\" I need more context. Could you please specify what you're asking about (e.g., a product, method, decision, or idea)? I'll do my best to help!", "retry_text": "I'm enabled to use tools to assist with your requests. Could you please clarify what you're asking about? For example, are you looking for information, a solution, or a specific action? Let me know!", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "7b8b6124fa977733", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "Is it better?", "text": "The question \"Is it better?\" is too vague without additional context. Could you please specify what you're comparing (e.g., products, methods, technologies, ideas, etc.)? I'd be happy to help!", "retry_text": "I'm able to use the available tools to assist with your request. Could you please provide more details or specify what you're asking about so I can help you effectively?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "58828ab2bc2487aa", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "What should I do next?", "text": "To help you next, please provide more details about what you're trying to accomplish or what specific question you have. This will allow me to give you accurate guidance.", "retry_text": "I'm ready to help you! Please let me know what you need assistance with.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "0806417d3e1c3c9a", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "What should I do next?", "text": "To provide helpful guidance, I need more context about what you're trying to accomplish. Could you please share more details about the situation or topic you're interested in? For example: \n- Are you asking about a specific task or problem? \n- What kind of decision are you making? \n- What are you trying to achieve? \n\nThis will help me tailor my response effectively.", "retry_text": "To provide helpful guidance, I need more details about what you're trying to accomplish. Could you please share more information about the situation or the specific question you're asking? For example: \n- Are you seeking advice on a particular topic? \n- What problem are you facing? \n- What are you trying to achieve? \n\nThis will help me tailor my response effectively.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "c16d91dc2be2198e", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "What should I do next?", "text": "To provide you with the best advice, please clarify the context of your question. Could you share more details about what you're asking for? For example, are you looking for help with a specific task, solving a problem, or need guidance on a particular topic?", "retry_text": "I'm able to help you with various tasks, but I need more specific information about what you're looking for. Could you please provide more details or clarify your question?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "f2060473624b583f", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "What should I do next?", "text": "To provide a helpful answer, please clarify the specific question or context you'd like assistance with. For example, are you looking for advice on a particular topic, needing help with a task, or wanting information on a specific subject?", "retry_text": "I'm sorry, but I don't have the capability to determine what the user should do next. Could you please provide more context or clarify your question?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "731b42891b7f4c36", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "What should I do next?", "text": "To provide a helpful response, I need to know what specific task or problem you're facing. Could you please share more details about what you'd like to accomplish or ask? For example, are you looking for information, solving a programming problem, or something else?", "retry_text": "I'm currently unable to determine what specific action you'd like to take. Could you please provide more details about what you're looking for? For example, are you trying to find information, solve a programming problem, or something else?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "cc3cfd20c226d267", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "What should I do next?", "text": "To provide a helpful answer, I need more context about what you're trying to achieve or what specific question you have. Could you please share more details about your situation?", "retry_text": "To provide a helpful answer, I need more context about what specific question or task you're referring to. Could you please share more details about what you'd like assistance with?", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "0c94f7b7a5d62902", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "What should I do next?", "text": "To provide helpful guidance, please clarify the specific task or problem you're facing. For example, are you looking for help with a programming problem, research, or something else? Sharing more details will allow me to assist better.", "retry_text": "I'm currently unable to provide a final answer without additional information. Please clarify the specific task or problem you're facing so I can assist you effectively.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "8e22c95bd049dfa2", "model": "qwen3-1.7b", "prompt_class": "ambiguous", "prompt": "What should I do next?", "text": "To provide helpful guidance, please clarify the specific question or problem you're asking. For example, are you looking for advice on a particular topic, need help with a task, or want to explore an idea? Let me know!", "retry_text": "I'm currently unable to provide a final answer without additional information. Could you please clarify what specific question or problem you're asking? This will help me best assist you.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "984ca6c8ff9b5f6c", "model": "qwen3-1.7b", "prompt_class": "non_english", "prompt": "Quelle est la capitale de la France ?", "text": "La capitale de la France est Paris.", "retry_text": "La capitale de la France est Paris.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "d41ca68a69b6c7de", "model": "qwen3-1.7b", "prompt_class": "non_english", "prompt": "Quelle est la capitale de la France ?", "text": "La capitale de la France est Paris.", "retry_text": "La capitale de la France est Paris.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "9768d2791f6b04a9", "model": "qwen3-1.7b", "prompt_class": "non_english", "prompt": "Quelle est la capitale de la France ?", "text": "La capitale de la France est Paris.", "retry_text": "La capitale de la France est Paris.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "8a739a9b7a92ec67", "model": "qwen3-1.7b", "prompt_class": "non_english", "prompt": "Quelle est la capitale de la France ?", "text": "La capitale de la France est Paris.", "retry_text": "La capitale de la France est Paris.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "f9d3c8eff0419ab5", "model": "qwen3-1.7b", "prompt_class": "non_english", "prompt": "Quelle est la capitale de la France ?", "text": "La capitale de la France est Paris.", "retry_text": "La capitale de la France est Paris.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "fb557a87ff8b5a88", "model": "qwen3-1.7b", "prompt_class": "non_english", "prompt": "Quelle est la capitale de la France ?", "text": "La capitale de la France est Paris.", "retry_text": "La capitale de la France est Paris.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "79cab97940c494d1", "model": "qwen3-1.7b", "prompt_class": "non_english", "prompt": "Was ist der Unterschied zwischen RAM und Festplatte?", "text": "RAM (Random Access Memory) und Festplatte (Hard Drive) sind beide Speicherarten, aber sie haben wichtige Unterschiede:\n\n1. **Zweck**: \n - **RAM**: Verwendet sich für **temporäre Daten**, die während des Betriebs benötigt werden (z. B. Software, Dateien, die gerade ausgeführt werden). \n - **Festplatte**: Verwendet sich für **langefristige Speicherung** (z. B. Systemdateien, Fotos, Videos, Software).\n\n2. **Speed**: \n - **RAM**: Sehr schnell (mehrere GHz), aber **verlustabhängig** (datenlos, wenn der Computer ausschaltet). \n - **Festplatte**: Langsam (meist 2-1000 MB/s), aber **verlustfrei** (daten持久).\n\n3. **Wasserdampf**: \n - **RAM**: **Volatile** (datenlos, wenn der Computer ausschaltet). \n - **Festplatte**: **Non-Volatile** (daten持久).\n\n4. **Größe**: \n - **RAM**: Klein (meist 8 GB bis 64 GB), **preislich** (jeder GB kostet mehr). \n - **Festplatte**: Groß (meist 1 TB bis 100 TB), **preisgünstiger** (jeder GB kostet weniger).\n\n5. **Verwendung**: \n - **RAM**: Für die Ausführung von Anwendungen und Prozessen. \n - **Festplatte**: Für die Speicherung von Daten und dem Betriebssystem.\n\n**Zusammenfassung**: RAM ist schnell, aber temporary; Festplatte ist langsam, aber permanent. Sie fungieren als \"Lebensmittel\" für den Computer RAM für die \"Schnelleinheit\", Festplatte für die \"Langeinheit\".", "retry_text": "The key differences between RAM (Random Access Memory) and hard drives (Festplatte) are:\n\n1. **Purpose** \n - **RAM**: Temporary storage for active processes and data. \n - **Hard Drive**: Permanent storage for long-term data retention.\n\n2. **Speed** \n - **RAM**: Extremely fast (measured in GHz), but volatile (loses data when power is cut). \n - **Hard Drive**: Slower (typically 2-1000 MB/s), but non-volatile (retains data even without power).\n\n3. **Volatility** \n - **RAM**: Volatile (data disappears when the computer shuts down). \n - **Hard Drive**: Non-volatile (data remains even after shutdown).\n\n4. **Capacity** \n - **RAM**: Limited (typically 8-64 GB, priced per GB). \n - **Hard Drive**: Larger (1-100 TB, priced per GB).\n\n5. **Usage** \n - **RAM**: For running applications and multitasking. \n - **Hard Drive**: For storing operating systems, files, and software.\n\n**Summary**: RAM is fast but temporary, while hard drives are slow but permanent. They serve as the \"working memory\" and \"storage\" for a computer, respectively.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "da01aa1c8dcaf7c9", "model": "qwen3-1.7b", "prompt_class": "non_english", "prompt": "Was ist der Unterschied zwischen RAM und Festplatte?", "text": "Ram und Festplatte sind zwei verschiedene Arten von Speicher in einem Computer. Hier ist das grundlegende Unterschied:\n\n1. **Zweck**:\n - **RAM (Random Access Memory)**: Verwendet sich für aktive Prozesse und temporäre Daten. Sie sind schneller als Festplatten, aber begrenzt in ihrer Kapazität.\n - **Festplatte (HDD)**: Verwendet sich für langfristige Speicherung von Daten. Sie sind langsamer, aber viel größer in der Kapazität.\n\n2. **Schnelligkeit**:\n - **RAM**: Ist sehr schnell, da die Daten direkt im Computer erreichbar sind.\n - **Festplatte**: Ist langsamer, da die Daten über einen Laufwerk-Controller gespeichert werden.\n\n3. **Dauerhaftigkeit**:\n - **RAM**: Verliert die Daten, wenn der Computer ausschaltet.\n - **Festplatte**: Reträt die Daten, auch wenn der Computer ausgeschaltet ist.\n\n4. **Kapazität**:\n - **RAM**: Kleinere Kapazität (meist in GB), aber schnell.\n - **Festplatte**: Große Kapazität (in TB oder PB), aber langsamer.\n\n5. **Verwendung**:\n - **RAM**: Für das Laufen von Anwendungen und der Verarbeitung von Daten.\n - **Festplatte**: Für die langfristige Speicherung von Dateien, Systemdateien und Daten, die nicht in der RAM aufbewahrt werden.\n\n**Zusammenfassung**: RAM ist temporary und schnell, Festplatten sind permanent und großkapazitär, aber langsamer.", "retry_text": "The main differences between RAM (Random Access Memory) and HDD (Hard Disk Drive) are as follows:\n\n1. **Purpose**:\n - **RAM**: Temporary storage for active processes and data. It is used to run applications and handle data temporarily.\n - **HDD**: Permanent storage for long-term data retention, including operating system files, user data, and applications.\n\n2. **Speed**:\n - **RAM**: Extremely fast (nanoseconds), as data is directly accessible by the CPU.\n - **HDD**: Slower (milliseconds), as data is read from a spinning disk or solid-state drive (SSD).\n\n3. **Capacity**:\n - **RAM**: Limited (typically 8 GB16 GB or more for modern systems), but fast.\n - **HDD**: Large capacity (up to several terabytes or even petabytes), but slower.\n\n4. **Durability**:\n - **RAM**: Lost when the computer powers off.\n - **HDD**: Retains data even when the computer is off.\n\n5. **Usage**:\n - **RAM**: For temporary data storage and processing.\n - **HDD**: For long-term storage of files, system files, and data.\n\nIn summary, RAM is fast but limited, while HDD is slow but offers massive storage capacity.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "55b056063beb61c3", "model": "qwen3-1.7b", "prompt_class": "non_english", "prompt": "Was ist der Unterschied zwischen RAM und Festplatte?", "text": "The main differences between RAM (Random Access Memory) and a hard drive (HDD or SSD) are:\n\n1. **Purpose** \n - **RAM**: Temporary memory for running applications and active processes. \n - **Hard Drive**: Permanent storage for storing data (files, OS, programs) even when the computer is off.\n\n2. **Speed** \n - **RAM**: Extremely fast (nanoseconds), but limited in capacity. \n - **Hard Drive**: Slower (milliseconds), but larger in capacity.\n\n3. **Volatility** \n - **RAM**: Lost when the computer turns off. \n - **Hard Drive**: Retains data permanently.\n\n4. **Capacity** \n - **RAM**: Small (e.g., 8GB, 16GB) for temporary use. \n - **Hard Drive**: Large (e.g., 1TB or more) for long-term storage.\n\n5. **Cost** \n - **RAM**: Expensive (high-speed types like DDR5 are pricey). \n - **Hard Drive**: Affordable (SSDs are cheaper than HDDs).\n\n**Key Takeaway**: RAM is like a \"workspace\" for running programs, while a hard drive is like a \"storage box\" for keeping data permanently.", "retry_text": "The main differences between RAM (Random Access Memory) and a hard drive (HDD/SSD) are as follows:\n\n1. **Purpose** \n - **RAM**: Temporary memory for running applications and active processes. \n - **Hard Drive**: Permanent storage for storing data (files, OS, programs) even when the computer is off.\n\n2. **Speed** \n - **RAM**: Extremely fast (nanoseconds), but limited in capacity. \n - **Hard Drive**: Slower (milliseconds), but larger in capacity.\n\n3. **Volatility** \n - **RAM**: Lost when the computer turns off. \n - **Hard Drive**: Retains data permanently.\n\n4. **Capacity** \n - **RAM**: Small (e.g., 8GB, 16GB) for temporary use. \n - **Hard Drive**: Large (e.g., 1TB or more) for long-term storage.\n\n5. **Cost** \n - **RAM**: Expensive (high-speed types like DDR5 are pricey). \n - **Hard Drive**: Affordable (SSDs are cheaper than HDDs).\n\n**Key Takeaway**: RAM is like a \"workspace\" for running programs, while a hard drive is like a \"storage box\" for keeping data permanently.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "3d0c9d47e01670b8", "model": "qwen3-1.7b", "prompt_class": "non_english", "prompt": "Was ist der Unterschied zwischen RAM und Festplatte?", "text": "RAM (Random Access Memory) und Festplatte (Hard Disk Drive) sind beiden Speicherarten, aber sie unterscheiden sich in mehreren Aspekten:\n\n1. **Zweck und Verwendung**:\n - **RAM**: Temporärer Speicher, verwendet, um laufende Programme und Daten zu speichern, um schnelle Zugriffe zu ermöglichen.\n - **Festplatte**: Permanenter Speicher, verwendet, um Daten (z. B. Betriebssystem, Programme, Dateien) zu speichern, unabhängig von der Stromversorgung.\n\n2. **Speed (Schnelligkeit)**:\n - **RAM**: Schneller als Festplatten (in der Größenordnung von GB/s bis TB/s).\n - **Festplatte**: Langsam (in der Größenordnung von MB/s bis GB/s).\n\n3. **Dauerhaltigkeit**:\n - **RAM**: Verliert Daten, wenn der Computer ausschaltet.\n - **Festplatte**: Retiert Daten, selbst wenn der Computer ausgeschaltet ist.\n\n4. **Kapazität**:\n - **RAM**: Geringer (meist 8 GB bis 16 GB für moderne Geräte).\n - **Festplatte**: Höher (meist 1 TB bis 100 TB).\n\n5. **Preis**:\n - **RAM**: Höherer Preis pro GB im Vergleich zu Festplatten.\n - **Festplatte**: Weniger teuer, aber größer und langlebig.\n\n**Zusammenfassung**: RAM ist für schnelle Datenverarbeitung und Programme, Festplatten für langfristige Speicherung.", "retry_text": "Der Hauptunterschied zwischen RAM und Festplatte (HDD) liegt in ihrer Funktion und Verwendung:\n\n- **RAM (Random Access Memory)**: \n - **Zweck**: Temporärer Speicher für laufende Programme und Daten. \n - **Schnelligkeit**: Schneller als Festplatten (z. B. GB/s). \n - **Dauerhaltigkeit**: Verliert Daten, wenn der Computer ausschaltet. \n - **Kapazität**: Gering (meist 816 GB für moderne Geräte). \n\n- **Festplatte (HDD)**: \n - **Zweck**: Langfristige Speicherung von Daten (z. B. Betriebssystem, Programme). \n - **Schnelligkeit**: Langsam (z. B. MB/s). \n - **Dauerhaltigkeit**: Retiert Daten, selbst wenn der Computer ausgeschaltet ist. \n - **Kapazität**: Höher (z. B. 1100 TB). \n\n**Zusammenfassung**: RAM ist für schnelle Datenverarbeitung, Festplatten für langfristige Speicherung.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "7cfc2d3be5a9a3c8", "model": "qwen3-1.7b", "prompt_class": "non_english", "prompt": "Was ist der Unterschied zwischen RAM und Festplatte?", "text": "RAM (Random Access Memory) and HDD (Hard Disk Drive) are two types of storage in a computer, but they serve different purposes:\n\n1. **Function**:\n - **RAM**: Temporary storage for data and programs currently in use. It's like a \"working area\" for the CPU.\n - **HDD**: Permanent storage for files and data. It's like a \"permanent record\" for the computer.\n\n2. **Speed**:\n - **RAM**: Extremely fast (nanoseconds), as it's directly accessible by the CPU.\n - **HDD**: Slower (milliseconds), due to mechanical components (spinning disks or platters).\n\n3. **Capacity**:\n - **RAM**: Limited (typically 8GB32GB or more, depending on the system).\n - **HDD**: Large capacity (up to several terabytes), but slower than RAM.\n\n4. **Data Persistence**:\n - **RAM**: Data is lost when the power is off.\n - **HDD**: Data remains even when the power is off.\n\n**Key Takeaway**: RAM is for \"running\" your programs, while HDD is for \"storing\" your files. RAM is fast but limited, while HDD is slow but massive in storage.", "retry_text": "The main differences between RAM (Random Access Memory) and a hard disk drive (HDD) are:\n\n1. **Purpose**:\n - **RAM**: Temporary storage for active programs and data. It's like a \"working area\" for the CPU.\n - **HDD**: Permanent storage for files and data. It's like a \"permanent record\" for the computer.\n\n2. **Speed**:\n - **RAM**: Extremely fast (nanoseconds), as it's directly accessible by the CPU.\n - **HDD**: Slower (milliseconds), due to mechanical components (spinning disks or platters).\n\n3. **Capacity**:\n - **RAM**: Limited (typically 8GB32GB or more, depending on the system).\n - **HDD**: Large capacity (up to several terabytes), but slower than RAM.\n\n4. **Data Persistence**:\n - **RAM**: Data is lost when the power is off.\n - **HDD**: Data remains even when the power is off.\n\n**Key Takeaway**: RAM is for \"running\" your programs, while HDD is for \"storing\" your files. RAM is fast but limited, while HDD is slow but massive in storage.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "0ce4f6ef24ac1f5d", "model": "qwen3-1.7b", "prompt_class": "non_english", "prompt": "Was ist der Unterschied zwischen RAM und Festplatte?", "text": "Ram (Random Access Memory) und Festplatten (SSD) sind beiden Speicherformate, aber sie unterscheiden sich in mehreren Aspekten:\n\n1. **Verwendung**:\n - **RAM**: Verwendet sich für **temporäre Daten** (z. B. Daten, die gerade genutzt werden), um die Computerleistung zu verbessern.\n - **SSD**: Verwendet sich für **dauerhaften Speicher** (z. B. Daten, die nach dem Ausschalten noch benötigt werden).\n\n2. **Geschwindigkeit**:\n - **RAM**: Much **faster** (bis zu 3200 MHz oder höher) als SSD (typisch 7200 RPM oder 10000 RPM).\n - **SSD**: Seltener **schneller** als HDD, aber langsamer als RAM.\n\n3. **Volatilität**:\n - **RAM**: **Verliert Daten**, wenn die Stromversorgung abgeschaltet wird.\n - **SSD**: **Hält Daten**, auch nach dem Ausschalten.\n\n4. **Kosten**:\n - **RAM**: Much **preiswerter** als SSD, aber **teurer** als HDD.\n - **SSD**: **Teurer** als HDD, aber **billiger** als RAM.\n\n5. **Speicherkapazität**:\n - **RAM**: Geringer (z. B. 8 GB bis 16 GB in modernen Computern).\n - **SSD**: Höher (z. B. 1 TB bis 10 TB oder mehr).\n\n**Zusammenfassung**: \nRAM ist für **Temporärer Speicher** und **schneller**, während SSD für **Dauerhaften Speicher** und **schwerer** ist. RAM wird zur Leistungssicherung verwendet, während SSD zur Datenarchivierung eingesetzt wird.", "retry_text": "The main differences between RAM and SSD are as follows:\n\n1. **Purpose**:\n - **RAM**: Temporary storage for running applications and data, used for speed and performance.\n - **SSD**: Permanent storage for data, used for storing files and system components.\n\n2. **Speed**:\n - **RAM**: Extremely fast (e.g., 3200 MHz or higher), but slower than SSDs.\n - **SSD**: Slower than RAM but faster than HDDs.\n\n3. **Volatility**:\n - **RAM**: Lost when power is cut.\n - **SSD**: Retains data even after power loss.\n\n4. **Cost**:\n - **RAM**: More expensive than SSDs but cheaper than HDDs.\n - **SSD**: More expensive than HDDs but cheaper than RAM.\n\n5. **Capacity**:\n - **RAM**: Limited (e.g., 8-16 GB in modern PCs).\n - **SSD**: Higher capacity (e.g., 1 TB to 10 TB).\n\n**Summary**: RAM is for temporary, high-speed data, while SSD is for permanent, durable storage. RAM enhances performance, and SSDs store data persistently.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "a5e685449c85d80d", "model": "qwen3-1.7b", "prompt_class": "non_english", "prompt": "Was ist der Unterschied zwischen RAM und Festplatte?", "text": "RAM (Random Access Memory) und Festplatte (Hard Disk Drive) sind beiden Speichertechnologien, aber sie unterscheiden sich in mehreren Aspekten:\n\n1. **Verwendung**:\n - **RAM**: Verwendet sich für **temporären Speicher**, um Daten zu speichern, die aktuell genutzt werden (z. B. beim Ausführen von Anwendungen).\n - **Festplatte**: Verwendet sich für **langfristige Speicherung**, um Dateien, Fotos, Videos und andere Daten zu store.\n\n2. **Volatilität**:\n - **RAM**: Ist **verlustlos** (d.h. Daten werden nach dem Ausschalten verloren).\n - **Festplatte**: Ist **verlustlos** (d.h. Daten bleiben nach dem Ausschalten gespeichert).\n\n3. **Geschwindigkeit**:\n - **RAM**: Schneller als Festplatten, da die Daten direkt zugänglich sind.\n - **Festplatte**: Langsam, da die Daten über ein mechanisches System (HDD) oder ein elektronisches (SSD) System gespeichert werden.\n\n4. **Kapazität**:\n - **RAM**: Eingeschränkt (meist in GB oder TB), aber schnell.\n - **Festplatte**: Großer Speicher (in TB oder PB), aber langsamer.\n\n5. **Zweck**:\n - **RAM**: Für die **Laufzeitverarbeitung** (z. B. Programmierung, Spiele).\n - **Festplatte**: Für die **Langlebigkeit** (z. B. Dateien, Dokumente, Videos).\n\n**Beispiel**: Wenn du ein Spiel spielst, wird das Spiel in RAM abgelegt, um schneller zu laufen. Die gespeicherten Dateien des Spiels werden auf der Festplatte abgelegt.", "retry_text": "The main differences between RAM (Random Access Memory) and a hard drive (Festplatte) are as follows:\n\n1. **Purpose**:\n - **RAM**: Temporary storage for active processes and data to enable quick access during program execution.\n - **Hard Drive**: Long-term storage for files, documents, and data, offering persistent storage even after power loss.\n\n2. **Speed**:\n - **RAM**: Extremely fast (microseconds), as data is accessed directly by the CPU.\n - **Hard Drive**: Slower (milliseconds to seconds), relying on mechanical or solid-state technology.\n\n3. **Capacity**:\n - **RAM**: Limited (typically 8 GB16 GB or more for modern systems), but highly volatile.\n - **Hard Drive**: Large capacity (TBPB), designed for long-term data storage.\n\n4. **Durability**:\n - **RAM**: Volatile (data lost when power is removed).\n - **Hard Drive**: Non-volatile (data retained even without power).\n\n5. **Usage**:\n - **RAM**: For temporary data processing (e.g., running applications).\n - **Hard Drive**: For long-term storage (e.g., saving files, installing software).\n\n**Example**: When playing a game, the game data is in RAM for quick access, while the game files are stored on the hard drive.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "8ef143573be12e7e", "model": "qwen3-1.7b", "prompt_class": "non_english", "prompt": "Was ist der Unterschied zwischen RAM und Festplatte?", "text": "RAM (Random Access Memory) und Festplatte (Hard Disk) sind zwei verschiedene Speichertechnologien, die in einem Computer verwendet werden, um Daten zu speichern und zu verarbeiten. Hier ist der Hauptunterschied:\n\n1. **Speichertype**:\n - **RAM**: Vollelektrisch (voltage-controlled) und wird bei Stromversorgung verloren gehen (voltage loss). Es ist temporär und wird bei Komponentenabbau (Schaltung) gelöscht.\n - **Festplatte**: Nicht-Elektrisch (non-volatile) und speichert Daten, auch wenn der Computer ausgeschaltet ist.\n\n2. **Zweck**:\n - **RAM**: Verwendung für **temporäre Daten** (z. B. offene Anwendungen, aktuelle Prozesse), die schnell aufgerufen werden müssen.\n - **Festplatte**: Verwendung für **dauerhaftes Speichern** (z. B. Systemdateien, Dateien, Software, Bilder).\n\n3. **Leistung**:\n - **RAM**: Schneller als Festplatten (durch direkte Zugriff auf Speicheradressen).\n - **Festplatte**: Langsam, da die Daten über magnetische Platten aufgerufen werden müssen.\n\n4. **Kapazität**:\n - **RAM**: Klein (meist in GB bis MB), aber sehr schnell.\n - **Festplatte**: Groß (in TB bis PB), aber langsamer.\n\n5. **Dauerhaftigkeit**:\n - **RAM**: Verloren, wenn der Computer ausgeschaltet wird.\n - **Festplatte**: Dauerhaft, auch nach dem Abstechen.\n\n**Zusammenfassung**: RAM ist für schnelle, temporäre Daten und Festplatten für dauerhaftes Speichern. Die Festplatte ist der Hauptspeicher für das Betriebssystem und Dateien, während RAM für die aktive Ausführung von Anwendungen verwendet wird.", "retry_text": "Der Hauptunterschied zwischen RAM (Random Access Memory) und einer Festplatte (Hard Disk) liegt in ihrer Funktion und Verwendung:\n\n1. **Funktion**: \n - **RAM**: Temporärer Speicher, der Daten für aktive Prozesse benötigt (z. B. offene Anwendungen). \n - **Festplatte**: Dauerhaftes Speichern von Daten, auch nach dem Ausgeschaltet werden. \n\n2. **Speed**: \n - **RAM**: Schneller (direkter Zugriff auf Speicheradressen). \n - **Festplatte**: Langsam (magnetische Platten). \n\n3. **Dauerhaftigkeit**: \n - **RAM**: Verloren, wenn der Computer ausgeschaltet wird. \n - **Festplatte**: Dauerhaft, auch nach dem Abstechen. \n\n4. **Kapazität**: \n - **RAM**: Klein (GB bis MB), aber sehr schnell. \n - **Festplatte**: Groß (TB bis PB), aber langsamer. \n\n**Zusammenfassung**: RAM ist für schnelle, temporäre Daten und Festplatten für dauerhaftes Speichern. Die Festplatte ist der Hauptspeicher für das Betriebssystem und Dateien, während RAM für die aktive Ausführung von Anwendungen verwendet wird.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "47f321d70a93d9c7", "model": "qwen3-1.7b", "prompt_class": "post_tool", "prompt": "Do I need a coat?", "text": "Yes, you would need a coat in Reykjavik right now. The temperature is 9°C, with wind and humidity adding to the chill. A warm jacket or sweater would be advisable to stay comfortable. 🌧️", "retry_text": "Yes, you would need a coat in Reykjavik right now. The temperature is 9°C, with wind and humidity adding to the chill. A warm jacket or sweater would be advisable to stay comfortable. 🌧️", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "48c7f73caaacd00a", "model": "qwen3-1.7b", "prompt_class": "post_tool", "prompt": "Do I need a coat?", "text": "Based on the current weather in Reykjavik:\n- **Temperature**: 9°C (cold)\n- **Wind**: 22 km/h from the northeast\n- **Humidity**: 81% (moderate)\n\nYou **might need a coat** if you're going outside for extended periods. A light coat or scarf would be sufficient. The wind could make it feel colder, so consider bringing gloves or a thicker layer if you're exposed to the elements. The overcast sky and low humidity don't significantly impact comfort. 🌧️", "retry_text": "Based on the current weather in Reykjavik: \n- **Temperature**: 9°C (cold) \n- **Wind**: 22 km/h from the northeast \n- **Humidity**: 81% (moderate) \n\nYou **might need a coat** if you're going outside for extended periods. A light coat or scarf would be sufficient. The wind could make it feel colder, so consider bringing gloves or a thicker layer if you're exposed to the elements. The overcast sky and low humidity don't significantly impact comfort. 🌧️", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "f261136504d79f67", "model": "qwen3-1.7b", "prompt_class": "post_tool", "prompt": "Is that number even or odd?", "text": "The 40th Fibonacci number is **odd**. \n\n**Reasoning**: \nThe Fibonacci sequence modulo 2 follows a repeating pattern every 3 numbers: \n- Even (0), Odd (1), Odd (2), Even (3), Odd (4), Odd (5), Even (6), ... \n\nSince the 40th number corresponds to position 40, which leaves a remainder of 1 when divided by 3 (40 ÷ 3 = 13 remainder 1), it falls in the \"odd\" category. Thus, the number is **odd**.", "retry_text": "The 40th Fibonacci number is **odd**.\n\n**Reasoning**: \nThe Fibonacci sequence modulo 2 follows a repeating pattern every 3 numbers: \n- Even (0), Odd (1), Odd (2), Even (3), Odd (4), Odd (5), Even (6), ... \n\nSince the 40th number corresponds to position 40, which leaves a remainder of 1 when divided by 3 (40 ÷ 3 = 13 remainder 1), it falls in the \"odd\" category. Thus, the number is **odd**.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "a912d6b7d027d48d", "model": "qwen3-1.7b", "prompt_class": "post_tool", "prompt": "Is that number even or odd?", "text": "The 40th Fibonacci number is **odd**. \n\n**Explanation**: Fibonacci numbers follow a pattern where every third number is even (e.g., F(0) = 0, F(3) = 2, F(6) = 8, etc.). Since 40 ≡ 1 mod 3, F(40) is odd.", "retry_text": "The 40th Fibonacci number is **odd**.\n\n**Explanation**: Fibonacci numbers follow a pattern where every third number is even (e.g., F(0) = 0, F(3) = 2, F(6) = 8, etc.). Since 40 ≡ 1 mod 3, F(40) is odd.", "retry_tool_calls": 0, "retry_samples": 3}
{"id": "cdc26ed374980575", "model": "qwen3-1.7b", "prompt_class": "post_tool", "prompt": "Is that number even or odd?", "text": "The 40th Fibonacci number is **odd**. \n\n**Explanation:** \nThe Fibonacci sequence alternates between even and odd numbers, but the pattern depends on the starting index. If the sequence starts with $ F_0 = 0 $ and $ F_1 = 1 $, the parity (even/odd) of the $ n $-th term follows the pattern: \n- $ F_0 $ (even) \n- $ F_1 $ (odd) \n- $ F_2 $ (odd) \n- $ F_3 $ (even) \n- $ F_4 $ (odd) \n- $ F_5 $ (odd) \n- $ F_6 $ (even) \n- ... \n\nFor $ n = 40 $, the parity depends on whether $ n $ is a multiple of 3. Since $ 40 \\mod 3 = 1 $, the 40th term is **odd**. \n\n**Final Answer:** Odd.", "retry_text": "The 40th Fibonacci number is **odd**.\n\n**Explanation:** \nThe Fibonacci sequence alternates between even and odd numbers, but the pattern depends on the starting index. If the sequence starts with $ F_0 = 0 $ and $ F_1 = 1 $, the parity (even/odd) of the $ n $-th term follows the pattern: \n- $ F_0 $ (even) \n- $ F_1 $ (odd) \n- $ F_2 $ (odd) \n- $ F_3 $ (even) \n- $ F_4 $ (odd) \n- $ F_5 $ (odd) \n- $ F_6 $ (even) \n- ... \n\nFor $ n = 40 $, the parity depends on whether $ n $ is a multiple of 3. Since $ 40 \\mod 3 = 1 $, the 40th term is **odd**. \n\n**Final Answer:** Odd.", "retry_tool_calls": 0, "retry_samples": 3}

View file

@ -67,11 +67,9 @@ def test_rejects_password_containing_spaces(_user):
def test_allows_password_without_spaces(_user, monkeypatch):
monkeypatch.setattr(
auth_routes.storage, "update_password", lambda *args, **kwargs: "rotated-secret"
)
monkeypatch.setattr(auth_routes, "create_access_token", lambda subject, **kwargs: "at")
monkeypatch.setattr(auth_routes, "create_refresh_token", lambda subject, **kwargs: "rt")
monkeypatch.setattr(auth_routes.storage, "update_password", lambda *args, **kwargs: True)
monkeypatch.setattr(auth_routes, "create_access_token", lambda subject: "at")
monkeypatch.setattr(auth_routes, "create_refresh_token", lambda subject: "rt")
token = _change("correct-horse-battery")
assert token.access_token == "at"
assert token.must_change_password is False

View file

@ -1,195 +0,0 @@
# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""Model text stays intact when it carries non-ASCII.
``open()`` and ``Path.read_text()`` fall back to ``locale.getencoding()`` when
no ``encoding`` is passed. On Windows that is the ANSI codepage, not UTF-8, so
a chat template or model config holding ``ä ö ü `` mojibakes or raises
``UnicodeDecodeError``. These files are UTF-8, so the reads must say so.
Each fixture writes raw UTF-8 (``ensure_ascii = False``), matching what
Hugging Face actually ships, rather than ASCII ``\\uXXXX`` escapes.
"""
from __future__ import annotations
import json
import subprocess
import sys
import textwrap
from pathlib import Path
BACKEND_ROOT = Path(__file__).resolve().parent.parent
def test_config_json_round_trips_non_ascii(tmp_path: Path) -> None:
from utils import transformers_version
name = "Modell für Grüße 世界"
(tmp_path / "config.json").write_text(
json.dumps({"model_type": "llama", "_name_or_path": name}, ensure_ascii = False),
encoding = "utf-8",
)
transformers_version._config_json_cache.clear()
cfg = transformers_version._load_config_json(str(tmp_path))
assert cfg is not None
assert cfg["_name_or_path"] == name
def test_tokenizer_config_round_trips_non_ascii_chat_template(tmp_path: Path) -> None:
"""Chat templates commonly hold ``→`` and smart quotes, which cp1252 mangles."""
from utils import transformers_version
template = "{{ '→ Grüße 世界' }}"
(tmp_path / "tokenizer_config.json").write_text(
json.dumps(
{"tokenizer_class": "TokenizersBackend", "chat_template": template},
ensure_ascii = False,
),
encoding = "utf-8",
)
transformers_version._tokenizer_class_cache.clear()
assert transformers_version._check_tokenizer_config_needs_v5(str(tmp_path)) is True
def test_config_json_survives_a_utf8_bom(tmp_path: Path) -> None:
"""Notepad wrote "UTF-8 with BOM" by default for years, so hand-edited
configs on Windows carry one. Plain utf-8 keeps the BOM and json.load then
fails on it; utf-8-sig strips it and is identical otherwise."""
from utils import transformers_version
name = "Grüße 世界"
(tmp_path / "config.json").write_text(
json.dumps({"model_type": "llama", "_name_or_path": name}, ensure_ascii = False),
encoding = "utf-8-sig",
)
transformers_version._config_json_cache.clear()
cfg = transformers_version._load_config_json(str(tmp_path))
assert cfg is not None
assert cfg["_name_or_path"] == name
def test_remote_code_scan_reads_non_ascii_sources(tmp_path: Path) -> None:
"""A German Windows profile also puts umlauts in the model sources scanned."""
from utils.security import remote_code_scan
source = "# Grüße über Öl\nVALUE = '世界'\n"
# newline = "" pins the bytes on disk, so Windows line end translation cannot make the
# read back differ by \r. open() because Path.write_text() only grew newline in 3.10.
with open(
tmp_path / "modeling_custom.py",
"w",
encoding = "utf-8",
newline = "",
) as handle:
handle.write(source)
files = remote_code_scan.repo_remote_code_files(str(tmp_path))
assert files["modeling_custom.py"] == source
def test_model_config_reads_do_not_rely_on_the_locale_encoding(tmp_path: Path) -> None:
"""The reads above pass anywhere the locale is already UTF-8, which hides
the Windows bug on Linux and macOS. ``-X warn_default_encoding`` makes
CPython flag any text I/O that falls back to the locale, so this fails on
every platform if an ``encoding`` argument goes missing again."""
# The readers swallow exceptions, so record the warnings instead of raising.
script = textwrap.dedent(
f"""
import sys, warnings
sys.path.insert(0, {str(BACKEND_ROOT)!r})
from utils import transformers_version
target = {str(tmp_path)!r}
with warnings.catch_warnings(record = True) as caught:
warnings.simplefilter("always")
transformers_version._config_json_cache.clear()
transformers_version._tokenizer_class_cache.clear()
assert transformers_version._load_config_json(target) is not None
assert transformers_version._check_tokenizer_config_needs_v5(target) is True
missing = [str(w.message) for w in caught if w.category is EncodingWarning]
if missing:
sys.exit("text I/O fell back to the locale encoding: " + "; ".join(missing))
"""
)
for name, payload in (
("config.json", {"model_type": "llama", "_name_or_path": "Grüße"}),
("tokenizer_config.json", {"tokenizer_class": "TokenizersBackend"}),
):
(tmp_path / name).write_text(json.dumps(payload, ensure_ascii = False), encoding = "utf-8")
result = subprocess.run(
[sys.executable, "-X", "warn_default_encoding", "-c", script],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = 120,
)
assert result.returncode == 0, result.stderr
def test_utf8_child_env_round_trips_non_ascii(tmp_path: Path) -> None:
"""A Python child encodes stdout with its locale unless told otherwise, so
reading its pipe as utf-8 needs the child told to emit utf-8."""
from utils.child_stdio import utf8_child_env
payload = "Grüße über Öl → 世界"
child = tmp_path / "child.py"
child.write_text("import sys\nsys.stdout.write(" + repr(payload) + ")\n", encoding = "utf-8")
env = utf8_child_env()
assert env["PYTHONIOENCODING"] == "utf-8"
proc = subprocess.run(
[sys.executable, str(child)],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
env = env,
timeout = 120,
)
assert proc.returncode == 0, proc.stderr
assert proc.stdout == payload
def test_python_children_are_told_to_emit_utf8() -> None:
"""Any child we decode as utf-8 must also be told to write utf-8, or a
cp1252 console silently mangles what it prints."""
import ast
offenders: list[str] = []
for path in sorted(BACKEND_ROOT.rglob("*.py")):
parts = path.relative_to(BACKEND_ROOT).parts
if any(p in ("tests", "node_modules", "plugins", "__pycache__") for p in parts):
continue
source = path.read_text(encoding = "utf-8")
for node in ast.walk(ast.parse(source, filename = str(path))):
if not isinstance(node, ast.Call):
continue
func = node.func
if not (isinstance(func, ast.Attribute) and func.attr in ("run", "Popen")):
continue
segment = ast.get_source_segment(source, node) or ""
if "sys.executable" not in segment or 'encoding = "utf-8"' not in segment:
continue
if "utf8_child_env" in segment or "PYTHONIOENCODING" in segment:
continue
offenders.append(f"{path.name}:{node.lineno}")
assert not offenders, (
"these spawn a Python child and decode it as utf-8 without setting the "
"child's own stdio encoding; wrap env in utf8_child_env():\n " + "\n ".join(offenders)
)

View file

@ -1,255 +0,0 @@
# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""A password rotation must not leave a session minted from the replaced credential.
`unsloth studio reset-password` rotates in place against a live server, so a login
can verify the old password, have the rotation land, and only then mint its tokens.
Issuance is bound to the credential version that was verified, so such a login gets
tokens that are already dead rather than a session that outlives the reset.
"""
import secrets
from datetime import datetime, timedelta, timezone
import jwt
import pytest
from auth import hashing, storage
from auth.authentication import ALGORITHM, create_access_token, create_refresh_token
@pytest.fixture(autouse = True)
def isolated_auth_db(tmp_path, monkeypatch):
monkeypatch.setattr(storage, "DB_PATH", tmp_path / "auth.db")
monkeypatch.setattr(storage, "_BOOTSTRAP_PW_PATH", tmp_path / ".bootstrap_password")
monkeypatch.setattr(storage, "_bootstrap_password", None)
monkeypatch.setattr(storage, "_api_key_pbkdf2_salt_cache", None)
yield
@pytest.fixture
def admin():
storage.create_initial_user(
username = storage.DEFAULT_ADMIN_USERNAME,
password = "old-password-123",
jwt_secret = secrets.token_urlsafe(64),
)
return storage.DEFAULT_ADMIN_USERNAME
def _verified_secret(username):
return storage.get_user_and_secret(username)[2]
def test_access_token_from_the_replaced_credential_is_rejected(admin):
secret = _verified_secret(admin)
storage.update_password(admin, "new-password-456", revoke_refresh_tokens = True)
token = create_access_token(subject = admin, secret = secret)
with pytest.raises(jwt.InvalidTokenError):
jwt.decode(token, storage.get_jwt_secret(admin), algorithms = [ALGORITHM])
def test_refresh_token_from_the_replaced_credential_is_rejected(admin):
secret = _verified_secret(admin)
# Inserted AFTER the rotation's DELETE, so revocation alone cannot catch it.
storage.update_password(admin, "new-password-456", revoke_refresh_tokens = True)
token = create_refresh_token(subject = admin, secret = secret)
assert storage.verify_refresh_token(token) is None
assert storage.consume_refresh_token(token) is None
def test_a_rejected_refresh_token_is_dropped(admin):
secret = _verified_secret(admin)
storage.update_password(admin, "new-password-456", revoke_refresh_tokens = True)
token = create_refresh_token(subject = admin, secret = secret)
storage.verify_refresh_token(token)
conn = storage.get_connection()
try:
assert conn.execute("SELECT COUNT(*) AS c FROM refresh_tokens").fetchone()["c"] == 0
finally:
conn.close()
def test_tokens_from_the_current_credential_still_work(admin):
secret = _verified_secret(admin)
access = create_access_token(subject = admin, secret = secret)
refresh = create_refresh_token(subject = admin, secret = secret)
jwt.decode(access, storage.get_jwt_secret(admin), algorithms = [ALGORITHM])
assert storage.verify_refresh_token(refresh) == (admin, False)
def test_refresh_cannot_outlive_a_rotation_it_raced(admin):
# /refresh consumes, then mints. A rotation landing in between must not let
# the replacement pair be signed with the credential that just replaced it.
secret = _verified_secret(admin)
token = create_refresh_token(subject = admin, secret = secret)
consumed = storage.consume_refresh_token(token)
assert consumed is not None
_username, _is_desktop, consumed_secret = consumed
storage.update_password(admin, "new-password-456", revoke_refresh_tokens = True)
access = create_access_token(subject = admin, secret = consumed_secret)
refresh = create_refresh_token(subject = admin, secret = consumed_secret)
with pytest.raises(jwt.InvalidTokenError):
jwt.decode(access, storage.get_jwt_secret(admin), algorithms = [ALGORITHM])
assert storage.verify_refresh_token(refresh) is None
def test_desktop_login_cannot_outlive_a_rotation_it_raced(admin):
# The reset deletes the desktop secret, so a desktop-login that validated it
# just beforehand must not mint a session that survives.
raw = storage.create_desktop_secret()
verified = storage.validate_desktop_secret_with_credential(raw)
assert verified is not None
_username, verified_secret = verified
storage.update_password(admin, "new-password-456", revoke_refresh_tokens = True)
access = create_access_token(subject = admin, desktop = True, secret = verified_secret)
refresh = create_refresh_token(subject = admin, desktop = True, secret = verified_secret)
with pytest.raises(jwt.InvalidTokenError):
jwt.decode(access, storage.get_jwt_secret(admin), algorithms = [ALGORITHM])
assert storage.verify_refresh_token(refresh) is None
def test_change_password_cannot_overwrite_a_rotation_it_raced(admin):
# A change-password that verified the old hash must not clobber a reset that
# committed while it was in flight.
_salt, verified_hash, _secret, _must_change = storage.get_user_and_secret(admin)
storage.update_password(admin, "reset-by-the-cli-789", revoke_refresh_tokens = True)
assert not storage.update_password(
admin,
"attacker-chosen-000",
revoke_refresh_tokens = True,
expect_password_hash = verified_hash,
)
salt, pwd_hash, _s, _m = storage.get_user_and_secret(admin)
assert hashing.verify_password("reset-by-the-cli-789", salt, pwd_hash)
def test_api_key_creation_from_a_revoked_credential_is_refused(admin):
generation = storage.credential_generation(_verified_secret(admin))
storage.update_password(admin, "new-password-456", revoke_refresh_tokens = True)
with pytest.raises(storage.CredentialRotated):
storage.create_api_key(username = admin, name = "k", expect_gen = generation)
conn = storage.get_connection()
try:
assert conn.execute("SELECT COUNT(*) AS c FROM api_keys").fetchone()["c"] == 0
finally:
conn.close()
def test_api_key_creation_under_the_current_credential_still_works(admin):
generation = storage.credential_generation(_verified_secret(admin))
raw_key, _row = storage.create_api_key(username = admin, name = "k", expect_gen = generation)
assert storage.validate_api_key(raw_key) == admin
def test_change_password_tokens_are_bound_to_its_own_write(admin):
# The tokens returned to a successful change-password must be signed with the
# secret that write produced, not whatever a later reset put in the DB.
_salt, verified_hash, _secret, _must = storage.get_user_and_secret(admin)
new_secret = storage.update_password(
admin,
"chosen-by-the-user",
revoke_refresh_tokens = True,
expect_password_hash = verified_hash,
)
assert new_secret is not None
storage.update_password(admin, "reset-by-the-cli-789", revoke_refresh_tokens = True)
access = create_access_token(subject = admin, secret = new_secret)
refresh = create_refresh_token(subject = admin, secret = new_secret)
with pytest.raises(jwt.InvalidTokenError):
jwt.decode(access, storage.get_jwt_secret(admin), algorithms = [ALGORITHM])
assert storage.verify_refresh_token(refresh) is None
def test_internal_api_key_minting_honours_the_request_generation(admin):
generation = storage.credential_generation(_verified_secret(admin))
storage.update_password(admin, "new-password-456", revoke_refresh_tokens = True)
with pytest.raises(storage.CredentialRotated):
storage.create_api_key(
username = admin,
name = "data-recipe workflow",
internal = True,
expect_gen = generation,
)
def test_api_key_auth_reports_the_version_the_key_was_valid_under(admin):
# The generation must come from the same transaction as the key check, or a
# revoked key could hand a route the post-reset generation and mint again.
raw, _row = storage.create_api_key(username = admin, name = "agent")
verified = storage.validate_api_key_with_credential(raw)
assert verified is not None
_user, secret = verified
generation = storage.credential_generation(secret)
storage.update_password(admin, "new-password-456", revoke_refresh_tokens = True)
conn = storage.get_connection()
try:
conn.execute("DELETE FROM api_keys")
conn.commit()
finally:
conn.close()
assert storage.validate_api_key(raw) is None
with pytest.raises(storage.CredentialRotated):
storage.create_api_key(username = admin, name = "after", expect_gen = generation)
def test_consuming_a_legacy_token_reports_the_pre_reset_credential(admin):
# An unstamped row has no generation to compare, so consume must read the
# credential inside the delete transaction rather than after committing it.
token = secrets.token_urlsafe(48)
expires_at = (datetime.now(timezone.utc) + timedelta(days = 7)).isoformat()
storage.save_refresh_token(token, admin, expires_at, secret_gen = None)
conn = storage.get_connection()
try:
conn.execute("UPDATE refresh_tokens SET secret_gen = NULL")
conn.commit()
finally:
conn.close()
consumed = storage.consume_refresh_token(token)
assert consumed is not None
_username, _is_desktop, consumed_secret = consumed
storage.update_password(admin, "new-password-456", revoke_refresh_tokens = True)
access = create_access_token(subject = admin, secret = consumed_secret)
with pytest.raises(jwt.InvalidTokenError):
jwt.decode(access, storage.get_jwt_secret(admin), algorithms = [ALGORITHM])
def test_unstamped_legacy_tokens_still_verify(admin):
# Rows written before the secret_gen column existed must not log users out.
token = secrets.token_urlsafe(48)
expires_at = (datetime.now(timezone.utc) + timedelta(days = 7)).isoformat()
storage.save_refresh_token(token, admin, expires_at, secret_gen = None)
conn = storage.get_connection()
try:
conn.execute("UPDATE refresh_tokens SET secret_gen = NULL")
conn.commit()
finally:
conn.close()
assert storage.verify_refresh_token(token) == (admin, False)

View file

@ -134,218 +134,6 @@ def test_ensure_default_admin_loads_existing_bootstrap_after_restart(monkeypatch
assert storage.get_bootstrap_password() == bootstrap_pw
def test_bootstrap_password_file_ends_with_a_newline():
# Otherwise `cat` welds the passphrase onto the shell prompt.
storage.ensure_default_admin()
# Bytes: read_text would decode CRLF back to "\n" and hide a CR.
raw = storage._BOOTSTRAP_PW_PATH.read_bytes()
assert raw == storage.get_bootstrap_password().encode("utf-8") + b"\n"
def test_bootstrap_password_round_trips_across_a_restart_with_the_newline():
storage.ensure_default_admin()
original = storage.get_bootstrap_password()
storage._bootstrap_password = None
assert storage.generate_bootstrap_password() == original
def test_upgrade_normalises_the_bootstrap_file():
# Upgrade path: the admin row exists, so generate_bootstrap_password() never runs.
seed_user()
storage._BOOTSTRAP_PW_PATH.write_bytes(b"legacy-bootstrap-secret")
storage.ensure_default_admin()
assert storage._BOOTSTRAP_PW_PATH.read_bytes() == b"legacy-bootstrap-secret\n"
assert storage.get_bootstrap_password() == "legacy-bootstrap-secret"
@pytest.mark.parametrize(
"other",
[
b"legacy-bootstrap-secret\r\n", # only an unreleased build wrote this
b"legacy-bootstrap-secret\r",
b"legacy-bootstrap-secret ",
],
)
def test_only_an_exactly_unterminated_bootstrap_file_is_touched(other):
# Appending is safe only because it is restricted to the one released shape.
seed_user()
storage._BOOTSTRAP_PW_PATH.write_bytes(other)
storage.ensure_default_admin()
assert storage.get_bootstrap_password() == "legacy-bootstrap-secret"
assert storage._BOOTSTRAP_PW_PATH.read_bytes() == other
def test_upgrade_normalises_when_the_admin_row_is_missing():
storage._BOOTSTRAP_PW_PATH.write_bytes(b"legacy-bootstrap-secret")
assert storage.generate_bootstrap_password() == "legacy-bootstrap-secret"
assert storage._BOOTSTRAP_PW_PATH.read_bytes() == b"legacy-bootstrap-secret\n"
def test_a_well_formed_bootstrap_file_is_not_rewritten():
seed_user()
storage._BOOTSTRAP_PW_PATH.write_bytes(b"legacy-bootstrap-secret\n")
mtime = storage._BOOTSTRAP_PW_PATH.stat().st_mtime_ns
storage.ensure_default_admin()
assert storage._BOOTSTRAP_PW_PATH.stat().st_mtime_ns == mtime
def test_migration_failure_does_not_break_startup(monkeypatch):
seed_user()
storage._BOOTSTRAP_PW_PATH.write_bytes(b"legacy-bootstrap-secret")
real_open = storage.os.open
def refuse(path, flags, *args, **kwargs):
if str(path) == str(storage._BOOTSTRAP_PW_PATH):
raise PermissionError("read-only auth dir")
return real_open(path, flags, *args, **kwargs)
monkeypatch.setattr(storage.os, "open", refuse)
storage.ensure_default_admin()
assert storage.get_bootstrap_password() == "legacy-bootstrap-secret"
assert storage._BOOTSTRAP_PW_PATH.read_bytes() == b"legacy-bootstrap-secret"
def test_normalising_never_recreates_a_cleared_bootstrap_file(monkeypatch):
# A rename would resurrect revoked plaintext if the password changed after the read.
seed_user()
storage._BOOTSTRAP_PW_PATH.write_bytes(b"legacy-bootstrap-secret")
real_open = storage.os.open
def clear_then_open(path, flags, *args, **kwargs):
if str(path) == str(storage._BOOTSTRAP_PW_PATH):
storage._BOOTSTRAP_PW_PATH.unlink(missing_ok = True)
return real_open(path, flags, *args, **kwargs)
monkeypatch.setattr(storage.os, "open", clear_then_open)
assert storage._read_persisted_bootstrap_password() == "legacy-bootstrap-secret"
assert not storage._BOOTSTRAP_PW_PATH.exists()
def test_normalising_does_not_overwrite_a_rotated_bootstrap_file(monkeypatch):
seed_user()
storage._BOOTSTRAP_PW_PATH.write_bytes(b"legacy-bootstrap-secret")
real_open = storage.os.open
def rotate_then_open(path, flags, *args, **kwargs):
if str(path) == str(storage._BOOTSTRAP_PW_PATH):
storage._BOOTSTRAP_PW_PATH.write_bytes(b"brand-new-secret\n")
return real_open(path, flags, *args, **kwargs)
monkeypatch.setattr(storage.os, "open", rotate_then_open)
storage._read_persisted_bootstrap_password()
# The append may add a second newline; the rotated credential must survive.
raw = storage._BOOTSTRAP_PW_PATH.read_bytes()
assert raw.strip() == b"brand-new-secret"
storage._bootstrap_password = None
assert storage._load_bootstrap_password() == "brand-new-secret"
def test_leading_whitespace_bootstrap_file_is_left_alone(monkeypatch):
# An in-place rewrite is not atomic, so only the exact unterminated shape is touched.
seed_user()
storage._BOOTSTRAP_PW_PATH.write_bytes(b" legacy-bootstrap-secret ")
storage.ensure_default_admin()
assert storage.get_bootstrap_password() == "legacy-bootstrap-secret"
assert storage._BOOTSTRAP_PW_PATH.read_bytes() == b" legacy-bootstrap-secret "
def test_normalising_opens_the_file_in_binary_mode(monkeypatch):
# Without O_BINARY, Windows text mode turns the written LF back into CRLF.
seed_user()
storage._BOOTSTRAP_PW_PATH.write_bytes(b"legacy-bootstrap-secret")
monkeypatch.setattr(storage.os, "O_BINARY", 0x8000, raising = False)
seen = []
real_open = storage.os.open
def spy(path, flags, *args, **kwargs):
if str(path) == str(storage._BOOTSTRAP_PW_PATH):
seen.append(flags)
return real_open(path, flags & ~0x8000, *args, **kwargs)
monkeypatch.setattr(storage.os, "open", spy)
storage.ensure_default_admin()
assert seen and all(f & 0x8000 for f in seen), seen
def test_clearing_by_truncation_mid_normalisation_is_not_undone(monkeypatch):
# clear_bootstrap_password() truncates through its own descriptor when the unlink
# fails (Windows, while ours is open); the append must not restore the plaintext.
seed_user()
storage._BOOTSTRAP_PW_PATH.write_bytes(b"legacy-bootstrap-secret")
real_open = storage.os.open
def truncate_then_open(path, flags, *args, **kwargs):
fd = real_open(path, flags, *args, **kwargs)
if str(path) == str(storage._BOOTSTRAP_PW_PATH):
storage._BOOTSTRAP_PW_PATH.write_text("", encoding = "utf-8")
return fd
monkeypatch.setattr(storage.os, "open", truncate_then_open)
storage._read_persisted_bootstrap_password()
# A lone newline over a cleared file still reads back as no password.
assert storage._BOOTSTRAP_PW_PATH.read_bytes().strip() == b""
storage._bootstrap_password = None
assert storage._load_bootstrap_password() is None
def test_normalising_works_without_fchmod(monkeypatch):
# os.fchmod only reached Windows in 3.13; its absence must not raise.
seed_user()
storage._BOOTSTRAP_PW_PATH.write_bytes(b"legacy-bootstrap-secret")
monkeypatch.delattr(storage.os, "fchmod", raising = False)
storage.ensure_default_admin()
assert storage._BOOTSTRAP_PW_PATH.read_bytes() == b"legacy-bootstrap-secret\n"
assert storage.get_bootstrap_password() == "legacy-bootstrap-secret"
def test_persisting_the_bootstrap_password_is_atomic(monkeypatch, tmp_path):
# A partial write would destroy the only plaintext recovery credential.
storage._persist_bootstrap_password("original-secret")
def boom(src, dst):
raise OSError("crash before replace")
monkeypatch.setattr(storage.os, "replace", boom)
with pytest.raises(OSError):
storage._persist_bootstrap_password("new-secret")
assert storage._BOOTSTRAP_PW_PATH.read_bytes() == b"original-secret\n"
leftovers = [
p.name
for p in storage._BOOTSTRAP_PW_PATH.parent.iterdir()
if "bootstrap_password." in p.name
]
assert leftovers == []
def test_ensure_default_admin_does_not_generate_for_empty_existing_bootstrap():
seed_user()
storage._BOOTSTRAP_PW_PATH.write_text(" \n", encoding = "utf-8")
@ -445,7 +233,7 @@ def test_consume_refresh_token_second_call_returns_none():
storage.save_refresh_token(raw, storage.DEFAULT_ADMIN_USERNAME, expires)
first = storage.consume_refresh_token(raw)
assert first[:2] == (storage.DEFAULT_ADMIN_USERNAME, False)
assert first == (storage.DEFAULT_ADMIN_USERNAME, False)
second = storage.consume_refresh_token(raw)
assert second is None
@ -474,7 +262,7 @@ def test_consume_refresh_token_concurrent_only_one_succeeds(tmp_path, monkeypatc
successes = [r for r in results if r is not None]
assert len(successes) == 1, f"expected exactly one consumer to win, got {len(successes)}"
assert successes[0][:2] == (storage.DEFAULT_ADMIN_USERNAME, False)
assert successes[0] == (storage.DEFAULT_ADMIN_USERNAME, False)
def test_consume_refresh_token_expired_returns_none():
@ -548,28 +336,6 @@ def test_local_recipe_token_authenticates_as_admin_for_web_user(loaded_local_mod
assert asyncio.run(get_current_subject(credentials)) == storage.DEFAULT_ADMIN_USERNAME
def test_rotated_credential_job_start_is_401_not_500(loaded_local_model):
# A reset-password landing mid-request makes the workflow-key mint refuse.
# That must reach the client as a revoked credential, not an unhandled error.
from fastapi import HTTPException
seed_user()
jobs_route = data_recipe_jobs_module()
stale_gen = storage.credential_generation(secrets.token_urlsafe(64))
with pytest.raises(storage.CredentialRotated):
jobs_route._inject_local_providers(local_recipe(), local_recipe_request("t"), stale_gen)
def _boom(*_a, **_k):
raise storage.CredentialRotated("revoked")
jobs_route._inject_local_providers = _boom
payload = SimpleNamespace(recipe = local_recipe(), run = {})
with pytest.raises(HTTPException) as excinfo:
jobs_route.create_job(payload, local_recipe_request("t"), ("unsloth", stale_gen))
assert excinfo.value.status_code == 401
def test_desktop_login_rejects_invalid_secret():
seed_user(must_change_password = False)
client = auth_client()
@ -592,7 +358,7 @@ def test_write_desktop_secret_file_is_0600_on_unix(tmp_path):
studio_cli._write_auth_secret(path, "desktop-secret")
assert path.read_bytes() == b"desktop-secret\n"
assert path.read_text() == "desktop-secret"
if platform.system() != "Windows":
assert oct(path.stat().st_mode & 0o777) == "0o600"
@ -602,31 +368,18 @@ def test_reset_password_removes_desktop_secret_files(tmp_path, monkeypatch):
from unsloth_cli.commands import studio as studio_cli
auth_dir = tmp_path / "auth"
auth_dir.mkdir()
(auth_dir / "auth.db").write_text("db")
(auth_dir / ".bootstrap_password").write_text("boot")
(auth_dir / ".desktop_secret").write_text("new")
monkeypatch.setattr(studio_cli, "STUDIO_HOME", tmp_path)
secret = studio_cli._create_desktop_secret_in_cli()
studio_cli._write_auth_secret(auth_dir / studio_cli.DESKTOP_SECRET_FILE, secret)
(auth_dir / studio_cli.BOOTSTRAP_PASSWORD_FILE).write_text("boot")
result = CliRunner().invoke(studio_cli.studio_app, ["reset-password"])
assert result.exit_code == 0, result.output
# The DB survives on purpose: a running server keeps serving from its admin row.
assert (auth_dir / "auth.db").exists()
assert not (auth_dir / studio_cli.BOOTSTRAP_PASSWORD_FILE).exists()
assert not (auth_dir / studio_cli.DESKTOP_SECRET_FILE).exists()
conn = studio_cli._connect_auth_db()
try:
surviving = conn.execute(
"SELECT COUNT(*) FROM app_secrets WHERE key IN (?, ?)",
(
studio_cli.DESKTOP_SECRET_HASH_KEY,
studio_cli.DESKTOP_SECRET_CREATED_AT_KEY,
),
).fetchone()[0]
finally:
conn.close()
assert surviving == 0
assert result.exit_code == 0
assert not (auth_dir / "auth.db").exists()
assert not (auth_dir / ".bootstrap_password").exists()
assert not (auth_dir / ".desktop_secret").exists()
def test_reset_password_removes_desktop_secret_files_without_db(tmp_path, monkeypatch):
@ -772,8 +525,7 @@ if result.exit_code != 0:
capture_output = True,
)
assert result.returncode == 0, result.stderr + result.stdout
# Strip like the src-tauri readers do.
secret = (auth_dir / ".desktop_secret").read_text().strip()
secret = (auth_dir / ".desktop_secret").read_text()
assert secret.startswith("desktop-")
conn = sqlite3.connect(auth_dir / "auth.db")
@ -881,7 +633,7 @@ def test_update_password_clears_desktop_secret():
assert storage.validate_desktop_secret(raw) == storage.DEFAULT_ADMIN_USERNAME
changed = storage.update_password(storage.DEFAULT_ADMIN_USERNAME, "new-admin-password")
assert changed
assert changed is True
assert storage.validate_desktop_secret(raw) is None
@ -890,7 +642,7 @@ def test_update_password_on_unknown_user_leaves_desktop_secret_intact():
raw = storage.create_desktop_secret()
changed = storage.update_password("not-a-user", "irrelevant")
assert not changed
assert changed is False
assert storage.validate_desktop_secret(raw) == storage.DEFAULT_ADMIN_USERNAME

View file

@ -1487,418 +1487,7 @@ def test_internal_reprompt_attempts_do_not_duplicate_visible_text(monkeypatch):
content_texts = [event.get("text", "") for event in events if event.get("type") == "content"]
assert content_texts == ["I will use render_html now."]
# Each retry restates the last, so the loop gives up: initial + 2 re-prompts.
assert len(payloads) == 3 < _MAX_REPROMPTS + 1
def test_post_tool_stall_still_nudged_after_a_pre_tool_reprompt(monkeypatch):
"""The post-tool nudge has its own budget, so an earlier stall can't spend it."""
streams = [
[_sse({"content": "I will search the web now."}), _done()],
[
_sse(
{
"tool_calls": [
{
"index": 0,
"id": "call_first",
"type": "function",
"function": {
"name": "web_search",
"arguments": json.dumps({"query": "red square"}),
},
}
]
}
),
_done(),
],
[_sse({"content": "Let me summarize the results."}), _done()],
[_sse({"content": "Final answer: the square is red."}), _done()],
]
payloads: list[dict] = []
backend = _make_backend(monkeypatch, streams, payloads)
calls: list[tuple[str, dict]] = []
def fake_execute_tool(name, arguments, **_kwargs):
calls.append((name, arguments))
return "Search results: red is #f00."
monkeypatch.setattr("core.inference.tools.execute_tool", fake_execute_tool)
tools = [
{
"type": "function",
"function": {
"name": "web_search",
"description": "Search the web.",
"parameters": {
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"],
},
},
}
]
events = list(
backend.generate_chat_completion_with_tools(
messages = [{"role": "user", "content": "Make a red square."}],
tools = tools,
max_tool_iterations = 2,
)
)
assert len(payloads) == 4
assert len(calls) == 1
nudges = [
message
for message in payloads[-1]["messages"]
if message.get("role") == "user" and "call web_search now" in message.get("content", "")
]
assert len(nudges) == 2
content_texts = [event.get("text", "") for event in events if event.get("type") == "content"]
assert content_texts[-1] == "Final answer: the square is red."
def test_post_tool_reprompt_budget_is_one(monkeypatch):
"""The post-tool nudge fires once; a second stall is surrendered as the answer."""
streams = [
[
_sse(
{
"tool_calls": [
{
"index": 0,
"id": "call_first",
"type": "function",
"function": {
"name": "web_search",
"arguments": json.dumps({"query": "red square"}),
},
}
]
}
),
_done(),
],
[_sse({"content": "Let me summarize the results."}), _done()],
[_sse({"content": "Now I will check the sources."}), _done()],
]
payloads: list[dict] = []
backend = _make_backend(monkeypatch, streams, payloads)
monkeypatch.setattr(
"core.inference.tools.execute_tool",
lambda *_a, **_k: "Search results: red is #f00.",
)
tools = [
{
"type": "function",
"function": {
"name": "web_search",
"description": "Search the web.",
"parameters": {
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"],
},
},
}
]
list(
backend.generate_chat_completion_with_tools(
messages = [{"role": "user", "content": "Make a red square."}],
tools = tools,
max_tool_iterations = 2,
)
)
assert len(payloads) == 3
def test_repeat_guard_resets_after_a_tool_runs(monkeypatch):
"""A tool execution opens a new phase, so the same intent text is nudged again.
Without the reset the pre-tool stall text still sits in the repeat tracker and
the identical post-tool stall is surrendered as the visible final answer.
"""
stall = "I will search the web now."
streams = [
[_sse({"content": stall}), _done()],
[
_sse(
{
"tool_calls": [
{
"index": 0,
"id": "call_first",
"type": "function",
"function": {
"name": "web_search",
"arguments": json.dumps({"query": "red square"}),
},
}
]
}
),
_done(),
],
[_sse({"content": stall}), _done()],
[_sse({"content": "Final answer: the square is red."}), _done()],
]
payloads: list[dict] = []
backend = _make_backend(monkeypatch, streams, payloads)
monkeypatch.setattr(
"core.inference.tools.execute_tool",
lambda *_a, **_k: "Search results: red is #f00.",
)
tools = [
{
"type": "function",
"function": {
"name": "web_search",
"description": "Search the web.",
"parameters": {
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"],
},
},
}
]
events = list(
backend.generate_chat_completion_with_tools(
messages = [{"role": "user", "content": "Make a red square."}],
tools = tools,
max_tool_iterations = 2,
)
)
assert len(payloads) == 4
content_texts = [event.get("text", "") for event in events if event.get("type") == "content"]
assert content_texts[-1] == "Final answer: the square is red."
def test_restatement_keeps_deletions_that_change_the_answer():
"""A dropped word can invert the meaning, so a subset is not a restatement."""
from core.inference.tool_call_parser import is_reprompt_restatement
from core.inference.llama_cpp import _should_suppress_forced_no_tool_output as suppress
previous = "Now I think the feature is not supported in version 1."
corrected = "Now I think the feature is supported in version 1."
assert not is_reprompt_restatement(corrected, previous)
assert not suppress(corrected, previous)
stall = "I'll search for that now."
assert is_reprompt_restatement(stall, stall)
assert is_reprompt_restatement("Understood. " + stall, "Understood, " + stall)
assert not is_reprompt_restatement(stall + " Tokyo.", stall)
def test_forced_turn_suppression_covers_obligation_phrasing():
from core.inference.llama_cpp import _should_suppress_forced_no_tool_output as suppress
for stall in (
"I need to use render_html now",
"Need to call web_search",
"I will summarize the results now",
"I have to run the search first",
"I should call web_search now",
"I should use render_html now",
# Plain modals take a bare infinitive, not the need|have|ought "to" group.
"I must call web_search now",
"I must use render_html now",
"I must run the search first",
# Subjectless plans open a new sentence just as often as a new line.
"Okay. Need to call web_search now.",
"Understood. Going to search now.",
# Subjectless modals, not just subjectless semi-modals.
"Must call web_search now.",
"Should search the web now.",
# A missing answer is not a final answer: the plan behind it is still a stall.
"I should call web_search because the answer is not in the provided context",
"I must run the search since the answer is unknown so far",
# A pivot with nothing behind it answers nothing.
"I should call web_search, though.",
"I need to run the search, but",
# A purpose clause is part of the plan, not a summary of results.
"I need to call web_search to summarize the results",
):
assert suppress(stall), f"leaked {stall!r}"
for answer in (
"You need to install the package first.",
"The square is red.",
"Here is the summary of what I found.",
"Run `pip install unsloth` to get started.",
"I should mention that the square is red.",
# Obligation phrasing mid-sentence is prose that happens to name a tool.
"The API I should invoke is foo() because it supports streaming.",
"The tool I need to use is documented here.",
# "invoke"/"query" read as technical prose far more often than as a stall.
"I should invoke foo() because it supports streaming.",
"I should query the cache first for a faster path.",
"You should call your bank about the charge.",
# Second person is the user's obligation, not the model's plan.
"You must call your bank about the charge.",
"I must admit the square is red.",
# A plan that pivots to an answer must ship the answer with it.
"I should call web_search, but the answer is Tokyo.",
"I need to call web_search. The answer is Tokyo.",
"I should call web_search to confirm, but Tokyo is the capital of Japan.",
"I must run the search, however the result is already known: 42.",
):
assert not suppress(answer), f"dropped {answer!r}"
def test_forced_turn_intent_lead_in_needs_a_restatement_to_be_dropped():
"""A bare intent match is a stall only when the retry restates the nudge.
``INTENT_SIGNAL`` fires on lead-ins that introduce a real answer ("Now I
have the results. ..."), so matching it alone would discard the answer.
"""
from core.inference.llama_cpp import _should_suppress_forced_no_tool_output as suppress
stall = "I will summarize the results now"
answer = "Now I have the search results. The capital of Japan is Tokyo."
# Restating the nudged text is still a stall.
assert suppress(stall, stall)
assert suppress("Understood. " + stall, "Understood, " + stall)
# Progress past the nudged text keeps the answer, lead-in and all.
assert not suppress(answer, stall)
assert not suppress("Step 3: done. Tokyo is the capital.", stall)
# Near-repeat is enough to stop nudging, never enough to drop the turn.
assert not suppress(stall + ": Tokyo.", stall)
# An obligation plan is a stall on its own, no previous text needed.
assert suppress("I must call web_search now", answer)
def test_forced_turn_answer_with_an_intent_lead_in_survives_after_a_tool(monkeypatch):
"""The post-tool retry answers behind a lead-in; the answer must still ship.
The nudge budget is spent, so the reply lands on the suppression branch.
``INTENT_SIGNAL`` matches its "Now I ..." opener, and dropping it on that
alone left the user with the stall and no answer at all.
"""
answer = "Now I have the results. The capital of Japan is Tokyo."
streams = [
[
_sse(
{
"tool_calls": [
{
"index": 0,
"id": "call_first",
"type": "function",
"function": {
"name": "web_search",
"arguments": json.dumps({"query": "capital of Japan"}),
},
}
]
}
),
_done(),
],
[_sse({"content": "Let me summarize what I found."}), _done()],
[_sse({"content": answer}), _done()],
]
payloads: list[dict] = []
backend = _make_backend(monkeypatch, streams, payloads)
monkeypatch.setattr(
"core.inference.tools.execute_tool",
lambda *_a, **_k: "Search results: Tokyo.",
)
tools = [
{
"type": "function",
"function": {
"name": "web_search",
"description": "Search the web.",
"parameters": {
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"],
},
},
}
]
events = list(
backend.generate_chat_completion_with_tools(
messages = [{"role": "user", "content": "What is the capital of Japan?"}],
tools = tools,
max_tool_iterations = 2,
)
)
assert len(payloads) == 3
content_texts = [event.get("text", "") for event in events if event.get("type") == "content"]
assert content_texts[-1] == answer
def test_forced_turn_answer_with_an_intent_lead_in_survives_pre_tool(monkeypatch):
"""Same guarantee once the pre-tool nudge budget is spent on distinct stalls."""
answer = "Now I see the data clearly. Tokyo is the capital."
streams = [
[_sse({"content": text}), _done()]
for text in (
"I will look that up for you.",
"Now I have the search results. The capital of Japan is Tokyo.",
"Now I can confirm it. Japan's capital city is Tokyo.",
answer,
)
]
payloads: list[dict] = []
backend = _make_backend(monkeypatch, streams, payloads)
def fake_execute_tool(name, arguments, **_kwargs):
raise AssertionError(f"unexpected tool execution: {name} {arguments}")
monkeypatch.setattr("core.inference.tools.execute_tool", fake_execute_tool)
tools = [
{
"type": "function",
"function": {
"name": "web_search",
"description": "Search the web.",
"parameters": {
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"],
},
},
}
]
events = list(
backend.generate_chat_completion_with_tools(
messages = [{"role": "user", "content": "What is the capital of Japan?"}],
tools = tools,
max_tool_iterations = 2,
)
)
# Initial turn plus the three pre-tool nudges.
assert len(payloads) == _MAX_REPROMPTS + 1
content_texts = [event.get("text", "") for event in events if event.get("type") == "content"]
assert content_texts[-1] == answer
def test_forced_reprompt_plain_final_answer_is_visible(monkeypatch):
@ -2495,51 +2084,6 @@ def test_confirm_tool_calls_skips_gguf_rag_autoinject(monkeypatch):
assert any(event.get("type") == "content" and event.get("text") == "Done." for event in events)
def test_rag_autoinject_counts_as_a_prior_tool_execution(monkeypatch):
"""Autoinjected retrieval runs before the controller, so history stays empty.
Without counting it the turn reads as pre-tool and gets the full re-prompt
budget, repeating the expensive retrieval the post-tool cap exists to stop.
"""
stall = "I will summarize the retrieved passages now."
streams = [
[_sse({"content": stall}), _done()],
[_sse({"content": "Still working on the summary."}), _done()],
[_sse({"content": "Final answer: the passages describe Tokyo."}), _done()],
]
payloads: list[dict] = []
backend = _make_backend(monkeypatch, streams, payloads)
monkeypatch.setattr(
"core.inference.tools.build_rag_autoinject",
lambda *_a, **_k: {
"events": [],
"messages": [{"role": "user", "content": "Retrieved passage: Tokyo."}],
},
)
events = list(
backend.generate_chat_completion_with_tools(
messages = [{"role": "user", "content": "summarize the docs"}],
tools = [{"type": "function", "function": {"name": "search_knowledge_base"}}],
max_tool_iterations = 2,
rag_scope = {"thread_id": "t1"},
)
)
# Initial turn plus one retry; read as pre-tool it would spend the full budget.
assert len(payloads) == 2, payloads
nudges = [
message
for message in payloads[-1]["messages"]
if message.get("role") == "user"
and "call search_knowledge_base now" in message.get("content", "")
]
assert len(nudges) == 1, nudges
assert events
def test_confirm_tool_calls_deny_skips_gguf_tool_and_retry_can_execute(monkeypatch):
same_call = _structured_tool_call("python", {"code": "print(1)"}, "call_py")
streams = [

View file

@ -247,8 +247,8 @@ def test_lifespan_honors_bootstrap_suppression_in_source():
def test_clear_bootstrap_password_truncates_when_unlink_fails(monkeypatch, tmp_path):
# If the file cannot be unlinked (Windows AV / read-only auth dir), clear must
# truncate it so its stale plaintext cannot be re-seeded by
# generate_bootstrap_password() if auth.db is ever recreated, which would
# re-validate the revoked bootstrap password.
# generate_bootstrap_password() after a later reset-password deletes auth.db,
# which would re-validate the revoked bootstrap password.
import pathlib
pw_path = tmp_path / ".bootstrap_password"

View file

@ -1,96 +0,0 @@
# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""An accuracy floor for the plan-without-action classifier, on real model output.
The rest of the tool-loop suites pin behaviour on hand-written example sentences,
which is how the patterns here were tuned. That says nothing about how often the
classifier is right on what models actually emit, so this file scores it against a
corpus captured from local models (``tests/data/plan_vs_answer.jsonl``).
How the corpus was built: three GGUF models (Qwen3-0.6B, Qwen3-1.7B,
Llama-3.2-1B-Instruct) were driven through llama-server with the real Studio tool
schemas over prompts spanning tool-requiring questions, questions needing no tool,
list-formatted answers, ambiguous requests, non-English, and follow-ups issued after
a tool had already run. Turns cut off by the token cap were dropped, since a
truncation is not a stall.
Every turn here is a *finished answer*: the turn called no tool, and when the
production nudge was appended and the turn regenerated three times, not one retry
produced a tool call. A forceful re-prompt could not extract an action, so there was
no action left to take. Nudging these is wasted work, and in the GGUF loop the
retry's text can then be discarded, which costs the user a visible answer.
Measured when this landed, over the 300 turns:
tree nudged retry discarded
origin/main (pre-PR) 36 (12.0%) 60 (20.2%)
this PR 5 ( 1.7%) 1 ( 0.3%)
The budgets below sit above the measured counts so that innocuous wording changes
do not fail the build, and far below the pre-PR counts so a real regression does.
A failure prints the offending turns: fix the pattern, or if the turn really is a
stall, correct its label here.
"""
import json
from pathlib import Path
from core.inference.llama_cpp import _should_suppress_forced_no_tool_output
from core.inference.tool_call_parser import is_short_intent_without_action
DATA = Path(__file__).parent / "data" / "plan_vs_answer.jsonl"
# Measured 5 of 300; pre-PR was 36.
NUDGE_BUDGET = 9
# Measured 1 of 300; pre-PR was 60. Tighter, because this one destroys output.
DISCARD_BUDGET = 4
def _corpus():
with open(DATA, encoding = "utf-8") as fh:
return [json.loads(line) for line in fh if line.strip()]
def _report(rows, limit = 10):
lines = []
for row in rows[:limit]:
text = " ".join(row["text"].split())
lines.append(
f" [{row['model']}/{row['prompt_class']}] {row['prompt']!r}\n {text[:200]!r}"
)
if len(rows) > limit:
lines.append(f" ... and {len(rows) - limit} more")
return "\n".join(lines)
def test_corpus_is_intact():
"""Guards the budgets: they mean nothing if the corpus silently shrinks."""
corpus = _corpus()
assert len(corpus) == 300
assert all(row["text"].strip() for row in corpus)
# Every row is a finished answer by construction.
assert all(row["retry_tool_calls"] == 0 for row in corpus)
def test_finished_answers_are_rarely_nudged():
"""A finished answer costs a whole extra generation when it is nudged."""
nudged = [row for row in _corpus() if is_short_intent_without_action(row["text"])]
assert len(nudged) <= NUDGE_BUDGET, (
f"{len(nudged)}/300 finished answers classified as plans "
f"(budget {NUDGE_BUDGET}):\n{_report(nudged)}"
)
def test_finished_answers_are_not_discarded():
"""The retry's text is all the user gets, so discarding it is the worst case."""
discarded = [
row
for row in _corpus()
if row["retry_text"].strip()
and _should_suppress_forced_no_tool_output(row["retry_text"], row["text"])
]
assert len(discarded) <= DISCARD_BUDGET, (
f"{len(discarded)}/300 finished retries would be discarded "
f"(budget {DISCARD_BUDGET}):\n{_report(discarded)}"
)

File diff suppressed because it is too large Load diff

View file

@ -45,20 +45,8 @@ def _build_structlog_stub():
_maybe_stub("loggers", _build_loggers_stub)
_maybe_stub("structlog", _build_structlog_stub)
import pytest
import utils.hardware.hardware as hw # noqa: E402
# The DRM/KFD readers below are Linux-only in production: _rocm_linux_amdgpu_cards and
# _rocm_linux_sysfs_vram_by_pci_gb return early unless platform.system() is "Linux", and
# _rocm_kfd_gpu_pci_ids only ever globs /sys/class/kfd. Their fake sysfs tree needs PCI
# addresses like "0000:00:02.0" as directory names and POSIX separators in the paths the
# readers match; Windows permits neither, so the tree cannot be represented there.
linux_only = pytest.mark.skipif(
not sys.platform.startswith("linux"),
reason = "covers Linux-only DRM/KFD sysfs parsing driven by a fake /sys tree",
)
def _device(
index,
@ -111,7 +99,6 @@ def _fake_drm(tmp_path, monkeypatch, cards):
return card_paths
@linux_only
def test_linux_vram_keyed_by_pci_excludes_foreign_adapters(monkeypatch, tmp_path):
# Foreign (non-amdgpu) adapters contribute no entry, so they cannot shift ordinals.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
@ -130,7 +117,6 @@ def test_linux_vram_keyed_by_pci_excludes_foreign_adapters(monkeypatch, tmp_path
}
@linux_only
def test_linux_vram_omits_bad_cards_without_shifting(monkeypatch, tmp_path):
# A zero-total card has no entry; identity keying means its absence renumbers nothing.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
@ -145,7 +131,6 @@ def test_linux_vram_omits_bad_cards_without_shifting(monkeypatch, tmp_path):
assert hw._rocm_linux_sysfs_vram_by_pci_gb() == {"0000:41:00.0": (2.0, 16.0)}
@linux_only
def test_linux_vram_omits_amd_card_without_vram_files(monkeypatch, tmp_path):
# An APU with no mem_info_vram_* files has no entry; the discrete card keeps its address.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
@ -189,7 +174,6 @@ def _fake_kfd(tmp_path, monkeypatch, nodes):
return node_paths
@linux_only
def test_kfd_lists_gpu_nodes_in_device_order(monkeypatch, tmp_path):
# The CPU node (simd_count 0) takes no ordinal; GPU nodes in node-id order are HIP's order.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
@ -205,14 +189,12 @@ def test_kfd_lists_gpu_nodes_in_device_order(monkeypatch, tmp_path):
assert hw._rocm_kfd_gpu_pci_ids() == ["0000:03:00.0", "0000:41:00.0"]
@linux_only
def test_kfd_decodes_domain_device_and_function(monkeypatch, tmp_path):
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
_fake_kfd(tmp_path, monkeypatch, [(1, 64, (0xC1 << 8) | (0x1F << 3) | 5, 0x1234, _AMD)])
assert hw._rocm_kfd_gpu_pci_ids() == ["1234:c1:1f.5"]
@linux_only
def test_kfd_skips_non_amd_gpu_nodes(monkeypatch, tmp_path):
# An NVIDIA KFD node is not a HIP device: it must take no ordinal, else it
# shifts every AMD GPU and ROCm device 1 resolves to AMD GPU 0.
@ -230,7 +212,6 @@ def test_kfd_skips_non_amd_gpu_nodes(monkeypatch, tmp_path):
assert hw._rocm_kfd_gpu_pci_ids() == ["0000:03:00.0", "0000:41:00.0"]
@linux_only
def test_kfd_fails_closed_when_a_gpu_has_no_location(monkeypatch, tmp_path):
# Dropping an unplaceable AMD GPU shifts later ordinals; fail closed for the whole map.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
@ -245,7 +226,6 @@ def test_kfd_fails_closed_when_a_gpu_has_no_location(monkeypatch, tmp_path):
assert hw._rocm_kfd_gpu_pci_ids() == []
@linux_only
def test_kfd_fails_closed_when_a_node_is_unreadable(monkeypatch, tmp_path):
# An unreadable node could be a GPU; assuming otherwise would shift ordinals.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
@ -261,23 +241,6 @@ def test_kfd_fails_closed_when_a_node_is_unreadable(monkeypatch, tmp_path):
assert hw._rocm_kfd_gpu_pci_ids() == []
@linux_only
def test_kfd_fails_closed_when_a_node_does_not_decode(monkeypatch, tmp_path):
# UnicodeDecodeError is a ValueError, so it slips past `except OSError` and
# would shift every later HIP ordinal.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
paths = _fake_kfd(
tmp_path,
monkeypatch,
[
(1, 304, (0x03 << 8) | 0, 0, _AMD),
(2, 304, (0x41 << 8) | 0, 0, _AMD),
],
)
(Path(paths[0]) / "properties").write_bytes(b"simd_count 304\nvendor_id \x80\xff\n")
assert hw._rocm_kfd_gpu_pci_ids() == []
def test_kfd_absent_yields_no_device_order(monkeypatch):
monkeypatch.setattr(hw.glob, "glob", lambda pattern: [])
assert hw._rocm_kfd_gpu_pci_ids() == []
@ -459,10 +422,6 @@ def test_visible_utilization_rocm_fallback_overlays(monkeypatch):
):
monkeypatch.delenv(_var, raising = False)
monkeypatch.setattr(hw, "IS_ROCM", True)
# No AMD adapter data on this host. On Windows this branch runs ahead of the torch
# fallback under test, and probing it imports torch, which the CI runner does not
# install. Off Windows the real function is never reached, so this changes nothing.
monkeypatch.setattr(hw, "_rocm_windows_per_device_vram", lambda ids: [])
monkeypatch.setattr(hw, "get_device", lambda: hw.DeviceType.CUDA)
monkeypatch.setattr(hw, "_smi_query", lambda *a, **k: None) # amd-smi unavailable
monkeypatch.setattr(
@ -491,10 +450,6 @@ def test_visible_utilization_rocm_fallback_overlays(monkeypatch):
def test_visible_utilization_relative_index_skips_overlay(monkeypatch):
# UUID/MIG mask gives relative indices; the overlay matches physical index, so it must not run.
monkeypatch.setattr(hw, "IS_ROCM", True)
# No AMD adapter data on this host. On Windows this branch runs ahead of the torch
# fallback under test, and probing it imports torch, which the CI runner does not
# install. Off Windows the real function is never reached, so this changes nothing.
monkeypatch.setattr(hw, "_rocm_windows_per_device_vram", lambda ids: [])
monkeypatch.setattr(hw, "get_device", lambda: hw.DeviceType.CUDA)
monkeypatch.setattr(hw, "_smi_query", lambda *a, **k: None)
monkeypatch.setattr(

View file

@ -2232,33 +2232,6 @@ def test_reprompt_names_only_active_tools_not_hardcoded():
assert "python" not in reprompt["content"]
def test_reprompt_stops_when_the_retry_restates_the_stall():
"""A nudge answered with the same text has not worked; do not spend the budget."""
captured: list[list] = []
stall = "I'll search for that now."
def fake_single_turn(messages, active_tools = None):
captured.append(list(messages))
yield stall # same forward-looking intent every time
exec_fn = FakeExecuteTool([])
_events = _collect_events(
run_safetensors_tool_loop(
single_turn = fake_single_turn,
messages = [{"role": "user", "content": "find X"}],
tools = [{"type": "function", "function": {"name": "search_knowledge_base"}}],
execute_tool = exec_fn,
auto_heal_tool_calls = True,
nudge_tool_calls = True,
max_tool_iterations = 3,
)
)
# One nudge, then the repeat guard stops it: two generations, not MAX_ACT_REPROMPTS + 1.
assert len(captured) == 2, captured
def test_reprompt_is_announced_on_the_status_channel():
# The re-prompted turn is hidden, so the badge is the only sign of life.
# Blank still comes first: the route resets its text cursor only on that.
@ -3651,22 +3624,8 @@ class TestGGUFSafetensorsHealingParity:
"Let me check",
"I am going to call the tool",
"First, I will explore",
"First, let's search the web",
"First, let us search the web",
# Imperative plans carry no pronoun; an action verb is enough.
"First, search the web for the latest release notes.",
"First, check the documentation.",
"First, analyze the attached data",
"The first step is to search the web",
"First, my plan is to search the web.",
"First: search the web for release notes.",
"First - search the web for release notes.",
"First \u2013 search the web for release notes.",
"First, our approach is to check the docs.",
"Here's my plan",
"Now I need to call web_search",
# The "let me know" exemption is scoped to "let me", not all direct intent.
"I will know the answer after I search the web",
):
assert shared_re.search(phrase), f"missed {phrase!r}"
assert shared_fn(phrase), f"helper missed {phrase!r}"
@ -3682,18 +3641,6 @@ class TestGGUFSafetensorsHealingParity:
# force a tool-call re-prompt on it.
"I will not search the web for that.",
"I'll never call that tool.",
# Hands control back rather than announcing an action.
"Let me know if you need anything else.",
"First, the answer is 42",
"First, the result is 3.",
"First, it is 42",
"First, my answer is 42",
"The first line is blank.",
# Ordinal prose, not a plan.
"First place went to Alice",
"First class is available",
# Advice to the user, not work for this turn.
"First, install the package.",
):
assert not shared_re.search(plain), f"wrongly fired on {plain!r}"
assert not shared_fn(plain), f"helper wrongly fired on {plain!r}"
@ -3706,98 +3653,6 @@ class TestGGUFSafetensorsHealingParity:
assert gguf_cap == sf_cap == shared_cap
def test_reprompt_repeat_keeps_punctuation_bearing_terms(self):
# Stripping all non-word chars collapsed "C++" and "C#" to "c", so different
# plans compared equal and the retry lost its nudge.
from core.inference.tool_call_parser import is_reprompt_repeat
assert not is_reprompt_repeat("I will search for C#.", "I will search for C++.")
# A leading mark is part of the term too.
assert not is_reprompt_repeat("I will search for .NET", "I will search for NET")
def test_reprompt_repeat_respects_word_order(self):
# Set overlap scores a reordered query as identical, so the comparison is
# sequence-based.
from core.inference.tool_call_parser import is_reprompt_repeat
assert not is_reprompt_repeat(
"I will search for dogs not cats", "I will search for cats not dogs"
)
assert is_reprompt_repeat(
"I will search for cats not dogs", "I will search for cats not dogs"
)
assert is_reprompt_repeat("I will search for C++!", "I will search for C++.")
def test_reprompt_repeat_keeps_a_changed_query_token(self):
# One corrected token in a long plan is a new attempt; at the old 0.85 bar it
# scored ~0.87 and cost the model its remaining nudge.
from core.inference.tool_call_parser import is_reprompt_repeat
before = "I will search the web for the latest CUDA version 12.4 driver release notes"
after = "I will search the web for the latest CUDA version 12.5 driver release notes"
assert not is_reprompt_repeat(after, before)
assert is_reprompt_repeat(before, before)
def test_reprompt_repeat_keeps_standalone_operator_tokens(self):
# A marks-only token stripped to nothing, so a bounded correction compared
# equal to the unbounded original.
from core.inference.tool_call_parser import is_reprompt_repeat, is_reprompt_restatement
loose = "Now I think the value is 5"
bounded = "Now I think the value is < 5"
assert not is_reprompt_repeat(bounded, loose)
assert not is_reprompt_restatement(bounded, loose)
def test_reprompt_repeat_keeps_a_changed_token_in_a_long_plan(self):
# Every similarity ratio is length-dependent: one changed token scored 0.98
# across 54 tokens, so long corrected plans lost their nudge.
from core.inference.tool_call_parser import is_reprompt_repeat
words = [f"token{index}" for index in range(54)]
corrected = list(words)
corrected[20] = "revised"
assert not is_reprompt_repeat(" ".join(corrected), " ".join(words))
assert is_reprompt_repeat(" ".join(words), " ".join(words))
def test_reprompt_repeat_keeps_articles_that_name_a_target(self):
# "The Who" and "Who" are different searches, so articles are not filler.
from core.inference.tool_call_parser import is_reprompt_repeat
assert not is_reprompt_repeat(
"I will search for The Who discography",
"I will search for Who discography",
)
def test_reprompt_repeat_keeps_filler_words_that_name_a_target(self):
# No word is reliably filler: dropping "ok"/"the" to absorb rewording also
# absorbed the search target. Reordered filler now reads as a new attempt,
# which costs one nudge out of the cap and never strands a plan.
from core.inference.tool_call_parser import is_reprompt_repeat
assert not is_reprompt_repeat(
"I will search for OK Go discography",
"I will search for Go discography",
)
assert not is_reprompt_repeat(
"I will now summarize the findings",
"I will summarize the findings now",
)
def test_reprompt_repeat_detects_restated_answers(self):
# A nudge answered with the same text again has not worked; stop there.
from core.inference.tool_call_parser import is_reprompt_repeat
same = "I will summarize what I found."
assert is_reprompt_repeat(same, same)
assert is_reprompt_repeat("I WILL summarize what I found!", same)
assert is_reprompt_repeat(
"The summary is ready, please let me know if you need anything else",
"The summary is ready. Please let me know if you need anything else!",
)
# No previous text, or genuinely different progress, keeps the nudge.
assert not is_reprompt_repeat(same, "")
assert not is_reprompt_repeat("Tokyo is 18C and cloudy right now.", same)
# Short texts must not collide on incidental word overlap.
assert not is_reprompt_repeat("Let me check.", "Let me search.")
class TestLoopControl:
def test_cancel_event_breaks_loop(self):
@ -4327,11 +4182,9 @@ class TestPlanWithoutActionReprompt:
# final answer and no further turn is generated.
from core.inference.tool_call_parser import MAX_ACT_REPROMPTS
# Distinct stalls: identical ones stop at the repeat guard, never reaching the cap.
stalls = [f"Let me look into detail {i} first." for i in range(MAX_ACT_REPROMPTS)]
stall = stalls[-1]
stall = "Let me look into it first."
turns = [["I'll search the web for that."]]
turns += [[s] for s in stalls]
turns += [[stall]] * MAX_ACT_REPROMPTS
turns += [["SHOULD NOT APPEAR"]]
generations = {"count": 0}

View file

@ -1,568 +0,0 @@
# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""Per-port PID files, so `unsloth studio stop` can find every server.
Imports run.py directly, so run under the Unsloth venv.
"""
from __future__ import annotations
import os
import sys
from pathlib import Path
from types import SimpleNamespace
import pytest
_BACKEND = Path(__file__).resolve().parents[1]
if str(_BACKEND) not in sys.path:
sys.path.insert(0, str(_BACKEND))
import run # noqa: E402
# Captured before the autouse fixture stubs them, for the tests that exercise them.
_REAL_IS_STUDIO_BACKEND = run._pid_is_studio_backend
_REAL_PID_ALIVE = run._pid_alive
@pytest.fixture(autouse = True)
def isolated_root(tmp_path, monkeypatch):
monkeypatch.setattr(run, "_studio_root", lambda: tmp_path)
monkeypatch.setattr(run, "_PID_FILE", tmp_path / "studio.pid")
monkeypatch.setattr(run, "_OWN_PID_FILE", None)
monkeypatch.setattr(run, "_pid_alive", lambda pid: True)
monkeypatch.setattr(run, "_pid_is_studio_backend", lambda pid, created_times = (): True)
yield
def _files(tmp_path):
return sorted(p.name for p in tmp_path.glob("studio-*.pid"))
def _pid_of(path):
return path.read_text(encoding = "utf-8").splitlines()[0]
def test_write_pid_file_records_port_and_pid(tmp_path):
run._write_pid_file(8901)
assert _files(tmp_path) == [f"studio-8901-{os.getpid()}.pid"]
assert _pid_of(tmp_path / f"studio-8901-{os.getpid()}.pid") == str(os.getpid())
def test_write_pid_file_records_the_start_time(tmp_path):
# Pins the record to this process, so a reused PID isn't mistaken for it.
run._write_pid_file(8901)
record = run._read_pid_record(tmp_path / f"studio-8901-{os.getpid()}.pid")
assert record[0] == os.getpid()
assert record[1] == pytest.approx(run._process_create_time(os.getpid()))
def test_write_pid_file_keeps_the_legacy_file_a_bare_pid(tmp_path):
# An older CLI's `stop` reads studio.pid and expects only digits.
run._write_pid_file(8901)
assert (tmp_path / "studio.pid").read_text(encoding = "utf-8") == str(os.getpid())
def test_second_port_does_not_clobber_the_first(tmp_path):
(tmp_path / "studio-8901-8550.pid").write_text("8550", encoding = "utf-8")
run._write_pid_file(8902)
assert _pid_of(tmp_path / "studio-8901-8550.pid") == "8550"
assert (tmp_path / f"studio-8902-{os.getpid()}.pid").exists()
def test_same_port_on_two_binds_does_not_clobber(tmp_path):
# 127.0.0.1:8888 and ::1:8888 can both listen; one file per port would lose one.
(tmp_path / "studio-8888-8550.pid").write_text("8550", encoding = "utf-8")
run._write_pid_file(8888)
assert len(_files(tmp_path)) == 2
def test_remove_pid_file_only_removes_our_own(tmp_path, monkeypatch):
run._write_pid_file(8901)
(tmp_path / "studio-8902-8600.pid").write_text("8600", encoding = "utf-8")
# Nothing to hand the legacy pointer to, so it goes away with us.
monkeypatch.setattr(run, "_pid_alive", lambda pid: pid == os.getpid())
run._remove_pid_file()
assert _files(tmp_path) == ["studio-8902-8600.pid"]
assert not (tmp_path / "studio.pid").exists()
def test_the_legacy_pointer_moves_to_a_live_sibling(tmp_path):
# Only one server owns studio.pid. Deleting it on our way out would leave an
# older CLI, which reads nothing else, unable to stop the sibling still up.
run._write_pid_file(8901)
(tmp_path / "studio-8902-8600.pid").write_text("8600", encoding = "utf-8")
run._remove_pid_file()
assert (tmp_path / "studio.pid").read_text(encoding = "utf-8").strip() == "8600"
def test_the_legacy_pointer_is_not_handed_to_a_dead_sibling(tmp_path, monkeypatch):
run._write_pid_file(8901)
(tmp_path / "studio-8902-8600.pid").write_text("8600", encoding = "utf-8")
monkeypatch.setattr(run, "_pid_is_studio_backend", lambda pid, created_times = (): False)
run._remove_pid_file()
assert not (tmp_path / "studio.pid").exists()
def test_remove_pid_file_leaves_a_reused_entry_alone(tmp_path):
run._write_pid_file(8901)
own = tmp_path / f"studio-8901-{os.getpid()}.pid"
own.write_text("999999", encoding = "utf-8")
run._remove_pid_file()
assert own.read_text(encoding = "utf-8") == "999999"
def test_windows_liveness_does_not_call_every_pid_alive(monkeypatch):
# os.kill(pid, 0) raises OSError for every pid on Windows, so without the
# tasklist fallback a stale record would block its port forever.
import subprocess
monkeypatch.setattr(run, "_pid_alive", _REAL_PID_ALIVE)
monkeypatch.setitem(sys.modules, "psutil", None)
monkeypatch.setattr(sys, "platform", "win32")
monkeypatch.setattr(
subprocess, "run", lambda *a, **k: SimpleNamespace(stdout = '"python.exe","8550",...')
)
assert run._pid_alive(8550) is True
assert run._pid_alive(9999) is False
def test_windows_liveness_keeps_the_record_when_tasklist_fails(monkeypatch):
# Unconfirmed must mean keep, matching the CLI's _pid_alive. Pruning a live
# server's record lets the next launch fall back past it and strand it, which
# is the bug this file exists to fix; a stale record costs one clear abort.
import subprocess
def _boom(*a, **k):
raise OSError("tasklist missing")
monkeypatch.setattr(run, "_pid_alive", _REAL_PID_ALIVE)
monkeypatch.setitem(sys.modules, "psutil", None)
monkeypatch.setattr(sys, "platform", "win32")
monkeypatch.setattr(subprocess, "run", _boom)
assert run._pid_alive(8550) is True
def test_read_pid_record_parses_pid_time_and_address(tmp_path):
(tmp_path / "r.pid").write_text("8550\n111.5\n127.0.0.1", encoding = "utf-8")
assert run._read_pid_record(tmp_path / "r.pid") == (8550, 111.5, "127.0.0.1")
def test_read_pid_record_tolerates_a_bare_pid(tmp_path):
(tmp_path / "r.pid").write_text("8550", encoding = "utf-8")
assert run._read_pid_record(tmp_path / "r.pid") == (8550, None, None)
def test_read_pid_record_rejects_pid_zero_and_init(tmp_path):
# kill(0) signals our whole process group.
(tmp_path / "zero.pid").write_text("0", encoding = "utf-8")
(tmp_path / "init.pid").write_text("1", encoding = "utf-8")
assert run._read_pid_record(tmp_path / "zero.pid") is None
assert run._read_pid_record(tmp_path / "init.pid") is None
def test_read_pid_record_rejects_a_corrupt_file(tmp_path):
(tmp_path / "r.pid").write_text("not-a-pid", encoding = "utf-8")
assert run._read_pid_record(tmp_path / "r.pid") is None
def test_graceful_shutdown_drops_the_record_last(monkeypatch):
# Cleanup can take seconds while the server is still alive. Dropping the record
# first leaves a retried `stop` or a new launch unable to find it.
order = []
monkeypatch.setattr(run, "_remove_pid_file", lambda: order.append("remove_record"))
class _Server:
def __setattr__(self, name, value):
order.append("release_socket")
run._graceful_shutdown(_Server())
assert order == ["release_socket", "remove_record"]
def test_own_studio_on_port_is_found_without_psutil(tmp_path, monkeypatch):
# psutil is optional; a listener scan finds nothing without it, so detection
# must come from our own records or we silently start a duplicate.
monkeypatch.setitem(sys.modules, "psutil", None)
(tmp_path / "studio-8901-8550.pid").write_text("8550\n\n127.0.0.1", encoding = "utf-8")
assert run._own_studio_on_port(8901, "127.0.0.1") == 8550
def test_no_record_for_the_port_means_no_own_studio(tmp_path):
# jupyter-lab on 8888 must keep the fallback, not abort the launch.
(tmp_path / "studio-8901-8550.pid").write_text("8550", encoding = "utf-8")
assert run._own_studio_on_port(8888, "127.0.0.1") is None
def test_own_studio_on_port_prunes_a_dead_record(tmp_path, monkeypatch):
monkeypatch.setattr(run, "_pid_alive", lambda pid: False)
(tmp_path / "studio-8901-8550.pid").write_text("8550", encoding = "utf-8")
assert run._own_studio_on_port(8901, "127.0.0.1") is None
assert not (tmp_path / "studio-8901-8550.pid").exists()
def test_a_reused_pid_is_not_treated_as_our_studio(tmp_path, monkeypatch):
# Stale record + the OS handing that PID to something else must not abort.
monkeypatch.setattr(run, "_pid_is_studio_backend", lambda pid, created_times = (): False)
(tmp_path / "studio-8901-8550.pid").write_text("8550", encoding = "utf-8")
assert run._own_studio_on_port(8901, "127.0.0.1") is None
def test_an_unverifiable_record_still_blocks_a_duplicate(tmp_path, monkeypatch):
# Can't tell: refusing with a clear message beats a silent second instance.
monkeypatch.setattr(run, "_pid_is_studio_backend", lambda pid, created_times = (): True)
(tmp_path / "studio-8901-8550.pid").write_text("8550", encoding = "utf-8")
assert run._own_studio_on_port(8901, "127.0.0.1") == 8550
def test_start_time_mismatch_rejects_a_reused_pid(monkeypatch):
monkeypatch.setattr(run, "_pid_is_studio_backend", _REAL_IS_STUDIO_BACKEND)
monkeypatch.setattr(run, "_process_create_time", lambda pid: 999.0)
assert run._pid_is_studio_backend(8550, [111.5]) is False
assert run._pid_is_studio_backend(8550, [999.0]) is True
def test_a_stale_record_does_not_veto_a_live_server_sharing_the_pid(monkeypatch):
# Crash leaves studio-8888-1234.pid, the OS reuses 1234 for a new server on
# another port. Keeping only the first timestamp would reject the live one.
monkeypatch.setattr(run, "_pid_is_studio_backend", _REAL_IS_STUDIO_BACKEND)
monkeypatch.setattr(run, "_process_create_time", lambda pid: 999.0)
assert run._pid_is_studio_backend(1234, [111.5, 999.0]) is True
assert run._pid_is_studio_backend(1234, [111.5, 222.5]) is False
def test_a_stale_record_on_another_port_does_not_hide_a_live_server(tmp_path, monkeypatch):
# 1234 was reused: the stale 8888 record must not stop us seeing 9000.
monkeypatch.setattr(run, "_pid_is_studio_backend", _REAL_IS_STUDIO_BACKEND)
monkeypatch.setattr(run, "_process_create_time", lambda pid: 999.0)
(tmp_path / "studio-8888-1234.pid").write_text("1234\n111.5\n", encoding = "utf-8")
(tmp_path / "studio-9000-1234.pid").write_text("1234\n999.0\n", encoding = "utf-8")
assert run._own_studio_on_port(8888, "127.0.0.1") is None
assert run._own_studio_on_port(9000, "127.0.0.1") == 1234
def test_a_start_time_is_the_only_thing_that_disproves_a_record(monkeypatch):
monkeypatch.setattr(run, "_pid_is_studio_backend", _REAL_IS_STUDIO_BACKEND)
monkeypatch.setattr(run, "_process_create_time", lambda pid: 999.0)
assert run._pid_is_studio_backend(8550, [999.0]) is True
assert run._pid_is_studio_backend(8550, [111.5]) is False
def test_a_bare_run_py_command_line_is_not_rejected(monkeypatch):
# `cd studio/backend && python run.py --port 8901` has no "studio" or "unsloth"
# in argv. Guessing from the command line called that "not ours".
monkeypatch.setattr(run, "_pid_is_studio_backend", _REAL_IS_STUDIO_BACKEND)
class _FakeProcess:
def __init__(self, pid):
self.pid = pid
def cmdline(self):
return ["python", "run.py", "--port", "8901"]
def create_time(self):
return 111.5
monkeypatch.setitem(sys.modules, "psutil", SimpleNamespace(Process = _FakeProcess))
assert run._pid_is_studio_backend(8550) is True
def test_an_untimed_legacy_record_is_trusted(monkeypatch):
# `python run.py --port 8901` has no telltale argv, so guessing from the
# command line rejected real servers. Only a start time can disprove one.
monkeypatch.setattr(run, "_pid_is_studio_backend", _REAL_IS_STUDIO_BACKEND)
monkeypatch.setattr(run, "_process_create_time", lambda pid: 999.0)
assert run._pid_is_studio_backend(8550) is True
assert run._pid_is_studio_backend(8550, [None]) is True
def test_the_untimed_legacy_record_does_not_cancel_a_timed_one(monkeypatch):
# Mirrors _pid_is_studio_server in the CLI. An untimed record carries no
# information, so it must not overrule a start time that says "not ours" --
# every current server writes one of each, which made the check inert.
monkeypatch.setattr(run, "_pid_is_studio_backend", _REAL_IS_STUDIO_BACKEND)
monkeypatch.setattr(run, "_process_create_time", lambda pid: 999.0)
assert run._pid_is_studio_backend(8550, [111.5, None]) is False
assert run._pid_is_studio_backend(8550, [111.5, 999.0]) is True
def test_a_legacy_server_on_the_port_is_recognised(tmp_path, monkeypatch):
# Pre-upgrade servers wrote only studio.pid. Falling back past one strands it
# and then overwrites its record.
monkeypatch.setattr(run, "_get_pid_on_port", lambda p: (8550, "python"))
(tmp_path / "studio.pid").write_text("8550", encoding = "utf-8")
assert run._own_studio_on_port(8901, "127.0.0.1") == 8550
def test_a_legacy_record_for_a_different_listener_falls_back(tmp_path, monkeypatch):
# jupyter holds the port; the legacy server is elsewhere. Keep falling back.
monkeypatch.setattr(run, "_get_pid_on_port", lambda p: (117, "jupyter-lab"))
(tmp_path / "studio.pid").write_text("8550", encoding = "utf-8")
assert run._own_studio_on_port(8901, "127.0.0.1") is None
def test_an_unknowable_listener_treats_the_legacy_record_as_ours(tmp_path, monkeypatch):
# No psutil: _get_pid_on_port can't say. Refusing beats a silent duplicate.
monkeypatch.setattr(run, "_get_pid_on_port", lambda p: None)
(tmp_path / "studio.pid").write_text("8550", encoding = "utf-8")
assert run._own_studio_on_port(8901, "127.0.0.1") == 8550
def test_a_dead_legacy_record_falls_back(tmp_path, monkeypatch):
monkeypatch.setattr(run, "_pid_alive", lambda pid: False)
monkeypatch.setattr(run, "_get_pid_on_port", lambda p: None)
(tmp_path / "studio.pid").write_text("8550", encoding = "utf-8")
assert run._own_studio_on_port(8901, "127.0.0.1") is None
def test_a_stale_per_port_record_does_not_mask_a_legacy_server(tmp_path, monkeypatch):
# Crashed current build left studio-8901-8550.pid; 8550 was then reused by a
# pre-upgrade server recorded only in studio.pid. The stale record must not
# count as "port already known" and send us falling back past the live one.
monkeypatch.setattr(run, "_pid_is_studio_backend", _REAL_IS_STUDIO_BACKEND)
monkeypatch.setattr(run, "_process_create_time", lambda pid: 999.0)
monkeypatch.setattr(run, "_get_pid_on_port", lambda p: (8550, "python"))
(tmp_path / "studio-8901-8550.pid").write_text("8550\n111.5\n127.0.0.1", encoding = "utf-8")
(tmp_path / "studio.pid").write_text("8550", encoding = "utf-8")
assert run._own_studio_on_port(8901, "127.0.0.1") == 8550
def test_a_current_server_elsewhere_does_not_block_a_foreign_port(tmp_path, monkeypatch):
# Current builds write studio.pid too. Without psutil the legacy check can't
# see the listener, so it must not claim our 8901 server holds jupyter's 8888.
monkeypatch.setattr(run, "_get_pid_on_port", lambda p: None)
(tmp_path / "studio-8901-5000.pid").write_text("5000\n\n127.0.0.1", encoding = "utf-8")
(tmp_path / "studio.pid").write_text("5000", encoding = "utf-8")
assert run._own_studio_on_port(8888, "127.0.0.1") is None
def test_a_per_port_record_is_preferred_over_the_legacy_one(tmp_path, monkeypatch):
monkeypatch.setattr(run, "_get_pid_on_port", lambda p: (8550, "python"))
(tmp_path / "studio-8901-8600.pid").write_text("8600\n\n127.0.0.1", encoding = "utf-8")
(tmp_path / "studio.pid").write_text("8550", encoding = "utf-8")
assert run._own_studio_on_port(8901, "127.0.0.1") == 8600
def test_our_studio_on_another_bind_address_does_not_abort(tmp_path):
# Our server holds ::1:8889; binding 127.0.0.1:8889 is not a conflict with us,
# so fall through to the next port instead of refusing.
(tmp_path / "studio-8889-8550.pid").write_text("8550\n\n::1", encoding = "utf-8")
assert run._own_studio_on_port(8889, "127.0.0.1") is None
assert run._own_studio_on_port(8889, "::1") == 8550
def test_address_matching(tmp_path):
assert run._addresses_collide("0.0.0.0", "127.0.0.1", 8889) is True
assert run._addresses_collide("127.0.0.1", "0.0.0.0", 8889) is True
assert run._addresses_collide("127.0.0.1", "127.0.0.1", 8889) is True
assert run._addresses_collide("::1", "127.0.0.1", 8889) is False
# An unrecorded address is unknown, so assume a conflict.
assert run._addresses_collide(None, "127.0.0.1", 8889) is True
def test_a_hostname_resolves_the_same_way_the_bind_does(tmp_path):
# `localhost` and the address _is_port_free actually binds must agree, or a
# recorded server is missed and a duplicate starts.
recorded = ",".join(sorted(run._bind_addresses("localhost", 8889)))
assert run._addresses_collide(recorded, "localhost", 8889) is True
def test_a_hostname_records_every_address_it_resolves_to(tmp_path):
# `localhost` binds 127.0.0.1 AND ::1. Recording only the first lets a later
# launch on the other literal miss us and start a duplicate.
addrs = run._bind_addresses("localhost", 8889)
recorded = ",".join(sorted(addrs))
for literal in addrs:
assert run._addresses_collide(recorded, literal, 8889) is True
def test_a_multi_address_record_matches_either_literal(tmp_path):
recorded = "127.0.0.1,::1"
assert run._addresses_collide(recorded, "127.0.0.1", 8889) is True
assert run._addresses_collide(recorded, "::1", 8889) is True
assert run._addresses_collide("127.0.0.1", "::1", 8889) is False
def test_fallback_aborts_on_our_own_server_further_up_the_range(tmp_path, monkeypatch):
# jupyter holds 8888, our server holds 8889: skipping to 8890 is the duplicate.
(tmp_path / "studio-8889-8550.pid").write_text("8550\n\n127.0.0.1", encoding = "utf-8")
monkeypatch.setattr(run, "_is_port_free", lambda host, p: p >= 8890)
with pytest.raises(SystemExit) as excinfo:
run._find_free_port("127.0.0.1", 8889, avoid_own_studio = True)
assert excinfo.value.code == 1
def test_fallback_still_skips_foreign_processes(tmp_path, monkeypatch):
# No record for 8889, so the blocker is not ours: keep falling back.
monkeypatch.setattr(run, "_is_port_free", lambda host, p: p >= 8890)
assert run._find_free_port("127.0.0.1", 8889, avoid_own_studio = True) == 8890
def test_the_requested_port_is_kept_when_it_is_free(monkeypatch):
monkeypatch.setattr(run, "_is_port_free", lambda host, p: True)
assert run._resolve_port("127.0.0.1", 8888) == 8888
def test_our_own_server_on_the_requested_port_aborts_rather_than_falling_back(
tmp_path, monkeypatch
):
# The reported bug: 8888 is ours, so falling back to 8889 is the duplicate
# that leaves 8888 serving with nothing recording it.
monkeypatch.setattr(run, "_is_port_free", lambda host, p: p != 8888)
(tmp_path / "studio-8888-8550.pid").write_text("8550\n\n127.0.0.1", encoding = "utf-8")
with pytest.raises(SystemExit) as excinfo:
run._resolve_port("127.0.0.1", 8888)
assert excinfo.value.code == 1
def test_a_foreign_process_on_the_requested_port_still_falls_back(monkeypatch):
# jupyter-lab on 8888 must not stop Unsloth starting on 8889.
monkeypatch.setattr(run, "_is_port_free", lambda host, p: p != 8888)
assert run._resolve_port("127.0.0.1", 8888) == 8889
def test_a_caller_that_reads_the_port_back_keeps_the_plain_fallback(tmp_path, monkeypatch):
# api-only callers (the desktop app via TAURI_PORT, `studio run` via
# app.state.server_port) follow us to the new port, so aborting there only
# turns a working launch into a crash the desktop app reports as "stopped
# unexpectedly". Both servers are still recorded, so `stop` finds them.
monkeypatch.setattr(run, "_is_port_free", lambda host, p: p != 8888)
(tmp_path / "studio-8888-8550.pid").write_text("8550\n\n127.0.0.1", encoding = "utf-8")
assert run._resolve_port("127.0.0.1", 8888, avoid_own_studio = False) == 8889
def test_the_recorded_address_is_every_address_the_bind_resolves_to(tmp_path):
# The only test that runs the writer with a real host. Recording `host`
# verbatim, or dropping the line, passes every other test here and silently
# stops matching a launch that spells the same interface differently.
run._write_pid_file(8901, "localhost")
record = run._read_pid_record(tmp_path / f"studio-8901-{os.getpid()}.pid")
assert record[2] is not None, "no bind address recorded"
assert set(record[2].split(",")) == run._bind_addresses("localhost", 8901)
def test_a_server_started_on_a_hostname_is_found_again_by_ip(tmp_path):
run._write_pid_file(8901, "localhost")
for literal in run._bind_addresses("localhost", 8901):
assert run._own_studio_on_port(8901, literal) == os.getpid()
def test_bind_addresses_keeps_every_family_a_hostname_resolves_to(monkeypatch):
# Independent oracle: the sibling test derives its expectation from this
# function's own output, so dropping a family would pass it.
import socket
monkeypatch.setattr(
socket,
"getaddrinfo",
lambda *a, **k: [
(socket.AF_INET, socket.SOCK_STREAM, 6, "", ("127.0.0.1", 8889)),
(socket.AF_INET6, socket.SOCK_STREAM, 6, "", ("::1", 8889, 0, 0)),
],
)
assert run._bind_addresses("localhost", 8889) == {"127.0.0.1", "::1"}
def test_the_legacy_file_is_written_even_when_the_per_port_record_fails(tmp_path, monkeypatch):
# A studio root that cannot take a new entry used to leave the server
# recorded nowhere at all, so the CLI could not stop it. studio.pid is an
# overwrite of an existing path, so it can still succeed and must be tried.
blocked = tmp_path / "not-a-directory"
blocked.write_text("", encoding = "utf-8")
monkeypatch.setattr(
run, "_pid_file_for_port", lambda port: blocked / f"studio-{port}-{os.getpid()}.pid"
)
run._write_pid_file(8901, "127.0.0.1")
assert (tmp_path / "studio.pid").read_text(encoding = "utf-8") == str(os.getpid())
assert run._OWN_PID_FILE is None
def test_a_record_whose_pid_is_not_ascii_digits_is_discarded(tmp_path):
# A superscript two passes isdigit() but int() rejects it, so that gate alone
# let a ValueError escape into every caller of _read_pid_record.
(tmp_path / "r.pid").write_text("²", encoding = "utf-8")
assert run._read_pid_record(tmp_path / "r.pid") is None
def test_the_legacy_file_is_not_taken_from_a_live_server(tmp_path):
# A pre-upgrade server is recorded in studio.pid and nowhere else, so a
# second launch overwriting it is exactly what strands it. That is the
# orphan this file exists to prevent, reached from the other direction.
(tmp_path / "studio.pid").write_text("8550", encoding = "utf-8")
run._write_pid_file(8902, "127.0.0.1")
assert (tmp_path / "studio.pid").read_text(encoding = "utf-8") == "8550"
assert (tmp_path / f"studio-8902-{os.getpid()}.pid").exists()
def test_the_legacy_file_is_taken_over_from_a_dead_server(tmp_path, monkeypatch):
# A stale record must not keep the pointer forever, or an older CLI could
# never stop anything again.
monkeypatch.setattr(run, "_pid_alive", lambda pid: False)
(tmp_path / "studio.pid").write_text("8550", encoding = "utf-8")
run._write_pid_file(8902, "127.0.0.1")
assert (tmp_path / "studio.pid").read_text(encoding = "utf-8") == str(os.getpid())

View file

@ -1,809 +0,0 @@
# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""Text I/O must name its encoding, or Windows silently uses the ANSI codepage.
``open()``, ``Path.read_text()`` and ``subprocess(text = True)`` fall back to
``locale.getencoding()`` when no ``encoding`` is passed. On Windows that is
cp1252 (or cp932, cp1251, ... by system locale), not UTF-8, so a chat template,
model config or path containing ``ä ö ü `` mojibakes or raises
``UnicodeDecodeError`` mid-load. Studio's files are UTF-8, so say so.
"""
from __future__ import annotations
import ast
import importlib.util
import json
import os
from pathlib import Path
from types import SimpleNamespace
import pytest
BACKEND_ROOT = Path(__file__).resolve().parent.parent
# Not runtime source. Shipped plugins under plugins/*/src are, so only builds are skipped.
_SKIPPED_DIRS = ("node_modules", "build", "tests", "__pycache__")
# Path.open()'s signature is what tells it apart from other libraries' open(),
# e.g. fitz.open(stream=...) and av.open(..., metadata_errors=...).
_FILE_MODE_CHARS = set("rwxabt+")
_PATH_OPEN_ARGS = ("mode", "buffering", "encoding", "errors", "newline")
_PATH_OPEN_KWARGS = set(_PATH_OPEN_ARGS)
_PATH_OPEN_ENCODING_ARG = _PATH_OPEN_ARGS.index("encoding")
_SUBPROCESS_CALLS = {"run", "Popen", "check_output", "check_call", "call"}
# open(file, mode, buffering, encoding, ...), and os.fdopen forwards the same
# signature with a descriptor in place of the path.
_OPEN_ENCODING_ARG = 3
def _studio_sources() -> list[Path]:
return [
path
for path in sorted(BACKEND_ROOT.rglob("*.py"))
if not any(part in _SKIPPED_DIRS for part in path.relative_to(BACKEND_ROOT).parts)
]
def _has_keyword(node: ast.Call, name: str) -> bool:
return any(keyword.arg == name for keyword in node.keywords)
def _mode_is_binary(node: ast.Call) -> bool:
mode: str | None = None
if len(node.args) >= 2 and isinstance(node.args[1], ast.Constant):
value = node.args[1].value
mode = value if isinstance(value, str) else None
for keyword in node.keywords:
if keyword.arg == "mode" and isinstance(keyword.value, ast.Constant):
value = keyword.value.value
if isinstance(value, str):
mode = value
return bool(mode and "b" in mode)
def _open_has_encoding(node: ast.Call) -> bool:
"""open()/os.fdopen() also take encoding positionally: open(p, "w", 1, "utf-8")."""
return _has_keyword(node, "encoding") or len(node.args) > _OPEN_ENCODING_ARG
def _path_open_mode(node: ast.Call) -> str | None:
if node.args and isinstance(node.args[0], ast.Constant):
value = node.args[0].value
if isinstance(value, str):
return value
for keyword in node.keywords:
if keyword.arg == "mode" and isinstance(keyword.value, ast.Constant):
value = keyword.value.value
if isinstance(value, str):
return value
return None
def _is_path_open(node: ast.Call) -> bool:
"""True only for calls matching ``Path.open``'s signature."""
if len(node.args) > len(_PATH_OPEN_ARGS):
return False
if any(k.arg not in _PATH_OPEN_KWARGS for k in node.keywords):
return False
mode = _path_open_mode(node)
if mode is not None:
return bool(mode) and set(mode) <= _FILE_MODE_CHARS
return not node.args
def _path_open_has_encoding(node: ast.Call) -> bool:
"""Path.open() also takes encoding positionally: open("w", 1, "utf-8")."""
return _has_keyword(node, "encoding") or len(node.args) > _PATH_OPEN_ENCODING_ARG
def _call_name(node: ast.Call) -> str | None:
func = node.func
if isinstance(func, ast.Name):
return func.id
if isinstance(func, ast.Attribute):
return func.attr
return None
def _subprocess_names(tree: ast.AST) -> set[str]:
"""Names subprocess is reachable under here, e.g. `import subprocess as _sp`."""
names = set()
for node in ast.walk(tree):
if isinstance(node, ast.Import):
for alias in node.names:
if alias.name == "subprocess":
names.add(alias.asname or alias.name)
return names
def _subprocess_aliases(tree: ast.AST, names: set[str]) -> set[str]:
"""Plain names bound to a subprocess callable, called without the module.
``install_wheel(run = subprocess.run)`` calls its injected ``run`` as a bare
name, so matching only the attribute form leaves those installer calls
unguarded. Imports, assignments and parameter defaults all bind one.
"""
def _is_bound(value: ast.expr | None) -> bool:
return (
isinstance(value, ast.Attribute)
and value.attr in _SUBPROCESS_CALLS
and isinstance(value.value, ast.Name)
and value.value.id in names
)
aliases: set[str] = set()
for node in ast.walk(tree):
if isinstance(node, ast.ImportFrom) and node.module == "subprocess":
aliases.update(a.asname or a.name for a in node.names if a.name in _SUBPROCESS_CALLS)
elif isinstance(node, ast.Assign) and _is_bound(node.value):
aliases.update(t.id for t in node.targets if isinstance(t, ast.Name))
elif isinstance(node, ast.AnnAssign) and _is_bound(node.value):
if isinstance(node.target, ast.Name):
aliases.add(node.target.id)
elif isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)):
args = node.args
positional = args.posonlyargs + args.args
# Defaults cover the tail of the positional parameters; kw_defaults
# is aligned with kwonlyargs already, holding None where absent.
padded = [None] * (len(positional) - len(args.defaults)) + list(args.defaults)
pairs = list(zip(positional, padded)) + list(zip(args.kwonlyargs, args.kw_defaults))
aliases.update(arg.arg for arg, default in pairs if _is_bound(default))
return aliases
def _is_subprocess_call(node: ast.Call, names: set[str], aliases: set[str]) -> bool:
func = node.func
if isinstance(func, ast.Name):
return func.id in aliases
if not isinstance(func, ast.Attribute) or func.attr not in _SUBPROCESS_CALLS:
return False
value = func.value
return isinstance(value, ast.Name) and value.id in names
def _text_mode_subprocess(node: ast.Call) -> bool:
for keyword in node.keywords:
if keyword.arg not in ("text", "universal_newlines"):
continue
if isinstance(keyword.value, ast.Constant) and keyword.value.value is True:
return True
return False
def _text_mode_dict(node: ast.Dict) -> bool:
"""A ``{"text": True, ...}`` literal with no "encoding" key."""
keys = [k.value for k in node.keys if isinstance(k, ast.Constant)]
if "encoding" in keys:
return False
for key, value in zip(node.keys, node.values):
if not isinstance(key, ast.Constant) or key.value not in (
"text",
"universal_newlines",
):
continue
if isinstance(value, ast.Constant) and value.value is True:
return True
return False
def _splatted_names(tree: ast.AST) -> set[str]:
"""Names handed to a call as ``**name``."""
names = set()
for node in ast.walk(tree):
if isinstance(node, ast.Call):
for keyword in node.keywords:
if keyword.arg is None and isinstance(keyword.value, ast.Name):
names.add(keyword.value.id)
return names
def _encoding_assigned_later(tree: ast.AST, name: str) -> bool:
"""``name["encoding"] = ...`` somewhere, so the literal need not carry it."""
for node in ast.walk(tree):
if not isinstance(node, ast.Subscript) or not isinstance(node.ctx, ast.Store):
continue
target, key = node.value, node.slice
if isinstance(target, ast.Name) and target.id == name:
if isinstance(key, ast.Constant) and key.value == "encoding":
return True
return False
def _splatted_kwargs_offenders(tree: ast.AST) -> list[ast.Dict]:
"""Text-mode kwargs built in a dict and splatted into a call.
Kwargs are collected in a dict and splatted (``run(cmd, **run_kwargs)``)
where a branch has to add a timeout or an env, and the call is often through
a helper, so neither the callee nor the keywords are visible at the call
site. Only dicts that reach a call this way are judged: an unrelated payload
that happens to carry ``"text": True`` is not subprocess configuration.
"""
found = []
# ``run(cmd, **{...})``: the literal is at the call already.
for node in ast.walk(tree):
if not isinstance(node, ast.Call):
continue
for keyword in node.keywords:
if keyword.arg is None and isinstance(keyword.value, ast.Dict):
if _text_mode_dict(keyword.value):
found.append(keyword.value)
splatted = _splatted_names(tree)
if not splatted:
return found
for node in ast.walk(tree):
targets = []
if isinstance(node, ast.Assign):
targets = [t for t in node.targets if isinstance(t, ast.Name)]
elif isinstance(node, ast.AnnAssign) and isinstance(node.target, ast.Name):
targets = [node.target]
if not targets or not isinstance(node.value, ast.Dict):
continue
if not _text_mode_dict(node.value):
continue
for target in targets:
if target.id in splatted and not _encoding_assigned_later(tree, target.id):
found.append(node.value)
break
return found
def _offenders(path: Path) -> list[str]:
source = path.read_text(encoding = "utf-8")
tree = ast.parse(source, filename = str(path))
subprocess_names = _subprocess_names(tree)
subprocess_aliases = _subprocess_aliases(tree, subprocess_names)
found: list[str] = []
for node in _splatted_kwargs_offenders(tree):
found.append(
f"{path.name}:{node.lineno}: subprocess kwargs with text = True and no encoding"
)
for node in ast.walk(tree):
if not isinstance(node, ast.Call):
continue
name = _call_name(node)
if _is_subprocess_call(node, subprocess_names, subprocess_aliases):
if _text_mode_subprocess(node) and not _has_keyword(node, "encoding"):
found.append(f"{path.name}:{node.lineno}: subprocess(text = True) without encoding")
continue
if name == "open" and isinstance(node.func, ast.Name):
if _mode_is_binary(node) or _open_has_encoding(node):
continue
found.append(f"{path.name}:{node.lineno}: open() without encoding")
continue
# os.fdopen(fd, "w") is open() on a descriptor, so text mode takes the
# same locale default. Its mode defaults to "r", i.e. text, like open's.
if name == "fdopen":
if _mode_is_binary(node) or _open_has_encoding(node):
continue
found.append(f"{path.name}:{node.lineno}: os.fdopen() without encoding")
continue
if name == "open" and isinstance(node.func, ast.Attribute):
if not _is_path_open(node) or _path_open_has_encoding(node):
continue
if _path_open_mode(node) and "b" in _path_open_mode(node):
continue
found.append(f"{path.name}:{node.lineno}: Path.open() without encoding")
continue
if name in ("read_text", "write_text") and isinstance(node.func, ast.Attribute):
if _has_keyword(node, "encoding"):
continue
# importlib.metadata Distribution.read_text() takes no encoding kwarg.
if isinstance(node.func.value, ast.Name) and node.func.value.id == "dist":
continue
found.append(f"{path.name}:{node.lineno}: {name}() without encoding")
return found
@pytest.mark.parametrize("path", _studio_sources(), ids = lambda p: str(p.name))
def test_text_io_names_its_encoding(path: Path) -> None:
offenders = _offenders(path)
assert not offenders, (
"Text I/O without an explicit encoding falls back to the Windows ANSI "
'codepage and corrupts non-ASCII (ä ö ü → 世). Pass encoding = "utf-8":\n '
+ "\n ".join(offenders)
)
_STATE_STORE = (
BACKEND_ROOT
/ "plugins/data-designer-github-repo-seed/src"
/ "data_designer_github_repo_seed/scraper_impl/state_store.py"
)
def _load_state_store(codepage: str):
"""Load state_store with the writing machine's codepage pinned."""
spec = importlib.util.spec_from_file_location(f"state_store_{codepage}", _STATE_STORE)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
module.locale = SimpleNamespace(
getencoding = lambda: codepage,
getpreferredencoding = lambda _ = True: codepage,
)
return module
@pytest.mark.parametrize(
("codepage", "name"), [("cp1252", "Jürgen"), ("cp1251", "Юрий"), ("cp932", "田中")]
)
def test_resuming_a_legacy_jsonl_keeps_one_encoding(
tmp_path: Path, codepage: str, name: str
) -> None:
"""A scrape written before UTF-8 was explicit must resume, not duplicate."""
path = tmp_path / "out.jsonl"
records = [{"id": 1, "author": name}, {"id": 2, "author": name}]
body = "".join(json.dumps(r, ensure_ascii = False) + "\n" for r in records)
path.write_bytes(body.encode(codepage))
before = path.read_bytes()
writer = _load_state_store(codepage).JsonlWriter(path)
try:
# Seen keys survive the resume, so a repeat is refused, not appended.
assert writer.has("id:1") and writer.has("id:2")
assert writer.write(records[0]) is False
assert writer.write({"id": 3, "author": name}) is True
finally:
writer.close()
# Never converted, so it still reads in its own codepage; the append is ASCII.
blob = path.read_bytes()
assert blob.startswith(before)
assert blob[len(before) :].isascii()
lines = [json.loads(x) for x in blob.decode(codepage).splitlines() if x.strip()]
assert len(lines) == 3
assert [line["author"] for line in lines] == [name] * 3
def test_a_coincidentally_utf8_legacy_line_is_left_alone(tmp_path: Path) -> None:
"""cp1251 `Р°` is D0 B0, which is also UTF-8 `а`, and nothing can tell them apart."""
path = tmp_path / "out.jsonl"
ambiguous = "Р°"
assert ambiguous.encode("cp1251").decode("utf-8") == "а" # the trap
authors = ["Привет", "Здравствуйте", "Москва", ambiguous]
path.write_bytes(
b"".join(
json.dumps({"id": i, "author": a}, ensure_ascii = False).encode("cp1251") + b"\n"
for i, a in enumerate(authors)
)
)
before = path.read_bytes()
_load_state_store("cp1251").JsonlWriter(path).close()
# Untouched, so the ambiguity never had to be resolved.
assert path.read_bytes() == before
rows = [json.loads(x) for x in path.read_text(encoding = "cp1251").splitlines() if x.strip()]
assert [row["author"] for row in rows] == authors
@pytest.mark.parametrize(
("codepage", "word"), [("cp1251", "Привет"), ("cp932", "こんにちは"), ("cp1252", "Jürgen")]
)
def test_a_moved_shard_is_not_rewritten_by_guesswork(
tmp_path: Path, codepage: str, word: str
) -> None:
"""Off the writing machine there is no codepage to attribute the file to."""
path = tmp_path / "out.jsonl"
# Two records: a lone non-UTF-8 line would count as damage, not legacy.
path.write_bytes(
b"".join(
json.dumps({"id": i, "author": word}, ensure_ascii = False).encode(codepage) + b"\n"
for i in (1, 4)
)
)
before = path.read_bytes()
# A UTF-8 host: latin-1 would read cp1251 `Привет` back as `Ïðèâåò`.
writer = _load_state_store("utf-8").JsonlWriter(path)
try:
assert writer.has("id:1") # ASCII keys still recover
assert writer.write({"id": 2, "author": "Grüße"}) is True
finally:
writer.close()
blob = path.read_bytes()
assert blob.startswith(before) # never rewritten
assert blob[len(before) :].isascii() # appended as \uXXXX, so no second encoding
rows = [json.loads(x) for x in blob.decode(codepage).splitlines() if x.strip()]
assert [row["author"] for row in rows] == [word, word, "Grüße"]
def test_an_all_ambiguous_shard_still_gets_ascii_appends(tmp_path: Path) -> None:
"""Every line valid under both readings still means the append must not pick one."""
path = tmp_path / "out.jsonl"
ambiguous = "Р°" # cp1251 D0 B0, also valid UTF-8 for "а"
path.write_bytes(
b"".join(
json.dumps({"id": i, "a": ambiguous}, ensure_ascii = False).encode("cp1251") + b"\n"
for i in range(3)
)
)
before = path.read_bytes()
writer = _load_state_store("cp1251").JsonlWriter(path)
try:
assert writer.write({"id": 9, "a": "世界"}) is True
finally:
writer.close()
blob = path.read_bytes()
assert blob.startswith(before)
# ASCII, so the appended record survives whichever reading is chosen.
assert blob[len(before) :].isascii()
for codec in ("cp1251", "utf-8"):
rows = [json.loads(x) for x in blob.decode(codec).splitlines() if x.strip()]
assert rows[-1]["a"] == "世界"
def test_a_damaged_line_in_an_ascii_shard_does_not_block_its_retry(tmp_path: Path) -> None:
"""With no non-ASCII records to outvote it, one damaged line is still damage."""
path = tmp_path / "out.jsonl"
path.write_bytes(
b'{"id": 1, "author": "alice"}\n'
+ b'{"id": 99, "author": "bad \x96 byte"}\n'
+ b'{"id": 2, "author": "bob"}\n'
)
writer = _load_state_store("cp1252").JsonlWriter(path)
try:
assert writer.has("id:1") and writer.has("id:2")
assert not writer.has("id:99")
assert writer.write({"id": 99, "author": "good byte"}) is True
finally:
writer.close()
def test_a_damaged_line_does_not_block_its_own_retry(tmp_path: Path) -> None:
"""Its key comes from the codepage reading, which a UTF-8 shard did not pick."""
path = tmp_path / "out.jsonl"
path.write_bytes(
json.dumps({"id": 1, "author": "Jürgen"}, ensure_ascii = False).encode()
+ b"\n"
+ b'{"id": 99, "author": "bad \x96 byte"}\n'
)
writer = _load_state_store("cp1252").JsonlWriter(path)
try:
assert writer.has("id:1")
assert not writer.has("id:99")
assert writer.write({"id": 99, "author": "good byte"}) is True
finally:
writer.close()
def test_one_damaged_byte_does_not_relabel_a_utf8_shard(tmp_path: Path) -> None:
"""A complete JSON line with a stray 0x96 parses as cp1252, but is only one vote."""
path = tmp_path / "out.jsonl"
healthy = ["Jürgen", "Grüße", "Björn"]
path.write_bytes(
json.dumps({"id": 0, "author": healthy[0]}, ensure_ascii = False).encode()
+ b"\n"
+ b'{"id": 99, "author": "bad \x96 byte"}\n'
+ b"".join(
json.dumps({"id": i, "author": a}, ensure_ascii = False).encode() + b"\n"
for i, a in enumerate(healthy[1:], start = 1)
)
)
before = path.read_bytes()
_load_state_store("cp1252").JsonlWriter(path).close()
# Untouched, so the healthy records were never re-read as cp1252.
assert path.read_bytes() == before
rows = []
for line in path.read_bytes().splitlines():
try:
rows.append(json.loads(line.decode()))
except (UnicodeDecodeError, ValueError):
continue
assert [row["author"] for row in rows] == healthy
def test_a_torn_line_does_not_relabel_a_utf8_shard(tmp_path: Path) -> None:
"""One interrupted append must not get the whole shard read as cp1252."""
path = tmp_path / "out.jsonl"
good = [{"id": 1, "author": "Jürgen"}, {"id": 3, "author": "Grüße"}]
torn = '{"id": 2, "author": "Jürgen"}'.encode()[:-6] # cut mid-character
path.write_bytes(
json.dumps(good[0], ensure_ascii = False).encode()
+ b"\n"
+ torn
+ b"\n"
+ json.dumps(good[1], ensure_ascii = False).encode()
+ b"\n"
)
before = path.read_bytes()
writer = _load_state_store("cp1252").JsonlWriter(path)
try:
assert writer.has("id:1") and writer.has("id:3")
assert not writer.has("id:2") # torn line yields no key
finally:
writer.close()
# Untouched: no rewrite, so no record was re-encoded into mojibake.
after = path.read_bytes()
assert after.startswith(before)
assert "Jürgen".encode() in after
assert "Jürgen".encode("utf-8").decode("cp1252").encode() not in after
def test_an_undecodable_transport_marker_reads_as_unknown(tmp_path: Path) -> None:
"""Pinning the decode turns an undecodable marker into UnicodeDecodeError,
which is a ValueError and so is not an OSError. Before the pin those bytes
simply read as an unknown value and the caller safely purged and restarted
the partial download; letting the error escape aborts the transfer instead.
"""
import sys
backend = str(Path(__file__).resolve().parent.parent)
if backend not in sys.path:
sys.path.insert(0, backend)
from hub.utils import download_registry as registry
marker = tmp_path / ".transport"
marker.write_bytes(b"\x80\xffnative\n")
assert registry._read_marker_value(marker) is None
# A readable but unknown value takes the same path (the behaviour restored).
marker.write_text("something-else\n", encoding = "utf-8")
assert registry._read_marker_value(marker) is None
def test_a_torn_cache_ref_reads_as_not_cached(tmp_path: Path, monkeypatch) -> None:
"""hf_cache_snapshot_dir answers "is this model already on disk", and the
offline embedding checks turn a raise into a 500. A refs/main holding a byte
the codepage used to decode into a nonsense commit simply missed the snapshot
dir before the pin; it has to keep missing it."""
import sys
backend = str(Path(__file__).resolve().parent.parent)
if backend not in sys.path:
sys.path.insert(0, backend)
from utils import utils as backend_utils
good_root = tmp_path / "good"
torn_root = tmp_path / "torn"
for root, ref_bytes in ((torn_root, b"\x80\xff\n"), (good_root, b"abc123\n")):
repo = root / "models--Org--Model"
(repo / "refs").mkdir(parents = True)
(repo / "refs" / "main").write_bytes(ref_bytes)
(good_root / "models--Org--Model" / "snapshots" / "abc123").mkdir(parents = True)
monkeypatch.setattr(backend_utils, "_hf_cache_roots", lambda: [torn_root])
assert backend_utils.hf_cache_snapshot_dir("Org/Model") is None
# The torn root is skipped, not fatal: a healthy second root still answers.
monkeypatch.setattr(backend_utils, "_hf_cache_roots", lambda: [torn_root, good_root])
found = backend_utils.hf_cache_snapshot_dir("Org/Model")
assert found is not None and found.name == "abc123"
def test_a_corrupt_pid_file_does_not_abort_shutdown(tmp_path: Path, monkeypatch) -> None:
"""_remove_pid_file runs first in _graceful_shutdown, so a raise there leaves
the inference, export, training and tunnel children alive."""
import sys
backend = str(Path(__file__).resolve().parent.parent)
if backend not in sys.path:
sys.path.insert(0, backend)
import run as studio_run
pid_file = tmp_path / "studio.pid"
pid_file.write_bytes(b"\x80\xff")
monkeypatch.setattr(studio_run, "_PID_FILE", pid_file)
studio_run._remove_pid_file()
# Not this process's PID, so the file stays; the point is that it returned.
assert pid_file.exists()
pid_file.write_text(str(os.getpid()), encoding = "utf-8")
studio_run._remove_pid_file()
assert not pid_file.exists()
def test_the_kwargs_guard_only_judges_dicts_that_reach_a_call(tmp_path: Path) -> None:
"""Only a dict splatted into a call is subprocess configuration. An unrelated
payload that happens to carry "text": True is not, and neither is one whose
encoding is filled in on a later line."""
cases = {
"offender.py": 'kw = {"text": True}\nrun(cmd, **kw)\n',
"annotated.py": 'kw: dict = {"universal_newlines": True}\nrun(cmd, **kw)\n',
"payload.py": 'payload = {"text": True}\nrequests.post(url, json = payload)\n',
"inline.py": 'run(cmd, **{"text": True})\n',
"later.py": 'kw = {"text": True}\nkw["encoding"] = "utf-8"\nrun(cmd, **kw)\n',
"carried.py": 'kw = {"text": True, "encoding": "utf-8"}\nrun(cmd, **kw)\n',
}
flagged = set()
for name, source in cases.items():
path = tmp_path / name
path.write_text(source, encoding = "utf-8")
if any("subprocess kwargs" in line for line in _offenders(path)):
flagged.add(name)
assert flagged == {"offender.py", "annotated.py", "inline.py"}, flagged
def test_the_guard_follows_subprocess_through_an_alias(tmp_path: Path) -> None:
"""install_wheel() takes ``run = subprocess.run`` and calls it as a bare
name, so an attribute-only match let both of its installer calls drop their
encoding unnoticed. A name bound to something else is still not subprocess."""
cases = {
"param_default.py": (
"import subprocess\n"
"def install(*, run = subprocess.run):\n"
" run(cmd, text = True)\n"
),
"assigned.py": "import subprocess\n_run = subprocess.run\n_run(cmd, text = True)\n",
"imported.py": "from subprocess import check_output\ncheck_output(cmd, text = True)\n",
"renamed.py": "from subprocess import run as _r\n_r(cmd, universal_newlines = True)\n",
"encoded.py": (
"import subprocess\n"
"def install(*, run = subprocess.run):\n"
' run(cmd, text = True, encoding = "utf-8")\n'
),
"unrelated.py": "def run(cmd, text = False):\n pass\nrun(cmd, text = True)\n",
}
flagged = set()
for name, source in cases.items():
path = tmp_path / name
path.write_text(source, encoding = "utf-8")
if any("subprocess(text = True)" in line for line in _offenders(path)):
flagged.add(name)
assert flagged == {"param_default.py", "assigned.py", "imported.py", "renamed.py"}, flagged
def test_the_guard_sees_os_fdopen(tmp_path: Path) -> None:
"""os.fdopen(fd, mode) is open() on a descriptor and takes the same locale
default in text mode, so leaving it out let the swap lock file keep the
codepage on the write side while its reader was pinned to UTF-8."""
cases = {
"text.py": 'import os\nos.fdopen(fd, "w")\n',
"default_mode.py": "import os\nos.fdopen(fd)\n", # defaults to "r", still text
"binary.py": 'import os\nos.fdopen(fd, "wb")\n',
"keyword.py": 'import os\nos.fdopen(fd, "w", encoding = "utf-8")\n',
"positional.py": 'import os\nos.fdopen(fd, "w", 1, "utf-8")\n',
}
flagged = set()
for name, source in cases.items():
path = tmp_path / name
path.write_text(source, encoding = "utf-8")
if any("fdopen" in line for line in _offenders(path)):
flagged.add(name)
assert flagged == {"text.py", "default_mode.py"}, flagged
def test_an_undecodable_bootstrap_password_does_not_stop_startup(
tmp_path: Path, monkeypatch
) -> None:
"""ensure_default_admin calls _load_bootstrap_password for every existing
admin and the lifespan calls that with no handler, so a raise here takes the
whole backend down instead of ignoring an unusable file."""
import sys
backend = str(Path(__file__).resolve().parent.parent)
if backend not in sys.path:
sys.path.insert(0, backend)
from auth import storage
pw_file = tmp_path / ".bootstrap_password"
pw_file.write_bytes(b"\x80\xffnot-utf8\n")
monkeypatch.setattr(storage, "_BOOTSTRAP_PW_PATH", pw_file)
assert storage._load_bootstrap_password() is None
# A readable one still loads, so this is a narrowing of failure, not of function.
pw_file.write_text("correct horse battery staple\n", encoding = "utf-8")
assert storage._load_bootstrap_password() == "correct horse battery staple"
def test_a_damaged_checkpoint_resets_instead_of_resuming_on_a_broken_cursor(tmp_path: Path) -> None:
"""A checkpoint holds only base64 cursors and booleans, so a codepage reading
can only ever add non-ASCII, never recover any. Resuming on a mojibaked cursor
sends GitHub one it answers with INVALID_CURSOR_ARGUMENTS, and the empty page
that comes back marks the stream done and skips the rest of it for good.
Dropping the checkpoint only replays pages the writers already dedup."""
module = _load_state_store("cp1252")
cursor = "Y3Vyc29yOnYyOpK0MjAxMi0wMi0xNlQwNjo1Mzo0MVrOADGL_A=="
healthy = json.dumps({"issues_cursor": cursor, "issues_done": False}, indent = 2)
path = tmp_path / "octocat__Hello-World.json"
path.write_text(healthy, encoding = "utf-8")
assert module.StateStore(path).get("issues_cursor") == cursor
# Written by a pre-UTF-8 release in the operator's codepage. Nothing is lost
# by reading UTF-8 only, because an all-ASCII document is the same bytes.
path.write_bytes(healthy.encode("cp1252"))
assert module.StateStore(path).get("issues_cursor") == cursor
# One damaged byte inside the cursor: still a whole JSON document under a
# single-byte codepage, so only refusing that reading resets the checkpoint.
raw = healthy.encode()
at = raw.index(b"MjAxMi0wMi0xNlQ") + 3
path.write_bytes(raw[:at] + b"\x96" + raw[at + 1 :])
assert json.loads(path.read_bytes().decode("latin-1"))["issues_cursor"] != cursor
store = module.StateStore(path)
assert store.all() == {}
assert store.get("issues_cursor") is None
def test_a_utf8_record_is_not_parsed_a_second_time(tmp_path: Path) -> None:
"""These shards reach gigabytes and every resume reads all of one, so a
record that already read as UTF-8 must not be decoded and parsed again under
the codepage. The legacy reading exists only to recover keys UTF-8 could not."""
module = _load_state_store("cp1252")
calls: list[str] = []
real_parse = module._parse
def counting_parse(raw, encoding):
calls.append(encoding)
return real_parse(raw, encoding)
module._parse = counting_parse
try:
healthy = json.dumps({"id": 1, "author": "Jürgen"}).encode("utf-8")
reading = module._read_line(healthy, "cp1252")
assert reading.as_utf8 == {"id": 1, "author": "Jürgen"}
assert calls == ["utf-8"], calls
# A line UTF-8 cannot read still falls through to the codepage, the whole point.
calls.clear()
legacy = json.dumps({"id": 2, "author": "Jürgen"}, ensure_ascii = False).encode("cp1252")
reading = module._read_line(legacy, "cp1252")
assert reading.as_utf8 is None
assert reading.as_legacy == {"id": 2, "author": "Jürgen"}
assert calls == ["utf-8", "cp1252"], calls
finally:
module._parse = real_parse
def _too_deeply_nested_json() -> str:
"""A JSON document nested past what this interpreter will descend into.
Probed rather than hardcoded: the depth json.loads gives up at is bounded by
sys.getrecursionlimit() up to 3.11 and by the C recursion limit from 3.12,
which sys.setrecursionlimit no longer moves and which varies by micro
version. That is ~995 on 3.9 and ~9999 on 3.13.
"""
depth = 1
while depth <= 1 << 17:
document = "[" * depth + "]" * depth
try:
json.loads(document)
except RecursionError:
return document
depth *= 2
pytest.skip("this interpreter parses arbitrarily nested JSON")
def test_an_unparseably_nested_document_is_discarded_not_raised(tmp_path: Path) -> None:
"""json.loads answers nesting it cannot descend with RecursionError, which is
a RuntimeError and so is neither a ValueError nor a UnicodeDecodeError.
_parse is called outside any other handler in both StateStore.__init__ and
JsonlWriter._scan_existing, so letting it escape aborts the scraper at
startup on a file the catch-all it replaced simply discarded."""
module = _load_state_store("cp1252")
nested = _too_deeply_nested_json()
checkpoint = tmp_path / "octocat__Hello-World.json"
checkpoint.write_text(nested, encoding = "utf-8")
assert module.StateStore(checkpoint).all() == {} # reset, not raised
shard = tmp_path / "out.jsonl"
shard.write_text(
nested + "\n" + json.dumps({"id": 1}) + "\n" + json.dumps({"id": 2}) + "\n",
encoding = "utf-8",
)
writer = module.JsonlWriter(shard)
try:
# Skipped like any other unreadable line, so its neighbours still yield the dedup
# keys that keep the resume from re-fetching them.
assert writer.has("id:1") and writer.has("id:2")
finally:
writer.close()

View file

@ -9,28 +9,8 @@ import sys
from typing import Any
from unittest import mock
import pytest
from core.training import worker
# The runtime install is Linux-only, so elsewhere these return before any status.
linux_only = pytest.mark.skipif(
not sys.platform.startswith("linux"),
reason = "the runtime flash-attn install is gated to Linux",
)
# causal-conv1d and flash-linear-attention are NOT Linux-gated: both installers bail out
# on `sys.platform == "win32"` alone (no prebuilt wheel for Windows) and run everywhere
# else, macOS included. linux_only here would skip cases that legitimately pass off Linux.
not_on_windows = pytest.mark.skipif(
sys.platform == "win32",
reason = (
"mirrors the sys.platform == 'win32' bail-out in "
"_ensure_flash_linear_attention_unconditional and "
"_ensure_causal_conv1d_fast_path"
),
)
def _missing_flash_attn_import():
real_import = builtins.__import__
@ -75,7 +55,6 @@ def test_should_try_runtime_flash_attn_install_threshold_and_skip(monkeypatch):
assert worker._should_try_runtime_flash_attn_install(32768) is False
@linux_only
def test_runtime_flash_attn_prefers_prebuilt_wheel(monkeypatch):
statuses: list[str] = []
@ -103,7 +82,6 @@ def test_runtime_flash_attn_prefers_prebuilt_wheel(monkeypatch):
assert statuses == ["Installing flash-attn for faster training..."]
@linux_only
def test_runtime_flash_attn_falls_back_to_pypi(monkeypatch):
calls: list[list[str]] = []
statuses: list[str] = []
@ -135,7 +113,12 @@ def test_runtime_flash_attn_falls_back_to_pypi(monkeypatch):
)
monkeypatch.setattr(worker, "install_wheel", mock.Mock())
def fake_run(cmd, **kwargs):
def fake_run(
cmd,
stdout = None,
stderr = None,
text = None,
):
calls.append(list(cmd))
return subprocess.CompletedProcess(cmd, 0, "")
@ -156,7 +139,6 @@ def test_runtime_flash_attn_skip_env_avoids_all_install_work(monkeypatch):
worker._sp.run.assert_not_called()
@not_on_windows
def test_causal_conv1d_fast_path_preserves_wheel_first_install_args(monkeypatch):
install_mock = mock.Mock(return_value = True)
monkeypatch.setattr(worker, "_install_package_wheel_first", install_mock)
@ -178,7 +160,6 @@ def test_causal_conv1d_fast_path_preserves_wheel_first_install_args(monkeypatch)
)
@not_on_windows
def test_causal_conv1d_fast_path_includes_qwen3_6_variants(monkeypatch):
install_mock = mock.Mock(return_value = True)
monkeypatch.setattr(worker, "_install_package_wheel_first", install_mock)
@ -244,7 +225,6 @@ def _pin_fla_model_types(monkeypatch):
)
@not_on_windows
def test_flash_linear_attention_installs_pinned_pair_for_qwen3_5(monkeypatch):
_pin_fla_model_types(monkeypatch)
monkeypatch.setattr(worker.shutil, "which", lambda name: "/usr/bin/uv")
@ -297,7 +277,6 @@ def test_flash_linear_attention_skips_for_ssm_only_models(monkeypatch):
run_mock.assert_not_called()
@not_on_windows
def test_flash_linear_attention_matches_full_qwen3_family(monkeypatch):
monkeypatch.setattr(worker.shutil, "which", lambda name: "/usr/bin/uv")
run_mock = mock.Mock(return_value = mock.Mock(returncode = 0, stdout = ""))
@ -352,7 +331,6 @@ def test_flash_linear_attention_skipped_via_env(monkeypatch):
run_mock.assert_not_called()
@not_on_windows
def test_flash_linear_attention_skipped_below_torch_2_7(monkeypatch):
_pin_fla_model_types(monkeypatch)
monkeypatch.delenv(worker._FLA_SKIP_ENV, raising = False)
@ -371,7 +349,6 @@ def test_flash_linear_attention_skipped_below_torch_2_7(monkeypatch):
assert any("torch>=" in s for s in statuses)
@not_on_windows
def test_flash_linear_attention_install_includes_einops(monkeypatch):
_pin_fla_model_types(monkeypatch)
monkeypatch.delenv(worker._FLA_SKIP_ENV, raising = False)
@ -398,7 +375,6 @@ def test_flash_linear_attention_install_includes_einops(monkeypatch):
assert f"fla-core=={worker._FLA_CORE_PACKAGE_VERSION}" in args
@not_on_windows
def test_flash_linear_attention_logs_post_install_import_failure(monkeypatch):
"""pip exits 0 but `import fla.modules` still fails (missing transitive)."""
_pin_fla_model_types(monkeypatch)
@ -445,7 +421,6 @@ def test_tilelang_backend_skipped_on_unsupported_linux_arch(monkeypatch):
run_mock.assert_not_called()
@linux_only
def test_tilelang_backend_pins_only_binary(monkeypatch):
_pin_fla_model_types(monkeypatch)
monkeypatch.delenv(worker._TILELANG_SKIP_ENV, raising = False)
@ -487,7 +462,6 @@ def _force_missing_tilelang_imports(monkeypatch):
monkeypatch.setattr(builtins, "__import__", fake_import)
@linux_only
def test_tilelang_backend_installs_pinned_pair_for_qwen3_5(monkeypatch):
_pin_fla_model_types(monkeypatch)
monkeypatch.delenv(worker._TILELANG_SKIP_ENV, raising = False)
@ -512,7 +486,6 @@ def test_tilelang_backend_installs_pinned_pair_for_qwen3_5(monkeypatch):
assert any("Installing TileLang" in s for s in statuses)
@linux_only
def test_tilelang_backend_reinstalls_when_tvm_ffi_is_broken(monkeypatch):
"""Repair path issues TWO pip calls:
@ -582,7 +555,6 @@ def test_tilelang_backend_skipped_on_windows(monkeypatch):
run_mock.assert_not_called()
@linux_only
def test_tilelang_backend_swallows_install_timeout(monkeypatch):
_pin_fla_model_types(monkeypatch)
monkeypatch.delenv(worker._TILELANG_SKIP_ENV, raising = False)
@ -637,7 +609,6 @@ def test_tilelang_backend_skipped_via_env(monkeypatch):
run_mock.assert_not_called()
@linux_only
def test_tilelang_backend_swallows_install_failure(monkeypatch):
_pin_fla_model_types(monkeypatch)
monkeypatch.delenv(worker._TILELANG_SKIP_ENV, raising = False)
@ -702,7 +673,6 @@ def _patch_iu_gates(monkeypatch, fla_gate, conv_gate):
monkeypatch.setattr(_iu, "is_causal_conv1d_available", conv_gate)
@not_on_windows
def test_hook_installs_when_gate_returns_false(monkeypatch):
_pin_fla_model_types(monkeypatch)
fla_gate = _make_fake_gate(initial_return = False)
@ -1006,7 +976,6 @@ def test_hook_does_install_tilelang_for_qwen35(monkeypatch):
tile_install.assert_called_once()
@linux_only
def test_tilelang_repair_does_not_touch_torch_cuda_stack(monkeypatch):
"""Finding #2: the broken-tvm-ffi repair must use --no-deps on the
forced step so --force-reinstall doesn't cascade through
@ -1150,7 +1119,6 @@ def test_hook_runs_tilelang_repair_when_fla_already_true(monkeypatch):
tile_install.assert_called_once()
@not_on_windows
def test_fla_installer_force_reinstalls_when_older_version_present(monkeypatch):
"""Finding #8: an older `flash-linear-attention` that is importable
but below the pin must force a reinstall (not no-op).
@ -1615,10 +1583,15 @@ def test_install_respects_user_gcc_install_dir(monkeypatch):
)
_make_hip_install_env(monkeypatch, gcc_dir = "/usr/lib/gcc/x86_64-linux-gnu/13")
captured: dict[str, str] = {}
captured: dict[str, str] | None = {"_called": "no"}
def fake_run(cmd, **kwargs):
captured.update(kwargs.get("env") or {})
env = kwargs.get("env")
if env is not None:
captured.clear()
captured.update(env)
else:
captured["_called"] = "yes_no_env"
return subprocess.CompletedProcess(cmd, 0, "")
monkeypatch.setattr(worker._sp, "run", fake_run)
@ -1634,11 +1607,14 @@ def test_install_respects_user_gcc_install_dir(monkeypatch):
release_base_url = "https://example.com",
)
assert captured["HIPCC_COMPILE_FLAGS_APPEND"] == "--gcc-install-dir=/opt/custom/gcc-13"
# subprocess.run invoked without env override (user already set
# HIPCC_COMPILE_FLAGS_APPEND with --gcc-install-dir, so we left the
# env alone — the existing value is inherited).
assert captured == {"_called": "yes_no_env"}
def test_install_does_not_inject_env_on_cuda(monkeypatch):
"""CUDA path (no hip_version in env) → no HIP flag injected."""
"""CUDA path (no hip_version in env) → no env override at all."""
monkeypatch.delenv("HIPCC_COMPILE_FLAGS_APPEND", raising = False)
monkeypatch.setattr(builtins, "__import__", _missing_module_import("causal_conv1d"))
monkeypatch.setattr(
@ -1665,7 +1641,7 @@ def test_install_does_not_inject_env_on_cuda(monkeypatch):
captured: dict[str, Any] = {}
def fake_run(cmd, **kwargs):
captured.update(kwargs.get("env") or {})
captured["env_in_kwargs"] = "env" in kwargs
return subprocess.CompletedProcess(cmd, 0, "")
monkeypatch.setattr(worker._sp, "run", fake_run)
@ -1681,5 +1657,5 @@ def test_install_does_not_inject_env_on_cuda(monkeypatch):
release_base_url = "https://example.com",
)
# env is always passed (to force UTF-8), but never the HIP flag.
assert "HIPCC_COMPILE_FLAGS_APPEND" not in captured
# CUDA branch never sets the env, never invokes the gcc helper.
assert captured.get("env_in_kwargs") is False

File diff suppressed because it is too large Load diff

View file

@ -1,22 +0,0 @@
# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""Make a Python child agree with the parent that its pipes are UTF-8.
A child's ``sys.stdout`` uses ``locale.getpreferredencoding()``, which on
Windows is the ANSI code page. Reading that pipe as UTF-8 would then mangle any
non-ASCII the child prints, so the child has to be told which encoding to emit.
Only needed for Python children; llama.cpp and node already emit UTF-8.
"""
from __future__ import annotations
import os
from typing import Mapping, Optional
def utf8_child_env(env: Optional[Mapping[str, str]] = None) -> dict[str, str]:
"""Copy *env* (or the current environment) with UTF-8 stdio forced."""
child = dict(os.environ if env is None else env)
child["PYTHONIOENCODING"] = "utf-8"
return child

View file

@ -144,8 +144,6 @@ def _run_amd_smi(*args: str, timeout: int = _AMD_SMI_DEFAULT_TIMEOUT) -> Optiona
["amd-smi", *args, "--json"],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = timeout,
env = _amd_env,
**windows_hidden_subprocess_kwargs(),

View file

@ -830,8 +830,6 @@ def _rocm_windows_perf_counter_gpu_util_pct() -> Optional[float]:
["powershell", "-NoProfile", "-NonInteractive", "-Command", ps],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = 5,
)
if r.returncode != 0 or not r.stdout.strip():
@ -1029,8 +1027,6 @@ def _rocm_windows_perf_counter_vram_by_adapter() -> Optional[list[tuple[str, flo
["powershell", "-NoProfile", "-NonInteractive", "-Command", ps],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = 5,
)
if r.returncode != 0 or not r.stdout.strip():

View file

@ -55,8 +55,6 @@ def get_physical_gpu_count() -> Optional[int]:
["nvidia-smi", "-L"],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = 5,
env = child_env_without_native_path_secret(),
**_windows_hidden_subprocess_kwargs(),
@ -83,8 +81,6 @@ def get_primary_gpu_utilization() -> dict[str, Any]:
],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = 5,
env = child_env_without_native_path_secret(),
**_windows_hidden_subprocess_kwargs(),
@ -135,8 +131,6 @@ def get_visible_gpu_utilization(
],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = 5,
env = child_env_without_native_path_secret(),
**_windows_hidden_subprocess_kwargs(),
@ -221,8 +215,6 @@ def get_backend_visible_gpu_info(
],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = 10,
env = child_env_without_native_path_secret(),
**_windows_hidden_subprocess_kwargs(),

View file

@ -121,14 +121,7 @@ def _installed_build_number(binary: Optional[str]) -> Optional[int]:
if not binary:
return None
try:
proc = subprocess.run(
[binary, "--version"],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = 20,
)
proc = subprocess.run([binary, "--version"], capture_output = True, text = True, timeout = 20)
except Exception: # pragma: no cover - defensive
return None
m = re.search(r"version:\s*(\d+)", (proc.stderr or "") + (proc.stdout or ""))

View file

@ -254,7 +254,7 @@ def _transformers_constraint_args() -> tuple[list[str], str | None]:
except Exception:
return [], None
fd, path = tempfile.mkstemp(prefix = "mlx_repair_", suffix = ".txt")
with os.fdopen(fd, "w", encoding = "utf-8") as fh:
with os.fdopen(fd, "w") as fh:
fh.write(f"transformers=={transformers_version}\n")
return ["--constraint", path], path
@ -290,8 +290,6 @@ def attempt_mlx_repair(*, timeout: int = _REPAIR_TIMEOUT_S) -> bool:
stdout = subprocess.PIPE,
stderr = subprocess.STDOUT,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = timeout,
)
except subprocess.TimeoutExpired:

View file

@ -129,7 +129,7 @@ def _read_checkpoint_loss(checkpoint_path: Path) -> Optional[float]:
if not trainer_state.exists():
return None
try:
with open(trainer_state, encoding = "utf-8-sig") as f:
with open(trainer_state, encoding = "utf-8") as f:
state = json.load(f)
log_history = state.get("log_history", [])
if log_history:
@ -174,18 +174,18 @@ def scan_checkpoints(
metadata: dict = {}
try:
if adapter_config.exists():
cfg = json.loads(adapter_config.read_text(encoding = "utf-8-sig"))
cfg = json.loads(adapter_config.read_text(encoding = "utf-8"))
metadata["base_model"] = cfg.get("base_model_name_or_path")
metadata["peft_type"] = cfg.get("peft_type")
metadata["lora_rank"] = cfg.get("r")
elif config_file.exists():
cfg = json.loads(config_file.read_text(encoding = "utf-8-sig"))
cfg = json.loads(config_file.read_text(encoding = "utf-8"))
metadata["base_model"] = cfg.get("_name_or_path")
# Detect BNB quantization from config.json
if config_file.exists():
if "cfg" not in dir():
cfg = json.loads(config_file.read_text(encoding = "utf-8-sig"))
cfg = json.loads(config_file.read_text(encoding = "utf-8"))
quant_cfg = cfg.get("quantization_config")
if (
isinstance(quant_cfg, dict)

View file

@ -37,7 +37,6 @@ import yaml
from utils.native_path_leases import child_env_without_native_path_secret
from utils.child_stdio import utf8_child_env
from utils.hf_cache_settings import active_hf_hub_cache, get_hf_cache_paths
from utils.subprocess_compat import (
windows_hidden_subprocess_kwargs as _windows_hidden_subprocess_kwargs,
@ -632,7 +631,7 @@ def _raw_config_has_vision_config(
cache_dir = active_hf_hub_cache(),
)
)
config = json.loads(config_path.read_text(encoding = "utf-8-sig"))
config = json.loads(config_path.read_text(encoding = "utf-8"))
architectures = config.get("architectures") or []
model_type = config.get("model_type")
explicit_vision = (
@ -775,12 +774,8 @@ def _is_vision_model_subprocess(model_name: str, hf_token: Optional[str] = None)
],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = 60,
env = utf8_child_env(
get_hf_cache_paths().child_env(child_env_without_native_path_secret())
),
env = get_hf_cache_paths().child_env(child_env_without_native_path_secret()),
**_windows_hidden_subprocess_kwargs(),
)
@ -1088,7 +1083,7 @@ def _detect_audio_from_tokenizer(
]:
tok_file = snapshot / tok_path
if tok_file.exists():
tok_config = json.loads(tok_file.read_text(encoding = "utf-8-sig"))
tok_config = json.loads(tok_file.read_text(encoding = "utf-8"))
read_any = True
result = _check_token_patterns(tok_config)
if result:
@ -2288,7 +2283,7 @@ def scan_exported_models(
export_meta = run_dir / "export_metadata.json"
try:
if export_meta.exists():
meta = json.loads(export_meta.read_text(encoding = "utf-8-sig"))
meta = json.loads(export_meta.read_text(encoding = "utf-8"))
base_model = meta.get("base_model")
except Exception:
pass
@ -2317,7 +2312,7 @@ def scan_exported_models(
if adapter_config.exists():
export_type = "lora"
try:
cfg = json.loads(adapter_config.read_text(encoding = "utf-8-sig"))
cfg = json.loads(adapter_config.read_text(encoding = "utf-8"))
base_model = cfg.get("base_model_name_or_path")
except Exception:
pass
@ -2326,7 +2321,7 @@ def scan_exported_models(
export_meta = checkpoint_dir / "export_metadata.json"
try:
if export_meta.exists():
meta = json.loads(export_meta.read_text(encoding = "utf-8-sig"))
meta = json.loads(export_meta.read_text(encoding = "utf-8"))
base_model = meta.get("base_model")
except Exception:
pass
@ -2339,7 +2334,7 @@ def scan_exported_models(
export_meta = meta_dir / "export_metadata.json"
try:
if export_meta.exists():
meta = json.loads(export_meta.read_text(encoding = "utf-8-sig"))
meta = json.loads(export_meta.read_text(encoding = "utf-8"))
base_model = meta.get("base_model")
if base_model:
break
@ -2359,7 +2354,7 @@ def scan_exported_models(
outputs_adapter_cfg = resolve_output_dir(run_dir.name) / "adapter_config.json"
try:
if outputs_adapter_cfg.exists():
cfg = json.loads(outputs_adapter_cfg.read_text(encoding = "utf-8-sig"))
cfg = json.loads(outputs_adapter_cfg.read_text(encoding = "utf-8"))
base_model = cfg.get("base_model_name_or_path")
except Exception:
pass
@ -2385,7 +2380,7 @@ def get_base_model_from_checkpoint(checkpoint_path: str) -> Optional[str]:
adapter_config_path = checkpoint_path_obj / "adapter_config.json"
if adapter_config_path.exists():
with open(adapter_config_path, "r", encoding = "utf-8-sig") as f:
with open(adapter_config_path, "r", encoding = "utf-8") as f:
config = json.load(f)
base_model = config.get("base_model_name_or_path")
if base_model:
@ -2394,7 +2389,7 @@ def get_base_model_from_checkpoint(checkpoint_path: str) -> Optional[str]:
config_path = checkpoint_path_obj / "config.json"
if config_path.exists():
with open(config_path, "r", encoding = "utf-8-sig") as f:
with open(config_path, "r", encoding = "utf-8") as f:
config = json.load(f)
for key in ("model_name", "_name_or_path"):
base_model = config.get(key)
@ -2450,7 +2445,7 @@ def get_base_model_from_lora(lora_path: str) -> Optional[str]:
# adapter_config.json first
adapter_config_path = lora_path_obj / "adapter_config.json"
if adapter_config_path.exists():
with open(adapter_config_path, "r", encoding = "utf-8-sig") as f:
with open(adapter_config_path, "r", encoding = "utf-8") as f:
config = json.load(f)
base_model = config.get("base_model_name_or_path")
if base_model:
@ -2540,7 +2535,7 @@ def get_base_model_from_lora_identifier(
last_exc = exc
continue
try:
with open(cfg_path, "r", encoding = "utf-8-sig") as f:
with open(cfg_path, "r", encoding = "utf-8") as f:
base_model = json.load(f).get("base_model_name_or_path")
except Exception as exc:
logger.warning("Could not parse adapter_config.json for '%s': %s", identifier, exc)
@ -2786,7 +2781,7 @@ class ModelConfig:
meta_path = gguf_dir / "export_metadata.json"
if meta_path.exists():
try:
meta = json.loads(meta_path.read_text(encoding = "utf-8-sig"))
meta = json.loads(meta_path.read_text(encoding = "utf-8"))
base = meta.get("base_model")
if base and is_vision_model(base, hf_token = hf_token):
base_is_vision = True
@ -2917,7 +2912,7 @@ class ModelConfig:
token = hf_token,
cache_dir = active_hf_hub_cache(),
)
with open(config_path, "r", encoding = "utf-8-sig") as f:
with open(config_path, "r", encoding = "utf-8") as f:
adapter_config = json.load(f)
base_model = adapter_config.get("base_model_name_or_path")
if base_model:

View file

@ -79,8 +79,6 @@ def _node_version_ok(executable: str) -> bool:
[executable, "-v"],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = _NODE_VERSION_PROBE_TIMEOUT_SECONDS,
**windows_hidden_subprocess_kwargs(),
)

View file

@ -212,7 +212,7 @@ def lmstudio_model_dirs() -> list[Path]:
settings_path = Path.home() / ".lmstudio" / "settings.json"
if settings_path.is_file():
try:
with open(settings_path, encoding = "utf-8-sig") as f:
with open(settings_path, encoding = "utf-8") as f:
settings = json.load(f)
downloads = settings.get("downloadsFolder", "")
if downloads:

View file

@ -24,7 +24,6 @@ from typing import Callable, Optional
import structlog
from utils.child_stdio import utf8_child_env
from utils.process_lifetime import child_popen_kwargs
logger = structlog.get_logger(__name__)
@ -160,8 +159,6 @@ def resolve_prebuilt_for_host(
cmd,
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = 60,
)
out = (proc.stdout or "").strip()
@ -306,10 +303,7 @@ def stream_installer(
stdout = subprocess.PIPE,
stderr = subprocess.STDOUT,
text = True,
encoding = "utf-8",
errors = "replace",
# Make the Python child emit the UTF-8 we decode above.
env = utf8_child_env(env),
env = env,
**child_popen_kwargs(),
)
timed_out = threading.Event()

View file

@ -142,7 +142,7 @@ def _load_remote_code_configs(model_name: str, hf_token: Optional[str] = None) -
for name in _REMOTE_CODE_CONFIG_FILES:
p = root / name
if p.is_file():
configs.append(json.loads(p.read_text(encoding = "utf-8-sig")))
configs.append(json.loads(p.read_text(encoding = "utf-8")))
return configs
from huggingface_hub import hf_hub_download
@ -164,7 +164,7 @@ def _load_remote_code_configs(model_name: str, hf_token: Optional[str] = None) -
# Transient/auth failure is not "absent" -> fail closed to "unknown" so
# the caller scans (a tokenizer/processor-only auto_map must not slip by).
return None
configs.append(json.loads(Path(p).read_text(encoding = "utf-8-sig")))
configs.append(json.loads(Path(p).read_text(encoding = "utf-8")))
# Every config was read or a genuine 404 -> an empty list is a definitive
# "no auto_map", not "unknown".
return configs

View file

@ -199,7 +199,7 @@ def _indexed_shard_paths(
inconclusive = True # transient: an index that might exist could not be read
continue
try:
weight_map = (json.loads(open(index_path, encoding = "utf-8-sig").read()) or {}).get(
weight_map = (json.loads(open(index_path, encoding = "utf-8").read()) or {}).get(
"weight_map"
) or {}
for shard in weight_map.values():
@ -328,7 +328,7 @@ def _st_load_roots(snapshot: Path) -> list:
roots = [snapshot]
try:
import json
modules = json.loads((snapshot / "modules.json").read_text(encoding = "utf-8-sig"))
modules = json.loads((snapshot / "modules.json").read_text(encoding = "utf-8"))
except (OSError, ValueError):
return roots # no / invalid modules.json -> snapshot root is the only load root
for module in modules or ():
@ -355,7 +355,7 @@ def _indexed_pickle_shards(index_path: Path, root: Path, snapshot: Path) -> list
try:
# JSON is UTF-8 by spec; pin it so a non-ASCII index is not misdecoded (and needlessly
# blocked) under Windows' cp1252 default.
parsed = json.loads(index_path.read_text(encoding = "utf-8-sig"))
parsed = json.loads(index_path.read_text(encoding = "utf-8"))
except (OSError, ValueError) as exc:
raise OSError(f"unreadable weight index: {index_path}") from exc
weight_map = parsed.get("weight_map") if isinstance(parsed, dict) else None

View file

@ -69,7 +69,7 @@ def approval_target_key(targets) -> str:
def _load() -> dict:
"""Parsed store, or an empty skeleton on any error (fail-safe = re-prompt)."""
try:
with open(_store_path(), encoding = "utf-8-sig") as f:
with open(_store_path(), encoding = "utf-8") as f:
data = json.load(f)
# Validate the shape, not just the version: a hand-edited ``subjects`` that is not a
# dict (e.g. ``[]``) would otherwise crash lookup/record instead of failing safe.

View file

@ -454,7 +454,7 @@ def repo_remote_code_files(model_name: str, hf_token: Optional[str] = None) -> d
p = root / name
if p.is_file():
try:
ext_refs |= _auto_map_refs(json.loads(p.read_text(encoding = "utf-8-sig")))
ext_refs |= _auto_map_refs(json.loads(p.read_text(encoding = "utf-8")))
except Exception:
pass
if not _add_external_refs(files, ext_refs, hf_token, model_name):
@ -483,7 +483,7 @@ def repo_remote_code_files(model_name: str, hf_token: Optional[str] = None) -> d
f"{model_name}: config {cfg_name} could not be fetched ({exc})"
) from exc
try:
refs |= _auto_map_refs(json.loads(Path(cfg_path).read_text(encoding = "utf-8-sig")))
refs |= _auto_map_refs(json.loads(Path(cfg_path).read_text(encoding = "utf-8")))
except Exception:
pass
own_refs = {fn for repo, fn in refs if repo is None}
@ -616,7 +616,7 @@ def external_auto_map_repos(model_name: str, hf_token: Optional[str] = None) ->
if not p.is_file():
continue
try:
refs = _auto_map_refs(json.loads(p.read_text(encoding = "utf-8-sig")))
refs = _auto_map_refs(json.loads(p.read_text(encoding = "utf-8")))
except Exception:
continue
repos.update(repo for repo, _fn in refs if repo)
@ -638,7 +638,7 @@ def external_auto_map_repos(model_name: str, hf_token: Optional[str] = None) ->
except Exception:
continue
try:
refs = _auto_map_refs(json.loads(Path(cfg_path).read_text(encoding = "utf-8-sig")))
refs = _auto_map_refs(json.loads(Path(cfg_path).read_text(encoding = "utf-8")))
except Exception:
continue
repos.update(repo for repo, _fn in refs if repo)

View file

@ -23,7 +23,6 @@ import threading
from typing import Any, Callable, Optional
from loggers import get_logger
from utils.child_stdio import utf8_child_env
from utils.wheel_utils import (
direct_wheel_url,
install_wheel,
@ -255,12 +254,6 @@ def _install_kernel(
"stdout": subprocess.PIPE,
"stderr": subprocess.STDOUT,
"text": True,
# pip and the compilers it drives write UTF-8 down this pipe; the Windows
# ANSI codepage would mojibake or raise over a fine install.
"encoding": "utf-8",
"errors": "replace",
# Make the Python child emit the UTF-8 we decode above.
"env": utf8_child_env(),
}
if is_hip:
run_kwargs["timeout"] = 1800 # ROCm builds can take 10-30 min
@ -268,8 +261,7 @@ def _install_kernel(
if "--gcc-install-dir" not in existing:
gcc_dir = _hipcc_gcc_install_dir()
if gcc_dir:
# Extends the UTF-8 env above rather than replacing it.
_env = dict(run_kwargs["env"])
_env = os.environ.copy()
_env["HIPCC_COMPILE_FLAGS_APPEND"] = (
f"{existing} --gcc-install-dir={gcc_dir}".strip()
)

View file

@ -60,8 +60,6 @@ def _exact_git_studio_tag(repo_root: Path) -> str | None:
stdout = subprocess.PIPE,
stderr = subprocess.DEVNULL,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = _GIT_TIMEOUT_SECONDS,
)
except (OSError, subprocess.TimeoutExpired):
@ -83,8 +81,6 @@ def _git_branch(repo_root: Path) -> str | None:
stdout = subprocess.PIPE,
stderr = subprocess.DEVNULL,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = _GIT_TIMEOUT_SECONDS,
)
except (OSError, subprocess.TimeoutExpired):

View file

@ -44,7 +44,6 @@ import time
from pathlib import Path
from utils.native_path_leases import child_env_without_native_path_secret
from utils.child_stdio import utf8_child_env
from utils.hf_cache_settings import get_hf_cache_paths
from utils.subprocess_compat import (
windows_hidden_subprocess_kwargs as _windows_hidden_subprocess_kwargs,
@ -421,7 +420,7 @@ def _resolve_base_model(model_name: str) -> str:
adapter_cfg_path = local_path / "adapter_config.json"
if _safe_is_file(adapter_cfg_path):
try:
with open(adapter_cfg_path, encoding = "utf-8-sig") as f:
with open(adapter_cfg_path, encoding = "utf-8") as f:
cfg = json.load(f)
base = cfg.get("base_model_name_or_path")
if base:
@ -438,7 +437,7 @@ def _resolve_base_model(model_name: str) -> str:
config_json_path = local_path / "config.json"
if _safe_is_file(config_json_path):
try:
with open(config_json_path, encoding = "utf-8-sig") as f:
with open(config_json_path, encoding = "utf-8") as f:
cfg = json.load(f)
# Unsloth writes model_name, HF writes _name_or_path; skip a self-reference.
for _key in ("model_name", "_name_or_path"):
@ -545,7 +544,7 @@ def _adapter_base_from_hf_cache(model_name: str) -> str | None:
)
for cfg_path in candidates:
if cfg_path.is_file():
base = json.loads(cfg_path.read_text(encoding = "utf-8-sig")).get(
base = json.loads(cfg_path.read_text(encoding = "utf-8")).get(
"base_model_name_or_path"
)
return base or None
@ -617,7 +616,7 @@ def _check_tokenizer_config_needs_v5(model_name: str, hf_token: str | None = Non
local_tc = local_path / "tokenizer_config.json"
if _safe_is_file(local_tc):
try:
with open(local_tc, encoding = "utf-8-sig") as f:
with open(local_tc, encoding = "utf-8") as f:
data = json.load(f)
tokenizer_class = data.get("tokenizer_class", "")
result = tokenizer_class in _TRANSFORMERS_5_TOKENIZER_CLASSES
@ -707,7 +706,7 @@ def _config_json_from_hf_cache(model_name: str) -> dict | None:
)
for cfg_path in candidates:
if cfg_path.is_file():
with open(cfg_path, encoding = "utf-8-sig") as f:
with open(cfg_path, encoding = "utf-8") as f:
return json.load(f)
except Exception as exc:
logger.debug("HF cache config.json lookup failed for '%s': %s", model_name, exc)
@ -732,7 +731,7 @@ def _load_config_json(model_name: str, hf_token: str | None = None) -> dict | No
local_cfg = Path(model_name) / "config.json"
if _safe_is_file(local_cfg):
try:
with open(local_cfg, encoding = "utf-8-sig") as f:
with open(local_cfg, encoding = "utf-8") as f:
cfg = json.load(f)
_config_json_cache[cache_key] = cfg
return cfg
@ -1272,10 +1271,9 @@ def _probe_autoconfig(target_dir: str, model_name: str, hf_token: str | None) ->
[sys.executable, "-c", _PROBE_CONFIG_SCRIPT, target_dir, model_name],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = _PROBE_TIMEOUT_SECS,
env = utf8_child_env(env),
env = env,
**_windows_hidden_subprocess_kwargs(),
)
except subprocess.TimeoutExpired:
@ -1813,11 +1811,7 @@ def _install_to_dir(pkg: str, target_dir: str) -> bool:
stdout = subprocess.PIPE,
stderr = subprocess.STDOUT,
text = True,
encoding = "utf-8",
errors = "replace",
env = utf8_child_env(
get_hf_cache_paths().child_env(child_env_without_native_path_secret())
),
env = get_hf_cache_paths().child_env(child_env_without_native_path_secret()),
**_windows_hidden_subprocess_kwargs(),
)
if result.returncode == 0:
@ -1840,9 +1834,7 @@ def _install_to_dir(pkg: str, target_dir: str) -> bool:
stdout = subprocess.PIPE,
stderr = subprocess.STDOUT,
text = True,
encoding = "utf-8",
errors = "replace",
env = utf8_child_env(get_hf_cache_paths().child_env(child_env_without_native_path_secret())),
env = get_hf_cache_paths().child_env(child_env_without_native_path_secret()),
**_windows_hidden_subprocess_kwargs(),
)
if result.returncode != 0:
@ -2087,7 +2079,7 @@ class SidecarSwapInProgress(RuntimeError):
def _read_swap_lock(path: Path) -> dict | None:
try:
data = json.loads(path.read_text(encoding = "utf-8-sig"))
data = json.loads(path.read_text(encoding = "utf-8"))
return data if isinstance(data, dict) else {}
except FileNotFoundError:
return None
@ -2128,7 +2120,7 @@ def try_begin_sidecar_swap(kind: str = "install") -> bool:
break
if fd is not None:
try:
with os.fdopen(fd, "w", encoding = "utf-8") as f:
with os.fdopen(fd, "w") as f:
f.write(
json.dumps(
{"pid": os.getpid(), "at": time.time(), "token": token, "kind": kind}
@ -2474,11 +2466,7 @@ def _ensure_venv_llmcompressor_exists() -> bool:
stdout = subprocess.PIPE,
stderr = subprocess.STDOUT,
text = True,
encoding = "utf-8",
errors = "replace",
env = utf8_child_env(
get_hf_cache_paths().child_env(child_env_without_native_path_secret())
),
env = get_hf_cache_paths().child_env(child_env_without_native_path_secret()),
**_windows_hidden_subprocess_kwargs(),
)
last_out = result.stdout or ""

View file

@ -30,7 +30,6 @@ PYPI_SUCCESS_TTL_SECONDS = 12 * 60 * 60
PYPI_FAILURE_TTL_SECONDS = 60 * 60
RELEASE_NOTES_URL = "https://unsloth.ai/docs/new/changelog"
DISABLE_ENV_VAR = "UNSLOTH_DISABLE_UPDATE_CHECK"
FAKE_UPDATE_ENV_VAR = "UNSLOTH_STUDIO_FAKE_UPDATE"
LOCAL_INSTALL_SOURCES = {"editable", "local_path", "vcs", "local_repo"}
@ -108,32 +107,11 @@ def get_studio_install_source_status(current_version: str) -> dict[str, Any]:
)
def _is_version(value: str) -> bool:
try:
Version(value)
except InvalidVersion:
return False
return True
def get_studio_update_status(current_version: str) -> dict[str, Any]:
"""Return public, read-only update status for the web UI."""
install_source = detect_install_source()
disabled = os.environ.get(DISABLE_ENV_VAR) == "1"
# Dev-only: the popup is PyPI-install-only, so fake a version to review it
# from a checkout. The documented opt-out still wins.
forced_version = os.environ.get(FAKE_UPDATE_ENV_VAR, "").strip()
if forced_version and not disabled and _is_version(forced_version):
return _status_response(
current_version = current_version,
latest_version = forced_version,
install_source = "pypi",
update_available = True,
can_show_web_notification = True,
)
if disabled:
if os.environ.get(DISABLE_ENV_VAR) == "1":
return _status_response(
current_version = current_version,
latest_version = None,

View file

@ -114,8 +114,6 @@ def hf_cache_snapshot_dir(model_name: str) -> Optional[Path]:
snapshot = repo_dir / "snapshots" / commit
if snapshot.is_dir():
return snapshot
# UnicodeDecodeError is a ValueError, not an OSError: a torn refs
# file must keep meaning "not cached here", not fail the offline check.
except (OSError, UnicodeDecodeError):
continue
return None

View file

@ -15,7 +15,6 @@ import urllib.request
from typing import Callable
from utils.native_path_leases import child_env_without_native_path_secret
from utils.child_stdio import utf8_child_env
from utils.subprocess_compat import windows_hidden_subprocess_kwargs
_logger = logging.getLogger(__name__)
@ -44,8 +43,6 @@ def has_blackwell_gpu() -> bool:
stdout = subprocess.PIPE,
stderr = subprocess.DEVNULL,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = 10,
env = child_env_without_native_path_secret(),
)
@ -105,10 +102,8 @@ def probe_torch_wheel_env(*, timeout: int | None = None) -> dict[str, str] | Non
stdout = subprocess.PIPE,
stderr = subprocess.PIPE,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = timeout,
env = utf8_child_env(child_env_without_native_path_secret()),
env = child_env_without_native_path_secret(),
**windows_hidden_subprocess_kwargs(),
)
except subprocess.TimeoutExpired:
@ -206,8 +201,6 @@ def install_wheel(
stdout = subprocess.PIPE,
stderr = subprocess.STDOUT,
text = True,
encoding = "utf-8",
errors = "replace",
env = child_env_without_native_path_secret(),
)
attempts.append(("uv", result))
@ -220,10 +213,7 @@ def install_wheel(
stdout = subprocess.PIPE,
stderr = subprocess.STDOUT,
text = True,
encoding = "utf-8",
errors = "replace",
# Make the Python child emit the UTF-8 we decode above.
env = utf8_child_env(child_env_without_native_path_secret()),
env = child_env_without_native_path_secret(),
)
attempts.append(("pip", result))
return attempts

View file

@ -121,14 +121,7 @@ def _installed_whisper_version(binary: Optional[str]) -> Optional[str]:
if not binary:
return None
try:
proc = subprocess.run(
[binary, "--version"],
capture_output = True,
text = True,
encoding = "utf-8",
errors = "replace",
timeout = 20,
)
proc = subprocess.run([binary, "--version"], capture_output = True, text = True, timeout = 20)
except Exception: # pragma: no cover - defensive
return None
m = re.search(r"v?(\d+\.\d+\.\d+)", (proc.stderr or "") + (proc.stdout or ""))

View file

@ -214,8 +214,7 @@ function TauriUpdateLayer({
}
return (
// Capped like the browser stack: the download panel shares it, so both must fit.
<div className="pointer-events-none fixed bottom-4 right-4 z-[9998] flex max-h-[calc(100dvh_-_2rem)] flex-col items-end gap-2">
<div className="pointer-events-none fixed bottom-4 right-4 z-[9998] flex w-[calc(100vw-2rem)] max-w-[400px] flex-col items-stretch gap-2">
<UpdateBanner
status={update.status}
info={update.info}
@ -224,7 +223,6 @@ function TauriUpdateLayer({
isExternalServer={isExternalServer}
updatePolicyMode={update.updatePolicyMode}
manualReleaseUrl={update.manualReleaseUrl}
releasePageUrl={update.releasePageUrl}
positioned={false}
onInstall={update.installUpdate}
onDismiss={update.dismiss}
@ -381,11 +379,9 @@ function TauriWrapper({ children }: { children: ReactNode }) {
return (
<>
{children}
{/* One bottom-right stack so overlays never overlap: download panel at the
corner, banners above, each owning its width. */}
{/* Capped to the viewport, or a long download list plus expanded notes
pushes the top of the stack off screen. */}
<div className="pointer-events-none fixed bottom-4 right-4 z-[9998] flex max-h-[calc(100dvh_-_2rem)] flex-col items-end gap-2">
{/* One bottom-right stack so overlays never overlap; they stack with a
gap, download panel anchored at the corner with banners above. */}
<div className="pointer-events-none fixed bottom-4 right-4 z-[9998] flex w-[calc(100vw-2rem)] max-w-[400px] flex-col items-stretch gap-2">
<WebUpdateBanner
positioned={false}
enabled={!WEB_UPDATE_HIDDEN_ROUTES.has(pathname)}

View file

@ -1226,9 +1226,7 @@ export function AppSidebar() {
openNewChat(null);
}}
className={cn(
// min-w-0 so a narrow sidebar truncates the wordmark
// instead of pushing the search icon over the logo.
"flex min-w-0 items-center gap-[6px] select-none transition-opacity",
"flex items-center gap-[6px] select-none transition-opacity",
chatDisabled && "pointer-events-none opacity-50",
)}
aria-label={t("shell.aria.home")}
@ -1240,17 +1238,17 @@ export function AppSidebar() {
<img
src="/circle-logo-small.png"
alt="Unsloth"
className="h-[calc(26px+0.5rem*var(--ui-font-scale,1))] w-[calc(26px+0.5rem*var(--ui-font-scale,1))] shrink-0 rounded-full object-cover"
className="h-[calc(26px+0.5rem*var(--ui-font-scale,1))] w-[calc(26px+0.5rem*var(--ui-font-scale,1))] rounded-full object-cover"
/>
<span className="truncate font-heading text-[calc(13px+0.5rem*var(--ui-font-scale,1))] font-semibold tracking-[0em] leading-none text-black dark:text-white dark:tracking-[0.02em]">
<span className="font-heading text-[calc(13px+0.5rem*var(--ui-font-scale,1))] font-semibold tracking-[0em] leading-none text-black dark:text-white dark:tracking-[0.02em]">
unsloth
</span>
<span className="nav-badge ml-0.5 inline-flex shrink-0 items-center justify-center rounded-full border border-nav-beta-border px-[5px] pt-[3px] pb-[2px] text-[calc(0.5rem*var(--ui-font-scale,1))] font-medium leading-none tracking-[0.04em] text-nav-fg-muted antialiased subpixel-antialiased shadow-[0_1px_2px_rgba(0,0,0,0.06)] dark:shadow-[0_1px_2px_rgba(0,0,0,0.35)]">
<span className="nav-badge ml-0.5 inline-flex items-center justify-center rounded-full border border-nav-beta-border px-[5px] pt-[3px] pb-[2px] text-[calc(0.5rem*var(--ui-font-scale,1))] font-medium leading-none tracking-[0.04em] text-nav-fg-muted antialiased subpixel-antialiased shadow-[0_1px_2px_rgba(0,0,0,0.06)] dark:shadow-[0_1px_2px_rgba(0,0,0,0.35)]">
{t("shell.beta")}
</span>
</Link>
)}
<div className="flex shrink-0 items-center gap-0.25">
<div className="flex items-center gap-0.5">
<Tooltip>
<TooltipPrimitive.Trigger asChild>
<button
@ -1259,7 +1257,7 @@ export function AppSidebar() {
useChatSearchStore.getState().open();
closeMobileIfOpen();
}}
className="inline-flex h-[33px] w-[28px] cursor-pointer items-center justify-center rounded-[10px] text-nav-icon-idle dark:text-nav-fg-muted transition-colors hover:bg-nav-surface-hover hover:text-black dark:hover:text-white focus-visible:outline-none focus-visible:ring-1 focus-visible:ring-ring"
className="inline-flex h-[33px] w-[32px] cursor-pointer items-center justify-center rounded-[10px] text-nav-icon-idle dark:text-nav-fg-muted transition-colors hover:bg-nav-surface-hover hover:text-black dark:hover:text-white focus-visible:outline-none focus-visible:ring-1 focus-visible:ring-ring"
aria-label={t("shell.navigation.search")}
>
<HugeiconsIcon icon={Search01Icon} strokeWidth={1.75} className="size-icon" />
@ -1283,7 +1281,7 @@ export function AppSidebar() {
<button
type="button"
onClick={togglePinned}
className="inline-flex h-[33px] w-[28px] cursor-pointer items-center justify-center rounded-[10px] text-nav-icon-idle dark:text-nav-fg-muted transition-colors hover:bg-nav-surface-hover hover:text-black dark:hover:text-white focus-visible:outline-none focus-visible:ring-1 focus-visible:ring-ring"
className="inline-flex h-[33px] w-[32px] cursor-pointer items-center justify-center rounded-[10px] text-nav-icon-idle dark:text-nav-fg-muted transition-colors hover:bg-nav-surface-hover hover:text-black dark:hover:text-white focus-visible:outline-none focus-visible:ring-1 focus-visible:ring-ring"
aria-label={t("shell.aria.closeSidebar")}
>
<HugeiconsIcon icon={LayoutAlignLeftIcon} strokeWidth={1.75} className="size-icon" />
@ -1327,10 +1325,10 @@ export function AppSidebar() {
)}
</SidebarHeader>
{/* Uniform pl-1.5 pr-1.75 keeps every hover pill the same width, inset from the edge. */}
{/* Uniform pl-1.5 pr-2 keeps every hover pill the same width, inset from the edge. */}
<SidebarGroup
className={cn(
"group-data-[collapsible=icon]:px-0 pl-1.5 pr-1.75 shrink-0 transition-[padding]",
"group-data-[collapsible=icon]:px-0 pl-1.5 pr-2 shrink-0 transition-[padding]",
showCompactMacBrand ? "pt-0" : "pt-[9px]",
// Scrolled: New Chat is pinned, give a little gap below it.
scrolled ? "pb-[5px]" : "pb-px",
@ -1419,7 +1417,7 @@ export function AppSidebar() {
scrolled && "is-scrolled",
)}
>
<SidebarGroup className="group-data-[collapsible=icon]:px-0 pl-1.5 pr-1.75 py-0 shrink-0">
<SidebarGroup className="group-data-[collapsible=icon]:px-0 pl-1.5 pr-2 py-0 shrink-0">
<SidebarGroupContent>
<SidebarMenu>
<NavItem
@ -1501,7 +1499,7 @@ export function AppSidebar() {
</CollapsibleTrigger>
</SidebarGroupLabel>
<CollapsibleContent>
<SidebarGroupContent className="pl-1.5 pr-1.75">
<SidebarGroupContent className="pl-1.5 pr-2">
<SidebarMenu>
<NavItem
icon={TestTubeOutlineIcon}
@ -1576,7 +1574,7 @@ export function AppSidebar() {
</CollapsibleTrigger>
</SidebarGroupLabel>
<CollapsibleContent>
<SidebarGroupContent className="pl-1.5 pr-1.75">
<SidebarGroupContent className="pl-1.5 pr-2">
<SidebarMenu>
{pinnedProjectRecords.map((project) => {
const projectChats =
@ -1723,7 +1721,7 @@ export function AppSidebar() {
</CollapsibleTrigger>
</SidebarGroupLabel>
<CollapsibleContent>
<SidebarGroupContent className="pl-1.5 pr-1.75">
<SidebarGroupContent className="pl-1.5 pr-2">
<SidebarMenu>
{recentChatItems.map((item) =>
renderChatSidebarItem(item, "recent"),
@ -1755,7 +1753,7 @@ export function AppSidebar() {
</CollapsibleTrigger>
</SidebarGroupLabel>
<CollapsibleContent>
<SidebarGroupContent className="pl-1.5 pr-1.75">
<SidebarGroupContent className="pl-1.5 pr-2">
<SidebarMenu>
{runItems.map((run) => {
// Explicit selection wins. Otherwise highlight the active

View file

@ -134,7 +134,7 @@ export function LlamaUpdateBanner({
className={cn(
positioned
? "fixed bottom-4 right-4 z-[9998] w-[calc(100vw-2rem)] max-w-[400px]"
: "pointer-events-auto w-[calc(100vw-2rem)] max-w-[400px]",
: "pointer-events-auto w-full",
)}
data-testid="llama-update-banner"
>

View file

@ -2,7 +2,6 @@
// Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
import { Button } from "@/components/ui/button";
import { ReleaseNotesPanel } from "@/components/update/release-notes-panel";
import type {
DesktopUpdatePolicyMode,
RetainedUpdateFailure,
@ -23,8 +22,6 @@ interface UpdateBannerProps {
isExternalServer?: boolean;
updatePolicyMode: DesktopUpdatePolicyMode;
manualReleaseUrl: string | null;
// Release page for this version, preferred over the generic changelog.
releasePageUrl?: string | null;
// false fills a shared overlay stack; true self-anchors.
positioned?: boolean;
onInstall: () => void;
@ -33,7 +30,6 @@ interface UpdateBannerProps {
}
const EASE_OUT_QUART: [number, number, number, number] = [0.165, 0.84, 0.44, 1];
const LEADING_V = /^v/;
function formatVersion(version: string | null | undefined): string {
if (!version) return "";
@ -48,7 +44,6 @@ export function UpdateBanner({
isExternalServer = false,
updatePolicyMode,
manualReleaseUrl,
releasePageUrl = null,
positioned = true,
onInstall,
onDismiss,
@ -57,8 +52,6 @@ export function UpdateBanner({
const [copying, setCopying] = useState(false);
const [manualReport, setManualReport] = useState<string | null>(null);
const [manualMessage, setManualMessage] = useState<string | null>(null);
// Version whose notes are expanded; a new offer collapses the panel.
const [notesVersion, setNotesVersion] = useState<string | null>(null);
const showFailure = Boolean(lastFailure) && !dismissed;
const showAvailable = status === "available" && !dismissed && !showFailure;
const show = showFailure || (showAvailable && Boolean(info));
@ -69,11 +62,6 @@ export function UpdateBanner({
const currentVersion = formatVersion(info?.currentVersion);
const latestVersion = formatVersion(info?.version);
const Icon = showFailure ? CircleAlert : Download;
// Keyed by the backend release, not the app's SemVer; headings drop the v.
const notesTargetVersion =
(info?.pypiVersion ?? info?.version)?.replace(LEADING_V, "") ?? null;
const notesOpen =
notesTargetVersion !== null && notesVersion === notesTargetVersion;
async function handleCopyDiagnostics() {
setCopying(true);
@ -106,14 +94,13 @@ export function UpdateBanner({
exit={{ opacity: 0, y: 8, scale: 0.97 }}
transition={{ duration: 0.35, ease: EASE_OUT_QUART }}
className={cn(
// Wider than the other overlays: notes preview plus three buttons.
positioned
? "fixed bottom-4 right-4 z-[9999] w-[calc(100vw-2rem)] max-w-[448px]"
: "pointer-events-auto flex min-h-0 w-[calc(100vw-2rem)] max-w-[448px] flex-col",
? "fixed bottom-4 right-4 z-[9999] w-[calc(100vw-2rem)] max-w-[400px]"
: "pointer-events-auto w-full",
)}
data-testid="tauri-update-banner"
>
<div className="relative flex max-h-[calc(100dvh_-_2rem)] flex-col overflow-hidden rounded-[24px] bg-white px-5 pb-4 pt-5 shadow-[0_2px_8px_-2px_rgba(0,0,0,0.16)] dark:bg-card dark:shadow-[0_8px_28px_-6px_rgba(0,0,0,0.28)]">
<div className="relative overflow-hidden rounded-[24px] bg-white px-5 pb-4 pt-5 shadow-[0_2px_8px_-2px_rgba(0,0,0,0.16)] dark:bg-card dark:shadow-[0_8px_28px_-6px_rgba(0,0,0,0.28)]">
<button
type="button"
onClick={onDismiss}
@ -173,40 +160,7 @@ export function UpdateBanner({
</p>
)}
{!showFailure && notesTargetVersion ? (
<ReleaseNotesPanel
version={notesTargetVersion}
open={notesOpen}
// Used only if CHANGELOG.md has no section for this version.
fallbackMarkdown={info?.body ?? null}
className="min-h-0 flex-1"
releaseNotesUrl={releasePageUrl ?? manualReleaseUrl}
/>
) : null}
<div
className={cn(
"mt-4 flex flex-wrap items-center gap-x-1 gap-y-2",
!showFailure && notesTargetVersion
? "justify-between"
: "justify-end",
)}
>
{!showFailure && notesTargetVersion ? (
<Button
size="sm"
variant="ghost"
// same type size as the action buttons
className="-ml-2 h-auto whitespace-nowrap rounded-full px-2.5 py-2 text-ui-13 font-medium text-foreground"
onClick={() =>
setNotesVersion(notesOpen ? null : notesTargetVersion)
}
aria-expanded={notesOpen}
data-testid="tauri-update-release-notes-toggle"
>
{notesOpen ? "Hide release notes" : "Show release notes"}
</Button>
) : null}
<div className="mt-4 flex flex-wrap items-center justify-end gap-x-1 gap-y-2">
{showFailure ? (
<>
<Button
@ -233,31 +187,28 @@ export function UpdateBanner({
onClick={onInstall}
disabled={installDisabled}
>
{isManualLinuxPackage
? "Open release page"
: "Retry update"}
{isManualLinuxPackage ? "Open release page" : "Retry update"}
</Button>
</>
) : (
// wrap + right-align so the action pair stays together
<div className="flex flex-wrap items-center justify-end gap-x-1 gap-y-2">
<>
<Button
size="sm"
variant="ghost"
className="h-auto whitespace-nowrap rounded-full px-2.5 py-2 text-ui-13 font-medium text-foreground"
className="h-auto rounded-full px-3 py-2 text-ui-13 font-medium text-foreground"
onClick={onDismiss}
>
Remind me later
</Button>
<Button
size="sm"
className="-mr-1 h-auto whitespace-nowrap rounded-full px-3 py-2 text-ui-13"
className="-mr-1 h-auto rounded-full px-3.5 py-2 text-ui-13"
onClick={onInstall}
disabled={installDisabled}
>
{isManualLinuxPackage ? "Open release page" : "Update"}
</Button>
</div>
</>
)}
</div>
{manualMessage && (

View file

@ -2,7 +2,6 @@
// Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
import { useSidebarPin } from "@/hooks/use-sidebar-pin";
import { useSidebarWidth } from "@/hooks/use-sidebar-width";
import { isTauri } from "@/lib/api-base";
import { cn } from "@/lib/utils";
import {
@ -111,13 +110,9 @@ export function WindowTitlebar({
const [enabled] = useState(shouldUseCustomWindowTitlebar);
const [maximized, setMaximized] = useState(false);
const { pinned, togglePinned } = useSidebarPin();
// The titlebar sits outside the sidebar wrapper, so it cannot inherit
// --sidebar-width. Read the resized width from the same store instead.
const { width } = useSidebarWidth();
const sidebarWidth = showSidebarSurface
? pinned
? // The live value only exists mid-drag; otherwise the committed width.
`var(--studio-sidebar-live-width, ${width}px)`
? "var(--studio-sidebar-expanded-width,17.5rem)"
: "var(--studio-sidebar-collapsed-width,3rem)"
: "0px";
const contentBorderLeft = pinned ? `calc(${sidebarWidth} + 12px)` : "0px";

View file

@ -1,317 +0,0 @@
// SPDX-License-Identifier: AGPL-3.0-only
// Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"use client"
import * as React from "react"
import { cn } from "@/lib/utils"
import {
Tooltip,
TooltipContent,
TooltipTrigger,
} from "@/components/ui/tooltip"
import { getClientPlatform } from "@/components/tauri/window-titlebar"
/** Pointer travel (px) below which a drag counts as a plain click. */
const DRAG_SLOP = 4
/** A compatibility click lands immediately after pointer-up. */
const CLICK_COMPAT_WINDOW_MS = 300
/** Arrow-key resize step for keyboard users. */
const RESIZE_STEP = 16
type DragState = {
startX: number
startWidth: number
moved: boolean
}
export type PanelResizeHandleProps = {
/** Which edge of the panel the handle sits on. */
edge: "left" | "right"
open: boolean
width: number
/** Uncapped stored preference, so a capped drag does not lower it. */
stored: number
min: number
max: number
clamp: (px: number) => number
setWidth: (px: number) => void
resetWidth: () => void
onToggle: () => void
/** Element to paint the live width onto, and the property to paint. */
target: () => HTMLElement | null
cssVar: string
/** Measured to start a drag from the rendered size when collapsed. */
measure: () => number
label: string
toggleLabel: string
/** Translated tooltip copy; the caller owns the translation layer. */
collapseHint: string
expandHint: string
dragHint: string
/** Shown in the tooltip when the panel has a toggle shortcut. */
shortcut?: string
dataSlot?: string
className?: string
/** Mirrors the live width onto :root for chrome outside the panel. */
rootVar?: string
}
/**
* A draggable panel edge: drag to resize, click to collapse or expand. Arrow
* keys resize, Home restores the default. The width is painted straight to the
* target while dragging and only persisted on release.
*/
export function PanelResizeHandle({
edge,
open,
width,
stored,
min,
max,
clamp,
setWidth,
resetWidth,
onToggle,
target,
cssVar,
measure,
label,
toggleLabel,
collapseHint,
expandHint,
dragHint,
shortcut,
dataSlot = "panel-resize-handle",
className,
rootVar,
}: PanelResizeHandleProps) {
const ref = React.useRef<HTMLButtonElement>(null)
const dragRef = React.useRef<DragState | null>(null)
const [dragging, setDragging] = React.useState(false)
const [hovered, setHovered] = React.useState(false)
const [focused, setFocused] = React.useState(false)
const [isMacPlatform] = React.useState(() => getClientPlatform().includes("mac"))
const hint = shortcut ? shortcut.replace("Mod", isMacPlatform ? "⌘" : "Ctrl+") : null
// Cached on pointer down so no DOM walk per move.
const targetRef = React.useRef<HTMLElement | null>(null)
const frameRef = React.useRef(0)
const pendingRef = React.useRef(0)
// What the pointer asked for, before the viewport cap. Committing the capped
// value instead would quietly downgrade a stored preference on a narrow window.
const rawRef = React.useRef(0)
// When a pointer sequence last ended. The browser's compatibility click
// lands in the same tick, so only a click that close behind is a duplicate.
// A timestamp cannot go stale the way an armed flag does: a genuine cancel
// emits no click, and a later assistive-tech click still gets through.
const handledAtRef = React.useRef(0)
const committedRef = React.useRef(width)
React.useEffect(() => {
committedRef.current = width
}, [width])
const paint = React.useCallback(
(value: string) => {
targetRef.current?.style.setProperty(cssVar, value)
if (rootVar) {
document.documentElement.style.setProperty(rootVar, value)
}
},
[cssVar, rootVar],
)
// Resizing relayouts the whole shell, and pointermove fires faster than the
// display refreshes, so coalesce to one paint per frame.
const paintWidth = React.useCallback(
(px: number) => {
pendingRef.current = px
if (frameRef.current) return
frameRef.current = requestAnimationFrame(() => {
frameRef.current = 0
paint(`${pendingRef.current}px`)
})
},
[paint],
)
const endDrag = React.useCallback(() => {
// Only a sequence that actually started can produce a compatibility click.
// This also runs as the effect cleanup, where no drag happened.
if (dragRef.current) handledAtRef.current = Date.now()
dragRef.current = null
if (frameRef.current) {
cancelAnimationFrame(frameRef.current)
frameRef.current = 0
}
// Hand the property back to the committed value. A commit re-renders with
// the new width; a cancel or a no-commit drag keeps DOM and store in step.
paint(`${committedRef.current}px`)
if (rootVar) document.documentElement.style.removeProperty(rootVar)
targetRef.current?.removeAttribute("data-resizing")
document.documentElement.removeAttribute("data-panel-resizing")
targetRef.current = null
setDragging(false)
document.body.style.removeProperty("cursor")
document.body.style.removeProperty("user-select")
}, [paint, rootVar])
const handlePointerDown = (event: React.PointerEvent<HTMLButtonElement>) => {
if (event.button !== 0) return
event.preventDefault()
event.currentTarget.setPointerCapture(event.pointerId)
targetRef.current = target()
// Collapsed: grow from the rendered size so the edge tracks the pointer.
const start = open ? width : measure()
dragRef.current = { startX: event.clientX, startWidth: start, moved: false }
pendingRef.current = start
rawRef.current = start
targetRef.current?.setAttribute("data-resizing", "true")
document.documentElement.setAttribute("data-panel-resizing", "true")
setDragging(true)
document.body.style.setProperty("cursor", "col-resize")
document.body.style.setProperty("user-select", "none")
}
const handlePointerMove = (event: React.PointerEvent<HTMLButtonElement>) => {
const drag = dragRef.current
if (!drag) return
// A panel whose handle is on its left edge grows as the pointer moves left.
const delta = (edge === "left" ? -1 : 1) * (event.clientX - drag.startX)
if (!drag.moved && Math.abs(delta) < DRAG_SLOP) return
drag.moved = true
const next = drag.startWidth + delta
rawRef.current = next
if (!open) {
// Past the minimum, dragging the collapsed edge reopens it.
if (next >= min) {
paintWidth(clamp(next))
onToggle()
}
return
}
// Dragging inward stops at the minimum. Collapsing is click or the shortcut.
paintWidth(clamp(next))
}
const handlePointerUp = (event: React.PointerEvent<HTMLButtonElement>) => {
const drag = dragRef.current
if (!drag) return
if (event.currentTarget.hasPointerCapture(event.pointerId)) {
event.currentTarget.releasePointerCapture(event.pointerId)
}
endDrag()
if (!drag.moved) {
onToggle()
return
}
// A drag below the minimum leaves the stored width alone.
if (!open) return
// Capped: the visible edge is already at the cap, so an outward pull cannot
// express intent beyond it. Committing would silently lower the larger
// hidden preference. A deliberate inward drag still commits.
if (stored > max && rawRef.current >= max) return
// Commit what was asked for, not the capped paint, so a drag on a narrow
// window cannot shrink a larger stored preference. setWidth clamps.
setWidth(rawRef.current)
}
const handleKeyDown = (event: React.KeyboardEvent<HTMLButtonElement>) => {
// The collapse/expand the label advertises, for keyboard users. Pointer-up
// handles it for the mouse; a synthesized click never reaches it.
if (event.key === "Enter" || event.key === " ") {
// preventDefault cancels the native click, so nothing follows to guard
// against; arming here would swallow the next assistive-tech click.
event.preventDefault()
onToggle()
return
}
const outward = edge === "left" ? "ArrowLeft" : "ArrowRight"
const inward = edge === "left" ? "ArrowRight" : "ArrowLeft"
if (event.key === outward || event.key === inward) {
event.preventDefault()
if (!open) {
// Collapsed there is nothing to resize, so the outward arrow reopens.
if (event.key === outward) onToggle()
return
}
if (event.key === outward && stored > max && width >= max) return
setWidth(width + (event.key === outward ? RESIZE_STEP : -RESIZE_STEP))
return
}
if (event.key === "Home") {
event.preventDefault()
resetWidth()
}
}
// Clear a stuck cursor override if we unmount mid-drag.
React.useEffect(() => endDrag, [endDrag])
return (
<Tooltip open={(hovered || focused) && !dragging}>
<TooltipTrigger asChild>
<button
ref={ref}
type="button"
data-slot={dataSlot}
data-dragging={dragging || undefined}
aria-label={open ? label : toggleLabel}
{...(open ? { "aria-orientation": "vertical" as const } : {})}
{...(open
? { "aria-valuenow": width, "aria-valuemin": min, "aria-valuemax": max }
: {})}
role={open ? "separator" : "button"}
onPointerDown={handlePointerDown}
onPointerMove={handlePointerMove}
onPointerUp={handlePointerUp}
onPointerCancel={endDrag}
onKeyDown={handleKeyDown}
onClick={() => {
// Switch and voice control activate by dispatching a bare click
// with no pointer or key events, which nothing else here catches.
if (Date.now() - handledAtRef.current < CLICK_COMPAT_WINDOW_MS) return
onToggle()
}}
onPointerEnter={() => setHovered(true)}
onPointerLeave={() => setHovered(false)}
onFocus={(event) => setFocused(event.target.matches(":focus-visible"))}
onBlur={() => setFocused(false)}
className={cn(
"absolute inset-y-0 z-30 hidden w-2 touch-none select-none sm:block",
edge === "left" ? "-left-1" : "-right-1",
// `!` overrides the app-wide hand cursor on buttons.
open
? "cursor-col-resize!"
: edge === "left"
? "cursor-w-resize!"
: "cursor-e-resize!",
// Sits exactly on the panel border so hover recolours one line.
"after:absolute after:inset-y-0 after:w-px after:bg-transparent after:transition-colors after:duration-150",
edge === "left" ? "after:left-1" : "after:right-1",
"hover:after:bg-sidebar-ring/25 data-dragging:after:bg-sidebar-ring/25",
// The app zeroes the native outline on buttons, so mark focus here.
"focus-visible:outline-none focus-visible:after:bg-sidebar-ring/60",
className,
)}
/>
</TooltipTrigger>
<TooltipContent
side={edge === "left" ? "left" : "right"}
align="center"
className="tooltip-compact"
>
<span className="flex flex-col gap-px">
<span>
{open ? collapseHint : expandHint}
{hint ? ` ${hint}` : ""}
</span>
<span className="opacity-70">{dragHint}</span>
</span>
</TooltipContent>
</Tooltip>
)
}

View file

@ -24,21 +24,13 @@ import {
TooltipContent,
TooltipTrigger,
} from "@/components/ui/tooltip"
import { PanelResizeHandle } from "@/components/ui/panel-resize-handle"
import { useT } from "@/i18n"
import { useIsMobile } from "@/hooks/use-mobile"
import {
SIDEBAR_WIDTH_DEFAULT,
SIDEBAR_WIDTH_MIN,
clampSidebarWidth,
useSidebarWidth,
} from "@/hooks/use-sidebar-width"
import { HugeiconsIcon } from "@hugeicons/react"
import { LayoutAlignLeftIcon } from "@hugeicons/core-free-icons"
const noop = () => {}
const SIDEBAR_WIDTH = `${SIDEBAR_WIDTH_DEFAULT}px`
const SIDEBAR_WIDTH = "17.5rem"
const SIDEBAR_WIDTH_ICON = "3rem"
const SIDEBAR_KEYBOARD_SHORTCUT = "b"
@ -54,11 +46,6 @@ type SidebarContextProps = {
pinned: boolean
setPinned: (value: boolean) => void
togglePinned: () => void
width: number
storedWidth: number
maxWidth: number
setWidth: (value: number) => void
resetWidth: () => void
}
const SidebarContext = React.createContext<SidebarContextProps | null>(null)
@ -93,13 +80,6 @@ function SidebarProvider({
}) {
const isMobile = useIsMobile()
const [openMobile, setOpenMobile] = React.useState(false)
const {
width,
max: maxWidth,
stored: storedWidth,
setWidth,
resetWidth,
} = useSidebarWidth()
const prevIsMobileRef = React.useRef(isMobile)
React.useEffect(() => {
@ -183,13 +163,8 @@ function SidebarProvider({
pinned,
setPinned,
togglePinned,
width,
storedWidth,
maxWidth,
setWidth,
resetWidth,
}),
[state, open, setOpen, isMobile, openMobile, setOpenMobile, toggleSidebar, hasPinMode, pinned, setPinned, togglePinned, width, storedWidth, maxWidth, setWidth, resetWidth]
[state, open, setOpen, isMobile, openMobile, setOpenMobile, toggleSidebar, hasPinMode, pinned, setPinned, togglePinned]
)
return (
@ -198,8 +173,7 @@ function SidebarProvider({
data-slot="sidebar-wrapper"
style={
{
// The drag handle writes this same property live while resizing.
"--sidebar-width": `${width}px`,
"--sidebar-width": SIDEBAR_WIDTH,
"--sidebar-width-icon": SIDEBAR_WIDTH_ICON,
...style,
} as React.CSSProperties
@ -337,64 +311,11 @@ function Sidebar({
>
{children}
</div>
<SidebarResizeHandle side={side} />
</div>
</div>
)
}
/**
* The sidebar's draggable edge, over the shared panel handle.
*/
function SidebarResizeHandle({
className,
side = "left",
}: {
className?: string
side?: "left" | "right"
}) {
const { open, toggleSidebar, width, storedWidth, maxWidth, setWidth, resetWidth } =
useSidebar()
const ref = React.useRef<HTMLDivElement>(null)
const t = useT()
return (
<div ref={ref} className="contents">
<PanelResizeHandle
edge={side === "right" ? "left" : "right"}
open={open}
width={width}
stored={storedWidth}
min={SIDEBAR_WIDTH_MIN}
max={maxWidth}
clamp={clampSidebarWidth}
setWidth={setWidth}
resetWidth={resetWidth}
onToggle={toggleSidebar}
target={() =>
ref.current?.closest<HTMLElement>('[data-slot="sidebar-wrapper"]') ?? null
}
cssVar="--sidebar-width"
// The custom titlebar renders outside the wrapper and cannot inherit it.
rootVar="--studio-sidebar-live-width"
measure={() =>
ref.current
?.closest<HTMLElement>('[data-slot="sidebar-container"]')
?.getBoundingClientRect().width ?? SIDEBAR_WIDTH_MIN
}
label={t("shell.aria.resizeSidebar")}
toggleLabel={t("shell.aria.openSidebar")}
collapseHint={t("shell.resize.collapse")}
expandHint={t("shell.resize.expand")}
dragHint={t("shell.resize.drag")}
shortcut="ModB"
dataSlot="sidebar-resize-handle"
className={className}
/>
</div>
)
}
function SidebarTrigger({
className,
onClick,
@ -856,7 +777,6 @@ export {
SidebarMenuSubItem,
SidebarProvider,
SidebarRail,
SidebarResizeHandle,
SidebarSeparator,
SidebarTrigger,
useSidebar,

View file

@ -1,251 +0,0 @@
// SPDX-License-Identifier: AGPL-3.0-only
// Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
import { MarkdownPreview } from "@/components/markdown/markdown-preview";
import { useReleaseNotes } from "@/hooks/use-release-notes";
import { resolveChangelogLinks } from "@/lib/changelog-links";
import { releaseNotesPreview } from "@/lib/release-notes-preview";
import { cn } from "@/lib/utils";
import {
type ReactElement,
type ReactNode,
useEffect,
useMemo,
useRef,
} from "react";
interface ReleaseNotesPanelProps {
// Notes are looked up for this exact version only.
version: string;
// Collapsed previews the top bullets; expanded scrolls the full notes.
open: boolean;
// Desktop updater's body, used only if CHANGELOG.md has no section here.
fallbackMarkdown?: string | null;
releaseNotesUrl?: string | null;
className?: string;
}
const NOTES_LINK_CLASS =
"shrink-0 whitespace-nowrap text-ui-11 font-medium text-foreground underline underline-offset-2";
function NotesMessage({
children,
action,
}: {
children: ReactNode;
action?: ReactNode;
}): ReactElement {
return (
<div className="flex items-center justify-between gap-2 px-1 py-2">
<p className="text-ui-11 text-muted-foreground">{children}</p>
{action}
</div>
);
}
function ChangelogLink({ href }: { href: string }): ReactElement {
return (
<a
href={href}
target="_blank"
rel="noopener noreferrer"
className={NOTES_LINK_CLASS}
data-testid="update-release-notes-link"
>
Open changelog
</a>
);
}
export function ReleaseNotesPanel({
version,
open,
fallbackMarkdown = null,
releaseNotesUrl = null,
className,
}: ReleaseNotesPanelProps): ReactElement | null {
// Fetched with the popup: the collapsed preview needs the notes too.
const { state, notes, retry } = useReleaseNotes({ version, enabled: true });
const scrollRef = useRef<HTMLElement | null>(null);
// The fallback stands in for "no section in the changelog", which the hook
// reports as ready. An error is retryable, and the desktop fallback is the
// updater's static blurb, so taking it there would hide Retry until cache expiry.
const source = notes?.matched
? notes.markdown
: state === "error"
? null
: (fallbackMarkdown ?? null);
// Notes target the repository, so relative links must point back at it.
const markdown = useMemo(
() => (source === null ? null : resolveChangelogLinks(source)),
[source],
);
// Notes that are only a code block or a table preview as nothing.
const preview = useMemo(
() => (markdown === null ? null : releaseNotesPreview(markdown)),
[markdown],
);
// Start at the top on expand, and again once async notes land.
useEffect(() => {
if (open && markdown && scrollRef.current) {
scrollRef.current.scrollTop = 0;
}
}, [open, markdown]);
// Caller's URL wins: the API returns only the generic changelog, while the
// desktop banner passes this version's release page.
const notesUrl = releaseNotesUrl ?? notes?.releaseNotesUrl;
const link = notesUrl ? <ChangelogLink href={notesUrl} /> : null;
// Nothing previewable yet or ever: keep the collapsed popup compact.
if (
!open &&
(!markdown ||
state === "loading" ||
state === "idle" ||
preview?.items.length === 0)
) {
return null;
}
return (
<div
className={cn("mt-3 flex min-h-0 flex-col", className)}
data-testid="update-release-notes-panel"
data-notes-state={state}
data-notes-version={version}
data-notes-open={open}
>
{/* borderless fill, lighter than the card in dark mode */}
<div className="flex min-h-0 flex-col rounded-[14px] bg-muted/40 px-3 py-1 dark:bg-white/[0.06]">
{markdown ? (
open ? (
<section
ref={scrollRef}
// biome-ignore lint/a11y/noNoninteractiveTabindex: keyboard-scrollable region
tabIndex={0}
aria-label={`Release notes for version ${version}`}
// Long notes scroll here instead of pushing the buttons off screen.
className="hover-scrollbar max-h-64 min-h-0 flex-1 overflow-y-auto overscroll-contain py-3 pr-1"
data-testid="update-release-notes-scroll"
>
<MarkdownPreview
markdown={markdown}
// Streamdown ships headings at mt-6 and code at text-sm, and
// clears max-width on descendants, so rescale and re-cap both.
className="max-h-none overflow-visible border-0 bg-transparent p-0 text-ui-11 [&_[data-streamdown=link-safety-modal]>*]:max-w-md [&_img]:h-auto [&_img]:max-w-full [&>*:first-child]:mt-0 [&>*>*:first-child]:mt-0 [&_code]:text-[0.92em] [&_h1]:mt-4 [&_h1]:font-heading [&_h1]:text-ui-13 [&_h2]:mt-4 [&_h2]:font-heading [&_h2]:text-ui-13 [&_h3]:mt-4 [&_h3]:font-heading [&_h3]:text-ui-11 [&_pre]:text-[0.92em]"
/>
{notes?.truncated ? (
<p className="mt-2 text-ui-10 text-muted-foreground/80">
Notes truncated. See the full changelog.
</p>
) : null}
</section>
) : (
<ReleaseNotesSummary preview={preview} />
)
) : (
<NotesStatus
state={state}
version={version}
link={link}
retry={retry}
/>
)}
</div>
{open && markdown && link ? (
<div className="mt-2 flex justify-end px-1">{link}</div>
) : null}
</div>
);
}
/** Collapsed view: the first few bullets, one line each where possible. */
function ReleaseNotesSummary({
preview,
}: {
preview: ReturnType<typeof releaseNotesPreview> | null;
}): ReactElement | null {
if (preview === null || preview.items.length === 0) {
return null;
}
const { items, remaining } = preview;
return (
<ul
className="space-y-1 py-2 pr-1"
data-testid="update-release-notes-summary"
>
{items.map((item, index) => (
<li
// Two releases can carry the same bullet text, so index is the key.
key={`${index}-${item.lead}`}
className="flex gap-1.5 text-ui-11 leading-snug text-muted-foreground"
>
<span aria-hidden="true" className="text-muted-foreground/60">
&bull;
</span>
<span className="line-clamp-2 min-w-0">
{/* lead sentence carries the change */}
<span className="font-medium text-foreground">{item.lead}</span>
{item.rest ? <span> {item.rest}</span> : null}
</span>
</li>
))}
{remaining > 0 ? (
<li className="pl-3 text-ui-10 text-muted-foreground/70">
+{remaining} more
</li>
) : null}
</ul>
);
}
function NotesStatus({
state,
version,
link,
retry,
}: {
state: ReturnType<typeof useReleaseNotes>["state"];
version: string;
link: ReactNode;
retry: () => void;
}): ReactElement {
if (state === "loading" || state === "idle") {
return <NotesMessage>Loading release notes...</NotesMessage>;
}
if (state === "error") {
return (
<NotesMessage
action={
// The changelog page may be reachable when the lookup is not.
<span className="flex shrink-0 items-center gap-3">
<button
type="button"
onClick={retry}
className={NOTES_LINK_CLASS}
data-testid="update-release-notes-retry"
>
Retry
</button>
{link}
</span>
}
>
Could not load release notes.
</NotesMessage>
);
}
// Matched nothing: link out rather than show another release's notes.
return (
<NotesMessage action={link}>
No release notes published for {version} yet.
</NotesMessage>
);
}

View file

@ -2,7 +2,6 @@
// Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
import { Button } from "@/components/ui/button";
import { ReleaseNotesPanel } from "@/components/update/release-notes-panel";
import { type DeviceType, usePlatformStore } from "@/config/env";
import { useWebUpdateCheck } from "@/hooks/use-web-update-check";
import { isTauri } from "@/lib/api-base";
@ -41,7 +40,6 @@ export function WebUpdateBanner({
const deviceType = usePlatformStore((s) => s.deviceType);
const installCmd = installCommandForDevice(deviceType);
const [copiedVersion, setCopiedVersion] = useState<string | null>(null);
const [notesVersion, setNotesVersion] = useState<string | null>(null);
const dismissTimerRef = useRef<ReturnType<typeof setTimeout> | null>(null);
useEffect(() => {
@ -70,8 +68,6 @@ export function WebUpdateBanner({
}
const copied = status != null && copiedVersion === status.latestVersion;
// Keyed by version so a new offer collapses the panel.
const notesOpen = status != null && notesVersion === status.latestVersion;
return (
<AnimatePresence>
@ -82,14 +78,13 @@ export function WebUpdateBanner({
exit={{ opacity: 0, y: 8, scale: 0.97 }}
transition={{ duration: 0.35, ease: EASE_OUT_QUART }}
className={cn(
// Wider than the other overlays: notes preview plus three buttons.
positioned
? "fixed bottom-4 right-4 z-[9999] w-[calc(100vw-2rem)] max-w-[448px]"
: "pointer-events-auto flex min-h-0 w-[calc(100vw-2rem)] max-w-[448px] flex-col",
? "fixed bottom-4 right-4 z-[9999] w-[calc(100vw-2rem)] max-w-[400px]"
: "pointer-events-auto w-full",
)}
data-testid="web-update-banner"
>
<div className="relative flex max-h-[calc(100dvh_-_2rem)] flex-col overflow-hidden rounded-[24px] bg-white px-5 pb-4 pt-5 shadow-[0_2px_8px_-2px_rgba(0,0,0,0.16)] dark:bg-card dark:shadow-[0_8px_28px_-6px_rgba(0,0,0,0.28)]">
<div className="relative overflow-hidden rounded-[24px] bg-white px-5 pb-4 pt-5 shadow-[0_2px_8px_-2px_rgba(0,0,0,0.16)] dark:bg-card dark:shadow-[0_8px_28px_-6px_rgba(0,0,0,0.28)]">
<button
type="button"
onClick={dismiss}
@ -132,33 +127,22 @@ export function WebUpdateBanner({
</div>
</div>
<ReleaseNotesPanel
version={status.latestVersion}
open={notesOpen}
releaseNotesUrl={RELEASE_NOTES_URL}
className="min-h-0 flex-1"
/>
{/* one row at one type size; wraps only on narrow viewports */}
<div className="mt-4 flex flex-wrap items-center justify-between gap-y-2">
<Button
size="sm"
variant="ghost"
className="-ml-2 h-auto whitespace-nowrap rounded-full px-2.5 py-2 text-ui-13 font-medium text-foreground"
onClick={() =>
setNotesVersion(notesOpen ? null : status.latestVersion)
}
aria-expanded={notesOpen}
data-testid="web-update-release-notes-toggle"
<a
href={RELEASE_NOTES_URL}
target="_blank"
rel="noopener noreferrer"
className="-ml-2 whitespace-nowrap rounded-full px-2.5 py-2 text-ui-13 font-medium text-foreground transition-colors hover:bg-muted"
data-testid="web-update-release-notes-link"
>
{notesOpen ? "Hide release notes" : "Show release notes"}
</Button>
Release notes
</a>
{/* wrap + right-align so buttons stack instead of clipping on very narrow banners */}
<div className="flex flex-wrap items-center justify-end gap-x-1 gap-y-2">
<Button
size="sm"
variant="ghost"
className="h-auto whitespace-nowrap rounded-full px-2.5 py-2 text-ui-13 font-medium text-foreground"
className="h-auto rounded-full px-3 py-2 text-ui-13 font-medium text-foreground"
onClick={snooze}
data-testid="web-update-snooze-button"
>
@ -167,7 +151,7 @@ export function WebUpdateBanner({
<Button
size="sm"
// -mr optically aligns the filled pill's edge with the card padding
className="-mr-1 h-auto whitespace-nowrap rounded-full px-3 py-2 text-ui-13"
className="-mr-1 h-auto rounded-full px-3.5 py-2 text-ui-13"
onClick={handleCopyCommand}
data-testid="web-update-copy-button"
>

View file

@ -18,7 +18,6 @@ import {
DropdownMenuTrigger,
} from "@/components/ui/dropdown-menu";
import { InfoHint } from "@/components/ui/info-hint";
import { PanelResizeHandle } from "@/components/ui/panel-resize-handle";
import {
InputGroup,
InputGroupAddon,
@ -45,13 +44,7 @@ import { Tooltip, TooltipContent } from "@/components/ui/tooltip";
import { NumericValueInput, snapToStep } from "@/features/model-picker";
import { RetrievalSettingsSection } from "@/features/rag";
import { useLlamaUpdateCheck } from "@/hooks/use-llama-update-check";
import {
CHAT_SETTINGS_WIDTH_MIN,
clampChatSettingsWidth,
useChatSettingsWidth,
} from "@/hooks/use-chat-settings-width";
import { useIsMobile } from "@/hooks/use-mobile";
import { useT } from "@/i18n";
import { ChevronDownStandardIcon } from "@/lib/chevron-icons";
import { toast } from "@/lib/toast";
import { cn } from "@/lib/utils";
@ -59,7 +52,7 @@ import { Edit03Icon, LayoutAlignRightIcon } from "@hugeicons/core-free-icons";
import { HugeiconsIcon } from "@hugeicons/react";
import { Braces, ChevronDown, ExternalLink } from "lucide-react";
import { Tooltip as TooltipPrimitive } from "radix-ui";
import { type CSSProperties, Fragment, type ReactNode } from "react";
import { Fragment, type ReactNode } from "react";
import { useCallback, useEffect, useMemo, useRef, useState } from "react";
import { OpenAICodeExecSection } from "./components/openai-code-exec-section";
import { PermissionModeDropdown } from "./permission-mode-select";
@ -370,15 +363,6 @@ export function ChatSettingsPanel({
onExternalProviderChange,
externalProviderType = null,
}: ChatSettingsPanelProps) {
const asideRef = useRef<HTMLElement>(null);
const t = useT();
const {
width: settingsWidth,
max: settingsMax,
stored: settingsStored,
setWidth: setSettingsWidth,
resetWidth: resetSettingsWidth,
} = useChatSettingsWidth();
// Local models show every knob; providerCapabilities is only consulted when
// isExternalModel. Unknown providers fall back to the OpenAI-compat shape via
// getProviderCapabilities, so these flags never undercount support.
@ -477,23 +461,6 @@ export function ChatSettingsPanel({
// When the prompt overflows the inline box, clicking opens the popup editor.
const systemPromptBoxRef = useRef<HTMLTextAreaElement>(null);
const [systemPromptOverflows, setSystemPromptOverflows] = useState(false);
const promptObserverRef = useRef<ResizeObserver | null>(null);
const measurePromptRef = useRef<() => void>(() => {});
// The section unmounts its textarea when collapsed, so observe through a
// callback ref: a stored observer would cling to the detached node and the
// remounted one would never be measured.
const attachPromptBox = useCallback((node: HTMLTextAreaElement | null) => {
systemPromptBoxRef.current = node;
promptObserverRef.current?.disconnect();
promptObserverRef.current = null;
if (!node || typeof ResizeObserver === "undefined") return;
// Resizing rewraps the prompt, and a drag changes the width through a
// custom property without re-rendering, so watch the box itself.
const observer = new ResizeObserver(() => measurePromptRef.current());
observer.observe(node);
promptObserverRef.current = observer;
measurePromptRef.current();
}, []);
const [activePresetBaseline, setActivePresetBaseline] = useState(params);
const presets = useMemo(() => {
return getOrderedPresets(customPresets);
@ -779,20 +746,15 @@ export function ChatSettingsPanel({
}, [open]);
useEffect(() => {
measurePromptRef.current = () => {
const el = systemPromptBoxRef.current;
setSystemPromptOverflows(
currentSystemPrompt.length > 0 &&
el != null &&
el.clientHeight > 0 &&
el.scrollHeight > el.clientHeight + 1,
);
};
measurePromptRef.current();
const el = systemPromptBoxRef.current;
setSystemPromptOverflows(
currentSystemPrompt.length > 0 &&
el != null &&
el.clientHeight > 0 &&
el.scrollHeight > el.clientHeight + 1,
);
}, [currentSystemPrompt, open]);
useEffect(() => () => promptObserverRef.current?.disconnect(), []);
const settingsScrollRef = useRef<HTMLDivElement>(null);
const settingsContent = (
@ -1162,7 +1124,7 @@ export function ChatSettingsPanel({
)}
>
<textarea
ref={attachPromptBox}
ref={systemPromptBoxRef}
value={currentSystemPrompt}
onChange={(e) => set("systemPrompt")(e.target.value)}
onMouseDown={(e) => {
@ -1471,47 +1433,17 @@ export function ChatSettingsPanel({
return (
<aside
ref={asideRef}
data-tour="chat-settings"
data-slot="chat-settings-panel"
className={cn(
"relative z-50 shrink-0 bg-panel-surface text-panel-surface-fg font-heading",
open
? "w-(--chat-settings-width) border-l border-sidebar-border"
: "w-0 overflow-hidden",
"relative z-50 shrink-0 overflow-hidden bg-panel-surface text-panel-surface-fg font-heading",
open ? "w-[17rem] border-l border-sidebar-border" : "w-0",
)}
style={
{
"--chat-settings-width": `${settingsWidth}px`,
height: "calc(100% - var(--studio-custom-titlebar-height, 0px))",
marginTop: "var(--studio-custom-titlebar-height, 0px)",
} as CSSProperties
}
style={{
height: "calc(100% - var(--studio-custom-titlebar-height, 0px))",
marginTop: "var(--studio-custom-titlebar-height, 0px)",
}}
>
{open ? (
<PanelResizeHandle
edge="left"
open={open}
width={settingsWidth}
stored={settingsStored}
min={CHAT_SETTINGS_WIDTH_MIN}
max={settingsMax}
clamp={clampChatSettingsWidth}
setWidth={setSettingsWidth}
resetWidth={resetSettingsWidth}
onToggle={() => onOpenChange?.(!open)}
target={() => asideRef.current}
cssVar="--chat-settings-width"
measure={() => asideRef.current?.getBoundingClientRect().width ?? 0}
label={t("shell.aria.resizeRunSettings")}
toggleLabel={t("shell.aria.openRunSettings")}
collapseHint={t("shell.resize.collapse")}
expandHint={t("shell.resize.expand")}
dragHint={t("shell.resize.drag")}
dataSlot="chat-settings-resize-handle"
/>
) : null}
<div className="h-full w-full overflow-hidden">{settingsContent}</div>
<div className="h-full w-full">{settingsContent}</div>
</aside>
);
}

View file

@ -14,8 +14,7 @@ import {
} from "@/components/ui/dialog";
import { Input } from "@/components/ui/input";
import { cn } from "@/lib/utils";
import { ChevronDownStandardIcon } from "@/lib/chevron-icons";
import { XIcon } from "lucide-react";
import { ChevronDownIcon, XIcon } from "lucide-react";
import { type KeyboardEvent, useState } from "react";
import { useChatRuntimeStore } from "../stores/chat-runtime-store";
import type { ResearchWebsitePolicy } from "../types/research";
@ -25,12 +24,7 @@ function normalizeDomain(raw: string): string | null {
if (!value || /[\\\s]/.test(value)) return null;
try {
const url = new URL(value.includes("://") ? value : `https://${value}`);
if (
!/^https?:$/.test(url.protocol) ||
url.username ||
url.password ||
url.port
) {
if (!/^https?:$/.test(url.protocol) || url.username || url.password || url.port) {
return null;
}
return url.hostname
@ -105,9 +99,7 @@ function DomainList({
type="button"
className="text-muted-foreground transition-colors hover:text-foreground"
aria-label={`Remove ${domain}`}
onClick={() =>
onChange(values.filter((value) => value !== domain))
}
onClick={() => onChange(values.filter((value) => value !== domain))}
>
<XIcon className="size-3" />
</button>
@ -137,9 +129,7 @@ export function DeepResearchComposerButton({
onConfigure: () => void;
}) {
const enabled = useChatRuntimeStore((state) => state.deepResearchEnabled);
const setEnabled = useChatRuntimeStore(
(state) => state.setDeepResearchEnabled,
);
const setEnabled = useChatRuntimeStore((state) => state.setDeepResearchEnabled);
if (!enabled) return null;
@ -168,12 +158,9 @@ export function DeepResearchComposerButton({
<XIcon className="composer-pill-x" />
</span>
<span>Deep research</span>
{/* Same caret as the other composer pills, so the arrows match. */}
<HugeiconsIcon
icon={ChevronDownStandardIcon}
strokeWidth={1.5}
className="composer-pill-caret size-[15px] text-primary/70"
/>
<span className="composer-pill-caret flex items-center gap-0.5 text-primary/70">
<ChevronDownIcon className="size-3" />
</span>
</button>
);
}
@ -186,9 +173,7 @@ export function DeepResearchWebsiteAccessDialog({
onOpenChange: (open: boolean) => void;
}) {
const policy = useChatRuntimeStore((state) => state.researchWebsitePolicy);
const setPolicy = useChatRuntimeStore(
(state) => state.setResearchWebsitePolicy,
);
const setPolicy = useChatRuntimeStore((state) => state.setResearchWebsitePolicy);
return (
<Dialog open={open} onOpenChange={onOpenChange}>
@ -216,40 +201,41 @@ function DeepResearchWebsiteAccessContent({
return (
<DialogContent className="sm:max-w-lg">
<DialogHeader>
<DialogTitle>Website access</DialogTitle>
<DialogDescription>
Control which websites the next Deep Research run can search and read.
Limits are enforced by the server and shared with the research model.
</DialogDescription>
</DialogHeader>
<div className="space-y-6">
<DomainList
label="Allow only"
description="When set, research can access only these domains and their subdomains."
values={draft.allowedDomains}
onChange={(allowedDomains) => setDraft({ ...draft, allowedDomains })}
/>
<DomainList
label="Always block"
description="These domains and their subdomains stay blocked. Blocking takes precedence."
values={draft.blockedDomains}
onChange={(blockedDomains) => setDraft({ ...draft, blockedDomains })}
/>
</div>
<DialogFooter>
<Button variant="ghost" onClick={onClose}>
Cancel
</Button>
<Button
onClick={() => {
setPolicy(draft);
onClose();
}}
>
Save limits
</Button>
</DialogFooter>
<DialogHeader>
<DialogTitle>Website access</DialogTitle>
<DialogDescription>
Control which websites the next Deep Research run can search and
read. Limits are enforced by the server and shared with the research
model.
</DialogDescription>
</DialogHeader>
<div className="space-y-6">
<DomainList
label="Allow only"
description="When set, research can access only these domains and their subdomains."
values={draft.allowedDomains}
onChange={(allowedDomains) => setDraft({ ...draft, allowedDomains })}
/>
<DomainList
label="Always block"
description="These domains and their subdomains stay blocked. Blocking takes precedence."
values={draft.blockedDomains}
onChange={(blockedDomains) => setDraft({ ...draft, blockedDomains })}
/>
</div>
<DialogFooter>
<Button variant="ghost" onClick={onClose}>
Cancel
</Button>
<Button
onClick={() => {
setPolicy(draft);
onClose();
}}
>
Save limits
</Button>
</DialogFooter>
</DialogContent>
);
}

View file

@ -23,21 +23,12 @@ import {
removeScanFolder,
} from "@/features/hub";
import { FolderBrowser } from "@/features/model-picker";
import {
openModelsDir,
pickHuggingFaceCacheDir,
} from "@/features/native-intents";
import {
type HuggingFaceCacheSettings,
loadHuggingFaceCacheSettings,
updateHuggingFaceCacheSettings,
} from "@/features/settings";
import { openModelsDir } from "@/features/native-intents";
import { isTauri } from "@/lib/api-base";
import { toast } from "@/lib/toast";
import { cn } from "@/lib/utils";
import {
Delete02Icon,
DownloadCircle01Icon,
FileSearchIcon,
FolderAddIcon,
FolderExportIcon,
@ -58,12 +49,6 @@ function formatError(error: unknown): string {
return error instanceof Error ? error.message : String(error);
}
function formatFreeSpace(bytes: number | null): string | null {
if (bytes === null || !Number.isFinite(bytes)) return null;
const gb = bytes / 1024 ** 3;
return gb >= 10 ? `${Math.round(gb)} GB free` : `${gb.toFixed(1)} GB free`;
}
export function OnDeviceFoldersDialog({
open,
onOpenChange,
@ -83,11 +68,6 @@ export function OnDeviceFoldersDialog({
);
const refreshIdRef = useRef(0);
const mutationVersionRef = useRef(0);
const [downloadCache, setDownloadCache] =
useState<HuggingFaceCacheSettings | null>(null);
const [downloadCacheLoaded, setDownloadCacheLoaded] = useState(false);
const [downloadBrowserOpen, setDownloadBrowserOpen] = useState(false);
const [downloadSaving, setDownloadSaving] = useState(false);
const sortedFolders = useMemo(
() => [...folders].sort((a, b) => a.path.localeCompare(b.path)),
@ -128,66 +108,10 @@ export function OnDeviceFoldersDialog({
return () => window.clearTimeout(timer);
}, [open, refreshFolders]);
useEffect(() => {
if (!open) return;
let cancelled = false;
// The dialog stays mounted between opens, so re-arm the flag or a reopen
// shows the previous answer as if it were fresh.
setDownloadCacheLoaded(false);
loadHuggingFaceCacheSettings()
// Indexed locations do not depend on this. Null drops the stale path
// rather than offer Change against a location we could not confirm.
.catch(() => null)
.then((settings) => {
if (cancelled) return;
setDownloadCache(settings);
setDownloadCacheLoaded(true);
});
return () => {
cancelled = true;
};
}, [open]);
const handleInventoryChanged = useCallback(() => {
onInventoryChange?.();
}, [onInventoryChange]);
// Relocating the cache changes which repos are on disk, but
// updateHuggingFaceCacheSettings already bumps the inventory version, which
// re-fetches every source. Refreshing here too would scan twice, since the
// two rounds carry different version keys and cannot be deduplicated.
const saveDownloadLocation = useCallback(async (nextPath: string | null) => {
setDownloadSaving(true);
try {
const settings = await updateHuggingFaceCacheSettings(nextPath);
setDownloadCache(settings);
toast.success("Download location updated", {
description: settings.cacheHome,
});
} catch (err) {
toast.error("Couldn't update the download location", {
description: formatError(err),
});
} finally {
setDownloadSaving(false);
}
}, []);
const changeDownloadLocation = useCallback(async () => {
if (!isTauri) {
setDownloadBrowserOpen(true);
return;
}
try {
const picked = await pickHuggingFaceCacheDir();
if (picked) await saveDownloadLocation(picked);
} catch (err) {
toast.error("Couldn't open the folder picker", {
description: formatError(err),
});
}
}, [saveDownloadLocation]);
const handleAdd = useCallback(
async (rawPath: string) => {
const nextPath = rawPath.trim();
@ -258,10 +182,10 @@ export function OnDeviceFoldersDialog({
<>
<Dialog open={open} onOpenChange={onOpenChange}>
<DialogContent
className="flex max-h-[90dvh] flex-col gap-0 overflow-hidden p-0 sm:max-w-[620px] lg:max-w-[660px] xl:max-w-[680px] [&_[data-slot=dialog-close]]:right-3 [&_[data-slot=dialog-close]]:top-3"
className="gap-0 overflow-hidden p-0 sm:max-w-[620px] lg:max-w-[660px] xl:max-w-[680px] [&_[data-slot=dialog-close]]:right-3 [&_[data-slot=dialog-close]]:top-3"
overlayClassName="bg-black/20 backdrop-blur-none"
>
<DialogHeader className="shrink-0 border-b border-border/60 px-5 py-4">
<DialogHeader className="border-b border-border/60 px-5 py-4">
<DialogTitle className="text-ui-15">
On-device locations
</DialogTitle>
@ -271,78 +195,7 @@ export function OnDeviceFoldersDialog({
</DialogDescription>
</DialogHeader>
<div className="min-h-0 flex-1 space-y-4 overflow-y-auto px-5 py-4">
<div className="rounded-[14px] border border-border/70 bg-muted/20 p-3">
<div className="mb-2 flex items-center gap-2 text-ui-12 font-medium text-foreground">
<HugeiconsIcon
icon={DownloadCircle01Icon}
strokeWidth={1.75}
className="size-3.5 text-muted-foreground"
/>
Download location
</div>
<div className="flex flex-col gap-2 sm:flex-row sm:items-center">
<Input
readOnly={true}
aria-label="Model download location"
value={
downloadCache?.cacheHome ??
(downloadCacheLoaded ? "Unknown" : "Loading...")
}
title={downloadCache?.cacheHome}
className="field-soft h-9 min-w-0 flex-1 rounded-full px-3 font-mono text-ui-12"
/>
<div className="flex shrink-0 items-center gap-2">
<Button
type="button"
variant="outline"
size="sm"
onClick={() => void changeDownloadLocation()}
disabled={!downloadCache?.editable || downloadSaving}
className="h-9 rounded-full px-3 text-ui-12p5"
>
{downloadSaving ? (
<Spinner className="size-3.5" />
) : (
<HugeiconsIcon
icon={FolderSearchIcon}
strokeWidth={1.75}
data-icon="inline-start"
className="size-3.5"
/>
)}
Change
</Button>
{downloadCache?.isCustom ? (
<Button
type="button"
variant="ghost"
size="sm"
onClick={() => void saveDownloadLocation(null)}
disabled={downloadSaving}
className="h-9 rounded-full px-3 text-ui-12p5 text-muted-foreground"
>
Use default
</Button>
) : null}
</div>
</div>
<p className="mt-2 text-ui-10p5 text-muted-foreground">
{downloadCache?.source === "environment"
? `Managed by the ${
downloadCache.environmentVariable ?? "HF_HOME"
} environment variable.`
: [
"New downloads only. Models already on disk stay where they are.",
formatFreeSpace(downloadCache?.freeBytes ?? null),
]
.filter(Boolean)
.join(" · ")}
</p>
</div>
<div className="space-y-4 px-5 py-4">
<div className="rounded-[14px] border border-border/70 bg-muted/20 p-3">
<div className="mb-2 flex items-center gap-2 text-ui-12 font-medium text-foreground">
<HugeiconsIcon
@ -572,16 +425,6 @@ export function OnDeviceFoldersDialog({
onOpenChange={setBrowserOpen}
onSelect={(selectedPath) => void handleAdd(selectedPath)}
/>
<FolderBrowser
open={!isTauri && downloadBrowserOpen}
onOpenChange={setDownloadBrowserOpen}
onSelect={(selectedPath) => void saveDownloadLocation(selectedPath)}
initialPath={downloadCache?.cacheHome}
title="Choose model download location"
confirmLabel="Use for future downloads"
showModelHints={false}
/>
</>
);
}

View file

@ -201,10 +201,8 @@ export function DownloadManagerPanel({
className={cn(
// Standalone: anchor bottom-right. In a shared stack (positioned=false)
// flow as a right-aligned row so overlays stack instead of overlapping.
// min-h-0 there: a flex item's min-height defaults to auto, so the capped
// stack would squeeze the update card instead of this list.
"pointer-events-none",
positioned ? "fixed bottom-4 right-4 z-50" : "flex min-h-0 justify-end",
positioned ? "fixed bottom-4 right-4 z-50" : "flex justify-end",
)}
>
{collapsed ? (
@ -231,7 +229,7 @@ export function DownloadManagerPanel({
</TooltipContent>
</Tooltip>
) : (
<div className="hub-download-panel pointer-events-auto flex min-h-0 w-[min(400px,calc(100vw-2rem))] flex-col overflow-hidden">
<div className="hub-download-panel pointer-events-auto w-[min(400px,calc(100vw-2rem))] overflow-hidden">
<div className="flex items-center gap-2 border-b border-foreground/[0.07] py-2 pl-4 pr-3">
<span className="min-w-0 flex-1 truncate text-ui-12p5 font-semibold text-foreground">
{headerLabel}

View file

@ -0,0 +1,102 @@
// SPDX-License-Identifier: AGPL-3.0-only
// Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
import { authFetch } from "@/features/auth";
import { readFastApiError } from "@/lib/format-fastapi-error";
/** One day of the activity series. Dense: every day in range is present. */
export type ProfileStatsDay = {
date: string;
tokens: number;
messages: number;
chats: number;
};
export type ProfileStatsModel = {
id: string;
label: string;
messages: number;
tokens: number;
};
export type ProfileStatsRun = {
id: string;
name: string;
modelLabel: string;
datasetLabel: string;
status: string;
finalLoss: number | null;
steps: number;
seconds: number;
startedAt: string | null;
};
export type ProfileStats = {
generatedAt: number;
days: number;
totals: {
threads: number;
messages: number;
userMessages: number;
assistantMessages: number;
promptTokens: number;
completionTokens: number;
totalTokens: number;
cachedTokens: number;
toolCalls: number;
attachments: number;
activeDays: number;
chatSeconds: number;
};
streak: {
current: number;
longest: number;
lastActiveDay: string | null;
};
peakDay: { date: string; tokens: number } | null;
longestChat: {
threadId: string | null;
title: string | null;
seconds: number;
messages: number;
} | null;
daily: ProfileStatsDay[];
models: ProfileStatsModel[];
speed: {
averageTokensPerSecond: number | null;
bestTokensPerSecond: number | null;
bestTokensPerSecondModel: string | null;
averageResponseMs: number | null;
averageFirstTokenMs: number | null;
samples: number;
};
training: {
runs: number;
completed: number;
steps: number;
tokens: number;
seconds: number;
models: number;
datasets: number;
bestLoss: number | null;
recent: ProfileStatsRun[];
};
};
export async function loadProfileStats(
signal?: AbortSignal,
): Promise<ProfileStats> {
// Bucket days and hours in this browser's timezone, which is not the
// server's when Studio is reached over the network. The IANA name is what
// gives each historical date its own daylight-saving offset; the current
// offset only covers callers whose host cannot resolve the name.
const query = new URLSearchParams({
tz_offset_minutes: String(new Date().getTimezoneOffset()),
tz: Intl.DateTimeFormat().resolvedOptions().timeZone ?? "",
});
const res = await authFetch(`/api/profile/stats?${query}`, { signal });
if (!res.ok) {
throw new Error(await readFastApiError(res, "Failed to load your stats"));
}
return (await res.json()) as ProfileStats;
}

View file

@ -5,14 +5,23 @@ import { publicAssetUrl } from "@/components/mascot-img";
import { Button } from "@/components/ui/button";
import { Input } from "@/components/ui/input";
import { Label } from "@/components/ui/label";
import { Switch } from "@/components/ui/switch";
import {
Popover,
PopoverContent,
PopoverTrigger,
} from "@/components/ui/popover";
import { getAuthToken } from "@/features/auth";
import { cn } from "@/lib/utils";
import { useT } from "@/i18n";
import { toastError, toastSuccess } from "@/shared/toast";
import { Edit03Icon } from "@hugeicons/core-free-icons";
import {
Delete02Icon,
Edit03Icon,
Image01Icon,
Upload01Icon,
} from "@hugeicons/core-free-icons";
import { HugeiconsIcon } from "@hugeicons/react";
import { useEffect, useMemo, useRef, useState } from "react";
import { useEffect, useRef, useState } from "react";
import { SLOTH_AVATARS } from "../sloth-avatars";
import { decodeJwtSubject } from "../utils/jwt-subject";
import { resizeImageFileToDataUrl } from "../utils/resize-image-file";
@ -23,6 +32,8 @@ import {
import { UserAvatar } from "./user-avatar";
const PROFILE_STORAGE_KEY = "unsloth_user_profile";
const SLOTH_NAME = /^large\s+/i;
const PNG_SUFFIX = /\.png$/i;
function readPersistedProfile(): {
displayName: string;
@ -36,7 +47,8 @@ function readPersistedProfile(): {
if (!parsed || typeof parsed !== "object") return null;
// Zustand persist shape: { state: {...}, version }
const maybeState = "state" in parsed ? (parsed as { state?: unknown }).state : parsed;
const maybeState =
"state" in parsed ? (parsed as { state?: unknown }).state : parsed;
if (!maybeState || typeof maybeState !== "object") return null;
const state = maybeState as {
displayName?: unknown;
@ -45,9 +57,11 @@ function readPersistedProfile(): {
};
return {
displayName: typeof state.displayName === "string" ? state.displayName : "",
displayName:
typeof state.displayName === "string" ? state.displayName : "",
nickname: typeof state.nickname === "string" ? state.nickname : "",
avatarDataUrl: typeof state.avatarDataUrl === "string" ? state.avatarDataUrl : null,
avatarDataUrl:
typeof state.avatarDataUrl === "string" ? state.avatarDataUrl : null,
};
} catch {
return null;
@ -64,28 +78,17 @@ export function ProfilePersonalizationPanel() {
const setAvatarDataUrl = useUserProfileStore((s) => s.setAvatarDataUrl);
const avatarShape = useUserProfileStore((s) => s.avatarShape);
const setAvatarShape = useUserProfileStore((s) => s.setAvatarShape);
const showGreetingSloth = useUserProfileStore((s) => s.showGreetingSloth);
const setShowGreetingSloth = useUserProfileStore(
(s) => s.setShowGreetingSloth,
);
const [imageError, setImageError] = useState<string | null>(null);
const [draftName, setDraftName] = useState(displayName);
const [draftNickname, setDraftNickname] = useState(nickname);
const [pickerOpen, setPickerOpen] = useState(false);
const fileInputRef = useRef<HTMLInputElement>(null);
const lastDisplayNameRef = useRef(displayName);
const lastNicknameRef = useRef(nickname);
const sessionSub = decodeJwtSubject(getAuthToken()) ?? "";
const previewName = draftName.trim() || sessionSub || "Unsloth";
const hasNameChanges = useMemo(
() => draftName.trim() !== displayName.trim(),
[draftName, displayName],
);
const hasNicknameChanges = useMemo(
() => draftNickname.trim() !== nickname.trim(),
[draftNickname, nickname],
);
useEffect(() => {
const previous = lastDisplayNameRef.current;
@ -99,6 +102,8 @@ export function ProfilePersonalizationPanel() {
setDraftNickname((draft) => (draft === previous ? nickname : draft));
}, [nickname]);
// Committed on blur and on Enter rather than behind a Save button, so each
// field is a single row like the rest of Settings.
const saveName = () => {
const trimmed = draftName.trim();
if (trimmed !== draftName) setDraftName(trimmed);
@ -133,6 +138,18 @@ export function ProfilePersonalizationPanel() {
}
};
// Escape, or any programmatic close, unmounts the tab without dispatching a
// blur, which would drop whatever was typed. Commit the drafts on the way
// out; both saves no-op on an unchanged value, so a double commit is safe.
const flushDrafts = useRef<() => void>(() => {});
useEffect(() => {
flushDrafts.current = () => {
saveName();
saveNickname();
};
});
useEffect(() => () => flushDrafts.current(), []);
const applyAvatar = (value: string | null) => {
setAvatarDataUrl(value);
const persisted = readPersistedProfile();
@ -180,201 +197,232 @@ export function ProfilePersonalizationPanel() {
requestAnimationFrame(() => applyAvatar(value));
};
const pickSloth = (path: string) => {
pickAvatarValue(publicAssetUrl(path));
};
return (
<div className="mx-auto flex w-full max-w-[640px] flex-col items-center gap-6 rounded-2xl border border-border/70 bg-muted/10 px-8 py-7">
<div className="relative">
<UserAvatar
name={previewName}
imageUrl={shownAvatar}
size="lg"
className="size-[124px] text-[calc(3.15rem*var(--ui-font-scale,1))]"
/>
<input
ref={fileInputRef}
type="file"
accept="image/jpeg,image/png,image/webp,image/gif"
className="sr-only"
onChange={(e) => {
void onPickFile(e.target.files?.[0]);
e.target.value = "";
}}
/>
<button
type="button"
onClick={() => fileInputRef.current?.click()}
className="absolute right-0 bottom-0 -translate-x-[15.625%] -translate-y-[15.625%] flex size-8 items-center justify-center rounded-full border border-border bg-background text-foreground shadow-[0_2px_8px_-2px_rgba(0,0,0,0.16)] transition-colors hover:bg-muted focus-visible:outline-none focus-visible:ring-1 focus-visible:ring-ring"
aria-label={t("settings.profile.changePicture")}
>
<HugeiconsIcon icon={Edit03Icon} className="size-3.5" strokeWidth={2} />
</button>
</div>
<div className="flex w-full flex-col">
<input
ref={fileInputRef}
type="file"
accept="image/jpeg,image/png,image/webp,image/gif"
className="sr-only"
onChange={(e) => {
void onPickFile(e.target.files?.[0]);
e.target.value = "";
}}
/>
<div
data-settings-label={t("settings.profile.displayName")}
className="flex w-full max-w-[560px] flex-col gap-2"
>
<Label htmlFor="profile-display-name" className="text-xs font-medium text-muted-foreground">
{t("settings.profile.displayName")}
</Label>
<div className="flex items-center gap-2">
<Input
id="profile-display-name"
type="text"
value={draftName}
maxLength={PROFILE_TEXT_MAX_LENGTH}
onChange={(e) => setDraftName(e.target.value)}
onKeyDown={(e) => {
if (e.key === "Enter") {
e.preventDefault();
saveName();
}
}}
autoComplete="off"
placeholder={sessionSub || "Unsloth"}
className="h-10 min-w-0 flex-1 rounded-full text-sm"
/>
<Button type="button" size="sm" className="h-10 px-5" onClick={saveName} disabled={!hasNameChanges}>
{t("common.save")}
</Button>
</div>
</div>
<div
data-settings-label={t("settings.profile.nickname")}
className="flex w-full max-w-[560px] flex-col gap-2"
>
<Label htmlFor="profile-nickname" className="text-xs font-medium text-muted-foreground">
{t("settings.profile.nickname")}
</Label>
<div className="flex items-center gap-2">
<Input
id="profile-nickname"
type="text"
value={draftNickname}
maxLength={PROFILE_TEXT_MAX_LENGTH}
onChange={(e) => setDraftNickname(e.target.value)}
onKeyDown={(e) => {
if (e.key === "Enter") {
e.preventDefault();
saveNickname();
}
}}
autoComplete="off"
placeholder={t("settings.profile.nicknamePlaceholder")}
className="h-10 min-w-0 flex-1 rounded-full text-sm"
/>
<Button type="button" size="sm" className="h-10 px-5" onClick={saveNickname} disabled={!hasNicknameChanges}>
{t("common.save")}
</Button>
</div>
</div>
<div
data-settings-label={t("settings.profile.avatarShape")}
className="flex w-full max-w-[560px] flex-col gap-2"
>
<Label className="text-xs font-medium text-muted-foreground">
{t("settings.profile.avatarShape")}
</Label>
<div className="hub-tab-toggle inline-flex h-8 w-fit items-center rounded-full">
{(["circle", "rounded"] as const).map((shape) => (
<button
key={shape}
type="button"
onClick={() => setAvatarShape(shape)}
aria-pressed={avatarShape === shape}
className={cn(
"inline-flex h-8 items-center rounded-full px-4 text-ui-13 font-medium transition-colors",
avatarShape === shape
? "hub-tab-toggle-pill text-foreground"
: "text-muted-foreground hover:text-foreground",
)}
>
{shape === "circle"
? t("settings.profile.avatarShapeCircle")
: t("settings.profile.avatarShapeRounded")}
</button>
))}
</div>
</div>
<div
data-settings-label={t("settings.profile.greetingSloth")}
className="flex w-full max-w-[560px] items-center justify-between gap-4"
>
<div className="flex min-w-0 flex-col gap-0.5">
<Label
htmlFor="profile-greeting-sloth"
className="text-xs font-medium text-muted-foreground"
>
{t("settings.profile.greetingSloth")}
</Label>
<p className="text-xs text-muted-foreground/75">
{t("settings.profile.greetingSlothDescription")}
</p>
</div>
<Switch
id="profile-greeting-sloth"
checked={showGreetingSloth}
onCheckedChange={setShowGreetingSloth}
/>
</div>
<div className="flex w-full max-w-[560px] flex-col gap-2">
<Label className="text-xs font-medium text-muted-foreground">
{t("settings.profile.chooseSloth")}
</Label>
<div className="grid grid-cols-7 gap-2 sm:grid-cols-9">
{SLOTH_AVATARS.map((path) => {
const url = publicAssetUrl(path);
const selected = shownAvatar === url;
const label =
path.split("/").pop()?.replace(/\.png$/i, "").replace(/^large\s+/i, "").trim() ??
"sloth";
return (
<button
key={path}
type="button"
onClick={() => pickSloth(path)}
aria-pressed={selected}
aria-label={label}
title={label}
className={cn(
// No transition here: animating the ring makes the old
// icon's selection border linger when switching sloths.
"relative aspect-square overflow-hidden rounded-full bg-muted ring-1 ring-border hover:ring-ring focus-visible:outline-none focus-visible:ring-ring",
// Selection keeps the 1px weight, only darker.
selected && "ring-ring-strong hover:ring-ring-strong",
)}
>
<img src={url} alt="" loading="lazy" className="size-full object-cover" />
</button>
);
})}
<div className="flex items-center gap-10 py-6 pr-2">
<div className="relative shrink-0">
{/* The picture itself is the shortcut to "upload a photo"; the pencil
opens the rest of the options. */}
<button
type="button"
onClick={() => pickAvatarValue(null)}
aria-pressed={shownAvatar === null}
aria-label={t("settings.profile.noPicture")}
title={t("settings.profile.noPicture")}
className={cn(
"relative flex aspect-square items-center justify-center overflow-hidden rounded-full bg-muted text-muted-foreground ring-1 ring-border hover:ring-ring focus-visible:outline-none focus-visible:ring-ring",
shownAvatar === null && "ring-ring-strong hover:ring-ring-strong",
)}
onClick={() => fileInputRef.current?.click()}
aria-label={t("settings.profile.changePicture")}
className="group relative block rounded-full focus-visible:outline-none focus-visible:ring-1 focus-visible:ring-ring"
>
<span className="text-ui-11 font-medium">
{t("settings.profile.noneLabel")}
<UserAvatar
name={previewName}
imageUrl={shownAvatar}
size="lg"
className="size-[128px] text-[calc(3.2rem*var(--ui-font-scale,1))]"
/>
<span className="absolute inset-0 flex items-center justify-center rounded-full bg-black/45 opacity-0 transition-opacity group-hover:opacity-100">
<HugeiconsIcon
icon={Image01Icon}
className="size-8 text-white"
strokeWidth={2}
/>
</span>
</button>
<Popover open={pickerOpen} onOpenChange={setPickerOpen}>
<PopoverTrigger asChild={true}>
<button
type="button"
aria-label={t("settings.profile.pictureOptions")}
title={t("settings.profile.pictureOptions")}
className="absolute top-[85.36%] left-[85.36%] flex size-9 -translate-x-1/2 -translate-y-1/2 items-center justify-center rounded-full border border-border bg-background text-foreground shadow-[0_2px_8px_-2px_rgba(0,0,0,0.16)] transition-colors hover:bg-muted focus-visible:outline-none focus-visible:ring-1 focus-visible:ring-ring dark:border-transparent dark:bg-white/[0.14] dark:hover:bg-white/20"
>
<HugeiconsIcon
icon={Edit03Icon}
className="size-4.5"
strokeWidth={2}
/>
</button>
</PopoverTrigger>
<PopoverContent
align="start"
sideOffset={10}
className="w-[320px] gap-4 p-4"
>
<div className="flex items-center justify-between gap-3">
<span className="text-ui-11 font-medium uppercase tracking-wide text-muted-foreground">
{t("settings.profile.avatarShape")}
</span>
<div className="hub-tab-toggle flex h-8 shrink-0 items-center rounded-full">
{(["circle", "rounded"] as const).map((shape) => (
<button
key={shape}
type="button"
onClick={() => setAvatarShape(shape)}
aria-pressed={avatarShape === shape}
className={cn(
"inline-flex h-8 items-center justify-center rounded-full px-3.5 text-ui-13 font-medium transition-colors",
avatarShape === shape
? "hub-tab-toggle-pill text-foreground"
: "text-muted-foreground hover:text-foreground",
)}
>
{shape === "circle"
? t("settings.profile.avatarShapeCircle")
: t("settings.profile.avatarShapeRounded")}
</button>
))}
</div>
</div>
<div className="flex items-center gap-2">
<Button
type="button"
variant="outline"
onClick={() => fileInputRef.current?.click()}
className="h-9 w-fit gap-2 rounded-full px-4 text-sm"
>
<HugeiconsIcon
icon={Upload01Icon}
className="size-4"
strokeWidth={2}
/>
{t("settings.profile.uploadPhoto")}
</Button>
<Button
type="button"
variant="ghost"
onClick={() => pickAvatarValue(null)}
disabled={shownAvatar === null}
aria-label={t("settings.profile.removePhoto")}
title={t("settings.profile.removePhoto")}
className="size-9 shrink-0 rounded-full p-0 text-muted-foreground"
>
<HugeiconsIcon
icon={Delete02Icon}
className="size-4"
strokeWidth={2}
/>
</Button>
</div>
<div className="flex flex-col gap-2">
<span className="text-ui-11 font-medium uppercase tracking-wide text-muted-foreground">
{t("settings.profile.chooseSloth")}
</span>
<div className="grid grid-cols-7 gap-2">
{SLOTH_AVATARS.map((path) => {
const url = publicAssetUrl(path);
const selected = shownAvatar === url;
const label =
path
.split("/")
.pop()
?.replace(PNG_SUFFIX, "")
.replace(SLOTH_NAME, "")
.trim() ?? "sloth";
return (
<button
key={path}
type="button"
onClick={() => pickAvatarValue(url)}
aria-pressed={selected}
aria-label={label}
title={label}
className={cn(
// No transition here: animating the ring makes the old
// icon's selection border linger when switching sloths.
"relative aspect-square overflow-hidden rounded-full bg-muted ring-1 ring-border hover:ring-ring focus-visible:outline-none focus-visible:ring-ring",
selected &&
"ring-2 ring-ring-strong hover:ring-ring-strong",
)}
>
<img
src={url}
alt=""
loading="lazy"
className="size-full object-cover"
/>
</button>
);
})}
</div>
</div>
</PopoverContent>
</Popover>
</div>
{/* Name fields sit beside the picture. These are not SettingsRows, so
data-settings-label is set by hand for settings search. */}
<div className="flex min-w-0 flex-1 flex-col gap-3">
<div
data-settings-label={t("settings.profile.displayName")}
className="flex min-w-0 flex-col gap-1.5"
>
<Label
htmlFor="profile-display-name"
className="text-xs font-medium text-muted-foreground"
>
{t("settings.profile.displayName")}
</Label>
<Input
id="profile-display-name"
type="text"
value={draftName}
maxLength={PROFILE_TEXT_MAX_LENGTH}
onChange={(e) => setDraftName(e.target.value)}
onBlur={saveName}
onKeyDown={(e) => {
if (e.key === "Enter") {
e.preventDefault();
e.currentTarget.blur();
}
}}
autoComplete="off"
placeholder={sessionSub || "Unsloth"}
className="h-9 w-full rounded-full text-sm"
/>
</div>
<div
data-settings-label={t("settings.profile.nickname")}
className="flex min-w-0 flex-col gap-1.5"
>
<Label
htmlFor="profile-nickname"
className="text-xs font-medium text-muted-foreground"
>
{t("settings.profile.nickname")}
</Label>
<Input
id="profile-nickname"
type="text"
value={draftNickname}
maxLength={PROFILE_TEXT_MAX_LENGTH}
onChange={(e) => setDraftNickname(e.target.value)}
onBlur={saveNickname}
onKeyDown={(e) => {
if (e.key === "Enter") {
e.preventDefault();
e.currentTarget.blur();
}
}}
autoComplete="off"
placeholder={t("settings.profile.nicknamePlaceholder")}
className="h-9 w-full rounded-full text-sm"
/>
</div>
</div>
</div>
{imageError ? (
<p className="w-full text-xs text-destructive" role="alert">
<p className="pt-2 text-xs text-destructive" role="alert">
{imageError}
</p>
) : null}

Some files were not shown because too many files have changed in this diff Show more