unsloth/studio/backend/routes/preview.py
Daniel Han 80d3434d61
Studio: require signed capability tokens for /p preview links (#6666)
* Studio: require signed capability tokens for /p preview links

The public /p preview routes added in #6486 run model load and chat
generation as the admin user with no authentication. The only gate is the
preview ref, a deterministic outputs-root path (run or run/checkpoint) that
is guessable rather than secret. On a network-reachable Studio (--secure
tunnel or -H 0.0.0.0), an unauthenticated caller who guesses a ref can
consume GPU and probe a private fine-tuned checkpoint.

Make the share link an unguessable, revocable capability:

- Sign the canonical ref with a dedicated server-side secret (HMAC-SHA256,
  stored in app_secrets, independent of the JWT/login secret).
- Require a valid token on every /p chat, models, and page request before
  resolving a checkpoint or loading a model; missing or invalid tokens get a
  generic 404 so the surface never confirms a ref exists.
- Accept the token via ?k= (browser link and preview page) or
  Authorization: Bearer (OpenAI-compatible clients).
- Rotate the secret to revoke every outstanding link
  (POST /api/settings/preview-links/rotate).
- Clamp preview generation (max_tokens/max_completion_tokens <= 1024, n = 1)
  and set Referrer-Policy: no-referrer on the page so the token is not
  leaked via Referer.

Training history hands the authenticated owner the signed token, and the
copy-link button builds /p/{ref}?k={sig}.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: honor a lower caller token limit in the preview clamp

Codex review: when only the legacy max_tokens was sent, the clamp left
max_completion_tokens at the 1024 default, and _effective_max_tokens prefers
max_completion_tokens, so a request like max_tokens=16 could still generate up
to 1024 tokens. Derive one effective limit (max_completion_tokens wins, else the
legacy max_tokens) and pin both fields to it so a caller's lower limit is kept.

* Studio: add preview kill switch, rate limit, and revoke-links UI

Follow-ups to the /p preview capability work:

- Public-sharing kill switch: a persisted setting (default on) gates the public
  /p surface. When off, every preview request 404s even with a valid token, and
  the owner UI stops offering share links. GET/PUT /api/settings/preview-sharing;
  enforced in _verify_or_404.
- Per-IP rate limit on the preview chat route: a coarse in-process sliding-window
  limiter (20 req/min/IP) returns 429 + Retry-After before the GPU lock is taken.
  Client IP honors X-Forwarded-For only when UNSLOTH_STUDIO_TRUST_FORWARDED is
  set, matching the login limiter's trust model.
- Settings UI: a "Preview sharing" section with the public-sharing toggle and a
  "Revoke all preview links" button (confirm dialog) that rotates the secret.

Tests cover the kill switch (404 when off), the 429 path, the sliding window,
client-IP trust behavior, and the setting default.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: fix preview-fields sharing arg and refresh sigs after revoke

Codex review:
- P1: get_training_run_detail and update_training_run called _preview_fields
  with only output_dir after it gained a required sharing_on parameter, raising
  a 500 TypeError once get_run succeeded. Pass get_preview_sharing_enabled() at
  both sites; add a detail-endpoint regression test.
- P2: after rotating the preview secret from settings, the history grid still
  held stale preview_sig values, so a freshly copied link would 404. Emit
  emitTrainingRunsChanged() after a successful revoke so the grid refetches
  freshly signed refs.

* Studio: harden preview sharing controls (Codex review)

- Fail closed: a read failure on the preview-sharing kill switch now returns
  False instead of defaulting to enabled, so an unavailable settings DB can't
  reopen the public surface. A missing key still defaults to enabled.
- Per-IP rate limit behind the managed Cloudflare tunnel: client_ip now honors
  CF-Connecting-IP when the socket peer is loopback, so tunneled visitors are
  keyed by their real IP instead of collapsing onto the local cloudflared peer.
- GET /p no longer mints key/share_url when sharing is disabled; it returns
  sharing_enabled=false so clients don't distribute links that 404.
- Settings UI: toggling public sharing emits the training-runs-changed event so
  the history grid shows/hides Copy preview link without a manual refresh.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: harden preview rate limiter and IP keying (Opus review)

From a two-agent review of the PR:

- Rate limiter no longer evicts an active bucket when the table is full: a flood
  of distinct keys could otherwise cycle out a throttled bucket and reset its
  counter. Evict only aged-out buckets; if the table is full of live clients,
  fail closed (deny the new key) instead.
- client_ip keys on the rightmost (proxy-appended) X-Forwarded-For hop when the
  trust env is set; the leftmost is client-spoofable. Documented the
  append/overwrite-proxy assumption.
- _verify_or_404 checks the capability token before the kill-switch DB read, so
  unauthenticated /p spam can't be used as an unbounded settings-DB sink and the
  response is identical regardless of the sharing on/off state.

Tests: nested run/checkpoint happy path + wrong-ref rejection, the eviction
fail-closed behavior, and route-level coverage for the rotate / preview-sharing
settings endpoints.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-06-25 21:40:48 -07:00

296 lines
11 KiB
Python

# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""Per-checkpoint preview endpoints: /p/{run}[/{checkpoint}]/v1/..."""
from __future__ import annotations
import asyncio
import html
from pathlib import Path
from urllib.parse import quote
from fastapi import APIRouter, Depends, HTTPException, Request
from fastapi.responses import FileResponse, HTMLResponse, StreamingResponse
from loggers import get_logger
from auth.authentication import get_current_subject
from auth.storage import DEFAULT_ADMIN_USERNAME
from models.inference import ChatCompletionRequest, LoadRequest
from routes.inference import load_model, openai_chat_completions
from state.tool_policy import tools_force_disabled
from utils.client_ip import client_ip
from utils.models.checkpoints import list_preview_targets, resolve_preview_checkpoint
from utils.preview_rate_limit import check_rate_limit
from utils.preview_sharing_settings import get_preview_sharing_enabled
from utils.preview_token import sign_preview_ref, verify_preview_ref
logger = get_logger(__name__)
router = APIRouter()
# A shared preview link is a public bearer capability; cap per-request generation
# so a single call can't tie up the (serialized) preview GPU indefinitely.
_PREVIEW_MAX_OUTPUT_TOKENS = 1024
# Capability-gated (signed ref required); resolve_preview_checkpoint pins `run`
# under outputs_root. One model loads at a time, so serialize load+generate.
_preview_lock = asyncio.Lock()
def _extract_token(request: Request) -> str | None:
"""Capability token from the ``?k=`` query (browser link + preview page) or an
``Authorization: Bearer`` header (OpenAI-compatible clients using it as api_key)."""
token = request.query_params.get("k")
if token:
return token
header = request.headers.get("authorization", "")
if header[:7].lower() == "bearer ":
return header[7:].strip() or None
return None
def _verify_or_404(run: str, checkpoint: str | None, request: Request) -> None:
"""Require a valid preview capability BEFORE any checkpoint resolve / model load.
Missing or invalid tokens get a generic 404 -- identical to a non-existent ref --
so the public surface never confirms whether a run/checkpoint exists. When an
admin has switched public sharing off, every public request 404s regardless of
token.
Verify the (cheap, no-I/O) capability first: an unauthenticated caller with a
bad/missing token is rejected without the kill-switch DB read, so spamming
``/p/...`` can't be used as an unbounded settings-DB sink, and the response is
identical whether or not sharing is enabled (no on/off oracle).
"""
ref = run if not checkpoint else f"{run}/{checkpoint}"
if not verify_preview_ref(ref, _extract_token(request)):
raise HTTPException(status_code = 404, detail = "Not found")
if not get_preview_sharing_enabled():
raise HTTPException(status_code = 404, detail = "Not found")
def _enforce_rate_limit(request: Request) -> None:
"""Throttle the GPU-backed preview chat per client IP (429 on exceed)."""
retry_after = check_rate_limit(client_ip(request))
if retry_after:
raise HTTPException(
status_code = 429,
detail = "Too many preview requests. Please slow down.",
headers = {"Retry-After": str(retry_after)},
)
def _resolve_or_4xx(run: str, checkpoint: str | None):
try:
return resolve_preview_checkpoint(run, checkpoint)
except ValueError as exc:
# Detail can carry the absolute install path on a symlink escape; log it,
# return a generic message on this public route.
logger.warning("preview path rejected: %s", exc)
raise HTTPException(status_code = 400, detail = "Invalid run or checkpoint")
except FileNotFoundError as exc:
raise HTTPException(status_code = 404, detail = str(exc))
def _sanitize_preview_payload(
payload: ChatCompletionRequest, is_lora: bool
) -> ChatCompletionRequest:
# Public surface: strip tools/MCP + provider routing (no host code / open proxy).
# Normalize use_adapter (never trust the caller): pin True for LoRA, None for
# merged. _apply_adapter_state mutates the shared model without restoring, so an
# unpinned `false` would persist to later visitors who omit the field.
#
# Cap generation cost on this public, GPU-backed surface. Derive one effective
# limit (mirroring _effective_max_tokens: max_completion_tokens wins, else the
# legacy max_tokens) and pin BOTH fields to it, so a caller's lower limit is
# honored and neither field can exceed the ceiling.
requested = (
payload.max_completion_tokens
if payload.max_completion_tokens is not None
else payload.max_tokens
)
capped_max_tokens = (
min(requested, _PREVIEW_MAX_OUTPUT_TOKENS)
if requested is not None
else _PREVIEW_MAX_OUTPUT_TOKENS
)
return payload.model_copy(
update = {
"tools": None,
"enable_tools": False,
"enabled_tools": None,
"mcp_enabled": False,
"bypass_permissions": False,
"confirm_tool_calls": False,
"session_id": None,
"rag_scope": None,
"openai_code_exec_container_id": None,
"anthropic_code_exec_container_id": None,
"provider_id": None,
"provider_type": None,
"external_model": None,
"encrypted_api_key": None,
"provider_base_url": None,
"use_adapter": True if is_lora else None,
"max_tokens": capped_max_tokens,
"max_completion_tokens": capped_max_tokens,
"n": 1,
}
)
async def _unlock_after(body_iterator):
# Hold the lock until the stream drains so another checkpoint can't swap mid-stream.
try:
async for chunk in body_iterator:
yield chunk
finally:
_preview_lock.release()
async def _serve_chat(
run: str, checkpoint: str | None, payload: ChatCompletionRequest, request: Request
):
path = _resolve_or_4xx(run, checkpoint)
is_lora = (path / "adapter_config.json").exists()
payload = _sanitize_preview_payload(payload, is_lora)
await _preview_lock.acquire()
keep_locked = False
try:
await load_model(LoadRequest(model_path = str(path)), request, DEFAULT_ADMIN_USERNAME)
# Beats a process-wide `--enable-tools` (enable_tools=False alone wouldn't).
with tools_force_disabled():
response = await openai_chat_completions(payload, request, DEFAULT_ADMIN_USERNAME)
if isinstance(response, StreamingResponse):
response.body_iterator = _unlock_after(response.body_iterator)
keep_locked = True
return response
finally:
if not keep_locked:
_preview_lock.release()
@router.get("")
async def list_previews(request: Request, current_subject: str = Depends(get_current_subject)):
base = str(request.base_url)
sharing_on = get_preview_sharing_enabled()
previews = []
for target in list_preview_targets():
ref = quote(target["ref"], safe = "/")
# Mint the capability for the authenticated owner: ``key`` for OpenAI
# clients (Bearer / api_key), ``share_url`` for the browser link. When
# public sharing is off, every public /p request 404s, so don't hand out
# dead credentials -- omit the capability and signal the disabled state.
token = sign_preview_ref(target["ref"]) if sharing_on else None
previews.append(
{
**target,
"url": f"{base}p/{ref}/v1",
"key": token,
"share_url": f"{base}p/{ref}?k={token}" if token else None,
}
)
return {"object": "list", "data": previews, "sharing_enabled": sharing_on}
@router.post("/{run}/v1/chat/completions")
async def preview_chat_latest(run: str, payload: ChatCompletionRequest, request: Request):
_verify_or_404(run, None, request)
_enforce_rate_limit(request)
return await _serve_chat(run, None, payload, request)
@router.post("/{run}/{checkpoint}/v1/chat/completions")
async def preview_chat_checkpoint(
run: str, checkpoint: str, payload: ChatCompletionRequest, request: Request
):
_verify_or_404(run, checkpoint, request)
_enforce_rate_limit(request)
return await _serve_chat(run, checkpoint, payload, request)
def _models_response(run: str, checkpoint: str | None):
path = _resolve_or_4xx(run, checkpoint)
model_id = run if not checkpoint else f"{run}/{checkpoint}"
return {
"object": "list",
"data": [
{
"id": model_id,
"object": "model",
"created": int(path.stat().st_mtime),
"owned_by": "unsloth-studio",
}
],
}
# The models/page GET routes only stat the checkpoint dir (no GPU), so they are
# token-gated but not rate-limited; only the GPU-backed chat path is throttled.
@router.get("/{run}/v1/models")
async def preview_models_latest(run: str, request: Request):
_verify_or_404(run, None, request)
return _models_response(run, None)
@router.get("/{run}/{checkpoint}/v1/models")
async def preview_models_checkpoint(run: str, checkpoint: str, request: Request):
_verify_or_404(run, checkpoint, request)
return _models_response(run, checkpoint)
# Serve logo/fonts here too: the SPA static mount is absent in --api-only (Tauri).
_FRONTEND_DIST = (Path(__file__).resolve().parents[2] / "frontend" / "dist").resolve()
_PREVIEW_ASSET_MEDIA_TYPES = {
".png": "image/png",
".woff": "font/woff",
".woff2": "font/woff2",
}
@router.get("/_assets/{asset_path:path}")
async def preview_asset(asset_path: str):
target = (_FRONTEND_DIST / asset_path).resolve()
media_type = _PREVIEW_ASSET_MEDIA_TYPES.get(target.suffix.lower())
if media_type is None or not target.is_relative_to(_FRONTEND_DIST) or not target.is_file():
raise HTTPException(status_code = 404, detail = "Not found")
return FileResponse(target, media_type = media_type)
# Self-contained public page; only the title is interpolated.
_PREVIEW_PAGE_HTML = (
Path(__file__).resolve().parent.parent / "assets" / "preview_page.html"
).read_text(encoding = "utf-8")
_PREVIEW_PAGE_CSP = (
"default-src 'self'; script-src 'unsafe-inline'; style-src 'unsafe-inline'; "
"img-src 'self'; font-src 'self'; connect-src 'self'; base-uri 'none'"
)
def _preview_page(run: str, checkpoint: str | None) -> HTMLResponse:
_resolve_or_4xx(run, checkpoint)
title = run if not checkpoint else f"{run}/{checkpoint}"
page = _PREVIEW_PAGE_HTML.replace("__TITLE__", html.escape(title))
# no-referrer: the capability token rides in the query string, so keep it out
# of the Referer header on any outbound navigation.
return HTMLResponse(
page,
headers = {
"Content-Security-Policy": _PREVIEW_PAGE_CSP,
"Referrer-Policy": "no-referrer",
},
)
@router.get("/{run}", response_class = HTMLResponse)
async def preview_page_latest(run: str, request: Request):
_verify_or_404(run, None, request)
return _preview_page(run, None)
@router.get("/{run}/{checkpoint}", response_class = HTMLResponse)
async def preview_page_checkpoint(run: str, checkpoint: str, request: Request):
_verify_or_404(run, checkpoint, request)
return _preview_page(run, checkpoint)