* Studio: require signed capability tokens for /p preview links The public /p preview routes added in #6486 run model load and chat generation as the admin user with no authentication. The only gate is the preview ref, a deterministic outputs-root path (run or run/checkpoint) that is guessable rather than secret. On a network-reachable Studio (--secure tunnel or -H 0.0.0.0), an unauthenticated caller who guesses a ref can consume GPU and probe a private fine-tuned checkpoint. Make the share link an unguessable, revocable capability: - Sign the canonical ref with a dedicated server-side secret (HMAC-SHA256, stored in app_secrets, independent of the JWT/login secret). - Require a valid token on every /p chat, models, and page request before resolving a checkpoint or loading a model; missing or invalid tokens get a generic 404 so the surface never confirms a ref exists. - Accept the token via ?k= (browser link and preview page) or Authorization: Bearer (OpenAI-compatible clients). - Rotate the secret to revoke every outstanding link (POST /api/settings/preview-links/rotate). - Clamp preview generation (max_tokens/max_completion_tokens <= 1024, n = 1) and set Referrer-Policy: no-referrer on the page so the token is not leaked via Referer. Training history hands the authenticated owner the signed token, and the copy-link button builds /p/{ref}?k={sig}. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: honor a lower caller token limit in the preview clamp Codex review: when only the legacy max_tokens was sent, the clamp left max_completion_tokens at the 1024 default, and _effective_max_tokens prefers max_completion_tokens, so a request like max_tokens=16 could still generate up to 1024 tokens. Derive one effective limit (max_completion_tokens wins, else the legacy max_tokens) and pin both fields to it so a caller's lower limit is kept. * Studio: add preview kill switch, rate limit, and revoke-links UI Follow-ups to the /p preview capability work: - Public-sharing kill switch: a persisted setting (default on) gates the public /p surface. When off, every preview request 404s even with a valid token, and the owner UI stops offering share links. GET/PUT /api/settings/preview-sharing; enforced in _verify_or_404. - Per-IP rate limit on the preview chat route: a coarse in-process sliding-window limiter (20 req/min/IP) returns 429 + Retry-After before the GPU lock is taken. Client IP honors X-Forwarded-For only when UNSLOTH_STUDIO_TRUST_FORWARDED is set, matching the login limiter's trust model. - Settings UI: a "Preview sharing" section with the public-sharing toggle and a "Revoke all preview links" button (confirm dialog) that rotates the secret. Tests cover the kill switch (404 when off), the 429 path, the sliding window, client-IP trust behavior, and the setting default. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: fix preview-fields sharing arg and refresh sigs after revoke Codex review: - P1: get_training_run_detail and update_training_run called _preview_fields with only output_dir after it gained a required sharing_on parameter, raising a 500 TypeError once get_run succeeded. Pass get_preview_sharing_enabled() at both sites; add a detail-endpoint regression test. - P2: after rotating the preview secret from settings, the history grid still held stale preview_sig values, so a freshly copied link would 404. Emit emitTrainingRunsChanged() after a successful revoke so the grid refetches freshly signed refs. * Studio: harden preview sharing controls (Codex review) - Fail closed: a read failure on the preview-sharing kill switch now returns False instead of defaulting to enabled, so an unavailable settings DB can't reopen the public surface. A missing key still defaults to enabled. - Per-IP rate limit behind the managed Cloudflare tunnel: client_ip now honors CF-Connecting-IP when the socket peer is loopback, so tunneled visitors are keyed by their real IP instead of collapsing onto the local cloudflared peer. - GET /p no longer mints key/share_url when sharing is disabled; it returns sharing_enabled=false so clients don't distribute links that 404. - Settings UI: toggling public sharing emits the training-runs-changed event so the history grid shows/hides Copy preview link without a manual refresh. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: harden preview rate limiter and IP keying (Opus review) From a two-agent review of the PR: - Rate limiter no longer evicts an active bucket when the table is full: a flood of distinct keys could otherwise cycle out a throttled bucket and reset its counter. Evict only aged-out buckets; if the table is full of live clients, fail closed (deny the new key) instead. - client_ip keys on the rightmost (proxy-appended) X-Forwarded-For hop when the trust env is set; the leftmost is client-spoofable. Documented the append/overwrite-proxy assumption. - _verify_or_404 checks the capability token before the kill-switch DB read, so unauthenticated /p spam can't be used as an unbounded settings-DB sink and the response is identical regardless of the sharing on/off state. Tests: nested run/checkpoint happy path + wrong-ref rejection, the eviction fail-closed behavior, and route-level coverage for the rotate / preview-sharing settings endpoints. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
296 lines
11 KiB
Python
296 lines
11 KiB
Python
# SPDX-License-Identifier: AGPL-3.0-only
|
|
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
|
|
|
"""Per-checkpoint preview endpoints: /p/{run}[/{checkpoint}]/v1/..."""
|
|
|
|
from __future__ import annotations
|
|
|
|
import asyncio
|
|
import html
|
|
from pathlib import Path
|
|
from urllib.parse import quote
|
|
|
|
from fastapi import APIRouter, Depends, HTTPException, Request
|
|
from fastapi.responses import FileResponse, HTMLResponse, StreamingResponse
|
|
from loggers import get_logger
|
|
|
|
from auth.authentication import get_current_subject
|
|
from auth.storage import DEFAULT_ADMIN_USERNAME
|
|
from models.inference import ChatCompletionRequest, LoadRequest
|
|
from routes.inference import load_model, openai_chat_completions
|
|
from state.tool_policy import tools_force_disabled
|
|
from utils.client_ip import client_ip
|
|
from utils.models.checkpoints import list_preview_targets, resolve_preview_checkpoint
|
|
from utils.preview_rate_limit import check_rate_limit
|
|
from utils.preview_sharing_settings import get_preview_sharing_enabled
|
|
from utils.preview_token import sign_preview_ref, verify_preview_ref
|
|
|
|
logger = get_logger(__name__)
|
|
|
|
router = APIRouter()
|
|
|
|
# A shared preview link is a public bearer capability; cap per-request generation
|
|
# so a single call can't tie up the (serialized) preview GPU indefinitely.
|
|
_PREVIEW_MAX_OUTPUT_TOKENS = 1024
|
|
|
|
# Capability-gated (signed ref required); resolve_preview_checkpoint pins `run`
|
|
# under outputs_root. One model loads at a time, so serialize load+generate.
|
|
_preview_lock = asyncio.Lock()
|
|
|
|
|
|
def _extract_token(request: Request) -> str | None:
|
|
"""Capability token from the ``?k=`` query (browser link + preview page) or an
|
|
``Authorization: Bearer`` header (OpenAI-compatible clients using it as api_key)."""
|
|
token = request.query_params.get("k")
|
|
if token:
|
|
return token
|
|
header = request.headers.get("authorization", "")
|
|
if header[:7].lower() == "bearer ":
|
|
return header[7:].strip() or None
|
|
return None
|
|
|
|
|
|
def _verify_or_404(run: str, checkpoint: str | None, request: Request) -> None:
|
|
"""Require a valid preview capability BEFORE any checkpoint resolve / model load.
|
|
|
|
Missing or invalid tokens get a generic 404 -- identical to a non-existent ref --
|
|
so the public surface never confirms whether a run/checkpoint exists. When an
|
|
admin has switched public sharing off, every public request 404s regardless of
|
|
token.
|
|
|
|
Verify the (cheap, no-I/O) capability first: an unauthenticated caller with a
|
|
bad/missing token is rejected without the kill-switch DB read, so spamming
|
|
``/p/...`` can't be used as an unbounded settings-DB sink, and the response is
|
|
identical whether or not sharing is enabled (no on/off oracle).
|
|
"""
|
|
ref = run if not checkpoint else f"{run}/{checkpoint}"
|
|
if not verify_preview_ref(ref, _extract_token(request)):
|
|
raise HTTPException(status_code = 404, detail = "Not found")
|
|
if not get_preview_sharing_enabled():
|
|
raise HTTPException(status_code = 404, detail = "Not found")
|
|
|
|
|
|
def _enforce_rate_limit(request: Request) -> None:
|
|
"""Throttle the GPU-backed preview chat per client IP (429 on exceed)."""
|
|
retry_after = check_rate_limit(client_ip(request))
|
|
if retry_after:
|
|
raise HTTPException(
|
|
status_code = 429,
|
|
detail = "Too many preview requests. Please slow down.",
|
|
headers = {"Retry-After": str(retry_after)},
|
|
)
|
|
|
|
|
|
def _resolve_or_4xx(run: str, checkpoint: str | None):
|
|
try:
|
|
return resolve_preview_checkpoint(run, checkpoint)
|
|
except ValueError as exc:
|
|
# Detail can carry the absolute install path on a symlink escape; log it,
|
|
# return a generic message on this public route.
|
|
logger.warning("preview path rejected: %s", exc)
|
|
raise HTTPException(status_code = 400, detail = "Invalid run or checkpoint")
|
|
except FileNotFoundError as exc:
|
|
raise HTTPException(status_code = 404, detail = str(exc))
|
|
|
|
|
|
def _sanitize_preview_payload(
|
|
payload: ChatCompletionRequest, is_lora: bool
|
|
) -> ChatCompletionRequest:
|
|
# Public surface: strip tools/MCP + provider routing (no host code / open proxy).
|
|
# Normalize use_adapter (never trust the caller): pin True for LoRA, None for
|
|
# merged. _apply_adapter_state mutates the shared model without restoring, so an
|
|
# unpinned `false` would persist to later visitors who omit the field.
|
|
#
|
|
# Cap generation cost on this public, GPU-backed surface. Derive one effective
|
|
# limit (mirroring _effective_max_tokens: max_completion_tokens wins, else the
|
|
# legacy max_tokens) and pin BOTH fields to it, so a caller's lower limit is
|
|
# honored and neither field can exceed the ceiling.
|
|
requested = (
|
|
payload.max_completion_tokens
|
|
if payload.max_completion_tokens is not None
|
|
else payload.max_tokens
|
|
)
|
|
capped_max_tokens = (
|
|
min(requested, _PREVIEW_MAX_OUTPUT_TOKENS)
|
|
if requested is not None
|
|
else _PREVIEW_MAX_OUTPUT_TOKENS
|
|
)
|
|
return payload.model_copy(
|
|
update = {
|
|
"tools": None,
|
|
"enable_tools": False,
|
|
"enabled_tools": None,
|
|
"mcp_enabled": False,
|
|
"bypass_permissions": False,
|
|
"confirm_tool_calls": False,
|
|
"session_id": None,
|
|
"rag_scope": None,
|
|
"openai_code_exec_container_id": None,
|
|
"anthropic_code_exec_container_id": None,
|
|
"provider_id": None,
|
|
"provider_type": None,
|
|
"external_model": None,
|
|
"encrypted_api_key": None,
|
|
"provider_base_url": None,
|
|
"use_adapter": True if is_lora else None,
|
|
"max_tokens": capped_max_tokens,
|
|
"max_completion_tokens": capped_max_tokens,
|
|
"n": 1,
|
|
}
|
|
)
|
|
|
|
|
|
async def _unlock_after(body_iterator):
|
|
# Hold the lock until the stream drains so another checkpoint can't swap mid-stream.
|
|
try:
|
|
async for chunk in body_iterator:
|
|
yield chunk
|
|
finally:
|
|
_preview_lock.release()
|
|
|
|
|
|
async def _serve_chat(
|
|
run: str, checkpoint: str | None, payload: ChatCompletionRequest, request: Request
|
|
):
|
|
path = _resolve_or_4xx(run, checkpoint)
|
|
is_lora = (path / "adapter_config.json").exists()
|
|
payload = _sanitize_preview_payload(payload, is_lora)
|
|
await _preview_lock.acquire()
|
|
keep_locked = False
|
|
try:
|
|
await load_model(LoadRequest(model_path = str(path)), request, DEFAULT_ADMIN_USERNAME)
|
|
# Beats a process-wide `--enable-tools` (enable_tools=False alone wouldn't).
|
|
with tools_force_disabled():
|
|
response = await openai_chat_completions(payload, request, DEFAULT_ADMIN_USERNAME)
|
|
if isinstance(response, StreamingResponse):
|
|
response.body_iterator = _unlock_after(response.body_iterator)
|
|
keep_locked = True
|
|
return response
|
|
finally:
|
|
if not keep_locked:
|
|
_preview_lock.release()
|
|
|
|
|
|
@router.get("")
|
|
async def list_previews(request: Request, current_subject: str = Depends(get_current_subject)):
|
|
base = str(request.base_url)
|
|
sharing_on = get_preview_sharing_enabled()
|
|
previews = []
|
|
for target in list_preview_targets():
|
|
ref = quote(target["ref"], safe = "/")
|
|
# Mint the capability for the authenticated owner: ``key`` for OpenAI
|
|
# clients (Bearer / api_key), ``share_url`` for the browser link. When
|
|
# public sharing is off, every public /p request 404s, so don't hand out
|
|
# dead credentials -- omit the capability and signal the disabled state.
|
|
token = sign_preview_ref(target["ref"]) if sharing_on else None
|
|
previews.append(
|
|
{
|
|
**target,
|
|
"url": f"{base}p/{ref}/v1",
|
|
"key": token,
|
|
"share_url": f"{base}p/{ref}?k={token}" if token else None,
|
|
}
|
|
)
|
|
return {"object": "list", "data": previews, "sharing_enabled": sharing_on}
|
|
|
|
|
|
@router.post("/{run}/v1/chat/completions")
|
|
async def preview_chat_latest(run: str, payload: ChatCompletionRequest, request: Request):
|
|
_verify_or_404(run, None, request)
|
|
_enforce_rate_limit(request)
|
|
return await _serve_chat(run, None, payload, request)
|
|
|
|
|
|
@router.post("/{run}/{checkpoint}/v1/chat/completions")
|
|
async def preview_chat_checkpoint(
|
|
run: str, checkpoint: str, payload: ChatCompletionRequest, request: Request
|
|
):
|
|
_verify_or_404(run, checkpoint, request)
|
|
_enforce_rate_limit(request)
|
|
return await _serve_chat(run, checkpoint, payload, request)
|
|
|
|
|
|
def _models_response(run: str, checkpoint: str | None):
|
|
path = _resolve_or_4xx(run, checkpoint)
|
|
model_id = run if not checkpoint else f"{run}/{checkpoint}"
|
|
return {
|
|
"object": "list",
|
|
"data": [
|
|
{
|
|
"id": model_id,
|
|
"object": "model",
|
|
"created": int(path.stat().st_mtime),
|
|
"owned_by": "unsloth-studio",
|
|
}
|
|
],
|
|
}
|
|
|
|
|
|
# The models/page GET routes only stat the checkpoint dir (no GPU), so they are
|
|
# token-gated but not rate-limited; only the GPU-backed chat path is throttled.
|
|
@router.get("/{run}/v1/models")
|
|
async def preview_models_latest(run: str, request: Request):
|
|
_verify_or_404(run, None, request)
|
|
return _models_response(run, None)
|
|
|
|
|
|
@router.get("/{run}/{checkpoint}/v1/models")
|
|
async def preview_models_checkpoint(run: str, checkpoint: str, request: Request):
|
|
_verify_or_404(run, checkpoint, request)
|
|
return _models_response(run, checkpoint)
|
|
|
|
|
|
# Serve logo/fonts here too: the SPA static mount is absent in --api-only (Tauri).
|
|
_FRONTEND_DIST = (Path(__file__).resolve().parents[2] / "frontend" / "dist").resolve()
|
|
_PREVIEW_ASSET_MEDIA_TYPES = {
|
|
".png": "image/png",
|
|
".woff": "font/woff",
|
|
".woff2": "font/woff2",
|
|
}
|
|
|
|
|
|
@router.get("/_assets/{asset_path:path}")
|
|
async def preview_asset(asset_path: str):
|
|
target = (_FRONTEND_DIST / asset_path).resolve()
|
|
media_type = _PREVIEW_ASSET_MEDIA_TYPES.get(target.suffix.lower())
|
|
if media_type is None or not target.is_relative_to(_FRONTEND_DIST) or not target.is_file():
|
|
raise HTTPException(status_code = 404, detail = "Not found")
|
|
return FileResponse(target, media_type = media_type)
|
|
|
|
|
|
# Self-contained public page; only the title is interpolated.
|
|
_PREVIEW_PAGE_HTML = (
|
|
Path(__file__).resolve().parent.parent / "assets" / "preview_page.html"
|
|
).read_text(encoding = "utf-8")
|
|
|
|
_PREVIEW_PAGE_CSP = (
|
|
"default-src 'self'; script-src 'unsafe-inline'; style-src 'unsafe-inline'; "
|
|
"img-src 'self'; font-src 'self'; connect-src 'self'; base-uri 'none'"
|
|
)
|
|
|
|
|
|
def _preview_page(run: str, checkpoint: str | None) -> HTMLResponse:
|
|
_resolve_or_4xx(run, checkpoint)
|
|
title = run if not checkpoint else f"{run}/{checkpoint}"
|
|
page = _PREVIEW_PAGE_HTML.replace("__TITLE__", html.escape(title))
|
|
# no-referrer: the capability token rides in the query string, so keep it out
|
|
# of the Referer header on any outbound navigation.
|
|
return HTMLResponse(
|
|
page,
|
|
headers = {
|
|
"Content-Security-Policy": _PREVIEW_PAGE_CSP,
|
|
"Referrer-Policy": "no-referrer",
|
|
},
|
|
)
|
|
|
|
|
|
@router.get("/{run}", response_class = HTMLResponse)
|
|
async def preview_page_latest(run: str, request: Request):
|
|
_verify_or_404(run, None, request)
|
|
return _preview_page(run, None)
|
|
|
|
|
|
@router.get("/{run}/{checkpoint}", response_class = HTMLResponse)
|
|
async def preview_page_checkpoint(run: str, checkpoint: str, request: Request):
|
|
_verify_or_404(run, checkpoint, request)
|
|
return _preview_page(run, checkpoint)
|