unsloth/studio/backend/models/__init__.py
Daniel Han f08aef1804 Studio (#4237)
* Rebuild Studio branch on top of main

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix security and code quality issues for Studio PR #4237

- Validate models_dir query param against allowed directory roots
  to prevent path traversal in /api/models/local endpoint
- Replace string startswith() with Path.is_relative_to() for
  frontend path traversal check in serve_frontend
- Sanitize SSE error messages to not leak exception details to
  clients (4 locations in inference.py)
- Bind port-discovery socket to 127.0.0.1 instead of all interfaces
  in llama_cpp backend
- Import datasets_root and resolve_output_dir in embedding training
  function to fix NameError and use managed output directory
- Remove stale .gitignore entries for package-lock.json and test
  directories so tests can be tracked in version control
- Add venv-reexecution logic to ui CLI command matching the studio
  command behavior

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Move models_dir path validation before try/except block

The HTTPException(403) was inside the try/except Exception handler,
so it would be caught and re-raised as a 500. Moving the validation
before the try block ensures the 403 is returned directly and also
makes the control flow clearer for static analysis (path is validated
before any filesystem operations).

* Use os.path.realpath + startswith for models_dir validation

CodeQL py/path-injection does not recognize Path.is_relative_to() as
a sanitizer. Switched to os.path.realpath + str.startswith which is
a recognized sanitizer pattern in CodeQL's taint analysis. The
startswith check uses root_str + os.sep to prevent prefix collisions
(e.g. /app/models_evil matching /app/models).

* Never pass user input to Path constructor in models_dir validation

CodeQL traces taint through Path(resolved) even after a startswith
barrier guard. Fix: the user-supplied models_dir is only used as a
string for comparison against allowed roots. The Path object passed
to _scan_models_dir comes from the trusted allowed_roots list, not
from user input. This fully breaks the taint chain.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-03-12 03:36:19 -07:00

120 lines
2.6 KiB
Python

# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""
Pydantic models for API request/response schemas
"""
from .training import (
TrainingStartRequest,
TrainingJobResponse,
TrainingStatus,
TrainingProgress,
)
from .models import (
CheckpointInfo,
ModelCheckpoints,
CheckpointListResponse,
ModelDetails,
LocalModelInfo,
LocalModelListResponse,
LoRAInfo,
LoRAScanResponse,
ModelListResponse,
)
from .auth import (
AuthSetupRequest,
AuthLoginRequest,
RefreshTokenRequest,
AuthStatusResponse,
)
from .export import (
LoadCheckpointRequest,
ExportStatusResponse,
ExportOperationResponse,
ExportMergedModelRequest,
ExportBaseModelRequest,
ExportGGUFRequest,
ExportLoRAAdapterRequest,
)
from .users import Token
from .datasets import (
CheckFormatRequest,
CheckFormatResponse,
)
from .inference import (
LoadRequest,
UnloadRequest,
GenerateRequest,
LoadResponse,
UnloadResponse,
InferenceStatusResponse,
)
from .responses import (
TrainingStopResponse,
TrainingMetricsResponse,
LoRABaseModelResponse,
VisionCheckResponse,
EmbeddingCheckResponse,
)
from .data_recipe import (
RecipePayload,
PreviewResponse,
ValidateError,
ValidateResponse,
JobCreateResponse,
)
__all__ = [
# Training schemas
"TrainingStartRequest",
"TrainingJobResponse",
"TrainingStatus",
"TrainingProgress",
# Model management schemas
"ModelDetails",
"LocalModelInfo",
"LocalModelListResponse",
"LoRAInfo",
"LoRAScanResponse",
"ModelListResponse",
# Auth schemas
"AuthSetupRequest",
"AuthLoginRequest",
"RefreshTokenRequest",
"AuthStatusResponse",
# Export schemas
"CheckpointInfo",
"ModelCheckpoints",
"CheckpointListResponse",
"LoadCheckpointRequest",
"ExportStatusResponse",
"ExportOperationResponse",
"ExportMergedModelRequest",
"ExportBaseModelRequest",
"ExportGGUFRequest",
"ExportLoRAAdapterRequest",
"Token",
# Dataset schemas
"CheckFormatRequest",
"CheckFormatResponse",
# Inference schemas
"LoadRequest",
"UnloadRequest",
"GenerateRequest",
"LoadResponse",
"UnloadResponse",
"InferenceStatusResponse",
# Response schemas
"TrainingStopResponse",
"TrainingMetricsResponse",
"LoRABaseModelResponse",
"VisionCheckResponse",
"EmbeddingCheckResponse",
# Data recipe
"RecipePayload",
"PreviewResponse",
"ValidateError",
"ValidateResponse",
"JobCreateResponse",
]