unsloth/studio/backend/loggers/config.py
Daniel Han f08aef1804 Studio (#4237)
* Rebuild Studio branch on top of main

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix security and code quality issues for Studio PR #4237

- Validate models_dir query param against allowed directory roots
  to prevent path traversal in /api/models/local endpoint
- Replace string startswith() with Path.is_relative_to() for
  frontend path traversal check in serve_frontend
- Sanitize SSE error messages to not leak exception details to
  clients (4 locations in inference.py)
- Bind port-discovery socket to 127.0.0.1 instead of all interfaces
  in llama_cpp backend
- Import datasets_root and resolve_output_dir in embedding training
  function to fix NameError and use managed output directory
- Remove stale .gitignore entries for package-lock.json and test
  directories so tests can be tracked in version control
- Add venv-reexecution logic to ui CLI command matching the studio
  command behavior

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Move models_dir path validation before try/except block

The HTTPException(403) was inside the try/except Exception handler,
so it would be caught and re-raised as a 500. Moving the validation
before the try block ensures the 403 is returned directly and also
makes the control flow clearer for static analysis (path is validated
before any filesystem operations).

* Use os.path.realpath + startswith for models_dir validation

CodeQL py/path-injection does not recognize Path.is_relative_to() as
a sanitizer. Switched to os.path.realpath + str.startswith which is
a recognized sanitizer pattern in CodeQL's taint analysis. The
startswith check uses root_str + os.sep to prevent prefix collisions
(e.g. /app/models_evil matching /app/models).

* Never pass user input to Path constructor in models_dir validation

CodeQL traces taint through Path(resolved) even after a startswith
barrier guard. Fix: the user-supplied models_dir is only used as a
string for comparison against allowed roots. The Path object passed
to _scan_models_dir comes from the trusted allowed_roots list, not
from user input. This fully breaks the taint chain.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-03-12 03:36:19 -07:00

76 lines
2.9 KiB
Python

# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""Logging configuration for structured logging with structlog.
This module provides centralized logging configuration with environment-specific
formats and processors. Supports both development and production environments
with consistent structured logging.
Key Features:
- Environment-specific formatting (JSON for production, console for development)
- Timestamp standardization (ISO format)
- Context variable integration
- Log level filtering
- Logger caching for performance
"""
import logging
import os
import sys
from typing import Optional
import structlog
class LogConfig:
"""Structured logging configuration for the application.
Provides static method to configure structlog with environment-specific
formatting and processors for consistent structured logging.
"""
@staticmethod
def setup_logging(
service_name: str = "unsloth-studio-backend", env: Optional[str] = None
) -> structlog.BoundLogger:
"""Configure structured logging for the application.
Args:
service_name: Name of the service for logging identification
env: Environment (development/production), affects logging format
"""
# Determine log level from environment
log_level_name = os.getenv("LOG_LEVEL", "INFO").upper()
# Fallback to INFO if an invalid level is provided
log_level = getattr(logging, log_level_name, logging.INFO)
structlog.configure(
processors = [
# Reorder processors to control field order
structlog.processors.TimeStamper(fmt = "iso"), # timestamp first
structlog.processors.add_log_level, # level second
structlog.contextvars.merge_contextvars,
# Custom processor to flatten the extra field
lambda logger, method_name, event_dict: {
"timestamp": event_dict.get("timestamp"),
"level": event_dict.get("level"),
"event": event_dict.get("event"),
**(event_dict.get("extra", {})), # Flatten extra into main dict
**{
k: v
for k, v in event_dict.items()
if k not in ["timestamp", "level", "event", "extra"]
},
},
(
structlog.processors.JSONRenderer(sort_keys = False) # Preserve order
if env == "production"
else structlog.dev.ConsoleRenderer()
),
],
wrapper_class = structlog.make_filtering_bound_logger(log_level),
logger_factory = structlog.PrintLoggerFactory(file = sys.stdout),
cache_logger_on_first_use = True,
)
return structlog.get_logger(service_name)