unsloth/studio/backend/utils/datasets/__init__.py
Daniel Han bb4eb88fdc
Studio: tools, thinking blocks, code execution and web search for safetensors (#5520)
Adds tools, thinking blocks, code execution, and web search support to the safetensors / transformers and MLX inference backends in Studio, bringing them to parity with the GGUF path.

What ships
- safetensors / transformers agentic tool loop with cumulative-text state machine, tool-call XML parser, and template kwarg forwarding (tools / enable_thinking / reasoning_effort / preserve_thinking).
- MLX backend: same kwargs accepted on Apple Silicon; chat_template_info shipped through worker IPC; pills enable for Qwen / Qwen3 / Qwen3.5 / Gemma reasoning.
- Capability classifier (_detect_safetensors_features) gates supports_tools on actual parser-compatible emission markers (<tool_call> / <function=) so Llama-3 / Mistral / Gemma 4 do not advertise toggles the parser cannot honour.
- gpt-oss override stays: reasoning on, tools off (Harmony channel, not <tool_call> XML).
- CWE-209 hygiene: safetensors SSE error path emits a constant message and logs the trace server-side.

Validation
- 256 unit tests green (43 tool-loop, 11 capability advertise, 7 MLX backend, 5 main-added, 190 adjacent inference / anthropic / openai regression).
- Cross-OS staging CI green on ubuntu-latest / macos-14 / windows-latest plus a dedicated MLX cartesian probe against real unsloth/Qwen3.5-0.8B on macos-14 (CI 26098107440).
- Capability parity verified across Qwen3 / Qwen3.5 / Llama-3 / Mistral / Gemma / DeepSeek-R1 / gpt-oss (incl. BF16).
- Manual confirmation from Imagineer99 on Qwen3.5-2B: think + search + code exec working.

Closes the safetensors / MLX gap with the GGUF backend.
2026-05-19 06:30:17 -07:00

107 lines
2.8 KiB
Python

# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""
Dataset utilities package.
This package provides utilities for dataset format detection, conversion,
and processing for LLM and VLM fine-tuning workflows.
Modules:
- format_detection: Detect dataset formats (Alpaca, ShareGPT, ChatML)
- format_conversion: Convert between dataset formats
- chat_templates: Apply chat templates to datasets
- vlm_processing: Vision-Language Model processing utilities
- data_collators: Custom data collators for training
- model_mappings: Model-to-template mapping constants
"""
# Format detection
from .format_detection import (
detect_dataset_format,
detect_custom_format_heuristic,
detect_multimodal_dataset,
detect_vlm_dataset_structure,
)
# Format conversion
from .format_conversion import (
standardize_chat_format,
convert_chatml_to_alpaca,
convert_alpaca_to_chatml,
convert_to_vlm_format,
convert_llava_to_vlm_format,
convert_sharegpt_with_images_to_vlm_format,
)
# Chat templates
from .chat_templates import (
apply_chat_template_to_dataset,
get_dataset_info_summary,
get_tokenizer_chat_template,
DEFAULT_ALPACA_TEMPLATE,
)
# VLM processing
from .vlm_processing import (
generate_smart_vlm_instruction,
)
# Data collators
from .data_collators import (
DataCollatorSpeechSeq2SeqWithPadding,
DeepSeekOCRDataCollator,
VLMDataCollator,
)
# Model mappings (constants)
from .model_mappings import (
TEMPLATE_TO_MODEL_MAPPER,
MODEL_TO_TEMPLATE_MAPPER,
TEMPLATE_TO_RESPONSES_MAPPER,
is_gpt_oss_model_name,
)
# Legacy imports from the original dataset_utils.py for backward compatibility
# These functions have not yet been refactored into separate modules
from .dataset_utils import (
check_dataset_format,
format_and_template_dataset,
format_dataset,
)
# Public API
__all__ = [
# Detection
"detect_dataset_format",
"detect_custom_format_heuristic",
"detect_multimodal_dataset",
"detect_vlm_dataset_structure",
# Conversion
"standardize_chat_format",
"convert_chatml_to_alpaca",
"convert_alpaca_to_chatml",
"convert_to_vlm_format",
"convert_llava_to_vlm_format",
"convert_sharegpt_with_images_to_vlm_format",
# Templates
"apply_chat_template_to_dataset",
"get_dataset_info_summary",
"get_tokenizer_chat_template",
"DEFAULT_ALPACA_TEMPLATE",
# VLM
"generate_smart_vlm_instruction",
# Collators
"DataCollatorSpeechSeq2SeqWithPadding",
"DeepSeekOCRDataCollator",
"VLMDataCollator",
# Mappings
"TEMPLATE_TO_MODEL_MAPPER",
"MODEL_TO_TEMPLATE_MAPPER",
"TEMPLATE_TO_RESPONSES_MAPPER",
"is_gpt_oss_model_name",
# Main entry points
"check_dataset_format",
"format_and_template_dataset",
"format_dataset",
]