unsloth/studio/backend/core/inference
Daniel Han 734cec9e7a
Studio STT: only load safetensors weights for custom dictation models (RCE fix) (#7364)
* Studio STT: only load safetensors weights for custom dictation models

The STT sidecar accepts arbitrary Hugging Face owner/model repos for
custom dictation models and, when safetensors were absent, downloaded
and loaded pytorch_model.bin through WhisperForConditionalGeneration
.from_pretrained. PyTorch checkpoints are pickles that execute code
during deserialization, and this path does not run the malware gate the
normal model loader applies, so an authenticated client on an exposed
Studio instance could load a crafted Whisper-looking repo and run code
in the backend.

Restrict custom STT repos to safetensors: the snapshot selector no
longer falls back to pytorch_model.bin(.index.json), the cached-snapshot
completeness check ignores pickle weights, and the load forces
use_safetensors so a stray cached pickle still cannot execute. The five
curated Whisper defaults already ship safetensors only, so this changes
nothing for the built-in models.

* STT: reject safetensors indexes that reference non-safetensors shards

A safetensors index (model.safetensors.index.json) is attacker-supplied
JSON and can name pytorch_model-*.bin shards in its weight_map.
Transformers dispatches shard loading per file by extension, so those
.bin shards still load through torch.load (pickle) even with
use_safetensors set. Require every weight_map value to end in
.safetensors in both the snapshot selector and the completeness check so
no pickle shard is downloaded or reused.
2026-07-23 03:15:45 -07:00
..
sandbox_site Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
__init__.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
_html_to_md.py Studio: stream live tool output with SSE heartbeats, fix web page extraction, and surface interrupted turns (#7083) 2026-07-15 08:41:00 -07:00
_vulkan_probe.py Studio: add Vulkan llama.cpp support (#5819) 2026-07-09 03:39:48 -07:00
anthropic_compat.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
api_monitor.py Studio: trim serving-log noise and surface llama-server engine stats (#6377) 2026-06-17 05:37:57 -07:00
audio_codecs.py Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
chat_eos.py Studio: stop chat generation on the assistant-turn-end token (fixes Qwen3.5 loop) (#6804) 2026-07-06 10:07:56 -07:00
chat_template_helpers.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
chat_templates.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
defaults.py Add DeepSeek-V4-Flash-GGUF to Studio with none/high/max reasoning (#6908) 2026-07-07 06:13:43 -07:00
external_provider.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
inference.py fix(studio): ignore reasoning in tool reprompts (#7134) 2026-07-17 20:22:11 -03:00
key_exchange.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
llama_admission.py Studio: queue local GGUF OpenAI-compatible requests before llama-server (#7047) 2026-07-10 17:05:48 -03:00
llama_cpp.py Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
llama_http.py fix(studio/llama_cpp): disable trust_env on the loopback health probe (#6750) (#6752) 2026-06-30 19:09:26 +02:00
llama_keepwarm.py persist llama.cpp KV cache across idle auto-unload (slot save/restore) (#7204) 2026-07-20 00:12:42 -07:00
llama_server_args.py persist llama.cpp KV cache across idle auto-unload (slot save/restore) (#7204) 2026-07-20 00:12:42 -07:00
llama_stats.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
local_model_resolver.py Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
mcp_client.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
mcp_config_import.py studio: show MCP "Import config" on the add-server form (#6030) 2026-06-11 16:17:22 +01:00
message_content.py fix(studio): handle multimodal list content in inference text paths (#4383) (#6480) 2026-06-23 01:26:11 -07:00
mlx_inference.py Studio: reuse MLX prompt cache across turns instead of re-prefilling (#7311) 2026-07-22 02:35:33 -07:00
model_ids.py studio: list the full local model catalog from /v1/models (#6519) 2026-06-26 20:42:06 -03:00
orchestrator.py Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
passthrough_healing.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
presence_penalty.py Studio: apply presence_penalty on the safetensors and MLX inference paths (#6923) 2026-07-06 22:24:47 -07:00
pricing.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
providers.py Allow API key for Ollama connections (#7173) 2026-07-18 22:47:00 -07:00
runtime_context.py Expose runtime context length for hub models (#6154) 2026-06-11 22:13:53 +03:00
safetensors_agentic.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
stt_ggml_sidecar.py Studio: add local speech-to-text dictation engine (#7095) 2026-07-23 01:39:03 -07:00
stt_sidecar.py Studio STT: only load safetensors weights for custom dictation models (RCE fix) (#7364) 2026-07-23 03:15:45 -07:00
tensor_fallback.py studio: deterministic VRAM auto-fit for GGUF (MTP reserve, compute buffer, total-based budget) (#6312) 2026-06-17 03:10:22 -07:00
tool_call_parser.py Studio: Inkling support fixes (#7153) 2026-07-15 11:22:38 -07:00
tool_loop_controller.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
tool_stream_exec.py Studio: stream live tool output with SSE heartbeats, fix web page extraction, and surface interrupted turns (#7083) 2026-07-15 08:41:00 -07:00
tools.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
worker.py fix(studio): honor MLX adapter state in compare mode (#7196) 2026-07-19 00:27:55 -07:00