unsloth/studio/backend/utils
Daniel Han c2114d64dd
Studio: fail closed on index-referenced nested pickle shards in the offline embedding gate (#7366)
* Studio: fail closed on index-referenced nested pickle shards in the offline embedding gate

The offline embedding security gate (HF_HUB_OFFLINE / TRANSFORMERS_OFFLINE)
only scanned the direct files of each SentenceTransformer load root and never
parsed local weight indexes, so a cached snapshot whose pytorch_model.bin.index.json
maps a weight to a nested shard (e.g. shards/pytorch_model-00001-of-00001.bin) was
treated as inert and allowed. The loader then follows the index into the subdir and
unpickles the shard. The online gate already blocks index-referenced subdir pickles,
so the offline path was strictly weaker.

Parse each local weight index in a load root and follow weight_map into nested dirs,
flagging any referenced pickle-extension shard. Paths resolve lexically (normpath),
never Path.resolve(), since HF cache snapshot files symlink into blobs/ and resolving
would leave the snapshot dir and false-block every sharded model offline. An absolute
path, a .. traversal that escapes the snapshot, or an unreadable/invalid index fails
closed. The existing safetensors-sibling suppression is kept.

* Studio: classify offline indexed shards by torch.load path, not pickle extension

load_state_dict picks safetensors vs torch.load per shard by the shard's own
suffix, so two offline-gate gaps remained:

- A model.safetensors.index.json whose weight_map points at a .bin shard was
  suppressed by has_base_safetensors (the index file itself matches the base
  safetensors regex), yet Transformers still torch.loads that shard. Only the
  pytorch index is superseded by a base safetensors now; a safetensors index is
  the chosen archive, so its non-safetensors targets are always flagged.

- A pytorch index can map weights to arbitrary names (shards/payload,
  weights.data); the loader torch.loads any target not ending in .safetensors.
  Flag indexed shards by that rule instead of a pickle-extension allowlist.

Restrict the scan to the two torch-family indexes (tf/flax load via non-pickle
loaders). Add regression tests for both cases.

* Studio: match offline weight-index filenames case-insensitively

The index-name check compared the on-disk filename exactly, while the
surrounding weight and safetensors matches use case-insensitive rules. On a
case-insensitive volume (Windows or macOS) from_pretrained opens an oddly-cased
cache file such as PYTORCH_MODEL.BIN.INDEX.JSON when it requests the canonical
lowercase name, so the exact-case check skipped it and a nested pickle shard it
referenced was allowed through. Lower-case the index name before matching, as
the rest of the gate does, and add a regression test.

* Studio: match load_state_dict format/selection exactly in the offline index scan

Two edge cases in the offline weight-index scan:

- load_state_dict decides safetensors vs torch.load with a case-sensitive
  endswith(".safetensors"), so a shard named payload.SAFETENSORS still
  deserializes via torch.load. Classify indexed shard suffixes case-sensitively
  to match, instead of lower-casing (which treated such a shard as inert).

- A complete direct model.safetensors is selected before either sharded index,
  so a stale model.safetensors.index.json referencing a .bin shard never loads.
  Skip both indexes when a direct model.safetensors is present, so an otherwise
  loadable model is not over-blocked.

Add regression tests for both.

* Studio: read the offline weight index as UTF-8

Path.read_text() uses the locale default, which is cp1252 on Windows, so a
UTF-8 weight index with non-ASCII bytes raised UnicodeDecodeError and the gate
blocked an otherwise loadable model. JSON is UTF-8 by spec (and how the loader
reads it), so pin the encoding.

* Studio: resolve safetensors alternatives via the loader's own filename lookup

The offline gate decided a safetensors alternative existed by case-folding the
directory listing. On a case-sensitive filesystem that let an uppercase decoy
such as MODEL.SAFETENSORS suppress the pickle scan, yet from_pretrained asks for
the canonical lowercase model.safetensors, does not find the decoy, and selects
the pickle (a direct pytorch_model.bin or the pytorch index) and deserializes it.

Probe each alternative with (root / name).is_file() instead, mirroring the
loader: is_file() honors the platform's case rules, so a decoy suppresses only
where the loader would truly open it. Suppression must never fail open; detection
stays case-insensitive (fail closed). Add regression tests for the direct and
indexed pickle decoys (skipped on case-insensitive volumes, where no bypass
exists).

* Studio: resolve indexes and shards exactly as from_pretrained does

Two more loader-fidelity gaps in the offline index scan:

- Shard lookup normalized backslashes to forward slashes. On POSIX a backslash
  is a literal filename character, so an index naming dir\payload.bin matches a
  real pickle of that exact name that Transformers joins and deserializes, while
  the normalized dir/payload.bin missed it. Join the raw weight_map value with
  os.path.join so the probe mirrors the loader on each platform.

- Index detection case-folded the directory listing, so on a case-sensitive
  filesystem an uppercase PYTORCH_MODEL.BIN.INDEX.JSON artifact the loader never
  opens was treated as live and its shard blocked. Probe the canonical name with
  the loader's own is_file lookup instead, so an index counts only where
  from_pretrained would actually load it.

Update the uppercase-index tests to assert the correct per-filesystem behavior
and add a POSIX backslash-shard regression test.

---------

Co-authored-by: danielhanchen <unslothai@gmail.com>
2026-07-23 20:06:30 -07:00
..
datasets Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
hardware studio: show system-wide VRAM in the multi-GPU System tab view on ROCm (#7216) 2026-07-22 03:55:35 -07:00
inference Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
models Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
paths Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
prebuilt Studio: add local speech-to-text dictation engine (#7095) 2026-07-23 01:39:03 -07:00
security Studio: fail closed on index-referenced nested pickle shards in the offline embedding gate (#7366) 2026-07-23 20:06:30 -07:00
.gitkeep root studio folder 2026-02-02 09:13:49 +00:00
__init__.py Final cleanup 2026-03-12 18:28:04 +00:00
_studio_release_build.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
api_errors.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
cache_cleanup.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
client_ip.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
coding_agents.py feat: detect installed coding agent CLIs in Studio settings (#6909) 2026-07-08 05:26:50 -07:00
cpu_threads.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
downsample.py Formatting: ruff line-length 100, kwarg-spacing passes, drop blank after short local imports (#6079) 2026-06-08 04:24:13 -07:00
embedding_model_settings.py Studio: customizable RAG embedding model with HF search, settings tab reorganization (#6800) 2026-07-02 05:26:33 -07:00
helper_precache_settings.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
hf_cache_settings.py Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
hf_token_validation.py Studio: validate Hugging Face tokens before use (#7261) 2026-07-20 14:40:14 +01:00
hf_xet_fallback.py Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
hidden_models.py Studio: add local speech-to-text dictation engine (#7095) 2026-07-23 01:39:03 -07:00
host_policy.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
lifespan_shutdown.py Studio: make lifespan shutdown resilient to a dead default executor (#6307) 2026-06-15 22:51:46 -07:00
llama_cpp_freshness.py Studio: add local speech-to-text dictation engine (#7095) 2026-07-23 01:39:03 -07:00
llama_cpp_update.py Studio: add local speech-to-text dictation engine (#7095) 2026-07-23 01:39:03 -07:00
mlx_repair.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
native_path_leases.py Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
node_runtime.py Studio: use an isolated Node.js for the frontend build instead of replacing the system Node/npm (#6533) 2026-06-21 21:17:29 -07:00
openai_auto_switch_settings.py persist llama.cpp KV cache across idle auto-unload (slot save/restore) (#7204) 2026-07-20 00:12:42 -07:00
personalization_settings.py studio: persist personalization (profile, avatar, theme) server-side (#6516) 2026-06-22 04:09:48 -07:00
preview_rate_limit.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
preview_sharing_settings.py Studio: require signed capability tokens for /p preview links (#6666) 2026-06-25 21:40:48 -07:00
preview_token.py Studio: require signed capability tokens for /p preview links (#6666) 2026-06-25 21:40:48 -07:00
process_lifetime.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
ssm_runtime.py Auto-install SSM kernels (causal-conv1d, mamba-ssm) for inference loads (#6535) 2026-06-22 04:48:29 -07:00
studio_version.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
subprocess_compat.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
training_runs.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
transformers_dtype.py Studio: Fix torch_dtype deprecation warning on startup and ASR load (#6999) 2026-07-13 17:34:25 -03:00
transformers_latest.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
transformers_version.py Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
update_status.py Studio: make code comments and docstrings more succinct (#6029) 2026-06-08 23:07:28 -07:00
upload_limits.py Studio: add local speech-to-text dictation engine (#7095) 2026-07-23 01:39:03 -07:00
utils.py Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
uv_path_safety.py Make _uv_safe_path space-safe on macOS/Linux (#6503) (#6534) 2026-06-24 04:02:24 -07:00
wheel_utils.py Studio: fix flash-attn and torchao install on Blackwell (sm_100+) GPUs (Closes #6961) (#6970) 2026-07-08 06:38:10 -07:00
whisper_cpp_freshness.py Studio: add local speech-to-text dictation engine (#7095) 2026-07-23 01:39:03 -07:00
whisper_cpp_update.py Studio: add local speech-to-text dictation engine (#7095) 2026-07-23 01:39:03 -07:00