* Studio: GPU memory dropdown — llama.cpp --fit on and manual gpu-layers/cpu-moe * Studio: simplify GPU memory changes (reuse ParamSlider, GPU_LAYERS_ALL, loadedGpuMemoryFields helper) * Studio: GPU picker — choose which GPUs a GGUF model loads on (gpu_ids) * Studio: simplify GPU picker (share /api/system fetch, validate gpu_ids) * Studio: GPU picker review fixes (gate relative indices, no cross-model leak, validate, types) * Studio: group GPU controls under a collapsible GPU section * Studio: GPU feature review fixes (fix fit-ctx test, behavior-test the floor, comment accuracy) * Studio: make GPU a top-level settings section (not nested under Model) * Studio: flatten GPU controls into the Model section, group by GPU/context/generation * Studio: move GPU Memory to the bottom of Model with its dependent controls beneath it * Studio: move GPU Memory below Tensor Parallelism and GPUs below GPU Memory * Studio: tighten GPU Memory and GPU Layers tooltip copy * Studio: fix fit-mode context slider track-click, restore GPU Memory tooltip, shorten fit dropdown label * Studio: GPU Memory tooltip one mode per line, briefer * Studio: note HIP_VISIBLE_DEVICES (ROCm) in the GPUs picker tooltip * Studio: narrow the GPU Memory dropdown to fit the shortened label * Studio: use 'llama.cpp --fit' in the GPU Memory tooltip for consistency * Studio: allow Tensor Parallelism in Manual GPU mode * Studio: graduated MoE-on-CPU offload (--n-cpu-moe) replacing the all-or-nothing toggle * Studio: size the MoE-offload slider for staged (deferred-load) models * Studio: share one GGUF header walk for the context-length and MoE-count readers * Studio: size the GPU Layers slider for staged models (one staged-header read) * Studio: move Tensor Parallelism below the GPUs picker * Studio: GPU split (--tensor-split) per-GPU model share in Manual mode * Studio: tolerate whitespace in GPU split input, move it below GPU Layers * Studio: rename the GPU split control to "Split ratio" * Studio: Split ratio sends explicit even input; fix blank=free-VRAM (not even) copy * Studio: tighten llama.cpp --fit VRAM margin with --fit-target 512 * Studio: GPU memory review fixes (rollback re-baseline, single-GPU TP gate, accurate copy) * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: move Split ratio below MoE Layers on CPU * Studio: address PR review (fix GPU-info hydration race, share fit context-length across load paths) * Studio: address codex review (manual single-GPU TP guard, GPU-aware spec defaults in fit/manual, GGUF-only context/preference) * Studio: address codex review round 2 (gpu_present seed, single-GPU tensor-split guard, staged manual-knob reset, strip inherited offload flags) * Studio: address codex review round 3 (strip inherited --n-cpu-moe, CPU-fallback warning in Manual mode) * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: address codex review round 4 (preserve pinned fit context across a later Apply) * Studio: address codex review round 5 (honor GPU picker for diffusion GGUFs, clear fit pin on cross-model switch) * Studio: preserve the pending GPU Memory mode when staging a model * Studio: pin diffusion GPU device order and reset GPU-memory state for diffusion loads * Studio: address codex review round 6 (fit-Auto rollback context, preserve manual non-tensor split modes, persist GPU mode on load not select) * Studio: persist the applied GPU Memory mode, not the requested one (skip diffusion loads) * Studio: replace Manual-mode split-ratio field with per-GPU layer sliders * Studio: clarify per-GPU layer split hint for tensor-parallel mode * Studio: address codex review round 7 (allow GGUF gpu_ids past the legacy guard, replay GPU-memory fields on respawn) * Studio: address codex review round 8 (size the validate preflight like the load in fit mode, across both load paths) * Studio: skip the training-OOM guard for llama.cpp --fit GGUF loads (they spill to RAM) * Studio: drop the now-redundant compare-path validate sizing (the --fit guard skip makes it moot) * Studio: address codex review round 9 (keep the training guard for fit loads, forward gpu_ids to validate, strip inherited manual tensor-split) * Studio: address codex review round 10 (gate GPU-memory adoption on is_gguf, record manual knobs only in Manual mode) * Studio: handle diffusion GGUFs symmetrically in the GPU Memory controls (preserve the standing mode preference, hide the inapplicable mode/TP controls) * Studio: remember the GPU Memory settings per model * Studio: consolidate --fit mode and Manual mode into a single Manual mode * Studio: preserve the per-GPU layer split across GPU Layers changes * Studio: trim overly long GPU Memory comments * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * address GPU memory config review comments * trim redundant GPU memory tests * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Reconcile manual-mode TP drops with the #6659 drop-site invariants * Preserve quantized KV in manual --fit, charge GGUF companions in full, reconcile GPU pick on load * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Clear stale GPU baseline on non-GGUF loads so it can't read as dirty * Fix no-context-shift test for the conditional -c flag * Credit manual GPU-layer offload for cached HF GGUFs * Reset per-model load knobs on GGUF quant switch * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Strip inherited tensor-split when manual ratio is cleared * Match auto-load validation to safetensors placement * Reset editable manual knobs after Auto GGUF loads * Record a single device for diffusion GPU picks * Reset per-model GPU knobs before applying saved settings * Address review comments * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Guard manual tensor splits and keep remembered context on auto-load * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Snapshot compare knobs, seed splits from free VRAM, flag zero-offload loads * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Exempt CPU-only loads from the guard floor and harden compare and reseed paths * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Reach full offload from the layers slider and charge extras drafters in the guard * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Warm the GPU device cache before pick reconciles and disable staged GPU controls * Align the training guard with inherited extras, spec mode, and compare targets * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Hide GPUs from companion-less zero-offload loads * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Size diffusion picks per device, own manual offload flags, reject XPU picks * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Drop tensor flags at zero layers and exempt CPU-pinned drafters * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Allowlist the zero-layer tensor parallel drop site * Keep validate and load guards on the same extras and refresh stale baselines * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Drop mismatched manual tensor splits before launch * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Gate XPU picks on the real backend field and harden split and hydration paths * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Weight full GPUs as zero, clamp split shares, and refine the zero-layer mask gate * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Carry fit context across mode changes and align drafter and picker gates * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Catch variant switches, uncached diffusion repos, and text-only mmproj skips * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Check companions on the first device and size native and remote zero-layer loads * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Replace the training guard's precise VRAM modeling with a conservative bound * Baseline context pins on non-GGUF hydration and reprobe list-seeded staged GGUFs * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Size manual splits by their largest share and preserve resolved context from Default * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Default-deny unsized required companions and price KV at the effective cache dtype * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Reserve MTP draft KV and MLA target-copy in the training guard * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Size tensor-parallel loads per device and show GPU controls for native GGUFs * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Reserve MTP overhead for uncached remote GGUFs and the mmproj runtime factor * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Drop the training-coexistence VRAM estimation this PR added * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Gate remembered load settings to GGUF picks * Lock the remaining load-time controls during a staged load * Clear the stale native-path token on compare loads * Drop a stale guard reference from the zero-offload masking comment * Seed GPU baselines from the rollback response and drop never-emitted offload flags * Match validate's training guard to load and keep the native reload token * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Trim verbose GPU-memory comments * Thread the variants header walk off the event loop, honor device pins on zero-offload, and hold staged GPU edits * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Honor manual placement and classify pinned zero-offload loads * Close diffusion admission and status hydration gaps * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Check the actual diffusion GPU during training * Align staged baselines and manual reload dedupe * Fix GGUF placement and rollback state * Harden manual GGUF placement boundaries * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Remove unused resolve_tensor_parallel import in llama_cpp.py The name is used only in llama_server_args.py, routes/inference.py, and tests, not in llama_cpp.py; the unused hoisted import trips the import-hoist verifier in the source-lint CI job. * Fix diffusion GPU dedup and training guard for non-numeric device tokens The diffusion runner drives only its single lowest device and the backend records that one device (self._gpu_ids = [sorted(gpu_ids)[0]]), but the reload dedupe compared it against the full requested list, so a multi-GPU pick that resolves to the same device forced a needless reload. Normalize the request the same way for a loaded diffusion model in both _already_in_target_state and the route _request_matches_loaded_settings. The chat-during-training coexistence guard called int() on the single-device token and hard-rejected when it could not parse. A non-numeric token (a CUDA UUID / MIG handle) now sizes against the whole visible pool like the GGUF guard instead of falsely blocking the load, and an empty token (a CPU-only runner such as a CPU diffusion GGUF) is allowed outright since it uses no GPU VRAM. * Tighten comments added by the GPU memory config changes * Harden GGUF placement from independent review: VRAM sizing, diffusion TP reset, tensor_split validation - Training coexistence guard: a single-device runner pinned through an unresolvable UUID/MIG token was sized against the aggregate visible-VRAM pool, so a load could pass on capacity it cannot use and then OOM active training. Size against the worst-case visible device (min free) instead, keeping the guard's documented default-deny contract. The empty-token (CPU-only runner) allow path is unchanged. - Diffusion startup: _start_diffusion_server now resets self._tensor_parallel to False alongside the other placement resets. A prior tensor-parallel chat load (process killed but not fully unload-reset) otherwise left /status misreporting tensor parallelism and made an identical diffusion re-Apply reload against the stale state. - tensor_split: reject negative / non-finite / all-zero splits up front. They were dropped at launch but still compared raw in the reload dedupe, so an identical Apply reloaded indefinitely. - Tests: the shared httpx stub was incomplete and, installed via setdefault before real httpx loaded, broke a combined pytest run (collection errors on httpx.Response). Import the real installed httpx instead. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: danielhanchen <unslothshared@gmail.com> Co-authored-by: danielhanchen <danielhanchen@gmail.com>
457 lines
16 KiB
Python
457 lines
16 KiB
Python
# SPDX-License-Identifier: AGPL-3.0-only
|
|
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
|
|
|
"""``general.*`` reader for GGUF headers, used by ``detect_mmproj_file`` to
|
|
pair weights and projectors via ``general.base_model.0.repo_url``. ~30 ms
|
|
per file, cached by (path, mtime, size)."""
|
|
|
|
from __future__ import annotations
|
|
|
|
import os
|
|
import struct
|
|
import threading
|
|
from pathlib import Path
|
|
from typing import Dict, Optional, Tuple
|
|
|
|
from loggers import get_logger
|
|
|
|
logger = get_logger(__name__)
|
|
|
|
|
|
_GGUF_MAGIC = 0x46554747 # b"GGUF" LE u32
|
|
|
|
_WANTED_GENERAL_KEYS: frozenset[str] = frozenset(
|
|
{
|
|
"general.architecture",
|
|
"general.type",
|
|
"general.name",
|
|
"general.basename",
|
|
"general.organization",
|
|
"general.size_label",
|
|
"general.finetune",
|
|
"general.base_model.0.name",
|
|
"general.base_model.0.organization",
|
|
"general.base_model.0.repo_url",
|
|
"general.repo_url",
|
|
"general.source.url",
|
|
"general.source.repo_url",
|
|
"general.source.huggingface.repository",
|
|
}
|
|
)
|
|
|
|
|
|
# Cache failed parses too so a broken file is not retried each scan.
|
|
_CacheKey = Tuple[str, int, int]
|
|
_METADATA_CACHE: Dict[_CacheKey, Optional[Dict[str, str]]] = {}
|
|
_CACHE_LOCK = threading.Lock()
|
|
_CACHE_MAX_ENTRIES = 4096
|
|
|
|
# Separate cache for single bool capability keys (e.g. clip.has_audio_encoder),
|
|
# keyed by (file cache key, wanted key). None = key absent / file unreadable.
|
|
_BOOL_CACHE: Dict[Tuple[_CacheKey, str], Optional[bool]] = {}
|
|
|
|
# GGUF header dims for the staged/deferred-load UI: context_length, layer_count
|
|
# (block_count), and moe_layer_count (block_count minus leading dense layers; 0
|
|
# if not MoE). One cached pass fills all three so the staged sheet can size every
|
|
# slider before the model loads. None = unreadable / not a GGUF.
|
|
_DIMS_CACHE: Dict[_CacheKey, Optional[Dict[str, Optional[int]]]] = {}
|
|
|
|
|
|
def _cache_key(path: str) -> Optional[_CacheKey]:
|
|
try:
|
|
st = os.stat(path)
|
|
except OSError:
|
|
return None
|
|
try:
|
|
resolved = str(Path(path).resolve())
|
|
except OSError:
|
|
resolved = str(path)
|
|
return (resolved, st.st_mtime_ns, st.st_size)
|
|
|
|
|
|
def read_gguf_general_metadata(path: str) -> Optional[Dict[str, str]]:
|
|
"""Return ``general.*`` strings from a GGUF header, or ``None`` if the
|
|
file is missing, unreadable, or not a GGUF. ``{}`` means valid but
|
|
carrying none of the wanted keys."""
|
|
key = _cache_key(path)
|
|
if key is None:
|
|
return None
|
|
with _CACHE_LOCK:
|
|
if key in _METADATA_CACHE:
|
|
return _METADATA_CACHE[key]
|
|
result = _parse_gguf_header(path)
|
|
with _CACHE_LOCK:
|
|
# Arbitrary eviction; header reads are cheap so true LRU is overkill.
|
|
while len(_METADATA_CACHE) >= _CACHE_MAX_ENTRIES:
|
|
try:
|
|
_METADATA_CACHE.pop(next(iter(_METADATA_CACHE)))
|
|
except StopIteration:
|
|
break
|
|
_METADATA_CACHE[key] = result
|
|
return result
|
|
|
|
|
|
def _parse_gguf_header(path: str) -> Optional[Dict[str, str]]:
|
|
out: Dict[str, str] = {}
|
|
try:
|
|
with open(path, "rb") as f:
|
|
head = f.read(24)
|
|
if len(head) < 24:
|
|
return None
|
|
magic, _version, _tcount, kv_count = struct.unpack("<IIQQ", head)
|
|
if magic != _GGUF_MAGIC:
|
|
return None
|
|
|
|
for _ in range(kv_count):
|
|
try:
|
|
klen_bytes = f.read(8)
|
|
if len(klen_bytes) < 8:
|
|
break
|
|
klen = struct.unpack("<Q", klen_bytes)[0]
|
|
if klen > 1 << 20: # 1 MB sanity bound
|
|
break
|
|
kbytes = f.read(klen)
|
|
if len(kbytes) < klen:
|
|
break
|
|
key = kbytes.decode("utf-8", "replace")
|
|
vt_bytes = f.read(4)
|
|
if len(vt_bytes) < 4:
|
|
break
|
|
vtype = struct.unpack("<I", vt_bytes)[0]
|
|
|
|
if vtype == 8 and key in _WANTED_GENERAL_KEYS:
|
|
slen_bytes = f.read(8)
|
|
if len(slen_bytes) < 8:
|
|
break
|
|
slen = struct.unpack("<Q", slen_bytes)[0]
|
|
if slen > 1 << 22: # 4 MB sanity bound
|
|
break
|
|
sbytes = f.read(slen)
|
|
if len(sbytes) < slen:
|
|
break
|
|
out[key] = sbytes.decode("utf-8", "replace")
|
|
else:
|
|
if not _skip_gguf_value(f, vtype):
|
|
break
|
|
except (struct.error, UnicodeDecodeError):
|
|
break
|
|
except OSError as e:
|
|
logger.debug(f"read_gguf_general_metadata: cannot open {path}: {e}")
|
|
return None
|
|
except Exception as e:
|
|
logger.debug(f"read_gguf_general_metadata: parse failure on {path}: {e}")
|
|
return None
|
|
return out
|
|
|
|
|
|
def read_gguf_staged_dims(path: str) -> Optional[Dict[str, Optional[int]]]:
|
|
"""GGUF header dims for the staged-load UI in one cached pass:
|
|
``{"context_length", "layer_count", "moe_layer_count"}``. Each may be None
|
|
when absent (moe_layer_count is 0 for a dense model). Returns ``None`` if not
|
|
a GGUF / unreadable. Cached by (path, mtime, size). Lets the staged sheet size
|
|
the context, GPU-layers and MoE sliders before the model loads."""
|
|
key = _cache_key(path)
|
|
if key is None:
|
|
return None
|
|
with _CACHE_LOCK:
|
|
if key in _DIMS_CACHE:
|
|
return _DIMS_CACHE[key]
|
|
result = _parse_gguf_staged_dims(path)
|
|
with _CACHE_LOCK:
|
|
while len(_DIMS_CACHE) >= _CACHE_MAX_ENTRIES:
|
|
try:
|
|
_DIMS_CACHE.pop(next(iter(_DIMS_CACHE)))
|
|
except StopIteration:
|
|
break
|
|
_DIMS_CACHE[key] = result
|
|
return result
|
|
|
|
|
|
def read_gguf_context_length(path: str) -> Optional[int]:
|
|
"""Native training context length (``{arch}.context_length``), or ``None``.
|
|
Thin accessor over read_gguf_staged_dims."""
|
|
dims = read_gguf_staged_dims(path)
|
|
return dims["context_length"] if dims else None
|
|
|
|
|
|
def _parse_gguf_arch_uints(path: str, wanted_suffixes: frozenset[str]) -> Optional[Dict[str, int]]:
|
|
"""Walk a GGUF header once and return the requested architecture-namespaced
|
|
uint (vtype 4/10) keys, e.g. ``{"block_count": 32}``. Keys are
|
|
``{arch}.<suffix>``; the arch is learned from ``general.architecture`` (GGUF
|
|
writes general.* before arch.* keys, matching the loader's own parser).
|
|
Returns ``None`` if not a GGUF / unreadable, else a dict (possibly empty or
|
|
partial when some keys are absent)."""
|
|
arch: Optional[str] = None
|
|
found: Dict[str, int] = {}
|
|
try:
|
|
with open(path, "rb") as f:
|
|
head = f.read(24)
|
|
if len(head) < 24:
|
|
return None
|
|
magic, _version, _tcount, kv_count = struct.unpack("<IIQQ", head)
|
|
if magic != _GGUF_MAGIC:
|
|
return None
|
|
|
|
for _ in range(kv_count):
|
|
try:
|
|
klen_bytes = f.read(8)
|
|
if len(klen_bytes) < 8:
|
|
break
|
|
klen = struct.unpack("<Q", klen_bytes)[0]
|
|
if klen > 1 << 20: # 1 MB sanity bound
|
|
break
|
|
kbytes = f.read(klen)
|
|
if len(kbytes) < klen:
|
|
break
|
|
key = kbytes.decode("utf-8", "replace")
|
|
vt_bytes = f.read(4)
|
|
if len(vt_bytes) < 4:
|
|
break
|
|
vtype = struct.unpack("<I", vt_bytes)[0]
|
|
|
|
if vtype == 8 and key == "general.architecture":
|
|
slen_bytes = f.read(8)
|
|
if len(slen_bytes) < 8:
|
|
break
|
|
slen = struct.unpack("<Q", slen_bytes)[0]
|
|
if slen > 1 << 22: # 4 MB sanity bound
|
|
break
|
|
sbytes = f.read(slen)
|
|
if len(sbytes) < slen:
|
|
break
|
|
arch = sbytes.decode("utf-8", "replace")
|
|
elif (
|
|
arch is not None
|
|
and vtype in (4, 10)
|
|
and key.startswith(f"{arch}.")
|
|
and key[len(arch) + 1 :] in wanted_suffixes
|
|
):
|
|
width = 4 if vtype == 4 else 8
|
|
n_bytes = f.read(width)
|
|
if len(n_bytes) < width:
|
|
break
|
|
found[key[len(arch) + 1 :]] = struct.unpack(
|
|
"<I" if vtype == 4 else "<Q", n_bytes
|
|
)[0]
|
|
if len(found) == len(wanted_suffixes):
|
|
break
|
|
else:
|
|
if not _skip_gguf_value(f, vtype):
|
|
break
|
|
except (struct.error, UnicodeDecodeError):
|
|
break
|
|
except OSError as e:
|
|
logger.debug(f"_parse_gguf_arch_uints: cannot open {path}: {e}")
|
|
return None
|
|
except Exception as e:
|
|
logger.debug(f"_parse_gguf_arch_uints: parse failure on {path}: {e}")
|
|
return None
|
|
return found
|
|
|
|
|
|
def _parse_gguf_staged_dims(path: str) -> Optional[Dict[str, Optional[int]]]:
|
|
vals = _parse_gguf_arch_uints(
|
|
path,
|
|
frozenset(
|
|
{
|
|
"context_length",
|
|
"block_count",
|
|
"expert_count",
|
|
"leading_dense_block_count",
|
|
}
|
|
),
|
|
)
|
|
if vals is None:
|
|
return None
|
|
ctx = vals.get("context_length")
|
|
block = vals.get("block_count")
|
|
# A real context/layer count is positive; treat 0/garbage as absent so the
|
|
# UI never builds a slider with max < min.
|
|
context_length = ctx if ctx and ctx > 0 else None
|
|
layer_count = block if block and block > 0 else None
|
|
# MoE layer count = block_count - leading dense layers, only when experts
|
|
# exist; else 0 (dense -> slider hidden). Mirrors n_moe_layers in
|
|
# core/inference/llama_cpp.py.
|
|
if not vals.get("expert_count") or not block:
|
|
moe_layer_count: Optional[int] = 0
|
|
else:
|
|
moe_layer_count = max(0, block - (vals.get("leading_dense_block_count") or 0))
|
|
return {
|
|
"context_length": context_length,
|
|
"layer_count": layer_count,
|
|
"moe_layer_count": moe_layer_count,
|
|
}
|
|
|
|
|
|
# Strings (8) and arrays (9) are handled inline.
|
|
_FIXED_VTYPE_SIZES: Dict[int, int] = {
|
|
0: 1, # uint8
|
|
1: 1, # int8
|
|
2: 2, # uint16
|
|
3: 2, # int16
|
|
4: 4, # uint32
|
|
5: 4, # int32
|
|
6: 4, # float32
|
|
7: 1, # bool
|
|
10: 8, # uint64
|
|
11: 8, # int64
|
|
12: 8, # float64
|
|
}
|
|
|
|
|
|
def _skip_gguf_value(f, vtype: int) -> bool:
|
|
"""Advance past one GGUF value. ``f.seek(.., 1)`` past EOF is legal on a
|
|
regular file, so truncation is caught on the next read; return False only
|
|
for unknown types or sanity-bound overflow."""
|
|
if vtype == 8: # STRING
|
|
slen_bytes = f.read(8)
|
|
if len(slen_bytes) < 8:
|
|
return False
|
|
slen = struct.unpack("<Q", slen_bytes)[0]
|
|
if slen > 1 << 30: # 1 GB sanity bound
|
|
return False
|
|
f.seek(slen, 1)
|
|
return True
|
|
if vtype == 9: # ARRAY
|
|
head = f.read(12)
|
|
if len(head) < 12:
|
|
return False
|
|
atype, alen = struct.unpack("<IQ", head)
|
|
if alen > 1 << 30:
|
|
return False
|
|
if atype == 8:
|
|
for _ in range(alen):
|
|
slen_bytes = f.read(8)
|
|
if len(slen_bytes) < 8:
|
|
return False
|
|
slen = struct.unpack("<Q", slen_bytes)[0]
|
|
if slen > 1 << 30:
|
|
return False
|
|
f.seek(slen, 1)
|
|
return True
|
|
sz = _FIXED_VTYPE_SIZES.get(atype)
|
|
if sz is None:
|
|
return False
|
|
f.seek(sz * alen, 1)
|
|
return True
|
|
sz = _FIXED_VTYPE_SIZES.get(vtype)
|
|
if sz is None:
|
|
return False
|
|
f.seek(sz, 1)
|
|
return True
|
|
|
|
|
|
def _parse_gguf_bool(path: str, wanted_key: str) -> Optional[bool]:
|
|
"""Bool value of ``wanted_key`` (GGUF vtype 7), or ``None`` if absent /
|
|
unreadable. Mirrors ``_parse_gguf_header`` for a single bool key."""
|
|
try:
|
|
with open(path, "rb") as f:
|
|
head = f.read(24)
|
|
if len(head) < 24:
|
|
return None
|
|
magic, _version, _tcount, kv_count = struct.unpack("<IIQQ", head)
|
|
if magic != _GGUF_MAGIC:
|
|
return None
|
|
|
|
for _ in range(kv_count):
|
|
try:
|
|
klen_bytes = f.read(8)
|
|
if len(klen_bytes) < 8:
|
|
break
|
|
klen = struct.unpack("<Q", klen_bytes)[0]
|
|
if klen > 1 << 20: # 1 MB sanity bound
|
|
break
|
|
kbytes = f.read(klen)
|
|
if len(kbytes) < klen:
|
|
break
|
|
key = kbytes.decode("utf-8", "replace")
|
|
vt_bytes = f.read(4)
|
|
if len(vt_bytes) < 4:
|
|
break
|
|
vtype = struct.unpack("<I", vt_bytes)[0]
|
|
|
|
if key == wanted_key and vtype == 7: # BOOL (1 byte)
|
|
bbyte = f.read(1)
|
|
if len(bbyte) < 1:
|
|
break
|
|
return bbyte[0] != 0
|
|
if not _skip_gguf_value(f, vtype):
|
|
break
|
|
except (struct.error, UnicodeDecodeError):
|
|
break
|
|
except OSError as e:
|
|
logger.debug(f"_parse_gguf_bool: cannot open {path}: {e}")
|
|
return None
|
|
except Exception as e:
|
|
logger.debug(f"_parse_gguf_bool: parse failure on {path}: {e}")
|
|
return None
|
|
return None
|
|
|
|
|
|
def _read_gguf_bool(path: str, wanted_key: str) -> Optional[bool]:
|
|
"""Cached single-bool-key read, keyed by (path, mtime, size, wanted_key)."""
|
|
fkey = _cache_key(path)
|
|
if fkey is None:
|
|
return None
|
|
ckey = (fkey, wanted_key)
|
|
with _CACHE_LOCK:
|
|
if ckey in _BOOL_CACHE:
|
|
return _BOOL_CACHE[ckey]
|
|
result = _parse_gguf_bool(path, wanted_key)
|
|
with _CACHE_LOCK:
|
|
while len(_BOOL_CACHE) >= _CACHE_MAX_ENTRIES:
|
|
try:
|
|
_BOOL_CACHE.pop(next(iter(_BOOL_CACHE)))
|
|
except StopIteration:
|
|
break
|
|
_BOOL_CACHE[ckey] = result
|
|
return result
|
|
|
|
|
|
def read_mmproj_audio_capability(path: str) -> Optional[bool]:
|
|
"""``clip.has_audio_encoder`` from an mmproj GGUF (e.g. Gemma 4's
|
|
gemma4ua): ``True``/``False`` if present, ``None`` if absent/unreadable.
|
|
Flags audio-input models independently of tokenizer token names."""
|
|
return _read_gguf_bool(path, "clip.has_audio_encoder")
|
|
|
|
|
|
def is_mmproj_by_metadata(meta: Optional[Dict[str, str]]) -> Optional[bool]:
|
|
"""True/False from ``general.type``; None means fall back to filename."""
|
|
if not meta:
|
|
return None
|
|
t = meta.get("general.type")
|
|
if t is None:
|
|
return None
|
|
return t.lower() == "mmproj"
|
|
|
|
|
|
def pairing_score(
|
|
weight_meta: Optional[Dict[str, str]], mmproj_meta: Optional[Dict[str, str]]
|
|
) -> int:
|
|
"""Pairing confidence: 100 = base_model URL match, 80 = basename + org,
|
|
60 = basename, -1 = definitive mismatch, 0 = decide from filename."""
|
|
if not weight_meta or not mmproj_meta:
|
|
return 0
|
|
|
|
w_url = weight_meta.get("general.base_model.0.repo_url")
|
|
p_url = mmproj_meta.get("general.base_model.0.repo_url")
|
|
if w_url and p_url:
|
|
return 100 if w_url.strip().rstrip("/") == p_url.strip().rstrip("/") else -1
|
|
|
|
w_base = weight_meta.get("general.basename")
|
|
p_base = mmproj_meta.get("general.basename")
|
|
w_org = weight_meta.get("general.base_model.0.organization") or weight_meta.get(
|
|
"general.organization"
|
|
)
|
|
p_org = mmproj_meta.get("general.base_model.0.organization") or mmproj_meta.get(
|
|
"general.organization"
|
|
)
|
|
if w_base and p_base and w_org and p_org:
|
|
if w_base.lower() == p_base.lower() and w_org.lower() == p_org.lower():
|
|
return 80
|
|
return -1
|
|
|
|
if w_base and p_base:
|
|
return 60 if w_base.lower() == p_base.lower() else -1
|
|
|
|
return 0
|