* Studio: cap GGUF context to unified memory on Apple Silicon
* Studio: tighten Apple ctx-cap comments and drop the overstated MLX-sync claim
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Studio: reserve flat MTP fraction and floor sparse-KV ctx in the Apple unified-memory cap
The Apple Silicon GGUF context cap mirrored the discrete-GPU auto-fit branch but
missed two protections the discrete path already applies:
- It passed the full unified-memory budget with budget_frac=1.0 without first
reserving the flat MTP fraction the discrete path takes off via _pin_fraction.
With an MTP draft whose KV cannot be byte-sized (e.g. Qwen3.6-MTP, #6529), the
cap filled the whole budget and left nothing for the draft, so unified memory
could still over-commit. Reserve _flat_mtp_reserve up front; this is a no-op
when MTP is not engaged.
- It required _can_estimate_kv(), so a GGUF with sparse KV metadata skipped the
cap entirely and launched at full native context. Mirror the discrete
file-size-only fallback and floor the auto context to 4096 when the cache
cannot be sized.
Adds regression tests for both paths.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Studio: tighten comments in the Apple unified-memory context cap
Condense the verbose comment blocks in the Apple budget helper, the no-GPU
Metal branch, and the context-fit tests. Comments only, no code change
(verified with ast-based comment_tools check); suite still green.
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: danielhanchen <danielhanchen@gmail.com>