unsloth/studio/backend/core/inference
Daniel Han fc04809bfe
Studio: warn when llama.cpp prebuilt is too old for MTP (#5528)
* Studio: warn when llama.cpp prebuilt is too old for MTP

Layered on #5527. Adds a one-shot llama-server --help capability probe
so users get a clear signal when their prebuilt is missing MTP support,
plus a graceful fallback if they load an MTP GGUF against an outdated
binary.

What's surfaced:

1. Startup log + stderr line in main.py:lifespan() if MTP isn't
   advertised:
     WARNING: llama.cpp prebuilt is missing MTP support
     (--spec-type mtp / draft-mtp). Run `unsloth studio update` to
     refresh it. MTP GGUFs will load without speculative decoding.
2. Load-time graceful fallback in load_model's spec block: skip the
   auto-emit and log a clear warning instead of letting llama-server
   fail with an unknown-flag error.
3. /api/inference/status now returns llama_cpp_supports_mtp: bool so
   the frontend can show a banner / popup.

Probe internals:

- Class-level cache keyed on (binary_path, mtime). One subprocess call
  the first time, instant thereafter. Touching the binary (e.g. via
  `unsloth studio update`) invalidates the cache automatically because
  the mtime changes, so the new build is picked up without restarting
  the server.
- Recognises both upstream naming forms: the original draft-mtp from
  llama.cpp PR #22673 and the renamed mtp variant in later commits.
- Spec block uses whichever token the binary accepts so we emit the
  right value regardless of which release the user has.

Tests:

- 6 new cases in test_llama_cpp_mtp_detection.py covering each probe
  variant (draft-mtp, renamed mtp, pre-MTP build, missing binary,
  mtime-based cache invalidation).
- Existing 38 MTP detection cases still pass; broader 188-test
  regression suite (server args, reload inheritance, gguf metadata,
  load progress, context fit, model validation) still green.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-18 00:19:47 -07:00
..
__init__.py Final cleanup 2026-03-12 18:28:04 +00:00
_html_to_md.py fix: studio web search SSL failures and empty page content (#4754) 2026-04-01 06:12:02 -07:00
anthropic_compat.py Studio: support images on /v1/messages (Anthropic-compat) (#5128) 2026-04-22 03:25:07 +04:00
audio_codecs.py Add native GGUF intake to Studio (#5246) 2026-05-04 11:46:18 +02:00
defaults.py Add Qwen3.6 inference defaults for Studio (#5065) 2026-04-16 11:42:42 -07:00
external_provider.py studio/chat: reuse Anthropic code_execution container across turns (#5519) 2026-05-17 17:49:38 +04:00
inference.py Pin bitsandbytes to continuous-release_main on ROCm (4-bit decode fix) (#4954) 2026-04-10 06:25:39 -07:00
key_exchange.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
llama_cpp.py Studio: warn when llama.cpp prebuilt is too old for MTP (#5528) 2026-05-18 00:19:47 -07:00
llama_server_args.py Studio: auto-enable MTP speculative decoding for MTP GGUFs (#5527) 2026-05-18 00:15:42 -07:00
mlx_inference.py MLX training support for Studio on Apple Silicon (#5340) 2026-05-14 05:24:20 -07:00
orchestrator.py Add native GGUF intake to Studio (#5246) 2026-05-04 11:46:18 +02:00
providers.py Polish/cloud to providers (#5450) 2026-05-15 19:29:21 +04:00
tools.py studio: tighten sandbox blocklist precision (bash, hf upload, NOFILE) (#5487) 2026-05-18 00:01:17 -07:00
worker.py studio: load cached GGUF models when fully offline (#5505) 2026-05-17 21:25:39 -07:00