unsloth/studio/backend/core
Daniel Han 068e129cb2 fix(mtp): gate legacy --draft-min with --draft-max on chained ngram
Codex flagged that the legacy chained mtp+ngram path emits --draft-min 48
while --draft-max is suppressed and later reused for MTP (spec_draft_n_max,
typically 2/3). That produces an inverted legacy ngram range
(--draft-min 48 --draft-max 2/3) on affected binaries, which can break or
effectively disable ngram-mod for auto CPU MTP loads and forced mtp+ngram
requests.

Both --draft-min and --draft-max are generic flags on legacy llama-server
builds, so either one would race with MTP's own values. Gate the pair
together: when chain_with_mtp=True on the legacy flavor we drop both flags
and rely on MTP's emission for the chained range. Standalone ngram still
emits both, preserving a valid min<=max window.

Updated test_build_ngram_mod_flags_legacy_chained_omits_draft_max (now
omits_draft_min_and_max) and added a min<=max guard on the standalone
case. Full suite (test_llama_cpp_mtp_detection + test_llama_server_args)
passes locally: 229 / 229.
2026-05-23 16:42:39 +00:00
..
data_recipe chore(deps): bump the npm-oxc-validator group across 1 directory with 2 updates (#5667) 2026-05-22 04:46:20 -07:00
export feat(studio): MLX training tab on Apple Silicon (LoRA / full FT, VLM, export) (#5265) 2026-05-05 23:54:58 -07:00
inference fix(mtp): gate legacy --draft-min with --draft-max on chained ngram 2026-05-23 16:42:39 +00:00
training studio: install flash-linear-attention and tilelang for Qwen3.5 family (#5434) 2026-05-18 03:49:06 -07:00
__init__.py [Studio] Show non exported models in chat UI (#4892) 2026-04-14 15:03:58 +04:00
tool_healing.py studio: extract tool-call XML parser into a reusable helper module (#5583) 2026-05-19 05:06:17 -07:00