_estimate_mtp_overhead_bytes is also reached for the separate-drafter spec modes (draft-simple / draft-eagle3) through _user_draft_via_extras. Those modes load a small distinct drafter with its own KV -- already counted in the draft KV + weights -- and keep no duplicated full target context; only MTP runs a second context over the target model's own KV geometry (llama.cpp ctx_tgt). Charging the ~main-KV-sized f16 copy there over-reserved by tens of GiB on an MLA model and needlessly shrank the advertised context, the same under-advertising #6312 set out to fix. Thread mtp_keeps_target_ctx through _estimate_mtp_overhead_bytes (True for MTP, False for separate-drafter modes) and derive _engaged_is_mtp at the fit call site so the target copy is added only when the engaged mode is actually MTP. MLA + MTP (GLM-5.2 / DeepSeek / Kimi) is unchanged, so the GLM-5.2 OOM fix is preserved; non-MLA and the draft-simple / draft-eagle3 paths no longer pay the copy. test_mtp_mla_target_ctx.py adds a case asserting the separate-drafter reserve collapses to the draft KV (no target copy) while the default MTP path keeps it. |
||
|---|---|---|
| .. | ||
| data_recipe | ||
| export | ||
| inference | ||
| rag | ||
| training | ||
| __init__.py | ||
| _torchao_stub.py | ||
| tool_healing.py | ||