* Studio (Windows): keep prompt caching on full GPU offload (#5692 follow-up) The #5692 full-offload tuning also added --no-cache-prompt, which disables in-VRAM prompt-prefix reuse. That is unrelated to the host-RAM KV checkpoints #5692 fixed (--cache-ram 0 / --ctx-checkpoints 0): a fully offloaded model keeps its KV cache in VRAM, so reusing a common prefix does not copy to system RAM and does not cause the PCI-E overhead. --no-cache-prompt only forces every request to re-prefill the whole prompt, which is small for short chats but severe for large stable system prompts reused across calls (coding agents, long multi-turn chats). Remove --no-cache-prompt; keep the checkpoint disables and the thread/OMP tuning. _prompt_cache_disabled stays False (its default), so slot save/restore is intact. Verified on a fully offloaded gemma GGUF: an identical repeated prompt reprefills 1 token instead of 2220. * Guard against re-adding --no-cache-prompt to any llama-server command Add a backend-wide test that AST-scans studio/backend and fails if --no-cache-prompt is appended/extended/+= into a command. This locks in the #7260 fix across every code path, not just load_model. Detecting the flag or honouring a user-supplied one stays allowed. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| src-tauri | ||
| __init__.py | ||
| install_llama_prebuilt.py | ||
| install_node_prebuilt.py | ||
| install_python_stack.py | ||
| LICENSE.AGPL-3.0 | ||
| MCP.md | ||
| node_prebuilt_pins.json | ||
| package-lock.json | ||
| package.json | ||
| setup.bat | ||
| setup.ps1 | ||
| setup.sh | ||
| Unsloth_Studio_Colab.ipynb | ||