Refactor command building (deduplicate HF/local paths) and add flags for better performance: - --parallel 1: studio is single-user, so only 1 inference slot is needed. The previous auto-detect picked 4 slots, wasting VRAM on 3 unused KV caches. - --flash-attn on: force flash attention for faster inference. Default is "auto" which may not always enable it. - --fit on: auto-adjust parameters to fit in available device memory. Already the default but now explicit. Also cleaned up the duplicated command building for HF vs local mode into a single block. |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| __init__.py | ||
| install_python_stack.py | ||
| LICENSE.AGPL-3.0 | ||
| setup.bat | ||
| setup.ps1 | ||
| setup.sh | ||
| Unsloth_Studio_Colab.ipynb | ||