Exclude host-pinned buffers (CUDA_Host etc.) from the GPU buffer-size signal so a binary that pins host memory but loads weights on CPU is not misread as GPU offload. Scan every 'offloaded N/M layers to GPU' line and accept if any N>0 so a speculative draft model logging 0/k before the main model's 33/33 is not flagged CPU-only. Same fix in the installer and runtime classifiers. |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| src-tauri | ||
| __init__.py | ||
| install_llama_prebuilt.py | ||
| install_python_stack.py | ||
| LICENSE.AGPL-3.0 | ||
| package-lock.json | ||
| package.json | ||
| setup.bat | ||
| setup.ps1 | ||
| setup.sh | ||
| Unsloth_Studio_Colab.ipynb | ||