Reviewers found a real CPU-only log shape that passed: an 'offloading 0 repeating layers to GPU' planning line, or a GPU KV/compute buffer, could read as GPU before the definitive 'offloaded 0/33' was seen. Check the explicit counted offload first (any N>0 wins, all zero is CPU-only), restrict the buffer-size signal to GPU model buffers (KV/compute on GPU with weights on CPU is still CPU inference), and add HIP/MUSA/CANN to the model-buffer markers so an older log naming those backends is not misread as CPU. Same in both classifiers. |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| src-tauri | ||
| __init__.py | ||
| install_llama_prebuilt.py | ||
| install_python_stack.py | ||
| LICENSE.AGPL-3.0 | ||
| package-lock.json | ||
| package.json | ||
| setup.bat | ||
| setup.ps1 | ||
| setup.sh | ||
| Unsloth_Studio_Colab.ipynb | ||