* Route CPU-only Linux x86_64 to ggml-org/llama.cpp prebuilts
setup.sh hard-coded _HELPER_RELEASE_REPO=unslothai/llama.cpp for every
non-Darwin host. unslothai/llama.cpp only publishes Linux CUDA bundles
(app-*-linux-x64-cuda*.tar.gz), so a CPU-only Linux host walked ~30
releases looking for a non-existent app-*-linux-x64-cpu asset, exited
the prebuilt planner with "no compatible Linux prebuilt asset was
found", and fell through to a source build. Free CI runners
(ubuntu-latest with no GPU) hit this on every install, and anyone
running Studio on a Linux laptop without an NVIDIA GPU paid the
~3 minute cmake+make cost on first install.
ggml-org publishes llama-<tag>-bin-ubuntu-x64.tar.gz on every release
and install_llama_prebuilt.py already knows how to fetch it: when
called with --published-repo ggml-org/llama.cpp, the Linux x86_64 +
not has_usable_nvidia branch in direct_upstream_release_plan picks up
that asset directly. The fix is purely on the routing side.
Tighten the gate so a Linux host routes to ggml-org only when it is
x86_64 and has no GPU detection tool installed (nvidia-smi, rocminfo,
amd-smi, hipconfig, hipinfo). Everything else stays on the current
path:
- macOS: already on ggml-org, unchanged
- Windows: already on ggml-org via setup.ps1, unchanged
- Linux CUDA: nvidia-smi present -> unslothai/llama.cpp, unchanged
- Linux ROCm: rocminfo / amd-smi / hipconfig / hipinfo present
-> unslothai/llama.cpp -> source build with HIP,
unchanged
- Linux Intel / Vulkan / SYCL: no NVIDIA / AMD tools, hits the new
ggml-org route, gets upstream CPU asset (same as
today's source-build CPU output, ~3 min faster)
- Linux arm64 / s390x: not x86_64 -> unslothai/llama.cpp ->
source build, unchanged
* Tighten routing comment in studio/setup.sh