* fix(rocm): prepend system ROCm libs on native Linux to avoid bundled HIP crash Prebuilt llama.cpp bundles ship their own ROCR/HIP runtime which can be incompatible with the host's amdkfd kernel driver, causing hsa_init() to crash or report zero devices. The llama-server then silently falls back to CPU while the UI reports GPU. The existing workaround (_wsl_system_rocm_lib_dirs) that prepends /opt/rocm/lib to LD_LIBRARY_PATH was gated on WSL (/dev/dxg) only, leaving native Linux AMD hosts unprotected. This commit adds _native_linux_system_rocm_lib_dirs(), a parallel helper gated on: - Linux platform (not WSL) - /dev/kfd present (bare-metal AMD compute) - Bundle contains bundled HIP libs (libggml-hip.so) - System has libhsa-runtime64.so(.1) It is called from both _llama_server_env_for_binary (serve-time) and binary_env (install-time validation), directly after the WSL block in both paths. Fixes #7208 Fixes #7208 * Add UNSLOTH_LLAMA_NO_SYSTEM_ROCM opt-out to native-Linux system ROCm preference for PR #7233 Lets a host where the bundled runtime works but system ROCm is mismatched keep the bundle. Mirrored in llama_cpp.py and install_llama_prebuilt.py. * Prefer env-configured ROCm root over /opt/rocm fallback for PR #7233 Put HIP_PATH/HIP_PATH_57/ROCM_PATH-derived roots before /opt/rocm so a stale /opt/rocm can't shadow the driver-matching install the env vars point at. Mirrored in llama_cpp.py and install_llama_prebuilt.py. * Match versioned libggml-hip.so via glob so the native-Linux ROCm fix fires for PR #7233 * Clarify native-Linux ROCm prepend uses the consistent system stack for PR #7233 * llama_cpp: tighten native-Linux ROCm prepend comments (no code change) --------- Co-authored-by: Daniel Han <danielhanchen@gmail.com> |
||
|---|---|---|
| .. | ||
| data_recipe | ||
| export | ||
| inference | ||
| rag | ||
| training | ||
| __init__.py | ||
| _torchao_stub.py | ||
| import_guards.py | ||
| tool_healing.py | ||