* Vulkan GPUs: real device names and selectable ordinals
Rebases the durable half of #7356 onto the inference_gpu transport #7476
landed on main. Those two PRs solve an overlapping problem and disagree on
the data model, so merging #7356 as-is would ship two parallel Vulkan
device concepts with different index semantics. This keeps main's transport
and adds what #7356 had that #7476 does not.
- _vulkan_probe.py emits a 5th column, ggml's device description, sanitized
for the tab protocol and UTF-8 safe. Reader tolerates 4- or 5-column
output so an older probe still parses.
- llama_cpp gains _run_vulkan_probe (shared parse) and
vulkan_device_inventory (names + is_igpu + real totals).
- get_vulkan_inference_gpu_info reports the real name and an explicit
is_igpu instead of "Vulkan<i>" and a total == 0 guess.
- index_kind becomes "vulkan", not "relative", and gpu_ids picks are
supported on Vulkan builds once the probe enumerated ordinals. The XPU ban
no longer applies to them: a Vulkan pick is a ggml ordinal, not a torch-xpu
index, so it works on an Intel host too.
- Frontend picker reads the Vulkan inventory as the pickable set.
Memory deliberately still comes from _get_gpu_memory, not the inventory.
That path applies _apply_igpu_host_reserve_mib and zeroes a shared total;
budgeting an APU off its raw shared total would hand out the whole machine's
RAM with no OS headroom. Identity is joined onto it by ordinal, so a probe
failure degrades to Vulkan<i> names with the memory readings intact.
Dropped from #7356 as superseded: validate_vulkan_gpu_ids (main's
resolve_requested_gpu_ids already rejects duplicates and
_resolve_gguf_gpu_ids_for_request already probes for existence), the
gguf_devices transport, and the iGPU budget fallback in 71619891e, which
main's aggregateGpuMemoryTotalGb handles better by counting a shared pool
once.
Also keeps #7356's removal of the late diffusion raise, so the graceful
gpu_ids drop stays reachable for a GGUF only classified as diffusion after
download. #7415's real guard, _reject_vulkan_diffusion_gpu_ids_before_
teardown, is untouched.
Verified on Windows + Strix Halo: backend Vulkan/GPU-selection suites at the
same 4 pre-existing failures as main, tests/studio 1671 passed with no new
failures, frontend typecheck clean. Hardware confirmation of the underlying
behavior is on #7356 from @Bebiv24 (RX 9070 XT + RX 480).
Co-authored-by: LeoBorcherding <borchborchmail@gmail.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: LeoBorcherding <borchborchmail@gmail.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
This commit is contained in:
parent
da447d47ba
commit
74295d93d8
8 changed files with 297 additions and 70 deletions
|
|
@ -1706,7 +1706,7 @@ def get_visible_gpu_utilization() -> Dict[str, Any]:
|
|||
"backend": _backend_label(device),
|
||||
"parent_visible_gpu_ids": [],
|
||||
"devices": [],
|
||||
"index_kind": "relative",
|
||||
"index_kind": "vulkan",
|
||||
}
|
||||
|
||||
|
||||
|
|
@ -2600,22 +2600,35 @@ def get_vulkan_inference_gpu_info() -> Optional[Dict[str, Any]]:
|
|||
"backend_cuda_visible_devices": None,
|
||||
"parent_visible_gpu_ids": [],
|
||||
"devices": [],
|
||||
"index_kind": "relative",
|
||||
"index_kind": "vulkan",
|
||||
}
|
||||
# Identity (real device description, explicit iGPU flag) comes from the
|
||||
# inventory; the memory numbers stay on _get_gpu_memory, which applies the
|
||||
# iGPU host reserve and zeroes a shared total. Budgeting an APU off the raw
|
||||
# shared total instead would hand out the whole machine's RAM with no OS
|
||||
# headroom. Join by ordinal; a probe failure just leaves names unresolved.
|
||||
identity: Dict[int, Dict[str, Any]] = {}
|
||||
try:
|
||||
identity = {row["index"]: row for row in LlamaCppBackend.vulkan_device_inventory()}
|
||||
except Exception as e:
|
||||
logger.debug("Vulkan device inventory failed, falling back to ordinals: %s", e)
|
||||
|
||||
try:
|
||||
for ordinal, free_mib, total_mib in LlamaCppBackend._get_gpu_memory():
|
||||
# Integrated Vulkan GPUs report total=0 because their memory is
|
||||
# shared. Publish the capped free value as their usable inference
|
||||
# budget and mark it so clients do not add system RAM again.
|
||||
shared_memory = total_mib == 0
|
||||
info = identity.get(ordinal, {})
|
||||
# _get_gpu_memory reports total 0 for a shared pool; prefer the
|
||||
# explicit flag when the inventory resolved this ordinal.
|
||||
shared_memory = bool(info["is_igpu"]) if "is_igpu" in info else total_mib == 0
|
||||
budget_mib = total_mib or free_mib
|
||||
used_mib = max(0, total_mib - free_mib) if total_mib else None
|
||||
result["devices"].append(
|
||||
{
|
||||
"index": ordinal,
|
||||
"index_kind": "relative",
|
||||
# ggml Vulkan ordinals are the space `--device Vulkan<i>` pins,
|
||||
# so unlike a torch-xpu relative ordinal these are selectable.
|
||||
"index_kind": "vulkan",
|
||||
"visible_ordinal": ordinal,
|
||||
"name": f"Vulkan{ordinal}",
|
||||
"name": info.get("name") or f"Vulkan{ordinal}",
|
||||
"memory_total_gb": round(budget_mib / 1024, 2),
|
||||
"vram_used_gb": round(used_mib / 1024, 2) if used_mib is not None else None,
|
||||
"vram_free_gb": round(free_mib / 1024, 2),
|
||||
|
|
@ -2727,7 +2740,7 @@ def get_backend_visible_gpu_info() -> Dict[str, Any]:
|
|||
"backend_cuda_visible_devices": os.environ.get("CUDA_VISIBLE_DEVICES"),
|
||||
"parent_visible_gpu_ids": [],
|
||||
"devices": [],
|
||||
"index_kind": "relative",
|
||||
"index_kind": "vulkan",
|
||||
}
|
||||
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue