studio: show system-wide VRAM in the multi-GPU System tab view on ROCm (#7216)
* studio: show system-wide VRAM in the multi-GPU System tab view on ROCm The System tab's per-GPU list comes from get_visible_gpu_utilization. When amd-smi is unavailable (always on Windows, minimal Linux installs) it fell back to torch, whose readings are process-local: on Windows WDDM hands each process its own budget, so a model held by the separate llama-server process read as ~0 VRAM used even with the GPU full (#7072). The primary-GPU endpoint already compensates with system-wide sources -- Windows Performance Counters (Task Manager's source) and Linux DRM sysfs -- but the multi-device endpoint never got those fallbacks. Add per-GPU variants of both sources and overlay them onto the torch fallback: _rocm_windows_perf_counter_vram_per_adapter_gb() attributes Dedicated Usage per physical adapter (phys_<N> in the counter instance name), and _rocm_linux_sysfs_vram_per_card_gb() reads mem_info_vram_{used,total} per DRM card. _overlay_system_wide_vram() applies them to the device list, ROCm-only, best-effort: unmatched adapters and ambiguous card counts keep the torch figures, and NVIDIA paths are untouched. Fixes #7072 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * studio: match VRAM overlay sources by device, honor unified memory, unblock the loop Five review fixes on the multi-GPU system-wide VRAM overlay: 1. Linux: match DRM cards to devices by PHYSICAL index instead of a positional zip, so a reordering visibility mask (HIP_VISIBLE_DEVICES=1,0) no longer swaps each card's figures onto the other GPU (which would mislead auto_select_gpu_ids and the coexistence checks). An index with no matching card keeps its torch figures. 2. Linux: skip the overlay for a device whose sysfs total is below torch's -- on unified-memory APUs (Strix Halo) mem_info_vram_total is only the small dedicated slice while torch sees the GTT-backed pool, and _apply_unified_memory_correction already defines larger-total-wins. 3. Windows: group counter instances by adapter LUID, not the phys_<N> suffix -- separate adapters each read phys_0, which collapsed every GPU into key 0. LUIDs are mapped to 0-based positions by ascending value as the closest stand-in for device order. 4. Windows: pair the system-wide usage with the physical capacity from get_device_properties (as the primary-GPU fallback does) -- under WDDM mem_get_info's "total" is the process budget, which misreported capacity and pushed utilization to 100%. 5. Run get_visible_gpu_utilization off the event loop in the /hardware/visible route (asyncio.to_thread, the repo's convention): the ROCm fallbacks can shell out to PowerShell with a 5s timeout, which would stall every other request while the System view polls. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * studio: skip the system-wide VRAM overlay for relative GPU indices The overlay matches its per-GPU sources (Windows perf counters, Linux sysfs) by physical device index, but under a UUID/MIG visibility mask the torch fallback enumerates ordinals and reports index_kind == "relative", where `index` is a visible ordinal, not a physical id. Applying the overlay there let card/adapter 0's system-wide VRAM overwrite the torch reading of a process that actually exposes physical GPU 1, misleading auto_select_gpu_ids and the coexistence checks. Gate the overlay on index_kind == "physical"; relative-index paths keep the torch fallback. * studio: drop the unreliable Windows VRAM overlay, keep the Linux one The multi-GPU system-wide VRAM overlay is now Linux-only. The Windows per-adapter Performance Counter path could not be made correct: the wildcard Get-Counter query also returns non-ROCm/iGPU adapters and LUID order is not the ROCm device order, so an adapter's usage could be overlaid onto the wrong GPU; and it read only Dedicated Usage, missing WDDM shared memory on unified-memory GPUs (Strix Halo), overstating free VRAM. Rather than misattribute VRAM and skew placement decisions, Windows keeps the process-local torch fallback (no regression vs before this PR); Linux DRM sysfs -- matched by physical index -- still fixes #7072 for the reporter's native-Linux ROCm case. Removes _rocm_windows_perf_counter_vram_per_adapter_gb and _torch_props_total_gb. * studio: key sysfs VRAM by DRM card number so filtering can't renumber cards _rocm_linux_sysfs_vram_per_card_gb dropped cards with a zero total or unreadable files and then the overlay enumerated the compacted list, so if card0 was dropped, card1's usage was assigned to physical GPU index 0 (equal-capacity GPUs slip past the unified-memory total guard). Return {card_number: (used, total)} and match a device to its card number directly: a hole stays a hole -- device 0 keeps its torch figures when card0 is absent, and card1 maps to device 1. * studio: key system-wide VRAM by ROCm ordinal, not raw DRM card number When a non-amdgpu adapter (Intel iGPU, a display-only card) owns an earlier DRM slot, DRM card numbers stop equalling ROCm device ordinals -- Intel card0 plus AMD card1/card2 gives ROCm devices 0/1, so keying the sysfs overlay by card number handed ROCm device 1 card1's data (AMD device 0) and left device 0 on stale torch figures, corrupting free-VRAM placement on equal-capacity GPUs. Only amdgpu cards expose mem_info_vram_*, so the glob already excludes foreign adapters; order the surviving cards by their PCI address (ROCm/HIP's default device order, read from each card's device symlink) and key by that position -- the ROCm physical ordinal, which is what the overlay matches against dev index. An unreadable / zero-total amdgpu card still consumes its ordinal so a later card is never renumbered onto its slot. * studio: skip the VRAM overlay under layered HIP-over-ROCR masks ROCR_VISIBLE_DEVICES filters physical GPUs at the HSA/ROCr layer, and a HIP_VISIBLE_DEVICES set on top selects WITHIN that already-filtered set (apply_gpu_ids sets HIP while leaving an inherited ROCR mask in place). When both are active _get_parent_visible_gpu_spec() prefers the HIP value, so the reported device index is a ROCR-relative ordinal, not a physical GPU id -- overlaying DRM-sysfs figures by that index would pull another GPU's usage (e.g. ROCR=2,3 + HIP=1 is physical GPU 3, but the overlay would read card 1), and equal-capacity cards bypass the total-size safeguard. Detect layered masks and keep torch's process-local figures there rather than risk misattribution; a single mask still leaves the index physical and is overlaid as before. * studio: only overlay whole-card VRAM onto 1:1 ROCm devices The overlay guard only skipped the case where sysfs total < torch total (unified-memory APUs), so a partitioned ROCm device (MI300 in CPX mode) -- where HIP exposes several logical devices per physical card but sysfs reports the whole card's aggregate -- passed the guard: the card total exceeds a partition's torch total, and the overlay overwrote the partition with whole-card usage and capacity, letting downstream selection think a partition had the entire card free. Require the sysfs card total to match the torch device total (within ~10%) so a mismatch in either direction -- unified memory (sysfs smaller) or partitioning (sysfs larger) -- keeps torch's figures. * studio: treat CUDA-over-ROCR as layered, enumerate AMD cards by driver Two remaining mismatches between the reported device index and the DRM card the overlay reads: - On ROCm the HIP layer honors CUDA_VISIBLE_DEVICES as well as HIP_VISIBLE_DEVICES, so a CUDA mask composed over ROCR layers identically: ROCR=2,3 with CUDA=1 is physical GPU 3, yet the spec reports the ROCR value [2,3] and the device was labeled index 2, overlaying card 2's usage onto GPU 3. The layered check now treats ROCR combined with either HIP or CUDA as layered. - The ROCm device set is now enumerated by bound driver (device/driver resolves to amdgpu) instead of by the presence of mem_info_vram_*. An AMD device with incomplete sysfs support (some APUs expose no VRAM files at all) was omitted by the glob entirely and shifted every later card down one ordinal, letting a similar-capacity GPU pass the total guard with another device's usage. Such a card now consumes its ordinal and simply yields no entry. * studio: honor GPU_DEVICE_ORDINAL and require an unambiguous card mapping Two remaining ways the reported device index could be matched to the wrong DRM card: - GPU_DEVICE_ORDINAL is a supported ROCm visibility variable that _get_parent_visible_gpu_spec() never consults, so GPU_DEVICE_ORDINAL=1 surfaces physical GPU 1 as torch ordinal 0 and it was mislabeled index 0, overlaying card 0's usage onto GPU 1. The mask check now covers it, and is renamed _rocm_device_index_unreliable() to say what it actually decides. - driver == amdgpu is only a SUPERSET of the ROCm-visible set: an amdgpu-bound adapter HIP cannot enumerate (an unsupported older AMD GPU beside a supported one) still took an ordinal and shifted every real compute device. There is no torch-side PCI identity to match against, so the overlay now requires the amdgpu card count to equal the device count -- exactly the condition under which position-in-PCI-order is a sound 1:1 mapping. Any disagreement keeps torch's process-local figures: less informative, never misattributed. * studio: keep the VRAM overlay working for masked GPU subsets The card-count guard compared the amdgpu card list against the VISIBLE device list, so any visibility mask disabled the overlay outright: HIP_VISIBLE_DEVICES=1,3 on a four-GPU host gives two devices against four cards. Those masked GPUs then kept reporting process-local torch usage, hiding VRAM held by llama-server and letting the training/chat placement checks overestimate free memory -- the exact problem the overlay exists to fix. The count check now applies only when no visibility mask is active, which is the case where the reported devices really are the whole host and a mismatch means an amdgpu adapter ROCm cannot enumerate is shifting the ordinals. Under a mask the subset is expected, so each device's physical index is validated individually instead: the per-card lookup bounds-checks it and the total-size guard rejects a card whose capacity does not match the device's. * studio: match GPUs to DRM cards by PCI identity, not by position Every mapping bug on this PR came from the same root cause: there was no authoritative link between a reported device index and a DRM card, so the overlay kept inferring one positionally and each heuristic broke on a new host shape -- foreign adapters on earlier DRM slots, cards with no VRAM sysfs, and most recently amdgpu-bound adapters HIP cannot enumerate, which the count guard could only catch on an unmasked host and therefore missed under any mask. Use the link ROCm itself enumerates from. KFD topology (/sys/class/kfd/kfd/topology/nodes/<N>/properties) lists exactly the GPUs HIP exposes -- GPU nodes in node-id order are HIP's device order -- and each carries its PCI location, so index N there IS physical device N with a stable identity. DRM sysfs now supplies system-wide VRAM keyed by that same PCI address, and the overlay is a join on it. Every previous skew becomes a failed join rather than a misattribution: an unenumerable adapter has no KFD node so it never takes an ordinal, a foreign adapter contributes no entry, and a masked subset resolves each physical index directly. That removes the count heuristic and its mask exception entirely. With no KFD topology there is no identity to join on, so the overlay is skipped rather than guessing positionally. * studio: require verified host visibility and AMD-only KFD nodes Three ways the identity map could still be built on a false premise: - The NVIDIA open kernel module registers KFD topology nodes with a positive SIMD count, so an earlier NVIDIA node shifted every AMD ordinal and ROCm device 1 resolved to AMD GPU 0. GPU nodes now require vendor_id 4098 (0x1002), the same filter install.sh already applies for this exact reason. - A GPU node with an unreadable properties file or no location_id was skipped, which silently shifted every later ordinal. Both now fail the whole map closed, so the overlay is disabled rather than misattributing. - A container exposing only some render devices through device cgroups sets no visibility variable, yet torch compacts what it can see to ordinals from zero while the host-mounted KFD and DRM trees still list every GPU. Nothing in the reported payload distinguishes that from a full host, and torch exposes no PCI id to check against, so the overlay now runs only when host visibility is positively verified: no visibility mask AND device count equal to the host GPU count. That also subsumes the previous layered-mask and GPU_DEVICE_ORDINAL checks, so _rocm_device_index_unreliable() is gone. This trades coverage for correctness: masked subsets and filtered containers now keep torch's process-local figures instead of a mapping that cannot be verified. * Fix the multi-GPU VRAM overlay docstring for PR #7216 The docstring claimed a reordering mask keeps each card on the right GPU, but the overlay skips any active visibility mask and keeps torch's figures. State the actual gating instead. * Tighten comments in the multi-GPU VRAM overlay and its tests Collapse the verbose docstrings and inline explanations added for the Linux ROCm system-wide VRAM overlay to succinct one-liners, keeping the non-obvious rationale (fail-closed KFD mapping, PCI-identity join, mask gating, the 10% whole-card guard). Comments only, no behavior change. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: danielhanchen <michaelhan2050@gmail.com>
This commit is contained in:
parent
f2f41bf9b1
commit
55433bd7b8
3 changed files with 767 additions and 1 deletions
|
|
@ -109,7 +109,9 @@ async def get_hardware_utilization(current_subject: str = Depends(get_current_su
|
|||
@router.get("/hardware/visible")
|
||||
async def get_visible_hardware_utilization(current_subject: str = Depends(get_current_subject)):
|
||||
from utils.hardware import get_visible_gpu_utilization
|
||||
return get_visible_gpu_utilization()
|
||||
|
||||
# Off the event loop: the ROCm fallbacks shell out (Windows perf counters, sysfs) and the System view polls this route.
|
||||
return await asyncio.to_thread(get_visible_gpu_utilization)
|
||||
|
||||
|
||||
@router.post("/start")
|
||||
|
|
|
|||
554
studio/backend/tests/test_rocm_multi_gpu_vram_system_wide.py
Normal file
554
studio/backend/tests/test_rocm_multi_gpu_vram_system_wide.py
Normal file
|
|
@ -0,0 +1,554 @@
|
|||
# SPDX-License-Identifier: AGPL-3.0-only
|
||||
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
||||
|
||||
"""The System tab's multi-GPU view must show system-wide VRAM on ROCm (#7072).
|
||||
|
||||
When amd-smi is unavailable, get_visible_gpu_utilization fell back to torch,
|
||||
whose readings are process-local: a model held by the separate llama-server
|
||||
process read as ~0 VRAM used even with the GPU full. These tests cover the
|
||||
per-GPU system-wide overlay the multi-device endpoint now applies, matched by
|
||||
physical device identity.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import importlib
|
||||
import sys
|
||||
import types
|
||||
from pathlib import Path
|
||||
|
||||
_BACKEND_DIR = Path(__file__).resolve().parent.parent
|
||||
if str(_BACKEND_DIR) not in sys.path:
|
||||
sys.path.insert(0, str(_BACKEND_DIR))
|
||||
|
||||
|
||||
def _maybe_stub(name: str, builder):
|
||||
# Stub only if the real module is missing, so we never shadow it for later tests.
|
||||
try:
|
||||
importlib.import_module(name)
|
||||
except ImportError:
|
||||
sys.modules[name] = builder()
|
||||
|
||||
|
||||
def _build_loggers_stub():
|
||||
m = types.ModuleType("loggers")
|
||||
m.get_logger = lambda name: __import__("logging").getLogger(name)
|
||||
return m
|
||||
|
||||
|
||||
def _build_structlog_stub():
|
||||
m = types.ModuleType("structlog")
|
||||
m.get_logger = lambda *a, **k: __import__("logging").getLogger("stub")
|
||||
return m
|
||||
|
||||
|
||||
_maybe_stub("loggers", _build_loggers_stub)
|
||||
_maybe_stub("structlog", _build_structlog_stub)
|
||||
|
||||
import utils.hardware.hardware as hw # noqa: E402
|
||||
|
||||
|
||||
def _device(
|
||||
index,
|
||||
used,
|
||||
total,
|
||||
*,
|
||||
ordinal = None,
|
||||
):
|
||||
return {
|
||||
"index": index,
|
||||
"index_kind": "physical",
|
||||
"visible_ordinal": index if ordinal is None else ordinal,
|
||||
"gpu_utilization_pct": None,
|
||||
"temperature_c": None,
|
||||
"vram_used_gb": used,
|
||||
"vram_total_gb": total,
|
||||
"vram_utilization_pct": round((used / total) * 100, 1) if total > 0 else None,
|
||||
"power_draw_w": None,
|
||||
"power_limit_w": None,
|
||||
"power_utilization_pct": None,
|
||||
}
|
||||
|
||||
|
||||
# ── Linux per-card sysfs ──
|
||||
|
||||
|
||||
def _fake_drm(tmp_path, monkeypatch, cards):
|
||||
"""Fake /sys/class/drm tree; glob returns cards REVERSED so the PCI sort must order them.
|
||||
|
||||
``cards``: (card_no, pci_bdf, driver, vram) tuples; vram is (used_gb, total_gb)
|
||||
or None for a device with no mem_info_vram_* files.
|
||||
"""
|
||||
drivers = tmp_path / "drivers"
|
||||
card_paths = []
|
||||
for card_no, bdf, driver, vram in cards:
|
||||
pci_dir = tmp_path / "pci" / bdf
|
||||
pci_dir.mkdir(parents = True, exist_ok = True)
|
||||
drv_dir = drivers / driver
|
||||
drv_dir.mkdir(parents = True, exist_ok = True)
|
||||
(pci_dir / "driver").symlink_to(drv_dir)
|
||||
if vram is not None:
|
||||
used, total = vram
|
||||
(pci_dir / "mem_info_vram_used").write_text(str(int(used * 1024**3)))
|
||||
(pci_dir / "mem_info_vram_total").write_text(str(int(total * 1024**3)))
|
||||
card_dir = tmp_path / "drm" / f"card{card_no}"
|
||||
card_dir.mkdir(parents = True, exist_ok = True)
|
||||
(card_dir / "device").symlink_to(pci_dir)
|
||||
card_paths.append(str(card_dir))
|
||||
monkeypatch.setattr(hw.glob, "glob", lambda pattern: list(reversed(card_paths)))
|
||||
return card_paths
|
||||
|
||||
|
||||
def test_linux_vram_keyed_by_pci_excludes_foreign_adapters(monkeypatch, tmp_path):
|
||||
# Foreign (non-amdgpu) adapters contribute no entry, so they cannot shift ordinals.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
_fake_drm(
|
||||
tmp_path,
|
||||
monkeypatch,
|
||||
[
|
||||
(0, "0000:00:02.0", "i915", (0.5, 2.0)), # foreign adapter: excluded
|
||||
(1, "0000:03:00.0", "amdgpu", (40, 48)), # AMD device 0
|
||||
(2, "0000:41:00.0", "amdgpu", (1, 8)), # AMD device 1
|
||||
],
|
||||
)
|
||||
assert hw._rocm_linux_sysfs_vram_by_pci_gb() == {
|
||||
"0000:03:00.0": (40.0, 48.0),
|
||||
"0000:41:00.0": (1.0, 8.0),
|
||||
}
|
||||
|
||||
|
||||
def test_linux_vram_omits_bad_cards_without_shifting(monkeypatch, tmp_path):
|
||||
# A zero-total card has no entry; identity keying means its absence renumbers nothing.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
_fake_drm(
|
||||
tmp_path,
|
||||
monkeypatch,
|
||||
[
|
||||
(0, "0000:03:00.0", "amdgpu", (0, 0)), # zero total -> no entry
|
||||
(1, "0000:41:00.0", "amdgpu", (2, 16)),
|
||||
],
|
||||
)
|
||||
assert hw._rocm_linux_sysfs_vram_by_pci_gb() == {"0000:41:00.0": (2.0, 16.0)}
|
||||
|
||||
|
||||
def test_linux_vram_omits_amd_card_without_vram_files(monkeypatch, tmp_path):
|
||||
# An APU with no mem_info_vram_* files has no entry; the discrete card keeps its address.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
_fake_drm(
|
||||
tmp_path,
|
||||
monkeypatch,
|
||||
[
|
||||
(0, "0000:03:00.0", "amdgpu", None), # APU: no VRAM sysfs files
|
||||
(1, "0000:41:00.0", "amdgpu", (2, 16)),
|
||||
],
|
||||
)
|
||||
assert hw._rocm_linux_sysfs_vram_by_pci_gb() == {"0000:41:00.0": (2.0, 16.0)}
|
||||
|
||||
|
||||
# ── KFD topology: the authoritative ROCm device order ──
|
||||
|
||||
|
||||
_AMD = 4098 # 0x1002
|
||||
_NVIDIA = 4318 # 0x10DE -- the open kernel module also registers KFD nodes
|
||||
|
||||
|
||||
def _fake_kfd(tmp_path, monkeypatch, nodes):
|
||||
"""Fake KFD topology nodes tree, returned out of node order so the sort must order it.
|
||||
|
||||
``nodes``: (node_id, simd_count, location_id, domain, vendor_id); simd_count 0
|
||||
marks a CPU node, location_id None omits the property.
|
||||
"""
|
||||
node_paths = []
|
||||
for node_id, simd_count, location_id, domain, vendor_id in nodes:
|
||||
d = tmp_path / "kfd" / str(node_id)
|
||||
d.mkdir(parents = True, exist_ok = True)
|
||||
lines = [f"cpu_cores_count {0 if simd_count else 8}", f"simd_count {simd_count}"]
|
||||
if location_id is not None:
|
||||
lines.append(f"location_id {location_id}")
|
||||
lines.append(f"domain {domain}")
|
||||
if vendor_id is not None:
|
||||
lines.append(f"vendor_id {vendor_id}")
|
||||
(d / "properties").write_text("\n".join(lines) + "\n")
|
||||
node_paths.append(str(d))
|
||||
monkeypatch.setattr(hw.glob, "glob", lambda pattern: list(reversed(node_paths)))
|
||||
return node_paths
|
||||
|
||||
|
||||
def test_kfd_lists_gpu_nodes_in_device_order(monkeypatch, tmp_path):
|
||||
# The CPU node (simd_count 0) takes no ordinal; GPU nodes in node-id order are HIP's order.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
_fake_kfd(
|
||||
tmp_path,
|
||||
monkeypatch,
|
||||
[
|
||||
(0, 0, None, 0, None), # CPU node
|
||||
(1, 304, (0x03 << 8) | (0x00 << 3) | 0, 0, _AMD), # 0000:03:00.0 -> dev 0
|
||||
(2, 304, (0x41 << 8) | (0x00 << 3) | 0, 0, _AMD), # 0000:41:00.0 -> dev 1
|
||||
],
|
||||
)
|
||||
assert hw._rocm_kfd_gpu_pci_ids() == ["0000:03:00.0", "0000:41:00.0"]
|
||||
|
||||
|
||||
def test_kfd_decodes_domain_device_and_function(monkeypatch, tmp_path):
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
_fake_kfd(tmp_path, monkeypatch, [(1, 64, (0xC1 << 8) | (0x1F << 3) | 5, 0x1234, _AMD)])
|
||||
assert hw._rocm_kfd_gpu_pci_ids() == ["1234:c1:1f.5"]
|
||||
|
||||
|
||||
def test_kfd_skips_non_amd_gpu_nodes(monkeypatch, tmp_path):
|
||||
# An NVIDIA KFD node is not a HIP device: it must take no ordinal, else it
|
||||
# shifts every AMD GPU and ROCm device 1 resolves to AMD GPU 0.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
_fake_kfd(
|
||||
tmp_path,
|
||||
monkeypatch,
|
||||
[
|
||||
(0, 0, None, 0, None), # CPU
|
||||
(1, 128, (0x01 << 8) | 0, 0, _NVIDIA), # NVIDIA: no ordinal
|
||||
(2, 304, (0x03 << 8) | 0, 0, _AMD), # AMD device 0
|
||||
(3, 304, (0x41 << 8) | 0, 0, _AMD), # AMD device 1
|
||||
],
|
||||
)
|
||||
assert hw._rocm_kfd_gpu_pci_ids() == ["0000:03:00.0", "0000:41:00.0"]
|
||||
|
||||
|
||||
def test_kfd_fails_closed_when_a_gpu_has_no_location(monkeypatch, tmp_path):
|
||||
# Dropping an unplaceable AMD GPU shifts later ordinals; fail closed for the whole map.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
_fake_kfd(
|
||||
tmp_path,
|
||||
monkeypatch,
|
||||
[
|
||||
(1, 304, None, 0, _AMD), # AMD GPU with no location_id
|
||||
(2, 304, (0x41 << 8) | 0, 0, _AMD),
|
||||
],
|
||||
)
|
||||
assert hw._rocm_kfd_gpu_pci_ids() == []
|
||||
|
||||
|
||||
def test_kfd_fails_closed_when_a_node_is_unreadable(monkeypatch, tmp_path):
|
||||
# An unreadable node could be a GPU; assuming otherwise would shift ordinals.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
paths = _fake_kfd(
|
||||
tmp_path,
|
||||
monkeypatch,
|
||||
[
|
||||
(1, 304, (0x03 << 8) | 0, 0, _AMD),
|
||||
(2, 304, (0x41 << 8) | 0, 0, _AMD),
|
||||
],
|
||||
)
|
||||
(Path(paths[0]) / "properties").unlink()
|
||||
assert hw._rocm_kfd_gpu_pci_ids() == []
|
||||
|
||||
|
||||
def test_kfd_absent_yields_no_device_order(monkeypatch):
|
||||
monkeypatch.setattr(hw.glob, "glob", lambda pattern: [])
|
||||
assert hw._rocm_kfd_gpu_pci_ids() == []
|
||||
|
||||
|
||||
# ── overlay ──
|
||||
|
||||
|
||||
def _patch_pci_map(monkeypatch, bdfs):
|
||||
"""Declare the ROCm device order by PCI address (index N is device N) and clear
|
||||
the visibility masks the overlay requires unset.
|
||||
"""
|
||||
for var in (
|
||||
"HIP_VISIBLE_DEVICES",
|
||||
"ROCR_VISIBLE_DEVICES",
|
||||
"CUDA_VISIBLE_DEVICES",
|
||||
"GPU_DEVICE_ORDINAL",
|
||||
):
|
||||
monkeypatch.delenv(var, raising = False)
|
||||
monkeypatch.setattr(hw, "_rocm_kfd_gpu_pci_ids", lambda: list(bdfs))
|
||||
|
||||
|
||||
def _pci(n):
|
||||
"""A distinct, well-formed PCI address for card n."""
|
||||
return f"0000:{n:02x}:00.0"
|
||||
|
||||
|
||||
def test_overlay_windows_is_noop_keeps_torch(monkeypatch):
|
||||
# Windows is intentionally not overlaid (perf counters can't map to ROCm ordinals): keep torch.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Windows")
|
||||
monkeypatch.setattr(
|
||||
hw,
|
||||
"_rocm_linux_sysfs_vram_by_pci_gb",
|
||||
lambda: (_ for _ in ()).throw(AssertionError("sysfs must not run on Windows")),
|
||||
)
|
||||
devices = [_device(0, used = 0.02, total = 8.0)]
|
||||
_patch_pci_map(monkeypatch, [_pci(0)])
|
||||
hw._overlay_system_wide_vram(devices)
|
||||
assert devices[0]["vram_used_gb"] == 0.02 # untouched
|
||||
|
||||
|
||||
def test_overlay_linux_matches_by_device_ordinal(monkeypatch):
|
||||
# Devices arriving as [index 1, index 0] each get their own GPU's figures by ordinal.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
monkeypatch.setattr(
|
||||
hw,
|
||||
"_rocm_linux_sysfs_vram_by_pci_gb",
|
||||
lambda: {_pci(0): (30.0, 45.0), _pci(1): (0.5, 8.0)}, # dev 0 big, dev 1 small
|
||||
)
|
||||
devices = [_device(1, used = 0.01, total = 8.0), _device(0, used = 0.02, total = 45.0)]
|
||||
_patch_pci_map(monkeypatch, [_pci(0), _pci(1)])
|
||||
hw._overlay_system_wide_vram(devices)
|
||||
assert devices[0]["vram_used_gb"] == 0.5 # index 1 -> device 1 (small)
|
||||
assert devices[0]["vram_total_gb"] == 8.0
|
||||
assert devices[1]["vram_used_gb"] == 30.0 # index 0 -> device 0 (big)
|
||||
assert devices[1]["vram_total_gb"] == 45.0
|
||||
|
||||
|
||||
def test_overlay_linux_ordinal_hole_does_not_shift(monkeypatch):
|
||||
# Device 0's card dropped: index 0 keeps torch, index 1 still maps to ordinal 1 (no compaction).
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
monkeypatch.setattr(hw, "_rocm_linux_sysfs_vram_by_pci_gb", lambda: {_pci(1): (0.5, 8.0)})
|
||||
devices = [_device(0, used = 0.02, total = 45.0), _device(1, used = 0.01, total = 8.0)]
|
||||
_patch_pci_map(monkeypatch, [_pci(0), _pci(1)])
|
||||
hw._overlay_system_wide_vram(devices)
|
||||
assert devices[0]["vram_used_gb"] == 0.02 # no ordinal 0 -> torch kept
|
||||
assert devices[1]["vram_used_gb"] == 0.5 # ordinal 1 -> device 1, not device 0
|
||||
|
||||
|
||||
def test_overlay_linux_skips_unified_memory_card(monkeypatch):
|
||||
# Unified-memory APU: the smaller sysfs total must not shrink torch's GTT-backed pool.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
monkeypatch.setattr(hw, "_rocm_linux_sysfs_vram_by_pci_gb", lambda: {_pci(0): (0.4, 1.0)})
|
||||
devices = [_device(0, used = 12.0, total = 96.0)] # torch's unified pool
|
||||
_patch_pci_map(monkeypatch, [_pci(0)])
|
||||
hw._overlay_system_wide_vram(devices)
|
||||
assert devices[0]["vram_used_gb"] == 12.0
|
||||
assert devices[0]["vram_total_gb"] == 96.0
|
||||
|
||||
|
||||
def test_overlay_linux_skips_partitioned_device(monkeypatch):
|
||||
# Partitioned MI300: the whole-card sysfs total dwarfs the partition, so the overlay must not overwrite it.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
monkeypatch.setattr(hw, "_rocm_linux_sysfs_vram_by_pci_gb", lambda: {_pci(0): (40.0, 192.0)})
|
||||
devices = [_device(0, used = 1.0, total = 24.0)] # torch partition
|
||||
_patch_pci_map(monkeypatch, [_pci(0)])
|
||||
hw._overlay_system_wide_vram(devices)
|
||||
assert devices[0]["vram_used_gb"] == 1.0 # partition figures kept
|
||||
assert devices[0]["vram_total_gb"] == 24.0
|
||||
|
||||
|
||||
def test_overlay_linux_out_of_range_index_untouched(monkeypatch):
|
||||
# A masked host exposing physical index 5 with no card 5: keep torch data.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
monkeypatch.setattr(
|
||||
hw, "_rocm_linux_sysfs_vram_by_pci_gb", lambda: {_pci(0): (30.0, 45.0), _pci(1): (0.5, 8.0)}
|
||||
)
|
||||
devices = [_device(5, used = 0.02, total = 45.0)]
|
||||
_patch_pci_map(monkeypatch, [_pci(0), _pci(1)])
|
||||
hw._overlay_system_wide_vram(devices)
|
||||
assert devices[0]["vram_used_gb"] == 0.02
|
||||
|
||||
|
||||
def test_overlay_ignores_adapters_rocm_cannot_enumerate(monkeypatch):
|
||||
# A HIP-unenumerable amdgpu adapter has no KFD node, so device 0 resolves to
|
||||
# the supported GPU's own address, never the display card's.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
monkeypatch.setattr(
|
||||
hw,
|
||||
"_rocm_linux_sysfs_vram_by_pci_gb",
|
||||
# Both in DRM sysfs with similar capacity -- what the total-size guard can't separate.
|
||||
lambda: {_pci(9): (30.0, 45.0), _pci(3): (12.0, 45.0)},
|
||||
)
|
||||
_patch_pci_map(monkeypatch, [_pci(3)]) # KFD lists only the supported GPU
|
||||
devices = [_device(0, used = 0.02, total = 45.0)] # torch sees that one GPU
|
||||
hw._overlay_system_wide_vram(devices)
|
||||
assert devices[0]["vram_used_gb"] == 12.0 # the supported GPU's own figures
|
||||
|
||||
|
||||
def test_overlay_skips_masked_subsets(monkeypatch):
|
||||
# Under a mask the index is not verifiably a host ordinal, so keep torch's figures.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
_patch_pci_map(monkeypatch, [_pci(0), _pci(1), _pci(2), _pci(3)])
|
||||
monkeypatch.setenv("HIP_VISIBLE_DEVICES", "1,3")
|
||||
monkeypatch.setattr(
|
||||
hw,
|
||||
"_rocm_linux_sysfs_vram_by_pci_gb",
|
||||
lambda: {_pci(1): (30.0, 48.0), _pci(3): (12.0, 48.0)},
|
||||
)
|
||||
devices = [_device(1, used = 0.02, total = 48.0), _device(3, used = 0.01, total = 48.0)]
|
||||
hw._overlay_system_wide_vram(devices)
|
||||
assert devices[0]["vram_used_gb"] == 0.02 # torch kept
|
||||
assert devices[1]["vram_used_gb"] == 0.01
|
||||
|
||||
|
||||
def test_overlay_skips_device_cgroup_filtered_container(monkeypatch):
|
||||
# A device-cgroup container sets no env var yet compacts torch's indices from
|
||||
# zero while KFD/DRM list every GPU, so the count mismatch must disable the overlay.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
_patch_pci_map(monkeypatch, [_pci(0), _pci(1), _pci(2), _pci(3)]) # host has 4
|
||||
monkeypatch.setattr(
|
||||
hw,
|
||||
"_rocm_linux_sysfs_vram_by_pci_gb",
|
||||
lambda: {_pci(0): (30.0, 48.0), _pci(2): (12.0, 48.0)},
|
||||
)
|
||||
devices = [_device(0, used = 0.02, total = 48.0)] # container sees 1, as index 0
|
||||
hw._overlay_system_wide_vram(devices)
|
||||
assert devices[0]["vram_used_gb"] == 0.02 # torch kept, not host GPU 0's 30.0
|
||||
|
||||
|
||||
def test_overlay_skips_without_kfd_topology(monkeypatch):
|
||||
# No KFD means no identity to join on; fall back to torch rather than guess.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
monkeypatch.setattr(hw, "_rocm_kfd_gpu_pci_ids", lambda: [])
|
||||
monkeypatch.setattr(
|
||||
hw,
|
||||
"_rocm_linux_sysfs_vram_by_pci_gb",
|
||||
lambda: (_ for _ in ()).throw(AssertionError("must not read sysfs without KFD")),
|
||||
)
|
||||
devices = [_device(0, used = 0.02, total = 45.0)]
|
||||
hw._overlay_system_wide_vram(devices)
|
||||
assert devices[0]["vram_used_gb"] == 0.02
|
||||
|
||||
|
||||
def test_overlay_empty_devices_is_noop(monkeypatch):
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
hw._overlay_system_wide_vram([]) # must not raise
|
||||
|
||||
|
||||
# ── integration: the ROCm torch fallback applies the overlay ──
|
||||
|
||||
|
||||
def test_visible_utilization_rocm_fallback_overlays(monkeypatch):
|
||||
for _var in (
|
||||
"HIP_VISIBLE_DEVICES",
|
||||
"ROCR_VISIBLE_DEVICES",
|
||||
"CUDA_VISIBLE_DEVICES",
|
||||
"GPU_DEVICE_ORDINAL",
|
||||
):
|
||||
monkeypatch.delenv(_var, raising = False)
|
||||
monkeypatch.setattr(hw, "IS_ROCM", True)
|
||||
monkeypatch.setattr(hw, "get_device", lambda: hw.DeviceType.CUDA)
|
||||
monkeypatch.setattr(hw, "_smi_query", lambda *a, **k: None) # amd-smi unavailable
|
||||
monkeypatch.setattr(
|
||||
hw,
|
||||
"_get_parent_visible_gpu_spec",
|
||||
lambda: {"raw": None, "numeric_ids": [0, 1], "supports_explicit_gpu_ids": True},
|
||||
)
|
||||
monkeypatch.setattr(hw, "get_parent_visible_gpu_ids", lambda: [0, 1])
|
||||
monkeypatch.setattr(
|
||||
hw,
|
||||
"_torch_get_per_device_info",
|
||||
lambda ids: [
|
||||
{"index": 0, "visible_ordinal": 0, "used_gb": 0.02, "total_gb": 45.0},
|
||||
{"index": 1, "visible_ordinal": 1, "used_gb": 0.01, "total_gb": 8.0},
|
||||
],
|
||||
)
|
||||
overlaid = []
|
||||
monkeypatch.setattr(
|
||||
hw, "_overlay_system_wide_vram", lambda devices: overlaid.append(len(devices))
|
||||
)
|
||||
result = hw.get_visible_gpu_utilization()
|
||||
assert result["available"] is True
|
||||
assert overlaid == [2]
|
||||
|
||||
|
||||
def test_visible_utilization_relative_index_skips_overlay(monkeypatch):
|
||||
# UUID/MIG mask gives relative indices; the overlay matches physical index, so it must not run.
|
||||
monkeypatch.setattr(hw, "IS_ROCM", True)
|
||||
monkeypatch.setattr(hw, "get_device", lambda: hw.DeviceType.CUDA)
|
||||
monkeypatch.setattr(hw, "_smi_query", lambda *a, **k: None)
|
||||
monkeypatch.setattr(
|
||||
hw,
|
||||
"_get_parent_visible_gpu_spec",
|
||||
lambda: {"raw": "GPU-uuid-a", "numeric_ids": None, "supports_explicit_gpu_ids": False},
|
||||
)
|
||||
monkeypatch.setattr(hw, "get_parent_visible_gpu_ids", lambda: []) # UUID mask
|
||||
monkeypatch.setattr(hw, "_torch_get_physical_gpu_count", lambda: 1)
|
||||
monkeypatch.setattr(
|
||||
hw,
|
||||
"_torch_get_per_device_info",
|
||||
lambda ids: [{"index": 0, "visible_ordinal": 0, "used_gb": 0.02, "total_gb": 8.0}],
|
||||
)
|
||||
called = []
|
||||
monkeypatch.setattr(hw, "_overlay_system_wide_vram", lambda devices: called.append(1))
|
||||
result = hw.get_visible_gpu_utilization()
|
||||
assert result["index_kind"] == "relative"
|
||||
assert called == []
|
||||
|
||||
|
||||
def test_visible_utilization_nvidia_fallback_skips_overlay(monkeypatch):
|
||||
monkeypatch.setattr(hw, "IS_ROCM", False)
|
||||
monkeypatch.setattr(hw, "get_device", lambda: hw.DeviceType.CUDA)
|
||||
monkeypatch.setattr(hw, "_smi_query", lambda *a, **k: None)
|
||||
monkeypatch.setattr(
|
||||
hw,
|
||||
"_get_parent_visible_gpu_spec",
|
||||
lambda: {"raw": None, "numeric_ids": [0], "supports_explicit_gpu_ids": True},
|
||||
)
|
||||
monkeypatch.setattr(hw, "get_parent_visible_gpu_ids", lambda: [0])
|
||||
monkeypatch.setattr(
|
||||
hw,
|
||||
"_torch_get_per_device_info",
|
||||
lambda ids: [{"index": 0, "visible_ordinal": 0, "used_gb": 1.0, "total_gb": 24.0}],
|
||||
)
|
||||
called = []
|
||||
monkeypatch.setattr(hw, "_overlay_system_wide_vram", lambda devices: called.append(1))
|
||||
result = hw.get_visible_gpu_utilization()
|
||||
assert result["available"] is True
|
||||
assert called == []
|
||||
|
||||
|
||||
def test_any_visibility_mask_is_detected(monkeypatch):
|
||||
# Any of these makes the index not a host-physical ordinal, so each must disable the overlay.
|
||||
for var in (
|
||||
"HIP_VISIBLE_DEVICES",
|
||||
"ROCR_VISIBLE_DEVICES",
|
||||
"CUDA_VISIBLE_DEVICES",
|
||||
"GPU_DEVICE_ORDINAL",
|
||||
):
|
||||
monkeypatch.delenv(var, raising = False)
|
||||
assert hw._rocm_visibility_mask_active() is False
|
||||
for var in (
|
||||
"HIP_VISIBLE_DEVICES",
|
||||
"ROCR_VISIBLE_DEVICES",
|
||||
"CUDA_VISIBLE_DEVICES",
|
||||
"GPU_DEVICE_ORDINAL",
|
||||
):
|
||||
monkeypatch.setenv(var, "1")
|
||||
assert hw._rocm_visibility_mask_active() is True, var
|
||||
monkeypatch.setenv(var, " ") # empty is not an active filter
|
||||
assert hw._rocm_visibility_mask_active() is False, var
|
||||
monkeypatch.delenv(var, raising = False)
|
||||
|
||||
|
||||
def test_overlay_skips_under_gpu_device_ordinal(monkeypatch):
|
||||
# GPU_DEVICE_ORDINAL=1 surfaces GPU 1 as torch ordinal 0, so index 0 is not GPU 0; overlay must not run.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
_patch_pci_map(monkeypatch, [_pci(0)])
|
||||
monkeypatch.setenv("GPU_DEVICE_ORDINAL", "1")
|
||||
monkeypatch.setattr(hw, "_rocm_linux_sysfs_vram_by_pci_gb", lambda: {_pci(0): (30.0, 45.0)})
|
||||
devices = [_device(0, used = 0.02, total = 45.0)]
|
||||
hw._overlay_system_wide_vram(devices)
|
||||
assert devices[0]["vram_used_gb"] == 0.02
|
||||
|
||||
|
||||
def test_visible_utilization_delegates_gating_to_the_overlay(monkeypatch):
|
||||
# The call site no longer pre-checks masks; the overlay gates itself, so a physical payload always reaches it.
|
||||
monkeypatch.setenv("ROCR_VISIBLE_DEVICES", "2,3")
|
||||
monkeypatch.setenv("HIP_VISIBLE_DEVICES", "1")
|
||||
monkeypatch.setattr(hw, "IS_ROCM", True)
|
||||
monkeypatch.setattr(hw, "get_device", lambda: hw.DeviceType.CUDA)
|
||||
monkeypatch.setattr(hw, "_smi_query", lambda *a, **k: None)
|
||||
monkeypatch.setattr(
|
||||
hw,
|
||||
"_get_parent_visible_gpu_spec",
|
||||
lambda: {"raw": "1", "numeric_ids": [1], "supports_explicit_gpu_ids": True},
|
||||
)
|
||||
monkeypatch.setattr(hw, "get_parent_visible_gpu_ids", lambda: [1])
|
||||
monkeypatch.setattr(
|
||||
hw,
|
||||
"_torch_get_per_device_info",
|
||||
lambda ids: [{"index": 1, "visible_ordinal": 0, "used_gb": 0.02, "total_gb": 8.0}],
|
||||
)
|
||||
# Real overlay + gating: the layered mask must leave torch's figures.
|
||||
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
|
||||
monkeypatch.setattr(hw, "_rocm_kfd_gpu_pci_ids", lambda: [_pci(0), _pci(1)])
|
||||
monkeypatch.setattr(hw, "_rocm_linux_sysfs_vram_by_pci_gb", lambda: {_pci(1): (30.0, 8.0)})
|
||||
result = hw.get_visible_gpu_utilization()
|
||||
assert result["index_kind"] == "physical"
|
||||
assert result["devices"][0]["vram_used_gb"] == 0.02 # untouched
|
||||
|
|
@ -734,6 +734,141 @@ def _rocm_linux_sysfs_vram_gb() -> tuple[Optional[float], Optional[float]]:
|
|||
return None, None
|
||||
|
||||
|
||||
# 0x1002. NVIDIA's open kernel module also registers KFD nodes (vendor_id 0x10DE);
|
||||
# a non-AMD node is not a HIP device and must never take an ordinal.
|
||||
_AMD_PCI_VENDOR_ID = 4098
|
||||
|
||||
|
||||
def _rocm_kfd_gpu_pci_ids() -> list[str]:
|
||||
"""PCI addresses of the GPUs ROCm enumerates, in HIP device order.
|
||||
|
||||
Reads /sys/class/kfd/kfd/topology/nodes/<N>/properties, the topology ROCm
|
||||
itself enumerates from: AMD GPU nodes (simd_count > 0 excludes CPUs,
|
||||
vendor_id == AMD excludes NVIDIA) in node-id order are HIP's device order, so
|
||||
position N is ROCm physical device N. Unlike DRM sysfs, an amdgpu adapter HIP
|
||||
cannot enumerate has no node here, so it never consumes an ordinal.
|
||||
|
||||
Returns [] (disabling the overlay) when KFD is absent, and FAILS CLOSED the
|
||||
same way on any unreadable node or an AMD node with no location_id: dropping
|
||||
one would shift every later ordinal and let a similar-capacity GPU pass the
|
||||
total-size guard while showing another card's usage.
|
||||
|
||||
location_id is the kernel's (bus << 8) | devfn; domain is separate.
|
||||
"""
|
||||
nodes: list[tuple[int, str]] = []
|
||||
try:
|
||||
node_dirs = glob.glob("/sys/class/kfd/kfd/topology/nodes/*")
|
||||
except Exception:
|
||||
return []
|
||||
for node_dir in node_dirs:
|
||||
m = re.fullmatch(r".*/(\d+)", node_dir)
|
||||
if m is None:
|
||||
continue
|
||||
props: dict[str, int] = {}
|
||||
try:
|
||||
with open(os.path.join(node_dir, "properties")) as f:
|
||||
for line in f:
|
||||
parts = line.split()
|
||||
if len(parts) == 2:
|
||||
try:
|
||||
props[parts[0]] = int(parts[1])
|
||||
except ValueError:
|
||||
continue
|
||||
except OSError:
|
||||
return [] # unreadable node could be a GPU: fail closed, don't shift
|
||||
if props.get("simd_count", 0) <= 0:
|
||||
continue # CPU node, not a GPU
|
||||
if props.get("vendor_id") != _AMD_PCI_VENDOR_ID:
|
||||
continue # non-AMD GPU node (NVIDIA open driver): not a HIP device
|
||||
location_id = props.get("location_id")
|
||||
if location_id is None:
|
||||
return [] # an AMD GPU we cannot place: fail closed for the whole map
|
||||
domain = props.get("domain", 0)
|
||||
bus = (location_id >> 8) & 0xFF
|
||||
devfn = location_id & 0xFF
|
||||
bdf = f"{domain:04x}:{bus:02x}:{(devfn >> 3) & 0x1F:02x}.{devfn & 0x7}"
|
||||
nodes.append((int(m.group(1)), bdf))
|
||||
nodes.sort(key = lambda n: n[0])
|
||||
return [bdf for _node_id, bdf in nodes]
|
||||
|
||||
|
||||
def _rocm_linux_amdgpu_cards() -> list[tuple[str, int, str]]:
|
||||
"""The amdgpu-bound DRM cards in PCI order: ``(pci_bdf, card_no, device_dir)``.
|
||||
|
||||
Membership is by the BOUND DRIVER, not the VRAM sysfs files: an AMD device
|
||||
with incomplete sysfs support (some APUs expose no mem_info_vram_*) still
|
||||
consumes a ROCm ordinal, and dropping it would shift every later card down.
|
||||
PCI order is HIP's default enumeration order, so list position is the ROCm
|
||||
ordinal; card_no is a stable tiebreak when the BDF cannot be resolved.
|
||||
|
||||
NOTE this is a superset of the ROCm-visible set (a HIP-unsupported amdgpu
|
||||
adapter appears too), so callers must check the counts agree before assuming
|
||||
a 1:1 mapping onto torch devices.
|
||||
"""
|
||||
if platform.system() != "Linux":
|
||||
return []
|
||||
amd_cards: list[tuple[str, int, str]] = []
|
||||
try:
|
||||
for card_path in glob.glob("/sys/class/drm/card*"):
|
||||
# Match card<N> exactly so connector nodes (card0-DP-1) are skipped.
|
||||
m = re.fullmatch(r".*/card(\d+)", card_path)
|
||||
if m is None:
|
||||
continue
|
||||
dev_dir = os.path.join(card_path, "device")
|
||||
try:
|
||||
driver = os.path.basename(os.path.realpath(os.path.join(dev_dir, "driver")))
|
||||
except OSError:
|
||||
continue
|
||||
if driver != "amdgpu":
|
||||
continue # foreign adapter: not a ROCm device, takes no ordinal
|
||||
try:
|
||||
bdf = os.path.basename(os.path.realpath(dev_dir))
|
||||
except OSError:
|
||||
bdf = ""
|
||||
amd_cards.append((bdf, int(m.group(1)), dev_dir))
|
||||
except Exception:
|
||||
return []
|
||||
amd_cards.sort(key = lambda c: (c[0], c[1]))
|
||||
return amd_cards
|
||||
|
||||
|
||||
def _rocm_linux_sysfs_vram_by_pci_gb() -> dict[str, tuple[float, float]]:
|
||||
"""System-wide AMD VRAM via Linux DRM sysfs, keyed by the card's PCI address.
|
||||
|
||||
Reads each card's mem_info_vram_{used,total} (kernel-updated across all
|
||||
processes) so every GPU gets its own figure, unlike _rocm_linux_sysfs_vram_gb
|
||||
which sums the host. Keyed by PCI address, not an ordinal, so the caller can
|
||||
join it to _rocm_kfd_gpu_pci_ids() by identity: DRM card numbers include
|
||||
foreign adapters and this set includes cards HIP does not enumerate, so any
|
||||
ordinal from this list alone can be shifted relative to ROCm's. A card with
|
||||
missing/unreadable/zero-total figures simply has no entry. Empty off Linux.
|
||||
"""
|
||||
if platform.system() != "Linux":
|
||||
return {}
|
||||
|
||||
try:
|
||||
by_pci: dict[str, tuple[float, float]] = {}
|
||||
for bdf, _card_no, dev_dir in _rocm_linux_amdgpu_cards():
|
||||
if not bdf:
|
||||
continue
|
||||
try:
|
||||
with open(os.path.join(dev_dir, "mem_info_vram_used")) as f:
|
||||
used_bytes = int(f.read().strip())
|
||||
with open(os.path.join(dev_dir, "mem_info_vram_total")) as f:
|
||||
total_bytes = int(f.read().strip())
|
||||
except (OSError, ValueError):
|
||||
continue
|
||||
if total_bytes <= 0:
|
||||
continue
|
||||
by_pci[bdf.lower()] = (
|
||||
round(used_bytes / (1024**3), 2),
|
||||
round(total_bytes / (1024**3), 2),
|
||||
)
|
||||
return by_pci
|
||||
except Exception:
|
||||
return {}
|
||||
|
||||
|
||||
# ── Windows AMD/ROCm per-adapter VRAM (issue #7072) ──────────────────────────
|
||||
# amd-smi is disabled and hipMemGetInfo reports free==total, so read used from the
|
||||
# per-LUID "GPU Adapter Memory" perf counters and take each total from torch, so
|
||||
|
|
@ -1222,6 +1357,75 @@ def _reconcile_primary_rocm_unified_memory(
|
|||
_apply_unified_memory_correction(utilization, torch_devices[0])
|
||||
|
||||
|
||||
def _rocm_visibility_mask_active() -> bool:
|
||||
"""True when any ROCm/CUDA visibility variable filters the device set."""
|
||||
for var in (
|
||||
"HIP_VISIBLE_DEVICES",
|
||||
"ROCR_VISIBLE_DEVICES",
|
||||
"CUDA_VISIBLE_DEVICES",
|
||||
"GPU_DEVICE_ORDINAL",
|
||||
):
|
||||
value = os.environ.get(var)
|
||||
if value and value.strip():
|
||||
return True
|
||||
return False
|
||||
|
||||
|
||||
def _overlay_system_wide_vram(devices: list[Dict[str, Any]]) -> None:
|
||||
"""Replace process-local torch VRAM with system-wide Linux ROCm figures.
|
||||
|
||||
The torch fallback is process-local, so a model served by the separate
|
||||
llama-server process reads as ~0 used even with the GPU full (#7072). DRM
|
||||
sysfs gives per-card figures the kernel updates across all processes. Sources
|
||||
are matched by the device's PHYSICAL index (never list position), and only
|
||||
when NO visibility mask is active and the device count equals the host GPU
|
||||
count; under any mask the index is not a verifiable host ordinal, so torch's
|
||||
figures are kept. Best-effort, in place: a device with no matching card, or a
|
||||
unified-memory APU whose sysfs total is below torch's GTT-backed total, keeps
|
||||
torch's (mirrors _apply_unified_memory_correction).
|
||||
|
||||
Windows is intentionally not overlaid: its per-adapter perf counters cannot be
|
||||
mapped to ROCm ordinals and miss WDDM shared memory, so the multi-GPU view
|
||||
keeps torch there rather than risk misattributing another adapter's usage.
|
||||
"""
|
||||
if not devices or platform.system() != "Linux":
|
||||
return
|
||||
# Match by PCI identity, never list position: index N in KFD topology is ROCm
|
||||
# physical device N and carries its PCI address, which DRM sysfs keys on too.
|
||||
# The two gates below verify ``index`` really is a host-physical ordinal
|
||||
# (torch exposes no PCI id to check directly):
|
||||
# * No visibility mask -- any mask makes ``index`` container/ROCR-relative
|
||||
# rather than a host ordinal.
|
||||
# * Device count == host GPU count -- rules out a device-cgroup container
|
||||
# that sets no env var yet compacts torch's indices from zero.
|
||||
pci_by_ordinal = _rocm_kfd_gpu_pci_ids()
|
||||
if not pci_by_ordinal:
|
||||
return
|
||||
if _rocm_visibility_mask_active() or len(devices) != len(pci_by_ordinal):
|
||||
return
|
||||
vram_by_pci = _rocm_linux_sysfs_vram_by_pci_gb()
|
||||
for dev in devices:
|
||||
index = dev.get("index")
|
||||
if not isinstance(index, int) or not (0 <= index < len(pci_by_ordinal)):
|
||||
continue
|
||||
entry = vram_by_pci.get(pci_by_ordinal[index].lower())
|
||||
if entry is None:
|
||||
continue
|
||||
used, total = entry
|
||||
dev_total = dev.get("vram_total_gb") or 0.0
|
||||
# Overlay only a device that maps 1:1 to the whole card: torch total must
|
||||
# match sysfs total within ~10%. A mismatch either way means a different
|
||||
# memory scope -- a unified-memory APU (sysfs sees only the dedicated
|
||||
# slice, torch the GTT pool) or a partitioned MI300 (sysfs reports the
|
||||
# whole card, dwarfing a partition) -- and overlaying would misstate free
|
||||
# VRAM (a partition would look like it has the whole card free).
|
||||
if dev_total <= 0 or abs(total - dev_total) > 0.1 * dev_total:
|
||||
continue
|
||||
dev["vram_used_gb"] = used
|
||||
dev["vram_total_gb"] = total
|
||||
dev["vram_utilization_pct"] = round((used / total) * 100, 1) if total > 0 else None
|
||||
|
||||
|
||||
def get_visible_gpu_utilization() -> Dict[str, Any]:
|
||||
device = get_device()
|
||||
|
||||
|
|
@ -1317,6 +1521,12 @@ def get_visible_gpu_utilization() -> Dict[str, Any]:
|
|||
"power_utilization_pct": None,
|
||||
}
|
||||
)
|
||||
if IS_ROCM and index_kind == "physical":
|
||||
# Swap process-local torch VRAM for system-wide sysfs so a model
|
||||
# held by the separate llama-server process shows up (#7072).
|
||||
# Physical-index only: a relative index (UUID/MIG mask) is not a
|
||||
# host GPU id. The overlay verifies the rest itself.
|
||||
_overlay_system_wide_vram(devices)
|
||||
return {
|
||||
"available": True,
|
||||
"backend": _backend_label(device),
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue