studio: show system-wide VRAM in the multi-GPU System tab view on ROCm (#7216)

* studio: show system-wide VRAM in the multi-GPU System tab view on ROCm

The System tab's per-GPU list comes from get_visible_gpu_utilization. When
amd-smi is unavailable (always on Windows, minimal Linux installs) it fell back
to torch, whose readings are process-local: on Windows WDDM hands each process
its own budget, so a model held by the separate llama-server process read as
~0 VRAM used even with the GPU full (#7072). The primary-GPU endpoint already
compensates with system-wide sources -- Windows Performance Counters (Task
Manager's source) and Linux DRM sysfs -- but the multi-device endpoint never
got those fallbacks.

Add per-GPU variants of both sources and overlay them onto the torch fallback:
_rocm_windows_perf_counter_vram_per_adapter_gb() attributes Dedicated Usage per
physical adapter (phys_<N> in the counter instance name), and
_rocm_linux_sysfs_vram_per_card_gb() reads mem_info_vram_{used,total} per DRM
card. _overlay_system_wide_vram() applies them to the device list, ROCm-only,
best-effort: unmatched adapters and ambiguous card counts keep the torch
figures, and NVIDIA paths are untouched.

Fixes #7072

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: match VRAM overlay sources by device, honor unified memory, unblock the loop

Five review fixes on the multi-GPU system-wide VRAM overlay:

1. Linux: match DRM cards to devices by PHYSICAL index instead of a positional
   zip, so a reordering visibility mask (HIP_VISIBLE_DEVICES=1,0) no longer
   swaps each card's figures onto the other GPU (which would mislead
   auto_select_gpu_ids and the coexistence checks). An index with no matching
   card keeps its torch figures.

2. Linux: skip the overlay for a device whose sysfs total is below torch's --
   on unified-memory APUs (Strix Halo) mem_info_vram_total is only the small
   dedicated slice while torch sees the GTT-backed pool, and
   _apply_unified_memory_correction already defines larger-total-wins.

3. Windows: group counter instances by adapter LUID, not the phys_<N> suffix --
   separate adapters each read phys_0, which collapsed every GPU into key 0.
   LUIDs are mapped to 0-based positions by ascending value as the closest
   stand-in for device order.

4. Windows: pair the system-wide usage with the physical capacity from
   get_device_properties (as the primary-GPU fallback does) -- under WDDM
   mem_get_info's "total" is the process budget, which misreported capacity
   and pushed utilization to 100%.

5. Run get_visible_gpu_utilization off the event loop in the /hardware/visible
   route (asyncio.to_thread, the repo's convention): the ROCm fallbacks can
   shell out to PowerShell with a 5s timeout, which would stall every other
   request while the System view polls.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio: skip the system-wide VRAM overlay for relative GPU indices

The overlay matches its per-GPU sources (Windows perf counters, Linux sysfs) by
physical device index, but under a UUID/MIG visibility mask the torch fallback
enumerates ordinals and reports index_kind == "relative", where `index` is a
visible ordinal, not a physical id. Applying the overlay there let card/adapter
0's system-wide VRAM overwrite the torch reading of a process that actually
exposes physical GPU 1, misleading auto_select_gpu_ids and the coexistence
checks. Gate the overlay on index_kind == "physical"; relative-index paths keep
the torch fallback.

* studio: drop the unreliable Windows VRAM overlay, keep the Linux one

The multi-GPU system-wide VRAM overlay is now Linux-only. The Windows
per-adapter Performance Counter path could not be made correct: the wildcard
Get-Counter query also returns non-ROCm/iGPU adapters and LUID order is not the
ROCm device order, so an adapter's usage could be overlaid onto the wrong GPU;
and it read only Dedicated Usage, missing WDDM shared memory on unified-memory
GPUs (Strix Halo), overstating free VRAM. Rather than misattribute VRAM and
skew placement decisions, Windows keeps the process-local torch fallback (no
regression vs before this PR); Linux DRM sysfs -- matched by physical index --
still fixes #7072 for the reporter's native-Linux ROCm case.

Removes _rocm_windows_perf_counter_vram_per_adapter_gb and _torch_props_total_gb.

* studio: key sysfs VRAM by DRM card number so filtering can't renumber cards

_rocm_linux_sysfs_vram_per_card_gb dropped cards with a zero total or unreadable
files and then the overlay enumerated the compacted list, so if card0 was
dropped, card1's usage was assigned to physical GPU index 0 (equal-capacity GPUs
slip past the unified-memory total guard). Return {card_number: (used, total)}
and match a device to its card number directly: a hole stays a hole -- device 0
keeps its torch figures when card0 is absent, and card1 maps to device 1.

* studio: key system-wide VRAM by ROCm ordinal, not raw DRM card number

When a non-amdgpu adapter (Intel iGPU, a display-only card) owns an earlier
DRM slot, DRM card numbers stop equalling ROCm device ordinals -- Intel card0
plus AMD card1/card2 gives ROCm devices 0/1, so keying the sysfs overlay by
card number handed ROCm device 1 card1's data (AMD device 0) and left device 0
on stale torch figures, corrupting free-VRAM placement on equal-capacity GPUs.

Only amdgpu cards expose mem_info_vram_*, so the glob already excludes foreign
adapters; order the surviving cards by their PCI address (ROCm/HIP's default
device order, read from each card's device symlink) and key by that position --
the ROCm physical ordinal, which is what the overlay matches against dev index.
An unreadable / zero-total amdgpu card still consumes its ordinal so a later
card is never renumbered onto its slot.

* studio: skip the VRAM overlay under layered HIP-over-ROCR masks

ROCR_VISIBLE_DEVICES filters physical GPUs at the HSA/ROCr layer, and a
HIP_VISIBLE_DEVICES set on top selects WITHIN that already-filtered set
(apply_gpu_ids sets HIP while leaving an inherited ROCR mask in place). When
both are active _get_parent_visible_gpu_spec() prefers the HIP value, so the
reported device index is a ROCR-relative ordinal, not a physical GPU id --
overlaying DRM-sysfs figures by that index would pull another GPU's usage
(e.g. ROCR=2,3 + HIP=1 is physical GPU 3, but the overlay would read card 1),
and equal-capacity cards bypass the total-size safeguard. Detect layered masks
and keep torch's process-local figures there rather than risk misattribution;
a single mask still leaves the index physical and is overlaid as before.

* studio: only overlay whole-card VRAM onto 1:1 ROCm devices

The overlay guard only skipped the case where sysfs total < torch total
(unified-memory APUs), so a partitioned ROCm device (MI300 in CPX mode) --
where HIP exposes several logical devices per physical card but sysfs reports
the whole card's aggregate -- passed the guard: the card total exceeds a
partition's torch total, and the overlay overwrote the partition with
whole-card usage and capacity, letting downstream selection think a partition
had the entire card free. Require the sysfs card total to match the torch
device total (within ~10%) so a mismatch in either direction -- unified memory
(sysfs smaller) or partitioning (sysfs larger) -- keeps torch's figures.

* studio: treat CUDA-over-ROCR as layered, enumerate AMD cards by driver

Two remaining mismatches between the reported device index and the DRM card the
overlay reads:

- On ROCm the HIP layer honors CUDA_VISIBLE_DEVICES as well as
  HIP_VISIBLE_DEVICES, so a CUDA mask composed over ROCR layers identically:
  ROCR=2,3 with CUDA=1 is physical GPU 3, yet the spec reports the ROCR value
  [2,3] and the device was labeled index 2, overlaying card 2's usage onto GPU 3.
  The layered check now treats ROCR combined with either HIP or CUDA as layered.

- The ROCm device set is now enumerated by bound driver (device/driver resolves
  to amdgpu) instead of by the presence of mem_info_vram_*. An AMD device with
  incomplete sysfs support (some APUs expose no VRAM files at all) was omitted
  by the glob entirely and shifted every later card down one ordinal, letting a
  similar-capacity GPU pass the total guard with another device's usage. Such a
  card now consumes its ordinal and simply yields no entry.

* studio: honor GPU_DEVICE_ORDINAL and require an unambiguous card mapping

Two remaining ways the reported device index could be matched to the wrong DRM
card:

- GPU_DEVICE_ORDINAL is a supported ROCm visibility variable that
  _get_parent_visible_gpu_spec() never consults, so GPU_DEVICE_ORDINAL=1
  surfaces physical GPU 1 as torch ordinal 0 and it was mislabeled index 0,
  overlaying card 0's usage onto GPU 1. The mask check now covers it, and is
  renamed _rocm_device_index_unreliable() to say what it actually decides.

- driver == amdgpu is only a SUPERSET of the ROCm-visible set: an amdgpu-bound
  adapter HIP cannot enumerate (an unsupported older AMD GPU beside a supported
  one) still took an ordinal and shifted every real compute device. There is no
  torch-side PCI identity to match against, so the overlay now requires the
  amdgpu card count to equal the device count -- exactly the condition under
  which position-in-PCI-order is a sound 1:1 mapping. Any disagreement keeps
  torch's process-local figures: less informative, never misattributed.

* studio: keep the VRAM overlay working for masked GPU subsets

The card-count guard compared the amdgpu card list against the VISIBLE device
list, so any visibility mask disabled the overlay outright: HIP_VISIBLE_DEVICES=1,3
on a four-GPU host gives two devices against four cards. Those masked GPUs then
kept reporting process-local torch usage, hiding VRAM held by llama-server and
letting the training/chat placement checks overestimate free memory -- the exact
problem the overlay exists to fix.

The count check now applies only when no visibility mask is active, which is the
case where the reported devices really are the whole host and a mismatch means an
amdgpu adapter ROCm cannot enumerate is shifting the ordinals. Under a mask the
subset is expected, so each device's physical index is validated individually
instead: the per-card lookup bounds-checks it and the total-size guard rejects a
card whose capacity does not match the device's.

* studio: match GPUs to DRM cards by PCI identity, not by position

Every mapping bug on this PR came from the same root cause: there was no
authoritative link between a reported device index and a DRM card, so the
overlay kept inferring one positionally and each heuristic broke on a new host
shape -- foreign adapters on earlier DRM slots, cards with no VRAM sysfs, and
most recently amdgpu-bound adapters HIP cannot enumerate, which the count guard
could only catch on an unmasked host and therefore missed under any mask.

Use the link ROCm itself enumerates from. KFD topology
(/sys/class/kfd/kfd/topology/nodes/<N>/properties) lists exactly the GPUs HIP
exposes -- GPU nodes in node-id order are HIP's device order -- and each carries
its PCI location, so index N there IS physical device N with a stable identity.
DRM sysfs now supplies system-wide VRAM keyed by that same PCI address, and the
overlay is a join on it.

Every previous skew becomes a failed join rather than a misattribution: an
unenumerable adapter has no KFD node so it never takes an ordinal, a foreign
adapter contributes no entry, and a masked subset resolves each physical index
directly. That removes the count heuristic and its mask exception entirely. With
no KFD topology there is no identity to join on, so the overlay is skipped rather
than guessing positionally.

* studio: require verified host visibility and AMD-only KFD nodes

Three ways the identity map could still be built on a false premise:

- The NVIDIA open kernel module registers KFD topology nodes with a positive
  SIMD count, so an earlier NVIDIA node shifted every AMD ordinal and ROCm
  device 1 resolved to AMD GPU 0. GPU nodes now require vendor_id 4098 (0x1002),
  the same filter install.sh already applies for this exact reason.

- A GPU node with an unreadable properties file or no location_id was skipped,
  which silently shifted every later ordinal. Both now fail the whole map
  closed, so the overlay is disabled rather than misattributing.

- A container exposing only some render devices through device cgroups sets no
  visibility variable, yet torch compacts what it can see to ordinals from zero
  while the host-mounted KFD and DRM trees still list every GPU. Nothing in the
  reported payload distinguishes that from a full host, and torch exposes no PCI
  id to check against, so the overlay now runs only when host visibility is
  positively verified: no visibility mask AND device count equal to the host GPU
  count. That also subsumes the previous layered-mask and GPU_DEVICE_ORDINAL
  checks, so _rocm_device_index_unreliable() is gone.

This trades coverage for correctness: masked subsets and filtered containers now
keep torch's process-local figures instead of a mapping that cannot be verified.

* Fix the multi-GPU VRAM overlay docstring for PR #7216

The docstring claimed a reordering mask keeps each card on the right GPU,
but the overlay skips any active visibility mask and keeps torch's figures.
State the actual gating instead.

* Tighten comments in the multi-GPU VRAM overlay and its tests

Collapse the verbose docstrings and inline explanations added for the Linux
ROCm system-wide VRAM overlay to succinct one-liners, keeping the non-obvious
rationale (fail-closed KFD mapping, PCI-identity join, mask gating, the 10%
whole-card guard). Comments only, no behavior change.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: danielhanchen <michaelhan2050@gmail.com>
This commit is contained in:
Hakan Baysal 2026-07-22 13:55:35 +03:00 committed by GitHub
commit 55433bd7b8
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
3 changed files with 767 additions and 1 deletions

View file

@ -109,7 +109,9 @@ async def get_hardware_utilization(current_subject: str = Depends(get_current_su
@router.get("/hardware/visible")
async def get_visible_hardware_utilization(current_subject: str = Depends(get_current_subject)):
from utils.hardware import get_visible_gpu_utilization
return get_visible_gpu_utilization()
# Off the event loop: the ROCm fallbacks shell out (Windows perf counters, sysfs) and the System view polls this route.
return await asyncio.to_thread(get_visible_gpu_utilization)
@router.post("/start")

View file

@ -0,0 +1,554 @@
# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""The System tab's multi-GPU view must show system-wide VRAM on ROCm (#7072).
When amd-smi is unavailable, get_visible_gpu_utilization fell back to torch,
whose readings are process-local: a model held by the separate llama-server
process read as ~0 VRAM used even with the GPU full. These tests cover the
per-GPU system-wide overlay the multi-device endpoint now applies, matched by
physical device identity.
"""
from __future__ import annotations
import importlib
import sys
import types
from pathlib import Path
_BACKEND_DIR = Path(__file__).resolve().parent.parent
if str(_BACKEND_DIR) not in sys.path:
sys.path.insert(0, str(_BACKEND_DIR))
def _maybe_stub(name: str, builder):
# Stub only if the real module is missing, so we never shadow it for later tests.
try:
importlib.import_module(name)
except ImportError:
sys.modules[name] = builder()
def _build_loggers_stub():
m = types.ModuleType("loggers")
m.get_logger = lambda name: __import__("logging").getLogger(name)
return m
def _build_structlog_stub():
m = types.ModuleType("structlog")
m.get_logger = lambda *a, **k: __import__("logging").getLogger("stub")
return m
_maybe_stub("loggers", _build_loggers_stub)
_maybe_stub("structlog", _build_structlog_stub)
import utils.hardware.hardware as hw # noqa: E402
def _device(
index,
used,
total,
*,
ordinal = None,
):
return {
"index": index,
"index_kind": "physical",
"visible_ordinal": index if ordinal is None else ordinal,
"gpu_utilization_pct": None,
"temperature_c": None,
"vram_used_gb": used,
"vram_total_gb": total,
"vram_utilization_pct": round((used / total) * 100, 1) if total > 0 else None,
"power_draw_w": None,
"power_limit_w": None,
"power_utilization_pct": None,
}
# ── Linux per-card sysfs ──
def _fake_drm(tmp_path, monkeypatch, cards):
"""Fake /sys/class/drm tree; glob returns cards REVERSED so the PCI sort must order them.
``cards``: (card_no, pci_bdf, driver, vram) tuples; vram is (used_gb, total_gb)
or None for a device with no mem_info_vram_* files.
"""
drivers = tmp_path / "drivers"
card_paths = []
for card_no, bdf, driver, vram in cards:
pci_dir = tmp_path / "pci" / bdf
pci_dir.mkdir(parents = True, exist_ok = True)
drv_dir = drivers / driver
drv_dir.mkdir(parents = True, exist_ok = True)
(pci_dir / "driver").symlink_to(drv_dir)
if vram is not None:
used, total = vram
(pci_dir / "mem_info_vram_used").write_text(str(int(used * 1024**3)))
(pci_dir / "mem_info_vram_total").write_text(str(int(total * 1024**3)))
card_dir = tmp_path / "drm" / f"card{card_no}"
card_dir.mkdir(parents = True, exist_ok = True)
(card_dir / "device").symlink_to(pci_dir)
card_paths.append(str(card_dir))
monkeypatch.setattr(hw.glob, "glob", lambda pattern: list(reversed(card_paths)))
return card_paths
def test_linux_vram_keyed_by_pci_excludes_foreign_adapters(monkeypatch, tmp_path):
# Foreign (non-amdgpu) adapters contribute no entry, so they cannot shift ordinals.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
_fake_drm(
tmp_path,
monkeypatch,
[
(0, "0000:00:02.0", "i915", (0.5, 2.0)), # foreign adapter: excluded
(1, "0000:03:00.0", "amdgpu", (40, 48)), # AMD device 0
(2, "0000:41:00.0", "amdgpu", (1, 8)), # AMD device 1
],
)
assert hw._rocm_linux_sysfs_vram_by_pci_gb() == {
"0000:03:00.0": (40.0, 48.0),
"0000:41:00.0": (1.0, 8.0),
}
def test_linux_vram_omits_bad_cards_without_shifting(monkeypatch, tmp_path):
# A zero-total card has no entry; identity keying means its absence renumbers nothing.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
_fake_drm(
tmp_path,
monkeypatch,
[
(0, "0000:03:00.0", "amdgpu", (0, 0)), # zero total -> no entry
(1, "0000:41:00.0", "amdgpu", (2, 16)),
],
)
assert hw._rocm_linux_sysfs_vram_by_pci_gb() == {"0000:41:00.0": (2.0, 16.0)}
def test_linux_vram_omits_amd_card_without_vram_files(monkeypatch, tmp_path):
# An APU with no mem_info_vram_* files has no entry; the discrete card keeps its address.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
_fake_drm(
tmp_path,
monkeypatch,
[
(0, "0000:03:00.0", "amdgpu", None), # APU: no VRAM sysfs files
(1, "0000:41:00.0", "amdgpu", (2, 16)),
],
)
assert hw._rocm_linux_sysfs_vram_by_pci_gb() == {"0000:41:00.0": (2.0, 16.0)}
# ── KFD topology: the authoritative ROCm device order ──
_AMD = 4098 # 0x1002
_NVIDIA = 4318 # 0x10DE -- the open kernel module also registers KFD nodes
def _fake_kfd(tmp_path, monkeypatch, nodes):
"""Fake KFD topology nodes tree, returned out of node order so the sort must order it.
``nodes``: (node_id, simd_count, location_id, domain, vendor_id); simd_count 0
marks a CPU node, location_id None omits the property.
"""
node_paths = []
for node_id, simd_count, location_id, domain, vendor_id in nodes:
d = tmp_path / "kfd" / str(node_id)
d.mkdir(parents = True, exist_ok = True)
lines = [f"cpu_cores_count {0 if simd_count else 8}", f"simd_count {simd_count}"]
if location_id is not None:
lines.append(f"location_id {location_id}")
lines.append(f"domain {domain}")
if vendor_id is not None:
lines.append(f"vendor_id {vendor_id}")
(d / "properties").write_text("\n".join(lines) + "\n")
node_paths.append(str(d))
monkeypatch.setattr(hw.glob, "glob", lambda pattern: list(reversed(node_paths)))
return node_paths
def test_kfd_lists_gpu_nodes_in_device_order(monkeypatch, tmp_path):
# The CPU node (simd_count 0) takes no ordinal; GPU nodes in node-id order are HIP's order.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
_fake_kfd(
tmp_path,
monkeypatch,
[
(0, 0, None, 0, None), # CPU node
(1, 304, (0x03 << 8) | (0x00 << 3) | 0, 0, _AMD), # 0000:03:00.0 -> dev 0
(2, 304, (0x41 << 8) | (0x00 << 3) | 0, 0, _AMD), # 0000:41:00.0 -> dev 1
],
)
assert hw._rocm_kfd_gpu_pci_ids() == ["0000:03:00.0", "0000:41:00.0"]
def test_kfd_decodes_domain_device_and_function(monkeypatch, tmp_path):
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
_fake_kfd(tmp_path, monkeypatch, [(1, 64, (0xC1 << 8) | (0x1F << 3) | 5, 0x1234, _AMD)])
assert hw._rocm_kfd_gpu_pci_ids() == ["1234:c1:1f.5"]
def test_kfd_skips_non_amd_gpu_nodes(monkeypatch, tmp_path):
# An NVIDIA KFD node is not a HIP device: it must take no ordinal, else it
# shifts every AMD GPU and ROCm device 1 resolves to AMD GPU 0.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
_fake_kfd(
tmp_path,
monkeypatch,
[
(0, 0, None, 0, None), # CPU
(1, 128, (0x01 << 8) | 0, 0, _NVIDIA), # NVIDIA: no ordinal
(2, 304, (0x03 << 8) | 0, 0, _AMD), # AMD device 0
(3, 304, (0x41 << 8) | 0, 0, _AMD), # AMD device 1
],
)
assert hw._rocm_kfd_gpu_pci_ids() == ["0000:03:00.0", "0000:41:00.0"]
def test_kfd_fails_closed_when_a_gpu_has_no_location(monkeypatch, tmp_path):
# Dropping an unplaceable AMD GPU shifts later ordinals; fail closed for the whole map.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
_fake_kfd(
tmp_path,
monkeypatch,
[
(1, 304, None, 0, _AMD), # AMD GPU with no location_id
(2, 304, (0x41 << 8) | 0, 0, _AMD),
],
)
assert hw._rocm_kfd_gpu_pci_ids() == []
def test_kfd_fails_closed_when_a_node_is_unreadable(monkeypatch, tmp_path):
# An unreadable node could be a GPU; assuming otherwise would shift ordinals.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
paths = _fake_kfd(
tmp_path,
monkeypatch,
[
(1, 304, (0x03 << 8) | 0, 0, _AMD),
(2, 304, (0x41 << 8) | 0, 0, _AMD),
],
)
(Path(paths[0]) / "properties").unlink()
assert hw._rocm_kfd_gpu_pci_ids() == []
def test_kfd_absent_yields_no_device_order(monkeypatch):
monkeypatch.setattr(hw.glob, "glob", lambda pattern: [])
assert hw._rocm_kfd_gpu_pci_ids() == []
# ── overlay ──
def _patch_pci_map(monkeypatch, bdfs):
"""Declare the ROCm device order by PCI address (index N is device N) and clear
the visibility masks the overlay requires unset.
"""
for var in (
"HIP_VISIBLE_DEVICES",
"ROCR_VISIBLE_DEVICES",
"CUDA_VISIBLE_DEVICES",
"GPU_DEVICE_ORDINAL",
):
monkeypatch.delenv(var, raising = False)
monkeypatch.setattr(hw, "_rocm_kfd_gpu_pci_ids", lambda: list(bdfs))
def _pci(n):
"""A distinct, well-formed PCI address for card n."""
return f"0000:{n:02x}:00.0"
def test_overlay_windows_is_noop_keeps_torch(monkeypatch):
# Windows is intentionally not overlaid (perf counters can't map to ROCm ordinals): keep torch.
monkeypatch.setattr(hw.platform, "system", lambda: "Windows")
monkeypatch.setattr(
hw,
"_rocm_linux_sysfs_vram_by_pci_gb",
lambda: (_ for _ in ()).throw(AssertionError("sysfs must not run on Windows")),
)
devices = [_device(0, used = 0.02, total = 8.0)]
_patch_pci_map(monkeypatch, [_pci(0)])
hw._overlay_system_wide_vram(devices)
assert devices[0]["vram_used_gb"] == 0.02 # untouched
def test_overlay_linux_matches_by_device_ordinal(monkeypatch):
# Devices arriving as [index 1, index 0] each get their own GPU's figures by ordinal.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
monkeypatch.setattr(
hw,
"_rocm_linux_sysfs_vram_by_pci_gb",
lambda: {_pci(0): (30.0, 45.0), _pci(1): (0.5, 8.0)}, # dev 0 big, dev 1 small
)
devices = [_device(1, used = 0.01, total = 8.0), _device(0, used = 0.02, total = 45.0)]
_patch_pci_map(monkeypatch, [_pci(0), _pci(1)])
hw._overlay_system_wide_vram(devices)
assert devices[0]["vram_used_gb"] == 0.5 # index 1 -> device 1 (small)
assert devices[0]["vram_total_gb"] == 8.0
assert devices[1]["vram_used_gb"] == 30.0 # index 0 -> device 0 (big)
assert devices[1]["vram_total_gb"] == 45.0
def test_overlay_linux_ordinal_hole_does_not_shift(monkeypatch):
# Device 0's card dropped: index 0 keeps torch, index 1 still maps to ordinal 1 (no compaction).
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
monkeypatch.setattr(hw, "_rocm_linux_sysfs_vram_by_pci_gb", lambda: {_pci(1): (0.5, 8.0)})
devices = [_device(0, used = 0.02, total = 45.0), _device(1, used = 0.01, total = 8.0)]
_patch_pci_map(monkeypatch, [_pci(0), _pci(1)])
hw._overlay_system_wide_vram(devices)
assert devices[0]["vram_used_gb"] == 0.02 # no ordinal 0 -> torch kept
assert devices[1]["vram_used_gb"] == 0.5 # ordinal 1 -> device 1, not device 0
def test_overlay_linux_skips_unified_memory_card(monkeypatch):
# Unified-memory APU: the smaller sysfs total must not shrink torch's GTT-backed pool.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
monkeypatch.setattr(hw, "_rocm_linux_sysfs_vram_by_pci_gb", lambda: {_pci(0): (0.4, 1.0)})
devices = [_device(0, used = 12.0, total = 96.0)] # torch's unified pool
_patch_pci_map(monkeypatch, [_pci(0)])
hw._overlay_system_wide_vram(devices)
assert devices[0]["vram_used_gb"] == 12.0
assert devices[0]["vram_total_gb"] == 96.0
def test_overlay_linux_skips_partitioned_device(monkeypatch):
# Partitioned MI300: the whole-card sysfs total dwarfs the partition, so the overlay must not overwrite it.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
monkeypatch.setattr(hw, "_rocm_linux_sysfs_vram_by_pci_gb", lambda: {_pci(0): (40.0, 192.0)})
devices = [_device(0, used = 1.0, total = 24.0)] # torch partition
_patch_pci_map(monkeypatch, [_pci(0)])
hw._overlay_system_wide_vram(devices)
assert devices[0]["vram_used_gb"] == 1.0 # partition figures kept
assert devices[0]["vram_total_gb"] == 24.0
def test_overlay_linux_out_of_range_index_untouched(monkeypatch):
# A masked host exposing physical index 5 with no card 5: keep torch data.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
monkeypatch.setattr(
hw, "_rocm_linux_sysfs_vram_by_pci_gb", lambda: {_pci(0): (30.0, 45.0), _pci(1): (0.5, 8.0)}
)
devices = [_device(5, used = 0.02, total = 45.0)]
_patch_pci_map(monkeypatch, [_pci(0), _pci(1)])
hw._overlay_system_wide_vram(devices)
assert devices[0]["vram_used_gb"] == 0.02
def test_overlay_ignores_adapters_rocm_cannot_enumerate(monkeypatch):
# A HIP-unenumerable amdgpu adapter has no KFD node, so device 0 resolves to
# the supported GPU's own address, never the display card's.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
monkeypatch.setattr(
hw,
"_rocm_linux_sysfs_vram_by_pci_gb",
# Both in DRM sysfs with similar capacity -- what the total-size guard can't separate.
lambda: {_pci(9): (30.0, 45.0), _pci(3): (12.0, 45.0)},
)
_patch_pci_map(monkeypatch, [_pci(3)]) # KFD lists only the supported GPU
devices = [_device(0, used = 0.02, total = 45.0)] # torch sees that one GPU
hw._overlay_system_wide_vram(devices)
assert devices[0]["vram_used_gb"] == 12.0 # the supported GPU's own figures
def test_overlay_skips_masked_subsets(monkeypatch):
# Under a mask the index is not verifiably a host ordinal, so keep torch's figures.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
_patch_pci_map(monkeypatch, [_pci(0), _pci(1), _pci(2), _pci(3)])
monkeypatch.setenv("HIP_VISIBLE_DEVICES", "1,3")
monkeypatch.setattr(
hw,
"_rocm_linux_sysfs_vram_by_pci_gb",
lambda: {_pci(1): (30.0, 48.0), _pci(3): (12.0, 48.0)},
)
devices = [_device(1, used = 0.02, total = 48.0), _device(3, used = 0.01, total = 48.0)]
hw._overlay_system_wide_vram(devices)
assert devices[0]["vram_used_gb"] == 0.02 # torch kept
assert devices[1]["vram_used_gb"] == 0.01
def test_overlay_skips_device_cgroup_filtered_container(monkeypatch):
# A device-cgroup container sets no env var yet compacts torch's indices from
# zero while KFD/DRM list every GPU, so the count mismatch must disable the overlay.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
_patch_pci_map(monkeypatch, [_pci(0), _pci(1), _pci(2), _pci(3)]) # host has 4
monkeypatch.setattr(
hw,
"_rocm_linux_sysfs_vram_by_pci_gb",
lambda: {_pci(0): (30.0, 48.0), _pci(2): (12.0, 48.0)},
)
devices = [_device(0, used = 0.02, total = 48.0)] # container sees 1, as index 0
hw._overlay_system_wide_vram(devices)
assert devices[0]["vram_used_gb"] == 0.02 # torch kept, not host GPU 0's 30.0
def test_overlay_skips_without_kfd_topology(monkeypatch):
# No KFD means no identity to join on; fall back to torch rather than guess.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
monkeypatch.setattr(hw, "_rocm_kfd_gpu_pci_ids", lambda: [])
monkeypatch.setattr(
hw,
"_rocm_linux_sysfs_vram_by_pci_gb",
lambda: (_ for _ in ()).throw(AssertionError("must not read sysfs without KFD")),
)
devices = [_device(0, used = 0.02, total = 45.0)]
hw._overlay_system_wide_vram(devices)
assert devices[0]["vram_used_gb"] == 0.02
def test_overlay_empty_devices_is_noop(monkeypatch):
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
hw._overlay_system_wide_vram([]) # must not raise
# ── integration: the ROCm torch fallback applies the overlay ──
def test_visible_utilization_rocm_fallback_overlays(monkeypatch):
for _var in (
"HIP_VISIBLE_DEVICES",
"ROCR_VISIBLE_DEVICES",
"CUDA_VISIBLE_DEVICES",
"GPU_DEVICE_ORDINAL",
):
monkeypatch.delenv(_var, raising = False)
monkeypatch.setattr(hw, "IS_ROCM", True)
monkeypatch.setattr(hw, "get_device", lambda: hw.DeviceType.CUDA)
monkeypatch.setattr(hw, "_smi_query", lambda *a, **k: None) # amd-smi unavailable
monkeypatch.setattr(
hw,
"_get_parent_visible_gpu_spec",
lambda: {"raw": None, "numeric_ids": [0, 1], "supports_explicit_gpu_ids": True},
)
monkeypatch.setattr(hw, "get_parent_visible_gpu_ids", lambda: [0, 1])
monkeypatch.setattr(
hw,
"_torch_get_per_device_info",
lambda ids: [
{"index": 0, "visible_ordinal": 0, "used_gb": 0.02, "total_gb": 45.0},
{"index": 1, "visible_ordinal": 1, "used_gb": 0.01, "total_gb": 8.0},
],
)
overlaid = []
monkeypatch.setattr(
hw, "_overlay_system_wide_vram", lambda devices: overlaid.append(len(devices))
)
result = hw.get_visible_gpu_utilization()
assert result["available"] is True
assert overlaid == [2]
def test_visible_utilization_relative_index_skips_overlay(monkeypatch):
# UUID/MIG mask gives relative indices; the overlay matches physical index, so it must not run.
monkeypatch.setattr(hw, "IS_ROCM", True)
monkeypatch.setattr(hw, "get_device", lambda: hw.DeviceType.CUDA)
monkeypatch.setattr(hw, "_smi_query", lambda *a, **k: None)
monkeypatch.setattr(
hw,
"_get_parent_visible_gpu_spec",
lambda: {"raw": "GPU-uuid-a", "numeric_ids": None, "supports_explicit_gpu_ids": False},
)
monkeypatch.setattr(hw, "get_parent_visible_gpu_ids", lambda: []) # UUID mask
monkeypatch.setattr(hw, "_torch_get_physical_gpu_count", lambda: 1)
monkeypatch.setattr(
hw,
"_torch_get_per_device_info",
lambda ids: [{"index": 0, "visible_ordinal": 0, "used_gb": 0.02, "total_gb": 8.0}],
)
called = []
monkeypatch.setattr(hw, "_overlay_system_wide_vram", lambda devices: called.append(1))
result = hw.get_visible_gpu_utilization()
assert result["index_kind"] == "relative"
assert called == []
def test_visible_utilization_nvidia_fallback_skips_overlay(monkeypatch):
monkeypatch.setattr(hw, "IS_ROCM", False)
monkeypatch.setattr(hw, "get_device", lambda: hw.DeviceType.CUDA)
monkeypatch.setattr(hw, "_smi_query", lambda *a, **k: None)
monkeypatch.setattr(
hw,
"_get_parent_visible_gpu_spec",
lambda: {"raw": None, "numeric_ids": [0], "supports_explicit_gpu_ids": True},
)
monkeypatch.setattr(hw, "get_parent_visible_gpu_ids", lambda: [0])
monkeypatch.setattr(
hw,
"_torch_get_per_device_info",
lambda ids: [{"index": 0, "visible_ordinal": 0, "used_gb": 1.0, "total_gb": 24.0}],
)
called = []
monkeypatch.setattr(hw, "_overlay_system_wide_vram", lambda devices: called.append(1))
result = hw.get_visible_gpu_utilization()
assert result["available"] is True
assert called == []
def test_any_visibility_mask_is_detected(monkeypatch):
# Any of these makes the index not a host-physical ordinal, so each must disable the overlay.
for var in (
"HIP_VISIBLE_DEVICES",
"ROCR_VISIBLE_DEVICES",
"CUDA_VISIBLE_DEVICES",
"GPU_DEVICE_ORDINAL",
):
monkeypatch.delenv(var, raising = False)
assert hw._rocm_visibility_mask_active() is False
for var in (
"HIP_VISIBLE_DEVICES",
"ROCR_VISIBLE_DEVICES",
"CUDA_VISIBLE_DEVICES",
"GPU_DEVICE_ORDINAL",
):
monkeypatch.setenv(var, "1")
assert hw._rocm_visibility_mask_active() is True, var
monkeypatch.setenv(var, " ") # empty is not an active filter
assert hw._rocm_visibility_mask_active() is False, var
monkeypatch.delenv(var, raising = False)
def test_overlay_skips_under_gpu_device_ordinal(monkeypatch):
# GPU_DEVICE_ORDINAL=1 surfaces GPU 1 as torch ordinal 0, so index 0 is not GPU 0; overlay must not run.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
_patch_pci_map(monkeypatch, [_pci(0)])
monkeypatch.setenv("GPU_DEVICE_ORDINAL", "1")
monkeypatch.setattr(hw, "_rocm_linux_sysfs_vram_by_pci_gb", lambda: {_pci(0): (30.0, 45.0)})
devices = [_device(0, used = 0.02, total = 45.0)]
hw._overlay_system_wide_vram(devices)
assert devices[0]["vram_used_gb"] == 0.02
def test_visible_utilization_delegates_gating_to_the_overlay(monkeypatch):
# The call site no longer pre-checks masks; the overlay gates itself, so a physical payload always reaches it.
monkeypatch.setenv("ROCR_VISIBLE_DEVICES", "2,3")
monkeypatch.setenv("HIP_VISIBLE_DEVICES", "1")
monkeypatch.setattr(hw, "IS_ROCM", True)
monkeypatch.setattr(hw, "get_device", lambda: hw.DeviceType.CUDA)
monkeypatch.setattr(hw, "_smi_query", lambda *a, **k: None)
monkeypatch.setattr(
hw,
"_get_parent_visible_gpu_spec",
lambda: {"raw": "1", "numeric_ids": [1], "supports_explicit_gpu_ids": True},
)
monkeypatch.setattr(hw, "get_parent_visible_gpu_ids", lambda: [1])
monkeypatch.setattr(
hw,
"_torch_get_per_device_info",
lambda ids: [{"index": 1, "visible_ordinal": 0, "used_gb": 0.02, "total_gb": 8.0}],
)
# Real overlay + gating: the layered mask must leave torch's figures.
monkeypatch.setattr(hw.platform, "system", lambda: "Linux")
monkeypatch.setattr(hw, "_rocm_kfd_gpu_pci_ids", lambda: [_pci(0), _pci(1)])
monkeypatch.setattr(hw, "_rocm_linux_sysfs_vram_by_pci_gb", lambda: {_pci(1): (30.0, 8.0)})
result = hw.get_visible_gpu_utilization()
assert result["index_kind"] == "physical"
assert result["devices"][0]["vram_used_gb"] == 0.02 # untouched

View file

@ -734,6 +734,141 @@ def _rocm_linux_sysfs_vram_gb() -> tuple[Optional[float], Optional[float]]:
return None, None
# 0x1002. NVIDIA's open kernel module also registers KFD nodes (vendor_id 0x10DE);
# a non-AMD node is not a HIP device and must never take an ordinal.
_AMD_PCI_VENDOR_ID = 4098
def _rocm_kfd_gpu_pci_ids() -> list[str]:
"""PCI addresses of the GPUs ROCm enumerates, in HIP device order.
Reads /sys/class/kfd/kfd/topology/nodes/<N>/properties, the topology ROCm
itself enumerates from: AMD GPU nodes (simd_count > 0 excludes CPUs,
vendor_id == AMD excludes NVIDIA) in node-id order are HIP's device order, so
position N is ROCm physical device N. Unlike DRM sysfs, an amdgpu adapter HIP
cannot enumerate has no node here, so it never consumes an ordinal.
Returns [] (disabling the overlay) when KFD is absent, and FAILS CLOSED the
same way on any unreadable node or an AMD node with no location_id: dropping
one would shift every later ordinal and let a similar-capacity GPU pass the
total-size guard while showing another card's usage.
location_id is the kernel's (bus << 8) | devfn; domain is separate.
"""
nodes: list[tuple[int, str]] = []
try:
node_dirs = glob.glob("/sys/class/kfd/kfd/topology/nodes/*")
except Exception:
return []
for node_dir in node_dirs:
m = re.fullmatch(r".*/(\d+)", node_dir)
if m is None:
continue
props: dict[str, int] = {}
try:
with open(os.path.join(node_dir, "properties")) as f:
for line in f:
parts = line.split()
if len(parts) == 2:
try:
props[parts[0]] = int(parts[1])
except ValueError:
continue
except OSError:
return [] # unreadable node could be a GPU: fail closed, don't shift
if props.get("simd_count", 0) <= 0:
continue # CPU node, not a GPU
if props.get("vendor_id") != _AMD_PCI_VENDOR_ID:
continue # non-AMD GPU node (NVIDIA open driver): not a HIP device
location_id = props.get("location_id")
if location_id is None:
return [] # an AMD GPU we cannot place: fail closed for the whole map
domain = props.get("domain", 0)
bus = (location_id >> 8) & 0xFF
devfn = location_id & 0xFF
bdf = f"{domain:04x}:{bus:02x}:{(devfn >> 3) & 0x1F:02x}.{devfn & 0x7}"
nodes.append((int(m.group(1)), bdf))
nodes.sort(key = lambda n: n[0])
return [bdf for _node_id, bdf in nodes]
def _rocm_linux_amdgpu_cards() -> list[tuple[str, int, str]]:
"""The amdgpu-bound DRM cards in PCI order: ``(pci_bdf, card_no, device_dir)``.
Membership is by the BOUND DRIVER, not the VRAM sysfs files: an AMD device
with incomplete sysfs support (some APUs expose no mem_info_vram_*) still
consumes a ROCm ordinal, and dropping it would shift every later card down.
PCI order is HIP's default enumeration order, so list position is the ROCm
ordinal; card_no is a stable tiebreak when the BDF cannot be resolved.
NOTE this is a superset of the ROCm-visible set (a HIP-unsupported amdgpu
adapter appears too), so callers must check the counts agree before assuming
a 1:1 mapping onto torch devices.
"""
if platform.system() != "Linux":
return []
amd_cards: list[tuple[str, int, str]] = []
try:
for card_path in glob.glob("/sys/class/drm/card*"):
# Match card<N> exactly so connector nodes (card0-DP-1) are skipped.
m = re.fullmatch(r".*/card(\d+)", card_path)
if m is None:
continue
dev_dir = os.path.join(card_path, "device")
try:
driver = os.path.basename(os.path.realpath(os.path.join(dev_dir, "driver")))
except OSError:
continue
if driver != "amdgpu":
continue # foreign adapter: not a ROCm device, takes no ordinal
try:
bdf = os.path.basename(os.path.realpath(dev_dir))
except OSError:
bdf = ""
amd_cards.append((bdf, int(m.group(1)), dev_dir))
except Exception:
return []
amd_cards.sort(key = lambda c: (c[0], c[1]))
return amd_cards
def _rocm_linux_sysfs_vram_by_pci_gb() -> dict[str, tuple[float, float]]:
"""System-wide AMD VRAM via Linux DRM sysfs, keyed by the card's PCI address.
Reads each card's mem_info_vram_{used,total} (kernel-updated across all
processes) so every GPU gets its own figure, unlike _rocm_linux_sysfs_vram_gb
which sums the host. Keyed by PCI address, not an ordinal, so the caller can
join it to _rocm_kfd_gpu_pci_ids() by identity: DRM card numbers include
foreign adapters and this set includes cards HIP does not enumerate, so any
ordinal from this list alone can be shifted relative to ROCm's. A card with
missing/unreadable/zero-total figures simply has no entry. Empty off Linux.
"""
if platform.system() != "Linux":
return {}
try:
by_pci: dict[str, tuple[float, float]] = {}
for bdf, _card_no, dev_dir in _rocm_linux_amdgpu_cards():
if not bdf:
continue
try:
with open(os.path.join(dev_dir, "mem_info_vram_used")) as f:
used_bytes = int(f.read().strip())
with open(os.path.join(dev_dir, "mem_info_vram_total")) as f:
total_bytes = int(f.read().strip())
except (OSError, ValueError):
continue
if total_bytes <= 0:
continue
by_pci[bdf.lower()] = (
round(used_bytes / (1024**3), 2),
round(total_bytes / (1024**3), 2),
)
return by_pci
except Exception:
return {}
# ── Windows AMD/ROCm per-adapter VRAM (issue #7072) ──────────────────────────
# amd-smi is disabled and hipMemGetInfo reports free==total, so read used from the
# per-LUID "GPU Adapter Memory" perf counters and take each total from torch, so
@ -1222,6 +1357,75 @@ def _reconcile_primary_rocm_unified_memory(
_apply_unified_memory_correction(utilization, torch_devices[0])
def _rocm_visibility_mask_active() -> bool:
"""True when any ROCm/CUDA visibility variable filters the device set."""
for var in (
"HIP_VISIBLE_DEVICES",
"ROCR_VISIBLE_DEVICES",
"CUDA_VISIBLE_DEVICES",
"GPU_DEVICE_ORDINAL",
):
value = os.environ.get(var)
if value and value.strip():
return True
return False
def _overlay_system_wide_vram(devices: list[Dict[str, Any]]) -> None:
"""Replace process-local torch VRAM with system-wide Linux ROCm figures.
The torch fallback is process-local, so a model served by the separate
llama-server process reads as ~0 used even with the GPU full (#7072). DRM
sysfs gives per-card figures the kernel updates across all processes. Sources
are matched by the device's PHYSICAL index (never list position), and only
when NO visibility mask is active and the device count equals the host GPU
count; under any mask the index is not a verifiable host ordinal, so torch's
figures are kept. Best-effort, in place: a device with no matching card, or a
unified-memory APU whose sysfs total is below torch's GTT-backed total, keeps
torch's (mirrors _apply_unified_memory_correction).
Windows is intentionally not overlaid: its per-adapter perf counters cannot be
mapped to ROCm ordinals and miss WDDM shared memory, so the multi-GPU view
keeps torch there rather than risk misattributing another adapter's usage.
"""
if not devices or platform.system() != "Linux":
return
# Match by PCI identity, never list position: index N in KFD topology is ROCm
# physical device N and carries its PCI address, which DRM sysfs keys on too.
# The two gates below verify ``index`` really is a host-physical ordinal
# (torch exposes no PCI id to check directly):
# * No visibility mask -- any mask makes ``index`` container/ROCR-relative
# rather than a host ordinal.
# * Device count == host GPU count -- rules out a device-cgroup container
# that sets no env var yet compacts torch's indices from zero.
pci_by_ordinal = _rocm_kfd_gpu_pci_ids()
if not pci_by_ordinal:
return
if _rocm_visibility_mask_active() or len(devices) != len(pci_by_ordinal):
return
vram_by_pci = _rocm_linux_sysfs_vram_by_pci_gb()
for dev in devices:
index = dev.get("index")
if not isinstance(index, int) or not (0 <= index < len(pci_by_ordinal)):
continue
entry = vram_by_pci.get(pci_by_ordinal[index].lower())
if entry is None:
continue
used, total = entry
dev_total = dev.get("vram_total_gb") or 0.0
# Overlay only a device that maps 1:1 to the whole card: torch total must
# match sysfs total within ~10%. A mismatch either way means a different
# memory scope -- a unified-memory APU (sysfs sees only the dedicated
# slice, torch the GTT pool) or a partitioned MI300 (sysfs reports the
# whole card, dwarfing a partition) -- and overlaying would misstate free
# VRAM (a partition would look like it has the whole card free).
if dev_total <= 0 or abs(total - dev_total) > 0.1 * dev_total:
continue
dev["vram_used_gb"] = used
dev["vram_total_gb"] = total
dev["vram_utilization_pct"] = round((used / total) * 100, 1) if total > 0 else None
def get_visible_gpu_utilization() -> Dict[str, Any]:
device = get_device()
@ -1317,6 +1521,12 @@ def get_visible_gpu_utilization() -> Dict[str, Any]:
"power_utilization_pct": None,
}
)
if IS_ROCM and index_kind == "physical":
# Swap process-local torch VRAM for system-wide sysfs so a model
# held by the separate llama-server process shows up (#7072).
# Physical-index only: a relative index (UUID/MIG mask) is not a
# host GPU id. The overlay verifies the rest itself.
_overlay_system_wide_vram(devices)
return {
"available": True,
"backend": _backend_label(device),