* fix(rocm): stop overwriting ROCR_VISIBLE_DEVICES in apply_gpu_ids ROCR_VISIBLE_DEVICES uses HSA agent-level indexing, not physical GPU indices. Setting it to a bare integer breaks multi-GPU ROCm systems where the parent already set ROCR_VISIBLE_DEVICES=0,1: narrowing to 1 causes torch.cuda.is_available() to return False in the training worker, producing a misleading 'no HIP accelerator' error even on a correctly configured ROCm host. HIP_VISIBLE_DEVICES is sufficient for GPU selection on ROCm. Leave ROCR_VISIBLE_DEVICES inherited from the parent environment. * test(rocm): update apply_gpu_ids test to assert ROCR_VISIBLE_DEVICES is not overwritten |
||
|---|---|---|
| .. | ||
| assets | ||
| auth | ||
| core | ||
| hub | ||
| loggers | ||
| models | ||
| plugins | ||
| requirements | ||
| routes | ||
| state | ||
| storage | ||
| tests | ||
| utils | ||
| __init__.py | ||
| _platform_compat.py | ||
| cloudflare_tunnel.py | ||
| colab.py | ||
| main.py | ||
| run.py | ||
| startup_banner.py | ||