unsloth/studio/backend/core/inference
Guerriero Riccardo c267895538
Studio: mask AMD GPU pins via ROCR so an unsupported iGPU can't crash llama-server (#7272)
* Studio: mask AMD GPU pins via ROCR so an unsupported iGPU can't crash llama-server

On a mixed AMD host (e.g. a discrete gfx1102 GPU next to a gfx1103 iGPU)
the bundled rocm-gfx110X llama.cpp build segfaults during HSA device
enumeration on the unsupported iGPU -- before llama-server prints a line,
so every model load fails with a bare signal and empty logs.

The GPU-subset pin masked visibility with HIP_VISIBLE_DEVICES, but HIP
filtering runs only after the HSA runtime has already enumerated (and
crashed on) every agent. Mask the subset via ROCR_VISIBLE_DEVICES (the
ROCr/HSA layer) instead, so a deselected/unsupported GPU is never
enumerated. Exactly one layer is masked (HIP cleared) to avoid the
double-mask reindex that would otherwise drop the child to CPU. The
whole-set tensor-split path and the CPU-only sentinel keep their existing
HIP behavior.

Also stop misreporting the resulting startup segfault as a vision
projector incompatibility: when the text-only mmproj retry also hard-
crashes with a signal, surface a GPU/driver init crash (with the ROCR
hint) instead of blaming the projector.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Tighten _emit_child_gpu_visibility comments for #7272

Comment and docstring only: condense the ROCR-vs-HIP masking rationale from ~22 to ~14 lines and the call-site note from 5 to 3, keeping every technical point (HSA enumeration segfault, physical ids, the -1 sentinel). Logic is unchanged, verified by an AST compare with docstrings stripped and by exercising _emit_child_gpu_visibility against a torch/HIP stub.

* Detect AMD SDK ROCm wheels (hip=None) in _emit_child_gpu_visibility (Codex P2)

The ROCm branch gated only on torch.version.hip, but AMD SDK wheels leave that unset while encoding 'rocm' in __version__ (detect_hardware handles this the same way). On such a wheel the masking was skipped entirely, leaving only CUDA_VISIBLE_DEVICES, so on a mixed AMD box the unsupported deselected iGPU still enumerated and could crash llama-server. Now the branch also treats 'rocm' in torch.__version__ as ROCm, mirroring detect_hardware. Adds tests for the hip=None SDK wheel (ROCR + default paths) and a CUDA guard so the version-string check can't false-positive.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Remap CUDA_VISIBLE_DEVICES to post-ROCR ordinals on prefer_rocr for PR #7272 (Codex P1)

On the prefer_rocr path _emit_child_gpu_visibility set ROCR_VISIBLE_DEVICES to the
physical id and cleared HIP_VISIBLE_DEVICES, but left CUDA_VISIBLE_DEVICES at the
physical id. ROCR re-indexes the visible agents from 0 and, with HIP cleared, HIP
honours CUDA_VISIBLE_DEVICES -- so a non-zero pick (e.g. GPU 1) pointed out of
range, HIP saw 0 devices, and the child fell back to CPU, defeating GPU-picker
selections other than physical GPU 0. Remap CUDA to the post-ROCR ordinals
(0..N-1); GPU 0 is unchanged, the default (HIP) path and the CPU sentinel are
untouched, and non-AMD wheels never enter this branch.

* Detect AMD SDK wheels in _resolve_visible_physical_ids for PR #7272 (Codex P2)

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Keep the HIP mask on Windows ROCm in prefer_rocr for PR #7272 (Codex P2)

* Ignore ROCR_VISIBLE_DEVICES in _resolve_visible_physical_ids on Windows for PR #7272 (Codex P2)

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Preserve inherited ROCR masks in the tensor-split pin for PR #7272 (Codex P2)

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Co-authored-by: Leo Borcherding <borchborchmail@gmail.com>
2026-07-22 20:16:25 -05:00
..
sandbox_site Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
__init__.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
_html_to_md.py Studio: stream live tool output with SSE heartbeats, fix web page extraction, and surface interrupted turns (#7083) 2026-07-15 08:41:00 -07:00
_vulkan_probe.py Studio: add Vulkan llama.cpp support (#5819) 2026-07-09 03:39:48 -07:00
anthropic_compat.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
api_monitor.py Studio: trim serving-log noise and surface llama-server engine stats (#6377) 2026-06-17 05:37:57 -07:00
audio_codecs.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
chat_eos.py Studio: stop chat generation on the assistant-turn-end token (fixes Qwen3.5 loop) (#6804) 2026-07-06 10:07:56 -07:00
chat_template_helpers.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
chat_templates.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
defaults.py Add DeepSeek-V4-Flash-GGUF to Studio with none/high/max reasoning (#6908) 2026-07-07 06:13:43 -07:00
external_provider.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
inference.py fix(studio): ignore reasoning in tool reprompts (#7134) 2026-07-17 20:22:11 -03:00
key_exchange.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
llama_admission.py Studio: queue local GGUF OpenAI-compatible requests before llama-server (#7047) 2026-07-10 17:05:48 -03:00
llama_cpp.py Studio: mask AMD GPU pins via ROCR so an unsupported iGPU can't crash llama-server (#7272) 2026-07-22 20:16:25 -05:00
llama_http.py fix(studio/llama_cpp): disable trust_env on the loopback health probe (#6750) (#6752) 2026-06-30 19:09:26 +02:00
llama_keepwarm.py persist llama.cpp KV cache across idle auto-unload (slot save/restore) (#7204) 2026-07-20 00:12:42 -07:00
llama_server_args.py persist llama.cpp KV cache across idle auto-unload (slot save/restore) (#7204) 2026-07-20 00:12:42 -07:00
llama_stats.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
local_model_resolver.py Studio: hide the RAG embedder and llama.cpp probe from the hub cached inventory (#7018) 2026-07-19 03:20:56 -07:00
mcp_client.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
mcp_config_import.py studio: show MCP "Import config" on the add-server form (#6030) 2026-06-11 16:17:22 +01:00
message_content.py fix(studio): handle multimodal list content in inference text paths (#4383) (#6480) 2026-06-23 01:26:11 -07:00
mlx_inference.py Studio: reuse MLX prompt cache across turns instead of re-prefilling (#7311) 2026-07-22 02:35:33 -07:00
model_ids.py studio: list the full local model catalog from /v1/models (#6519) 2026-06-26 20:42:06 -03:00
orchestrator.py Fix local CLI streamed generation error handling (#7135) 2026-07-20 23:14:58 -03:00
passthrough_healing.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
presence_penalty.py Studio: apply presence_penalty on the safetensors and MLX inference paths (#6923) 2026-07-06 22:24:47 -07:00
pricing.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
providers.py Allow API key for Ollama connections (#7173) 2026-07-18 22:47:00 -07:00
runtime_context.py Expose runtime context length for hub models (#6154) 2026-06-11 22:13:53 +03:00
safetensors_agentic.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
tensor_fallback.py studio: deterministic VRAM auto-fit for GGUF (MTP reserve, compute buffer, total-based budget) (#6312) 2026-06-17 03:10:22 -07:00
tool_call_parser.py Studio: Inkling support fixes (#7153) 2026-07-15 11:22:38 -07:00
tool_loop_controller.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
tool_stream_exec.py Studio: stream live tool output with SSE heartbeats, fix web page extraction, and surface interrupted turns (#7083) 2026-07-15 08:41:00 -07:00
tools.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
worker.py fix(studio): honor MLX adapter state in compare mode (#7196) 2026-07-19 00:27:55 -07:00