Six review findings, three of them evict-then-fail orderings: - The chat load reclaimed the GPU without telling the arbiter it existed. A chat load holds no llama-server process until its GGUF has downloaded, which is minutes, so a competing Images/Video acquire in that window found nothing to cancel, took the GPU, and the chat load then spawned onto the same device. It now registers an in-flight marker through acquire_for's register hook (under the arbiter lock, as the image and video loads do), the evictor cancels a marked load, and the route undoes itself if ownership moved while it loaded. - The Hub-download conflict check ran after that handoff, so a GGUF the download manager already owns destroyed the resident Images/Video pipeline and then 409'd, having loaded nothing. It moves above the handoff, together with the marker it handshakes with. - The image load released the engine router's transition lock before registering the load, so a second load choosing the other engine could unload the still-idle engine this one captured; the load then landed on a deactivated engine, where generate, status, unload and the arbiter's evictor can no longer reach it. Registration now happens under that lock and refuses if the engine changed. - Training a DiT family on a host with no GPU was accepted: nf4 is not a CPU fallback, its 4-bit load goes through bitsandbytes, which requires CUDA, XPU or MPS. The start unloaded the working Images pipeline, pulled the text encoders, and only then died in the child. Rejected before the teardown now, and /info stops advertising a precision that always 400s. SDXL keeps its documented fp32-on-CPU path. - Both diffusion pages kept the routed-pick marker forever, so re-picking the same checkpoint (after chat evicted it) neither loaded nor cleared the query string. The marker is released once the query is gone. The Images key also carried a stray NUL byte, which made the file read as binary to grep and other tooling. - diffusers dropped Python 3.9 in 0.38, so the unconditional >=0.39.0 pin left pip no candidate at all on 3.9 and made every install that composes the huggingface extras unresolvable there. The floor is conditional now. Also fixes tests that were already red on the branch: two hand-built request fakes had gone stale against fields this branch added, and the handoff-ordering test only failed on a host with fewer than two GPUs.
116 lines
4.6 KiB
Python
116 lines
4.6 KiB
Python
# SPDX-License-Identifier: AGPL-3.0-only
|
|
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
|
|
|
"""Single-GPU arbiter for Studio's heavy GPU consumers.
|
|
|
|
The chat backends, diffusion, and video share one GPU. Before taking it each calls
|
|
``acquire_for(owner)``, which evicts the current other owner so two large models never sit in VRAM
|
|
at once. The arbiter only sequences ownership (freeing is each backend's teardown); eviction runs
|
|
under the lock, so a transfer is atomic vs other acquires.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import threading
|
|
from typing import Any, Callable, Optional
|
|
|
|
from loggers import get_logger
|
|
|
|
logger = get_logger(__name__)
|
|
|
|
CHAT = "chat"
|
|
DIFFUSION = "diffusion"
|
|
VIDEO = "video"
|
|
|
|
_lock = threading.Lock()
|
|
_owner: Optional[str] = None
|
|
|
|
|
|
def _evict_chat() -> None:
|
|
import time
|
|
|
|
from core.inference import get_inference_backend
|
|
from routes.inference import get_llama_cpp_backend
|
|
|
|
from core.inference.llama_cpp import chat_load_active
|
|
|
|
llama = get_llama_cpp_backend()
|
|
# is_active (process exists), not is_loaded (exists AND healthy): a chat model still starting
|
|
# up holds VRAM but isn't healthy, so is_loaded would skip it and let the load race diffusion.
|
|
# chat_load_active too: an HF load has no process until its GGUF downloaded, so is_active
|
|
# alone found nothing to cancel and let that load spawn onto the GPU we just granted away.
|
|
# unload_model sets the cancel event the download loop polls, so the pending load aborts.
|
|
if llama.is_active or chat_load_active():
|
|
llama.unload_model()
|
|
orchestrator = get_inference_backend()
|
|
if orchestrator.active_model_name:
|
|
orchestrator.unload_model(orchestrator.active_model_name)
|
|
# Kill the subprocess too: its base CUDA context holds VRAM diffusion needs.
|
|
orchestrator._shutdown_subprocess(timeout = 5.0)
|
|
# The driver reclaims the killed VRAM asynchronously; wait for it to settle before diffusion
|
|
# allocates, else a warm chat->diffusion handoff can transiently OOM.
|
|
llama._wait_for_vram_settle(since_kill = time.monotonic())
|
|
|
|
|
|
def _evict_diffusion() -> None:
|
|
# Unload whichever engine the router has active (diffusers or native sd.cpp), so a
|
|
# chat acquire frees the right one.
|
|
from core.inference.diffusion_engine_router import get_active_diffusion_engine
|
|
get_active_diffusion_engine().unload()
|
|
|
|
|
|
def _evict_video() -> None:
|
|
from core.inference.video import get_video_backend
|
|
get_video_backend().unload()
|
|
|
|
|
|
# Patchable in tests via monkeypatch.setitem. Ownership is exclusive, so acquire_for's
|
|
# evict-the-current-owner generalises to any number of registered owners.
|
|
_EVICTORS = {CHAT: _evict_chat, DIFFUSION: _evict_diffusion, VIDEO: _evict_video}
|
|
|
|
|
|
def acquire_for(owner: str, register: Optional[Callable[[], Any]] = None) -> Any:
|
|
"""Make ``owner`` the sole GPU owner, evicting the other if it holds it.
|
|
|
|
``register``, if given, runs under the arbiter lock right after ownership transfers and its
|
|
return value is returned. Marking the in-flight load HERE (not after ``acquire_for`` returns)
|
|
closes the window where a competing acquire could evict this owner before its load is in-flight,
|
|
letting both loaders allocate VRAM at once. It must be quick and not re-enter the arbiter; if it
|
|
raises, ownership stays with ``owner``.
|
|
"""
|
|
global _owner
|
|
if owner not in _EVICTORS:
|
|
raise ValueError(f"unknown GPU owner: {owner!r}")
|
|
with _lock:
|
|
if _owner is not None and _owner != owner:
|
|
logger.info("gpu_arbiter: evicting %s for %s", _owner, owner)
|
|
_EVICTORS[_owner]()
|
|
_owner = owner
|
|
return register() if register is not None else None
|
|
|
|
|
|
def release(owner: str) -> None:
|
|
"""Drop ``owner``'s claim (no-op if it isn't the current owner)."""
|
|
global _owner
|
|
with _lock:
|
|
if _owner == owner:
|
|
_owner = None
|
|
|
|
|
|
def release_if(owner: str, predicate: Callable[[], bool]) -> bool:
|
|
"""Drop ``owner``'s claim only if it still holds it AND ``predicate()`` is true, atomically.
|
|
|
|
A slow unload's idle check and its ``release`` must not straddle a concurrent same-owner load
|
|
whose ``acquire_for(register=...)`` re-registers ownership under this lock; evaluating the
|
|
predicate under the lock keeps them atomic so ``release`` never clears the newer claim.
|
|
``predicate`` must be quick and not re-enter the arbiter. Returns True iff ownership was dropped."""
|
|
global _owner
|
|
with _lock:
|
|
if _owner != owner or not predicate():
|
|
return False
|
|
_owner = None
|
|
return True
|
|
|
|
|
|
def current_owner() -> Optional[str]:
|
|
return _owner
|