Text-to-video lands as a SIBLING of the image diffusion backend, not a mode of
it: video pipelines take frame/fps arguments, return frame stacks plus, for
LTX-2, synchronized audio, and persist MP4s -- none of the image module's
img2img/inpaint/ControlNet/LoRA surface applies. The image backend's hardware
and optimisation layers are imported unchanged (device/dtype resolution, memory
planning + offload tiers, attention backends, speed profiles, FBCache), and the
load-token/cancel-event concurrency skeleton is copied verbatim so lifecycle
behaviour cannot diverge.
core/inference/video_families.py: a pure VideoFamily registry (no torch) with
the ltx-2 entry -- LTX2Pipeline + LTX2VideoTransformer3DModel, base
Lightricks/LTX-2, unsloth/LTX-2.3-GGUF as the curated GGUF source, audio on,
frame lattice k*8+1, /32 resolutions with a vertical preset, and measured bf16
component sizes (the Gemma3-27B text encoder outweighs the 19B DiT itself).
MoE fields (transformer_2, guidance_scale_2) are declared now so the Wan2.2
A14B family lands later without churning the schema.
core/inference/video.py: VideoBackend with async begin_load + cache-scan
download progress, GGUF / single-file / full-pipeline loads (the GGUF DiT
assembles onto the base repo exactly like the image path), generation with
frame/size snapping BEFORE latents allocate, per-step progress + ETA and
cooperative cancel via the standard diffusers callback, and MP4 (H.264) export
through diffusers' PyAV encoder with the audio track muxed when the family
produces one. VAE tiling is always on: decoding a 100+ frame clip is the
memory peak, and the frames-aware estimate_video_runtime_mib (new, in
diffusion_memory) feeds the planner where the pixel-area image estimate would
badly undershoot. Loads are gated to unsloth/*, the official Lightricks base
repos, or local paths; PyAV availability is checked at load time so a missing
encoder cannot fail a clip after a multi-minute denoise.
core/inference/video_gallery.py: {id}.mp4 + {id}.json recipe sidecar pairs
under studio_root()/videos (an MP4 has no PNG text chunk to embed the recipe
in), with the image gallery's id/containment guards, newest-first listing that
skips orphans, delete/clear.
gpu_arbiter gains the VIDEO owner: ownership is exclusive, so the existing
evict-the-current-owner already generalises to chat/image/video all evicting
each other. The av (PyAV) dependency joins requirements/studio.txt.
Tests: video family detection/snapping/defaults, backend lifecycle on a faked
torch/diffusers runtime (GGUF assembly, shape snapping, distilled defaults,
cancel/progress, sentinel), gallery roundtrip/containment/orphans. 52 new
tests green plus the arbiter suite.
98 lines
3.5 KiB
Python
98 lines
3.5 KiB
Python
# SPDX-License-Identifier: AGPL-3.0-only
|
|
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
|
|
|
"""Single-GPU arbiter for Studio's two heavy GPU consumers.
|
|
|
|
The chat backends (llama-server + the Unsloth subprocess) and the diffusion
|
|
backend share one GPU. Before either takes the GPU it calls ``acquire_for(owner)``,
|
|
which evicts the current *other* owner so two large models never sit in VRAM at
|
|
once. The arbiter only sequences ownership; the actual freeing is delegated to
|
|
each backend's existing teardown.
|
|
|
|
Eviction runs under the arbiter lock, so an ownership transfer is atomic with
|
|
respect to other acquires.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import threading
|
|
from typing import Optional
|
|
|
|
from loggers import get_logger
|
|
|
|
logger = get_logger(__name__)
|
|
|
|
CHAT = "chat"
|
|
DIFFUSION = "diffusion"
|
|
VIDEO = "video"
|
|
|
|
_lock = threading.Lock()
|
|
_owner: Optional[str] = None
|
|
|
|
|
|
def _evict_chat() -> None:
|
|
import time
|
|
|
|
from core.inference import get_inference_backend
|
|
from routes.inference import get_llama_cpp_backend
|
|
|
|
llama = get_llama_cpp_backend()
|
|
# is_active (process exists), not is_loaded (process exists AND healthy): a
|
|
# chat model still starting up holds/keeps allocating VRAM but isn't healthy
|
|
# yet, so gating on is_loaded would skip it and let the load race the
|
|
# diffusion pipeline. unload_model() sets _cancel_event and kills the process.
|
|
if llama.is_active:
|
|
llama.unload_model()
|
|
orchestrator = get_inference_backend()
|
|
if orchestrator.active_model_name:
|
|
orchestrator.unload_model(orchestrator.active_model_name)
|
|
# Kill the subprocess too, not just the model: its base CUDA context holds
|
|
# VRAM the diffusion pipeline needs.
|
|
orchestrator._shutdown_subprocess(timeout = 5.0)
|
|
# The driver reclaims the killed process's VRAM asynchronously; wait for free
|
|
# memory to settle before the diffusion pipeline allocates, mirroring the chat
|
|
# reload path — otherwise a warm chat→diffusion handoff can transiently OOM.
|
|
llama._wait_for_vram_settle(since_kill = time.monotonic())
|
|
|
|
|
|
def _evict_diffusion() -> None:
|
|
# Unload whichever engine the router has active (diffusers or native sd.cpp), so a
|
|
# chat acquire frees the right one.
|
|
from core.inference.diffusion_engine_router import get_active_diffusion_engine
|
|
get_active_diffusion_engine().unload()
|
|
|
|
|
|
def _evict_video() -> None:
|
|
from core.inference.video import get_video_backend
|
|
get_video_backend().unload()
|
|
|
|
|
|
# Patchable in tests via monkeypatch.setitem. Ownership is exclusive -- only one
|
|
# owner holds the GPU at a time -- so acquire_for's evict-the-current-owner
|
|
# already generalises to any number of registered owners (chat / image / video
|
|
# all evict whichever of the others currently holds the GPU).
|
|
_EVICTORS = {CHAT: _evict_chat, DIFFUSION: _evict_diffusion, VIDEO: _evict_video}
|
|
|
|
|
|
def acquire_for(owner: str) -> None:
|
|
"""Make ``owner`` the sole GPU owner, evicting the other if it holds it."""
|
|
global _owner
|
|
if owner not in _EVICTORS:
|
|
raise ValueError(f"unknown GPU owner: {owner!r}")
|
|
with _lock:
|
|
if _owner is not None and _owner != owner:
|
|
logger.info("gpu_arbiter: evicting %s for %s", _owner, owner)
|
|
_EVICTORS[_owner]()
|
|
_owner = owner
|
|
|
|
|
|
def release(owner: str) -> None:
|
|
"""Drop ``owner``'s claim (no-op if it isn't the current owner)."""
|
|
global _owner
|
|
with _lock:
|
|
if _owner == owner:
|
|
_owner = None
|
|
|
|
|
|
def current_owner() -> Optional[str]:
|
|
return _owner
|