Bound the gallery blob cache, and three interlock fixes

Four review findings, all reproduced first:

- The gallery object-URL caches were unbounded. A clip runs from a few MB to a
  few hundred, both pages stay mounted after their first visit, and entries were
  only dropped on delete, so scrolling pinned everything for the session. Both
  pages now share a byte-budgeted LRU (512 MB video / 192 MB images) keyed off
  the visibility signal the near-viewport fetching already provides. On-screen
  media, the selected clip or image, and the item just fetched are never
  evicted, so eviction is invisible and a single item larger than the whole
  budget cannot evict itself into a refetch loop.

- The image, video and chat load guards ran two independent training probes but
  returned early when the FIRST one raised, so an unreadable LLM backend
  disabled the diffusion interlock and a load could proceed straight into an
  active diffusion trainer on the same GPU. The probes are independent now.

- An engine switch swallowed a failed teardown and published the new engine
  anyway, which is exactly the leak the unload exists to prevent: the arbiter's
  evictor, /images/unload and the next load all resolve through
  get_active_diffusion_engine(), so the still-resident pipeline (or a live
  sd-server) became unreachable and the next load allocated on top of it. The
  switch now fails and leaves the old engine published, so it stays reclaimable.

- The native generation timeout was 30 minutes while the Images page waits up to
  6 hours (SETTLE_MAX_MS), so slow-but-progressing CPU jobs died deterministically
  at the deadline. Measured on GPU-less runners, a 512x512 4-step Q2_K generation
  took 900 s on Linux and 1465 s on Windows, so larger images or step counts clear
  half an hour easily. The ceiling now matches the page's window and applies to
  the whole request: chunks of a split batch share one deadline instead of each
  getting a full budget. Cancellation is unchanged.

Declined: gating the huggingfacenotorch extra off Python 3.9 over the
conditional diffusers marker. The marker is deliberate and its comment says why:
diffusers dropped 3.9 in 0.38, so pinning >=0.39 outright leaves pip no candidate
and the whole extra unresolvable there. The pipelines it names live in
studio/backend, which cannot install on 3.9 anyway (studio.txt pins
matplotlib==3.10.9 and fastmcp>=3.0.2, both requires_python >=3.10), and the
extra is the general core one, so the alternative drops 3.9 for library users who
never touch Studio.
This commit is contained in:
Daniel Han 2026-07-27 05:08:45 +00:00
commit 3f6057a2b2
16 changed files with 359 additions and 50 deletions

View file

@ -10,6 +10,7 @@ PNG -- no real ``sd-cli``, no GPU.
from __future__ import annotations
import inspect
import os
import sys
import time
@ -491,3 +492,21 @@ def test_routing_prefer_native_overrides_gpu():
select_diffusion_engine("cuda", native_available = False, prefer_native = True)
== ENGINE_DIFFUSERS
)
def test_native_generation_timeout_matches_the_ui_settle_window():
# The native engine exists for slow CPU hosts: measured on GPU-less CI runners a 512x512 4-step
# Q2_K generation took 900 s (Linux) and 1465 s (Windows), so the old 30-minute default killed
# still-progressing jobs at higher resolutions or step counts while the Images page waited hours
# for them. The ceiling now matches that page's SETTLE_MAX_MS.
from core.inference.sd_cpp_engine import NATIVE_GENERATION_TIMEOUT_S, SdCppEngine
from core.inference import sd_cpp_backend
assert NATIVE_GENERATION_TIMEOUT_S == 6 * 60 * 60
for fn in (SdCppEngine.generate, SdCppEngine.upscale):
assert (
inspect.signature(fn).parameters["timeout"].default == NATIVE_GENERATION_TIMEOUT_S
), fn.__name__
# The resident-server path shares the same ceiling (applied per request, see
# test_server_generate_splits_batches_above_server_limit).
assert sd_cpp_backend.NATIVE_GENERATION_TIMEOUT_S == NATIVE_GENERATION_TIMEOUT_S