Four review findings, all reproduced first:
- The gallery object-URL caches were unbounded. A clip runs from a few MB to a
few hundred, both pages stay mounted after their first visit, and entries were
only dropped on delete, so scrolling pinned everything for the session. Both
pages now share a byte-budgeted LRU (512 MB video / 192 MB images) keyed off
the visibility signal the near-viewport fetching already provides. On-screen
media, the selected clip or image, and the item just fetched are never
evicted, so eviction is invisible and a single item larger than the whole
budget cannot evict itself into a refetch loop.
- The image, video and chat load guards ran two independent training probes but
returned early when the FIRST one raised, so an unreadable LLM backend
disabled the diffusion interlock and a load could proceed straight into an
active diffusion trainer on the same GPU. The probes are independent now.
- An engine switch swallowed a failed teardown and published the new engine
anyway, which is exactly the leak the unload exists to prevent: the arbiter's
evictor, /images/unload and the next load all resolve through
get_active_diffusion_engine(), so the still-resident pipeline (or a live
sd-server) became unreachable and the next load allocated on top of it. The
switch now fails and leaves the old engine published, so it stays reclaimable.
- The native generation timeout was 30 minutes while the Images page waits up to
6 hours (SETTLE_MAX_MS), so slow-but-progressing CPU jobs died deterministically
at the deadline. Measured on GPU-less runners, a 512x512 4-step Q2_K generation
took 900 s on Linux and 1465 s on Windows, so larger images or step counts clear
half an hour easily. The ceiling now matches the page's window and applies to
the whole request: chunks of a split batch share one deadline instead of each
getting a full budget. Cancellation is unchanged.
Declined: gating the huggingfacenotorch extra off Python 3.9 over the
conditional diffusers marker. The marker is deliberate and its comment says why:
diffusers dropped 3.9 in 0.38, so pinning >=0.39 outright leaves pip no candidate
and the whole extra unresolvable there. The pipelines it names live in
studio/backend, which cannot install on 3.9 anyway (studio.txt pins
matplotlib==3.10.9 and fastmcp>=3.0.2, both requires_python >=3.10), and the
extra is the general core one, so the alternative drops 3.9 for library users who
never touch Studio.