img2img and inpaint take their output size from the uploaded image and only
snap it to a multiple of 16, so an ordinary phone photo (up to the 4096/side
decode cap, 4x the txt2img 2048 ceiling and ~16x the area) drove an OOM-scale
latent and an opaque 500 on a normal card, while txt2img, upscale, edit, and
FLUX.2-klein inpaint are all already megapixel-bounded. Clamp the init longest
side to 2048 (the txt2img ceiling) before deriving width/height; edit is exempt
since its pipeline resizes to ~1MP internally.
_cast_nvfp4 quantized every nn.Linear with no filter, unlike the int8 and fp8
torchao text-encoder modes which exclude the VLM vision tower / lm_head / T5 wo.
On qwen-image / qwen-image-edit that 4-bit quantized the Qwen2.5-VL image tower,
degrading the edit/image conditioning the sibling schemes protect. Apply the
same make_filter_fn exclusion (require_bf16, mirroring _cast_fp8_dynamic).
Two evict/OOM fixes on the diffusion load paths:
- The video load moved a pipeline onto the GPU (apply_memory_plan) and
committed it while holding no lock, so an unload / GPU-arbiter eviction --
which bumps the load token and then barriers on _generate_lock before
freeing -- could hand VIDEO to chat/images and let the new owner allocate
concurrently with the in-flight placement, OOMing. Hold _generate_lock
across placement + the locked commit, mirroring the image backend, so an
evicting owner waits until this worker's placement is torn down or
committed. Lock order stays _generate_lock -> _lock (unload takes _lock
then releases it before the barrier), so there is no deadlock.
- resolve_local_single_file reinterpreted an On-Device folder as a base
single_file load whenever it held exactly one .safetensors, so a PEFT LoRA
adapter folder (adapter_config.json + adapter_model.safetensors) with a
family-token name was picked as a base checkpoint, evicting the resident
model before from_single_file failed on the adapter weights. Skip adapter
folders (adapter_config.json) and the adapter_model basename so the pick
stays a pipeline load and 400s in validation, before the GPU handoff.
Adds regression tests for both.
Fold PR #6872's image-generation fixes into the branch, deduped against the
round-12 dataset-upload and gallery integrity work already on image-generation.
Fixes carried forward from #6872:
- fp8 single-file transformer memory estimate: an fp8 checkpoint loads with no
quantization_config and diffusers upcasts it to bf16 (~2x resident), so budget
it accordingly in _plan_memory and estimate_safetensors_dense_mib.
- dense-quant OOM-evict preflight: when the GGUF fits resident but the dense bf16
transformer this path materializes does not, skip the fast path up front rather
than evict the current pipeline and OOM in finalization. Combined with the
existing offload->resident candidate re-plan so both the family-table estimate
and the on-disk shard measurement gate engagement (unified on the
transformer_resident_override_mib plan override).
- ControlNet: evict the previous module and its from_pipe wrapper before loading a
new one so swapping ControlNets within a base-model load cannot accumulate to OOM.
- ControlNet union_control_mode: raise on an unknown control type instead of
silently defaulting to canny.
- edit-family mask rejection: raise instead of silently dropping a mask on an
image-editing model that has no inpaint pipeline.
- companion cache: walk the snapshot dir and exclude transformer/ so the
dense-quant prefetch's cached shards do not inflate the companion total and
wrongly force offload.
- training: drop piecewise_constant from the LR scheduler enum and force bf16 for
fp16-incompatible families.
- dataset upload: batch-atomic staging with the same-stem duplicate guard.
- images page: guard negative-prompt restore on guidance>0, clear stale ControlNet
selection on restore, and revert an optimistic quant label when a pipeline load
never starts.
- uninstall (sh + ps1): keep the owner-marker guard on sd.cpp removal.
Conflicts resolved in favour of image-generation's evolved memory system,
loadSpecFor catalog, and stop-and-save (lora_path) run detection; #6872's fp8 and
dense-preflight fixes carried forward on top. All affected backend tests pass
(test_diffusion_backend, test_diffusion_training, test_diffusion_lora_trainer,
test_video_gallery, test_diffusion_controlnet).
Several image/video/training preflights ran before the route acquires the GPU or
frees resident models, but let a doomed local pick through and only failed deep in
the background load, after the user's chat/Images/Video model was already evicted.
- Local base_repo / base_model: _is_trusted_diffusion_repo accepts any existing
local path, but the base loads via from_pretrained (needs model_index.json). A
local dir that is not a diffusers pipeline passed the trust gate, evicted the
resident model, then failed. Add a shared _assert_local_base_is_pipeline check
and call it in the image, video, and training preflights.
- Dataset images: discover_image_caption_pairs only checked filenames, so a
corrupt or zero-byte upload passed the start-route preflight, freed the GPU, then
crashed the spawned trainer in PIL. Add an opt-in verify_images decode probe
(cheap PIL header check) that the start route enables; the trainers leave it off
since they decode every image anyway.
- Local single-file safetensors: the On-Device scanner advertises a bare
.safetensors directory (no model_index.json) as a text-to-image model, but the
picker starts it as a pipeline with no filename, so every click 400s. Reinterpret
such a pick as a single_file load of the sole checkpoint (resolve_local_single_file)
so the advertised model is actually loadable.
Regression tests for each: local non-pipeline base (image/video/training), the
verify_images decode gate, and resolve_local_single_file.
_reset_step_cache looked up reset_stateful_hooks on the transformer, but on a
diffusers CacheMixin transformer (Flux, QwenImage) that method lives only on the
HookRegistry; the transformer-level entry point is _reset_stateful_cache. So with
FBCache engaged on an image model the reset was a silent no-op, and the next
generation reused the previous request's first-block residual: a tensor-shape
mismatch (crash) when the resolution or batch changed, or stale cached output
otherwise. Prefer _reset_stateful_cache and fall back to reset_stateful_hooks,
matching the video backend. Update the tests to the real hook name.
- The companion base for a GGUF/single-file image load is resolved from the GGUF
repo's base_model card tag when no base_repo is passed, and that value loads via
from_pretrained. The explicit base_repo is already trust-gated, but the card tag is
attacker-controlled metadata on any remote repo, so it now clears the same
unsloth/allowlist/local trust bar; an untrusted tag is dropped in favour of the
curated family default and never reaches from_pretrained. This closes a pickle
deserialization vector on the normal GGUF load path (a user loading an attacker's
GGUF repo whose card points base_model at a malicious pipeline), matching the
trust discipline the ControlNet path already applies via evaluate_file_security.
The allowlist already contains every legitimate variant base, so variant
resolution for the supported unsloth GGUFs is unchanged.
- The images/unload route ran the slow VRAM-freeing unload on a thread and then
released the DIFFUSION arbiter owner unconditionally. release() is owner-guarded
and identity-less, so a concurrent /images/load that re-acquired DIFFUSION while
the unload ran would have its ownership cleared by the trailing release, and a
later chat load would then see no owner, skip eviction, and OOM against the newly
resident pipeline. The route now releases only when nothing is resident again.
- _apply_group_offload placed the resident companions before attaching the
transformer's group-offload hooks so a companion OOM returns with no hooks and the
whole-module fallback stays valid, but the streamed loop itself installs hooks on
each DiT in turn. On a dual-DiT pipeline where the second tower failed after the
first got its hooks, it returned False with hooks already installed, and the
caller's enable_model_cpu_offload fallback then crashed (diffusers rejects it on a
partially group-offloaded pipe). It now propagates the real failure once any hook
is installed, so the load fails with its actual cause instead of a misleading crash.
Adds regression tests: the untrusted card tag dropped to the family default (trusted
tag still honoured, explicit base still wins), unload keeping ownership when a model
is still resident (and releasing when not), and the partial dual-DiT hook set
propagating rather than falling through to a crashing whole-module offload.
- Image load now trust-gates a client-supplied base_repo. validate_load_request
rejects a base_repo that is not an unsloth/* repo, an allowlisted official base, or a
local path, mirroring the repo_id gate and the video loader. The route passes
base_repo into that pre-eviction validation, so an authenticated client can no longer
keep model_path on a trusted GGUF while pointing base_repo at an arbitrary remote repo
that the server would download and deserialize (a from_pretrained pickle/config path),
and no resident model is evicted for the rejected load.
- The keepwarm middleware now tracks the image and video generation routes
(/images/generate, /images/generations, /video/generate), so
other_inference_request_count() sees an in-flight generation and an API-key training
start is refused (409) before its unload would cancel that generation. endswith keeps
the GET *-progress and */cancel variants untracked.
- The OpenAI-compatible /v1 surface is now blanket body-capped like /api/inference,
instead of only /v1/chat/completions and /v1/completions. Every /v1 POST route
(images/generations, audio, embeddings, responses, messages, ...) buffers a JSON body
and none is a multipart-upload passthrough, so an unbounded ImageGenerationRequest
prompt on /v1/images/generations can no longer be buffered outside the request limit.
Adds regression tests: the base_repo trust gate at both the backend (untrusted remote
raises, local passes) and the route (untrusted base_repo returns 400 with no load), the
keepwarm tracking of the image/video generation paths (and not the progress/cancel
variants), and the /v1 surface being body-protected.
* Auto policies: deferred dense compile, video compile default, step cache and precision auto
Image dense loads with speed unset no longer sit at plain off: the load stays
bit-identical eager, and the 3rd generation in a session engages the default
compile profile plus the cuDNN attention upgrade mid-session (a one-off image
never pays the warmup, repeated use amortises it). Video dense loads resolve
straight to the default profile since a clip denoise amortises the compile
within a single run, and never to max.
Video also gains the image backend's tri-state auto policies: unset step cache
now decides from the default schedule and re-checks the actual step count per
generation, and unset precision (transformer_quant) hands the decision to the
hardware ladder instead of staying off. Memory badge reason now says plainly
that everything fits when no offload is planned.
* Rename Dtype to Precision, add the video Precision control, step cache Auto option
The images Advanced panel's Dtype row is now Precision (same control, clearer
name), and the video Advanced panel gains the matching Precision select wired
to the load route's existing transformer_quant field, gated to full-pipeline
loads the way the image control gates to GGUF. Step cache selects on both
pages gain an explicit Auto option as the default (the previous Off default
silently behaved as auto and never let anyone pin off), and the Speed and
Attention tooltips now state the deferred dense compile and the SageAttention
black-frame caveat.
* Model catalog: canonical diffusion model groups with device-aware routing
One canonical name per image/video model, its published artifacts (GGUF, FP8,
bnb-4bit, official BF16) as data, and pure routing helpers: suffix-stripped
canonical keys (owner-preserving; cross-owner merges only via explicit
aliases), group/artifact lookups, a flat back-compat options shim, load-spec
resolution replacing the pages' lookup tables, search matching over old ids
and format tokens, the GGUF fit ladder extracted from the variant expander,
and pickDefaultArtifact/pickDefaultQuant deciding what a bare group click
loads (downloaded first, then the best quality that fits 70 percent of VRAM,
GGUF as the safe fallback). Checked by npm run catalog:check, following the
i18n:check pattern.
* Picker: one canonical row per diffusion model with a format second level
The Images and Video pickers now render the curated catalog as one row per
model in Recommended: clicking loads the best artifact for the device (the
routed GGUF quant, a prequant FP8/bnb-4bit that fits, or the official BF16),
and a chevron opens the per-format list, with the GGUF row nesting the usual
quant expander. Live HF listing rows that belong to a group are deduplicated,
search collapses member repos into their group (old ids and format tokens
still match), and the On Device sections group cached member repos under the
same canonical name with the per-repo rows inside. Curated groups render from
the catalog rather than the HF listing, which finally surfaces LTX-2.3 in the
video Recommended list (its hub pipeline_tag is image-to-video, so the
text-to-video listing always missed it) and exposes the HunyuanVideo 720p
repack next to 480p.
Backend: /cached-models now tags trusted video-family repos text-to-video
instead of blanket text-to-image, and the pickers admit catalog-known
non-unsloth repos On Device, so cached Lightricks/Wan/Hunyuan pipelines
finally appear in the Video picker. Chat pickers pass no catalog and are
unchanged.
* Download formats, tab icons, plain-language train tips, 3-loop autoplay
The image Download button becomes a menu: PNG saves the original bytes with
the embedded recipe, JPEG and WebP re-encode client-side from the fetched
blob (JPEG flattened onto white). The video Download button gains MP4
(original, keeps audio), WebM and GIF; the latter two transcode server-side
from the stored MP4 via PyAV (VP9 realtime profile for WebM, ~12 fps adaptive
palette for GIF) behind a new gallery export route that 501s with a readable
message when a codec is missing.
Generated clips no longer loop forever: the player replays a clip three times
per selection, then pauses with controls up; a new generation or a refresh
gets its own three plays. The Create/Train tabs reuse the sidebar's New Chat
and Train icons (TestTubeOutlineIcon moved to a shared lib module), and every
Train tab helper text is now one plain sentence.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Keep the Create/Train tab icon and label on one line
TabsTrigger renders its children inside a plain inline span and the
Tailwind preflight gives svg display:block, so the HugeiconsIcon forced
the label onto a second line. Wrap icon plus label in their own
inline flex row inside each trigger.
* Strip -int8 and -nvfp4 prequant suffixes in the model catalog key
canonicalKeyFor already lowercases before matching, so -GGUF/-FP8 in any
case were covered; -int8 and -nvfp4 were not in the suffix table, so
such repos rendered as standalone rows in Recommended and On Device
instead of standardizing into their base-name group and routing through
pickDefaultArtifact. Added both suffixes plus case-insensitivity and
routing assertions to the catalog check.
* Standardize non-catalog picker rows to their base model name
The curated catalog already collapses its own groups, but hub listing
rows and cached repos outside the catalog (ERNIE-Image, FLUX.2-klein,
Qwen-Image-Edit-2509, FLUX.2-dev) still rendered raw ids with -GGUF /
-FP8 style suffixes in Recommended and On Device.
- model-catalog.ts: new stripArtifactSuffixesForDisplay, a
case-preserving twin of canonicalKeyFor's stripping that keeps the
owner prefix and original casing for display.
- pickers.tsx: recommended hub rows and the downloaded GGUF/model rows
pass their labels through it when a catalog is present, so only the
diffusion pickers change; chat rows keep raw ids. Click targets keep
the full repo id, and the format badge still shows the artifact kind.
- Catalog check covers the new helper across GGUF/FP8/int8/nvfp4 in
both cases plus no-op and suffix-only names.
* Offer official BF16/FP8 artifacts per model group and fix gallery label clipping
Model picker changes so groups are not limited to unsloth quant repos:
- model-catalog.ts: each image group that has an official vendor pipeline
now carries its BF16 (official) artifact as the top (highest quality)
entry - Tongyi-MAI/Z-Image-Turbo, Qwen/Qwen-Image, Qwen/Qwen-Image-2512,
Qwen/Qwen-Image-Edit-2511, black-forest-labs/FLUX.1-dev, FLUX.1-schnell
and FLUX.1-Kontext-dev. The LTX-2.3 video group now lists Lightricks'
own bf16 and fp8 distilled single-file checkpoints alongside the GGUF.
Resident sizes are set from the actual weight totals (FLUX ships a
duplicate single-file that from_pretrained ignores, so FLUX bf16 is ~32
GB not 54). The repos that used to be aliases are now real artifacts.
- The router already prefers the highest-quality artifact that fits the
0.7 x GPU budget, so a datacenter GPU now defaults to official BF16
while consumer GPUs still route to the fitting quant or GGUF. That is
why bnb-4bit was the Z-Image-Turbo default before: it was the only
non-GGUF artifact and it was already downloaded.
- diffusion.py: allowlist the four official image repos not previously
trusted (qwen/qwen-image-2512, qwen/qwen-image-edit-2511,
black-forest-labs/flux.1-schnell, flux.1-kontext-dev). All verified as
safetensors-only diffusers model_index pipelines. The LTX-2.3
checkpoints are already on the video trust list.
- catalog check: BF16-wins-on-datacenter, quant-wins-on-consumer, and the
single-file load specs for the LTX-2.3 checkpoints.
Also fixes the video gallery thumbnail caption: the leading duration was
clipped by the rounded corner and selection border, so the strip now has
enough left/bottom padding to clear the curve.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* video gallery: guard export transcode against a stream-less clip
_transcode_webm and _transcode_gif indexed src.streams.video[0] before
checking the stream list, so a container with no video stream raised a bare
IndexError that the broad handlers then re-labeled as a missing libvpx or
decoder. Raise an explicit RuntimeError naming the real cause in both the
WebM and GIF paths.
* Studio: honor explicit attention/format choices, fix distilled-LTX defaults and On Device catalog routing
* Remove stray planning notes accidentally committed to the branch
* video: add transformerQuant to the load callback deps
handleLoad reads transformerQuant but omitted it from the useCallback dep array,
so after the user changes only Precision and then selects a model or clicks
Reapply, the memoized callback keeps the stale closure and loads the previous
precision. The image page's equivalent callback already lists it.
* model picker: honor the format filter when routing catalog clicks; add catalog rows to the roving list
- routedArtifactFor now scopes a group's artifacts to the active format filter
(the same matchesFormatFilter predicate the visibility check uses) before
pickDefaultArtifact, so a group shown only because it owns a GGUF no longer
routes a click to a large non-GGUF download. Covers both the Recommended and
On Device grouped paths.
- hubOptionKeys now includes the catalog-group, search-catalog-group, and grouped
On Device row keys in exact render order, so arrow/Home/End roving reaches the
catalog rows instead of giving them a duplicate missing id and skipping them.
* model picker: don't treat a partial base cache as downloaded
A partially-cached base repo (a cancelled download that left only some weights)
was counted as downloaded, so an On Device click routed to a fresh multi-GB
re-download instead of the complete GGUF. The picker's endpoint (/api/models/
cached-models) did not carry a partial flag at all, so a frontend-only guard
could not see it. Surface partial from that endpoint by reusing the hub inventory
scan's snapshot-partial detector, plumb it through CachedModelRepo (backend +
frontend types), and skip partial base repos when building the downloaded set.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* model picker + diffusion: drop partial/unloadable cached rows, skip defer-compile before a LoRA gen
- On Device (cached non-GGUF) rows filtered partial-download snapshots back in: sortedCachedModels
gated on passesTaskGate + a groupForRepoId key match but, unlike downloadedSet, never checked
c.partial, so an incomplete unsloth snapshot showed as a loadable On Device row (click errors or
silently re-fetches multi-GB). It also admitted repos that only match the catalog by group KEY
(a base / uncurated-quant sibling like Qwen/Qwen-Image-2512) which have no loadable artifact and
dead-end at the trust gate. Add !c.partial and gate on artifactForRepoId (what loadSpecFor
resolves) instead of groupForRepoId, so a cached row shows only when the backend can load it.
- Deferred speed-auto engaged the compile profile on the 3rd generation BEFORE _apply_loras. A
compiled transformer rejects LoRA (supports_lora is False) and _apply_loras raises before its
unchanged-selection no-op, so once compile engaged every LoRA generation on that load failed
permanently. Skip the deferral when a LoRA is requested (compile and LoRA are mutually exclusive)
and let it engage on a later LoRA-free generation.
* Scope the cached-model partial probe to the listed snapshot dir
list_cached_models builds each row from the largest/complete copy across HF cache
roots, but _cached_repo_partial probed is_snapshot_partial with no repo_cache_dir,
so the scan spanned every root: a stale .incomplete copy in one root would flag a
complete copy in another as partial and hide the usable model from the picker (the
click then routes to a re-download). Forward the winning snapshot's repo_path so all
three partial signals are scoped to that copy, matching the sibling inventory paths
(models/dataset cache_inventory, local_inventory).
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Do not auto-route to gated repos, prefer complete cached copies, defer compile past attached LoRA, scope group expand keys
Four fixes:
- pickDefaultArtifact's not-downloaded ladder returned the gated BF16 FLUX.1-dev / Kontext-dev
before the open GGUF on a large GPU, so a bare group click routed to a repo the user may lack
license/token access to. Add a gated flag and skip gated artifacts in the not-downloaded ladder
(an already-downloaded gated artifact is still returned).
- list_cached_models picked the largest duplicate cache copy and computed partial only on it, so a
larger partial copy shadowed a smaller complete one; since partial rows are dropped from the
picker the usable model vanished. Prefer completeness, then size.
- the deferred-speed compile engaged on a no-LoRA generation while an adapter from a prior
generation was still attached, baking it into the compiled graph (the later unload is swallowed
on a compiled pipe); also defer while adapters remain attached.
- routeGroupClick's GGUF fallback toggled the context-free canonicalId while the chevron toggles
the context-scoped expandKey, leaving the format list un-collapsible in one context, dead in the
other, and risking cross-context expansion; thread expandKey through.
* Guard video pipeline repos from deletion, drop the always-failing LTX FP8 artifact, prefer 720p Hunyuan
Three round-6 fixes:
- cached non-GGUF video repos now surface in the Video On-Device picker with the normal delete
action, but /delete-cached only guarded chat + the Images engine, so a loaded/loading Wan / LTX /
Hunyuan pipeline could have its HF snapshot removed from under it. Add a VideoBackend
loading_repo_ids accessor and a video loaded/loading guard mirroring the Images one.
- the catalog advertised Lightricks/LTX-2.3-fp8 as loadable, but the LTX-2.3 loader refuses the
official scaled-FP8 single file (.weight_scale/.input_scale) and points to GGUF/BF16, so a pick
routed to a ~76 GB download that always fails on load. Remove the FP8 artifact.
- pickDefaultArtifact only sorts by format, so the HunyuanVideo group's 480p (listed first) beat
the 720p even on GPUs where 720p fits the budget. List 720p first so the fit loop prefers it and
falls back to 480p only on smaller cards.
* diffusion: add compute int8/fp8_dynamic text-encoder quant, wire into video
Add two torchao compute text-encoder quant modes to the diffusion precision
engine, alongside the existing layerwise fp8 and weight-only nvfp4:
- int8: per-token activation + per-channel weight (torch._int_mm), with per-layer
keep-bf16 selection. int8 degrades on large encoders unless the most
quant-sensitive decoder blocks stay bf16, so it engages only for families with
a measured keep-bf16 schedule (qwen-image / qwen-image-edit keep first+last 6,
flux.2-dev keeps first 3); a family without one falls back to fp8.
- fp8_dynamic: per-row fp8 compute (torch._scaled_mm), keeping the matmul in fp8
on the tensor cores instead of upcasting each forward like the layerwise fp8.
The selective int8 caster reuses the committed transformer-quant factory
(_make_quant_config / make_filter_fn / exclude_tokens_for_scheme) plus a small
structural first/last-N block skip, so it depends only on committed APIs.
Wire text-encoder quant into the video backend, which previously loaded the
companion encoder (Gemma3 / UMT5 / Qwen2.5-VL) dense bf16 while quantising only
the DiT. text_encoder_quant is plumbed through the load request, validation, the
load chain, the resolved record, and status, mirroring the image backend; it
applies for every load kind (the encoder is dense regardless of how the DiT was
sourced). Widen the image and video load request Literals and add the video
status field.
Tests: int8 family-schedule routing and fp8 fallback, fp8_dynamic routing,
hardware gates (int8 sm_80+, fp8_dynamic sm_89+), the structural block selection,
the real int8 filter closure (keeps the first blocks plus the vision tower /
lm_head / T5 wo dense), and the video route threading and 422 validation.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* text-encoder quant: skip the torchao modes under offload (both backends)
quantize_text_encoders applied int8-with-schedule / fp8_dynamic / nvfp4 (all torchao) to the
text encoder regardless of the offload policy. An offload placement then moves the quantized encoder
with Module.to(), which torchao tensor subclasses reject (aten._has_compatible_shallow_copy_type is
unimplemented) -- a hard crash, the same one the DiT path already skips torchao quant under offload to
avoid. Add offload_active to quantize_text_encoders and skip the torchao modes when set; layerwise fp8
is not torchao and still streams under offload. Both the video and image loaders pass
offload_active = (offload policy != none).
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* diffusion: skip non-bf16 linears for scaled_mm quant schemes
The fp8 / mxfp8 / nvfp4 schemes run on torch._scaled_mm and the fp4 / mx GEMMs,
which assert a bfloat16 input weight. On a mixed-precision DiT that keeps some
linears in fp32 for numerical stability (the Wan and Hunyuan video transformers
do this), quantize_ hits the first fp32 linear, raises, and the best-effort
wrapper swallows it to None, so the whole transformer stays dense with no error
and no speedup or memory saving.
Add a require_bf16 gate to make_filter_fn and pass it for the scaled_mm schemes
in quantize_transformer (and the fp8_dynamic text-encoder caster). The gate
skips non-bf16 linears so the scheme engages on the bf16 ones. int8 uses
torch._int_mm, which quantizes fp32/fp16 weights fine, so it leaves the gate off
and keeps its current coverage.
Verified on Wan2.2-TI2V-5B: fp8 and mxfp8 now quantize 303 linears via the
committed quantize_transformer path where they previously engaged 0.
* prequant builder: mirror the scaled-mm bf16 gate offline
The runtime DiT quantizer skips non-bf16 Linears for the scaled_mm schemes (fp8,
nvfp4, mxfp8) so the scheme engages on a mixed-precision transformer instead of
aborting on the first fp32 Linear. The offline prequant builder reused make_filter_fn
without that gate, so building an fp8/nvfp4/mxfp8 checkpoint for a mixed-precision DiT
(Wan, Hunyuan keep _keep_in_fp32_modules in fp32 even under torch_dtype=bf16) would hit
the same fp32 Linear and abort, breaking the builder's stated offline == runtime,
LPIPS-0 invariant. Thread require_bf16 = scheme in _SCALED_MM_SCHEMES through the builder,
record it in the checkpoint metadata, and verify it on load (mirrors the existing
exclude_name_tokens guard) so a future _SCALED_MM_SCHEMES change cannot silently load a
checkpoint built under the old filter.
* Keep nvfp4 fp32 linears quantised (bf16 gate is fp8/mxfp8 only)
Verified on torchao 0.17 / B200: fp8 per-row asserts 'PerRow quantization only
works for bfloat16 precision input weight' and mxfp8 asserts 'Only supporting bf16
out dtype', but NVFP4's high-precision conversion quantises an fp32 weight fine
(forward included). So the bf16 skip-gate must be fp8/mxfp8 only, not all scaled_mm
schemes -- otherwise nvfp4 leaves large fp32 projections dense, losing the intended
memory/speed gain. Rename _SCALED_MM_SCHEMES -> _REQUIRE_BF16_SCHEMES = (fp8, mxfp8)
and thread it through the runtime filter, the offline builder, and the loader
require_bf16 verification (offline == runtime preserved).
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
The auto FBCache policy keyed on the full step count whenever strength was omitted, but the
loader only passes the strength kwarg when it is set, so an img2img/inpaint pipe then runs its
OWN signature default (< 1, e.g. FluxImg2ImgPipeline's 0.6). FBCache would engage on the full
28 steps while the pipe actually denoises ~16, degrading the image on exactly the short
trajectory the policy exists to keep uncached. Thread the pipe's signature default into the
policy via a new effective_request_strength helper (unit-tested), so the effective denoise count
matches what the pipe runs.
_dense_quant_prefetch_needed widened the prefetch to pull the base repo's transformer/
shards whenever a dense-quant candidate resolved, but balanced/low_vram (and the legacy
cpu_offload flag) force load_pipeline onto offload unconditionally in plan_diffusion_memory,
so its re-plan never flips to OFFLOAD_NONE and the dense build never runs. The offloaded GGUF
path then never touches those shards, so the widened prefetch only wastes a multi-GB download,
and a disk-full on that begin_load pull has no GGUF fallback (unlike the in-load_pipeline dense
failure). Mirror plan_diffusion_memory's definite-offload gates so the prefetch stays scoped.
- _dense_quant_prefetch_needed widened the transformer/ prefetch to pull the base
repo's full dense bf16 shards even when a prequant checkpoint is configured
(candidate.prequant), contradicting its own docstring. That both defeats the
prequant download savings and can hard-fail begin_load on a disk-full (no GGUF
fallback there). Only widen for a real dense build (candidate is not None and
not candidate.prequant).
- DiffusionStatusResponse declared no 'resolved' field, so Pydantic's default
extra='ignore' silently dropped the per-control auto-policy provenance the
backend records (build_resolved_record / state.resolved) -- the plumbing never
reached any client. Declare the field so it round-trips.
For a narrow (fp8) Ideogram-4 base the planner reserves the family's known bf16
component total, but the reservation was gated on model_dense_mib being non-None.
On a first-time load an empty blob cache (or a best-effort download probe that
swallowed a transient HF error) leaves model_dense_mib None, so the guard skipped
the reservation exactly when it was needed: the planner then read 'size unknown ->
stay resident' and the ~54 GB pipeline OOMed a card that offload would have fit.
family_bf16_components_gb is a network-free constant, so reserve it whenever the
cache signal is absent (use it directly when None, else take the max).
- _dense_quant_prefetch_needed widened the prefetch to pull the base repo's bf16 transformer/ shards
whenever a dense-quant scheme could resolve, with no disk check. On the offload path that can fill
the cache volume mid-download and hard-fail the load in a spot unload/cancel cannot preempt, instead
of the disk guard falling back to running the GGUF as-is (the Dtype hint's documented disk fallback).
Defer to resolve_dense_quant_candidate, the same disk-aware resolver load_pipeline re-plans against,
so the prefetch widens only when the dense build would really run.
- An explicit Speed=off (bit-exact) load with an unset dtype was promoted to auto-quant by the Dtype
default, silently engaging int8/fp8 + compile and breaking the bit-exact request (an auto DEFAULT
overriding an EXPLICIT control). Suppress the auto-dtype default when speed is explicitly off, in both
load_pipeline and the prefetch.
The dense-quant re-plan passes transformer_resident_override_mib (the bf16 build
peak) AND computes companions via _companion_cache_bytes(base), which sums every
flat blob in the HF cache. Because the dense path prefetches the base transformer/
shards into that same cache before load_pipeline runs, the transformer is counted
twice, inflating the footprint (~44 GB instead of ~20 GB in the reproduction) and
wrongly forcing offload for models that fit resident -- the case this path exists
to enable. Add companion_override_mib and pass the auto-policy's own text-encoder
plus VAE estimate on the re-plan so the cache (with its prefetched transformer) is
not read for this artifact.
- Step cache 'Off' is preserved: the frontend defaulted to 'off' and mapped it
to an omitted transformer_cache, which the backend now reads as 'auto', so
leaving the control at Off silently enabled FBCache on 20+ step families.
Default the control to Auto, add an explicit Auto option, and send
auto -> omitted so Off maps to an explicit cache-off.
- Kernel auto-install adds --no-deps: 'pip install --only-binary :all: xformers'
resolves xformers' pinned torch and replaces the running torch/triton. --no-deps
installs only the best-effort kernel wheel; an ABI mismatch just fails to import
and falls back to native, never clobbering core deps.
- Do not retry a failed kernel install under the load lock: the pre-install runs
outside the locks, then the in-lock resolve re-attempts pip (up to 600s) while
holding _generate_lock/_lock and blocking unload/cancel/new loads. Record the
attempt in a process-level set so the in-lock call short-circuits to native.
- Cache auto-toggle keys on effective denoise steps: an image-conditioned run with
strength < 1 (upscale default 0.35) denoises a fraction of the requested steps,
so a 28-step request runs ~10 steps. Compute the effective count the way diffusers
get_timesteps does and gate FBCache on it, only when strength is actually applied.
Ideogram 4 assembles two DiTs per-component (a conditional transformer plus a
separate unconditional_transformer), so there is no transformer-only single-file
or GGUF artifact that could supply both. Add a pipeline_only family flag and
reject the gguf/single_file kinds in validate_load_request, before a load evicts
the current model, instead of assembling a pipeline missing its second DiT.
Extend the fp8 bf16-resident size override to a LOCAL directory mirror of the
ideogram-4-fp8 base: such a path never string-matches base_repo, so detect the
fp8 layout from the transformer shard headers (a *.weight_scale marker) and
reserve the bf16 footprint, matching the remote-base behaviour. A local nf4
mirror has no fp8 scales and correctly stays planned against its compressed bytes.
Review follow-ups on the image-generation PR:
- ControlNet: resolve_controlnet accepts a bare owner/name repo without the
non-GGUF base trust gate, and _controlnet_pipe hands it straight to
from_pretrained. A malicious pickle .bin would deserialize on load, so run
the same Hugging Face malware preflight (evaluate_file_security) the chat and
export loaders use before any remote ControlNet load; local dirs are exempt.
- Dataset thumbnails: key the cache on the full filename instead of the stem so
sample.png and sample.jpg no longer collide on one .thumbs file (which could
serve or delete the wrong image); the delete cleanup globs the same key.
- Diffusion training start: mirror start_training's API-key guard so an API
client cannot start training (which frees VRAM by unloading chat) while an
inference request is streaming; it now returns 409 before any GPU is freed.
- Model picker: include the curated safetensors row keys in the recommended
roving key list so arrow-key navigation reaches those rows instead of hitting
the duplicate option-missing id.
Tests: ControlNet malware gate (remote blocked before from_pretrained, local
skipped), thumbnail same-stem cache separation, API-key diffusion-start 409
before GPU free. Full diffusion suites green.
- Run the trainer's caption discovery in the start route BEFORE freeing GPU
residents, so a missing or uncaptionable dataset 400s without evicting the
loaded chat/Images model.
- sd.cpp unload now waits out a cancelled one-shot generation on the generate
lock before reporting the device free, matching the diffusers backend.
- Clearing a caption that came from metadata.jsonl writes an empty sidecar
tombstone instead of unlinking (both readers treat an existing sidecar as
authoritative), so the cleared label cannot resurface.
- The ControlNet wrapper pipe is only cached while its load is still current,
closing the unload race the model cache already handled.
The prefetch already scopes the file list (no packaged root singles, no
dtype-variant twins, no ONNX/Flax exports), but from_pretrained was then
called with the hub id, and its own snapshot sweep re-downloaded the
skipped files anyway: 24 GB per FLUX.1 repo and 65 GB on FLUX.2-dev, as
found in the blob cache. Return the snapshot dir from the prefetch (keyed
on the pipeline manifest) and hand it to every pipeline-assembly
from_pretrained site; any prefetch failure keeps the hub id and the old
behavior.