Comment-only pass over the Python this PR touches: drop what the code already
says, collapse multi-line explanations that still read on one line, and keep
the reasoning that is not recoverable from the code. No code, docstring
semantics or behaviour changes; verified with an AST comparison against the
previous revision, and the backend suite is unchanged (same 37 environment
failures as before: the API integration tests that need a live keyed server,
the flash-attn install hooks, and the GPU memory fields).
Three from the latest review.
The video download plan always asked for the wide base file list, so an LTX-2.3
pick staged the 2.0 base's VAEs, vocoder and connectors that the checkpoint
supplies itself, while the companion files the 2.3 assembly does read were left
out of the plan and pulled inline at load, outside the panel's progress, cancel
and disk preflight. The plan now recognises a 2.3 pick by name (the load keeps
the authoritative header probe, and under-guessing only falls back to the
load-time pull), narrows the base list, and stages the extras in the same entry
as the checkpoint so one repo stays one scoped job.
A pick routed from the chat picker arrives as ?model= and ?quant= with no picker
metadata, so a bare local .gguf or .safetensors was loaded as a pipeline: an
explicit model_kind wins over the backend's filename sniffing, so it evicted the
resident model and then failed on the missing model_index.json. Both pages now
derive the load kind from the path, the same way their own picker handlers do.
A torchao int8/fp8 build takes adapters only at load time. Switching artifact
inside one family keeps the LoRA selection, since the family did not change,
but the load did not bake it, so the next generation was rejected with 'reload
the model with the adapter selection' while the picker still showed the adapter
as active. The selection is now dropped once per resident build, with a message
saying to pick and load again.
Both read huggingface_hub's import-time HF_HUB_CACHE, which changing the cache
folder does not update: progress counted the old root while the download wrote to
the new one, and from_pretrained could split one model across both.
Mount-time recovery handled only phase=completed, so reloading the page
after a multi-minute generation failed left an idle view with no
diagnosis: the backend keeps the terminal failed record only until the
next job, and nothing else survives the reload. Surface it the same way
the poll does, filtering the cancelled sentinel.
The 22B distilled DiT was trained against ltx_core's fixed
DISTILLED_SIGMA_VALUES, but the diffusers scheduler derives 8-step
spacing from resolution-shifted flow matching and lands far off at
every reachable mu (second sigma 0.945-0.981 vs 0.99375, tail
0.37-0.61 -> 0.1 vs 0.725 -> 0.42 -> 0). At the distilled default step
count the backend now passes the list verbatim, neutralising the
scheduler's dynamic shift and terminal stretch for the call (they
distort even explicit sigmas) and restoring them afterwards. Other
step counts and the dev/base DiT keep the scheduler's own spacing.
Live-verified on B200 through the video branch backend: the scheduler
holds the exact curve after an 8-step distilled GGUF generation, config
restored, healthy clip. Also reword the transformer_quant resolved
reason to the measured reality: quant halves resident weights and
hosted checkpoints cut load time, while per-step speed is roughly bf16
parity.
Two live-test findings on the video progress endpoints:
- load-progress downloaded_bytes froze mid-download: the counter used
scan_cache_dir, which skips in-flight *.incomplete blobs, so it sat at the
last completed blob for the whole multi-GB shard pull while the disk kept
filling. Count the repo's cache directory directly (completed plus incomplete
blobs, snapshot symlinks skipped so nothing is double-counted).
- generate-progress reported total_steps=null / fraction=0 while step advanced:
the video API only carried the native total field while the image API exposes
total_steps and fraction, so one poller could not work against both. Derive
the image-compatible aliases in generate_progress and declare them on the
response model; the native total stays for back-compat.
- diffusion_cache: do not engage FBCache when the selected pipeline opens no cache_context.
A CacheMixin transformer is necessary but not sufficient -- Flux Kontext / img2img /
inpaint / controlnet reuse the CacheMixin FluxTransformer2DModel yet their __call__ never
opens a cache_context, so the First-Block-Cache hook raised 'No context is set' on the
first forward, crashing every default FLUX.1-Kontext edit (28 steps, above the FBCache
threshold). Detect it from the pipeline __call__ source, resolved off the instance so the
per-expert proxy view delegates to the real pipe.
- diffusion_attention: honor an explicit aiter backend on ROCm/AMD targets instead of
dropping it via the NVIDIA-only guard (aiter is the AMD ROCm kernel; it only works there).
- video: clear the CUDA cache on a failed load so a partially built pipeline's reserved VRAM
does not OOM the next load (mirrors the image backend), and re-check cancellation after the
export/mux so a clip cancelled during the blocking encode is discarded, not persisted.
- diffusion_auto_policy / diffusion_prequant: validate a request-supplied prequant path
override (present AND allowlisted) before budgeting the small prequant plan, so the loader
does not skip the dense shards and then rebuild dense after evicting the resident pipeline.
- diffusion_controlnet: family-gate a curated ControlNet addressed by its full repo id, not
only its short catalog id, so a cross-family repo id 400s up front instead of downloading
and loading through the wrong ControlNet class.
Video: a local FILE picked for a gguf/single_file load is handed straight to the
loader (the resolver returns the file itself, ignoring gguf_filename), so its own
suffix must match the kind. A .gguf picked as single_file (or a .safetensors picked
as gguf) slipped past the gguf_filename checks, evicting the resident GPU owner
before failing in from_single_file. Reject the mismatch in validate, before the handoff.
Engine router: publish the newly selected engine only after the old one finishes
unloading. The arbiter's diffusion evictor unloads the active engine, so flipping the
active name to the new (empty) engine first let a concurrent chat/video acquire evict
that empty engine and take the GPU while the old model was still freeing VRAM.
Two evict/OOM fixes on the diffusion load paths:
- The video load moved a pipeline onto the GPU (apply_memory_plan) and
committed it while holding no lock, so an unload / GPU-arbiter eviction --
which bumps the load token and then barriers on _generate_lock before
freeing -- could hand VIDEO to chat/images and let the new owner allocate
concurrently with the in-flight placement, OOMing. Hold _generate_lock
across placement + the locked commit, mirroring the image backend, so an
evicting owner waits until this worker's placement is torn down or
committed. Lock order stays _generate_lock -> _lock (unload takes _lock
then releases it before the barrier), so there is no deadlock.
- resolve_local_single_file reinterpreted an On-Device folder as a base
single_file load whenever it held exactly one .safetensors, so a PEFT LoRA
adapter folder (adapter_config.json + adapter_model.safetensors) with a
family-token name was picked as a base checkpoint, evicting the resident
model before from_single_file failed on the adapter weights. Skip adapter
folders (adapter_config.json) and the adapter_model basename so the pick
stays a pipeline load and 400s in validation, before the GPU handoff.
Adds regression tests for both.
The video load-request preflight gated its local-pipeline check on
root.is_dir(), so a bare local file (e.g. /models/ltx-2.safetensors) sent
with model_kind=pipeline skipped it, passed validation, and the route then
evicted the resident GPU model before from_pretrained failed on the
non-directory path. Gate on root.exists() instead (mirroring the image
loader's diffusion.validate_load_request), so a local file is rejected up
front. Add a regression test.
Three gates the image backend has were missing on the video path:
- kind/extension mismatch: a model_kind 'single_file' with a .gguf name (or 'gguf'
with a non-.gguf) passed the video preflight, so the route acquired VIDEO and
evicted the resident GPU owner before the wrong single-file loader failed in the
background. Reject the mismatch up front, mirroring the image loader.
- Windows-shaped missing local pick: the missing-path check only matched POSIX
prefixes (/ ~ ./ ../), so a missing Windows path (C:\ / C:/ or any backslash path)
was treated as a Hub repo and only failed after the GPU handoff. Use the image
loader's is_absolute()/backslash path-shaped check.
- speed=off auto-quant: a full-pipeline load with Speed=off but Precision=auto still
promoted the unset precision to auto-quant, so on a dense-capable GPU it engaged
torchao quantization and forced the speed back to default -- silently breaking the
user's bit-exact request. Suppress the auto promotion when speed_mode is off, as
the image loader does.
Regression tests for each.
Several image/video/training preflights ran before the route acquires the GPU or
frees resident models, but let a doomed local pick through and only failed deep in
the background load, after the user's chat/Images/Video model was already evicted.
- Local base_repo / base_model: _is_trusted_diffusion_repo accepts any existing
local path, but the base loads via from_pretrained (needs model_index.json). A
local dir that is not a diffusers pipeline passed the trust gate, evicted the
resident model, then failed. Add a shared _assert_local_base_is_pipeline check
and call it in the image, video, and training preflights.
- Dataset images: discover_image_caption_pairs only checked filenames, so a
corrupt or zero-byte upload passed the start-route preflight, freed the GPU, then
crashed the spawned trainer in PIL. Add an opt-in verify_images decode probe
(cheap PIL header check) that the start route enables; the trainers leave it off
since they decode every image anyway.
- Local single-file safetensors: the On-Device scanner advertises a bare
.safetensors directory (no model_index.json) as a text-to-image model, but the
picker starts it as a pipeline with no filename, so every click 400s. Reinterpret
such a pick as a single_file load of the sole checkpoint (resolve_local_single_file)
so the advertised model is actually loadable.
Regression tests for each: local non-pipeline base (image/video/training), the
verify_images decode gate, and resolve_local_single_file.
- _local_model_task now tags a local diffusers pipeline that resolves to a video
family (LTX / Wan / Hunyuan) as text-to-video, mirroring the cached-repo
_cached_repo_task, so supported local video pipelines surface in the Video
On-Device picker instead of being routed to the Images picker where the image
loader rejects them. Gated on _local_is_diffusers so only a real loadable
pipeline dir reaches the video check.
- Video validate_load_request now rejects a local pipeline pick whose directory has
no model_index.json before the GPU handoff, mirroring the image loader, so a bad
local pipeline can no longer evict the resident model and only then fail deep in
from_pretrained.
- The custom/env-mode uninstall now removes a sibling stable-diffusion.cpp only when
it carries the Studio owner marker. install_sd_cpp_prebuilt writes the canonical
.unsloth-studio-owned marker on install; uninstall.sh and uninstall.ps1 keep any
unowned checkout (a user's own git clone of stable-diffusion.cpp beside a custom
Studio root is no longer deleted). A pre-marker Studio build is left behind rather
than a user file removed.
Adds regression tests: local video pipeline tagged text-to-video (and a video-named
non-pipeline dir stays untagged so it can never trigger a doomed pipeline load), the
video local-pipeline preflight rejection, the install ownership marker, and the
uninstall keeping an unowned sibling while removing an owned one.
* Auto policies: deferred dense compile, video compile default, step cache and precision auto
Image dense loads with speed unset no longer sit at plain off: the load stays
bit-identical eager, and the 3rd generation in a session engages the default
compile profile plus the cuDNN attention upgrade mid-session (a one-off image
never pays the warmup, repeated use amortises it). Video dense loads resolve
straight to the default profile since a clip denoise amortises the compile
within a single run, and never to max.
Video also gains the image backend's tri-state auto policies: unset step cache
now decides from the default schedule and re-checks the actual step count per
generation, and unset precision (transformer_quant) hands the decision to the
hardware ladder instead of staying off. Memory badge reason now says plainly
that everything fits when no offload is planned.
* Rename Dtype to Precision, add the video Precision control, step cache Auto option
The images Advanced panel's Dtype row is now Precision (same control, clearer
name), and the video Advanced panel gains the matching Precision select wired
to the load route's existing transformer_quant field, gated to full-pipeline
loads the way the image control gates to GGUF. Step cache selects on both
pages gain an explicit Auto option as the default (the previous Off default
silently behaved as auto and never let anyone pin off), and the Speed and
Attention tooltips now state the deferred dense compile and the SageAttention
black-frame caveat.
* Model catalog: canonical diffusion model groups with device-aware routing
One canonical name per image/video model, its published artifacts (GGUF, FP8,
bnb-4bit, official BF16) as data, and pure routing helpers: suffix-stripped
canonical keys (owner-preserving; cross-owner merges only via explicit
aliases), group/artifact lookups, a flat back-compat options shim, load-spec
resolution replacing the pages' lookup tables, search matching over old ids
and format tokens, the GGUF fit ladder extracted from the variant expander,
and pickDefaultArtifact/pickDefaultQuant deciding what a bare group click
loads (downloaded first, then the best quality that fits 70 percent of VRAM,
GGUF as the safe fallback). Checked by npm run catalog:check, following the
i18n:check pattern.
* Picker: one canonical row per diffusion model with a format second level
The Images and Video pickers now render the curated catalog as one row per
model in Recommended: clicking loads the best artifact for the device (the
routed GGUF quant, a prequant FP8/bnb-4bit that fits, or the official BF16),
and a chevron opens the per-format list, with the GGUF row nesting the usual
quant expander. Live HF listing rows that belong to a group are deduplicated,
search collapses member repos into their group (old ids and format tokens
still match), and the On Device sections group cached member repos under the
same canonical name with the per-repo rows inside. Curated groups render from
the catalog rather than the HF listing, which finally surfaces LTX-2.3 in the
video Recommended list (its hub pipeline_tag is image-to-video, so the
text-to-video listing always missed it) and exposes the HunyuanVideo 720p
repack next to 480p.
Backend: /cached-models now tags trusted video-family repos text-to-video
instead of blanket text-to-image, and the pickers admit catalog-known
non-unsloth repos On Device, so cached Lightricks/Wan/Hunyuan pipelines
finally appear in the Video picker. Chat pickers pass no catalog and are
unchanged.
* Download formats, tab icons, plain-language train tips, 3-loop autoplay
The image Download button becomes a menu: PNG saves the original bytes with
the embedded recipe, JPEG and WebP re-encode client-side from the fetched
blob (JPEG flattened onto white). The video Download button gains MP4
(original, keeps audio), WebM and GIF; the latter two transcode server-side
from the stored MP4 via PyAV (VP9 realtime profile for WebM, ~12 fps adaptive
palette for GIF) behind a new gallery export route that 501s with a readable
message when a codec is missing.
Generated clips no longer loop forever: the player replays a clip three times
per selection, then pauses with controls up; a new generation or a refresh
gets its own three plays. The Create/Train tabs reuse the sidebar's New Chat
and Train icons (TestTubeOutlineIcon moved to a shared lib module), and every
Train tab helper text is now one plain sentence.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Keep the Create/Train tab icon and label on one line
TabsTrigger renders its children inside a plain inline span and the
Tailwind preflight gives svg display:block, so the HugeiconsIcon forced
the label onto a second line. Wrap icon plus label in their own
inline flex row inside each trigger.
* Strip -int8 and -nvfp4 prequant suffixes in the model catalog key
canonicalKeyFor already lowercases before matching, so -GGUF/-FP8 in any
case were covered; -int8 and -nvfp4 were not in the suffix table, so
such repos rendered as standalone rows in Recommended and On Device
instead of standardizing into their base-name group and routing through
pickDefaultArtifact. Added both suffixes plus case-insensitivity and
routing assertions to the catalog check.
* Standardize non-catalog picker rows to their base model name
The curated catalog already collapses its own groups, but hub listing
rows and cached repos outside the catalog (ERNIE-Image, FLUX.2-klein,
Qwen-Image-Edit-2509, FLUX.2-dev) still rendered raw ids with -GGUF /
-FP8 style suffixes in Recommended and On Device.
- model-catalog.ts: new stripArtifactSuffixesForDisplay, a
case-preserving twin of canonicalKeyFor's stripping that keeps the
owner prefix and original casing for display.
- pickers.tsx: recommended hub rows and the downloaded GGUF/model rows
pass their labels through it when a catalog is present, so only the
diffusion pickers change; chat rows keep raw ids. Click targets keep
the full repo id, and the format badge still shows the artifact kind.
- Catalog check covers the new helper across GGUF/FP8/int8/nvfp4 in
both cases plus no-op and suffix-only names.
* Offer official BF16/FP8 artifacts per model group and fix gallery label clipping
Model picker changes so groups are not limited to unsloth quant repos:
- model-catalog.ts: each image group that has an official vendor pipeline
now carries its BF16 (official) artifact as the top (highest quality)
entry - Tongyi-MAI/Z-Image-Turbo, Qwen/Qwen-Image, Qwen/Qwen-Image-2512,
Qwen/Qwen-Image-Edit-2511, black-forest-labs/FLUX.1-dev, FLUX.1-schnell
and FLUX.1-Kontext-dev. The LTX-2.3 video group now lists Lightricks'
own bf16 and fp8 distilled single-file checkpoints alongside the GGUF.
Resident sizes are set from the actual weight totals (FLUX ships a
duplicate single-file that from_pretrained ignores, so FLUX bf16 is ~32
GB not 54). The repos that used to be aliases are now real artifacts.
- The router already prefers the highest-quality artifact that fits the
0.7 x GPU budget, so a datacenter GPU now defaults to official BF16
while consumer GPUs still route to the fitting quant or GGUF. That is
why bnb-4bit was the Z-Image-Turbo default before: it was the only
non-GGUF artifact and it was already downloaded.
- diffusion.py: allowlist the four official image repos not previously
trusted (qwen/qwen-image-2512, qwen/qwen-image-edit-2511,
black-forest-labs/flux.1-schnell, flux.1-kontext-dev). All verified as
safetensors-only diffusers model_index pipelines. The LTX-2.3
checkpoints are already on the video trust list.
- catalog check: BF16-wins-on-datacenter, quant-wins-on-consumer, and the
single-file load specs for the LTX-2.3 checkpoints.
Also fixes the video gallery thumbnail caption: the leading duration was
clipped by the rounded corner and selection border, so the strip now has
enough left/bottom padding to clear the curve.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* video gallery: guard export transcode against a stream-less clip
_transcode_webm and _transcode_gif indexed src.streams.video[0] before
checking the stream list, so a container with no video stream raised a bare
IndexError that the broad handlers then re-labeled as a missing libvpx or
decoder. Raise an explicit RuntimeError naming the real cause in both the
WebM and GIF paths.
* Studio: honor explicit attention/format choices, fix distilled-LTX defaults and On Device catalog routing
* Remove stray planning notes accidentally committed to the branch
* video: add transformerQuant to the load callback deps
handleLoad reads transformerQuant but omitted it from the useCallback dep array,
so after the user changes only Precision and then selects a model or clicks
Reapply, the memoized callback keeps the stale closure and loads the previous
precision. The image page's equivalent callback already lists it.
* model picker: honor the format filter when routing catalog clicks; add catalog rows to the roving list
- routedArtifactFor now scopes a group's artifacts to the active format filter
(the same matchesFormatFilter predicate the visibility check uses) before
pickDefaultArtifact, so a group shown only because it owns a GGUF no longer
routes a click to a large non-GGUF download. Covers both the Recommended and
On Device grouped paths.
- hubOptionKeys now includes the catalog-group, search-catalog-group, and grouped
On Device row keys in exact render order, so arrow/Home/End roving reaches the
catalog rows instead of giving them a duplicate missing id and skipping them.
* model picker: don't treat a partial base cache as downloaded
A partially-cached base repo (a cancelled download that left only some weights)
was counted as downloaded, so an On Device click routed to a fresh multi-GB
re-download instead of the complete GGUF. The picker's endpoint (/api/models/
cached-models) did not carry a partial flag at all, so a frontend-only guard
could not see it. Surface partial from that endpoint by reusing the hub inventory
scan's snapshot-partial detector, plumb it through CachedModelRepo (backend +
frontend types), and skip partial base repos when building the downloaded set.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* model picker + diffusion: drop partial/unloadable cached rows, skip defer-compile before a LoRA gen
- On Device (cached non-GGUF) rows filtered partial-download snapshots back in: sortedCachedModels
gated on passesTaskGate + a groupForRepoId key match but, unlike downloadedSet, never checked
c.partial, so an incomplete unsloth snapshot showed as a loadable On Device row (click errors or
silently re-fetches multi-GB). It also admitted repos that only match the catalog by group KEY
(a base / uncurated-quant sibling like Qwen/Qwen-Image-2512) which have no loadable artifact and
dead-end at the trust gate. Add !c.partial and gate on artifactForRepoId (what loadSpecFor
resolves) instead of groupForRepoId, so a cached row shows only when the backend can load it.
- Deferred speed-auto engaged the compile profile on the 3rd generation BEFORE _apply_loras. A
compiled transformer rejects LoRA (supports_lora is False) and _apply_loras raises before its
unchanged-selection no-op, so once compile engaged every LoRA generation on that load failed
permanently. Skip the deferral when a LoRA is requested (compile and LoRA are mutually exclusive)
and let it engage on a later LoRA-free generation.
* Scope the cached-model partial probe to the listed snapshot dir
list_cached_models builds each row from the largest/complete copy across HF cache
roots, but _cached_repo_partial probed is_snapshot_partial with no repo_cache_dir,
so the scan spanned every root: a stale .incomplete copy in one root would flag a
complete copy in another as partial and hide the usable model from the picker (the
click then routes to a re-download). Forward the winning snapshot's repo_path so all
three partial signals are scoped to that copy, matching the sibling inventory paths
(models/dataset cache_inventory, local_inventory).
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Do not auto-route to gated repos, prefer complete cached copies, defer compile past attached LoRA, scope group expand keys
Four fixes:
- pickDefaultArtifact's not-downloaded ladder returned the gated BF16 FLUX.1-dev / Kontext-dev
before the open GGUF on a large GPU, so a bare group click routed to a repo the user may lack
license/token access to. Add a gated flag and skip gated artifacts in the not-downloaded ladder
(an already-downloaded gated artifact is still returned).
- list_cached_models picked the largest duplicate cache copy and computed partial only on it, so a
larger partial copy shadowed a smaller complete one; since partial rows are dropped from the
picker the usable model vanished. Prefer completeness, then size.
- the deferred-speed compile engaged on a no-LoRA generation while an adapter from a prior
generation was still attached, baking it into the compiled graph (the later unload is swallowed
on a compiled pipe); also defer while adapters remain attached.
- routeGroupClick's GGUF fallback toggled the context-free canonicalId while the chevron toggles
the context-scoped expandKey, leaving the format list un-collapsible in one context, dead in the
other, and risking cross-context expansion; thread expandKey through.
* Guard video pipeline repos from deletion, drop the always-failing LTX FP8 artifact, prefer 720p Hunyuan
Three round-6 fixes:
- cached non-GGUF video repos now surface in the Video On-Device picker with the normal delete
action, but /delete-cached only guarded chat + the Images engine, so a loaded/loading Wan / LTX /
Hunyuan pipeline could have its HF snapshot removed from under it. Add a VideoBackend
loading_repo_ids accessor and a video loaded/loading guard mirroring the Images one.
- the catalog advertised Lightricks/LTX-2.3-fp8 as loadable, but the LTX-2.3 loader refuses the
official scaled-FP8 single file (.weight_scale/.input_scale) and points to GGUF/BF16, so a pick
routed to a ~76 GB download that always fails on load. Remove the FP8 artifact.
- pickDefaultArtifact only sorts by format, so the HunyuanVideo group's 480p (listed first) beat
the 720p even on GPUs where 720p fits the budget. List 720p first so the fit loop prefers it and
falls back to 480p only on smaller cards.
* diffusion: add compute int8/fp8_dynamic text-encoder quant, wire into video
Add two torchao compute text-encoder quant modes to the diffusion precision
engine, alongside the existing layerwise fp8 and weight-only nvfp4:
- int8: per-token activation + per-channel weight (torch._int_mm), with per-layer
keep-bf16 selection. int8 degrades on large encoders unless the most
quant-sensitive decoder blocks stay bf16, so it engages only for families with
a measured keep-bf16 schedule (qwen-image / qwen-image-edit keep first+last 6,
flux.2-dev keeps first 3); a family without one falls back to fp8.
- fp8_dynamic: per-row fp8 compute (torch._scaled_mm), keeping the matmul in fp8
on the tensor cores instead of upcasting each forward like the layerwise fp8.
The selective int8 caster reuses the committed transformer-quant factory
(_make_quant_config / make_filter_fn / exclude_tokens_for_scheme) plus a small
structural first/last-N block skip, so it depends only on committed APIs.
Wire text-encoder quant into the video backend, which previously loaded the
companion encoder (Gemma3 / UMT5 / Qwen2.5-VL) dense bf16 while quantising only
the DiT. text_encoder_quant is plumbed through the load request, validation, the
load chain, the resolved record, and status, mirroring the image backend; it
applies for every load kind (the encoder is dense regardless of how the DiT was
sourced). Widen the image and video load request Literals and add the video
status field.
Tests: int8 family-schedule routing and fp8 fallback, fp8_dynamic routing,
hardware gates (int8 sm_80+, fp8_dynamic sm_89+), the structural block selection,
the real int8 filter closure (keeps the first blocks plus the vision tower /
lm_head / T5 wo dense), and the video route threading and 422 validation.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* text-encoder quant: skip the torchao modes under offload (both backends)
quantize_text_encoders applied int8-with-schedule / fp8_dynamic / nvfp4 (all torchao) to the
text encoder regardless of the offload policy. An offload placement then moves the quantized encoder
with Module.to(), which torchao tensor subclasses reject (aten._has_compatible_shallow_copy_type is
unimplemented) -- a hard crash, the same one the DiT path already skips torchao quant under offload to
avoid. Add offload_active to quantize_text_encoders and skip the torchao modes when set; layerwise fp8
is not torchao and still streams under offload. Both the video and image loaders pass
offload_active = (offload policy != none).
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* diffusion: skip non-bf16 linears for scaled_mm quant schemes
The fp8 / mxfp8 / nvfp4 schemes run on torch._scaled_mm and the fp4 / mx GEMMs,
which assert a bfloat16 input weight. On a mixed-precision DiT that keeps some
linears in fp32 for numerical stability (the Wan and Hunyuan video transformers
do this), quantize_ hits the first fp32 linear, raises, and the best-effort
wrapper swallows it to None, so the whole transformer stays dense with no error
and no speedup or memory saving.
Add a require_bf16 gate to make_filter_fn and pass it for the scaled_mm schemes
in quantize_transformer (and the fp8_dynamic text-encoder caster). The gate
skips non-bf16 linears so the scheme engages on the bf16 ones. int8 uses
torch._int_mm, which quantizes fp32/fp16 weights fine, so it leaves the gate off
and keeps its current coverage.
Verified on Wan2.2-TI2V-5B: fp8 and mxfp8 now quantize 303 linears via the
committed quantize_transformer path where they previously engaged 0.
* prequant builder: mirror the scaled-mm bf16 gate offline
The runtime DiT quantizer skips non-bf16 Linears for the scaled_mm schemes (fp8,
nvfp4, mxfp8) so the scheme engages on a mixed-precision transformer instead of
aborting on the first fp32 Linear. The offline prequant builder reused make_filter_fn
without that gate, so building an fp8/nvfp4/mxfp8 checkpoint for a mixed-precision DiT
(Wan, Hunyuan keep _keep_in_fp32_modules in fp32 even under torch_dtype=bf16) would hit
the same fp32 Linear and abort, breaking the builder's stated offline == runtime,
LPIPS-0 invariant. Thread require_bf16 = scheme in _SCALED_MM_SCHEMES through the builder,
record it in the checkpoint metadata, and verify it on load (mirrors the existing
exclude_name_tokens guard) so a future _SCALED_MM_SCHEMES change cannot silently load a
checkpoint built under the old filter.
* Keep nvfp4 fp32 linears quantised (bf16 gate is fp8/mxfp8 only)
Verified on torchao 0.17 / B200: fp8 per-row asserts 'PerRow quantization only
works for bfloat16 precision input weight' and mxfp8 asserts 'Only supporting bf16
out dtype', but NVFP4's high-precision conversion quantises an fp32 weight fine
(forward included). So the bf16 skip-gate must be fp8/mxfp8 only, not all scaled_mm
schemes -- otherwise nvfp4 leaves large fp32 projections dense, losing the intended
memory/speed gain. Rename _SCALED_MM_SCHEMES -> _REQUIRE_BF16_SCHEMES = (fp8, mxfp8)
and thread it through the runtime filter, the offline builder, and the loader
require_bf16 verification (offline == runtime preserved).
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Two follow-ups to the video-tab review fixes:
- the cached-hub arch fallback only probed the active HF cache, but the cached-gguf picker
scans the active, legacy, and default cache roots; a GGUF cached in a non-active root was
offered yet 400d on load. Probe all three roots.
- on a mount/refresh with a model already resident, status.loaded made the Reapply button render
but lastLoad (set only by our own loads) was null, so clicking it silently did nothing. Track a
session reapply descriptor and hide the button when it is absent rather than offer a dead control.
Three video-tab fixes:
- delete-cached refused a loaded video repo but not one a background load is still downloading
(status().loaded is False in that window); add VideoBackend.loading_repo_ids() mirroring the
image backend and check it in the route so deleting mid-download can no longer yank blobs.
- the cached-gguf picker tags a cached HUB GGUF by its general.architecture, but the loader's
arch fallback only read a LOCAL file, so an opaquely-named cached hub LTX GGUF the picker
offered 400d on load; read the arch from the cached blob (network-free) too.
- on a mount with a model already loaded (refresh / load from another client) steps and guidance
stuck at the pre-load default, so a base checkpoint silently generated a degraded clip; seed
them from the backend-authoritative status.defaults once per newly-loaded model.
Verified against diffusers 0.39 source + the HF configs/safetensors headers:
- The Wan VAE decodes in float32 (WanPipeline loads AutoencoderKLWan at torch.float32
while the pipe runs bf16); the loader cast every component to bf16, degrading every
clip. Add vae_force_fp32 (both Wan families) and pin pipe.vae back to fp32 after build.
- The Wan transformers ship FP32 on disk (safetensors headers are F32; A14B index =
57.15 GB per expert = 14.3B x 4, TI2V = 20.0 GB = 5B x 4), so bf16_components_gb held
the fp32 on-disk sums (114.3 / 20.0) instead of the documented bf16-resident sizes.
Halve to 57.2 (two A14B experts) and 10.0 (TI2V), so the plan no longer over-budgets
the DiTs ~2x and forces needless offload on an 80 GB GPU.
- TI2V-5B's VAE is 16x spatial (vae/config.json), so WanPipeline floors H/W to 16*2 = 32.
Snap TI2V to /32 (was /16) so the recorded size matches the generated clip (a 720
request was recorded but rendered at 704). A14B keeps /16 (Wan2.1 8x VAE).
- delete_cached_model guarded chat (llama.cpp/transformers) and Images but not the
Video backend, so a loaded video GGUF (which shares the On-Device GGUF delete UI)
could be removed from under a live pipeline. Add the mirror guard on
get_video_backend().status().
- The Video picker admits a local GGUF by its general.architecture, but the loader
detected the family only from path/name tokens, so a renamed ltxv file (e.g.
model.gguf) was offered yet failed validate_load_request with a 400. _detect_load_family
now falls back to reading the arch (its string is a family alias) when name detection
misses, so the loader accepts exactly what the picker offered; an unsupported video arch
(wan) still resolves to None and 400s as before.
apply_step_cache on the video load path omitted quant_active, so a quantized video
transformer (an engaged dense transformer_quant, or a GGUF checkpoint) that also
enabled First-Block-Cache without an explicit threshold used the dense bf16 threshold
(0.08) instead of the higher quantized threshold (0.12) the cache helper documents as
needed for quantized transformers to trigger. The advertised quant plus FBCache path
therefore cached far less than intended. Thread quant_active through exactly as the
image path (diffusion.py) does: an engaged transformer_quant or a GGUF transformer both
count as quant-active here.
Review follow-ups on the video inference backend:
- validate_load_request now rejects a -GGUF repo picked as a diffusers
pipeline (no gguf_filename) up front, instead of failing minutes later
in from_pretrained after the GPU owner was already evicted.
- New _detect_load_family helper shared by validate_load_request and
_run_load: when the repo id alone does not carry the family, fall back
to detecting it from the picked GGUF filename, so both paths agree.
- routes/video.py now threads base_repo into validate_load_request so an
untrusted companion repo is refused before the arbiter handoff.
- unload() now drains _generate_lock before _teardown_state so a
cancelled clip actually exits the denoise loop before the VRAM is
reported free.
- load_pipeline re-checks the load token after the generate-lock barrier
and raises if the load was superseded while waiting.
- Pre-commit global mutations (backend flags, gguf compile installs) are
registered per load token and rolled back in _run_load's error path
via _rollback_precommit_globals, so a failed load no longer leaks
process-wide state.
- fp32 memory estimates now apply a 2x dtype scale on non-CPU devices
for pipeline, single-file and companion sizes (bf16 tables assume
2 bytes/param); GGUF quant estimates stay unscaled.
Tests: GGUF-repo-as-pipeline rejection, _detect_load_family fallback and
override semantics; fake route backend accepts base_repo. 66 passed
across test_video_backend, test_video_routes, test_video_families,
test_video_gallery.
The load tail re-ran the already-filtered speed_optims tuple through
.items() as if it were still the raw applied dict from apply_speed_optims.
An empty tuple short-circuited to {} so CPU test runs passed, but on a real
GPU at least channels_last engages, the tuple is truthy, and every load
failed with 'tuple' object has no attribute 'items'. Store the filtered
tuple directly and add a regression test that forces one optimisation to
engage.
tokenizer/chat_template.jinja ships as its own file in the LTX-2 and
HunyuanVideo-1.5 repos and apply_chat_template reads it at generation
time, so a scoped snapshot without it loads fine and then crashes the
first generation.
The trusted-repo allowlist admitted the 720p t2v repack while the only
Hunyuan family entry carried 480p presets, so a 720p load silently
defaulted to 832x480. It now resolves a dedicated entry whose repo-id
alias outranks the generic token (same guider config, verified: both
repos ship guidance 6.0).
A scheduler-wrapped cancel unwinds pipe call by exception and skips the
pipeline's end-of-call maybe_free_model_hooks, leaving onloaded offload
modules on the GPU until the next request; generate() now frees them
before surfacing the cancelled sentinel.
A warm-cache predownload sweep never consults the cancel event (each cached
file returns instantly), so an unload during it was ignored until a cold
file hit the network. The 2.3 detection also only probed bare-file local
repos; resolve directory repos through the same child resolver the loader
uses so their base pull is scoped too.
A 2.3 checkpoint (GGUF or single file) carries the DiT and, with its extras
files, the connectors, both VAEs and the vocoder; only the 2.0 base repo's
scheduler, text encoder and tokenizer are read. Detect 2.3 from the
checkpoint header after the pull, re-estimate, and scope the base
pre-download accordingly (about 6 GB less per fresh install).
A bare from_pretrained snapshot of Lightricks/LTX-2 pulls the whole 314 GB
repo: 170 GB of packaged root checkpoints and a second 50 GB text-encoder
shard set, when the pipeline reads about 93 GB. Build the needed file list
once (shared with the progress estimate so the two cannot disagree),
download it per file with cancellation, and hand from_pretrained the local
snapshot dir. Clamp the progress counter to the estimate so stale cache
blobs can no longer report over 100 percent.
Re-plan memory with the quant steady factor when the bf16 table forces
offload a quantised DiT would not need, mirroring the image dense-quant
path, and fall back to the bf16 plan when quant does not engage. Stream
the second expert under group offload (model and sequential already hook
every module). Fail the load cleanly when quant engages on only one
expert instead of running mixed precision with quant reported off.
Persist guidance_2 in the gallery recipe so A14B clips are reproducible.