Krea's release guidance is to train on Krea-2-Raw and run adapters on
Turbo. Raw now leads the krea-2 training bases (Turbo stays available),
both vendor repos are trust-listed, and load_krea2_pipeline fails fast
with an upgrade hint on diffusers older than 0.39 instead of a bare
AttributeError mid-load
Inference:
- krea-2 DiffusionFamily (Krea2Pipeline / Krea2Transformer2DModel, base
krea/Krea-2-Turbo, bf16 only, no GGUF/sd.cpp mapping yet)
- Per-component pipeline loader (core/inference/diffusion_krea2.py): the krea
repo is exported with transformers 5.2, so the tokenizer config
(extra_special_tokens as a list, no slow-tokenizer vocab files) and the
text encoder rope settings (rope_parameters vs rope_scaling) need explicit
compat on the 4.x line; values are copied verbatim and equal the 4.x
Qwen3-VL defaults, so the math is unchanged. from_pretrained also
type-checks the tokenizer against the declared slow class, so the pipeline
is assembled through its constructor with the model_index init config
(is_distilled carries Turbo's fixed mu=1.15 schedule)
- Trust allowlist entry, curated picker entry + 8 step / cfg 0 defaults,
int8 exclusion token for the M=1 Krea2TimestepEmbedding projection
Training:
- krea-2 _FamilySpec in the DiT trainer: phased conditioning/transformer
load through the compat loader, shared Qwen-Image VAE latent path,
fixed-512 text embeds (static shapes, plain concat collate), inline 2x2
latent packing + shared position grid, the authors' recommended LoRA
target set and rank/alpha 32, lr 3e-4, 512px presets
- GPU smokes on B200: nf4 2.9 steps/s at 11.5 GB, bf16 3.4 steps/s at
30.1 GB, bf16 + regional compile 5.2 steps/s; adapter round-trip
generation verified
Split the SDXL trainer into a shared, architecture-agnostic layer so more model
families can be trained without duplicating the plumbing:
- New core/training/diffusion_train_common.py holds the config + validation, dataset
discovery, event emission, stop protocol, adapter publishing, and a lazy trainer
registry (get_trainer). diffusion_lora_trainer.py keeps the SDXL-specific loop and
re-exports the moved names so existing imports are unchanged.
- The SDXL-only base-model blocklist becomes a positive check: the family is resolved
from the base model (or an explicit model_family) via the diffusion family registry,
and a known-but-not-yet-trainable family is refused with a clear message. Unknown
custom names still default to the SDXL trainer.
- DiffusionFamily gains a trainable flag and train_base_repos; SDXL is marked trainable.
DiT families flip on when their trainers land.
- Trained adapters now write a <name>.json metadata sidecar (family, base model, rank,
trigger prompt, ...) that the LoRA scanner reads to family-gate the adapter in the
picker instead of showing it as unknown for every model.
- The training base-model trust allowlist adds the official FLUX.1-dev, Z-Image-Turbo,
and Qwen-Image repos (safetensors-only, no remote code).
- Persist the actual output image size in the gallery recipe instead of the
request sliders: Transform/Inpaint/Edit derive the size from the uploaded
image, Extend grows the canvas, and Upscale resizes it, so the sliders
recorded (and later restored) the wrong dimensions for those workflows.
- Reject a remote '*-GGUF' repo loaded as a full pipeline (no single-file
name) in validate_load_request, so the unloadable pick fails before chat is
evicted rather than deep in from_pretrained.
- Only publish an image-conditioned from_pipe wrapper to the shared aux cache
when the load is still current: from_pipe runs under the generate lock but
not the state lock, so an unload racing its construction could otherwise
cache a wrapper over torn-down modules that a later load would reuse.
- Verify the Windows CUDA runtime archive checksum before extracting it, like
the main sd-cli archive, so a corrupt or tampered runtime is rejected rather
than extracted next to the binary.
Check cancellation immediately after a ControlNet from_pretrained and before
any device placement, so an unload/eviction that raced the download does not
allocate several GB onto the GPU after the load was already cleared.
Require a loadable weight or shard index (not just config.json) before a local
ControlNet folder is advertised, so an interrupted copy is hidden instead of
failing deep in from_pretrained as a generic 500.
Do not record a strength-0 ControlNet in the gallery recipe: it is treated as
disabled and skipped, so the image is unconditioned and the metadata must not
claim a ControlNet was applied.
Build the control-type picker from the selected ControlNet's advertised
control_types instead of a hardcoded passthrough/canny pair, so a union model
with a precomputed depth or pose map sends the correct control_mode.
Reject LoRA on a torch.compile'd diffusers transformer (Speed=default/max):
diffusers requires the adapter loaded before compilation, so applying one to
the already-compiled module fails with adapter-key mismatches. The status
gate now hides the picker and generate raises a clear message instead.
Convert a cancelled Hub LoRA download (RuntimeError Cancelled) to the
diffusion cancellation sentinel in resolve_specs, so an unload/superseding
load during resolution maps to a 409 instead of a generic server error.
Drop weight-0 LoRA rows before the native support gate so a request carrying
only disabled adapters stays a no-op on families where native LoRA is
unsupported, matching the diffusers path.
Reject duplicate LoRA ids in the request model: both apply paths suffix
colliding names, so a repeated id would stack the same adapter past its
per-adapter weight bound.
Strip all user-typed <lora:...> prompt tags on the native path (only the
selected adapters are materialized in the managed lora-model-dir, so an
unselected tag can never resolve), and restore saved LoRA selections from a
gallery recipe so restore reproduces a LoRA image.
Keep diffusion.py importable without torch: the compile/arch patch modules
import torch at module level, so import them lazily at their load/unload
call sites instead of at module load. This restores the torchless contract
so get_diffusion_backend() works on a CPU/native sd.cpp install.
Match family reject keywords and aliases as whole path/name segments, not
raw substrings, so an unrelated word like edited, edition, or kontextual no
longer misroutes or hides a valid base image model, while supported edit
families (Qwen-Image-Edit, FLUX Kontext) still resolve. Mirror the same
segment matching in the picker task filter.
Route FLUX.2-dev native guidance through --guidance like the other FLUX
families rather than --cfg-scale. Reject native upscale requests that have
no input image. Read image header dimensions and reject over-limit inputs
before decoding pixels, so a crafted small-payload image cannot spike
memory. Reject an upscale that would shrink the source below its input
size. Validate the model_kind against the filename extension before the
GPU handoff. Estimate a local diffusers pipeline's size from its on-disk
weights so auto memory planning does not skip offload and OOM. Report
workflows: [txt2img] from the native backend status so the Create tab
stays enabled for a loaded native model. Clamp the outpaint canvas to the
backend's 4096px decode limit.
Adds regression tests for segment matching and kind/extension validation.
Backend:
- validate_load_request rejects a non-.gguf single-file name before the GPU
handoff, so a family-looking repo paired with README.md no longer evicts the
chat model and only fails in the background load.
- detect_family scopes the edit/kontext/inpaint keyword check to the model id
or filename basename, not arbitrary parent directories, so a valid
text-to-image file under a folder named edit is no longer rejected.
- the images gallery listing skips records that fail schema validation, so one
corrupt or hand-dropped PNG can no longer 500 the whole endpoint.
- _terminate reaps the killed sd-cli child so cancellation and timeout paths do
not leak zombie process-table entries.
- the images load route gates the chat-eviction handoff on the resolved device
being non-CPU, so a CPU-only diffusers fallback no longer evicts a resident
chat model for a load that cannot use the GPU.
Frontend:
- treat Images as a chat-like full-height route (no outer padding or scroll) so
its picker is not pushed down and the gallery is not clipped.
- allow /images under the chat-only guard so the native CPU/MPS image path is
reachable on the no-GPU hosts it was built for.
- roll the optimistic quant label back when a same-repo swap fails after the
load started, so the selector never advertises a quant that is not loaded.
A GGUF-quantized transformer's leading parameters are packed uint8 storage,
so reading next(parameters()).dtype handed nn.Module.to() an integer dtype
and every image-conditioned generation on a GGUF model (Qwen-Image-Edit)
failed with a 500. Probe the parameters for the first floating dtype, treat
an all-integer module as a no-op, and also catch TypeError so an unexpected
dtype can never break generation. Regression test included.
plan_diffusion_memory only applies the legacy cpu_offload override when no
memory_mode was supplied, matching the documented API contract that
memory_mode overrides cpu_offload when set; an explicit fast request now
stays resident even if the old flag is also enabled.
The transformer-quant dense path fetches the base repo's transformer/
shards inside the locked finalize phase, where unload and cancellation
cannot preempt the multi-GB download. The load worker now widens the
preemptible prefetch to include those shards when that path can actually
run: quant requested and supported for the device, scheme resolvable, and
no pre-quantized checkpoint shortcutting the dense build.
Memory planning and dense-quant path: size a local diffusers base's
resident companions from its on-disk VAE and text-encoder weights instead
of folding them to zero, feed the distilled variant hint into the runtime
headroom estimate so turbo and schnell models are not over-reserved, place
group-offload companions resident before attaching the transformer hooks
so a failed placement falls back to whole-module offload instead of
crashing, and bail out of the dense transformer download before it starts
when the requested quant scheme is unsupported so the load falls back to
GGUF cleanly.
sd.cpp stack: scrub the native path lease secret from sd-cli child env,
redact native load-progress errors, forward the resolved accelerator when
auto-installing a forced-native binary, release stale diffusion GPU
ownership on CPU-native loads, and remove the sd.cpp install tree on
uninstall.
Prequant and scripts: reject prequant artifacts missing base_model_id
when a base is requested, expanduser before checkpoint existence checks,
record and validate the int8 exclusion filter and fp8 fast-accum in
checkpoint metadata, make verify_prequant_backend allowlist its local
checkpoint and fail on missing or bad LPIPS and on load-peak regressions,
average only finite PSNR values in diffusion_quality, and reset the
process-wide attention backend between perf probe variants.
API and UI: normalize attention_backend casing before Literal validation,
close hidden popovers when leaving the Images page, and clear the stale
quant label when loading a direct local GGUF file.
Backend:
- Sanitize a blank hf_token to None in begin_load and load_pipeline, so the
default empty Studio token loads anonymously instead of 401ing as an explicit
empty credential.
- Free the ACTIVE diffusion engine before LLM training and in the delete-cached
guard: on a native (sd_cpp) selection the diffusers singleton reports
unloaded, so training could start against a live sd-cli generation and
delete-cached could remove a GGUF the native engine is using. Both now go
through diffusion_engine_router.get_active_diffusion_engine().
- Refuse delete-cached while a background image load is downloading the repo
(or its companion base): status().loaded is False in that window, but the
delete would yank blobs from under the in-flight assembly. Both engines
expose the in-flight ids via a new loading_repo_ids().
- Cap request seeds at 2**53-1: seeds round-trip through JSON gallery recipes,
where JavaScript rounds larger integers, so a restored recipe generated a
different image. Random seeds were already masked to this range.
- Add the task field to CachedModelRepo: the handler sets it for cached
diffusers image repos but response_model silently dropped it, letting
image-only repos pass the chat picker's task gate.
Frontend:
- Offset sequential run seeds by the batch size: the native engine seeds image
j of a run at seed+j, so a +1 run offset regenerated the previous run's
batch-mates.
- Revert the optimistic quant selection when a load fails to start.
- Stop disabling the Images page on chat-only hosts: the native sd.cpp engine
exists exactly for the no-GPU route.
Re-reviewed each merged phase PR against this branch's tip and fixed what is
still real:
- A superseded background load no longer cancels the current model's in-flight
generation: the load-token check now runs BEFORE the cancel signal, with a
re-check under the generate lock (Phase 1 review).
- enable_model_cpu_offload / enable_sequential_cpu_offload now forward the
resolved target device; diffusers defaults to CUDA, which broke offloaded
loads on non-CUDA accelerators such as Intel XPU (Phase 2 review).
- build_sd_cpp_command rejects a None prompt (str(None) previously slipped
into argv as the literal "None") and a mask without an init image, which
is an invalid sd-cli inpaint invocation (Phase 4/6 review).
- The dense-quant OOM fallback drops the caught exception before
clear_gpu_cache(): the traceback pinned the partially built dense
transformer, so the VRAM this cleanup exists to reclaim stayed allocated
through the GGUF rebuild (Phase 8 review).
- Pre-quantized transformers (built via from_config) are eval()'d to match
the from_pretrained paths, so train-mode layers cannot make prequant
inference nondeterministic (Phase 9 review).
- FBCache state is reset before each generation when a step cache is engaged:
diffusers never clears the stateful first-block residuals on the resident
transformer, so a resolution or batch change on the next request hit a
shape mismatch, and an unchanged request could reuse stale residuals
(Phase 12 review).
Each fix carries a regression test; the full diffusion battery passes.
Addresses review findings on the SDXL family:
- Reject a GGUF load for single_file_is_pipeline families (SDXL) in validate_load_request,
before the route evicts the current model; SDXL has no transformer-only GGUF variant.
- Skip base-repo weight files when a whole-pipeline single file is loaded: from_single_file
(config=base) needs only the base config/tokenizer/scheduler, so a local .safetensors no
longer triggers a multi-GB base download.
- Remove the SDXL refiner from the non-GGUF trust allowlist: it is an img2img-only pipeline
but this backend loads every sdxl repo as the base txt2img pipeline.
- Normalize a blank/whitespace hf_token to None once in load_pipeline so every load branch
degrades to anonymous instead of erroring on a malformed token.
- Read the denoiser dtype from a parameter (compile-wrapped modules may lack .dtype) and
access state.family.denoiser_attr directly.
Adds/updates regression tests for the trust allowlist, GGUF rejection, and base-config filter.
- _is_trusted_diffusion_repo: wrap Path.exists() so a repo id with invalid
characters (or a bare owner/name id) can't raise OSError; treat any failure as
not-a-local-path and fall through to the unsloth/ allowlist. validate_load_request
still raises the clear FileNotFoundError for a genuinely missing local pick.
- generate(): reject mask_image / upscale / reference_images supplied without an
input image, and reject reference_images on a family that does not support
reference conditioning, instead of silently degrading to txt2img / img2img.
Address review findings on the LoRA path:
- resolve_one: normalise a blank/whitespace hf_token to None (anonymous access)
and reject a client-supplied weight file with traversal / absolute path.
- resolve_specs: convert FileNotFoundError from an unknown/stale id to ValueError
so the route returns 400 instead of a generic 500.
- _scan_local: disambiguate local adapters that share a stem (foo.safetensors vs
foo.gguf) so each is uniquely addressable.
- inject_prompt_tags: the backend-validated weight now wins over a user-typed
<lora:ALIAS:...> for a selected adapter; unselected user tags are left alone.
- diffusers _apply_loras: reject a .gguf adapter with a clear error before touching
the pipe (diffusers loads safetensors only).
- _unload_locked: drop the explicit unload_lora_weights() on teardown; the pipe is
dropped wholesale (freeing adapters), so the previous call could race an in-flight
denoise on the same pipe.
- Images page: use a stable LoRA key and clear the selection (not just the options)
when the catalog refresh fails.
- resolve_controlnet enforces catalog family compatibility so a direct API call
cannot load a ControlNet built for another family through the wrong pipeline.
- Unknown ControlNet ids now surface as a 400 (call site maps FileNotFoundError
to ValueError) instead of a generic 500.
- strength 0 disables ControlNet entirely, so a no-op selection never pays the
download / VRAM cost; the control image is decoded and validated BEFORE the
ControlNet is resolved or built, so a malformed image fails fast for the same reason.
- ControlNet loads use the base compute dtype (state.dtype is a display string,
not a torch.dtype, so it silently fell back to float32) and honor the base
offload policy via group offloading instead of forcing the module resident.
- Empty/malformed HF token coerced to anonymous access.
- Flux Union ControlNet control_mode mapped from the selected control type.
- resolve_controlnet drops the unused hf_token/cancel_event params.
- ControlNetSpec validates guidance_start <= guidance_end (clean 422).
- Images UI ControlNet Select shows its placeholder when nothing is selected.
Adds regression tests for family enforcement and the union control-mode map.
A full-pipeline prefetch kept every repo file outside assets/, so an official
repo that ships multiple formats (SDXL Base: fp16 variants, ONNX, OpenVINO,
Flax, a top-level single-file twin) downloaded tens of GB from_pretrained never
loads. Skip non-torch exports and dtype-variant twins in
_pipeline_file_downloaded, and drop a component .bin when the same directory
carries a picked safetensors weight (diffusers' own preference).
Two review findings on the ControlNet path:
- resolve_controlnet's bare-repo fallback accepted any id with a slash, so a
path-shaped id (/tmp/x, ../x) reached from_pretrained as a local directory.
Restrict the fallback to a strict owner/name HF repo id shape.
- _controlnet_pipe now re-checks the cancel event after the blocking
from_pretrained: an unload that raced the download had already cleared the
caches, so caching the late module would pin it past the unload.
Three chained bugs that made Z-Image (and other GGUF DiTs) crash at generation
on anything but a huge, fully-idle GPU. Verified end to end on an RTX 6000 Ada:
Q2_K now plans resident and generates a real 1024x1024 PNG on both the resident
and forced-group-offload paths.
- Memory planner over-estimated the GGUF transformer's resident size. diffusers
keeps GGUF weights PACKED (uint8 GGUFParameter) and dequantises per-matmul
transiently, so resident VRAM is ~= the on-disk size, not the unpacked bf16
size (measured: Q2_K 3.64->3.68 GiB, Q8_0 7.22->7.25 GiB). The old per-quant
expansion (x8 for Q2) over-estimated ~7.6x, so a 3.6 GB model on a 48 GB-free
card was judged a "tight fit" and forced into group offload. Replace the
multiplier table with estimate_gguf_resident_mib = storage * 1.05 (matches
diffusers' own get_memory_footprint of a loaded GGUF model).
- torch.compile with fullgraph=True crashed under CPU offload: group/model/
sequential offload installs a @torch.compiler.disable'd ModuleGroup.onload_
hook, which graph-breaks. Drop fullgraph when offloading is planned, same as
the existing step-cache case (fullgraph = not (cache_active or offload_active)).
This mirrors diffusers' documented compile+offload guidance.
- compile_repeated_blocks compiles one graph per distinct block shape, but
Z-Image's "repeated" blocks are heterogeneous (~11 variants), above dynamo's
default recompile_limit of 8, so a resident load hard-errored under fullgraph.
Raise the limit (diffusers' documented fix for regional-compile recompilation).
Confirmed force_parameter_static_shapes=False is the wrong lever: same variant
count, ~6x slower compile.
Also drops the now-dead infer_gguf_quant_label / gguf_filename plumbing and adds
regression tests for the estimate and the offload fullgraph drop.