New base_precision config for the DiT trainers: nf4 (unchanged default) |
bf16 | int8 | fp8 | auto, advertised per family + per machine through
/api/train/diffusion/info (precision_modes, recommended_precision,
supports_compile) so the UI can gate the selector.
- bf16: dense transformer + regional torch.compile (auto-armed). The
measured speed mode: 2.3x nf4 on FLUX (1.81 -> 4.12 steps/s), 2.6x on
Z-Image (2.5 -> 6.38 steps/s) on B200, at dense-weight VRAM
(FLUX 24.7 GB / Z-Image 13.6 GB peak vs 10.4 / 4.7 for nf4).
- int8: torchao weight-only int8 on the frozen base, quantized AFTER
add_adapter (quantizing first trips peft 0.18's TorchaoLoraLinear,
which is incompatible with the torchao 0.16 config API). Runs eager:
inductor rejects the int8 subclass training graph (aliased subclass
outputs), so compile is force-disabled for it.
- fp8: torchao convert_to_float8_training on the frozen linears
(filter skips lora_ modules, proj_out, non-divisible-by-16 dims,
pad_inner_dim), applied after add_adapter, compile auto-armed.
Works and round-trips, but measured SLOWER than compiled bf16 at
LoRA-training shapes (FLUX 3.15 vs 4.12 steps/s; Z-Image similar),
so it is an explicit opt-in and auto never picks it.
- auto: free VRAM (measured before load) + dense-size table -> bf16
when it fits with headroom, int8 in the middle band, else nf4.
Prequant bnb repos always resolve to nf4; dense modes on them are
rejected at validation with a pointer to the family's dense base.
Two crashes found and fixed along the way:
- The cuDNN SDPA backend's training graph fails on the FLUX attention
shapes (torch 2.10 + cu130, B200): mha_graph.execute errors, then the
context degrades into illegal memory accesses. The perf-flag guard now
pins flash/mem-efficient SDPA for the run (mathematically equivalent,
snapshot/restored). nf4 escaped it by routing attention differently.
- Regional compile now uses dynamic=True (the inference layer's proven
default): dynamic=False specialisation fused a gemm_and_bias epilogue
that failed with CUBLAS_STATUS_EXECUTION_FAILED on the FLUX training
graph; dynamic=True is also faster (Z-Image 3.84 -> 6.38 steps/s).
Verified: 98 backend tests green (new test_diffusion_base_precision.py:
validation, auto policy table, fp8 filter, compile gating, /info fields);
per-mode 40-step runs on FLUX + Z-Image with loss means inside the nf4
envelope and adapter round-trip generation through the normal LoRA path
for bf16-, fp8-, and int8-trained adapters.
Perf core for the diffusion trainers, defaults preserving the training math:
- Phased model loading: the pipeline now loads without its transformer
(conditioning only), captions are encoded and the text encoders freed,
the VAE latent cache is built and the VAE freed, and only then does the
transformer load. The multi-GB denoiser never shares VRAM with the
encoders, cutting measured peak VRAM on B200: FLUX 17.1 -> 10.4 GB,
Qwen-Image 19.1 -> 12.8 GB, Z-Image 7.3 -> 4.7 GB.
- Latent cache (cache_latents, default on): per-image crop/flip variants
(cache_variants, default 4 vs the single frozen variant of the diffusers
--cache_latents) store the VAE posterior's affine parameters, so every
step still draws a fresh VAE sample; a cached center-crop Z-Image run
matches the uncached one at the bf16 nondeterminism floor.
- True batching: train_batch_size now actually batches the transformer
forward (it was silently 1). nf4 dequant dominates the step cost, so
batch 4 lands near batch-1 step time: 4.0x samples/s on Qwen-Image,
3.1x on FLUX, 2.1x on Z-Image, with multi-seed loss envelopes
overlapping batch-1.
- LR scheduler support in the DiT loop (lr_scheduler / lr_warmup_steps
were accepted but ignored); progress events now report the real
per-step LR.
- TF32 + high fp32 matmul precision under enable_tf32 (default on),
snapshot/restored around the run. cudnn.benchmark is scoped to a
caller opt-in only: autotuning the fp32 VAE convs doubled peak VRAM
on the DiT families for zero steady-state gain.
- Vectorized sigma gathering (drops a per-step Python search loop),
cached FLUX img_ids/guidance, fused torch AdamW fallback, steady-state
samples_per_second (excludes the first-step warmup).
- Regional torch.compile plumbing (compile_transformer off/on/auto with
eager fallback): auto stays off over a bitsandbytes base where compile
is a net loss (27 s warmup, slightly slower steady on Z-Image); it
arms automatically for the dense/quantized speed modes that follow.
- Stop parity with the LLM trainer: /api/train/diffusion/stop accepts an
optional {save} body and the service forwards save=False as a
no-save cancel; a new preparing event surfaces cache-build progress.
- SDXL trainer gets the same latent cache, perf flags, and fused
fallback; its batching, LR schedule, and min-SNR stay as they were.
Verified: 83 backend tests green; per-family 30-40 step runs with
adapter round-trip generation through the normal LoRA path (FLUX,
Qwen-Image, Z-Image all pass).
The base-model select's state could briefly hold the previous family's
repo after a family switch (the reseed effect runs a beat later, and a
value with no matching option makes the browser display the first option
anyway). The request then carried the stale repo: picking Qwen or Z-Image
still sent black-forest-labs/FLUX.1-dev and surfaced FLUX's gated-repo
error under the wrong family. Derive an effectiveBase clamped to the
current family's repos and use it for the select value, the start
request, and the deploy fallback.
Also move the Trigger prompt above Adapter name: the trigger describes
the dataset, the name only labels the output.
Add an Examples group to the training-images dropdown that imports a
curated dataset in one pick, alongside the existing cards. Cards now show
up to three preview thumbnails pulled from the public HF datasets-server
so the set is visible before download. Hide the trigger prompt when every
image already has a caption (a captioned style set needs no trigger), and
turn the training-settings toggle into a ghost button with a rotating
chevron.
Large example datasets (100+ images) rendered every tile at once, so the
caption review grid grew unbounded. Show 24 images per page with < >
chevrons and an x-y of N indicator; a new dataset or refresh resets to
the first page.
Two permissive ~100-image sets for the Train tab: huggan/smithsonian_butterflies_subset
(CC0, the classic diffusers-docs training set, imported as a subject set with a trigger
prompt since its metadata columns are species names not captions) and m1guelpf/nouns
(CC0, captioned pixel-art avatars via the text column). Both cap at 100 images.
The example-dataset cards still overran the ~340px config column: the
license used the Badge component whose baked-in w-fit and whitespace-nowrap
ignored the max-width and truncate, and the grid children had the default
min-width auto so wide content pushed past the column edge and clipped the
Import buttons. Replace the badge with a plain truncating pill span, and
give the config column min-w-0 with overflow-x-hidden so nothing escapes
its width.
When a dataset with images is selected, show a strip of up to 8 sampled
thumbnails with a +N more tile, so users can see what is in the folder
before training. Clicking the strip opens the existing caption review
grid. Samples are drawn evenly across the folder and refresh on dataset
change or after an upload/import.
The Train tab reused the LLM charts section, which also rendered an empty
Grad Norm card and an Eval Loss card showing an Evaluation not configured
placeholder with a red smear. Neither applies to diffusion LoRA training.
Add a diffusion-only two-card view that reuses the loss and learning-rate
cards directly with fixed presentation defaults, and note under the loss
chart that per-step loss is noisy by design so users read the smoothed
line for the trend rather than the raw jitter.
The example-dataset cards used a two-column grid in the ~340px config
column, which wrapped titles one word per line and let the long license
text overrun into the neighbouring card. Switch to one card per row with a
horizontal layout: title with a compact truncated license badge (full text
in the tooltip), a two-line clamped description, and the Import button on
the right.
The Create/Train switch had an icon inside the Train trigger that overhung
the pill corner. Drop the icon, make both triggers a fixed equal width so
the active pill sits flush in the top bar.
The fp32 LoRA parameters and the bnb 4-bit base matmuls need a single
compute dtype during the forward, exactly like the diffusers dreambooth
scripts run under accelerator.autocast. Without it the 4-bit backward on
FLUX.1-dev fails with an illegal-address CUBLAS error partway into the
first step. Z-Image and Qwen-Image smokes are unaffected and the SDXL
path (its own trainer) is untouched.
Replaces the Train LoRA dialog with a top-bar Create | Train segmented control next to
the model selector. Create renders the existing generation workspace unchanged; Train
renders the full-page training panel (unmounted in Create so its polling stops while the
backend run and its retained metric history survive a tab switch). Adds a deploy handler:
loading the trained adapter's base as a pipeline, queueing the adapter so the LoRA
discovery effect applies it once the base is loaded and LoRA-capable for the matching
family (with a mismatch warning), seeding the prompt with the trigger, and switching back
to Create. Removes the now-unused dialog.
New full-page training workspace for the Images tab. Left column configures the run:
model family (FLUX.1-dev, Qwen-Image, Z-Image, SDXL in popularity order, with per-family
VRAM/license notes and defaults, backfilled from the backend families list when present),
base repo, dataset (existing folder, browser upload, or one-click example import), an
in-browser caption labeling grid (per-image thumbnail + caption saved on blur, delete,
uncaptioned highlight), adapter name, trigger prompt, and collapsed training settings.
Right column shows the live run: progress + loss/avg/speed/peak-VRAM readouts, the reused
training loss/LR charts fed from metric_history, and a completion card that deploys the
adapter into Create or starts another run.
Extends the Images training client for the Train tab: the status type now carries
metric_history (step/loss/lr) plus catalog_path/family/base_model/samples_per_second/
peak_memory_gb; the start request gains model_family; and info gains an optional
families list (per-family bases + defaults). Adds typed calls for the dataset
labeling and one-click example endpoints: list images with captions, thumbnail URL,
write/clear a caption, delete an image, list example datasets, and import an example.
Cover the DiT spec table, the QLoRA prequant heuristic, the Z-Image bf16-only
guard, the gated-repo name check, family resolution now that FLUX/Qwen/Z-Image
are trainable (and GGUF repos are rejected as inference-only), the families list
in /diffusion/info, and the gated-base 400 preflight that leaves the GPU
untouched.
The training info endpoint now returns the trainable model families (name,
label, default + allowed base repos, recommended defaults, and a VRAM/access
note) so the Train UI can offer a base picker with realistic guidance. The start
route preflights a gated base repo (HEAD model_index.json with the user's token)
BEFORE freeing resident GPU workloads, so a missing FLUX.1-dev license/token
fails fast with an actionable 400 instead of evicting the loaded model and then
hitting a confusing mid-load 401.
SDXL re-encoded every caption with both CLIP text encoders on every step (pure
waste, since captions are constant) and kept the encoders resident. Precompute
each unique caption's embeddings once, then free the text encoders before the
loop: numerically identical (embeddings are deterministic and this consumes no
torch RNG, so the noise/timestep stream is unchanged) but faster and ~1.5 GB
lighter. Default the optimizer to 8-bit AdamW (bitsandbytes) with an fp32
fallback, halving optimizer state with no meaningful LoRA quality cost. Env
toggles (UNSLOTH_DIFFUSION_NO_PRECOMPUTE / _FP32_OPTIM) let the accuracy guard
A/B the paths.
Extends diffusion LoRA training beyond SDXL to the three popular DiT families
via a single shared flow-matching loop parameterised by small per-family specs
(loading, prompt/latent encoding, transformer forward, save). Verified against
diffusers 0.38.0:
- FLUX.1-dev: 2x2 latent packing + image ids, guidance-embed forward, on-the-fly
nf4 QLoRA of the 12B transformer (the dev repo is gated, so training needs the
user's HF token).
- Qwen-Image: 5D VAE latents normalised by the per-channel latents_mean/std,
img_shapes forward, prequant nf4 base by default (on-the-fly nf4 for the bf16
base).
- Z-Image: list I/O with the reversed timestep convention and a negated
prediction, bf16 only.
The registry (get_trainer) and DiffusionFamily.trainable / train_base_repos now
route these families to the DiT trainer; the SDXL blocklist guard is replaced by
a positive family resolution that also rejects GGUF repos (inference-only) and
still-unsupported families. Per-family defaults + labels + VRAM notes are exposed
via family_train_infos for the Train UI.
Memory: caption embeddings are precomputed once and the text encoders freed
before the loop; gradient checkpointing (non-reentrant, required for bnb 4-bit)
and 8-bit AdamW are on by default.
Cover caption precedence, thumbnail generation and .thumbs exclusion,
caption write/clear, image delete cleanup, path-traversal rejection on
names and filenames, and example import with a mocked datasets.load_dataset
(files plus sidecars written, idempotent second call, cap respected, load
failure mapped to 502).
The Train tab needs to let users caption small datasets in the browser and
pull in a ready-made set to see training work end to end, neither of which
the upload-only endpoint supported.
Add, under /api/train/diffusion/dataset:
- GET {name}/images lists every image with its resolved caption (metadata
beats a per-image sidecar, matching the trainer's discovery order) so
uncaptioned images are visible and flaggable.
- GET {name}/image/{filename} serves an image, with ?thumb=<px> returning a
cached downscaled JPEG kept in a hidden .thumbs subdir (regenerated when
the source is newer) so the labeling grid stays light.
- PUT {name}/caption/{filename} writes, or when blank clears, the .txt
sidecar; DELETE {name}/image/{filename} removes the image plus its
sidecars and thumbnails.
- GET dataset-examples lists a curated, license-labelled registry, and
POST dataset/import-example materializes one into a dataset folder as
numbered images + .txt captions. Two loaders cover the shapes seen in the
wild: streaming rows from datasets.load_dataset (dog-example, Tuxemon) and
a snapshot + jsonl walk for imagefolder repos whose captions live in a
non-standard *.jsonl (the public-domain tarot set). Imports are idempotent
and cap the image count.
Filenames and dataset names are validated against path traversal and pinned
inside the datasets root.
Cover the trainer registry (get_trainer resolves SDXL, unknown family raises),
family resolution (explicit model_family validation, resolved_family on the config),
the metadata sidecar write + scan read with family gating, and the service loss-history
folding (append, bad-point skipping, decimation at cap, family/perf fields) plus the
status route nesting metric_history.
The training service kept only the latest loss, so a live loss chart could show a
single point. Fold each progress event into bounded (step, loss, lr) history arrays
(capped at 4000 points, decimated when full) plus the latest throughput and peak VRAM,
and record the family / base model / catalog path on completion. The status endpoint
returns these as a nested metric_history object the UI can chart directly, and the
start request accepts an optional model_family override.
Split the SDXL trainer into a shared, architecture-agnostic layer so more model
families can be trained without duplicating the plumbing:
- New core/training/diffusion_train_common.py holds the config + validation, dataset
discovery, event emission, stop protocol, adapter publishing, and a lazy trainer
registry (get_trainer). diffusion_lora_trainer.py keeps the SDXL-specific loop and
re-exports the moved names so existing imports are unchanged.
- The SDXL-only base-model blocklist becomes a positive check: the family is resolved
from the base model (or an explicit model_family) via the diffusion family registry,
and a known-but-not-yet-trainable family is refused with a clear message. Unknown
custom names still default to the SDXL trainer.
- DiffusionFamily gains a trainable flag and train_base_repos; SDXL is marked trainable.
DiT families flip on when their trainers land.
- Trained adapters now write a <name>.json metadata sidecar (family, base model, rank,
trigger prompt, ...) that the LoRA scanner reads to family-gate the adapter in the
picker instead of showing it as unknown for every model.
- The training base-model trust allowlist adds the official FLUX.1-dev, Z-Image-Turbo,
and Qwen-Image repos (safetensors-only, no remote code).
The start route freed resident GPU workloads (export, Images pipeline, chat)
before the service validated the config, so a start that was then refused,
now including a non-SDXL base model, tore down the user's loaded model for
nothing. Run the same cheap normalise pass first; the LLM path already
follows this rule via its before_spawn hook.
The dialog assumed users knew the Studio home layout and that only SDXL is
trainable, and hid both facts behind free-text fields. Restructure it around
the three real decisions:
- Base model is a dropdown of the trainable SDXL picks (Base 1.0, Turbo, the
loaded SDXL pipeline when there is one) with a custom repo/path escape
hatch, instead of a bare text field defaulting to a repo id.
- Training images come from an in-browser upload (new dataset endpoints) or
a picker over existing dataset folders with image/caption counts. No shell
access or knowledge of the datasets root is needed any more, and the
captioning rules are explained inline.
- The output field is now Adapter name and the instance prompt is labelled
as the trigger prompt, with a no-captions warning wired to the selected
dataset's actual caption count.
Hyperparameters collapse behind a training settings toggle since the
defaults suit a first run. A completed run says where the adapter went and
offers Done / Train another, and the top-bar button gets an icon and a
plainer description. The dialog title states the SDXL-only scope and that
other families load LoRAs but cannot train them yet.
Training an image LoRA required knowing the Studio home layout and copying
files onto the server by hand, which is the most confusing step of the whole
flow. Two small endpoints fix that:
- GET /api/train/diffusion/info reports the datasets and outputs roots plus
every dataset folder that contains images (with image/caption counts), so
the UI can offer a picker instead of a blind free-text path.
- POST /api/train/diffusion/dataset uploads images and optional caption
.txt / metadata.jsonl files into a named folder under the datasets root,
creating it on first use and accumulating on repeat uploads so large sets
can arrive in batches. Names are validated to a single path component and
files stream to disk under the same per-upload size cap as LLM dataset
uploads. The returned name is a valid data_dir for /diffusion/start.
The trainer only supports the SDXL U-Net, but a FLUX / Qwen-Image / Z-Image
repo or a GGUF filename passed as base_model was accepted and then failed
minutes later inside StableDiffusionXLPipeline.from_pretrained with an
unrelated-looking error. Add a name-based guard in normalized() so known
DiT-family names and .gguf checkpoints are rejected up front, which the API
start route surfaces as an immediate 400 with a message that says exactly
which bases are trainable. Unrecognisable names still pass through so custom
local SDXL checkpoints keep working.