First diffusion training path in Studio: train a LoRA on the SDXL U-Net from an
image + caption dataset and export it as a diffusers .safetensors that the existing
diffusion LoRA loader (and any diffusers pipeline) can load.
core/training/diffusion_lora_trainer.py:
- DiffusionLoraConfig with validation/defaults (rank, alpha, targets, lr, steps, grad
accumulation, resolution, min-SNR gamma, gradient checkpointing, lr scheduler, seed,
mixed precision).
- discover_image_caption_pairs: captions from metadata.jsonl / captions.jsonl, per-image
.txt/.caption sidecars, or a dreambooth instance_prompt fallback (pure, unit-tested).
- run_diffusion_lora_training: the loop -- freeze base, PEFT-wrap the U-Net attention
projections, VAE-encode (fp32 VAE to avoid the SDXL fp16 overflow), sample noise +
timesteps, predict, MSE loss with optional min-SNR weighting (epsilon / v-prediction),
AdamW + get_scheduler + grad accumulation + grad clipping, then export via
save_lora_weights. Emits worker-protocol events (model_load_*, progress, complete) and
polls should_stop for a clean stop with a partial save.
- run_diffusion_training_process: mp.Queue subprocess adapter (event_queue / stop_queue),
so the training worker can spawn it; plus a CLI entry point.
Only SDXL (U-Net) is trained here; DiT families and the Studio UI form + route wiring are
follow-ups. The trainer is decoupled and worker-ready.
Tests: test_diffusion_lora_trainer.py covers caption discovery (metadata / sidecar /
instance prompt / skip-uncaptioned / errors), config normalisation + validation, the SDXL
add-time-ids, and the dict->config adapter. Verified live on GPU: a 60-step SDXL LoRA run
lowers the loss, exports a ~45 MB adapter, and loading it back shifts generation from
baseline (mean abs pixel diff ~55/255).