Commit graph

632 commits

Author SHA1 Message Date
Daniel Han
176a544e1c Merge branch 'diffusion-train-tab-2' into diffusion-krea2
# Conflicts:
#	studio/backend/core/inference/diffusion.py
2026-07-05 04:43:41 +00:00
Daniel Han
40723c3465 Merge branch 'diffusion-train-precision' into diffusion-train-tab-2 2026-07-05 04:41:58 +00:00
Daniel Han
7e3abefddd Merge branch 'diffusion-train-perf' into diffusion-train-precision 2026-07-05 04:41:57 +00:00
Daniel Han
39f9fa0fee Merge branch 'image-generation' into diffusion-train-perf 2026-07-05 04:41:56 +00:00
Daniel Han
62cae5fe71 Load image pipelines from the prefetched snapshot instead of re-sweeping the hub
The prefetch already scopes the file list (no packaged root singles, no
dtype-variant twins, no ONNX/Flax exports), but from_pretrained was then
called with the hub id, and its own snapshot sweep re-downloaded the
skipped files anyway: 24 GB per FLUX.1 repo and 65 GB on FLUX.2-dev, as
found in the blob cache. Return the snapshot dir from the prefetch (keyed
on the pipeline manifest) and hand it to every pipeline-assembly
from_pretrained site; any prefetch failure keeps the hub id and the old
behavior.
2026-07-05 04:41:43 +00:00
Daniel Han
86b12b2213 Merge branch 'diffusion-train-precision' into diffusion-train-tab-2 2026-07-05 02:38:57 +00:00
Daniel Han
ad8213ba70 Merge branch 'image-generation' into diffusion-train-perf 2026-07-05 02:38:55 +00:00
pre-commit-ci[bot]
20a3650108 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 02:12:27 +00:00
Daniel Han
dc36983562 Merge branch 'diffusion-train-tab-2' into diffusion-krea2 2026-07-05 02:11:25 +00:00
Daniel Han
a99b951c33 Merge branch 'diffusion-train-precision' into diffusion-train-tab-2
# Conflicts:
#	studio/backend/tests/test_diffusion_training.py
2026-07-05 02:11:15 +00:00
pre-commit-ci[bot]
91d7297d41 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 02:10:58 +00:00
Daniel Han
32a77623ba Merge branch 'diffusion-train-perf' into diffusion-train-precision
# Conflicts:
#	studio/backend/core/training/diffusion_train_common.py
2026-07-05 02:10:23 +00:00
Daniel Han
c2ab1a0e61 Merge branch 'image-generation' into diffusion-train-perf
# Conflicts:
#	studio/backend/core/training/diffusion_dit_trainer.py
#	studio/backend/core/training/diffusion_train_common.py
2026-07-05 02:09:39 +00:00
pre-commit-ci[bot]
f6f198fd5f [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 02:07:49 +00:00
Daniel Han
25d9cf9604 Stub diffusers.hooks too in the no-diffusers cache test 2026-07-05 02:04:16 +00:00
Daniel Han
e605075508 Fix diffusion training validation and honor lr_scheduler and batch size in the DiT trainer 2026-07-05 01:51:34 +00:00
Daniel Han
098809d2fe Reject sd-cli batch runs and clear stale output targets before a run 2026-07-05 01:49:19 +00:00
Daniel Han
76eee534ea Gate explicit attention kernels on NVIDIA CUDA and roll back partial FBCache hooks 2026-07-05 01:48:12 +00:00
Daniel Han
f1d9c88606 Validate load modes before eviction and wait out a cancelled denoise on unload 2026-07-05 01:47:02 +00:00
Daniel Han
99eb248607 Merge branch 'diffusion-train-tab-2' into diffusion-krea2 2026-07-05 00:31:33 +00:00
Daniel Han
71c20ded19 Merge branch 'diffusion-train-precision' into diffusion-train-tab-2 2026-07-05 00:31:31 +00:00
Daniel Han
79b97e9ad0 Merge branch 'diffusion-train-perf' into diffusion-train-precision 2026-07-05 00:31:30 +00:00
Daniel Han
ad7d5c6827 Coerce cache_latents and enable_tf32 string flags in the config dict path
The generic Studio config dict path can deliver these flags as strings, and a
non-empty string like "false" is truthy, so an opt-out silently no-ops (the
latent cache still builds, TF32 stays on). Coerce them the same way
gradient_checkpointing already is.
2026-07-05 00:12:47 +00:00
Daniel Han
8c9439fe8f Merge remote-tracking branch 'origin/diffusion-train-tab-2' into diffusion-krea2 2026-07-04 13:49:21 +00:00
Daniel Han
97f90e609c Merge remote-tracking branch 'origin/diffusion-train-precision' into diffusion-train-tab-2 2026-07-04 13:49:19 +00:00
Daniel Han
5b090bc0f0 Merge remote-tracking branch 'origin/diffusion-train-perf' into diffusion-train-precision 2026-07-04 13:49:18 +00:00
Daniel Han
da5f4232e3 Merge remote-tracking branch 'origin/image-generation' into diffusion-train-perf 2026-07-04 13:49:17 +00:00
Daniel Han
c241886c67 Merge branch 'image-generation' of https://github.com/unslothai/unsloth into image-generation 2026-07-04 13:47:00 +00:00
Daniel Han
24de50062c Enable conv-direct in the default native speed profile
Measured on the fresh linux x64 prebuilt (z-image Q8_0, sd-cli, 192 CPU
threads, 512x512, 9 steps, steady state): sampling 56.1s vs 51.3s (about
9 percent faster), VAE decode unchanged, peak RSS identical. The sd.cpp
engine only serves the no-GPU tier, so the default profile now matches
max: --diffusion-fa plus --diffusion-conv-direct.
2026-07-04 13:46:59 +00:00
Daniel Han
ef03dd2780 Merge diffusion-train-tab-2: qwen dense-quant family deny (black frames, measured) 2026-07-04 08:52:32 +00:00
Daniel Han
200543daca Merge diffusion-train-precision: qwen dense-quant family deny (black frames, measured) 2026-07-04 08:52:22 +00:00
Daniel Han
7aac4770cd Merge diffusion-train-perf: qwen dense-quant family deny (black frames, measured) 2026-07-04 08:52:13 +00:00
Daniel Han
92ac400bbb Merge image-generation: qwen dense-quant family deny (black frames, measured) 2026-07-04 08:52:03 +00:00
pre-commit-ci[bot]
ab56d81935 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 08:51:46 +00:00
Daniel Han
7bf80f6a4e Deny fp8/mxfp8/nvfp4 dense quant for the Qwen DiT (black frames, measured)
A 28-pair accuracy gate on a B200 (same-seed vs the dense bf16 reference)
found per-row fp8 dynamic quant renders EVERY qwen-image frame black
(mean luma 0.0000, SSIM 0.016), reproduced identically with on-the-fly
quantize_ on the dense transformer, so it is the model's activation range,
not a checkpoint artifact. mxfp8 shows real semantic damage at 1024px
(CLIP delta mean 0.0146, worst cases 0.064/0.102) and nvfp4 measures
LPIPS mean 0.51. int8 dynamic (per-token scales) is excellent on Qwen:
LPIPS mean 0.069, SSIM 0.958.

The per-scheme smoke probe only proves the GEMM kernel runs, so it cannot
catch model-level breakage. Add _FAMILY_SCHEME_DENY consulted by
select_transformer_quant_scheme: auto skips denied schemes (Qwen lands on
int8) and an explicit denied request returns None, the same GGUF-fallback
contract as an unsupported scheme. Family is threaded from the three
diffusion.py call sites; existing behavior is unchanged for every other
family. 4 new tests; 529 diffusion tests green; CI-sim green.
2026-07-04 08:51:10 +00:00
Daniel Han
466f86c45a Skip the krea-2 forward roundtrip test on hosts without diffusers
spec.forward imports Krea2Pipeline for prepare_position_ids, so the test needs a
real diffusers install; the backend CI matrix runs without one and failed on the
import. Same importorskip guard the sigmas gather test already uses.
2026-07-04 07:26:56 +00:00
Daniel Han
8324cdc407 Merge diffusion-train-tab-2: CI test fixes (diffusers import order, arbiter device pin, sigma-gather skip) 2026-07-04 05:08:18 +00:00
Daniel Han
3fcf218614 Merge diffusion-train-precision: CI test fixes (diffusers import order, arbiter device pin, sigma-gather skip) 2026-07-04 05:07:46 +00:00
Daniel Han
f5d5b09ae5 Merge diffusion-train-perf: CI test fixes (diffusers import order, arbiter device pin, sigma-gather skip) 2026-07-04 05:06:59 +00:00
Daniel Han
c63df7d918 Merge image-generation: CI test fixes (diffusers import order, arbiter device pin) 2026-07-04 05:05:54 +00:00
Daniel Han
f1007fb466 Skip the sigma-gather test when diffusers is not installed
CI runs the backend suite without diffusers; the test checks our index math against
the scheduler's own gather, so it skips rather than fails there.
2026-07-04 05:05:54 +00:00
Daniel Han
1d3aa53d1f Validate the training config before importing diffusers and pin the arbiter test's device
The fp16-on-bf16-family refusal in run_dit_lora_training now fires before the heavy
imports, so a host without diffusers gets the real validation error instead of
ModuleNotFoundError. test_in_progress_returns_409_after_validation_passes pins the
resolved device to cuda because the load route only takes the GPU arbiter for non-CPU
loads, which made the ownership assert host-dependent.
2026-07-04 05:01:58 +00:00
pre-commit-ci[bot]
48e1cbc242 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 03:31:24 +00:00
pre-commit-ci[bot]
d8495d058f [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 03:30:53 +00:00
pre-commit-ci[bot]
f2af2874df [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 03:30:20 +00:00
pre-commit-ci[bot]
32556949ce [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 03:29:49 +00:00
pre-commit-ci[bot]
07f27b23e8 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 03:29:17 +00:00
Daniel Han
2eded64b25 Merge diffusion-train-tab-2: run history robustness, epoch sentinel, torchao probe, image-generation review fixes 2026-07-04 03:25:30 +00:00
Daniel Han
7a8f363c52 Merge diffusion-train-precision: torchao functional probe + image-generation review fixes 2026-07-04 03:25:08 +00:00
Daniel Han
a4a38b5672 Merge diffusion-train-perf: image-generation review fixes 2026-07-04 03:24:44 +00:00