Commit graph

6,649 commits

Author SHA1 Message Date
Daniel Han
e1dd2dda6b Merge remote-tracking branch 'origin/diffusion-train-perf2' into fold-integration
# Conflicts:
#	studio/backend/core/training/diffusion_dit_trainer.py
#	studio/backend/core/training/diffusion_train_common.py
2026-07-07 01:06:42 +00:00
Daniel Han
2dbfd3c4a9 Merge remote-tracking branch 'origin/diffusion-krea2' into fold-integration
# Conflicts:
#	studio/backend/core/training/diffusion_train_common.py
2026-07-07 01:03:21 +00:00
Daniel Han
b52e7a5cc2 Merge remote-tracking branch 'origin/diffusion-train-tab-2' into fold-integration 2026-07-07 01:01:00 +00:00
pre-commit-ci[bot]
a8494e25f3 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-06 15:24:16 +00:00
Daniel Han
0f2f4334e2 Reject dense DiT precisions on a CUDA-absent host before eviction; stabilize family-info tests
The start-route preflight caught the bf16-GPU and int8-torchao requirements but not the dense
precisions' CUDA requirement: on a GPU-less host bf16_unsupported_reason exempts CPU-only, so a
bf16/fp8 (or int8-with-torchao) DiT request passed the preflight, evicted resident workloads, then
raised only in the trainer child. Add the dense-mode CUDA gate mirroring _resolve_base_precision so
the doomed run is rejected up front. Also pin bf16_unsupported_reason in the two positive-path
family-info tests so they are deterministic across GPU types (a non-bf16 CUDA box would otherwise
empty every DiT family's advertised modes).
2026-07-06 15:22:58 +00:00
Daniel Han
9ab68825ee Give Krea-2-Raw its undistilled 52-step recipe instead of the distilled default
Krea-2-Raw is in _TRUSTED_NON_GGUF_REPOS, so it is inference-loadable, but the generic
"krea" generation-defaults key matched it too and applied Turbo's distilled 8-step / no-CFG
recipe, producing degraded output on the undistilled base. Add a more specific krea-2-raw key
(52 steps, guidance 3.5 per the model card) ahead of the generic one, in both the backend
table and the frontend MODEL_DEFAULTS so the Studio UI and the OpenAI images route agree.
2026-07-06 14:00:39 +00:00
Daniel Han
aa54a062ec Gate DiT training on functional torchao for explicit int8; hide always-400 DiT modes on non-bf16 GPUs
The start route preflight only rejected non-bf16 GPUs; an explicit int8 request on
a host with a missing or stub torchao passed the preflight, evicted resident GPU
workloads, then died in the trainer child (its int8 base quantizer has no fallback).
Fold both gates into training_precision_preflight_error so int8-without-torchao fails
fast before eviction. Also empty the advertised DiT precision_modes (and surface the
reason in vram_note, drop compile) whenever the bf16 preflight would reject the family,
so /info never offers an nf4 DiT option the route always 400s.
2026-07-06 13:35:43 +00:00
Daniel Han
a1114bfdc3 Align LR-schedule copy with the diffusion two-card chart layout
DiffusionCharts deliberately renders only Training Loss + Gradient Norm (the LR
curve is the deterministic schedule the user picked), but the settings copy and two
comments still promised a live LR chart. Reword them so the UI no longer references a
chart that was intentionally dropped.
2026-07-06 11:21:06 +00:00
Daniel Han
2d974219bf Deploy Krea adapters on Turbo and use its distilled recipe over the API
- DiffusionFamily gains deploy_base_repo (krea/Krea-2-Turbo): deploying a LoRA
  trained on Raw now previews it on Turbo, not the non-distilled Raw checkpoint.
  Scoped to a same-precision override so it never turns an nf4 train base into a
  larger bf16 deploy load; exposed through family_train_infos -> the Train UI's
  onDeployClick / historical-run deploy resolve the deploy base.
- _GENERATION_DEFAULTS gains a Krea entry (8 steps, 0 CFG) so the OpenAI
  /v1/images/generations route matches the Create UI's documented distilled recipe
  instead of falling through to the generic (9, 0.0).
2026-07-06 11:16:25 +00:00
pre-commit-ci[bot]
cc6d7c96a6 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-06 11:04:54 +00:00
Daniel Han
6bd3e87c6f Gate DiT training precision: deny fp8 for Qwen, gate explicit int8 on torchao, gate advertised dense modes + route on bf16
- normalized() + family_train_infos() mirror the inference fp8 deny for
  Qwen-Image (activation outliers exceed fp8's range and corrupt the trained
  result); int8 stays allowed and the UI no longer advertises fp8 for it.
- _resolve_base_precision() gates an explicit int8 on a FUNCTIONAL torchao, the
  same gate auto and /info already apply, so a missing/stub torchao fails fast
  instead of silently loading dense with compile disabled.
- train_precision_modes() gates the dense modes (bf16/int8/fp8/auto) on
  torch.cuda.is_bf16_supported(), so a non-bf16 CUDA GPU (T4/V100/RTX 20xx) is
  offered only nf4 instead of a start that evicts resident models and then fails.
- start_diffusion_training preflights bf16 support for the DiT families BEFORE
  _free_gpu_for_diffusion_training(), so any DiT start (nf4 included, since the
  trainer requires bf16 unconditionally on CUDA) fails fast without eviction.
2026-07-06 11:04:07 +00:00
pre-commit-ci[bot]
477757eac7 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-06 10:42:55 +00:00
Daniel Han
6384fea272 dit trainer: preserve biases under mxfp8, gate explicit mxfp8 to Blackwell
- The torchao 0.17 MX training path swaps a matched frozen Linear's weight for a wrapper tensor
  whose linear override computes input @ weight_t and drops the bias, so mxfp8'ing a biased frozen
  linear silently loses its bias and corrupts the base output the LoRA regresses against (verified
  on Blackwell: the bias term is fully dropped). Skip biased linears in _mx_module_filter.
- _resolve_base_precision re-checked explicit dense modes against the live device but only rejected
  CPU, so an explicit mxfp8 request on a non-Blackwell CUDA GPU passed and then crashed at the first
  MX GEMM after a full dense-transformer load. /info only advertises mxfp8 on sm100+; mirror that
  gate here and fail fast for a stale or direct client below Blackwell.
2026-07-06 10:41:02 +00:00
Daniel Han
d6ce610324 Merge remote-tracking branch 'origin/diffusion-krea2' into diffusion-train-perf2 2026-07-05 11:52:30 +00:00
Daniel Han
099357cf40 Merge remote-tracking branch 'origin/diffusion-train-tab-2' into diffusion-krea2 2026-07-05 11:52:29 +00:00
Daniel Han
303d38cf98 Merge remote-tracking branch 'origin/diffusion-train-precision' into diffusion-train-tab-2 2026-07-05 11:52:28 +00:00
Daniel Han
78a6ad3fff Merge remote-tracking branch 'origin/diffusion-train-perf' into diffusion-train-precision 2026-07-05 11:52:27 +00:00
Daniel Han
59bf54e975 Merge remote-tracking branch 'origin/image-generation' into diffusion-train-perf 2026-07-05 11:52:25 +00:00
pre-commit-ci[bot]
83bc33d04a [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 11:43:58 +00:00
pre-commit-ci[bot]
293b770c78 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 11:43:23 +00:00
pre-commit-ci[bot]
9caab52bfd [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 11:42:50 +00:00
pre-commit-ci[bot]
1a72a06b1d [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 11:42:17 +00:00
pre-commit-ci[bot]
d6703962d3 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 11:41:43 +00:00
pre-commit-ci[bot]
e800675128 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 11:40:35 +00:00
Daniel Han
66387c1533 Merge branch 'diffusion-krea2' into diffusion-train-perf2 2026-07-05 11:39:16 +00:00
Daniel Han
f543af6d27 Merge branch 'diffusion-train-tab-2' into diffusion-krea2 2026-07-05 11:39:15 +00:00
Daniel Han
5979d68a03 Merge branch 'diffusion-train-precision' into diffusion-train-tab-2 2026-07-05 11:39:14 +00:00
Daniel Han
2d466df322 Merge branch 'diffusion-train-perf' into diffusion-train-precision 2026-07-05 11:39:12 +00:00
Daniel Han
b88d0d49b8 Merge branch 'image-generation' into diffusion-train-perf 2026-07-05 11:39:11 +00:00
Daniel Han
fc1e099124 Harden ControlNet loads, thumbnail cache keys, API training guard, and picker roving keys
Review follow-ups on the image-generation PR:

- ControlNet: resolve_controlnet accepts a bare owner/name repo without the
  non-GGUF base trust gate, and _controlnet_pipe hands it straight to
  from_pretrained. A malicious pickle .bin would deserialize on load, so run
  the same Hugging Face malware preflight (evaluate_file_security) the chat and
  export loaders use before any remote ControlNet load; local dirs are exempt.
- Dataset thumbnails: key the cache on the full filename instead of the stem so
  sample.png and sample.jpg no longer collide on one .thumbs file (which could
  serve or delete the wrong image); the delete cleanup globs the same key.
- Diffusion training start: mirror start_training's API-key guard so an API
  client cannot start training (which frees VRAM by unloading chat) while an
  inference request is streaming; it now returns 409 before any GPU is freed.
- Model picker: include the curated safetensors row keys in the recommended
  roving key list so arrow-key navigation reaches those rows instead of hitting
  the duplicate option-missing id.

Tests: ControlNet malware gate (remote blocked before from_pretrained, local
skipped), thumbnail same-stem cache separation, API-key diffusion-start 409
before GPU free. Full diffusion suites green.
2026-07-05 11:36:58 +00:00
Daniel Han
d329e5db37 Merge branch 'diffusion-krea2' into diffusion-train-perf2 2026-07-05 08:55:40 +00:00
Daniel Han
369a792b04 Merge branch 'diffusion-train-tab-2' into diffusion-krea2 2026-07-05 08:55:39 +00:00
Daniel Han
1f8315c698 Merge branch 'diffusion-train-precision' into diffusion-train-tab-2 2026-07-05 08:55:38 +00:00
Daniel Han
dd6be63d1a Merge branch 'diffusion-train-perf' into diffusion-train-precision 2026-07-05 08:55:37 +00:00
pre-commit-ci[bot]
1c40fa855f [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 07:59:24 +00:00
pre-commit-ci[bot]
d0f7dad7ec [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 07:58:51 +00:00
pre-commit-ci[bot]
b1fdefb43d [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 07:57:51 +00:00
pre-commit-ci[bot]
f343eadcd7 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 07:57:16 +00:00
Daniel Han
466f853f14 Merge branch 'diffusion-krea2' into diffusion-train-perf2 2026-07-05 07:56:51 +00:00
Daniel Han
4b93558ad4 Merge branch 'diffusion-train-tab-2' into diffusion-krea2 2026-07-05 07:56:47 +00:00
pre-commit-ci[bot]
bc38ca397e [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 07:56:43 +00:00
Daniel Han
69f32b9586 Merge branch 'diffusion-train-precision' into diffusion-train-tab-2 2026-07-05 07:56:20 +00:00
Daniel Han
3925aea07f Merge branch 'diffusion-train-perf' into diffusion-train-precision 2026-07-05 07:56:14 +00:00
Daniel Han
8c00f81a5d Size-gate the automatic diffusion latent cache
The latent cache holds two fp32 posterior tensors per crop/flip variant per
image, pinned on CUDA hosts, so datasets with thousands of images can exhaust
host or pinned memory with no fallback. Estimate the cache size from the first
real encoded latent and fall back to per-step VAE encoding when it exceeds a
4 GiB budget. UNSLOTH_DIFFUSION_FORCE_LATENT_CACHE bypasses the gate; the
existing UNSLOTH_DIFFUSION_NO_LATENT_CACHE opt-out is unchanged.
2026-07-05 07:53:12 +00:00
Daniel Han
da3a79468e Use permutation-cycle index sampling in diffusion trainers and guard non-object run records
Replace the with-replacement per-batch index draw in the SDXL and DiT LoRA
trainers with a shared PermutationBatchSampler that visits every image once per
cycle before repeating, so short runs cover the whole dataset. The sampler
reshuffles from the run's rng so the index stream stays seed-deterministic.

Guard the diffusion run detail route against a valid-JSON non-object record,
which previously raised TypeError and returned a 500; it now 404s like the list
path's shape check.

Add regression tests for both.
2026-07-05 07:49:30 +00:00
Daniel Han
715d4965c8 Merge branch 'diffusion-krea2' into diffusion-train-perf2 2026-07-05 07:41:32 +00:00
Daniel Han
866e03d4ce Merge branch 'diffusion-train-tab-2' into diffusion-krea2 2026-07-05 07:41:31 +00:00
Daniel Han
882b5354c4 Merge branch 'diffusion-train-precision' into diffusion-train-tab-2 2026-07-05 07:41:30 +00:00
Daniel Han
8c4cdcd385 Merge branch 'diffusion-train-perf' into diffusion-train-precision 2026-07-05 07:41:29 +00:00
Daniel Han
3977f1a71d Merge branch 'image-generation' into diffusion-train-perf 2026-07-05 07:41:28 +00:00