Commit graph

6,520 commits

Author SHA1 Message Date
Daniel Han
5df56c0804 Stub the torchao probe in the precision-mode capability tests
train_precision_modes gates int8/fp8/mxfp8 on has_functional_torchao, and the
Backend CI runner does not install torchao, so the three capability-gating
tests collapsed to nf4/bf16/auto and failed. They exercise the CAPABILITY
gate, not torchao presence: stub the probe functional alongside the CUDA
capability patch. Validated with a torchao-blocked run (22 passed).
2026-07-04 08:21:07 +00:00
Daniel Han
0393e5c398 Merge branch 'diffusion-krea2' into diffusion-train-perf2 2026-07-04 07:27:05 +00:00
Daniel Han
466f86c45a Skip the krea-2 forward roundtrip test on hosts without diffusers
spec.forward imports Krea2Pipeline for prepare_position_ids, so the test needs a
real diffusers install; the backend CI matrix runs without one and failed on the
import. Same importorskip guard the sigmas gather test already uses.
2026-07-04 07:26:56 +00:00
Daniel Han
e39ae8fb58 Merge diffusion-krea2: Dtype rename and empty-state copy 2026-07-04 06:17:58 +00:00
Daniel Han
2459bdbbf1 Merge diffusion-train-tab-2: Dtype rename and empty-state copy 2026-07-04 06:17:57 +00:00
Daniel Han
b219118ff0 Merge diffusion-train-precision: Dtype rename and empty-state copy 2026-07-04 06:17:56 +00:00
Daniel Han
be9aa7c9b3 Merge diffusion-train-perf: Dtype rename and empty-state copy 2026-07-04 06:17:55 +00:00
Daniel Han
168afdf6f1 Merge image-generation: Dtype rename and empty-state copy 2026-07-04 06:17:54 +00:00
Daniel Han
d146209f88 Rename the GGUF compute control to Dtype and simplify the empty-state copy
The always-visible description under the select is gone (the hint tooltip keeps the
full detail) and the no-model gallery placeholder now reads 'Select a diffusion model
to load'.
2026-07-04 06:17:45 +00:00
Daniel Han
8e589f8b27 Merge diffusion-krea2: CI test fixes (diffusers import order, arbiter device pin, sigma-gather skip) 2026-07-04 05:08:51 +00:00
pre-commit-ci[bot]
e6d775d0fc [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 05:08:43 +00:00
Daniel Han
8324cdc407 Merge diffusion-train-tab-2: CI test fixes (diffusers import order, arbiter device pin, sigma-gather skip) 2026-07-04 05:08:18 +00:00
Daniel Han
3fcf218614 Merge diffusion-train-precision: CI test fixes (diffusers import order, arbiter device pin, sigma-gather skip) 2026-07-04 05:07:46 +00:00
Daniel Han
f5d5b09ae5 Merge diffusion-train-perf: CI test fixes (diffusers import order, arbiter device pin, sigma-gather skip) 2026-07-04 05:06:59 +00:00
Daniel Han
c63df7d918 Merge image-generation: CI test fixes (diffusers import order, arbiter device pin) 2026-07-04 05:05:54 +00:00
Daniel Han
f1007fb466 Skip the sigma-gather test when diffusers is not installed
CI runs the backend suite without diffusers; the test checks our index math against
the scheduler's own gather, so it skips rather than fails there.
2026-07-04 05:05:54 +00:00
Daniel Han
1d3aa53d1f Validate the training config before importing diffusers and pin the arbiter test's device
The fp16-on-bf16-family refusal in run_dit_lora_training now fires before the heavy
imports, so a host without diffusers gets the real validation error instead of
ModuleNotFoundError. test_in_progress_returns_409_after_validation_passes pins the
resolved device to cuda because the load route only takes the GPU arbiter for non-CPU
loads, which made the ownership assert host-dependent.
2026-07-04 05:01:58 +00:00
Daniel Han
630689032e Merge branch 'image-generation' of https://github.com/unslothai/unsloth into image-generation 2026-07-04 05:01:57 +00:00
pre-commit-ci[bot]
3204874f7e [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 04:42:28 +00:00
pre-commit-ci[bot]
6ecbb2b8a3 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 04:41:56 +00:00
pre-commit-ci[bot]
96cecf9c47 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 04:41:23 +00:00
pre-commit-ci[bot]
d13ce4c74a [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 04:40:51 +00:00
pre-commit-ci[bot]
285c8fbd20 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 04:40:19 +00:00
Daniel Han
67f8f6cfae Merge diffusion-krea2: grad norm reconciliation + review fixes 2026-07-04 04:38:24 +00:00
Daniel Han
a82fc89d03 Merge diffusion-train-tab-2: grad norm reconciliation + review fixes 2026-07-04 04:38:05 +00:00
Daniel Han
325ae8e3ce Merge branch 'diffusion-krea2' of https://github.com/unslothai/unsloth into diffusion-krea2 2026-07-04 04:38:05 +00:00
Daniel Han
a346a0eb20 Merge diffusion-train-precision: grad norm chart + review fixes
# Conflicts:
#	studio/backend/core/training/diffusion_dit_trainer.py
#	studio/backend/core/training/diffusion_lora_trainer.py
#	studio/backend/core/training/diffusion_training_service.py
#	studio/backend/models/training.py
#	studio/frontend/src/features/images/api.ts
#	studio/frontend/src/features/images/train/diffusion-charts.tsx
#	studio/frontend/src/features/images/train/diffusion-train-panel.tsx
2026-07-04 04:37:58 +00:00
Daniel Han
e9b9bdc5f4 Merge branch 'diffusion-train-tab-2' of https://github.com/unslothai/unsloth into diffusion-train-tab-2 2026-07-04 04:33:44 +00:00
Daniel Han
84a661b363 Merge diffusion-train-perf: grad norm chart + review fixes 2026-07-04 04:33:37 +00:00
Daniel Han
223a546cd8 Merge image-generation: grad norm chart, completion state, Windows caption keys, GGUF compute copy
# Conflicts:
#	studio/backend/core/training/diffusion_dit_trainer.py
#	studio/backend/core/training/diffusion_lora_trainer.py
#	studio/backend/core/training/diffusion_training_service.py
2026-07-04 04:33:30 +00:00
Daniel Han
ff59b43533 Fail clearly when a local Krea 2 dir lacks model_index.json
Falling through to hf_hub_download with a filesystem path as the repo id
raised an opaque HFValidationError; a local dir without the file now
raises FileNotFoundError naming the directory
2026-07-04 04:31:09 +00:00
Daniel Han
2c5955bda8 Coerce num_epochs in normalized() and use utf-8 for run records
num_epochs was only int-coerced for the range check, so a string value
from a dict-built config would reach resolve_train_steps' arithmetic;
normalized() now stores the coerced int. Run record reads/writes pass
encoding utf-8 explicitly so non-ASCII prompts survive on Windows
2026-07-04 04:31:07 +00:00
Daniel Han
8ad8a58742 Add grad norm chart, clearer completion state, Windows caption keys, GGUF compute copy
- Trainers emit the pre-clip gradient norm; the service keeps a bounded
  grad_norm history and the Train tab renders a Grad Norm chart next to
  Loss and LR
- Completed runs show 'Training complete' with a celebratory marker in
  the success color instead of a plain status word
- metadata.jsonl caption keys now match on Windows (as_posix relative
  paths) in both the trainer discovery and the dataset image records
- RMSNorm eager patch skips installation on torch builds without
  F.rms_norm instead of failing at forward time
- GGUF compute description no longer says the GGUF is dequantised: the
  INT8/FP8/FP4 modes load the base model's bf16 transformer and quantise
  that directly; label no longer wraps in the Advanced panel
2026-07-04 04:31:04 +00:00
pre-commit-ci[bot]
003e28730c [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 03:31:57 +00:00
pre-commit-ci[bot]
48e1cbc242 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 03:31:24 +00:00
pre-commit-ci[bot]
d8495d058f [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 03:30:53 +00:00
pre-commit-ci[bot]
f2af2874df [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 03:30:20 +00:00
pre-commit-ci[bot]
32556949ce [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 03:29:49 +00:00
pre-commit-ci[bot]
07f27b23e8 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-04 03:29:17 +00:00
Daniel Han
95f783ae1e Merge diffusion-krea2: Raw training default + stacked review fixes
# Conflicts:
#	studio/backend/core/training/diffusion_train_common.py
#	studio/backend/tests/test_diffusion_dit_trainer.py
#	studio/backend/tests/test_diffusion_training.py
2026-07-04 03:28:32 +00:00
Daniel Han
2eded64b25 Merge diffusion-train-tab-2: run history robustness, epoch sentinel, torchao probe, image-generation review fixes 2026-07-04 03:25:30 +00:00
Daniel Han
7a8f363c52 Merge diffusion-train-precision: torchao functional probe + image-generation review fixes 2026-07-04 03:25:08 +00:00
Daniel Han
a4a38b5672 Merge diffusion-train-perf: image-generation review fixes 2026-07-04 03:24:44 +00:00
Daniel Han
7d6c022489 Merge image-generation: review fixes (engine unload, caption precedence, dataset counts, local diffusers tagging, family LoRA targets)
# Conflicts:
#	studio/backend/core/training/diffusion_dit_trainer.py
2026-07-04 03:24:38 +00:00
Daniel Han
5ea8eb958c Support the torchao 0.17 mxfp8 recipe API
torchao 0.17 removed MXLinearConfig from prototype.mx_formats in favour of
MXFP8TrainingOpConfig.from_recipe shared with MoE training. _mxfp8_training_config
tries the 0.16 API first and falls back to the 0.17 one; both feed quantize_.
mxfp8 still degrades to bf16 with a warning when neither import resolves
2026-07-04 03:23:33 +00:00
Daniel Han
dc290bdf71 Train Krea 2 LoRAs on the undistilled Raw checkpoint by default
Krea's release guidance is to train on Krea-2-Raw and run adapters on
Turbo. Raw now leads the krea-2 training bases (Turbo stays available),
both vendor repos are trust-listed, and load_krea2_pipeline fails fast
with an upgrade hint on diffusers older than 0.39 instead of a bare
AttributeError mid-load
2026-07-04 03:23:30 +00:00
Daniel Han
51de9da488 Fix review findings: run history robustness, epoch-mode sentinel, seed 0, history refresh race
- list_diffusion_runs skips wrong-shape records and the runs route tolerates
  per-record ValidationError so one bad file never breaks the panel
- max_steps: 0 epoch-mode sentinel no longer trips train_steps validation
  before epochs are resolved
- numberField keeps an explicit 0 (Seed, LR warmup) instead of falling back
- previous-runs list refetches once more shortly after a terminal status so
  the just-finished run appears even if the record write races the fetch
2026-07-04 03:23:20 +00:00
Daniel Han
89d99e31af Gate int8 and fp8 on a functional torchao import, not find_spec
The Windows ROCm torchao import stub satisfies find_spec and even lets
from torchao.quantization import quantize_ succeed, but its quantize_ is a
no-op: auto would pick int8, leave the transformer dense, and disable
compile as if it were quantized. has_functional_torchao imports the exact
symbols the int8 path uses and rejects the stub via its sentinel; both the
auto picker and the /info advertised modes now use it
2026-07-04 03:23:16 +00:00
Daniel Han
6f9d8b356c Merge branch 'image-generation' of https://github.com/unslothai/unsloth into image-generation 2026-07-04 03:23:05 +00:00
Daniel Han
bfbb902610 Fix review findings: engine unload before training, caption precedence, dataset caption counts, local diffusers tagging, family LoRA targets
- Unload the ACTIVE image engine (sd_cpp or diffusers) before diffusion training starts, not just the diffusers singleton
- Count metadata.jsonl captions in dataset summaries so metadata-captioned datasets are not reported as uncaptioned
- Sidecar captions now override metadata rows everywhere (grid edits win); trainer and dataset API agree
- Tag local diffusers image checkpoints with text-to-image so they appear in the Images picker
- Family LoRA targets (_FLUX_TARGETS etc) apply when the config carries the generic defaults; explicit overrides still win
2026-07-04 03:23:05 +00:00