unsloth/studio/backend/core/training
Daniel Han db163e55e3 Namespace the trainer conditioning cache per checkpoint, bound the learning rate
- The trainer keyed its persistent conditioning cache on family and
  resolution only, while the keys themselves carry just the caption or
  image content and crop variant. One cache directory reused for two
  checkpoints, or for the same repo at a new revision, let a warm run
  skip loading its encoders and train on the other model's embeddings
  and latent statistics. Namespace on the base checkpoint and its
  resolved revision as well. The revision helper now lives beside the
  cache in diffusion_train_extras and the inference wrapper delegates to
  it, so the two cannot disagree about what counts as the same source.
- The diffusion learning rate only checked positivity, but 1e309 floats
  to inf and satisfies gt, so the route evicted the resident models and
  started AdamW with an infinite rate: the first step destroys the
  adapter while progress looks normal and the result is saved. Bound it
  below 1.0, matching the LLM schema, which rejects inf for the same
  reason.
2026-07-26 11:51:13 +00:00
..
__init__.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
diffusion_dit_trainer.py Namespace the trainer conditioning cache per checkpoint, bound the learning rate 2026-07-26 11:51:13 +00:00
diffusion_lora_trainer.py Tighten comments across the remaining image stack files 2026-07-12 11:46:23 +00:00
diffusion_train_common.py Version the conditioning cache key and reject non-finite flow_shift 2026-07-26 10:51:01 +00:00
diffusion_train_extras.py Namespace the trainer conditioning cache per checkpoint, bound the learning rate 2026-07-26 11:51:13 +00:00
diffusion_training_service.py Tighten comments across the remaining image stack files 2026-07-12 11:46:23 +00:00
fsdp2_design_notes.md Tighten torchao configs and note the FSDP2 design for the DiT trainer 2026-07-20 07:22:49 +00:00
resume.py Fix resume training crash recovery and MLX checkpoints (#6796) 2026-07-21 02:34:58 -07:00
s3_dataset.py feat(studio): implement S3 dataset loading (completes #5951) (#6222) 2026-06-12 14:52:04 +02:00
trainer.py feat(studio): add DoRA support to studio (#7315) 2026-07-24 03:24:16 -07:00
training.py Fix training start NameError, the load-order guard test and CPU-only diffusion tests 2026-07-25 19:18:35 -07:00
worker.py feat(studio): add DoRA support to studio (#7315) 2026-07-24 03:24:16 -07:00