unsloth/studio/backend/core/training
Daniel Han 3b16e317b8 studio: fix 'Standard' gradient checkpointing leaving the smart offloader active
The training UI's 'Standard' option sends the string "true". The worker
converted "false"/"none"/"" to False but passed "true" through as a string.
unsloth's _configure_gradient_checkpointing only unpatches the smart
offloader on the (True, False) boolean branch; the string "true" matches
neither that nor "unsloth", so it returns without unpatching and the
offloader stays active. On memory-constrained / unified-memory GPUs that
offload then OOM-crashes even though the user asked for standard
checkpointing (reported on gfx1201 R9700, applies to Strix Halo too).

Normalize "true"/"1"/"yes" -> True so unsloth gets a real bool and
unpatches, matching the existing normalization in trainer.py. "unsloth"
and "mlx" stay strings; booleans pass through unchanged.
2026-07-21 03:52:21 -07:00
..
__init__.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
resume.py Fix resume training crash recovery and MLX checkpoints (#6796) 2026-07-21 02:34:58 -07:00
s3_dataset.py feat(studio): implement S3 dataset loading (completes #5951) (#6222) 2026-06-12 14:52:04 +02:00
trainer.py Fix text-only VLM CPT packing truncation (#7211) 2026-07-20 00:23:37 -07:00
training.py Fix resume training crash recovery and MLX checkpoints (#6796) 2026-07-21 02:34:58 -07:00
worker.py studio: fix 'Standard' gradient checkpointing leaving the smart offloader active 2026-07-21 03:52:21 -07:00