fix(studio): prevent ModuleNotFoundError in dataset.map() on Windows (#4473)
* fix(studio): prevent ModuleNotFoundError in dataset.map() on Windows On Windows, dataset.map() uses "spawn", which requires workers to import compiled modules from disk. Previously, clear_unsloth_compiled_cache() deleted the entire directory, causing workers to crash when looking for UnslothSFTTrainer.py. Changes: 1. Added `preserve_patterns` to cache cleanup to keep `Unsloth*Trainer.py` on Windows while clearing model-specific files. 2. Added the cache directory to PYTHONPATH for spawn workers. Linux/macOS behavior is unchanged. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix spawn-platform coverage, CWD path mismatch, and race condition for PR #4473 - Extend platform guard from win32-only to include macOS (also uses spawn since Python 3.8, same ModuleNotFoundError would occur) - Replace fragile CWD-based PYTHONPATH registration with centralized register_compiled_cache_on_path() that uses the same __file__-relative _CACHE_DIRS already used by cache_cleanup -- fixes path mismatch when studio is launched from a directory other than the repo root - Move PYTHONPATH registration to the top of _train_worker(), before any dataset.map() call (previously it ran late in config assembly, after dataset formatting which also calls dataset.map()) - Update inference.py model-unload to preserve trainer files on spawn platforms, preventing a race where unloading a model via inference tab would delete UnslothSFTTrainer.py while training workers are importing it * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix cache-dir precedence reversal in register_compiled_cache_on_path() Iterating _CACHE_DIRS in forward order while calling insert(0) each time reverses the declared priority: later entries shadow earlier ones. When multiple compiled-cache directories exist, spawned workers could import a stale trainer from the wrong cache. Fix: iterate in reverse so that the highest-priority entry (first in _CACHE_DIRS) is inserted last and ends up at position 0 in sys.path and PYTHONPATH. * fix: harden worker-count helpers against cpu_count=None and desired<=0 - safe_num_proc: guard os.cpu_count() with `or 1`, clamp multi-GPU path with max(1, min(4, desired)), clamp return with max(1, desired) - safe_thread_num_proc: same os.cpu_count() guard and return clamp - Add regression tests (31 L1 unit + 10 sandbox edge-case tests) * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * remove regression tests from PR --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Daniel Han <danielhanchen@users.noreply.github.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com>
This commit is contained in:
parent
62e3de181f
commit
4cedeba8c2
4 changed files with 93 additions and 10 deletions
|
|
@ -521,14 +521,14 @@ def safe_num_proc(desired: Optional[int] = None) -> int:
|
|||
|
||||
visible = get_visible_gpu_count()
|
||||
if visible > 1:
|
||||
capped = min(4, desired)
|
||||
capped = max(1, min(4, desired))
|
||||
logger.info(
|
||||
f"Multi-GPU detected ({visible} visible GPUs) -- "
|
||||
f"capping num_proc {desired} -> {capped} to avoid fork deadlocks"
|
||||
)
|
||||
return capped
|
||||
|
||||
return desired
|
||||
return max(1, desired)
|
||||
|
||||
|
||||
def safe_thread_num_proc(desired: Optional[int] = None) -> int:
|
||||
|
|
@ -551,7 +551,7 @@ def safe_thread_num_proc(desired: Optional[int] = None) -> int:
|
|||
if desired is None or not isinstance(desired, int):
|
||||
desired = max(1, (os.cpu_count() or 1) // 3)
|
||||
|
||||
return desired
|
||||
return max(1, desired)
|
||||
|
||||
|
||||
def dataset_map_num_proc(desired: Optional[int] = None) -> Optional[int]:
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue