DGX Spark / N1X: enable expandable_segments to reduce UMA fragmentation
Adds patch_dgx_spark_memory_config() (models/_utils.py), applied at import: on Spark-class machines it sets PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True so the CUDA caching allocator grows segments in virtual address space instead of fragmenting the shared unified-memory pool. More of the pool stays usable for weights/activations -> fewer fragmentation OOMs and headroom for larger models / longer sequences. Pure memory management: computed values are unchanged, so accuracy is unaffected (verified: gemma-3-270m losses identical with/without it). Regression-safe: gated by is_dgx_spark() (strict no-op on x86 NVIDIA, AMD/ROCm, Intel, Mac/MLX, discrete aarch64). Uses setdefault semantics -- only appends when expandable_segments is absent, never overrides a user's PYTORCH_CUDA_ALLOC_CONF; opt out with UNSLOTH_NO_EXPANDABLE_SEGMENTS=1. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
parent
4308d254f5
commit
1144d972d0
1 changed files with 25 additions and 0 deletions
|
|
@ -1009,7 +1009,32 @@ def patch_dgx_spark_caching_allocator_warmup():
|
|||
_mu.caching_allocator_warmup = _noop
|
||||
|
||||
|
||||
def patch_dgx_spark_memory_config():
|
||||
"""Memory-efficiency default for Spark UMA (accuracy-neutral, gated).
|
||||
|
||||
Enables the CUDA caching allocator's `expandable_segments` mode so segments can
|
||||
grow in virtual address space instead of fragmenting the shared unified-memory
|
||||
pool -- more of the pool stays usable for weights/activations (fewer
|
||||
fragmentation OOMs; headroom for larger models / longer sequences). Pure memory
|
||||
management: it never changes any computed value, so accuracy is unaffected.
|
||||
|
||||
Strictly no-op off-Spark (gated by `is_dgx_spark()`). Respects an existing
|
||||
PYTORCH_CUDA_ALLOC_CONF (only appends `expandable_segments` when absent, never
|
||||
overrides a user's setting) and an explicit opt-out
|
||||
(UNSLOTH_NO_EXPANDABLE_SEGMENTS=1). Must run before the first CUDA allocation;
|
||||
`import unsloth` precedes model load, so it is set in time for normal use.
|
||||
"""
|
||||
if not is_dgx_spark():
|
||||
return
|
||||
if os.environ.get("UNSLOTH_NO_EXPANDABLE_SEGMENTS") == "1":
|
||||
return
|
||||
conf = os.environ.get("PYTORCH_CUDA_ALLOC_CONF", "")
|
||||
if "expandable_segments" in conf:
|
||||
return # user already configured it -- do not override
|
||||
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = (conf + "," if conf else "") + "expandable_segments:True"
|
||||
|
||||
|
||||
patch_dgx_spark_memory_config()
|
||||
patch_dgx_spark_caching_allocator_warmup()
|
||||
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue