The A14B gate run crashed mid-denoise: accelerate offload hooks move modules with Module.to(), and torchao quantized tensors reject that (aten._has_compatible_shallow_copy_type is unimplemented, raised from compute_should_use_set_data). The 114 GB dual DiT plans model offload by default, so int8 there was a guaranteed crash. The loader now skips quant whenever the plan resolves to any offload policy, logs why, and surfaces the reason in the resolved record; a user who wants both can pin a resident memory mode. A dense DiT under offload beats a crashed one. |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| src-tauri | ||
| __init__.py | ||
| install_llama_prebuilt.py | ||
| install_node_prebuilt.py | ||
| install_python_stack.py | ||
| install_sd_cpp_prebuilt.py | ||
| LICENSE.AGPL-3.0 | ||
| node_prebuilt_pins.json | ||
| package-lock.json | ||
| package.json | ||
| setup.bat | ||
| setup.ps1 | ||
| setup.sh | ||
| Unsloth_Studio_Colab.ipynb | ||