Re-plan memory with the quant steady factor when the bf16 table forces offload a quantised DiT would not need, mirroring the image dense-quant path, and fall back to the bf16 plan when quant does not engage. Stream the second expert under group offload (model and sequential already hook every module). Fail the load cleanly when quant engages on only one expert instead of running mixed precision with quant reported off. Persist guidance_2 in the gallery recipe so A14B clips are reproducible. |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| src-tauri | ||
| __init__.py | ||
| install_llama_prebuilt.py | ||
| install_node_prebuilt.py | ||
| install_python_stack.py | ||
| install_sd_cpp_prebuilt.py | ||
| LICENSE.AGPL-3.0 | ||
| node_prebuilt_pins.json | ||
| package-lock.json | ||
| package.json | ||
| setup.bat | ||
| setup.ps1 | ||
| setup.sh | ||
| Unsloth_Studio_Colab.ipynb | ||