The loader used to plan memory from the GGUF file size and only offer the dense transformer-quant fast path when that plan was already resident, so on a card where the GGUF forced offload the int8/fp8 build (roughly half the bf16 bytes, or exactly the quantised size when a pre-quantized checkpoint exists) was never attempted. diffusion_auto_policy.py is a pure decision layer: a bf16-resident component table per family (transformer / text encoders / VAE, with base-repo overrides for the multi-size families), per-scheme size factors with separate steady and transient (build peak) numbers, and resolve_dense_quant_candidate which the loader now uses to re-plan memory against the candidate artifact before settling for offload. The engaged plan is adopted only when the dense build succeeds; the GGUF fallback keeps its own plan. Status now carries a resolved provenance record per Advanced control (value, source auto or explicit, reason) so the UI can label backend decisions. |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| src-tauri | ||
| __init__.py | ||
| install_llama_prebuilt.py | ||
| install_node_prebuilt.py | ||
| install_python_stack.py | ||
| install_sd_cpp_prebuilt.py | ||
| LICENSE.AGPL-3.0 | ||
| node_prebuilt_pins.json | ||
| package-lock.json | ||
| package.json | ||
| setup.bat | ||
| setup.ps1 | ||
| setup.sh | ||
| Unsloth_Studio_Colab.ipynb | ||