Review follow ups on the more-families branch: the per channel scale now broadcasts rank aware instead of assuming 2D (all shipped tensors are 2D, verified across all three fp8 components, but a future non 2D quantized tensor would have mis broadcast silently), the fused qkv split asserts the expected 3x hidden row count so a GQA style export fails loudly, fp8 detection scans every shard header rather than the first, and the excluded model match uses the segment aware token helper with a hunyuanimage-3 token so a future HunyuanImage 2.x is not blocked with a 3.0 reason. |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| src-tauri | ||
| __init__.py | ||
| install_llama_prebuilt.py | ||
| install_node_prebuilt.py | ||
| install_python_stack.py | ||
| install_sd_cpp_prebuilt.py | ||
| LICENSE.AGPL-3.0 | ||
| node_prebuilt_pins.json | ||
| package-lock.json | ||
| package.json | ||
| setup.bat | ||
| setup.ps1 | ||
| setup.sh | ||
| Unsloth_Studio_Colab.ipynb | ||