The studio was disabling flex attention entirely on Blackwell+ GPUs (sm_120 and above) by setting UNSLOTH_ENABLE_FLEX_ATTENTION=0 at startup. This was a workaround for the flex_attention backward kernel exceeding shared memory limits on these GPUs. The root cause is now fixed in unsloth-zoo (PR #542) which patches the backward kernel config selection to generate safe fallback configs that fit within the GPU's shared memory limit. With that fix, flex attention works correctly on Blackwell GPUs and provides a ~1.3x speedup over the SDPA fallback. |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| __init__.py | ||
| install_python_stack.py | ||
| LICENSE.AGPL-3.0 | ||
| setup.bat | ||
| setup.ps1 | ||
| setup.sh | ||
| Unsloth_Studio_Colab.ipynb | ||