Fuses the block's two inline modulations (1 + scale) * norm(x) + shift and two gated residuals x + gate * out to torch.addcmul, matching the existing qwen / z-image / flux fusions (compile-safe, 1-ULP more accurate, body-drift guarded). Stock-vs-patched equivalence test included; install count is now 7. |
||
|---|---|---|
| .. | ||
| data_recipe | ||
| export | ||
| inference | ||
| rag | ||
| training | ||
| __init__.py | ||
| _torchao_stub.py | ||
| import_guards.py | ||
| tool_healing.py | ||