Measured on the fresh linux x64 prebuilt (z-image Q8_0, sd-cli, 192 CPU threads, 512x512, 9 steps, steady state): sampling 56.1s vs 51.3s (about 9 percent faster), VAE decode unchanged, peak RSS identical. The sd.cpp engine only serves the no-GPU tier, so the default profile now matches max: --diffusion-fa plus --diffusion-conv-direct. |
||
|---|---|---|
| .. | ||
| assets | ||
| auth | ||
| core | ||
| hub | ||
| loggers | ||
| models | ||
| plugins | ||
| requirements | ||
| routes | ||
| state | ||
| storage | ||
| tests | ||
| utils | ||
| __init__.py | ||
| _platform_compat.py | ||
| cloudflare_tunnel.py | ||
| colab.py | ||
| main.py | ||
| run.py | ||
| startup_banner.py | ||