img2img and inpaint take their output size from the uploaded image and only snap it to a multiple of 16, so an ordinary phone photo (up to the 4096/side decode cap, 4x the txt2img 2048 ceiling and ~16x the area) drove an OOM-scale latent and an opaque 500 on a normal card, while txt2img, upscale, edit, and FLUX.2-klein inpaint are all already megapixel-bounded. Clamp the init longest side to 2048 (the txt2img ceiling) before deriving width/height; edit is exempt since its pipeline resizes to ~1MP internally. _cast_nvfp4 quantized every nn.Linear with no filter, unlike the int8 and fp8 torchao text-encoder modes which exclude the VLM vision tower / lm_head / T5 wo. On qwen-image / qwen-image-edit that 4-bit quantized the Qwen2.5-VL image tower, degrading the edit/image conditioning the sibling schemes protect. Apply the same make_filter_fn exclusion (require_bf16, mirroring _cast_fp8_dynamic). |
||
|---|---|---|
| .. | ||
| backend | ||
| frontend | ||
| src-tauri | ||
| __init__.py | ||
| install_llama_prebuilt.py | ||
| install_node_prebuilt.py | ||
| install_python_stack.py | ||
| install_sd_cpp_prebuilt.py | ||
| LICENSE.AGPL-3.0 | ||
| node_prebuilt_pins.json | ||
| package-lock.json | ||
| package.json | ||
| setup.bat | ||
| setup.ps1 | ||
| setup.sh | ||
| Unsloth_Studio_Colab.ipynb | ||