img2img and inpaint take their output size from the uploaded image and only snap it to a multiple of 16, so an ordinary phone photo (up to the 4096/side decode cap, 4x the txt2img 2048 ceiling and ~16x the area) drove an OOM-scale latent and an opaque 500 on a normal card, while txt2img, upscale, edit, and FLUX.2-klein inpaint are all already megapixel-bounded. Clamp the init longest side to 2048 (the txt2img ceiling) before deriving width/height; edit is exempt since its pipeline resizes to ~1MP internally. _cast_nvfp4 quantized every nn.Linear with no filter, unlike the int8 and fp8 torchao text-encoder modes which exclude the VLM vision tower / lm_head / T5 wo. On qwen-image / qwen-image-edit that 4-bit quantized the Qwen2.5-VL image tower, degrading the edit/image conditioning the sibling schemes protect. Apply the same make_filter_fn exclusion (require_bf16, mirroring _cast_fp8_dynamic). |
||
|---|---|---|
| .. | ||
| assets | ||
| auth | ||
| core | ||
| hub | ||
| loggers | ||
| models | ||
| plugins | ||
| requirements | ||
| routes | ||
| state | ||
| storage | ||
| tests | ||
| utils | ||
| __init__.py | ||
| _platform_compat.py | ||
| cloudflare_tunnel.py | ||
| colab.py | ||
| main.py | ||
| run.py | ||
| startup_banner.py | ||