* Studio: account for mmproj VRAM in GGUF fit budget (#5825) Vision GGUFs load the mmproj projector onto the GPU via --mmproj alongside the weights, but the context auto-sizing / GPU-selection budget sized off _get_gguf_size_bytes(model_path), which counts only the weight file(s). The projector was never added, so the budget was too optimistic: context got mis-estimated and tight vision loads spilled to system RAM / OOM'd. Resolve the launch projector once before GPU selection and fold its size into the fit budget. The same resolved path feeds both the budget and the --mmproj launch flag, so the two cannot disagree. The summary log now reports the projector size separately, keeping "GGUF size" accurate. Adds _mmproj_vram_bytes() + unit tests (no GPU / network / subprocess). * Studio: simplify mmproj summary-log concatenation (#5825) Address review: the summary log mixed explicit `+` with implicit f-string concatenation. Extract the optional projector fragment into `mmproj_note` so the logger.info uses uniform implicit concatenation. No behavioral change. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: trim mmproj VRAM comments --------- Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com> Co-authored-by: imagineer99 <samleejackson0@gmail.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| assets | ||
| auth | ||
| core | ||
| hub | ||
| loggers | ||
| models | ||
| plugins | ||
| requirements | ||
| routes | ||
| state | ||
| storage | ||
| tests | ||
| utils | ||
| __init__.py | ||
| _platform_compat.py | ||
| cloudflare_tunnel.py | ||
| colab.py | ||
| main.py | ||
| run.py | ||
| startup_banner.py | ||