Commit graph

6,179 commits

Author SHA1 Message Date
oobabooga
9ffb28b923 Merge remote-tracking branch 'origin/diffusion-image-workflows' into diffusion-lora 2026-07-01 13:30:26 -03:00
oobabooga
30f40ed44c Merge remote-tracking branch 'origin/diffusion-phase16-native-engine-routing' into diffusion-image-workflows 2026-07-01 13:15:13 -03:00
oobabooga
7d9a59c722 Merge remote-tracking branch 'origin/diffusion-phase15-int8-prequant' into diffusion-phase16-native-engine-routing 2026-07-01 12:54:08 -03:00
oobabooga
707983642a Merge remote-tracking branch 'origin/diffusion-phase14-int8-modulation' into diffusion-phase15-int8-prequant 2026-07-01 12:44:49 -03:00
oobabooga
761cc21157 Merge remote-tracking branch 'origin/diffusion-phase12-fbcache' into diffusion-phase14-int8-modulation 2026-07-01 12:34:16 -03:00
oobabooga
ff9dcf847c Merge remote-tracking branch 'origin/diffusion-phase11-consumer-int8' into diffusion-phase12-fbcache 2026-07-01 12:28:36 -03:00
oobabooga
ed4336dbf7 Merge remote-tracking branch 'origin/diffusion-phase10-attention' into diffusion-phase11-consumer-int8 2026-07-01 12:19:17 -03:00
oobabooga
6ebaf64dc6 Merge remote-tracking branch 'origin/diffusion-phase9-prequant' into diffusion-phase10-attention 2026-07-01 12:13:04 -03:00
oobabooga
731fb20bde Merge remote-tracking branch 'origin/diffusion-phase8-quant' into diffusion-phase9-prequant 2026-07-01 12:07:42 -03:00
oobabooga
b02eacd64f Merge remote-tracking branch 'origin/diffusion-phase7-perf' into diffusion-phase8-quant 2026-07-01 12:02:11 -03:00
oobabooga
6adbb45eb9 Merge remote-tracking branch 'origin/diffusion-phase6-features' into diffusion-phase7-perf 2026-07-01 11:56:04 -03:00
oobabooga
1fa8e307ea Merge remote-tracking branch 'origin/diffusion-phase4-native' into diffusion-phase6-features 2026-07-01 11:50:02 -03:00
oobabooga
d285a4d250 Merge remote-tracking branch 'origin/image-generation' into diffusion-phase4-native 2026-07-01 11:36:04 -03:00
pre-commit-ci[bot]
5728670f6e [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-01 11:57:22 +00:00
Daniel Han
9637e85311 Merge remote-tracking branch 'origin/diffusion-image-workflows' into diffusion-lora 2026-07-01 11:56:47 +00:00
Daniel Han
38ed3ce5b5 Merge remote-tracking branch 'origin/diffusion-phase16-native-engine-routing' into diffusion-image-workflows
# Conflicts:
#	studio/backend/core/inference/diffusion.py
#	studio/backend/core/inference/diffusion_families.py
#	studio/backend/tests/test_sd_cpp_install.py
#	studio/frontend/src/components/assistant-ui/model-selector/pickers.tsx
#	studio/frontend/src/features/images/api.ts
#	studio/frontend/src/features/images/images-page.tsx
#	studio/install_sd_cpp_prebuilt.py
2026-07-01 11:48:33 +00:00
pre-commit-ci[bot]
44beb54df5 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-01 11:36:08 +00:00
pre-commit-ci[bot]
db081d67de [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-01 11:34:22 +00:00
pre-commit-ci[bot]
ffe5b69773 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-01 11:33:51 +00:00
pre-commit-ci[bot]
a8e87ac77b [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-01 11:33:20 +00:00
pre-commit-ci[bot]
c1b4ed1233 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-01 11:32:45 +00:00
pre-commit-ci[bot]
1ab9db02ff [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-01 11:32:13 +00:00
Daniel Han
91edf06748 Merge branch 'diffusion-phase15-int8-prequant' into diffusion-phase16-native-engine-routing
# Conflicts:
#	studio/backend/routes/inference.py
2026-07-01 11:31:54 +00:00
pre-commit-ci[bot]
ab27ea8520 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-01 11:31:18 +00:00
pre-commit-ci[bot]
394f7985df [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-01 11:30:40 +00:00
Daniel Han
f899927652 Merge branch 'diffusion-phase14-int8-modulation' into diffusion-phase15-int8-prequant 2026-07-01 11:30:29 +00:00
Daniel Han
5951dbe145 Merge branch 'diffusion-phase12-fbcache' into diffusion-phase14-int8-modulation 2026-07-01 11:30:28 +00:00
Daniel Han
2a5713aff5 Merge branch 'diffusion-phase11-consumer-int8' into diffusion-phase12-fbcache 2026-07-01 11:30:26 +00:00
Daniel Han
2f26a785b1 Merge branch 'diffusion-phase10-attention' into diffusion-phase11-consumer-int8 2026-07-01 11:30:24 +00:00
Daniel Han
b84d7ea7f3 Merge branch 'diffusion-phase9-prequant' into diffusion-phase10-attention
# Conflicts:
#	studio/backend/tests/test_diffusion_routes.py
2026-07-01 11:30:23 +00:00
Daniel Han
4fec0bb79c Merge branch 'diffusion-phase8-quant' into diffusion-phase9-prequant
# Conflicts:
#	scripts/compare_engines.py
2026-07-01 11:29:36 +00:00
Daniel Han
5d86a7685a Merge branch 'diffusion-phase7-perf' into diffusion-phase8-quant
# Conflicts:
#	studio/backend/core/inference/diffusion.py
2026-07-01 11:29:00 +00:00
pre-commit-ci[bot]
7410f8fa1c [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-01 11:28:58 +00:00
Daniel Han
114baa6607 Merge branch 'diffusion-phase6-features' into diffusion-phase7-perf
# Conflicts:
#	studio/backend/core/inference/diffusion_memory.py
#	studio/backend/core/inference/diffusion_speed.py
2026-07-01 11:27:02 +00:00
pre-commit-ci[bot]
91780a250e [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-01 11:26:13 +00:00
Daniel Han
172a4eb642 Merge branch 'diffusion-phase4-native' into diffusion-phase6-features
# Conflicts:
#	studio/backend/core/inference/sd_cpp_engine.py
#	studio/backend/tests/test_sd_cpp_engine.py
2026-07-01 11:24:14 +00:00
Daniel Han
48628252bd Merge remote-tracking branch 'origin/image-generation' into diffusion-phase4-native
# Conflicts:
#	scripts/diffusion_bench.py
#	scripts/diffusion_quality.py
#	studio/backend/core/inference/diffusion.py
#	studio/backend/core/inference/diffusion_device.py
#	studio/backend/core/inference/diffusion_families.py
#	studio/backend/core/inference/diffusion_memory.py
#	studio/backend/core/inference/diffusion_precision.py
#	studio/backend/core/inference/diffusion_speed.py
#	studio/backend/models/inference.py
#	studio/backend/routes/inference.py
#	studio/backend/tests/test_diffusion_backend.py
#	studio/backend/tests/test_diffusion_device.py
#	studio/backend/tests/test_diffusion_memory.py
#	studio/backend/tests/test_diffusion_precision.py
#	studio/backend/tests/test_diffusion_speed.py
2026-07-01 11:19:15 +00:00
pre-commit-ci[bot]
387cfd268b [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-01 10:47:53 +00:00
Daniel Han
2cae1c1e7c Merge remote-tracking branch 'origin/main' into image-generation
# Conflicts:
#	studio/backend/routes/models.py
#	studio/frontend/src/components/assistant-ui/model-selector/pickers.tsx
2026-07-01 10:47:21 +00:00
Daniel Han
237d07b035 Studio Images: clearer error for an unsupported diffusion model
When a repo id resolves to no diffusion family the load raised 'Could not infer a
diffusion family... Pass family_override (z-image)', which points at an unrelated
family and doesn't say what is supported. Replace it with a message that lists the
supported families (from a new supported_family_names helper) and notes that video
models and image models whose diffusers transformer has no single-file loader are
not supported. Applies to both the diffusers and native sd.cpp load paths. Also
refreshes two stale family-registry comments that still called FLUX.2-dev omitted.
2026-07-01 10:17:35 +00:00
pre-commit-ci[bot]
4f8295e829 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-01 09:58:52 +00:00
Daniel Han
dfa9777fe3 Studio Images: add the FLUX.2-dev model family
Loading unsloth/FLUX.2-dev-GGUF failed because detect_family knew only the
Qwen3-based FLUX.2-klein, so FLUX.2-dev (the full, Mistral-based Flux2Pipeline)
resolved to nothing and the load errored. Add a flux.2-dev family: Flux2Pipeline
+ Flux2Transformer2DModel over the black-forest-labs/FLUX.2-dev base repo (gated,
reachable with an HF token), with its FLUX.2 32-channel VAE and Mistral text
encoder wired for the sd-cli path from the open Comfy-Org/flux2-dev mirror.
text-to-image only: diffusers 0.38 ships no Flux2 img2img / inpaint pipeline for
dev. Frontend gets sensible dev defaults (28 steps, guidance 4), distinct from
klein's turbo defaults. Verified live: GGUF load resolves the family + gated base
repo and generates a real 1024x1024 image on GPU.
2026-07-01 09:58:09 +00:00
pre-commit-ci[bot]
de729c0193 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-01 09:30:27 +00:00
Daniel Han
c846c9896a Studio: hide single-file image checkpoints from the chat model picker
The chat picker treats a cached repo as an image model, and hides it, only
when it ships a diffusers model_index.json. Single-file, ComfyUI, and
ControlNet image checkpoints (an FP8 Qwen-Image, a z-image safetensors, a
Qwen-Image ControlNet) carry none, so they surfaced as loadable chat models.
Fall back to resolving the repo id against the known diffusion families, the
same resolver the Images backend loads from, so these checkpoints are tagged
text-to-image and stay in the Images picker only.
2026-07-01 09:29:51 +00:00
Daniel Han
604949390b Studio Images: list on-device unsloth diffusion models in the picker
The Images picker's On Device tab hid every non-GGUF cached repo whenever a
task filter was active, so downloaded unsloth diffusion pipelines (bnb-4bit
and FP8 safetensors) never showed up there. List cached repos that pass the
task gate, limited under a filter to unsloth-hosted ones so base repos (which
fail the diffusion load trust gate) don't appear only to dead-end on click.
Chat behavior is unchanged: the task gate still drops image repos there.
2026-07-01 09:29:51 +00:00
Daniel Han
38a69ab154 Studio Images: clarify the GGUF transformer-quant Advanced control
Renamed the confusing "Transformer quant / GGUF default" control to "GGUF speed mode"
with an "Off (run the GGUF)" default, and reworded the hint to state plainly that FP8/INT8/
FP4 load the FULL base model (larger download + more VRAM) rather than re-packing the GGUF,
falling back to the GGUF if it can't fit. Behavior unchanged; labels/hint only.
2026-07-01 08:12:34 +00:00
Daniel Han
bc69dfad08
MLX CI: find llama-cli where save_pretrained_gguf actually installs it (#6777)
* MLX CI: find llama-cli where save_pretrained_gguf actually installs it

The GGUF reload step hardcoded the CWD-relative paths llama.cpp/llama-cli and
llama.cpp/build/bin/llama-cli, but save_pretrained_gguf builds and installs llama.cpp
under unsloth_zoo's LLAMA_CPP_DEFAULT_DIR ($UNSLOTH_LLAMA_CPP_PATH, else
~/.unsloth/llama.cpp), so the reload could not find the binary and failed the Mac M1
job with "llama-cli not found". _find_llama_cli now searches that install directory
(and honors the env override) before falling back to the old CWD layout, with a
recursive glob as a last resort. The search is a strict superset of the previous
paths, so it cannot regress a layout that already worked.

* MLX CI: return an absolute llama-cli path from the locator

Resolve the located binary to an absolute path. If UNSLOTH_LLAMA_CPP_PATH is a
relative directory (e.g. "."), Path(".") / "llama-cli" normalizes to the bare name
"llama-cli", and subprocess.run treats a separator-less argument as a PATH lookup
rather than a file to execute, raising FileNotFoundError. resolve() makes the returned
path absolute so it always runs the intended binary.

* MLX CI: give llama-cli EOF on stdin so GGUF reload cannot hang

With the binary now found, the GGUF reload actually invokes llama-cli and it timed
out after 300s generating 24 tokens on a 270m model, which is a stdin block rather
than slow generation: subprocess.run captured stdout/stderr but left stdin inherited,
so -no-cnv still left llama-cli waiting for interactive input. Pass
stdin=subprocess.DEVNULL so it receives an immediate EOF and runs the single prompt to
completion.

---------

Co-authored-by: danielhanchen <michaelhan2050@gmail.com>
2026-07-01 00:49:23 -07:00
Daniel Han
db5746acc7 Studio diffusion LoRA: sanitize dots out of adapter aliases
The LoRA alias is used as the diffusers PEFT adapter name, and PEFT rejects names
containing "." (module name can't contain "."). sanitize_alias kept dots, so a LoRA whose
filename carries a version tag (e.g. Qwen-Image-2512-Lightning-8steps-V1.0-bf16) failed to
apply with a 400. Replace dots too; the alias stays a valid native <lora:NAME:w> filename
stem. Adds regression coverage for internal dots.
2026-07-01 07:34:01 +00:00
Daniel Han
c42efd87ec Studio Images: keep curated safetensors models in Recommended after download
The curated bnb-4bit / fp8 diffusion rows were filtered out of the Images picker's
Recommended list once cached (curatedSafetensorsRows dropped anything in downloadedSet),
so they vanished from the picker after the first load and could only be found by typing an
exact search. The row already renders a downloaded badge, matching how GGUF Recommended
rows stay visible when cached. Drop the exclusion so the curated safetensors always list.
2026-07-01 06:35:14 +00:00
Daniel Han
8cc05ac89c
Reduce comments across recent fixes (#6776)
Condense the verbose comments and docstrings added by the recent
chat template, GPT-OSS detection, PEFT tensor-parallel, and Studio
inference proxy fixes. Comments and whitespace only; no code changes.
2026-06-30 23:13:36 -07:00