The base-model select's state could briefly hold the previous family's
repo after a family switch (the reseed effect runs a beat later, and a
value with no matching option makes the browser display the first option
anyway). The request then carried the stale repo: picking Qwen or Z-Image
still sent black-forest-labs/FLUX.1-dev and surfaced FLUX's gated-repo
error under the wrong family. Derive an effectiveBase clamped to the
current family's repos and use it for the select value, the start
request, and the deploy fallback.
Also move the Trigger prompt above Adapter name: the trigger describes
the dataset, the name only labels the output.
Add an Examples group to the training-images dropdown that imports a
curated dataset in one pick, alongside the existing cards. Cards now show
up to three preview thumbnails pulled from the public HF datasets-server
so the set is visible before download. Hide the trigger prompt when every
image already has a caption (a captioned style set needs no trigger), and
turn the training-settings toggle into a ghost button with a rotating
chevron.
Large example datasets (100+ images) rendered every tile at once, so the
caption review grid grew unbounded. Show 24 images per page with < >
chevrons and an x-y of N indicator; a new dataset or refresh resets to
the first page.
The example-dataset cards still overran the ~340px config column: the
license used the Badge component whose baked-in w-fit and whitespace-nowrap
ignored the max-width and truncate, and the grid children had the default
min-width auto so wide content pushed past the column edge and clipped the
Import buttons. Replace the badge with a plain truncating pill span, and
give the config column min-w-0 with overflow-x-hidden so nothing escapes
its width.
When a dataset with images is selected, show a strip of up to 8 sampled
thumbnails with a +N more tile, so users can see what is in the folder
before training. Clicking the strip opens the existing caption review
grid. Samples are drawn evenly across the folder and refresh on dataset
change or after an upload/import.
The Train tab reused the LLM charts section, which also rendered an empty
Grad Norm card and an Eval Loss card showing an Evaluation not configured
placeholder with a red smear. Neither applies to diffusion LoRA training.
Add a diffusion-only two-card view that reuses the loss and learning-rate
cards directly with fixed presentation defaults, and note under the loss
chart that per-step loss is noisy by design so users read the smoothed
line for the trend rather than the raw jitter.
The example-dataset cards used a two-column grid in the ~340px config
column, which wrapped titles one word per line and let the long license
text overrun into the neighbouring card. Switch to one card per row with a
horizontal layout: title with a compact truncated license badge (full text
in the tooltip), a two-line clamped description, and the Import button on
the right.
The Create/Train switch had an icon inside the Train trigger that overhung
the pill corner. Drop the icon, make both triggers a fixed equal width so
the active pill sits flush in the top bar.
Replaces the Train LoRA dialog with a top-bar Create | Train segmented control next to
the model selector. Create renders the existing generation workspace unchanged; Train
renders the full-page training panel (unmounted in Create so its polling stops while the
backend run and its retained metric history survive a tab switch). Adds a deploy handler:
loading the trained adapter's base as a pipeline, queueing the adapter so the LoRA
discovery effect applies it once the base is loaded and LoRA-capable for the matching
family (with a mismatch warning), seeding the prompt with the trigger, and switching back
to Create. Removes the now-unused dialog.
New full-page training workspace for the Images tab. Left column configures the run:
model family (FLUX.1-dev, Qwen-Image, Z-Image, SDXL in popularity order, with per-family
VRAM/license notes and defaults, backfilled from the backend families list when present),
base repo, dataset (existing folder, browser upload, or one-click example import), an
in-browser caption labeling grid (per-image thumbnail + caption saved on blur, delete,
uncaptioned highlight), adapter name, trigger prompt, and collapsed training settings.
Right column shows the live run: progress + loss/avg/speed/peak-VRAM readouts, the reused
training loss/LR charts fed from metric_history, and a completion card that deploys the
adapter into Create or starts another run.
Extends the Images training client for the Train tab: the status type now carries
metric_history (step/loss/lr) plus catalog_path/family/base_model/samples_per_second/
peak_memory_gb; the start request gains model_family; and info gains an optional
families list (per-family bases + defaults). Adds typed calls for the dataset
labeling and one-click example endpoints: list images with captions, thumbnail URL,
write/clear a caption, delete an image, list example datasets, and import an example.
The dialog assumed users knew the Studio home layout and that only SDXL is
trainable, and hid both facts behind free-text fields. Restructure it around
the three real decisions:
- Base model is a dropdown of the trainable SDXL picks (Base 1.0, Turbo, the
loaded SDXL pipeline when there is one) with a custom repo/path escape
hatch, instead of a bare text field defaulting to a repo id.
- Training images come from an in-browser upload (new dataset endpoints) or
a picker over existing dataset folders with image/caption counts. No shell
access or knowledge of the datasets root is needed any more, and the
captioning rules are explained inline.
- The output field is now Adapter name and the instance prompt is labelled
as the trigger prompt, with a no-captions warning wired to the selected
dataset's actual caption count.
Hyperparameters collapse behind a training settings toggle since the
defaults suit a first run. A completed run says where the adapter went and
offers Done / Train another, and the top-bar button gets an icon and a
plainer description. The dialog title states the SDXL-only scope and that
other families load LoRAs but cannot train them yet.
Check cancellation immediately after a ControlNet from_pretrained and before
any device placement, so an unload/eviction that raced the download does not
allocate several GB onto the GPU after the load was already cleared.
Require a loadable weight or shard index (not just config.json) before a local
ControlNet folder is advertised, so an interrupted copy is hidden instead of
failing deep in from_pretrained as a generic 500.
Do not record a strength-0 ControlNet in the gallery recipe: it is treated as
disabled and skipped, so the image is unconditioned and the metadata must not
claim a ControlNet was applied.
Build the control-type picker from the selected ControlNet's advertised
control_types instead of a hardcoded passthrough/canny pair, so a union model
with a precomputed depth or pose map sends the correct control_mode.
Reject LoRA on a torch.compile'd diffusers transformer (Speed=default/max):
diffusers requires the adapter loaded before compilation, so applying one to
the already-compiled module fails with adapter-key mismatches. The status
gate now hides the picker and generate raises a clear message instead.
Convert a cancelled Hub LoRA download (RuntimeError Cancelled) to the
diffusion cancellation sentinel in resolve_specs, so an unload/superseding
load during resolution maps to a 409 instead of a generic server error.
Drop weight-0 LoRA rows before the native support gate so a request carrying
only disabled adapters stays a no-op on families where native LoRA is
unsupported, matching the diffusers path.
Reject duplicate LoRA ids in the request model: both apply paths suffix
colliding names, so a repeated id would stack the same adapter past its
per-adapter weight bound.
Strip all user-typed <lora:...> prompt tags on the native path (only the
selected adapters are materialized in the managed lora-model-dir, so an
unselected tag can never resolve), and restore saved LoRA selections from a
gallery recipe so restore reproduces a LoRA image.
Keep diffusion.py importable without torch: the compile/arch patch modules
import torch at module level, so import them lazily at their load/unload
call sites instead of at module load. This restores the torchless contract
so get_diffusion_backend() works on a CPU/native sd.cpp install.
Match family reject keywords and aliases as whole path/name segments, not
raw substrings, so an unrelated word like edited, edition, or kontextual no
longer misroutes or hides a valid base image model, while supported edit
families (Qwen-Image-Edit, FLUX Kontext) still resolve. Mirror the same
segment matching in the picker task filter.
Route FLUX.2-dev native guidance through --guidance like the other FLUX
families rather than --cfg-scale. Reject native upscale requests that have
no input image. Read image header dimensions and reject over-limit inputs
before decoding pixels, so a crafted small-payload image cannot spike
memory. Reject an upscale that would shrink the source below its input
size. Validate the model_kind against the filename extension before the
GPU handoff. Estimate a local diffusers pipeline's size from its on-disk
weights so auto memory planning does not skip offload and OOM. Report
workflows: [txt2img] from the native backend status so the Create tab
stays enabled for a loaded native model. Clamp the outpaint canvas to the
backend's 4096px decode limit.
Adds regression tests for segment matching and kind/extension validation.
Backend:
- validate_load_request rejects a non-.gguf single-file name before the GPU
handoff, so a family-looking repo paired with README.md no longer evicts the
chat model and only fails in the background load.
- detect_family scopes the edit/kontext/inpaint keyword check to the model id
or filename basename, not arbitrary parent directories, so a valid
text-to-image file under a folder named edit is no longer rejected.
- the images gallery listing skips records that fail schema validation, so one
corrupt or hand-dropped PNG can no longer 500 the whole endpoint.
- _terminate reaps the killed sd-cli child so cancellation and timeout paths do
not leak zombie process-table entries.
- the images load route gates the chat-eviction handoff on the resolved device
being non-CPU, so a CPU-only diffusers fallback no longer evicts a resident
chat model for a load that cannot use the GPU.
Frontend:
- treat Images as a chat-like full-height route (no outer padding or scroll) so
its picker is not pushed down and the gallery is not clipped.
- allow /images under the chat-only guard so the native CPU/MPS image path is
reachable on the no-GPU hosts it was built for.
- roll the optimistic quant label back when a same-repo swap fails after the
load started, so the selector never advertises a quant that is not loaded.
The dataset and output placeholders showed /path/to/... examples, but the
training routes resolve those fields inside the Studio home and reject
absolute paths outside the approved roots, so following the placeholder
produced a 400. Use folder-name placeholders and say in the labels and the
dialog description where each folder resolves.
Memory planning and dense-quant path: size a local diffusers base's
resident companions from its on-disk VAE and text-encoder weights instead
of folding them to zero, feed the distilled variant hint into the runtime
headroom estimate so turbo and schnell models are not over-reserved, place
group-offload companions resident before attaching the transformer hooks
so a failed placement falls back to whole-module offload instead of
crashing, and bail out of the dense transformer download before it starts
when the requested quant scheme is unsupported so the load falls back to
GGUF cleanly.
sd.cpp stack: scrub the native path lease secret from sd-cli child env,
redact native load-progress errors, forward the resolved accelerator when
auto-installing a forced-native binary, release stale diffusion GPU
ownership on CPU-native loads, and remove the sd.cpp install tree on
uninstall.
Prequant and scripts: reject prequant artifacts missing base_model_id
when a base is requested, expanduser before checkpoint existence checks,
record and validate the int8 exclusion filter and fp8 fast-accum in
checkpoint metadata, make verify_prequant_backend allowlist its local
checkpoint and fail on missing or bad LPIPS and on load-peak regressions,
average only finite PSNR values in diffusion_quality, and reset the
process-wide attention backend between perf probe variants.
API and UI: normalize attention_backend casing before Literal validation,
close hidden popovers when leaving the Images page, and clear the stale
quant label when loading a direct local GGUF file.
Backend:
- Sanitize a blank hf_token to None in begin_load and load_pipeline, so the
default empty Studio token loads anonymously instead of 401ing as an explicit
empty credential.
- Free the ACTIVE diffusion engine before LLM training and in the delete-cached
guard: on a native (sd_cpp) selection the diffusers singleton reports
unloaded, so training could start against a live sd-cli generation and
delete-cached could remove a GGUF the native engine is using. Both now go
through diffusion_engine_router.get_active_diffusion_engine().
- Refuse delete-cached while a background image load is downloading the repo
(or its companion base): status().loaded is False in that window, but the
delete would yank blobs from under the in-flight assembly. Both engines
expose the in-flight ids via a new loading_repo_ids().
- Cap request seeds at 2**53-1: seeds round-trip through JSON gallery recipes,
where JavaScript rounds larger integers, so a restored recipe generated a
different image. Random seeds were already masked to this range.
- Add the task field to CachedModelRepo: the handler sets it for cached
diffusers image repos but response_model silently dropped it, letting
image-only repos pass the chat picker's task gate.
Frontend:
- Offset sequential run seeds by the batch size: the native engine seeds image
j of a run at seed+j, so a +1 run offset regenerated the previous run's
batch-mates.
- Revert the optimistic quant selection when a load fails to start.
- Stop disabling the Images page on chat-only hosts: the native sd.cpp engine
exists exactly for the no-GPU route.
The catalog-refresh .catch from the lower branch clears the selected adapters
too, which is right for its catalog-only picker but wrong here: this picker
holds free-text HF repo ids that are valid without being in the catalog, so a
transient refresh failure must not wipe them. Family swaps still clear the
selection and hidden LoRAs are never sent.
Nine review findings on the SDXL training dialog:
- Forward the saved Hub token so a gated/private SDXL base can be trained (the
image load flow already sends it).
- Re-seed the base-model field from the current default each time the dialog
opens; the keep-alive dialog otherwise kept its mount-time default after a
model loaded.
- Prefill from base_repo (the diffusers pipeline) rather than repo_id, which for
a GGUF/single-file SDXL load is the checkpoint path from_pretrained can't open.
- Add client-side validation of steps/rank/resolution/batch/learning-rate before
the request.
- Expose a precision selector (bf16/fp16/fp32) so non-bf16 GPUs can train from
the UI, not only the API.
- Gate the dialog on the active Images route (active && trainOpen) so switching
tabs closes it and stops its polling.
- Rescan the LoRA picker when a run completes, so a freshly-trained adapter
appears without a model reload.
- Cap the dialog height and scroll the body so the Start/Stop footer stays
reachable on short viewports.
- Correct the copy to not over-promise picker auto-discovery.
Freeing the resident Images pipeline before training is handled backend-side in
the diffusion training start route.
- The LoRA effect cleared the selection on every load->capable transition, which
wiped adapters restored from a gallery recipe before the model finished loading.
Track the previously-loaded family in a ref and clear only on a real family swap;
keep the selection on the initial load and on unload.
- Gate the generate payload's loras on loraCapable so a restored selection that is
hidden (loaded model does not support LoRA) is never sent to the backend.
Address review findings on the LoRA path:
- resolve_one: normalise a blank/whitespace hf_token to None (anonymous access)
and reject a client-supplied weight file with traversal / absolute path.
- resolve_specs: convert FileNotFoundError from an unknown/stale id to ValueError
so the route returns 400 instead of a generic 500.
- _scan_local: disambiguate local adapters that share a stem (foo.safetensors vs
foo.gguf) so each is uniquely addressable.
- inject_prompt_tags: the backend-validated weight now wins over a user-typed
<lora:ALIAS:...> for a selected adapter; unselected user tags are left alone.
- diffusers _apply_loras: reject a .gguf adapter with a clear error before touching
the pipe (diffusers loads safetensors only).
- _unload_locked: drop the explicit unload_lora_weights() on teardown; the pipe is
dropped wholesale (freeing adapters), so the previous call could race an in-flight
denoise on the same pipe.
- Images page: use a stable LoRA key and clear the selection (not just the options)
when the catalog refresh fails.
- resolve_controlnet enforces catalog family compatibility so a direct API call
cannot load a ControlNet built for another family through the wrong pipeline.
- Unknown ControlNet ids now surface as a 400 (call site maps FileNotFoundError
to ValueError) instead of a generic 500.
- strength 0 disables ControlNet entirely, so a no-op selection never pays the
download / VRAM cost; the control image is decoded and validated BEFORE the
ControlNet is resolved or built, so a malformed image fails fast for the same reason.
- ControlNet loads use the base compute dtype (state.dtype is a display string,
not a torch.dtype, so it silently fell back to float32) and honor the base
offload policy via group offloading instead of forcing the module resident.
- Empty/malformed HF token coerced to anonymous access.
- Flux Union ControlNet control_mode mapped from the selected control type.
- resolve_controlnet drops the unused hf_token/cancel_event params.
- ControlNetSpec validates guidance_start <= guidance_end (clean 422).
- Images UI ControlNet Select shows its placeholder when nothing is selected.
Adds regression tests for family enforcement and the union control-mode map.