Commit graph

6,652 commits

Author SHA1 Message Date
pre-commit-ci[bot]
f7fdb4e9e6 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 03:18:31 +00:00
Daniel Han
d442e00407 Scope the video base repo download to the files the pipeline loads
A bare from_pretrained snapshot of Lightricks/LTX-2 pulls the whole 314 GB
repo: 170 GB of packaged root checkpoints and a second 50 GB text-encoder
shard set, when the pipeline reads about 93 GB. Build the needed file list
once (shared with the progress estimate so the two cannot disagree),
download it per file with cancellation, and hand from_pretrained the local
snapshot dir. Clamp the progress counter to the estimate so stale cache
blobs can no longer report over 100 percent.
2026-07-05 03:17:55 +00:00
Daniel Han
92207a35eb Merge branch 'diffusion-more-families' into video-inference 2026-07-05 02:39:07 +00:00
Daniel Han
36b92961eb Merge branch 'diffusion-auto-badges' into diffusion-more-families 2026-07-05 02:39:06 +00:00
Daniel Han
3863dc60e5 Merge branch 'diffusion-auto-install' into diffusion-auto-badges 2026-07-05 02:39:04 +00:00
Daniel Han
68f33b1e0c Merge branch 'diffusion-fp16-accum' into diffusion-auto-install 2026-07-05 02:39:03 +00:00
Daniel Han
ce364cd507 Merge branch 'diffusion-auto-policy' into diffusion-fp16-accum 2026-07-05 02:39:02 +00:00
Daniel Han
26eeed371a Merge branch 'diffusion-train-perf2' into diffusion-auto-policy 2026-07-05 02:39:01 +00:00
Daniel Han
9416fbd73e Merge branch 'diffusion-krea2' into diffusion-train-perf2 2026-07-05 02:39:00 +00:00
Daniel Han
740ff8454e Merge branch 'diffusion-train-tab-2' into diffusion-krea2 2026-07-05 02:38:59 +00:00
Daniel Han
86b12b2213 Merge branch 'diffusion-train-precision' into diffusion-train-tab-2 2026-07-05 02:38:57 +00:00
Daniel Han
aaa9007c37 Merge branch 'diffusion-train-perf' into diffusion-train-precision 2026-07-05 02:38:56 +00:00
Daniel Han
ad8213ba70 Merge branch 'image-generation' into diffusion-train-perf 2026-07-05 02:38:55 +00:00
Daniel Han
add623d186 Merge branch 'diffusion-more-families' into video-inference 2026-07-05 02:19:37 +00:00
Daniel Han
5cccb44ca2 Merge branch 'diffusion-auto-badges' into diffusion-more-families
# Conflicts:
#	studio/backend/core/inference/diffusion_families.py
2026-07-05 02:17:52 +00:00
Daniel Han
4cb52c0399 Merge branch 'diffusion-auto-install' into diffusion-auto-badges 2026-07-05 02:17:27 +00:00
Daniel Han
32e68b4fd0 Merge branch 'diffusion-fp16-accum' into diffusion-auto-install
# Conflicts:
#	studio/backend/core/inference/diffusion.py
2026-07-05 02:17:17 +00:00
Daniel Han
0f9f184676 Merge branch 'diffusion-auto-policy' into diffusion-fp16-accum 2026-07-05 02:16:20 +00:00
Daniel Han
53af22cf1e Merge branch 'diffusion-auto-policy' of https://github.com/unslothai/unsloth into diffusion-auto-policy 2026-07-05 02:14:36 +00:00
Daniel Han
68735819cd Keep transformer_quant tri-state through pre-eviction validation 2026-07-05 02:14:31 +00:00
pre-commit-ci[bot]
78abe26387 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 02:13:58 +00:00
pre-commit-ci[bot]
1bc77d0fd8 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 02:12:59 +00:00
Daniel Han
e049560a4a Merge branch 'diffusion-train-perf2' into diffusion-auto-policy
# Conflicts:
#	studio/backend/core/inference/diffusion.py
2026-07-05 02:12:30 +00:00
pre-commit-ci[bot]
20a3650108 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 02:12:27 +00:00
Daniel Han
551c38bd4a Merge branch 'diffusion-krea2' into diffusion-train-perf2 2026-07-05 02:11:26 +00:00
Daniel Han
dc36983562 Merge branch 'diffusion-train-tab-2' into diffusion-krea2 2026-07-05 02:11:25 +00:00
Daniel Han
a99b951c33 Merge branch 'diffusion-train-precision' into diffusion-train-tab-2
# Conflicts:
#	studio/backend/tests/test_diffusion_training.py
2026-07-05 02:11:15 +00:00
pre-commit-ci[bot]
91d7297d41 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 02:10:58 +00:00
Daniel Han
32a77623ba Merge branch 'diffusion-train-perf' into diffusion-train-precision
# Conflicts:
#	studio/backend/core/training/diffusion_train_common.py
2026-07-05 02:10:23 +00:00
Daniel Han
c2ab1a0e61 Merge branch 'image-generation' into diffusion-train-perf
# Conflicts:
#	studio/backend/core/training/diffusion_dit_trainer.py
#	studio/backend/core/training/diffusion_train_common.py
2026-07-05 02:09:39 +00:00
pre-commit-ci[bot]
f6f198fd5f [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 02:07:49 +00:00
Daniel Han
25d9cf9604 Stub diffusers.hooks too in the no-diffusers cache test 2026-07-05 02:04:16 +00:00
Daniel Han
e605075508 Fix diffusion training validation and honor lr_scheduler and batch size in the DiT trainer 2026-07-05 01:51:34 +00:00
Daniel Han
098809d2fe Reject sd-cli batch runs and clear stale output targets before a run 2026-07-05 01:49:19 +00:00
Daniel Han
76eee534ea Gate explicit attention kernels on NVIDIA CUDA and roll back partial FBCache hooks 2026-07-05 01:48:12 +00:00
Daniel Han
f1d9c88606 Validate load modes before eviction and wait out a cancelled denoise on unload 2026-07-05 01:47:02 +00:00
Daniel Han
6a8b0b47e7 Fix review findings on image generation: failed-load VRAM, API defaults, preflights
- Free reserved VRAM in the diffusion load worker's failure path: a load-time OOM
  never commits _state and the next load's _unload_locked early-returns, so nothing
  else reclaimed the half-built pipeline's memory
- Use a monotonic clock for the denoise ETA rate
- Sync _GENERATION_DEFAULTS with the UI table: kontext, flux.2-dev, sdxl-turbo and
  SDXL base rows so /v1/images/generations stops falling back to 9 steps / CFG 0
- 400 (not sanitized 500) when /v1/images/generations hits an edit-only model
- Fail fast on pre-Ampere CUDA in the DiT trainer instead of dying in model load
- Run the trainer trust gate in the diffusion training route before freeing GPU
  residents so an untrusted base cannot tear down loaded chat/Images models
- Protect native sd.cpp companion VAE/text-encoder repos from cache deletion while
  a load is downloading them
- Exempt the task-scoped Images picker from the chat-only GGUF/MLX format gate so
  local diffusers pipelines stay selectable on no-GPU hosts
2026-07-05 01:00:47 +00:00
pre-commit-ci[bot]
fbcf1ade58 [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
2026-07-05 00:36:34 +00:00
Daniel Han
4d97574c4f Merge branch 'diffusion-more-families' into video-inference 2026-07-05 00:31:51 +00:00
Daniel Han
45fce770cc Merge branch 'diffusion-auto-badges' into diffusion-more-families 2026-07-05 00:31:50 +00:00
Daniel Han
b374b612a1 Merge branch 'diffusion-auto-install' into diffusion-auto-badges 2026-07-05 00:31:48 +00:00
Daniel Han
39ba65d163 Merge branch 'diffusion-fp16-accum' into diffusion-auto-install 2026-07-05 00:31:47 +00:00
Daniel Han
447113f5ca Merge branch 'diffusion-auto-policy' into diffusion-fp16-accum 2026-07-05 00:31:46 +00:00
Daniel Han
b4201e6390 Merge branch 'diffusion-train-perf2' into diffusion-auto-policy 2026-07-05 00:31:36 +00:00
Daniel Han
08674a297e Merge branch 'diffusion-krea2' into diffusion-train-perf2 2026-07-05 00:31:34 +00:00
Daniel Han
99eb248607 Merge branch 'diffusion-train-tab-2' into diffusion-krea2 2026-07-05 00:31:33 +00:00
Daniel Han
71c20ded19 Merge branch 'diffusion-train-precision' into diffusion-train-tab-2 2026-07-05 00:31:31 +00:00
Daniel Han
79b97e9ad0 Merge branch 'diffusion-train-perf' into diffusion-train-precision 2026-07-05 00:31:30 +00:00
Daniel Han
baf6a3c832 Fix video load planning, lifecycle, and trust gaps from review
Review round on the video backend:
- the resident memory check now budgets transformer plus companions like the
  image backend, instead of letting auto pick a resident placement that OOMs
  while the LTX text encoder and VAEs load
- a new load waits for the signalled in flight generation to exit before
  tearing down the old pipeline, so two models never share VRAM during a swap
- the load worker rechecks its token right before placement, narrowing the
  window where a cancelled load could put weights on a GPU the arbiter
  already handed to another backend
- the step cache installs before the speed profile and compile now keys
  fullgraph off an active cache, matching the image order; compiling
  fullgraph first crashed the first cached generation
- teardown uninstalls the process wide compiled GGUF dequantizer so a later
  speed off load gets the bit identical path
- explicit base_repo goes through the same trust gate as non GGUF repo ids,
  and local checkpoint paths are verified during validation, before the
  route evicts a resident model
- status reports only the speed optimisations that actually engaged
- the gallery file route streams via FileResponse with range support instead
  of buffering whole clips
- the build step reuses the checkpoint path resolved during planning
2026-07-05 00:24:16 +00:00
Daniel Han
85395e3b94 Harden ideogram fp8 dequant and scope the HunyuanImage exclusion
Review follow ups on the more-families branch: the per channel scale now
broadcasts rank aware instead of assuming 2D (all shipped tensors are 2D,
verified across all three fp8 components, but a future non 2D quantized
tensor would have mis broadcast silently), the fused qkv split asserts the
expected 3x hidden row count so a GQA style export fails loudly, fp8
detection scans every shard header rather than the first, and the excluded
model match uses the segment aware token helper with a hunyuanimage-3
token so a future HunyuanImage 2.x is not blocked with a 3.0 reason.
2026-07-05 00:16:14 +00:00