18 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d6d8b8b1a6 |
Fix the lost-generation proof set, the settle timeout and the hub inventory's diffusion gates
Seven fixes from the latest review round on the Images page and the hub cache inventory. Images page: - The lost-POST settle path built its "already seen" gallery id set inside the catch, after the request failed. By then the earlier runs of the same batch had already prepended their records, so run 2 could accept run 1's image as proof that its own request reached the backend. The set is now captured once before the first POST and grows with every record the batch produces. - settleLostGeneration fell out of its SETTLE_MAX_MS loop and returned normally, so a wedged generation was counted as done and the next run started against a busy backend. It now throws on timeout. - Restoring a recipe cleared the ControlNet selection but left the workflow tab and the init / mask / reference images pointing at whatever was loaded, so the next Generate conditioned on an unrelated image. It now clears all of them and returns to Create. - The download plan omitted the adapter selection the load itself bakes in. A baked LoRA forces the dense build path, so the plan described a different file set than the load that followed and the rest was pulled inline, outside the download manager. Both now derive the list from one helper. Hub cache inventory: - A download for a repo an Images or Video load is staging was allowed to start: only the llama.cpp loader was consulted. Both diffusion backends already expose loading_repo_ids for the delete guard, and the download guard now reads them too. - A companion-only prefetch (pipeline manifest plus VAE and text encoder, no transformer) passed the snapshot-partial check, since every file its manifest expected did arrive, and was advertised as on-device although from_pretrained cannot load it. - The single-file flag never reached the picker through the hub inventory path, so a checkpoint-only diffusion repo read as a full pipeline and failed after the handoff. The two pipeline-shape helpers now live in hub/utils/inventory_scan.py so /api/models/cached and the hub inventory classify the same repos the same way. |
||
|
|
36df317293 |
Trim the comments across the diffusion backend
Comment-only pass over the Python this PR touches: drop what the code already says, collapse multi-line explanations that still read on one line, and keep the reasoning that is not recoverable from the code. No code, docstring semantics or behaviour changes; verified with an AST comparison against the previous revision, and the backend suite is unchanged (same 37 environment failures as before: the API integration tests that need a live keyed server, the flash-attn install hooks, and the GPU memory fields). |
||
|
|
a7d415262a |
Keep the scoped download key derivable, and stop the hidden page hijacking a route
Four review findings, the first a regression from my own last commit: - Keying scoped download jobs by a digest of the file set broke the download manager: it builds that key client-side (it polls and cancels before any response tells it a key), so it watched and cancelled a key no worker owned and never fired its ready callback. Keep the derivable "@scope" key and refuse the second request instead when a live job on the slot is fetching a different file set -- decided inside the registry claim, under the lock, so a concurrent claim cannot slip past it. The manager records the file set on the job as well, so a sibling quant's transfer is not adopted locally either. - Both diffusion pages read the route query through a loose useSearch and both stay mounted once visited, so the hidden one consumed the other's ?model=: it navigated back to its own route and tried to load, say, an image checkpoint as a video model. Only the visible page consumes it. - The staged download plan was built without the configured HF token or the Advanced values the load itself sends. The token matters most: the backend's Hub metadata lookup is best-effort, so a gated base silently planned no companion entry and the load pulled those multi-GB files inline, outside the manager. The memory/quant controls decide whether the base transformer/ shards are needed at all, and the route dropped memory_mode, cpu_offload, the prequant path and the LoRA selection before asking for the plan. - The video preview kept playing after leaving the page: the keep-alive layout only hides it, and display:none does not pause a media element, so a clip the user unmuted kept its audio going over the next page. Pause on the active transition and do not auto-replay while hidden. Also completes the hand-built request bodies in the hub download tests: the scoped-files field this branch added to the route read as an AttributeError against them, failing five tests. |
||
|
|
032561ae21 |
Add a file-scoped flavour to the Hub download job
Lets a consumer that reads a deliberate subset of a repo stage it through the normal download manager. Keyed as "@scope" so it never collides with a quant or with the repo's full snapshot, and the file list rides the registry so an XET to HTTP retry respawns the same scoped job. |
||
|
|
dbb06ff60e |
Studio: add configurable model download location (#7274)
Adds a configurable Hugging Face model download cache location to Unsloth Studio, selectable from Settings, with per-cache download manifests, scoped deletion, and read-only inventory of previously selected caches. |
||
|
|
6d8c18cd1a |
Replace standalone Studio wording with Unsloth (#7221)
* Replace standalone Studio wording with Unsloth Replace the single word Studio with Unsloth wherever it is used as shorthand for Unsloth Studio in docs, CLI output, UI strings, i18n locales, workflow display names, comments and docstrings. Kept unchanged: the full name Unsloth Studio, third party product names (LM Studio, Visual Studio, Mac Studio), feature names (Recipe Studio, Fine-tuning Studio and its translations), and all identifiers such as env vars, commands, paths and filenames. * Address review feedback on the Studio wording rename Use "an" before Unsloth where the rename left the article as "a". Restore the split brand where Unsloth and Studio render as two halves of the full product name: the onboarding sidebar subtitle and the IPv6 localhost warning. Scope two messages to the full name Unsloth Studio where plain Unsloth was misleading: the AMD README bullet and the CLI studio setup error. |
||
|
|
2139200b3f |
Studio: don't re-download updated GGUFs on load (#7209) | ||
|
|
49f2879cf8 |
fix(studio): recover stalled Hub downloads over HTTP (#6858)
* fix(studio): recover stalled Hub downloads over HTTP * fix(studio): preserve retry generation and progress baseline * fix(studio): keep XET retry handoff nonterminal * fix(studio): preserve retry cancellation on claim failure * fix(studio): make retry failure cancellation atomic * fix(studio): close skipped retry state gaps * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Stabilize chat-only export gate detection on Windows * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Retrigger CI on a user-authored head * fix(studio): serialize XET HTTP retry handoff * List XET to HTTP retries that are briefly released from the repo guard as active downloads for PR #6858 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Settle no-process active downloads on shutdown so a parked XET retry cannot spawn after cleanup for PR #6858 * Settle exited-error and no-process downloads on shutdown and persist their cancel markers for PR #6858 * Keep terminal HTTP failures uncancelled and block companion deletion for released retry peers for PR #6858 * Trim download lifecycle test coverage * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com> Co-authored-by: Etherll <61019402+Etherll@users.noreply.github.com> |
||
|
|
b5aef63c03 |
Studio: resolve the repo-root MTP drafter after the MTP/ GGUF rename (#7031)
* Studio: resolve the repo-root MTP drafter after the MTP/ GGUF rename The Gemma 4 QAT GGUF repos renamed the higher-precision MTP/ subdir copies from gemma-4-...-<quant>-MTP.gguf to mtp-gemma-4-...-<quant>.gguf, so their basenames now start with the same mtp- prefix as the small repo-root drafter (mtp-gemma-4-E4B-it.gguf). The drafter selectors filtered candidates by a mtp- basename prefix and took the first in sort order. With the new names the MTP/ copies also match, and because MTP/ (uppercase) sorts before the lowercase root file, selection flipped to the large BF16 copy under MTP/ instead of the root drafter both functions document they should pick. Restrict both selectors, and the companion byte estimate, to root-level mtp-*.gguf so the MTP/ copies stay explicit-selection only: - core/inference/llama_cpp.py _pick_mtp (loader auto-download) - hub/utils/gguf_plan.py preferred_mtp_sibling (Hub variant plans) - routes/inference.py _remote_gguf_companion_bytes (VRAM headroom) Also reuse a drafter already in the local cache before downloading, so a device that already holds a copy on disk does not re-fetch it. Old-scheme names keep working (they have no root-level mtp- sibling to mis-select). Adds regression tests for the new naming, both selection paths, and the on-disk reuse. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: gate MTP drafter cache reuse to offline mode Reuse the cached drafter only when HF is offline. Online, route back through _download_companion_gguf/hf_hub_download so the current revision is checked (etag) and a changed drafter is refetched, matching the offline-only cross-snapshot reuse already used for the main GGUF. This avoids pairing freshly downloaded weights with a stale cached draft. Make the reuse tests offline and add an online-skips-reuse test. * Studio: prefer a root MTP drafter across all cached snapshots Offline reuse scanned snapshots one at a time and returned the first snapshot that held any drafter, only preferring root within it. A newer partial snapshot with just the MTP/ copy could shadow the small root drafter in an older snapshot. Collect drafters across all snapshots and prefer any repo-root file before an MTP/ copy. * Studio: keep newest-first snapshot order when reusing cached drafters Collecting root candidates and sorting by absolute snapshot path could pick a drafter from an older snapshot. _iter_hf_cache_snapshots yields newest first and the main GGUF is resolved in that order, so preserve it (root still preferred over MTP/ copies) to avoid pairing a fresh main weight with a stale drafter revision. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> |
||
|
|
d0c8d550a6 |
fix(studio/hub): apply repo_id length limit per segment, not whole string (#6946) (#6953)
* fix(studio/hub): apply repo_id length limit per segment, not whole string is_valid_repo_id() applied the 96-char limit to the full "namespace/repo_name" string, so a repo with a valid (<=96 char) name but a long combined id was falsely rejected. Match huggingface_hub.validate_repo_id by checking the length per segment instead. Fixes #6946. * Fix long repo id state filenames * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: Etherll <61019402+Etherll@users.noreply.github.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> |
||
|
|
d0f8d40c36 |
studio: allow updating HF models through UI (#5388)
* add models for /update endpoint * add logic for identifying out of date hf models * add endpoint for updating hf models * add relevant field to GgufVariantDetail * make exception handling better * add update_available flag for cached_models, and moved /update endpoint from inference -> models * hook up /update endpoint on the frontend * implement update scenarios for the model picker * fix bug where downloaded flag for an older revision was being wrongly set to false * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix import and make hf calls async * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * remove has_vision from UpdateRequest * fix ci * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * clear cancel event before updating gguf variant * set _cancel_event back if it was set initially * add hf_token to get_paths_info * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * studio: harden model update endpoint and update checks - update_hf_model: pass snapshot_download local_dir (local_path is not a valid kwarg and 500s when updating bicodec audio models) - get_gguf_variants: wrap the remote update check so a network, rate-limit, gated, or offline failure degrades to "no update info" instead of failing the whole variant listing, matching list_cached_models - add regression tests for both paths * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: HF model update detection and Update action for cached models Surface an "Update available" cue and a managed Update action for cached on-device models. /api/hub/update-status compares each cached main GGUF file's local blobs against the remote main revision using set membership across all cached revisions, so a repo that was already updated (and still holds the old snapshot alongside the new one) is not falsely flagged. The Update action re-downloads through the download manager so it shows in the Downloads panel with progress and cancel. The frontend wires the Update button into the GGUF, on-device, and model-selector cards and keeps the quant label fully visible when the action buttons crowd the row. Adds regression tests for the multi-revision update check. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: accept force_download kwarg in hf_xet_fallback test double The download seam now passes force_download to the attempt callable; the _FakeAttempt mock did not accept it, failing 6 tests with TypeError. Add the keyword (default False) so the scripted-results double matches the seam. * Fix Studio model update regressions * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Address Studio update review feedback * Address Studio update edge cases * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Share GGUF update status helper * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix GGUF update detection and cache cleanup * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix cached GGUF update badges --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: shimmyshimmer <107991372+shimmyshimmer@users.noreply.github.com> Co-authored-by: Etherll <61019402+Etherll@users.noreply.github.com> Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com> |
||
|
|
c72da05741 |
Studio: clean up empty leftover quant folders so they can be deleted (#6616)
* Studio: clean up empty leftover quant folders so they can be deleted An interrupted or cancelled split GGUF download leaves snapshots/<rev>/<quant>/ behind with no shards. Such a folder is neither a completed download nor a tracked partial (no .incomplete blobs, no manifest), so it was invisible in the variant list and a per-variant delete returned 404, leaving it on disk forever. - list_empty_gguf_variant_dirs: detect quant folders that are empty in every snapshot, excluding any quant that has shards in another snapshot. - get_gguf_variants_response: surface those quants as partial (cleanable) so the UI shows a delete affordance. - _delete_gguf_variant_from_repos: remove the empty (or just-emptied) quant subfolder and count it toward the result so the delete succeeds instead of 404. Adds hub/tests/test_empty_variant_folder.py. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: simplify empty-dir check to any(iterdir()) * Studio: tighten comments on empty-quant-folder cleanup * Studio: surface empty-folder removal failures and cleanables on local/offline paths Address review feedback on the empty leftover quant folder cleanup: - _remove_empty_variant_dirs now returns removal failures (read-only cache or a locked dir), and the variant delete raises 409 instead of a misleading 404; a concurrent download refilling the dir (ENOTEMPTY) is still treated as a skip. - Empty leftover folders are surfaced as cleanable on every variant-listing path (prefer_local_cache / offline / HF-fallback), not just a remote listing, via a single post-process that flips a listed quant to partial or appends an unlisted one. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: surface empty-folder cleanables even when metadata fetch fails When the cache holds only an empty leftover snapshots/<rev>/<quant>/ folder from an interrupted split download and the client is offline or the HF metadata request fails, _compute() re-raised before cleanables were marked, leaving the folder undeletable. Now fall back to marking cleanables against an empty response and return them if any; otherwise re-raise the original error. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> |
||
|
|
c42c1d56e8 |
Studio: free chat model VRAM at training start only when the GPU is tight (#6243)
* Studio: free chat model VRAM at training start only when the GPU is tight The training start route unconditionally tore down the transformers/MLX inference subprocess before training, and never stopped the llama.cpp GGUF server at all, so a loaded GGUF chat model kept holding VRAM for the whole run. Conversely the HF model was always unloaded even when there was plenty of room to keep it. Make the unload VRAM aware and cover every inference backend: - Add routes/training_vram.py with summarize_resident_chat(), can_keep_chat_during_training() and free_chat_models_for_training(). The keep/unload decision reuses the same estimator and live per device free VRAM reader the training GPU selection already uses (auto_select_gpu_ids, estimate_required_model_memory_gb, get_visible_gpu_utilization), so the probe agrees with the placement computed later in start_training. - When a chat model is resident and training fits alongside it with a conservative margin (required_gb * 1.15 + 4 GB), keep it loaded so the user can train and chat at the same time; on a multi GPU box training lands on a different GPU and both coexist. Otherwise unload the HF/MLX orchestrator and the llama.cpp GGUF server before training starts. - The export subprocess shutdown stays unconditional and now runs first so its freed VRAM is reflected in the decision. Default deny: non CUDA backends, unestimable models, or any probe error fall back to the previous always unload behavior. Adds tests/test_training_vram_coexistence.py and updates two existing route tests in test_gpu_selection.py. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: per-GPU floor for explicit GPU lists + don't unload chat on invalid gpu_ids Address review feedback on the chat coexistence probe: - Explicit gpu_ids mode now enforces a per-GPU floor in addition to the aggregate free-VRAM check, mirroring auto_select_gpu_ids' min_per_gpu_N. Without it, an uneven split such as free [45, 10] for a 40 GB job passed the aggregate threshold and kept chat loaded even though the 10 GB GPU could not hold its training shard, risking an OOM. - Invalid explicit gpu_ids (ids outside the visible set, or a UUID/MIG mask) make resolve_requested_gpu_ids raise. That request is rejected with a 400 before training starts, so leave the resident chat model untouched instead of unloading it. - Tighten the target_modules / gpu_ids type hints to List[str] / List[int]. Adds tests for the per-GPU floor (uneven split unloads, even split keeps) and for invalid gpu_ids keeping the chat model loaded. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: only free chat VRAM once training will start; handle in-flight and CPU-only chat Address the second review pass on the chat-coexistence path: - Run the chat/export VRAM teardown as a before_spawn hook inside TrainingBackend.start_training, fired only after the start guards pass. Previously the route freed chat VRAM before calling start_training, so a refused start (e.g. a lingering pump thread) would tear down the resident chat model even though no training job began. - Treat an in-flight HF chat load (loading_models set, no active model yet) as not safely sizeable: free it rather than risk both OOMing as the load keeps allocating after training starts. - Do not count or tear down a GGUF llama-server confirmed to run entirely on CPU (_gpu_offload_active is False): it holds no VRAM, so killing it cannot help training fit. Adds tests for the before_spawn hook (runs on start, skipped when a subprocess is alive or a pump thread will not die, survives a hook error), the in-flight load flag, and the CPU-only GGUF exclusion in both the resident summary and the unload path. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: treat any in-flight chat load (HF swap / mid-start GGUF) as unsafe to keep Tighten the in-flight detection in summarize_resident_chat so the keep check never sizes a load that is still allocating: - Flag loading on ANY non-empty loading_models, not only when active_model_name is empty. load_model adds the new model to loading_models before clearing the old active_model_name, so a replacement load during a swap was previously sized as a normal resident and could OOM as the new model finishes loading. - Flag a GGUF server that is active but not yet healthy (is_loaded False) as in-flight: it is still mmaping/offloading layers, so its final VRAM footprint is unknown. Consolidates the signal into a single resident["loading"] flag; the route frees the chat model whenever it is set. Adds tests for the replacement HF load and the mid-start GGUF cases. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: tighten comments in chat/training VRAM coexistence (comments only) * Studio: run before_spawn VRAM hook only after GPU-selection validation Reviewers found the before_spawn hook fired before prepare_gpu_selection validated gpu_ids (and before config build), so a refused start (invalid gpu_ids -> 400, or a bad grad-clip value) could still tear down chat/export VRAM. Move the hook to immediately before proc.start(), once all synchronous validation and process construction have passed. This also fixes the route's in-flight-chat loading branch, since that teardown runs inside the same hook. Add test_hook_skipped_when_gpu_selection_rejects. * Studio: recompute GPU auto-selection after the before_spawn VRAM hook Codex P2: with before_spawn moved after prepare_gpu_selection, placement was frozen against the pre-teardown VRAM state while the hook freed export/chat afterward. Auto-selection could pin training onto a GPU the hook then cleared (or onto a kept chat model). Split validation from placement: explicit gpu_ids are still validated before the hook (raise -> 400, no teardown; explicit placement is VRAM-independent), but VRAM-dependent auto-selection now runs after the hook so it sees the freed memory. Add test_auto_placement_runs_after_hook and test_explicit_placement_validated_before_hook. * Studio: allow chatting during training (lift sidebar gate + VRAM-aware load guard) (#6335) * Studio: allow chatting during training (lift sidebar gate + VRAM-aware load guard) The sidebar disabled New Chat, project, and home navigation while a training run was active, so users could not chat during training even though the backend serves inference fine alongside a run. This removes that gate and adds a backend guard so the one genuinely risky operation, loading a new local chat model mid-training, is refused with a clear 409 when it would not fit beside the run. Frontend (app-sidebar.tsx): drop the chatDisabled = isTrainingRunning gate and its consumers. Navigation triggers no model load on its own, so chat stays usable during training. Backend (routes/training_vram.py, routes/inference.py): add can_load_chat_during_training plus a load/validate guard that sizes the same effective load the loader performs (LoRA 4-bit to 16-bit resolved first, HF auto placement via auto_select_gpu_ids, explicit multi-GPU per-GPU floor, GGUF sized from on-disk shards and companions or the selected remote variant). It is a no-op when training is inactive, never blocks external providers or already-resident models, and default-denies only on a CUDA sizing failure so a load can never OOM the run. Validate refuses early with the real settings so the frontend does not unload the resident chat model for a load that would be rejected. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: address review feedback for chat-during-training load guard - Run the load/validate VRAM guard via asyncio.to_thread so the sync nvidia-smi + HF metadata work never blocks the event loop. - Size the GGUF KV cache at the requested context (_estimate_gguf_kv_gb) and add it to the local GGUF estimate so large-context picks are not under-counted. - Keep the requested quantization when adapter_config.json is malformed (not a JSON object) instead of raising in _effective_load_in_4bit. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: size the training load guard at the launcher's effective GGUF context The GGUF KV-cache estimate used max_seq_length only, but the llama.cpp launcher honors a user --ctx-size/-c in llama_extra_args. A load such as max_seq_length=4096 with --ctx-size 131072 was sized against a 4k cache while the server allocates 131k, so the guard could approve a long-context GGUF load that then OOMs training. Size the guard's KV at the larger of max_seq_length and the parsed --ctx-size (reusing the launcher's own parse_ctx_override), keeping the conservative f16 cache so the estimate is never smaller than what the server allocates. The chat model picker also validated with the raw max_seq_length while /load sizes with resolveLoadMaxSeqLength, so validate could pass, unload the current model, then have /load reject the native-context load. Validate now uses the same effective context; the load path is unchanged. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: size the GGUF training guard at the server parallel-slot count The KV-cache estimate assumed a single slot, but llama-server allocates the cache across --parallel slots (app.state.llama_parallel_slots). On a Studio launched with --parallel N>1 the guard under-sized the cache N-fold and could approve a GGUF chat load that then OOMs training. Thread the same slot count the loader uses into the guard's KV estimate; default 1 leaves single-slot setups unchanged. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Trim comments for chat-during-training guard * Studio: keep chat generation alive across navigation; Train spinner + Return to Chat Hoist the base chat runtime above the routed outlet so navigating to Train (or any tab) no longer aborts an in-flight generation; only an explicit Stop cancels. Add a Train sidebar spinner and swap New Chat to Return to Chat while a run is active, with a lightweight completion watch so the spinner clears from any tab. Also respawn a chat llama-server killed mid-session and guard unreadable HF cache dirs that 500'd the hub model list. * Studio: show Return to Chat on the Train tab whenever a chat is live Previously the top sidebar item only swapped to Return to Chat while training was running; on the Train tab with an idle/just-finished run it stayed New Chat, which started a fresh thread and cancelled an in-flight generation. Show Return to Chat (and navigate back, preserving the run) whenever a generation is running or its thread is still active, or training is in progress. * Studio: keep a running chat alive when starting a New Chat Starting a New Chat (or switching threads) while a generation was in flight remounted the single-chat runtime provider, which detached the in-flight run and cut the previous chat off (it showed up frozen / empty when reopened). Key the single-chat view by project instead of by thread or new-chat nonce so the provider stays mounted and assistant-ui switches to a fresh thread in place. The previous generation keeps streaming in the background and autosaves on completion, and returning to that thread reattaches the live run instead of reloading a half-saved one. Also: - "Return to Chat" now lands on the thread that is still generating rather than the empty new chat that became active after New Chat. - Skip the explicit /inference/cancel POST when an abort comes from a runtime detach (navigation / background switch) rather than an explicit Stop, so a backgrounded generation is never cancelled behind the scenes. * Studio: make model export non-blocking and inline The Export tab opened a full-screen modal that trapped focus, could not be closed or cancelled while running, and showed no progress. It also stopped training and unloaded the chat model before loading, so export could not run alongside them. Export now mirrors the training runtime pattern: - Inline panel embedded where the Export Model button was, with no modal or backdrop, so the rest of the UI stays usable during an export. - Global export runtime store plus an app-root lifecycle hook, so a run keeps going and streaming across navigation and is reflected on the Export nav item from any tab. - The worker log stream now stays connected across the load to export phase boundary instead of stranding on "Waiting for worker output". - Progress bar driven by phase and quant index (quant N of M for GGUF), with elapsed time and a working Cancel. - load-checkpoint no longer stops training or unloads inference; export loads in its own subprocess in parallel and surfaces out-of-memory as a clear error. - Add POST /api/export/cancel and is_export_active on /api/export/status. * Studio: show Return to Chat on the Export tab too Extend the New Chat to Return to Chat swap to the Export route so leaving a running chat for Export offers a way back to the live generation, matching the Train tab. * Studio: smooth out Export animations and polish the panel - Drop the height-based reveal animations (source switch, run panel, quant picker, hub fields) that caused flashing and reflow; use instant swaps and quick opacity fades instead. - Method and quant cards now transition colors only, with no transition-all or hover lift, so selecting a method or quant is crisp instead of jumpy. - Auto-scroll the export panel into view when it opens and add a scroll-to-bottom button when its output is below the fold, like Chat. - Show Return to Chat on the Export tab while an export is running, matching how training drives it on the Train tab. - Surface the current phase or stage in the live output before the first worker line arrives so the panel never looks stuck while progress is advancing. * Studio: show Return to Chat on every non-chat tab Generalize the Return to Chat swap from just Train/Export to any non-chat route (Recipes, Projects, Hub, ...) so a running or active chat is always one click away, instead of showing New Chat there. * Studio: stream export logs over the Cloudflare tunnel; drop janky export animations Exporting over a --secure Cloudflare quick tunnel showed "connecting..." with no logs while the progress bar advanced. Cloudflare buffers text/event-stream and only flushes when the stream closes, so the SSE log stream never reached the browser during the run (direct localhost is unaffected, which is why this only showed up over the tunnel). Add a tunnel-safe JSON poll fallback (GET /api/export/logs?since=) that the runtime lifecycle hook polls while a run is active. Short JSON responses are not buffered by the proxy, so logs show up in near real time over the tunnel. It shares the orchestrator's monotonic seq cursor with the SSE stream and the store de-dupes by seq, so the two transports run together (SSE on localhost, poll over the tunnel) without double-printing. A successful poll marks the panel "streaming" instead of leaving it stuck on "connecting...". Also remove the framer-motion AnimatePresence reveals from the export config and run panel (quant picker, hub fields, the inline run panel, and the live log section). The expand/slide animations flashed and felt clunky; the sections now render in place. * Studio: recover export over the Cloudflare tunnel when the blocking POST times out (524) A model export over a --secure Cloudflare quick tunnel showed "Request failed (524)" even though the export succeeded on the backend (the GGUF was written). Cloudflare returns 524 when a single request takes longer than ~100s to respond, and a GGUF conversion routinely runs for minutes, so the blocking per-method export POST is cut off while the backend keeps going. Confirm completion via short status polls instead of relying on the long POST response (the same approach that fixed log streaming): - The orchestrator records each finished op's outcome (status / output_path / error) with a monotonic seq, exposed on GET /api/export/status. - parseJson now preserves the HTTP status; a 524/520/522/523/502/503 or a status-less network drop is classified as a recoverable transport error. - runExport wraps each phase (load, every export method, each GGUF quant): on a recoverable failure it keeps the run alive (logs keep streaming, the panel shows "reconnecting...") and polls status until the still-running op finishes, then settles from the recorded result, recovering the output path for the success banner. A real 4xx still fails immediately; localhost still uses the fast POST response. applyBackendStatus also settles a reloaded run from the last-op record. Verified over the tunnel: a 3m14s gemma-4-E4B-it GGUF export now ends on the success banner with the output path instead of 524. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: keep the export method + logs visible after navigating away mid-export While an export was running, navigating to another tab and back to Export remounted the page and reset the local form state (exportMethod, quant levels), so the method card showed unselected and the run panel's log area was hidden until the card was re-clicked. The run itself lives in the global store and was unaffected. Seed exportMethod / quantLevels from the active run's summary via lazy useState initializers on (re)mount, and gate the panel's log area on the live run (isExporting / logLines / the run's method) rather than only the local form selection. The card stays selected and the logs/progress stay visible across navigation; nothing changes when no run is active. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Studio: address export/training review findings - Export: guard Start against an empty GGUF quant selection so an inline-panel run with no quant can't settle as success with no file produced. - Export: thread the source HF token into the background load so gated/private HF source exports (and gated bases) authenticate, matching the consent path. - Export: only settle a recovered (non-owned) run as a finished export when the last backend op was an export, not a standalone load_checkpoint. - Training: free the export subprocess whenever an export is active, not only once a checkpoint is loaded, so an in-flight export load can't race training for VRAM (current_checkpoint is unset during the load phase). --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> |
||
|
|
c6cf53759b |
Studio: add 'Load on selection' toggle to configure load options before loading (#6348)
* Studio: add 'Load on selection' toggle to configure load options before loading * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: seed staged speculative decoding from the standing default * Studio: address PR review for load-on-selection staging * Studio: handle direct GGUF staging and stale-stage edge cases from load-on-selection review * Studio: cancel replaced staged downloads and keep staged pick on load failure * Studio: centralize staged-download cancel and guard staged-load restore * fix: address staged GGUF load review * fix: honor staged GGUF load metadata * fix: clarify load-on-selection tooltip Keep the load-on-selection hint visually anchored to the control and make the on/off behavior explicit without changing the broader deferred-load flow. * Studio: reset orphaned staged knobs on abandon and cap Max Tokens to staged context * Studio: remove dead code and cancel staged download when loading a different model * fix: surface staged model in run settings before deferred load --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com> Co-authored-by: imagineer99 <samleejackson0@gmail.com> |
||
|
|
048f34e8f2 |
Fix GGUF variant file selection (#6342)
* Fix GGUF variant resolution * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Address GGUF variant review feedback * Harden GGUF endian filtering * Address GGUF endian review comments * Mirror GGUF endian filter in local resolver * Fix GGUF route import test stub * Apply GGUF endian filtering across load paths * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> |
||
|
|
3427e3fd62 |
Studio: fix Downloaded model list disappearing and order it by last download (#6247)
* Studio: fix Downloaded model list disappearing and order it by last download The chat model picker scan for cached GGUF and safetensors models aborted whenever an auxiliary Hugging Face cache dir (such as ~/.cache/huggingface/hub) was unreadable, returning an empty list. That hid the Downloaded section and let already downloaded models appear under Recommended. Isolate each cache probe so an inaccessible directory is skipped instead of failing the scan. Also order Downloaded newest-first using cached blob mtimes (multi-quant repos group by their most recent quant), keep the section visible while searching, and make the per-quant downloaded check per-snapshot and mmproj aware so a Recommended quant is never falsely marked downloaded. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: harden gguf-variants scan and dedupe by newest timestamp Guard f.stat() per file so a broken symlink or unreadable file in a snapshot no longer aborts the downloaded check early, and match quant labels case-insensitively. When the same repo is present in multiple caches with equal size, keep the newest last_modified so Downloaded ordering reflects the most recent copy. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: apply cache-scan guards to sibling endpoints found in review Extend the inaccessible-cache guard and mmproj/stat hardening to the parallel HF cache code paths flagged in review: - list_local_models and the Hub inventory scan now skip an unreadable auxiliary cache instead of returning 500. - The GGUF download-progress endpoint excludes mmproj adapters and guards f.stat() so one bad file does not zero a repo's progress. - The offline snapshot scanner guards its is_dir() probes. - The chat-only picker no longer renders a blank list when a search matches only cached non-GGUF models. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: danielhanchen <michaelhan2050@gmail.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> |
||
|
|
2b319e8d3a |
Studio: support separate-file MTP GGUF drafters (Gemma 4) (#6125)
* Studio: support separate-file MTP GGUF drafters (Gemma 4) * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: fix review findings for separate-file MTP drafters * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: pair local MTP drafters by name and include them in reload dedup * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: manage --model-draft in extras and reject MTP/ copies as models * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> |
||
|
|
aec41d17ed |
feat(studio): Hub + Download Manager (#5916)
Adds the Studio Hub and download manager: browse Hugging Face models and datasets, download GGUF and safetensors with live progress and cancellation, and manage on-device inventory. The Hub does not require a GPU, so it is available on chat-only hosts. CI: all substantive checks pass, including the three Core jobs after unsloth-zoo#736. The two red checks are non-code flakes, a transient npm-registry DNS resolution failure in the package scan and one quantized vision-model output assertion whose sibling shards passed. |