meta-pytorch/OpenEnv was transferred to huggingface/OpenEnv. The old
URL still works via GitHub's redirect from a clean clone, but uv's git
cache can fail to follow it on some machines — surfacing as a
'could not read Username for github.com' credential prompt mid-install.
Point the requirement at the canonical huggingface/OpenEnv.git (same
HEAD), which also sidesteps any stale cache entry keyed on the old URL.
The completion toast reported only document count. Capture each job's
num_chunks from its complete event into the index-progress store and sum
across the batch, so the toast reads 'N documents and M chunks indexed'.
Already-indexed (deduped) files contribute 0 new chunks.
Uploading several documents (or a folder) produced a separate toast per
file, which piled up. Replace the per-job toast stack with a single
aggregate toast driven by a new index-progress-store that's populated at
addDoc time — so it counts queued files (held by the concurrency
semaphore) in the denominator, which the per-job rag-store map can't see.
- index-progress-store: one entry per file in the batch
(queued/indexing/ready/error + 0..1 progress).
- Both addDoc paths register each file on entry and update it through
the lifecycle (setIndexing after acquiring a slot; setProgress on job
progress events; setReady/setError on terminal).
- ingestion-toast-stack: renders ONE toast — 'Indexing document(s) · X/Y
· Z%' with a progress bar while in flight (overall = completed files +
in-flight fractions, over total), 'RAG index ready · N documents
indexed' (+ failures) when the last finishes, then auto-dismisses.
Single-file uploads still read naturally ('Indexing document' / '1
document indexed').
IngestionProgress / rag-store jobs are untouched (still used by the
KB detail panel). Not build/UI verified here (no bun).
Uploading many docs (or a folder) previously spawned an ingestion
subprocess per file all at once, thrashing the GPU/CPU. Add a
configurable concurrency limit and a folder picker.
- ragIndexConcurrency setting (default 1) in the chat runtime store,
persisted like the other RAG scalar settings; exposed as a 'Parallel
indexing' slider (1-8) at the bottom of the sidebar Retrieval section.
- New rag-index-queue.ts semaphore: each document upload acquires a slot
before it starts and releases it once its ingestion job finishes
(complete / error / already-indexed), so bulk uploads drain at the
configured rate. Wired into both composer upload paths
(use-thread-doc-uploads + shared-composer).
- Folder upload: a second 'Attach a folder' button on the RAG attach
control uses a webkitdirectory input; every compatible file is routed
through the same queue. Multi-file select already worked (the input has
'multiple' and loops addDoc).
- Content-hash dedup (shipped earlier) means re-scanning a folder skips
already-indexed files.
Not build/UI verified here (no bun); needs bun typecheck + a browser
check of bulk/folder upload draining at the set concurrency.
External providers (OpenAI/Anthropic/Gemini) can't run the local
search_knowledge_base tool loop, so give them RAG by prefetching:
studio retrieves before calling the provider, injects the chunks into
the user prompt, and surfaces it as a synthetic tool call. Local models
are untouched (they keep tool-based RAG + decomposition).
Backend:
- New POST /api/rag/prefetch: momentarily loads the pre-cached helper
(gemma-4-E2B-it-GGUF) via LlamaCppBackend(kill_orphans=False) to
decompose the question into up to 3 queries, retrieves+merges+dedups
per query, unloads the helper. Raw single-query fallback if the helper
can't load. New core/rag/query_decompose.py owns the helper lifecycle.
- Factored the retrieval body of /search into _execute_search, reused by
both endpoints.
Frontend:
- prefetchRag() client.
- chat-adapter external branch: gated on isExternalRequest + ragToolEnabled
+ scope!=off + ragScopeHasDocs (no docs -> no prefetch, prior behavior
preserved). Formats hits as <chunk id=N> (parseChunks shape), injects
into the last user message (send-only; not shown in the user bubble),
seeds a synthetic search_knowledge_base tool-call part so the existing
chunk-card UI + [N] citations + source badges all work unchanged.
- Extends PR #5674's disabled-tool guard: when RAG is off, reinforce
'no document search (RAG) capabilities'; when prefetch ran, point the
model at the injected excerpts instead.
- RAG pill enabled for external providers regardless of supports_tools.
Not build/UI verified here (no bun/GPU/keys); needs bun typecheck+test
and a browser round-trip with real provider keys.
RAG retrieval runs entirely through the local search_knowledge_base
tool. If the loaded model doesn't support tool calling (e.g. a
safetensors model whose template advertises tools in an unparsable
emission format, so supports_tools is suppressed), enabling RAG does
nothing — the model never calls the tool. The pill stayed lit and
clickable, which was misleading.
Gate the RAG pill on supportsTools (in addition to modelLoaded), the
same condition web/code use when there's no provider builtin. Applied
to both composer surfaces (shared-composer and the in-thread
RagToggle), with a 'RAG needs a model that supports tool calling'
tooltip on the disabled state.
The backend dedups re-uploads and the sidepanel shows the doc once, but
the composer's pending-doc chips are created per addDoc call, so each
re-upload of an already-indexed file appended another 'Ready' chip for
the same document. In the already_indexed branch, if a chip with the
returned documentId already exists, drop the chip we just added instead
of marking it ready — so the composer shows each document only once.
Re-uploading the same file into the same scope (KB or thread) used to
parse, chunk, caption and embed it all over again, creating a duplicate
set of chunks. Dedup by content hash instead:
- schema: add rag_documents.content_hash (sha256 of the bytes) via the
standard PRAGMA/ALTER migration, plus (scope, content_hash) indexes.
- upload: _save_upload now streams the bytes through sha256 and returns
the digest alongside path/name/size.
- _start_ingestion: before inserting, look for a COMPLETED row in the
same scope with the same hash. If found, delete the redundant upload
from disk and return the existing document_id with already_indexed=
true and an empty job_id — no ingestion job is started. Only
'completed' rows dedup, so a failed/in-flight prior attempt can still
retry. Scope-local: the same file in two KBs is indexed in each.
- frontend: UploadResponse.already_indexed flows through the rag-store
(skips job subscription) into both upload paths, which mark the chip
ready immediately and toast '<file> is already indexed'.
Pre-existing rows have NULL content_hash and won't dedup until
re-uploaded once under the new path. Not build/UI-verified here (no bun
in this env); needs typecheck + browser check.
docx-preview defaults to blob: URLs for embedded images, which DOMPurify
strips from img src (blob: isn't in its default allowed-URI list), so
figures vanished after sanitize. Switch docx-preview to useBase64URL so
images inline as data: URIs, and add ADD_DATA_URI_TAGS: ['img'] to the
DOMPurify config so those data: URIs survive sanitization. Script /
handler / javascript: stripping is unchanged.
Previously a DOCX citation only showed the extracted snippet + a
Download button (Risk #3: never render a user-supplied .docx inline).
Add a faithful inline render that keeps that guarantee:
- New PreviewDocxView renders the .docx with docx-preview into an
off-screen element, then injects DOMPurify-sanitized HTML into the
live DOM (keeping <style> for docx-preview's scoped layout CSS).
Script tags, event handlers and javascript: URLs are stripped, so
a malicious .docx can't execute in the app origin.
- preview-store now fetches the raw bytes for docx and exposes them
via previewBlob, but deliberately keeps previewBlobUrl = null — no
object URL is created, so the 'open raw original inline' path stays
disabled (Risk #3) and Download remains the only raw-file path.
- preview-panel routes docx -> PreviewDocxView when a blob is present,
falling back to the text-view snippet otherwise. isInlineBlobAllowed
still returns false for docx, so html/unknown behaviour is unchanged.
- Deps: docx-preview + dompurify added to package.json.
- Tests updated: docx now asserts bytes-fetched-without-object-URL.
Not build/UI-verified in this environment (deps not installed here);
needs bun install + browser check.
- Add human-readable stage labels for caption_images ('Captioning
images') and extract_images ('Extracting images') so the raw
underscore stage names no longer leak into the progress toast.
- On completion, the toast title is now 'RAG index ready' (was
'Indexed') and the body reads '1 document and N chunk(s) indexed'
(was 'Indexed N chunks'), with chunk pluralization.
CodeQL flagged information exposure through an exception in the /warmup
and /reranker/precache endpoints: both returned str(exc) in the JSON
body, exposing internal paths and stack details to the client. Keep
the full exception in the server-side warning log and return a generic
error message ('Failed to load embedder' / 'Failed to download
reranker') to the caller instead. The frontend only surfaces the
message in a toast, so a generic string is sufficient.
Two changes to the RAG captioning log output:
- Drop the noisy per-image and path-selection info lines
(using-chat-VLM, loading-helper, per-image done). Only the
'caption_images: invoked' and 'caption_images: complete' lines
remain; warnings for genuine failures (helper load, per-image
request, helper unload) are kept.
- Configure structlog at the top of the ingestion subprocess worker
with the same env the parent uses. The worker runs in a spawned
process where structlog was never set up, so its logs fell back to
structlog's dev ConsoleRenderer ([info] ...) instead of the JSON
renderer the rest of the app uses. Now captioner/parser logs from
the subprocess match the parent's JSON format.
The toast stack and the chat settings sheet were both z-50, so an open
side panel (rendered later in the DOM) covered the indexing toast. Bump
the stack to z-[9999] — comfortably above the sheet's z-50 — so the
ingestion toast stays visible like the Sonner reranker toast does.
The thumbnail-rail refactor hoisted <Document> to wrap both the rail and
the main page so the PDF loads once. That moved the width-measuring scroll
container INSIDE <Document>, which only renders its children after the PDF
finishes loading. The old `useEffect(..., [])` ran on component mount —
when the container was still absent — so the ResizeObserver never attached,
`width` stayed null, and the main <Page> collapsed to width 0.
Replace the mount-effect measurement with a callback ref: the
ResizeObserver now attaches the instant the container node mounts,
regardless of when that happens relative to PDF load. Disconnects cleanly
on unmount / re-attach.
Verified: tsc clean, vite build succeeds, preview-pdf-smoke (incl. the
resize/debounce case) passes.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: honor --ctx-size and other forwarded args from `unsloth studio run` in Studio's context-fit logic
* refactor: extract resolve_requested_ctx as single source of truth
The test helper was reimplementing the two-line
'ctx_override = parse_ctx_override(...); requested_ctx = ctx_override
if ctx_override is not None else n_ctx' pattern locally, so the test
asserted against its own reimplementation rather than production logic.
Extract the conditional into resolve_requested_ctx and have both the
production caller and the test use it.
* fix(studio): honor pass-through cache type flags in KV VRAM estimate
Studio's KV cache VRAM estimate computed from the first-class
cache_type_kv even when the user passed -ctk/--cache-type-k/-ctv/
--cache-type-v via extras. Those flags reached llama-server fine
(last-wins on the CLI) but the pre-launch estimate kept using the
default f16 bytes-per-element, so GPU placement decisions could be
off when the user lowered cache precision via pass-through.
Adds parse_cache_override + resolve_cache_type_kv in llama_server_args.py
(mirroring parse_ctx_override / resolve_requested_ctx), wires both into
load_model alongside the existing ctx resolution, and adds focused
unit tests for the parser + resolver.
Follow-up to @rolandtannous review on #5815.
---------
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com>
assistant-ui mints a fresh __LOCALID_* draft id on every page load, so
RAG docs uploaded under the previous draft id were orphaned after a
refresh — the doc panel queries useThreadDocuments(activeThreadId) and
the new id had nothing.
Persist activeThreadId in localStorage and, on the first settled render
after load, have ActiveThreadSync ask aui to switchToThread(persisted)
when it differs from the freshly-minted draft. Because the draft was
already persisted to the backend by initialize()/ensureThreadRecord
when its first doc was uploaded, the adapter's fetch() resolves it and
aui adopts it as mainThreadId. That keeps aui's mainThreadId and our
activeThreadId unified, so the earlier divergence (uploads under the
persisted id vs chat-completion reading aui's fresh id) can't recur —
unlike the reverted localStorage-only attempt, the chat-adapter's
unstable_threadId now equals the persisted id after the switch.
A one-shot ref ensures we only re-adopt on initial load; user-driven
new-chat / thread switches still flow through normally. If the
persisted draft was never initialized (no doc/message, not in the
backend), switchToThread rejects and we fall back to the fresh draft.
Follow-ups after merging origin/main into feature/rag:
* chat-adapter.ts: drop `score` field from DocumentSourcePart and the
`chunk.score` copy — the remote "hide RAG retrieval scores from chunks,
citations, and side panel" commit removed `score` from ParsedChunk.
* chat-settings-sheet.tsx: remove the duplicate `const activeThreadId =`
introduced by the merge (kept the HEAD-side declaration at line 497).
* chat-settings-sheet.tsx: drop the "Min relevance" Slider that referenced
`ragMinScore` / `setRagMinScore` — same intent as the hide-scores commit
(these are still on the runtime store but the side-panel UI is gone).
* i18n locales (en, zh-CN): add `settings.tabs.knowledgeBases` translation
key so the new TabDef entry passes the TranslationKey union check.
Verified: `tsc --noEmit` clean, `vite build` succeeds.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Snapshot taken before fast-forwarding feature/rag to origin and merging main.
Bundles in-flight work so the merge has a clean tree:
Frontend
- PDF preview panel (preview-panel, preview-pdf-view, preview-text-view,
preview-unavailable) with lazy-rendered page thumbnail rail
- Resizable preview slot via useResizablePanelWidth hook (drag handle,
localStorage persistence, viewport clamping)
- Neutral scrollbar + Source Excerpt card restyle (no brand-coloured rail)
- Preview-store + chat-adapter / rag-api / kb-detail wiring
- Frontend test harness (vitest.config, setupTests, biome update) and the
paired __tests__ suites for preview, sources, document-row, chat-adapter,
rag-api, knowledge-bases-tab, search-knowledge-base-tool-ui
Backend
- RAG locator + authorization modules with chunking / retrieval / tool /
vector_store / studio_db updates
- Paired test_rag_* suites (authorization, locators, locator_backfill,
locator_migration, preview_routes, preview_target_locators, source_identity)
Other
- tests/fixtures/rag-preview for preview route fixtures (sample.pdf,
sample.txt, make_fixture_pdf.py)
- .gitignore + package(-lock).json adjustments for the new test runner
Will be squashed/reworked via interactive rebase after main is merged.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The two previous commits (41e43b6c7 and 3d3a00c2b) persisted
activeThreadId in localStorage so RAG docs would survive a page
reload. That broke fresh chats: stale localStorage values from a
prior session pinned activeThreadId to an old draft id, but the
chat-completion path reads aui's current mainThreadId via
unstable_threadId. The two diverged — uploads went under the stale
persisted id, the chat-completion turn looked up docs under the new
aui id, and nothing matched.
Revert the persistence + ActiveThreadSync guard. We're back to the
pre-fix behaviour where uploads-in-the-same-session work, and a
proper fix for the reload case (promote drafts to real chat_threads
rows on first doc upload so the id never changes) will land next.
ActiveThreadSync was reacting to aui's mainThreadId === null on mount
(aui hasn't booted yet) by calling setActiveThreadId(null), which
wiped the just-restored persisted draft id from localStorage and
emptied the doc panel for the user's thread. The previous fix only
covered the 'aui minted a different LOCALID' branch; it missed the
'mainThreadId is null while aui boots' branch.
Treat a null mainThreadId as a no-op for the sync. Explicit clears
(new chat, sidebar delete) keep going through setActiveThreadId(null)
directly, so this guard doesn't trap stale state — it just gives the
persisted draft id a chance to survive until aui finishes booting.