Backend:
- Deterministic SQLite connection cleanup. The RAG code used bare
`with get_connection() as conn:`, which commits but never closes, leaning
on GC to release handles (the rest of studio_db closes explicitly). Add a
closing_connection() context manager that commits/rolls back like sqlite3's
own manager and always closes, and route all 30 RAG call sites through it.
- filter_by_min_score no longer drops BM25-only and figure-ref hits. min_score
is a cosine floor, so it now gates only hits that carry a dense_score;
lexical and figure-ref hits (dense_score is None) pass through instead of
being silently discarded when the floor is raised.
- Fix two tests that could not pass against the production code: the RRF
fusion test asserted the wrong winner (c edges out b: 0.032266 vs 0.032258),
and two tool-handler scope tests stubbed retrieve_hybrid without accepting
the embedder_model kwarg the handler now passes (TypeError was swallowed,
leaving captured["scope"] unset).
Frontend:
- Removing an in-flight upload chip now routes through the teardown thunk
already registered for the aggregate-progress toast (abort, unsubscribe,
release the index slot, delete the backend doc with the correct kb/thread
scope key it closed over) and clears the toast entry. Deleting directly
leaked the concurrency slot and hardcoded the thread scope, mis-targeting
KB-scoped docs. Applied in both the composer hook and the compare-view
composer; drop the now-vestigial chip-scope-key tracking and unused
activeThreadId selectors. Add index-progress-store.remove(id).
Remove code with no live references, each confirmed dead via AST reference
analysis (no production callers and no importers), not just text search:
- chunk_belongs_to_document plus its dedicated tests and the now-orphaned
_insert_chunk test helper. The preview-target route already does a
single-query membership check and deliberately never called this helper.
- ingestion-progress.tsx and use-ingestion-events.ts (its only importer).
Superseded by the aggregate ingestion toast stack; zero importers.
- Unreferenced tests/fixtures/rag-preview sample files and their generator.
No behavior change. The only non-deletion edits reword two comments that
referenced the removed helper.
Shorten and condense comments across the RAG backend, frontend, and
tests for readability. Comment text only; no code, strings, identifiers,
or logic changed. License headers and lint/type pragmas are preserved.
* Studio: keep web search/code pills off on model load if user disabled them
* Studio: avoid redundant localStorage reads when resolving tool pills on load
The completion toast reported only document count. Capture each job's
num_chunks from its complete event into the index-progress store and sum
across the batch, so the toast reads 'N documents and M chunks indexed'.
Already-indexed (deduped) files contribute 0 new chunks.
Uploading several documents (or a folder) produced a separate toast per
file, which piled up. Replace the per-job toast stack with a single
aggregate toast driven by a new index-progress-store that's populated at
addDoc time — so it counts queued files (held by the concurrency
semaphore) in the denominator, which the per-job rag-store map can't see.
- index-progress-store: one entry per file in the batch
(queued/indexing/ready/error + 0..1 progress).
- Both addDoc paths register each file on entry and update it through
the lifecycle (setIndexing after acquiring a slot; setProgress on job
progress events; setReady/setError on terminal).
- ingestion-toast-stack: renders ONE toast — 'Indexing document(s) · X/Y
· Z%' with a progress bar while in flight (overall = completed files +
in-flight fractions, over total), 'RAG index ready · N documents
indexed' (+ failures) when the last finishes, then auto-dismisses.
Single-file uploads still read naturally ('Indexing document' / '1
document indexed').
IngestionProgress / rag-store jobs are untouched (still used by the
KB detail panel). Not build/UI verified here (no bun).
Uploading many docs (or a folder) previously spawned an ingestion
subprocess per file all at once, thrashing the GPU/CPU. Add a
configurable concurrency limit and a folder picker.
- ragIndexConcurrency setting (default 1) in the chat runtime store,
persisted like the other RAG scalar settings; exposed as a 'Parallel
indexing' slider (1-8) at the bottom of the sidebar Retrieval section.
- New rag-index-queue.ts semaphore: each document upload acquires a slot
before it starts and releases it once its ingestion job finishes
(complete / error / already-indexed), so bulk uploads drain at the
configured rate. Wired into both composer upload paths
(use-thread-doc-uploads + shared-composer).
- Folder upload: a second 'Attach a folder' button on the RAG attach
control uses a webkitdirectory input; every compatible file is routed
through the same queue. Multi-file select already worked (the input has
'multiple' and loops addDoc).
- Content-hash dedup (shipped earlier) means re-scanning a folder skips
already-indexed files.
Not build/UI verified here (no bun); needs bun typecheck + a browser
check of bulk/folder upload draining at the set concurrency.
External providers (OpenAI/Anthropic/Gemini) can't run the local
search_knowledge_base tool loop, so give them RAG by prefetching:
studio retrieves before calling the provider, injects the chunks into
the user prompt, and surfaces it as a synthetic tool call. Local models
are untouched (they keep tool-based RAG + decomposition).
Backend:
- New POST /api/rag/prefetch: momentarily loads the pre-cached helper
(gemma-4-E2B-it-GGUF) via LlamaCppBackend(kill_orphans=False) to
decompose the question into up to 3 queries, retrieves+merges+dedups
per query, unloads the helper. Raw single-query fallback if the helper
can't load. New core/rag/query_decompose.py owns the helper lifecycle.
- Factored the retrieval body of /search into _execute_search, reused by
both endpoints.
Frontend:
- prefetchRag() client.
- chat-adapter external branch: gated on isExternalRequest + ragToolEnabled
+ scope!=off + ragScopeHasDocs (no docs -> no prefetch, prior behavior
preserved). Formats hits as <chunk id=N> (parseChunks shape), injects
into the last user message (send-only; not shown in the user bubble),
seeds a synthetic search_knowledge_base tool-call part so the existing
chunk-card UI + [N] citations + source badges all work unchanged.
- Extends PR #5674's disabled-tool guard: when RAG is off, reinforce
'no document search (RAG) capabilities'; when prefetch ran, point the
model at the injected excerpts instead.
- RAG pill enabled for external providers regardless of supports_tools.
Not build/UI verified here (no bun/GPU/keys); needs bun typecheck+test
and a browser round-trip with real provider keys.
RAG retrieval runs entirely through the local search_knowledge_base
tool. If the loaded model doesn't support tool calling (e.g. a
safetensors model whose template advertises tools in an unparsable
emission format, so supports_tools is suppressed), enabling RAG does
nothing — the model never calls the tool. The pill stayed lit and
clickable, which was misleading.
Gate the RAG pill on supportsTools (in addition to modelLoaded), the
same condition web/code use when there's no provider builtin. Applied
to both composer surfaces (shared-composer and the in-thread
RagToggle), with a 'RAG needs a model that supports tool calling'
tooltip on the disabled state.
The backend dedups re-uploads and the sidepanel shows the doc once, but
the composer's pending-doc chips are created per addDoc call, so each
re-upload of an already-indexed file appended another 'Ready' chip for
the same document. In the already_indexed branch, if a chip with the
returned documentId already exists, drop the chip we just added instead
of marking it ready — so the composer shows each document only once.
Re-uploading the same file into the same scope (KB or thread) used to
parse, chunk, caption and embed it all over again, creating a duplicate
set of chunks. Dedup by content hash instead:
- schema: add rag_documents.content_hash (sha256 of the bytes) via the
standard PRAGMA/ALTER migration, plus (scope, content_hash) indexes.
- upload: _save_upload now streams the bytes through sha256 and returns
the digest alongside path/name/size.
- _start_ingestion: before inserting, look for a COMPLETED row in the
same scope with the same hash. If found, delete the redundant upload
from disk and return the existing document_id with already_indexed=
true and an empty job_id — no ingestion job is started. Only
'completed' rows dedup, so a failed/in-flight prior attempt can still
retry. Scope-local: the same file in two KBs is indexed in each.
- frontend: UploadResponse.already_indexed flows through the rag-store
(skips job subscription) into both upload paths, which mark the chip
ready immediately and toast '<file> is already indexed'.
Pre-existing rows have NULL content_hash and won't dedup until
re-uploaded once under the new path. Not build/UI-verified here (no bun
in this env); needs typecheck + browser check.
docx-preview defaults to blob: URLs for embedded images, which DOMPurify
strips from img src (blob: isn't in its default allowed-URI list), so
figures vanished after sanitize. Switch docx-preview to useBase64URL so
images inline as data: URIs, and add ADD_DATA_URI_TAGS: ['img'] to the
DOMPurify config so those data: URIs survive sanitization. Script /
handler / javascript: stripping is unchanged.
Previously a DOCX citation only showed the extracted snippet + a
Download button (Risk #3: never render a user-supplied .docx inline).
Add a faithful inline render that keeps that guarantee:
- New PreviewDocxView renders the .docx with docx-preview into an
off-screen element, then injects DOMPurify-sanitized HTML into the
live DOM (keeping <style> for docx-preview's scoped layout CSS).
Script tags, event handlers and javascript: URLs are stripped, so
a malicious .docx can't execute in the app origin.
- preview-store now fetches the raw bytes for docx and exposes them
via previewBlob, but deliberately keeps previewBlobUrl = null — no
object URL is created, so the 'open raw original inline' path stays
disabled (Risk #3) and Download remains the only raw-file path.
- preview-panel routes docx -> PreviewDocxView when a blob is present,
falling back to the text-view snippet otherwise. isInlineBlobAllowed
still returns false for docx, so html/unknown behaviour is unchanged.
- Deps: docx-preview + dompurify added to package.json.
- Tests updated: docx now asserts bytes-fetched-without-object-URL.
Not build/UI-verified in this environment (deps not installed here);
needs bun install + browser check.
- Add human-readable stage labels for caption_images ('Captioning
images') and extract_images ('Extracting images') so the raw
underscore stage names no longer leak into the progress toast.
- On completion, the toast title is now 'RAG index ready' (was
'Indexed') and the body reads '1 document and N chunk(s) indexed'
(was 'Indexed N chunks'), with chunk pluralization.
The toast stack and the chat settings sheet were both z-50, so an open
side panel (rendered later in the DOM) covered the indexing toast. Bump
the stack to z-[9999] — comfortably above the sheet's z-50 — so the
ingestion toast stays visible like the Sonner reranker toast does.
The thumbnail-rail refactor hoisted <Document> to wrap both the rail and
the main page so the PDF loads once. That moved the width-measuring scroll
container INSIDE <Document>, which only renders its children after the PDF
finishes loading. The old `useEffect(..., [])` ran on component mount —
when the container was still absent — so the ResizeObserver never attached,
`width` stayed null, and the main <Page> collapsed to width 0.
Replace the mount-effect measurement with a callback ref: the
ResizeObserver now attaches the instant the container node mounts,
regardless of when that happens relative to PDF load. Disconnects cleanly
on unmount / re-attach.
Verified: tsc clean, vite build succeeds, preview-pdf-smoke (incl. the
resize/debounce case) passes.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
assistant-ui mints a fresh __LOCALID_* draft id on every page load, so
RAG docs uploaded under the previous draft id were orphaned after a
refresh — the doc panel queries useThreadDocuments(activeThreadId) and
the new id had nothing.
Persist activeThreadId in localStorage and, on the first settled render
after load, have ActiveThreadSync ask aui to switchToThread(persisted)
when it differs from the freshly-minted draft. Because the draft was
already persisted to the backend by initialize()/ensureThreadRecord
when its first doc was uploaded, the adapter's fetch() resolves it and
aui adopts it as mainThreadId. That keeps aui's mainThreadId and our
activeThreadId unified, so the earlier divergence (uploads under the
persisted id vs chat-completion reading aui's fresh id) can't recur —
unlike the reverted localStorage-only attempt, the chat-adapter's
unstable_threadId now equals the persisted id after the switch.
A one-shot ref ensures we only re-adopt on initial load; user-driven
new-chat / thread switches still flow through normally. If the
persisted draft was never initialized (no doc/message, not in the
backend), switchToThread rejects and we fall back to the fresh draft.