Drops backfill_document_locators + the /documents/{id}/locators/backfill
route and its response model, the BackfillResult dataclass and the
backfill-only helpers (_scope_for_document, _update_vector_payloads), the
frontend backfillDocumentLocators client, and the backfill/migration
tests. Live preview-highlight locators (pdf_regions_for_chunks, computed
at ingest) are untouched.
Matches #5910's text-only footprint. Removes image-vector embedding
(encode_images, _stream_image_chunks, the _BGEVLAdapter CLIP shim), the
multimodal `mode`/KBMode concept + VL embedders (single text embedder
now), the mode selector UI across the KB dialogs + thread settings, the
MM badges, the /images serving route, and the dead image rendering in
the search tool card. Captioning (figure text spliced into markdown)
stays — #5910 keeps it too. DB mode/image columns left dormant (no
migration). RagDefaultsSection dropped (no controls left).
Backend:
- Deterministic SQLite connection cleanup. The RAG code used bare
`with get_connection() as conn:`, which commits but never closes, leaning
on GC to release handles (the rest of studio_db closes explicitly). Add a
closing_connection() context manager that commits/rolls back like sqlite3's
own manager and always closes, and route all 30 RAG call sites through it.
- filter_by_min_score no longer drops BM25-only and figure-ref hits. min_score
is a cosine floor, so it now gates only hits that carry a dense_score;
lexical and figure-ref hits (dense_score is None) pass through instead of
being silently discarded when the floor is raised.
- Fix two tests that could not pass against the production code: the RRF
fusion test asserted the wrong winner (c edges out b: 0.032266 vs 0.032258),
and two tool-handler scope tests stubbed retrieve_hybrid without accepting
the embedder_model kwarg the handler now passes (TypeError was swallowed,
leaving captured["scope"] unset).
Frontend:
- Removing an in-flight upload chip now routes through the teardown thunk
already registered for the aggregate-progress toast (abort, unsubscribe,
release the index slot, delete the backend doc with the correct kb/thread
scope key it closed over) and clears the toast entry. Deleting directly
leaked the concurrency slot and hardcoded the thread scope, mis-targeting
KB-scoped docs. Applied in both the composer hook and the compare-view
composer; drop the now-vestigial chip-scope-key tracking and unused
activeThreadId selectors. Add index-progress-store.remove(id).
Remove code with no live references, each confirmed dead via AST reference
analysis (no production callers and no importers), not just text search:
- chunk_belongs_to_document plus its dedicated tests and the now-orphaned
_insert_chunk test helper. The preview-target route already does a
single-query membership check and deliberately never called this helper.
- ingestion-progress.tsx and use-ingestion-events.ts (its only importer).
Superseded by the aggregate ingestion toast stack; zero importers.
- Unreferenced tests/fixtures/rag-preview sample files and their generator.
No behavior change. The only non-deletion edits reword two comments that
referenced the removed helper.
Shorten and condense comments across the RAG backend, frontend, and
tests for readability. Comment text only; no code, strings, identifiers,
or logic changed. License headers and lint/type pragmas are preserved.
External providers (OpenAI/Anthropic/Gemini) can't run the local
search_knowledge_base tool loop, so give them RAG by prefetching:
studio retrieves before calling the provider, injects the chunks into
the user prompt, and surfaces it as a synthetic tool call. Local models
are untouched (they keep tool-based RAG + decomposition).
Backend:
- New POST /api/rag/prefetch: momentarily loads the pre-cached helper
(gemma-4-E2B-it-GGUF) via LlamaCppBackend(kill_orphans=False) to
decompose the question into up to 3 queries, retrieves+merges+dedups
per query, unloads the helper. Raw single-query fallback if the helper
can't load. New core/rag/query_decompose.py owns the helper lifecycle.
- Factored the retrieval body of /search into _execute_search, reused by
both endpoints.
Frontend:
- prefetchRag() client.
- chat-adapter external branch: gated on isExternalRequest + ragToolEnabled
+ scope!=off + ragScopeHasDocs (no docs -> no prefetch, prior behavior
preserved). Formats hits as <chunk id=N> (parseChunks shape), injects
into the last user message (send-only; not shown in the user bubble),
seeds a synthetic search_knowledge_base tool-call part so the existing
chunk-card UI + [N] citations + source badges all work unchanged.
- Extends PR #5674's disabled-tool guard: when RAG is off, reinforce
'no document search (RAG) capabilities'; when prefetch ran, point the
model at the injected excerpts instead.
- RAG pill enabled for external providers regardless of supports_tools.
Not build/UI verified here (no bun/GPU/keys); needs bun typecheck+test
and a browser round-trip with real provider keys.
Two changes to the RAG captioning log output:
- Drop the noisy per-image and path-selection info lines
(using-chat-VLM, loading-helper, per-image done). Only the
'caption_images: invoked' and 'caption_images: complete' lines
remain; warnings for genuine failures (helper load, per-image
request, helper unload) are kept.
- Configure structlog at the top of the ingestion subprocess worker
with the same env the parent uses. The worker runs in a spawned
process where structlog was never set up, so its logs fell back to
structlog's dev ConsoleRenderer ([info] ...) instead of the JSON
renderer the rest of the app uses. Now captioner/parser logs from
the subprocess match the parent's JSON format.
Snapshot taken before fast-forwarding feature/rag to origin and merging main.
Bundles in-flight work so the merge has a clean tree:
Frontend
- PDF preview panel (preview-panel, preview-pdf-view, preview-text-view,
preview-unavailable) with lazy-rendered page thumbnail rail
- Resizable preview slot via useResizablePanelWidth hook (drag handle,
localStorage persistence, viewport clamping)
- Neutral scrollbar + Source Excerpt card restyle (no brand-coloured rail)
- Preview-store + chat-adapter / rag-api / kb-detail wiring
- Frontend test harness (vitest.config, setupTests, biome update) and the
paired __tests__ suites for preview, sources, document-row, chat-adapter,
rag-api, knowledge-bases-tab, search-knowledge-base-tool-ui
Backend
- RAG locator + authorization modules with chunking / retrieval / tool /
vector_store / studio_db updates
- Paired test_rag_* suites (authorization, locators, locator_backfill,
locator_migration, preview_routes, preview_target_locators, source_identity)
Other
- tests/fixtures/rag-preview for preview route fixtures (sample.pdf,
sample.txt, make_fixture_pdf.py)
- .gitignore + package(-lock).json adjustments for the new test runner
Will be squashed/reworked via interactive rebase after main is merged.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
get_reranker() acquires the module-level _lock and then, on first load,
calls unload() to clear any stale state before _load() instantiates the
CrossEncoder. unload() acquires the same _lock — but threading.Lock is
non-reentrant, so the second acquisition by the holding thread blocked
forever. Symptom: rerank=True hung the search_knowledge_base tool with
no further log output past 'rerank entered'.
Switch to threading.RLock so the same thread can re-enter without
blocking. unload()'s independent callers still work the same way; the
only behaviour change is that re-entrant acquisition from one thread
now succeeds.
When the reranker hung on rerank=True there were zero log lines after
'retrieved=N (no threshold)', which made it impossible to tell whether
the hang was in _load (CrossEncoder construction), in get_reranker's
lock acquisition, or in predict. Structlog routing may also be the
culprit since we never saw the 'Loading RAG reranker' info line.
Add unconditional stderr prints at each milestone — entered, device
resolved, before CrossEncoder, after CrossEncoder, rerank entered,
predict starting, predict done. These bypass any logger config and
show up directly in /tmp/studio.log next to the rest of the captured
stdout/stderr. Leaving structlog logger.info calls in place too so
the structured stream still gets the same data when routing works.
The reranker model (BAAI/bge-reranker-base by default, ~1.1 GB) was
never precached, so the first user-facing rerank call paid the full
download cost — which on slow connections looked like a hang and got
retried by upstream timeouts. The deprecation warning that surfaced
during the hang was actually from sentence-transformers internals
firing while the download was still in flight.
Mirror the precache_helper_gguf pattern: add precache_reranker() that
calls snapshot_download in a daemon thread at FastAPI startup. The
first opt-in rerank now finds the weights already on disk and only
pays the in-process model load.
Also tighten the loader:
- explicit device selection (cuda when torch.cuda.is_available,
else cpu) so we don't rely on sentence-transformers auto-detect
behaviour that has historically picked cpu under odd
CUDA_VISIBLE_DEVICES configs;
- structlog-shaped logs with elapsed_seconds around load + predict
so a real runtime hang is visible in /tmp/studio.log with
'RAG reranker predict starting' / 'RAG reranker predict done'.
Scores were only ever useful for debugging; surfacing them in chunk
cards (score X.XXX · dense Y.YYY) and citation hovers made the UI
noisy without giving the user anything actionable. Drop them in three
places:
- Backend search_knowledge_base no longer emits score / dense_score
attributes on the <chunk> tags fed to the LLM; the tool description
is updated to match.
- Chunk-card metadata in the assistant-ui tool result strips score /
dense lines.
- Source-badge hover tooltips drop the 'score N' meta line.
Also remove the 'Min relevance' slider from the chat settings sheet.
The backend min_score field stays plumbed (default 0 = no filter) so
the threshold can be re-exposed later or driven programmatically.
Captions were appended at the bottom of the page text, so the chunk
containing 'Figure 1: Asymmetries ...' got chunked separately from
'**Figure**: Flowchart with ...' on the same page. Retrieval surfaced
the caption-text chunk but the VLM description landed in a different
chunk, leaving the LLM without the visual content right next to the
figure label.
Splice each VLM caption right after the matching 'Figure N:' (or
'Table N:') line as '**Figure N description**: ...', so:
- The figure-boundary chunker now keeps both the original in-PDF
caption AND the VLM description in the same chunk (which starts
with 'Figure N:').
- Multi-figure pages get per-figure attribution — the prefix
'Figure N description' lets the LLM tell two figures on the same
page apart, even though the bbox renderer still emits one image
per page today (multi-figure clustering is a follow-up).
- When the page text has no figure lines (DOCX/HTML/TXT or rare
PDF layouts) the old end-of-page appendix is kept as a fallback.
page.get_images() only returns raster blobs embedded in the PDF's
resource dictionary, so vector schematics like Figure 1 — drawn purely
with paths/lines — were never extracted, and the VLM only ever saw
incidental embedded photos that happened to live near figures.
Replace the xref-based extraction with bbox rendering: union the
bounding rects of all vector drawings and raster image_info entries on
each page, expand a few points, and render the region with
get_pixmap(clip=bbox, matrix=2x). The captioner now receives the
actual figure — schematic arrows, box labels, legend text, and any
inset photos — and produces a caption that describes the figure as a
whole, not just one embedded sub-image.
Also sharpen the captioner prompt: explicitly tell the VLM the image
is a single figure cropped from a PDF page, and not to describe page
chrome or body paragraphs.
Dense vectors don't preserve numbers (BGE-small treats 'Figure 1' and
'Figure 10' as nearly identical), so a query like 'what does Figure 1
show' got out-ranked by chunks describing other figures that share more
vocabulary with the question — even after the figure-boundary chunker
ensured Figure 1's chunk started with the literal caption.
Detect 'Figure N' / 'Table N' (numbered, decimal, appendix-style)
references in the query, look up chunks that start with those captions
directly, and feed the result as a third RRF source. RRF gives them
rank-0 in the third ranking and the fused score lifts them above the
dense-vocabulary noise. No-ops when the query has no figure ref.
Dense embedders mean-pool over a whole chunk, so a 'Figure 1:' caption
buried at the end of a 500-token body chunk gets washed out by the
surrounding theory text and never surfaces for queries about that
figure. Pre-split each page's markdown at the start of every
Figure/Table caption line so the caption anchors its own chunk, which
gives both BM25 and the dense vector a focused, figure-dominated
target. Handles numbered, decimal, and appendix-style labels
(Figure 1, Figure 1.2, Figure B.1, Table 4, Fig./Tab. abbreviations).
Reasoning models (gemma-4, qwen3-thinking) burn the entire max_tokens
budget on <thinking> output and return empty visible content, so the
captioner produced zero captions for every image. Pass
chat_template_kwargs={enable_thinking: false} per-request to skip the
reasoning phase, and bump max_tokens 120 -> 200 as headroom.
Both modules used stdlib logging.getLogger which is not bridged to the
project's structlog config, so every probe / captioner log was silently
dropped. Switch to loggers.get_logger and convert %-format calls to
structlog kwargs so the captioning path becomes observable.
Four valid review comments from gemini-code-assist[bot] on #5759:
1. core/rag/bm25.py:_load — wrap bm25s.BM25.load + json.loads with
specific exception handlers (FileNotFoundError, OSError,
JSONDecodeError, ValueError) and log a warning instead of
propagating a 500. Corrupt/partial bm25 dirs now degrade to
empty-search rather than crashing the request.
2. core/rag/tool.py was importing _resolve_scope_embedder from
routes/rag.py — a layering violation (core depending on
routes). Move the resolver into a new core/rag/scope.py module
along with the chat-settings key constants; routes/rag.py
now re-imports it under the same name. Same behaviour, no
cycle, one source of truth for the resolution logic.
3. core/rag/bm25.py:rebuild_index — call delete_scope before saving
the new index so stale files from a previous build (or a
bm25s naming change) never coexist with current files. The
library's save() doesn't unlink files it doesn't write.
4. routes/rag.py:_save_upload was running f.write() synchronously
inside an async def. Switch to anyio.open_file() so each chunk
write runs in a worker thread instead of blocking the event
loop on multi-MB uploads. Cleanup unlink happens after the
async-with closes the handle so Windows is happy.
Skipped one (vector_store.py:133 'hasattr query_points' redundancy)
— that comment was on the pre-rewrite Qdrant code; the file is
now sqlite-vec backed and the hasattr check is gone.
core/rag/db.py, vector_store.py, tool.py, bm25.py, and reranker.py
all run only in the FastAPI parent process. Switch their loggers
from Python stdlib to studio's structlog get_logger so their output
shows up in the same JSON stream as the rest of the backend (the
request_completed / RAG search lines).
embeddings.py and ingestion.py stay on stdlib because they execute
inside the mp.spawn ingestion subprocess, which doesn't inherit the
parent's structlog configuration.
asg017/sqlite-vec is Apache-2.0 and OSI-approved. Replaces
qdrant-client (~30 MB) with a small SQLite extension loaded into a
dedicated rag.db file. Single file holds RAG vectors; bm25s indexes
and chat-side studio.db are unaffected.
- New core/rag/db.py owns the rag.db connection and sqlite-vec load.
Extension load runs once at first open. Process-wide singleton
protected by a lock; check_same_thread=False + WAL handles the
FastAPI thread pool.
- core/rag/vector_store.py keeps the same public API
(ensure_collection / upsert_chunks / search / collection_exists /
delete_scope / delete_document) so callers in routes/rag.py,
core/rag/ingestion.py, core/rag/tool.py, and core/rag/retrieval.py
don't change. ensure_collection is now a no-op; collection_exists
returns True iff the scope has at least one indexed vector.
- search uses sqlite-vec's vec_distance_cosine and converts distance
to similarity in [0, 1] so the per-scope min_score threshold
semantics stay identical.
- Mixed-dim scopes coexist behind WHERE scope = ? — the per-scope
embedder resolver guarantees one embedder per scope.
- requirements/rag.txt swaps qdrant-client for sqlite-vec.
- utils/paths/storage_roots.py drops rag_vectordb_root() (the old
qdrant directory); rag.db lives directly under rag_root().
- Rewritten tests/python/test_rag_vector_store.py for the new
semantics (collection_exists tracks populated scopes; new tests
for filtered search and upsert conflict resolution).
Python build requirement: connection.enable_load_extension(True)
must be available. install.sh creates the venv via uv-managed
python-build-standalone, which is compiled with
--enable-loadable-sqlite-extensions, so this works on standard
installs. core/rag/db.py raises an actionable error on the rare
custom-interpreter case.
The query was always going through the default text embedder
(bge-small, 384-d) regardless of how the scope was ingested. A
multimodal thread indexed by Qwen3-VL (2048-d) crashed at cosine
similarity with 'shapes (104,2048) and (384,) not aligned'.
Resolve the scope's embedder at search time:
- kb_<id> -> rag_knowledge_bases.embedding_model column
- thread_<id> -> thread settings (with fallback to defaults +
RAG_EMBEDDER_MATRIX matrix lookup)
Pass it through retrieve_hybrid/retrieve_dense to embeddings.encode
so the query lands in the same vector space as the docs. Both the
/api/rag/search route and the search_knowledge_base tool use the
same resolver.
BGE-VL inherits CLIP's 77-token text positional embedding table —
longer chunks crash inside the text model with a shape mismatch.
Pre-tokenize with truncation=True, max_length=77 and call
get_text_features directly so the high-level encode() (which does
not truncate) is bypassed. Log when truncation happens — text
chunks beyond the cap are silently cut, so multimodal mode is
lossy on the text channel. Image channel is unaffected.