Shorten and condense comments across the RAG backend, frontend, and
tests for readability. Comment text only; no code, strings, identifiers,
or logic changed. License headers and lint/type pragmas are preserved.
asg017/sqlite-vec is Apache-2.0 and OSI-approved. Replaces
qdrant-client (~30 MB) with a small SQLite extension loaded into a
dedicated rag.db file. Single file holds RAG vectors; bm25s indexes
and chat-side studio.db are unaffected.
- New core/rag/db.py owns the rag.db connection and sqlite-vec load.
Extension load runs once at first open. Process-wide singleton
protected by a lock; check_same_thread=False + WAL handles the
FastAPI thread pool.
- core/rag/vector_store.py keeps the same public API
(ensure_collection / upsert_chunks / search / collection_exists /
delete_scope / delete_document) so callers in routes/rag.py,
core/rag/ingestion.py, core/rag/tool.py, and core/rag/retrieval.py
don't change. ensure_collection is now a no-op; collection_exists
returns True iff the scope has at least one indexed vector.
- search uses sqlite-vec's vec_distance_cosine and converts distance
to similarity in [0, 1] so the per-scope min_score threshold
semantics stay identical.
- Mixed-dim scopes coexist behind WHERE scope = ? — the per-scope
embedder resolver guarantees one embedder per scope.
- requirements/rag.txt swaps qdrant-client for sqlite-vec.
- utils/paths/storage_roots.py drops rag_vectordb_root() (the old
qdrant directory); rag.db lives directly under rag_root().
- Rewritten tests/python/test_rag_vector_store.py for the new
semantics (collection_exists tracks populated scopes; new tests
for filtered search and upsert conflict resolution).
Python build requirement: connection.enable_load_extension(True)
must be available. install.sh creates the venv via uv-managed
python-build-standalone, which is compiled with
--enable-loadable-sqlite-extensions, so this works on standard
installs. core/rag/db.py raises an actionable error on the rare
custom-interpreter case.
BGE-VL inherits CLIP's 77-token text positional embedding table —
longer chunks crash inside the text model with a shape mismatch.
Pre-tokenize with truncation=True, max_length=77 and call
get_text_features directly so the high-level encode() (which does
not truncate) is bypassed. Log when truncation happens — text
chunks beyond the cap are silently cut, so multimodal mode is
lossy on the text channel. Image channel is unaffected.
BGE-VL's sentence-transformers shim (bge_vl_clip_transformer.py)
is coupled to specific ST internals and broke across both the 5.2
and 5.3 pin attempts. The canonical load path documented at
bge-model.com is transformers.AutoModel with trust_remote_code +
model.set_processor() + model.encode(text=/images=).
- Wrap that path in _BGEVLAdapter exposing the slice of
SentenceTransformer API the ingester uses (encode for text or PIL
images, get_sentence_embedding_dimension, best-effort tokenize).
- Restore BAAI/BGE-VL-base in RAG_EMBEDDER_MATRIX.
- Revert sentence_transformers pin to 5.2.0 — no longer relevant
since the multimodal path no longer touches ST.
BGE-VL (multimodal) and nomic-embed-text-v1.5 (late chunking) both
ship custom modeling code in their model repos; sentence-transformers
refuses to import the referenced module (e.g. bge_vl_clip_transformer)
without trust_remote_code=True. Safe to enable because the embedder
matrix is config-pinned — users don't supply arbitrary names.
When a KB has mode = 'multimodal', ingestion extracts images alongside
text and embeds both into a shared 512-d vector space via BGE-VL-base.
Image hits become first-class search results — useful for slides,
reports, and diagrams where text-only retrieval loses ~30-50% of the
content.
Backend
- embeddings.py: new encode_images(image_bytes_list) — opens bytes via
PIL and routes to the SentenceTransformer (BGE-VL accepts PIL images
in the same encode call as text).
- ingestion.py: _subprocess_worker gains document_id arg and a new
_stream_image_chunks() helper. For multimodal KBs the standard text
chunking runs first, then images are saved to
rag_uploads_root() / 'images' / <document_id> / img-NNNN.<ext> and
embedded; for each image with an adjacent caption, both an
'image'-kind chunk (vector = encoded image) and a 'caption'-kind
chunk (vector = encoded caption text) are streamed back with a
shared pair_group field.
- ingestion.py parent: _insert_chunks_and_collect_for_bm25 now reads
kind / image_path / pair_group from the subprocess message,
populates the new rag_chunks columns, and runs a second pass that
sets linked_chunk_id for each image ↔ caption pair. BM25 indexes
text + caption chunks only — image chunks have no tokenisable body.
- retrieval.py: Hit gains a `kind` field plumbed through bm25, dense,
RRF, and rerank paths.
- reranker.py: image-kind hits skip CrossEncoder rerank (text-only
model) but are appended back in their original relative position
rather than dropped.
- routes/rag.py: new GET /api/rag/images/{document_id}/{filename}
static-file route with realpath containment check. SearchHit gains
`kind` and `image_url` fields so the chat UI can render image
thumbnails alongside text hits. KB-doc upload threads kind/mode
through to ingestion.
Frontend
- rag-api.ts: SearchHit gains optional `kind` and `image_url`.
- kb-create-dialog.tsx: new Mode select (Text / Multimodal) alongside
the existing Chunking strategy select. The forbidden
(multimodal + late) combo is enforced in the UI — each side
disables the conflicting option on the other side with a tooltip
explaining why. Embedding-model placeholder cycles through the
three valid defaults (bge-small / nomic / BGE-VL).
- kb-list.tsx + chat-settings-sheet.tsx: 🖼️ MM badge alongside the
⚡ Late one so multimodal KBs are obvious at a glance.
Tests
- test_rag_multimodal.py: parser returns images when want_images=True
and skips them when False; _validate_mode_combo rejects the
forbidden (multimodal, late) pair with 400; RAG_EMBEDDER_MATRIX
contains the three valid combos and excludes the forbidden one;
image URL construction shape is verified. A server-marked test
loads BGE-VL-base end-to-end and confirms image + text vectors
share the same dimension.
Phase 3 of the plan is now feature-complete on the backend; the
remaining items (re-ingest UX for changing strategy on existing KBs)
are tracked under "Backfill UX" and can land separately.
When a KB has chunking_strategy = 'late', ingestion takes a separate
code path that embeds the full document in a single forward pass and
mean-pools token embeddings per chunk span. Each chunk vector carries
full-document context via the encoder's bidirectional attention —
Jina's published technique, ~+6.5 nDCG@10 on long docs.
Backend
- chunking.py: new chunk_pages_with_spans() that joins all pages into a
single full_doc, runs the existing recursive splitter, and returns
per-chunk (char_start, char_end) offsets. Page-number metadata is
recovered by overlap with the original page ranges so PDF citations
still work. Existing chunk_pages() unchanged.
- embeddings.py: new late_chunk_encode(doc_text, char_spans). Tokenizes
the doc with return_offsets_mapping, runs the underlying transformer
to get per-token last_hidden_state, then mean-pools per chunk span.
When the doc exceeds the embedder's context, falls back to windowed
late chunking with a 512-token overlap so cross-window context is
partially preserved.
- ingestion.py _subprocess_worker: branches on chunking_strategy.
'late' path: chunk_pages_with_spans -> late_chunk_encode -> one big
chunks_batch message. 'standard' path unchanged. Both reuse the same
parent-side pump.
- ingestion.enqueue_ingestion: new chunking_strategy + mode kwargs;
defaults to 'standard' / 'text' for legacy callers. embedder model
resolved via resolve_embedder() from the (mode, strategy) matrix.
- routes/rag.py: KB-doc upload reads chunking_strategy + mode from the
KB row (defensive .get for pre-Phase-3 schemas) and threads them
through _start_ingestion.
Frontend
- kb-create-dialog.tsx: new "Chunking strategy" select with Standard /
Late options. Embedding-model placeholder switches to nomic when
Late is picked. createKB request now carries chunking_strategy.
- kb-list.tsx + chat-settings-sheet.tsx: small "⚡ Late" badge next to
late-chunking KB names in the settings KB list and the chat sidebar
dropdown so users see the mode at a glance.
Tests
- test_rag_late_chunking.py: pure-python tests for chunk_pages_with_spans
(chunks index back into full_doc; page numbers inherited by overlap;
pages joined with blank line). A server-marked test loads
all-MiniLM-L6-v2 to exercise late_chunk_encode end-to-end.
No multimodal yet; that's Phase 3B-multimodal (next PR).