Studio: add RAG with hybrid search, reranker, chat integration
Backend (studio/backend/):
- core/rag/: parsers (PDF/TXT/MD/DOCX/HTML via pypdf/python-docx/bs4),
recursive token-aware chunker, embeddings singleton via
FastSentenceTransformer.from_pretrained(for_inference=True), Qdrant
local vector store, bm25s lexical index, RRF hybrid retrieval,
spawn-subprocess ingestion job with SSE progress, optional
CrossEncoder reranker (off-by-default).
- routes/rag.py: KB CRUD, doc upload (KB + per-thread), doc list/delete,
ingestion SSE, hybrid+rerank search, thread-index list/clear.
- routes/chat_history.py: purge thread RAG artifacts on thread delete
and clear-all (rag_documents has no FK cascade to chat_threads so
uploads work on un-persisted threads).
- studio.db gains 4 RAG tables; storage_roots gains rag_*() helpers.
- auth/authentication.py: get_current_subject_sse accepts ?token=... so
EventSource can stream ingestion progress.
Frontend (studio/frontend/):
- features/rag/: api client, Zustand store, hooks, dropzone, KB list,
doc rows, ingestion-progress, thread-index list components.
- Settings dialog gains a Knowledge Bases tab (master/detail + thread
documents list); /knowledge-bases deep-links to it.
- features/chat/: per-thread ragSource/enableRerank/ragTopK state in
chat-runtime-store; Retrieval section in chat-settings-sheet with KB
DropdownMenu (active highlight + per-row trash), thread doc list with
Clear-thread-index button, RAG Top K slider, reranker toggle;
chat-adapter retrieves before /v1/chat/completions and injects hits
as a system block; shared-composer + button routes documents into
pendingDocs (auto-uploads, send blocked while indexing).