unsloth/studio/backend/storage
Roland Tannous 266342a64e Studio: skip re-indexing an already-indexed document (content-hash dedup)
Re-uploading the same file into the same scope (KB or thread) used to
parse, chunk, caption and embed it all over again, creating a duplicate
set of chunks. Dedup by content hash instead:

  - schema: add rag_documents.content_hash (sha256 of the bytes) via the
    standard PRAGMA/ALTER migration, plus (scope, content_hash) indexes.
  - upload: _save_upload now streams the bytes through sha256 and returns
    the digest alongside path/name/size.
  - _start_ingestion: before inserting, look for a COMPLETED row in the
    same scope with the same hash. If found, delete the redundant upload
    from disk and return the existing document_id with already_indexed=
    true and an empty job_id — no ingestion job is started. Only
    'completed' rows dedup, so a failed/in-flight prior attempt can still
    retry. Scope-local: the same file in two KBs is indexed in each.
  - frontend: UploadResponse.already_indexed flows through the rag-store
    (skips job subscription) into both upload paths, which mark the chip
    ready immediately and toast '<file> is already indexed'.

Pre-existing rows have NULL content_hash and won't dedup until
re-uploaded once under the new path. Not build/UI-verified here (no bun
in this env); needs typecheck + browser check.
2026-05-28 15:19:06 +04:00
..
__init__.py feat(studio): training history persistence and past runs viewer (#4501) 2026-03-25 00:58:55 -07:00
mcp_servers_db.py Studio: add remote MCP server support (#5750) 2026-05-27 07:01:11 -07:00
providers_db.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
studio_db.py Studio: skip re-indexing an already-indexed document (content-hash dedup) 2026-05-28 15:19:06 +04:00