Re-uploading the same file into the same scope (KB or thread) used to
parse, chunk, caption and embed it all over again, creating a duplicate
set of chunks. Dedup by content hash instead:
- schema: add rag_documents.content_hash (sha256 of the bytes) via the
standard PRAGMA/ALTER migration, plus (scope, content_hash) indexes.
- upload: _save_upload now streams the bytes through sha256 and returns
the digest alongside path/name/size.
- _start_ingestion: before inserting, look for a COMPLETED row in the
same scope with the same hash. If found, delete the redundant upload
from disk and return the existing document_id with already_indexed=
true and an empty job_id — no ingestion job is started. Only
'completed' rows dedup, so a failed/in-flight prior attempt can still
retry. Scope-local: the same file in two KBs is indexed in each.
- frontend: UploadResponse.already_indexed flows through the rag-store
(skips job subscription) into both upload paths, which mark the chip
ready immediately and toast '<file> is already indexed'.
Pre-existing rows have NULL content_hash and won't dedup until
re-uploaded once under the new path. Not build/UI-verified here (no bun
in this env); needs typecheck + browser check.