Re-uploading the same file into the same scope (KB or thread) used to
parse, chunk, caption and embed it all over again, creating a duplicate
set of chunks. Dedup by content hash instead:
- schema: add rag_documents.content_hash (sha256 of the bytes) via the
standard PRAGMA/ALTER migration, plus (scope, content_hash) indexes.
- upload: _save_upload now streams the bytes through sha256 and returns
the digest alongside path/name/size.
- _start_ingestion: before inserting, look for a COMPLETED row in the
same scope with the same hash. If found, delete the redundant upload
from disk and return the existing document_id with already_indexed=
true and an empty job_id — no ingestion job is started. Only
'completed' rows dedup, so a failed/in-flight prior attempt can still
retry. Scope-local: the same file in two KBs is indexed in each.
- frontend: UploadResponse.already_indexed flows through the rag-store
(skips job subscription) into both upload paths, which mark the chip
ready immediately and toast '<file> is already indexed'.
Pre-existing rows have NULL content_hash and won't dedup until
re-uploaded once under the new path. Not build/UI-verified here (no bun
in this env); needs typecheck + browser check.
|
||
|---|---|---|
| .. | ||
| assets | ||
| auth | ||
| core | ||
| loggers | ||
| models | ||
| plugins | ||
| requirements | ||
| routes | ||
| state | ||
| storage | ||
| tests | ||
| utils | ||
| __init__.py | ||
| _platform_compat.py | ||
| colab.py | ||
| main.py | ||
| run.py | ||
| startup_banner.py | ||