The reranker model (BAAI/bge-reranker-base by default, ~1.1 GB) was
never precached, so the first user-facing rerank call paid the full
download cost — which on slow connections looked like a hang and got
retried by upstream timeouts. The deprecation warning that surfaced
during the hang was actually from sentence-transformers internals
firing while the download was still in flight.
Mirror the precache_helper_gguf pattern: add precache_reranker() that
calls snapshot_download in a daemon thread at FastAPI startup. The
first opt-in rerank now finds the weights already on disk and only
pays the in-process model load.
Also tighten the loader:
- explicit device selection (cuda when torch.cuda.is_available,
else cpu) so we don't rely on sentence-transformers auto-detect
behaviour that has historically picked cpu under odd
CUDA_VISIBLE_DEVICES configs;
- structlog-shaped logs with elapsed_seconds around load + predict
so a real runtime hang is visible in /tmp/studio.log with
'RAG reranker predict starting' / 'RAG reranker predict done'.
|
||
|---|---|---|
| .. | ||
| data_recipe | ||
| export | ||
| inference | ||
| rag | ||
| training | ||
| __init__.py | ||
| tool_healing.py | ||