Adds document extraction to the chat composer with a small footprint that
reuses the existing RAG preview UI and is fast by default.
Backend
- Adaptive PDF extraction: born-digital pages produce layout-aware Markdown
via pymupdf4llm and render no page images, so a text PDF issues no VLM
calls. Only pages without a text layer are detected as scanned and
rendered for OCR.
- Scanned pages are transcribed (not summarized) through the already loaded
vision model over /v1/chat/completions. No dedicated OCR model is loaded
and the chat model is never swapped out.
- Adaptive render DPI (120, env override) and bounded async caption
concurrency (2 local, 3 gguf, env override).
- /chat/document-support and /chat/extract-document endpoints with NDJSON
streaming progress, cancellation, multipart size guards, and token-budget
truncation.
Frontend
- Reuses the existing RAG DocumentPreviewSheet and MarkdownPreview to render
an extracted document inline (a new markdown preview target), instead of a
separate preview panel.
- Extraction uses whatever model is loaded; there is no OCR model picker,
cross-tab lock, or custom-code consent step.
- Compact document chips in the composer and transcript; image data is
stripped from persisted attachments.
- Document settings expose a mode (fast text, auto, scanned), a caption
toggle, a token budget, and concurrency. Unknown settings keys are ignored.
Adds backend tests for the adaptive path, the support probe, NDJSON
streaming, cancellation, error mapping, and the scanned-page dedup.