unsloth/studio/backend/core/rag
Roland Tannous 0be7ca39a4 Studio: render figure regions (vector + raster) for RAG captioning
page.get_images() only returns raster blobs embedded in the PDF's
resource dictionary, so vector schematics like Figure 1 — drawn purely
with paths/lines — were never extracted, and the VLM only ever saw
incidental embedded photos that happened to live near figures.

Replace the xref-based extraction with bbox rendering: union the
bounding rects of all vector drawings and raster image_info entries on
each page, expand a few points, and render the region with
get_pixmap(clip=bbox, matrix=2x). The captioner now receives the
actual figure — schematic arrows, box labels, legend text, and any
inset photos — and produces a caption that describes the figure as a
whole, not just one embedded sub-image.

Also sharpen the captioner prompt: explicitly tell the VLM the image
is a single figure cropped from a PDF page, and not to describe page
chrome or body paragraphs.
2026-05-27 16:36:54 +04:00
..
parsers Studio: render figure regions (vector + raster) for RAG captioning 2026-05-27 16:36:54 +04:00
__init__.py Studio: add RAG with hybrid search, reranker, chat integration 2026-05-23 18:46:15 +04:00
bm25.py Studio: trim verbose comments/docstrings across RAG code 2026-05-26 13:51:11 +04:00
captioner.py Studio: render figure regions (vector + raster) for RAG captioning 2026-05-27 16:36:54 +04:00
chunking.py Studio: break RAG chunks at figure/table caption boundaries 2026-05-27 16:09:23 +04:00
db.py Studio: trim verbose comments/docstrings across RAG code 2026-05-26 13:51:11 +04:00
embeddings.py Studio: pass BytesIO (not PIL Image) to BGE-VL encode so model.data_process can re-open 2026-05-27 05:44:40 +04:00
ingestion.py Studio: route RAG ingestion + captioner loggers through structlog 2026-05-27 15:20:22 +04:00
reranker.py Studio: trim verbose comments/docstrings across RAG code 2026-05-26 13:51:11 +04:00
retrieval.py Studio: add figure-reference retrieval source to RAG hybrid search 2026-05-27 16:22:08 +04:00
scope.py Studio: text-mode default with VLM-captioned figure splicing; helper VLM fallback 2026-05-27 11:16:56 +04:00
tool.py Studio: VLM-caption figures at ingest + pass image hits to LLM + render in card 2026-05-26 20:17:52 +04:00
vector_store.py Studio: trim verbose comments/docstrings across RAG code 2026-05-26 13:51:11 +04:00