Shorten and condense comments across the RAG backend, frontend, and
tests for readability. Comment text only; no code, strings, identifiers,
or logic changed. License headers and lint/type pragmas are preserved.
Two changes to the RAG captioning log output:
- Drop the noisy per-image and path-selection info lines
(using-chat-VLM, loading-helper, per-image done). Only the
'caption_images: invoked' and 'caption_images: complete' lines
remain; warnings for genuine failures (helper load, per-image
request, helper unload) are kept.
- Configure structlog at the top of the ingestion subprocess worker
with the same env the parent uses. The worker runs in a spawned
process where structlog was never set up, so its logs fell back to
structlog's dev ConsoleRenderer ([info] ...) instead of the JSON
renderer the rest of the app uses. Now captioner/parser logs from
the subprocess match the parent's JSON format.
page.get_images() only returns raster blobs embedded in the PDF's
resource dictionary, so vector schematics like Figure 1 — drawn purely
with paths/lines — were never extracted, and the VLM only ever saw
incidental embedded photos that happened to live near figures.
Replace the xref-based extraction with bbox rendering: union the
bounding rects of all vector drawings and raster image_info entries on
each page, expand a few points, and render the region with
get_pixmap(clip=bbox, matrix=2x). The captioner now receives the
actual figure — schematic arrows, box labels, legend text, and any
inset photos — and produces a caption that describes the figure as a
whole, not just one embedded sub-image.
Also sharpen the captioner prompt: explicitly tell the VLM the image
is a single figure cropped from a PDF page, and not to describe page
chrome or body paragraphs.
Reasoning models (gemma-4, qwen3-thinking) burn the entire max_tokens
budget on <thinking> output and return empty visible content, so the
captioner produced zero captions for every image. Pass
chat_template_kwargs={enable_thinking: false} per-request to skip the
reasoning phase, and bump max_tokens 120 -> 200 as headroom.
Both modules used stdlib logging.getLogger which is not bridged to the
project's structlog config, so every probe / captioner log was silently
dropped. Switch to loggers.get_logger and convert %-format calls to
structlog kwargs so the captioning path becomes observable.