Dense embedders mean-pool over a whole chunk, so a 'Figure 1:' caption buried at the end of a 500-token body chunk gets washed out by the surrounding theory text and never surfaces for queries about that figure. Pre-split each page's markdown at the start of every Figure/Table caption line so the caption anchors its own chunk, which gives both BM25 and the dense vector a focused, figure-dominated target. Handles numbered, decimal, and appendix-style labels (Figure 1, Figure 1.2, Figure B.1, Table 4, Fig./Tab. abbreviations). |
||
|---|---|---|
| .. | ||
| assets | ||
| auth | ||
| core | ||
| loggers | ||
| models | ||
| plugins | ||
| requirements | ||
| routes | ||
| state | ||
| storage | ||
| tests | ||
| utils | ||
| __init__.py | ||
| _platform_compat.py | ||
| colab.py | ||
| main.py | ||
| run.py | ||
| startup_banner.py | ||