Advanced Multimodal RAG

Standard multimodal RAG retrieves the wrong images because captions lack document context. Two fixes: context-aware summaries at ingestion, and using the generated answer to select images.