[1] Yang, X. et al. “CRAG — Comprehensive RAG Benchmark.” arXiv:2406.04744 (2024). [2] Xia, Y., Chen, J., Gao, J. “Winning Solution For Meta KDD Cup’24.” arXiv:2410.00005 (2024). [3] Ouyang, J. et al. “Revisiting the Solution of Meta KDD Cup 2024: CRAG.” arXiv:2409.15337 (2024). [4] “Multi-Stage Verification-Centric Framework for Mitigating Hallucination in Multi-Modal RAG”(KDD Cup 2025 CRAG-MM). arXiv:2507.20136 (2025). [5] Yao, S. et al. “ReAct: Synergizing Reasoning and Acting in Language Models.” arXiv:2210.03629 (2022). [6] Singh, A. et al. “Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG.” arXiv:2501.09136 (2025). [7] Liang, J. et al. “Reasoning RAG via System 1 or System 2: A Survey on Reasoning Agentic RAG for Industry Challenges.” arXiv:2506.10408 (2025). [8] Gao, L. et al. “PAL: Program-aided Language Models.” arXiv:2211.10435 (2022). [9] Chen, W. et al. “Program of Thoughts Prompting.” arXiv:2211.12588 (2022). [10] Faysse, M. et al. “ColPali: Efficient Document Retrieval with Vision Language Models.” arXiv:2407.01449, ICLR 2025. [11] Yu, S. et al. “VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.” arXiv:2410.10594, ICLR 2025. [12] Cho, J. et al. “M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding.” arXiv:2411.04952 (2024). [13] Wang, Q. et al. “ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents.” arXiv:2502.18017, EMNLP 2025. [14] Zhang, J. et al. “OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation.” arXiv:2412.02592, ICCV 2025. [15] “SERVAL” — VLM生成キャプションをテキストエンコーダで索引する generate-then-encode 方式 (2025). [16] Poznanski, J. et al. “olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models.” arXiv:2502.18443 (2025). [17] Niu, J. et al. “MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing.” arXiv:2509.22186 (2025). [18] Ouyang, L. et al. “OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations.” arXiv:2412.07626, CVPR 2025. [19] Dong, H. et al. “SpreadsheetLLM: Encoding Spreadsheets for Large Language Models.” arXiv:2407.09025 (2024). [20] Anthropic. “Introducing Contextual Retrieval.” Engineering blog (2024). [21] Cormack, G., Clarke, C., Buettcher, S. “Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods.” SIGIR 2009. [22] Asai, A. et al. “Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection.” arXiv:2310.11511 (2023). [23] Yan, S. et al. “Corrective Retrieval Augmented Generation.” arXiv:2401.15884 (2024). [24] Saad-Falcon, J. et al. “PDFTriage: Question Answering over Long, Structured Documents.” arXiv:2309.08872, EMNLP 2024 Industry. [25] “Scaling Beyond Context: A Survey of Multimodal RAG for Document Understanding.” arXiv:2510.15253 (2025). [26] Edge, D. et al. “From Local to Global: A Graph RAG Approach to Query-Focused Summarization.” arXiv:2404.16130 (2024). [27] Sarthi, P. et al. “RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval.” arXiv:2401.18059 (2024).