Document Type : Original Article
Authors
1
PhD student in artificial intelligence and robotics, Faculty of artificial intelligence and cognitive sciences, Imam Hossein Comprehensive University
2
Assistant Professor, Faculty of artificial intelligence and cognitive sciences, Imam Hossein Comprehensive University
10.22067/cke.2026.99374.1199
Abstract
Retrieval-Augmented Generation (RAG) is often deployed under the assumption that retrieving more documents improves generation by increasing evidence coverage. However, larger retrieved sets may introduce redundancy, distractors, and higher inference cost, making retrieval quantity different from retrieval utility. This study examines retrieval saturation—the point at which increasing retrieval depth yields diminishing, negligible, or negative returns when answer quality and operational cost are jointly considered. In a controlled design, only retrieval depth is varied, while the corpus, retriever family, generator, prompts, and decoding policy remain fixed. Experiments are conducted on four benchmarks representing knowledge-intensive generation: Natural Questions, HotpotQA, ASQA, and QAMPARI.
Three consistent patterns emerge. First, compact and multi-hop question answering tasks exhibit early saturation: performance improves only up to a limited retrieval depth, after which gains plateau or decline. Second, long-form and many-answer tasks show delayed saturation, as broader retrieval supports answer completeness, though later improvements are increasingly offset by higher cost. Third, dense retrieval with Contriever outperforms sparse retrieval with BM25 near the apparent optimum, yet both show clear saturation as depth increases. Qualitative analysis indicates that unsupported elaboration, redundancy, and context inefficiency become more common beyond the useful retrieval range, particularly in compact QA tasks.
Overall, retrieving more documents does not necessarily mean retrieving more useful evidence. Instead of treating retrieval depth as a monotonic variable, practical RAG design should frame it as a retrieval-budgeting problem that jointly optimizes answer quality, factual reliability, and inference cost.
Keywords
Subjects